MedScribe (Vals AI): published results
Vals AI (dataset with Protege) · 100 rubric-scored SOAP-note cases · index updated September 28, 2026
Claude Opus 5.5 has the highest indexed numerical score on MedScribe (Vals AI), 91.43% as of 2026-09-26, per Vals AI MedScribe leaderboard. Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance.
Published results
Showing top 20 of 104 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.932 pp; $1.154156/test; source snapshot 2026-09-26; run date not published | 91.43% | 2026-09-26 | |
| 2 | Claude Fable 5.1 Anthropic model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.953 pp; $0.963500/test; source snapshot 2026-09-26; run date not published | 91.29% | 2026-09-26 | |
| 3 | Claude Sonnet 5.5 Anthropic model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.96 pp; $0.508604/test; source snapshot 2026-09-26; run date not published | 91.10% | 2026-09-26 | |
| 4 | Claude Opus 5 Anthropic model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.236275/test; source snapshot 2026-09-26; run date not published | 90.98% | 2026-09-26 | |
| 5 | Muse Spark 1.2 Meta model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.037779/test; source snapshot 2026-09-26; run date not published | 90.06% | 2026-09-26 | |
| 6 | S | Grok 4.7 SpaceXAI model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not published | 89.38% | 2026-09-26 |
| 7 | Z | GLM 5.3 Flash zAI model ID zai/glm-5.3-flash; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.002235/test; source snapshot 2026-09-26; run date not published | 88.94% | 2026-09-26 |
| 8 | Muse Spark 1.1 Meta model ID meta/muse_spark_1_1; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.95 pp; $0.034628/test; source snapshot 2026-09-26; run date not published | 88.89% | 2026-09-26 | |
| 9 | Z | GLM 5.3 zAI model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.999 pp; $0.057236/test; source snapshot 2026-09-26; run date not published | 88.81% | 2026-09-26 |
| 10 | Claude Fable 5 Anthropic model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.945 pp; $0.583239/test; source snapshot 2026-09-26; run date not published | 88.52% | 2026-09-26 | |
| 11 | X | MiMo V2.6 Pro Xiaomi model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.938 pp; $0.009880/test; source snapshot 2026-09-26; run date not published | 88.31% | 2026-09-26 |
| 12 | GPT 5.1 OpenAI model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 1.942 pp; $0.096508/test; source snapshot 2026-09-26; run date not published | 88.09% | 2026-09-26 | |
| 13 | Kimi K3 Moonshot AI model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not published | 87.96% | 2026-09-26 | |
| 14 | GPT-6 Astra OpenAI model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 1.938 pp; $0.581991/test; source snapshot 2026-09-26; run date not published | 87.91% | 2026-09-26 | |
| 15 | MiniMax-M3 MiniMax model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.957 pp; $0.013748/test; source snapshot 2026-09-26; run date not published | 87.25% | 2026-09-26 | |
| 16 | S | Grok 4.5 SpaceXAI model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.944 pp; $0.033208/test; source snapshot 2026-09-26; run date not published | 86.88% | 2026-09-26 |
| 17 | GPT 5.5 OpenAI model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.142988/test; source snapshot 2026-09-26; run date not published | 86.87% | 2026-09-26 | |
| 18 | Claude Opus 4.6 (Nonthinking) Anthropic model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.942 pp; $0.115121/test; source snapshot 2026-09-26; run date not published | 86.74% | 2026-09-26 | |
| 19 | Grok 4.6 xAI model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.956 pp; $0.037190/test; source snapshot 2026-09-26; run date not published | 86.53% | 2026-09-26 | |
| 20 | Claude Opus 4.6 (Thinking) Anthropic model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.944 pp; $0.224735/test; source snapshot 2026-09-26; run date not published | 86.13% | 2026-09-26 | |
| 21 | Muse Spark Meta model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not published | 85.90% | 2026-09-26 | |
| 22 | Claude Opus 4.8 Anthropic model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.928 pp; $0.259121/test; source snapshot 2026-09-26; run date not published | 85.75% | 2026-09-26 | |
| 23 | D | DeepSeek V4.1 Flash DeepSeek model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 1.918 pp; $0.015397/test; source snapshot 2026-09-26; run date not published | 85.50% | 2026-09-26 |
| 24 | TM | Inkling Thinking Machines model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.844 pp; $0.165561/test; source snapshot 2026-09-26; run date not published | 85.41% | 2026-09-26 |
| 25 | Claude Opus 4.5 (Thinking) Anthropic model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.896 pp; $0.410224/test; source snapshot 2026-09-26; run date not published | 85.32% | 2026-09-26 | |
| 26 | X | MiMo V2.6 Flash Xiaomi model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.986 pp; $0.002184/test; source snapshot 2026-09-26; run date not published | 85.28% | 2026-09-26 |
| 27 | GPT-5.6 Sol OpenAI model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not published | 85.23% | 2026-09-26 | |
| 28 | Claude Haiku 4.5 (Thinking) Anthropic model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.042375/test; source snapshot 2026-09-26; run date not published | 85.23% | 2026-09-26 | |
| 29 | A | Qwen 3.8 Max Alibaba model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 1.999 pp; $0.089616/test; source snapshot 2026-09-26; run date not published | 84.95% | 2026-09-26 |
| 30 | Claude Sonnet 4.5 (Nonthinking) Anthropic model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.054649/test; source snapshot 2026-09-26; run date not published | 84.52% | 2026-09-26 | |
| 31 | Gemini 3.8 Flash Google model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.943 pp; $0.025238/test; source snapshot 2026-09-26; run date not published | 84.50% | 2026-09-26 | |
| 32 | GPT-5.6 Luna OpenAI model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published | 84.39% | 2026-09-26 | |
| 33 | GPT 5.2 OpenAI model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not published | 84.39% | 2026-09-26 | |
| 34 | TM | Inkling Small Thinking Machines model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.87 pp; $0.019018/test; source snapshot 2026-09-26; run date not published | 84.11% | 2026-09-26 |
| 35 | Claude Sonnet 4.5 (Thinking) Anthropic model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.873 pp; $0.082281/test; source snapshot 2026-09-26; run date not published | 84.10% | 2026-09-26 | |
| 36 | Gemini 3.7 Flash Google model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.004 pp; $0.058736/test; source snapshot 2026-09-26; run date not published | 83.94% | 2026-09-26 | |
| 37 | A | Qwen 3.8 27B Alibaba model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.977 pp; $0.046732/test; source snapshot 2026-09-26; run date not published | 83.85% | 2026-09-26 |
| 38 | X | MiMo V2.5 Pro Xiaomi model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.063 pp; $0.006030/test; source snapshot 2026-09-26; run date not published | 83.73% | 2026-09-26 |
| 39 | GPT-6 Luna OpenAI model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not published | 83.71% | 2026-09-26 | |
| 40 | GPT 5 OpenAI model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not published | 83.65% | 2026-09-26 | |
| 41 | T | Hy4 Preview Tencent model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.065 pp; $0.053243/test; source snapshot 2026-09-26; run date not published | 83.60% | 2026-09-26 |
| 42 | Z | GLM 5.2 zAI model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.002 pp; $0.044912/test; source snapshot 2026-09-26; run date not published | 83.53% | 2026-09-26 |
| 43 | Claude Opus 4.5 (Nonthinking) Anthropic model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.926 pp; $0.281674/test; source snapshot 2026-09-26; run date not published | 83.25% | 2026-09-26 | |
| 44 | Gemini 2.5 Flash (7/17) (Thinking) Google model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.908 pp; $0.014824/test; source snapshot 2026-09-26; run date not published | 82.98% | 2026-09-26 | |
| 45 | Claude Opus 4.7 Anthropic model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.977 pp; $0.177841/test; source snapshot 2026-09-26; run date not published | 82.95% | 2026-09-26 | |
| 46 | Gemini 2.5 Flash (7/17) (Nonthinking) Google model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.909 pp; $0.014869/test; source snapshot 2026-09-26; run date not published | 82.87% | 2026-09-26 | |
| 47 | GPT-5.6 Terra OpenAI model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.948 pp; $0.062588/test; source snapshot 2026-09-26; run date not published | 82.87% | 2026-09-26 | |
| 48 | GPT-6 Sol OpenAI model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 1.942 pp; $0.083138/test; source snapshot 2026-09-26; run date not published | 82.03% | 2026-09-26 | |
| 49 | S | Grok 4 Fast (Reasoning) SpaceXAI model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.137 pp; $0.002535/test; source snapshot 2026-09-26; run date not published | 81.63% | 2026-09-26 |
| 50 | AG | Ling 3.0 Flash Ant Group model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.039 pp; $0.001358/test; source snapshot 2026-09-26; run date not published | 80.90% | 2026-09-26 |
| 51 | MiniMax-M2.1 MiniMax model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.831 pp; $0.005087/test; source snapshot 2026-09-26; run date not published | 80.78% | 2026-09-26 | |
| 52 | GPT 5 Mini OpenAI model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.924 pp; $0.033478/test; source snapshot 2026-09-26; run date not published | 80.58% | 2026-09-26 | |
| 53 | D | DeepSeek V4 Flash 0731 DeepSeek model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 1.973 pp; $0.014247/test; source snapshot 2026-09-26; run date not published | 80.36% | 2026-09-26 |
| 54 | D | DeepSeek V4 Pro 0813 DeepSeek model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.004 pp; $0.041127/test; source snapshot 2026-09-26; run date not published | 80.17% | 2026-09-26 |
| 55 | MiniMax-M2.7 MiniMax model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.86 pp; $0.005124/test; source snapshot 2026-09-26; run date not published | 79.87% | 2026-09-26 | |
| 56 | S | Grok 4 Fast (Non-Reasoning) SpaceXAI model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.871 pp; $0.002056/test; source snapshot 2026-09-26; run date not published | 79.72% | 2026-09-26 |
| 57 | Gemini 3.6 Flash Google model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.861 pp; $0.073178/test; source snapshot 2026-09-26; run date not published | 79.66% | 2026-09-26 | |
| 58 | A | Qwen 3.7 Max Alibaba model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 1.907 pp; $0.069070/test; source snapshot 2026-09-26; run date not published | 79.40% | 2026-09-26 |
| 59 | S | Grok 4.1 Fast (Reasoning) SpaceXAI model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.866 pp; $0.002387/test; source snapshot 2026-09-26; run date not published | 78.73% | 2026-09-26 |
| 60 | Gemini 2.5 Flash Preview (9/25) (Thinking) Google model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.014526/test; source snapshot 2026-09-26; run date not published | 78.50% | 2026-09-26 | |
| 61 | S | Grok 4 SpaceXAI model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.084 pp; $0.063955/test; source snapshot 2026-09-26; run date not published | 78.15% | 2026-09-26 |
| 62 | Kimi K2.6 Moonshot AI model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 1.792 pp; $0.055962/test; source snapshot 2026-09-26; run date not published | 78.15% | 2026-09-26 | |
| 63 | Gemini 2.5 Flash Preview (9/25) (Nonthinking) Google model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.014385/test; source snapshot 2026-09-26; run date not published | 77.95% | 2026-09-26 | |
| 64 | GPT 5.4 (xhigh) OpenAI model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 3.316 pp; $0.639282/test; source snapshot 2026-09-26; run date not published | 77.55% | 2026-09-26 | |
| 65 | S | Grok 4.1 Fast Non-Reasoning SpaceXAI model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.04 pp; $0.001782/test; source snapshot 2026-09-26; run date not published | 77.46% | 2026-09-26 |
| 66 | A | Qwen 3 VL Plus Alibaba model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.020220/test; source snapshot 2026-09-26; run date not published | 77.13% | 2026-09-26 |
| 67 | GPT 5.4 Nano OpenAI model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.001800/test; source snapshot 2026-09-26; run date not published | 77.09% | 2026-09-26 | |
| 68 | A | Qwen 3.6 Plus Alibaba model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 1.917 pp; $0.029294/test; source snapshot 2026-09-26; run date not published | 76.96% | 2026-09-26 |
| 69 | o3 OpenAI model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.871 pp; $0.040334/test; source snapshot 2026-09-26; run date not published | 76.65% | 2026-09-26 | |
| 70 | Gemini 3.5 Flash Google model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.166341/test; source snapshot 2026-09-26; run date not published | 76.57% | 2026-09-26 | |
| 71 | Kimi K2.5 Moonshot AI model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 1.986 pp; $0.024890/test; source snapshot 2026-09-26; run date not published | 76.44% | 2026-09-26 | |
| 72 | Gemini 3.1 Pro Preview (02/26) Google model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not published | 76.11% | 2026-09-26 | |
| 73 | Claude Sonnet 5 Anthropic model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not published | 76.05% | 2026-09-26 | |
| 74 | Gemini 2.5 Flash Lite (9/25) (Nonthinking) Google model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.851 pp; $0.001332/test; source snapshot 2026-09-26; run date not published | 75.82% | 2026-09-26 | |
| 75 | AG | Ling 3.0 Flash Fin Ant Group model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 2.026 pp; $0.001618/test; source snapshot 2026-09-26; run date not published | 75.59% | 2026-09-26 |
| 76 | D | DeepSeek V4 DeepSeek model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.002 pp; $0.053954/test; source snapshot 2026-09-26; run date not published | 75.14% | 2026-09-26 |
| 77 | S | Grok 4.3 SpaceXAI model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.019 pp; $0.015293/test; source snapshot 2026-09-26; run date not published | 74.40% | 2026-09-26 |
| 78 | Claude Opus 4.1 (Thinking) Anthropic model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 1.965 pp; $0.263427/test; source snapshot 2026-09-26; run date not published | 73.90% | 2026-09-26 | |
| 79 | Gemini 2.5 Pro Google model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 1.91 pp; $0.046379/test; source snapshot 2026-09-26; run date not published | 73.55% | 2026-09-26 | |
| 80 | GPT 5 Nano OpenAI model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.006961/test; source snapshot 2026-09-26; run date not published | 72.86% | 2026-09-26 | |
| 81 | Gemini 2.5 Flash Lite (Nonthinking) Google model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.982 pp; $0.001211/test; source snapshot 2026-09-26; run date not published | 72.83% | 2026-09-26 | |
| 82 | A | Qwen 3 Max Thinking Alibaba model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.905 pp; $0.085327/test; source snapshot 2026-09-26; run date not published | 72.71% | 2026-09-26 |
| 83 | Claude Sonnet 4 (Nonthinking) Anthropic model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.038973/test; source snapshot 2026-09-26; run date not published | 72.41% | 2026-09-26 | |
| 84 | Z | GLM 5.1 zAI model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.064 pp; $0.023717/test; source snapshot 2026-09-26; run date not published | 72.27% | 2026-09-26 |
| 85 | X | MiMo V2.5 Xiaomi model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.851 pp; $0.001351/test; source snapshot 2026-09-26; run date not published | 72.15% | 2026-09-26 |
| 86 | Gemini 3 Pro (11/25) Google model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.9 pp; $0.061162/test; source snapshot 2026-09-26; run date not published | 72.04% | 2026-09-26 | |
| 87 | Claude Opus 4.1 (Nonthinking) Anthropic model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 2.021 pp; $0.187162/test; source snapshot 2026-09-26; run date not published | 71.75% | 2026-09-26 | |
| 88 | Gemini 3.5 Flash Lite Google model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.031 pp; $0.019867/test; source snapshot 2026-09-26; run date not published | 70.89% | 2026-09-26 | |
| 89 | A | Qwen 3.5 Flash Alibaba model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 2.09 pp; $0.004425/test; source snapshot 2026-09-26; run date not published | 70.62% | 2026-09-26 |
| 90 | Gemini 3 Flash (12/25) Google model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.014379/test; source snapshot 2026-09-26; run date not published | 69.92% | 2026-09-26 | |
| 91 | Claude Sonnet 4 (Thinking) Anthropic model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 2.212 pp; $0.053443/test; source snapshot 2026-09-26; run date not published | 69.35% | 2026-09-26 | |
| 92 | o4 Mini OpenAI model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.957 pp; $0.040605/test; source snapshot 2026-09-26; run date not published | 69.14% | 2026-09-26 | |
| 93 | Z | GLM 4.7 zAI model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.123 pp; $0.019082/test; source snapshot 2026-09-26; run date not published | 68.63% | 2026-09-26 |
| 94 | Mistral Medium 3.5 Mistral model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.011 pp; $0.157641/test; source snapshot 2026-09-26; run date not published | 67.73% | 2026-09-26 | |
| 95 | Gemini 2.5 Flash Lite (9/25) (Thinking) Google model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.002567/test; source snapshot 2026-09-26; run date not published | 66.88% | 2026-09-26 | |
| 96 | P | Laguna M.1 Poolside model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 2.007 pp; $0.002202/test; source snapshot 2026-09-26; run date not published | 65.91% | 2026-09-26 |
| 97 | Gemini 3.1 Flash Lite Preview Google model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.823 pp; $0.002195/test; source snapshot 2026-09-26; run date not published | 63.90% | 2026-09-26 | |
| 98 | S | Grok 4.20 (Reasoning) SpaceXAI model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.095 pp; $0.031303/test; source snapshot 2026-09-26; run date not published | 63.41% | 2026-09-26 |
| 99 | P | Laguna XS.2 Poolside model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 2.349 pp; $0.001126/test; source snapshot 2026-09-26; run date not published | 61.43% | 2026-09-26 |
| 100 | C | Command A+ Cohere model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 3.646 pp; $0.140316/test; source snapshot 2026-09-26; run date not published | 55.68% | 2026-09-26 |
| 101 | I | Mercury 2.5 Inception model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.095 pp; $0.004476/test; source snapshot 2026-09-26; run date not published | 55.09% | 2026-09-26 |
| 102 | Llama 4 Maverick Meta model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.871 pp; $0.002460/test; source snapshot 2026-09-26; run date not published | 54.22% | 2026-09-26 | |
| 103 | Llama 4 Scout Meta model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not published | 50.59% | 2026-09-26 | |
| 104 | Nemotron 3.5 Lightning NVIDIA model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 0.492 pp; $0.006241/test; source snapshot 2026-09-26; run date not published | 4.27% | 2026-09-26 | |
Scores preserve their source precision, with any scale conversion documented (independently run). Vals AI runs this benchmark. Full Overall data contains 104 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Vals AI (dataset with Protege) |
|---|---|
| category | documentation and coding benchmarks |
| released | 2026-02 |
| size | 100 rubric-scored SOAP-note cases |
| scale | percentage accuracy 0-100, higher better |
| result basis | independently run |
| source | Vals AI MedScribe leaderboard |
| last frontier result | 2026-09-26 |
What is MedScribe (Vals AI)?
MedScribe (Vals AI) is a documentation and coding benchmark from Vals AI, released 2026-02: 100 rubric-scored SOAP-note cases, scored on a percentage accuracy 0-100 scale. Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance.
Which model leads MedScribe (Vals AI)?
Claude Opus 5.5 (Anthropic) has the highest indexed numerical score on MedScribe (Vals AI) at 91.43% (evaluation setups may differ), per Vals AI MedScribe leaderboard, as of 2026-09-26.
Where do the MedScribe (Vals AI) numbers come from?
From Vals AI MedScribe leaderboard (independently run). Vals AI runs this benchmark. Full Overall data contains 104 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.