Health Evals

MedScribe (Vals AI): published results

Vals AI (dataset with Protege) · 100 rubric-scored SOAP-note cases · index updated September 28, 2026

Claude Opus 5.5 has the highest indexed numerical score on MedScribe (Vals AI), 91.43% as of 2026-09-26, per Vals AI MedScribe leaderboard. Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance.

Published results

#modelscoreas of
1Anthropic logoClaude Opus 5.5 Anthropic
model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.932 pp; $1.154156/test; source snapshot 2026-09-26; run date not published
91.43%2026-09-26
2Anthropic logoClaude Fable 5.1 Anthropic
model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.953 pp; $0.963500/test; source snapshot 2026-09-26; run date not published
91.29%2026-09-26
3Anthropic logoClaude Sonnet 5.5 Anthropic
model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.96 pp; $0.508604/test; source snapshot 2026-09-26; run date not published
91.10%2026-09-26
4Anthropic logoClaude Opus 5 Anthropic
model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.236275/test; source snapshot 2026-09-26; run date not published
90.98%2026-09-26
5Meta logoMuse Spark 1.2 Meta
model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.037779/test; source snapshot 2026-09-26; run date not published
90.06%2026-09-26
6SGrok 4.7 SpaceXAI
model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not published
89.38%2026-09-26
7ZGLM 5.3 Flash zAI
model ID zai/glm-5.3-flash; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.002235/test; source snapshot 2026-09-26; run date not published
88.94%2026-09-26
8Meta logoMuse Spark 1.1 Meta
model ID meta/muse_spark_1_1; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.95 pp; $0.034628/test; source snapshot 2026-09-26; run date not published
88.89%2026-09-26
9ZGLM 5.3 zAI
model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.999 pp; $0.057236/test; source snapshot 2026-09-26; run date not published
88.81%2026-09-26
10Anthropic logoClaude Fable 5 Anthropic
model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.945 pp; $0.583239/test; source snapshot 2026-09-26; run date not published
88.52%2026-09-26
11XMiMo V2.6 Pro Xiaomi
model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.938 pp; $0.009880/test; source snapshot 2026-09-26; run date not published
88.31%2026-09-26
12OpenAI logoGPT 5.1 OpenAI
model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 1.942 pp; $0.096508/test; source snapshot 2026-09-26; run date not published
88.09%2026-09-26
13Moonshot AI logoKimi K3 Moonshot AI
model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not published
87.96%2026-09-26
14OpenAI logoGPT-6 Astra OpenAI
model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 1.938 pp; $0.581991/test; source snapshot 2026-09-26; run date not published
87.91%2026-09-26
15MiniMax logoMiniMax-M3 MiniMax
model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.957 pp; $0.013748/test; source snapshot 2026-09-26; run date not published
87.25%2026-09-26
16SGrok 4.5 SpaceXAI
model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.944 pp; $0.033208/test; source snapshot 2026-09-26; run date not published
86.88%2026-09-26
17OpenAI logoGPT 5.5 OpenAI
model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.142988/test; source snapshot 2026-09-26; run date not published
86.87%2026-09-26
18Anthropic logoClaude Opus 4.6 (Nonthinking) Anthropic
model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.942 pp; $0.115121/test; source snapshot 2026-09-26; run date not published
86.74%2026-09-26
19xAI logoGrok 4.6 xAI
model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.956 pp; $0.037190/test; source snapshot 2026-09-26; run date not published
86.53%2026-09-26
20Anthropic logoClaude Opus 4.6 (Thinking) Anthropic
model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.944 pp; $0.224735/test; source snapshot 2026-09-26; run date not published
86.13%2026-09-26
21Meta logoMuse Spark Meta
model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not published
85.90%2026-09-26
22Anthropic logoClaude Opus 4.8 Anthropic
model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.928 pp; $0.259121/test; source snapshot 2026-09-26; run date not published
85.75%2026-09-26
23DDeepSeek V4.1 Flash DeepSeek
model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 1.918 pp; $0.015397/test; source snapshot 2026-09-26; run date not published
85.50%2026-09-26
24TMInkling Thinking Machines
model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.844 pp; $0.165561/test; source snapshot 2026-09-26; run date not published
85.41%2026-09-26
25Anthropic logoClaude Opus 4.5 (Thinking) Anthropic
model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.896 pp; $0.410224/test; source snapshot 2026-09-26; run date not published
85.32%2026-09-26
26XMiMo V2.6 Flash Xiaomi
model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.986 pp; $0.002184/test; source snapshot 2026-09-26; run date not published
85.28%2026-09-26
27OpenAI logoGPT-5.6 Sol OpenAI
model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not published
85.23%2026-09-26
28Anthropic logoClaude Haiku 4.5 (Thinking) Anthropic
model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.042375/test; source snapshot 2026-09-26; run date not published
85.23%2026-09-26
29AQwen 3.8 Max Alibaba
model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 1.999 pp; $0.089616/test; source snapshot 2026-09-26; run date not published
84.95%2026-09-26
30Anthropic logoClaude Sonnet 4.5 (Nonthinking) Anthropic
model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.054649/test; source snapshot 2026-09-26; run date not published
84.52%2026-09-26
31Google logoGemini 3.8 Flash Google
model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.943 pp; $0.025238/test; source snapshot 2026-09-26; run date not published
84.50%2026-09-26
32OpenAI logoGPT-5.6 Luna OpenAI
model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published
84.39%2026-09-26
33OpenAI logoGPT 5.2 OpenAI
model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not published
84.39%2026-09-26
34TMInkling Small Thinking Machines
model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.87 pp; $0.019018/test; source snapshot 2026-09-26; run date not published
84.11%2026-09-26
35Anthropic logoClaude Sonnet 4.5 (Thinking) Anthropic
model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.873 pp; $0.082281/test; source snapshot 2026-09-26; run date not published
84.10%2026-09-26
36Google logoGemini 3.7 Flash Google
model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.004 pp; $0.058736/test; source snapshot 2026-09-26; run date not published
83.94%2026-09-26
37AQwen 3.8 27B Alibaba
model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.977 pp; $0.046732/test; source snapshot 2026-09-26; run date not published
83.85%2026-09-26
38XMiMo V2.5 Pro Xiaomi
model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.063 pp; $0.006030/test; source snapshot 2026-09-26; run date not published
83.73%2026-09-26
39OpenAI logoGPT-6 Luna OpenAI
model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not published
83.71%2026-09-26
40OpenAI logoGPT 5 OpenAI
model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not published
83.65%2026-09-26
41THy4 Preview Tencent
model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.065 pp; $0.053243/test; source snapshot 2026-09-26; run date not published
83.60%2026-09-26
42ZGLM 5.2 zAI
model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.002 pp; $0.044912/test; source snapshot 2026-09-26; run date not published
83.53%2026-09-26
43Anthropic logoClaude Opus 4.5 (Nonthinking) Anthropic
model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.926 pp; $0.281674/test; source snapshot 2026-09-26; run date not published
83.25%2026-09-26
44Google logoGemini 2.5 Flash (7/17) (Thinking) Google
model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.908 pp; $0.014824/test; source snapshot 2026-09-26; run date not published
82.98%2026-09-26
45Anthropic logoClaude Opus 4.7 Anthropic
model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.977 pp; $0.177841/test; source snapshot 2026-09-26; run date not published
82.95%2026-09-26
46Google logoGemini 2.5 Flash (7/17) (Nonthinking) Google
model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.909 pp; $0.014869/test; source snapshot 2026-09-26; run date not published
82.87%2026-09-26
47OpenAI logoGPT-5.6 Terra OpenAI
model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.948 pp; $0.062588/test; source snapshot 2026-09-26; run date not published
82.87%2026-09-26
48OpenAI logoGPT-6 Sol OpenAI
model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 1.942 pp; $0.083138/test; source snapshot 2026-09-26; run date not published
82.03%2026-09-26
49SGrok 4 Fast (Reasoning) SpaceXAI
model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.137 pp; $0.002535/test; source snapshot 2026-09-26; run date not published
81.63%2026-09-26
50AGLing 3.0 Flash Ant Group
model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.039 pp; $0.001358/test; source snapshot 2026-09-26; run date not published
80.90%2026-09-26
51MiniMax logoMiniMax-M2.1 MiniMax
model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.831 pp; $0.005087/test; source snapshot 2026-09-26; run date not published
80.78%2026-09-26
52OpenAI logoGPT 5 Mini OpenAI
model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.924 pp; $0.033478/test; source snapshot 2026-09-26; run date not published
80.58%2026-09-26
53DDeepSeek V4 Flash 0731 DeepSeek
model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 1.973 pp; $0.014247/test; source snapshot 2026-09-26; run date not published
80.36%2026-09-26
54DDeepSeek V4 Pro 0813 DeepSeek
model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.004 pp; $0.041127/test; source snapshot 2026-09-26; run date not published
80.17%2026-09-26
55MiniMax logoMiniMax-M2.7 MiniMax
model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.86 pp; $0.005124/test; source snapshot 2026-09-26; run date not published
79.87%2026-09-26
56SGrok 4 Fast (Non-Reasoning) SpaceXAI
model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.871 pp; $0.002056/test; source snapshot 2026-09-26; run date not published
79.72%2026-09-26
57Google logoGemini 3.6 Flash Google
model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.861 pp; $0.073178/test; source snapshot 2026-09-26; run date not published
79.66%2026-09-26
58AQwen 3.7 Max Alibaba
model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 1.907 pp; $0.069070/test; source snapshot 2026-09-26; run date not published
79.40%2026-09-26
59SGrok 4.1 Fast (Reasoning) SpaceXAI
model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.866 pp; $0.002387/test; source snapshot 2026-09-26; run date not published
78.73%2026-09-26
60Google logoGemini 2.5 Flash Preview (9/25) (Thinking) Google
model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.014526/test; source snapshot 2026-09-26; run date not published
78.50%2026-09-26
61SGrok 4 SpaceXAI
model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.084 pp; $0.063955/test; source snapshot 2026-09-26; run date not published
78.15%2026-09-26
62Moonshot AI logoKimi K2.6 Moonshot AI
model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 1.792 pp; $0.055962/test; source snapshot 2026-09-26; run date not published
78.15%2026-09-26
63Google logoGemini 2.5 Flash Preview (9/25) (Nonthinking) Google
model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.014385/test; source snapshot 2026-09-26; run date not published
77.95%2026-09-26
64OpenAI logoGPT 5.4 (xhigh) OpenAI
model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 3.316 pp; $0.639282/test; source snapshot 2026-09-26; run date not published
77.55%2026-09-26
65SGrok 4.1 Fast Non-Reasoning SpaceXAI
model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.04 pp; $0.001782/test; source snapshot 2026-09-26; run date not published
77.46%2026-09-26
66AQwen 3 VL Plus Alibaba
model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.020220/test; source snapshot 2026-09-26; run date not published
77.13%2026-09-26
67OpenAI logoGPT 5.4 Nano OpenAI
model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.001800/test; source snapshot 2026-09-26; run date not published
77.09%2026-09-26
68AQwen 3.6 Plus Alibaba
model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 1.917 pp; $0.029294/test; source snapshot 2026-09-26; run date not published
76.96%2026-09-26
69OpenAI logoo3 OpenAI
model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.871 pp; $0.040334/test; source snapshot 2026-09-26; run date not published
76.65%2026-09-26
70Google logoGemini 3.5 Flash Google
model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.166341/test; source snapshot 2026-09-26; run date not published
76.57%2026-09-26
71Moonshot AI logoKimi K2.5 Moonshot AI
model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 1.986 pp; $0.024890/test; source snapshot 2026-09-26; run date not published
76.44%2026-09-26
72Google logoGemini 3.1 Pro Preview (02/26) Google
model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not published
76.11%2026-09-26
73Anthropic logoClaude Sonnet 5 Anthropic
model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not published
76.05%2026-09-26
74Google logoGemini 2.5 Flash Lite (9/25) (Nonthinking) Google
model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.851 pp; $0.001332/test; source snapshot 2026-09-26; run date not published
75.82%2026-09-26
75AGLing 3.0 Flash Fin Ant Group
model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 2.026 pp; $0.001618/test; source snapshot 2026-09-26; run date not published
75.59%2026-09-26
76DDeepSeek V4 DeepSeek
model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.002 pp; $0.053954/test; source snapshot 2026-09-26; run date not published
75.14%2026-09-26
77SGrok 4.3 SpaceXAI
model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.019 pp; $0.015293/test; source snapshot 2026-09-26; run date not published
74.40%2026-09-26
78Anthropic logoClaude Opus 4.1 (Thinking) Anthropic
model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 1.965 pp; $0.263427/test; source snapshot 2026-09-26; run date not published
73.90%2026-09-26
79Google logoGemini 2.5 Pro Google
model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 1.91 pp; $0.046379/test; source snapshot 2026-09-26; run date not published
73.55%2026-09-26
80OpenAI logoGPT 5 Nano OpenAI
model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.006961/test; source snapshot 2026-09-26; run date not published
72.86%2026-09-26
81Google logoGemini 2.5 Flash Lite (Nonthinking) Google
model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.982 pp; $0.001211/test; source snapshot 2026-09-26; run date not published
72.83%2026-09-26
82AQwen 3 Max Thinking Alibaba
model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.905 pp; $0.085327/test; source snapshot 2026-09-26; run date not published
72.71%2026-09-26
83Anthropic logoClaude Sonnet 4 (Nonthinking) Anthropic
model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.038973/test; source snapshot 2026-09-26; run date not published
72.41%2026-09-26
84ZGLM 5.1 zAI
model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.064 pp; $0.023717/test; source snapshot 2026-09-26; run date not published
72.27%2026-09-26
85XMiMo V2.5 Xiaomi
model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.851 pp; $0.001351/test; source snapshot 2026-09-26; run date not published
72.15%2026-09-26
86Google logoGemini 3 Pro (11/25) Google
model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.9 pp; $0.061162/test; source snapshot 2026-09-26; run date not published
72.04%2026-09-26
87Anthropic logoClaude Opus 4.1 (Nonthinking) Anthropic
model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 2.021 pp; $0.187162/test; source snapshot 2026-09-26; run date not published
71.75%2026-09-26
88Google logoGemini 3.5 Flash Lite Google
model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.031 pp; $0.019867/test; source snapshot 2026-09-26; run date not published
70.89%2026-09-26
89AQwen 3.5 Flash Alibaba
model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 2.09 pp; $0.004425/test; source snapshot 2026-09-26; run date not published
70.62%2026-09-26
90Google logoGemini 3 Flash (12/25) Google
model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.014379/test; source snapshot 2026-09-26; run date not published
69.92%2026-09-26
91Anthropic logoClaude Sonnet 4 (Thinking) Anthropic
model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 2.212 pp; $0.053443/test; source snapshot 2026-09-26; run date not published
69.35%2026-09-26
92OpenAI logoo4 Mini OpenAI
model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.957 pp; $0.040605/test; source snapshot 2026-09-26; run date not published
69.14%2026-09-26
93ZGLM 4.7 zAI
model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.123 pp; $0.019082/test; source snapshot 2026-09-26; run date not published
68.63%2026-09-26
94Mistral logoMistral Medium 3.5 Mistral
model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.011 pp; $0.157641/test; source snapshot 2026-09-26; run date not published
67.73%2026-09-26
95Google logoGemini 2.5 Flash Lite (9/25) (Thinking) Google
model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.002567/test; source snapshot 2026-09-26; run date not published
66.88%2026-09-26
96PLaguna M.1 Poolside
model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 2.007 pp; $0.002202/test; source snapshot 2026-09-26; run date not published
65.91%2026-09-26
97Google logoGemini 3.1 Flash Lite Preview Google
model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.823 pp; $0.002195/test; source snapshot 2026-09-26; run date not published
63.90%2026-09-26
98SGrok 4.20 (Reasoning) SpaceXAI
model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.095 pp; $0.031303/test; source snapshot 2026-09-26; run date not published
63.41%2026-09-26
99PLaguna XS.2 Poolside
model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 2.349 pp; $0.001126/test; source snapshot 2026-09-26; run date not published
61.43%2026-09-26
100CCommand A+ Cohere
model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 3.646 pp; $0.140316/test; source snapshot 2026-09-26; run date not published
55.68%2026-09-26
101IMercury 2.5 Inception
model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.095 pp; $0.004476/test; source snapshot 2026-09-26; run date not published
55.09%2026-09-26
102Meta logoLlama 4 Maverick Meta
model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.871 pp; $0.002460/test; source snapshot 2026-09-26; run date not published
54.22%2026-09-26
103Meta logoLlama 4 Scout Meta
model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not published
50.59%2026-09-26
104NVIDIA logoNemotron 3.5 Lightning NVIDIA
model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 0.492 pp; $0.006241/test; source snapshot 2026-09-26; run date not published
4.27%2026-09-26

Scores preserve their source precision, with any scale conversion documented (independently run). Vals AI runs this benchmark. Full Overall data contains 104 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherVals AI (dataset with Protege)
categorydocumentation and coding benchmarks
released2026-02
size100 rubric-scored SOAP-note cases
scalepercentage accuracy 0-100, higher better
result basisindependently run
sourceVals AI MedScribe leaderboard
last frontier result2026-09-26

What is MedScribe (Vals AI)?

MedScribe (Vals AI) is a documentation and coding benchmark from Vals AI, released 2026-02: 100 rubric-scored SOAP-note cases, scored on a percentage accuracy 0-100 scale. Clinical documentation support: quality of SOAP notes generated from clinical visits, scored against rubrics for documentation quality and compliance.

Which model leads MedScribe (Vals AI)?

Claude Opus 5.5 (Anthropic) has the highest indexed numerical score on MedScribe (Vals AI) at 91.43% (evaluation setups may differ), per Vals AI MedScribe leaderboard, as of 2026-09-26.

Where do the MedScribe (Vals AI) numbers come from?

From Vals AI MedScribe leaderboard (independently run). Vals AI runs this benchmark. Full Overall data contains 104 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.