Health Evals

MedCode (Vals AI): published results

Vals AI (dataset with Protege) · 2,755 patient records · index updated September 28, 2026

Claude Opus 5 has the highest indexed numerical score on MedCode (Vals AI), 63.57% as of 2026-09-26, per Vals AI MedCode leaderboard. ICD-10-CM diagnosis coding for entire hospital stays: models assign primary and secondary codes from discharge summaries plus progress/consult notes; ground truth double-annotated by certified professional coders.

Published results

#modelscoreas of
1Anthropic logoClaude Opus 5 Anthropic
model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.156845/test; source snapshot 2026-09-26; run date not published
63.57%2026-09-26
2Google logoGemini 3.1 Pro Preview (02/26) Google
model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not published
59.06%2026-09-26
3Anthropic logoClaude Fable 5 Anthropic
model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.203 pp; $0.591071/test; source snapshot 2026-09-26; run date not published
56.07%2026-09-26
4Google logoGemini 3 Flash (12/25) Google
model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.112 pp; $0.006187/test; source snapshot 2026-09-26; run date not published
55.92%2026-09-26
5Google logoGemini 3.5 Flash Google
model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.073716/test; source snapshot 2026-09-26; run date not published
55.83%2026-09-26
6Anthropic logoClaude Opus 4.7 Anthropic
model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.205 pp; $0.226314/test; source snapshot 2026-09-26; run date not published
54.86%2026-09-26
7Anthropic logoClaude Fable 5.1 Anthropic
model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.165 pp; $1.116862/test; source snapshot 2026-09-26; run date not published
53.51%2026-09-26
8Google logoGemini 3.7 Flash Google
model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.12 pp; $0.038331/test; source snapshot 2026-09-26; run date not published
53.39%2026-09-26
9Anthropic logoClaude Opus 4.8 Anthropic
model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.165 pp; $0.350925/test; source snapshot 2026-09-26; run date not published
53.22%2026-09-26
10Google logoGemini 3.6 Flash Google
model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.157 pp; $0.044216/test; source snapshot 2026-09-26; run date not published
53.15%2026-09-26
11Anthropic logoClaude Sonnet 5.5 Anthropic
model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.119 pp; $0.391592/test; source snapshot 2026-09-26; run date not published
52.92%2026-09-26
12OpenAI logoGPT 5.1 OpenAI
model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 2.151 pp; $0.014371/test; source snapshot 2026-09-26; run date not published
52.73%2026-09-26
13Google logoGemini 3 Pro (11/25) Google
model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.073 pp; $0.028248/test; source snapshot 2026-09-26; run date not published
52.20%2026-09-26
14Meta logoMuse Spark Meta
model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not published
51.31%2026-09-26
15Google logoGemini 2.5 Pro Google
model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.015389/test; source snapshot 2026-09-26; run date not published
50.59%2026-09-26
16Anthropic logoClaude Opus 5.5 Anthropic
model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.273 pp; $0.658237/test; source snapshot 2026-09-26; run date not published
49.80%2026-09-26
17OpenAI logoGPT 5.2 OpenAI
model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not published
49.75%2026-09-26
18OpenAI logoGPT 5 OpenAI
model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not published
49.63%2026-09-26
19SGrok 4.7 SpaceXAI
model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not published
49.55%2026-09-26
20Meta logoMuse Spark 1.2 Meta
model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.187 pp; $0.039023/test; source snapshot 2026-09-26; run date not published
49.35%2026-09-26
21Anthropic logoClaude Opus 4.5 (Thinking) Anthropic
model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 2.012 pp; $0.095846/test; source snapshot 2026-09-26; run date not published
49.16%2026-09-26
22Anthropic logoClaude Opus 4.6 (Thinking) Anthropic
model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.085 pp; $0.244127/test; source snapshot 2026-09-26; run date not published
49.13%2026-09-26
23OpenAI logoGPT 5.5 OpenAI
model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.188 pp; $0.160759/test; source snapshot 2026-09-26; run date not published
49.10%2026-09-26
24Moonshot AI logoKimi K3 Moonshot AI
model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not published
48.88%2026-09-26
25OpenAI logoGPT-6 Astra OpenAI
model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 2.131 pp; $0.451358/test; source snapshot 2026-09-26; run date not published
48.49%2026-09-26
26Anthropic logoClaude Opus 4.6 (Nonthinking) Anthropic
model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.05 pp; $0.006180/test; source snapshot 2026-09-26; run date not published
48.24%2026-09-26
27Google logoGemini 3.8 Flash Google
model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.18 pp; $0.017700/test; source snapshot 2026-09-26; run date not published
48.13%2026-09-26
28Google logoGemini 3.1 Flash Lite Preview Google
model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.071 pp; $0.002029/test; source snapshot 2026-09-26; run date not published
47.60%2026-09-26
29Anthropic logoClaude Sonnet 5 Anthropic
model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not published
47.54%2026-09-26
30OpenAI logoo3 OpenAI
model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.161 pp; $0.029818/test; source snapshot 2026-09-26; run date not published
47.29%2026-09-26
31Anthropic logoClaude Opus 4.1 (Thinking) Anthropic
model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 2.067 pp; $0.269254/test; source snapshot 2026-09-26; run date not published
47.23%2026-09-26
32OpenAI logoGPT-6 Sol OpenAI
model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 2.119 pp; $0.085827/test; source snapshot 2026-09-26; run date not published
47.07%2026-09-26
33MiniMax logoMiniMax-M3 MiniMax
model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.104 pp; $0.012125/test; source snapshot 2026-09-26; run date not published
46.29%2026-09-26
34Anthropic logoClaude Opus 4.5 (Nonthinking) Anthropic
model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.006826/test; source snapshot 2026-09-26; run date not published
45.17%2026-09-26
35XMiMo V2.6 Pro Xiaomi
model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.097 pp; $0.009528/test; source snapshot 2026-09-26; run date not published
44.97%2026-09-26
36SGrok 4.6 SpaceXAI
model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.256 pp; $0.050335/test; source snapshot 2026-09-26; run date not published
44.71%2026-09-26
37OpenAI logoGPT-6 Luna OpenAI
model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not published
44.69%2026-09-26
38Anthropic logoClaude Sonnet 4.5 (Thinking) Anthropic
model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.101495/test; source snapshot 2026-09-26; run date not published
44.13%2026-09-26
39OpenAI logoGPT-5.6 Sol OpenAI
model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not published
43.97%2026-09-26
40Google logoGemini 3.5 Flash Lite Google
model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.951 pp; $0.008093/test; source snapshot 2026-09-26; run date not published
43.49%2026-09-26
41OpenAI logoGPT-5.6 Terra OpenAI
model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.173 pp; $0.046558/test; source snapshot 2026-09-26; run date not published
43.41%2026-09-26
42SGrok 4.5 SpaceXAI
model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.313 pp; $0.048461/test; source snapshot 2026-09-26; run date not published
43.29%2026-09-26
43THy4 Preview Tencent
model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.134 pp; $0.062053/test; source snapshot 2026-09-26; run date not published
43.25%2026-09-26
44OpenAI logoGPT 5 Mini OpenAI
model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.045 pp; $0.005560/test; source snapshot 2026-09-26; run date not published
43.05%2026-09-26
45ZGLM 5.3 zAI
model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.111 pp; $0.081905/test; source snapshot 2026-09-26; run date not published
42.86%2026-09-26
46DDeepSeek V4 Pro 0813 DeepSeek
model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.16 pp; $0.061060/test; source snapshot 2026-09-26; run date not published
42.47%2026-09-26
47OpenAI logoGPT-5.6 Luna OpenAI
model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published
42.39%2026-09-26
48ZGLM 5.1 zAI
model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.024196/test; source snapshot 2026-09-26; run date not published
41.60%2026-09-26
49DDeepSeek V4 Flash 0731 DeepSeek
model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 2.15 pp; $0.019659/test; source snapshot 2026-09-26; run date not published
41.41%2026-09-26
50Anthropic logoClaude Opus 4.1 (Nonthinking) Anthropic
model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.206270/test; source snapshot 2026-09-26; run date not published
41.37%2026-09-26
51OpenAI logoGPT 5.4 (xhigh) OpenAI
model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.148 pp; $0.212108/test; source snapshot 2026-09-26; run date not published
41.29%2026-09-26
52TMInkling Thinking Machines
model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.23 pp; $0.127025/test; source snapshot 2026-09-26; run date not published
41.19%2026-09-26
53DDeepSeek V4.1 Flash DeepSeek
model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 2.042 pp; $0.012450/test; source snapshot 2026-09-26; run date not published
41.17%2026-09-26
54XMiMo V2.6 Flash Xiaomi
model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.032 pp; $0.002865/test; source snapshot 2026-09-26; run date not published
41.06%2026-09-26
55OpenAI logoGPT 5.4 Nano OpenAI
model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 2.256 pp; $0.000844/test; source snapshot 2026-09-26; run date not published
41.03%2026-09-26
56ZGLM 5.2 zAI
model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.166 pp; $0.045010/test; source snapshot 2026-09-26; run date not published
40.77%2026-09-26
57AQwen 3.8 Max Alibaba
model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 2.029 pp; $0.120827/test; source snapshot 2026-09-26; run date not published
40.67%2026-09-26
58Anthropic logoClaude Sonnet 4.5 (Nonthinking) Anthropic
model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.995 pp; $0.042403/test; source snapshot 2026-09-26; run date not published
40.57%2026-09-26
59Google logoGemini 2.5 Flash Preview (9/25) (Nonthinking) Google
model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.003692/test; source snapshot 2026-09-26; run date not published
40.54%2026-09-26
60DDeepSeek V4 DeepSeek
model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.122 pp; $0.060710/test; source snapshot 2026-09-26; run date not published
40.45%2026-09-26
61Google logoGemini 2.5 Flash (7/17) (Thinking) Google
model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.952 pp; $0.003661/test; source snapshot 2026-09-26; run date not published
40.36%2026-09-26
62Google logoGemini 2.5 Flash Preview (9/25) (Thinking) Google
model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.003653/test; source snapshot 2026-09-26; run date not published
40.33%2026-09-26
63Moonshot AI logoKimi K2.6 Moonshot AI
model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 2.041 pp; $0.041295/test; source snapshot 2026-09-26; run date not published
40.14%2026-09-26
64Moonshot AI logoKimi K2.5 Moonshot AI
model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 2.119 pp; $0.017275/test; source snapshot 2026-09-26; run date not published
39.32%2026-09-26
65AQwen 3.7 Max Alibaba
model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 2.196 pp; $0.042362/test; source snapshot 2026-09-26; run date not published
38.75%2026-09-26
66NVIDIA logoNemotron 3 Ultra NVIDIA
model ID nvidia/nemotron-3-ultra-550b-a55b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.001 pp; source snapshot 2026-09-26; run date not published
38.62%2026-09-26
67Google logoGemini 2.5 Flash (7/17) (Nonthinking) Google
model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.003698/test; source snapshot 2026-09-26; run date not published
38.42%2026-09-26
68SGrok 4 SpaceXAI
model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.206 pp; $0.034103/test; source snapshot 2026-09-26; run date not published
38.08%2026-09-26
69SGrok 4.3 SpaceXAI
model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.081 pp; $0.022202/test; source snapshot 2026-09-26; run date not published
38.07%2026-09-26
70TMInkling Small Thinking Machines
model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.206 pp; $0.016656/test; source snapshot 2026-09-26; run date not published
37.89%2026-09-26
71SGrok 4 Fast (Reasoning) SpaceXAI
model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.941 pp; $0.002143/test; source snapshot 2026-09-26; run date not published
37.38%2026-09-26
72AQwen 3.6 Plus Alibaba
model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 2.017 pp; $0.015673/test; source snapshot 2026-09-26; run date not published
36.89%2026-09-26
73Meta logoLlama 4 Maverick Meta
model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.994 pp; $0.002888/test; source snapshot 2026-09-26; run date not published
36.51%2026-09-26
74Anthropic logoClaude Sonnet 4 (Thinking) Anthropic
model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 1.939 pp; $0.069896/test; source snapshot 2026-09-26; run date not published
34.96%2026-09-26
75MiniMax logoMiniMax-M2.7 MiniMax
model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.985 pp; $0.007424/test; source snapshot 2026-09-26; run date not published
34.44%2026-09-26
76Google logoGemini 2.5 Flash Lite (9/25) (Thinking) Google
model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.736 pp; $0.001182/test; source snapshot 2026-09-26; run date not published
34.19%2026-09-26
77MiniMax logoMiniMax-M2.1 MiniMax
model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.943 pp; source snapshot 2026-09-26; run date not published
34.08%2026-09-26
78Anthropic logoClaude Sonnet 4 (Nonthinking) Anthropic
model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.906 pp; $0.039460/test; source snapshot 2026-09-26; run date not published
33.94%2026-09-26
79OpenAI logoo4 Mini OpenAI
model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.021 pp; $0.017605/test; source snapshot 2026-09-26; run date not published
33.79%2026-09-26
80Mistral logoMistral Medium 3.5 Mistral
model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.148 pp; $0.053370/test; source snapshot 2026-09-26; run date not published
33.75%2026-09-26
81AQwen 3.5 Flash Alibaba
model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 1.787 pp; $0.003934/test; source snapshot 2026-09-26; run date not published
33.00%2026-09-26
82ZGLM 4.7 zAI
model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.996 pp; $0.006710/test; source snapshot 2026-09-26; run date not published
32.77%2026-09-26
83Anthropic logoClaude Haiku 4.5 (Thinking) Anthropic
model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.020099/test; source snapshot 2026-09-26; run date not published
32.68%2026-09-26
84XMiMo V2.5 Pro Xiaomi
model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.006719/test; source snapshot 2026-09-26; run date not published
32.48%2026-09-26
85AGLing 3.0 Flash Ant Group
model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.908 pp; $0.001645/test; source snapshot 2026-09-26; run date not published
32.27%2026-09-26
86SGrok 4.20 (Reasoning) SpaceXAI
model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.036189/test; source snapshot 2026-09-26; run date not published
32.16%2026-09-26
87XMiMo V2.5 Xiaomi
model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.025 pp; $0.002162/test; source snapshot 2026-09-26; run date not published
31.89%2026-09-26
88AQwen 3 VL Plus Alibaba
model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.845 pp; $0.002519/test; source snapshot 2026-09-26; run date not published
31.65%2026-09-26
89AQwen 3 Max Thinking Alibaba
model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.014776/test; source snapshot 2026-09-26; run date not published
31.37%2026-09-26
90IMercury 2.5 Inception
model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.953 pp; $0.004535/test; source snapshot 2026-09-26; run date not published
31.33%2026-09-26
91OpenAI logoGPT 5 Nano OpenAI
model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.948 pp; $0.001729/test; source snapshot 2026-09-26; run date not published
30.44%2026-09-26
92SGrok 4 Fast (Non-Reasoning) SpaceXAI
model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.974 pp; $0.002149/test; source snapshot 2026-09-26; run date not published
30.04%2026-09-26
93AGLing 3.0 Flash Fin Ant Group
model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 1.943 pp; $0.001354/test; source snapshot 2026-09-26; run date not published
29.30%2026-09-26
94AQwen 3.8 27B Alibaba
model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.971 pp; $0.050145/test; source snapshot 2026-09-26; run date not published
28.70%2026-09-26
95SGrok 4.1 Fast Non-Reasoning SpaceXAI
model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.921 pp; $0.002193/test; source snapshot 2026-09-26; run date not published
28.35%2026-09-26
96SGrok 4.1 Fast (Reasoning) SpaceXAI
model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.992 pp; $0.002108/test; source snapshot 2026-09-26; run date not published
28.08%2026-09-26
97Google logoGemini 2.5 Flash Lite (Nonthinking) Google
model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.843 pp; $0.001342/test; source snapshot 2026-09-26; run date not published
27.11%2026-09-26
98Google logoGemini 2.5 Flash Lite (9/25) (Nonthinking) Google
model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.911 pp; $0.001440/test; source snapshot 2026-09-26; run date not published
27.08%2026-09-26
99Meta logoLlama 4 Scout Meta
model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not published
23.31%2026-09-26
100PLaguna M.1 Poolside
model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 1.693 pp; $0.002973/test; source snapshot 2026-09-26; run date not published
23.11%2026-09-26
101PLaguna XS.2 Poolside
model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 1.703 pp; $0.001445/test; source snapshot 2026-09-26; run date not published
21.25%2026-09-26
102CCommand A+ Cohere
model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 1.835 pp; $0.055695/test; source snapshot 2026-09-26; run date not published
19.72%2026-09-26

Scores preserve their source precision, with any scale conversion documented (independently run). Vals AI runs this benchmark. Full Overall data contains 102 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherVals AI (dataset with Protege)
categorydocumentation and coding benchmarks
released2026-02
size2,755 patient records
scalepercentage accuracy 0-100, higher better
result basisindependently run
sourceVals AI MedCode leaderboard
last frontier result2026-09-26

What is MedCode (Vals AI)?

MedCode (Vals AI) is a documentation and coding benchmark from Vals AI, released 2026-02: 2,755 patient records, scored on a percentage accuracy 0-100 scale. ICD-10-CM diagnosis coding for entire hospital stays: models assign primary and secondary codes from discharge summaries plus progress/consult notes; ground truth double-annotated by certified professional coders.

Which model leads MedCode (Vals AI)?

Claude Opus 5 (Anthropic) has the highest indexed numerical score on MedCode (Vals AI) at 63.57% (evaluation setups may differ), per Vals AI MedCode leaderboard, as of 2026-09-26.

Where do the MedCode (Vals AI) numbers come from?

From Vals AI MedCode leaderboard (independently run). Vals AI runs this benchmark. Full Overall data contains 102 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.