MedCode (Vals AI): published results
Vals AI (dataset with Protege) · 2,755 patient records · index updated September 28, 2026
Claude Opus 5 has the highest indexed numerical score on MedCode (Vals AI), 63.57% as of 2026-09-26, per Vals AI MedCode leaderboard. ICD-10-CM diagnosis coding for entire hospital stays: models assign primary and secondary codes from discharge summaries plus progress/consult notes; ground truth double-annotated by certified professional coders.
Published results
Showing top 20 of 102 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.156845/test; source snapshot 2026-09-26; run date not published | 63.57% | 2026-09-26 | |
| 2 | Gemini 3.1 Pro Preview (02/26) Google model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not published | 59.06% | 2026-09-26 | |
| 3 | Claude Fable 5 Anthropic model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.203 pp; $0.591071/test; source snapshot 2026-09-26; run date not published | 56.07% | 2026-09-26 | |
| 4 | Gemini 3 Flash (12/25) Google model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.112 pp; $0.006187/test; source snapshot 2026-09-26; run date not published | 55.92% | 2026-09-26 | |
| 5 | Gemini 3.5 Flash Google model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.073716/test; source snapshot 2026-09-26; run date not published | 55.83% | 2026-09-26 | |
| 6 | Claude Opus 4.7 Anthropic model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.205 pp; $0.226314/test; source snapshot 2026-09-26; run date not published | 54.86% | 2026-09-26 | |
| 7 | Claude Fable 5.1 Anthropic model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.165 pp; $1.116862/test; source snapshot 2026-09-26; run date not published | 53.51% | 2026-09-26 | |
| 8 | Gemini 3.7 Flash Google model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.12 pp; $0.038331/test; source snapshot 2026-09-26; run date not published | 53.39% | 2026-09-26 | |
| 9 | Claude Opus 4.8 Anthropic model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.165 pp; $0.350925/test; source snapshot 2026-09-26; run date not published | 53.22% | 2026-09-26 | |
| 10 | Gemini 3.6 Flash Google model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.157 pp; $0.044216/test; source snapshot 2026-09-26; run date not published | 53.15% | 2026-09-26 | |
| 11 | Claude Sonnet 5.5 Anthropic model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.119 pp; $0.391592/test; source snapshot 2026-09-26; run date not published | 52.92% | 2026-09-26 | |
| 12 | GPT 5.1 OpenAI model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 2.151 pp; $0.014371/test; source snapshot 2026-09-26; run date not published | 52.73% | 2026-09-26 | |
| 13 | Gemini 3 Pro (11/25) Google model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.073 pp; $0.028248/test; source snapshot 2026-09-26; run date not published | 52.20% | 2026-09-26 | |
| 14 | Muse Spark Meta model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not published | 51.31% | 2026-09-26 | |
| 15 | Gemini 2.5 Pro Google model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.015389/test; source snapshot 2026-09-26; run date not published | 50.59% | 2026-09-26 | |
| 16 | Claude Opus 5.5 Anthropic model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.273 pp; $0.658237/test; source snapshot 2026-09-26; run date not published | 49.80% | 2026-09-26 | |
| 17 | GPT 5.2 OpenAI model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not published | 49.75% | 2026-09-26 | |
| 18 | GPT 5 OpenAI model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not published | 49.63% | 2026-09-26 | |
| 19 | S | Grok 4.7 SpaceXAI model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not published | 49.55% | 2026-09-26 |
| 20 | Muse Spark 1.2 Meta model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.187 pp; $0.039023/test; source snapshot 2026-09-26; run date not published | 49.35% | 2026-09-26 | |
| 21 | Claude Opus 4.5 (Thinking) Anthropic model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 2.012 pp; $0.095846/test; source snapshot 2026-09-26; run date not published | 49.16% | 2026-09-26 | |
| 22 | Claude Opus 4.6 (Thinking) Anthropic model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.085 pp; $0.244127/test; source snapshot 2026-09-26; run date not published | 49.13% | 2026-09-26 | |
| 23 | GPT 5.5 OpenAI model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.188 pp; $0.160759/test; source snapshot 2026-09-26; run date not published | 49.10% | 2026-09-26 | |
| 24 | Kimi K3 Moonshot AI model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not published | 48.88% | 2026-09-26 | |
| 25 | GPT-6 Astra OpenAI model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 2.131 pp; $0.451358/test; source snapshot 2026-09-26; run date not published | 48.49% | 2026-09-26 | |
| 26 | Claude Opus 4.6 (Nonthinking) Anthropic model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.05 pp; $0.006180/test; source snapshot 2026-09-26; run date not published | 48.24% | 2026-09-26 | |
| 27 | Gemini 3.8 Flash Google model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.18 pp; $0.017700/test; source snapshot 2026-09-26; run date not published | 48.13% | 2026-09-26 | |
| 28 | Gemini 3.1 Flash Lite Preview Google model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.071 pp; $0.002029/test; source snapshot 2026-09-26; run date not published | 47.60% | 2026-09-26 | |
| 29 | Claude Sonnet 5 Anthropic model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not published | 47.54% | 2026-09-26 | |
| 30 | o3 OpenAI model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.161 pp; $0.029818/test; source snapshot 2026-09-26; run date not published | 47.29% | 2026-09-26 | |
| 31 | Claude Opus 4.1 (Thinking) Anthropic model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 2.067 pp; $0.269254/test; source snapshot 2026-09-26; run date not published | 47.23% | 2026-09-26 | |
| 32 | GPT-6 Sol OpenAI model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 2.119 pp; $0.085827/test; source snapshot 2026-09-26; run date not published | 47.07% | 2026-09-26 | |
| 33 | MiniMax-M3 MiniMax model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.104 pp; $0.012125/test; source snapshot 2026-09-26; run date not published | 46.29% | 2026-09-26 | |
| 34 | Claude Opus 4.5 (Nonthinking) Anthropic model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.006826/test; source snapshot 2026-09-26; run date not published | 45.17% | 2026-09-26 | |
| 35 | X | MiMo V2.6 Pro Xiaomi model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.097 pp; $0.009528/test; source snapshot 2026-09-26; run date not published | 44.97% | 2026-09-26 |
| 36 | S | Grok 4.6 SpaceXAI model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.256 pp; $0.050335/test; source snapshot 2026-09-26; run date not published | 44.71% | 2026-09-26 |
| 37 | GPT-6 Luna OpenAI model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not published | 44.69% | 2026-09-26 | |
| 38 | Claude Sonnet 4.5 (Thinking) Anthropic model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.101495/test; source snapshot 2026-09-26; run date not published | 44.13% | 2026-09-26 | |
| 39 | GPT-5.6 Sol OpenAI model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not published | 43.97% | 2026-09-26 | |
| 40 | Gemini 3.5 Flash Lite Google model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.951 pp; $0.008093/test; source snapshot 2026-09-26; run date not published | 43.49% | 2026-09-26 | |
| 41 | GPT-5.6 Terra OpenAI model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.173 pp; $0.046558/test; source snapshot 2026-09-26; run date not published | 43.41% | 2026-09-26 | |
| 42 | S | Grok 4.5 SpaceXAI model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.313 pp; $0.048461/test; source snapshot 2026-09-26; run date not published | 43.29% | 2026-09-26 |
| 43 | T | Hy4 Preview Tencent model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.134 pp; $0.062053/test; source snapshot 2026-09-26; run date not published | 43.25% | 2026-09-26 |
| 44 | GPT 5 Mini OpenAI model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.045 pp; $0.005560/test; source snapshot 2026-09-26; run date not published | 43.05% | 2026-09-26 | |
| 45 | Z | GLM 5.3 zAI model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.111 pp; $0.081905/test; source snapshot 2026-09-26; run date not published | 42.86% | 2026-09-26 |
| 46 | D | DeepSeek V4 Pro 0813 DeepSeek model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.16 pp; $0.061060/test; source snapshot 2026-09-26; run date not published | 42.47% | 2026-09-26 |
| 47 | GPT-5.6 Luna OpenAI model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published | 42.39% | 2026-09-26 | |
| 48 | Z | GLM 5.1 zAI model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.024196/test; source snapshot 2026-09-26; run date not published | 41.60% | 2026-09-26 |
| 49 | D | DeepSeek V4 Flash 0731 DeepSeek model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 2.15 pp; $0.019659/test; source snapshot 2026-09-26; run date not published | 41.41% | 2026-09-26 |
| 50 | Claude Opus 4.1 (Nonthinking) Anthropic model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.206270/test; source snapshot 2026-09-26; run date not published | 41.37% | 2026-09-26 | |
| 51 | GPT 5.4 (xhigh) OpenAI model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.148 pp; $0.212108/test; source snapshot 2026-09-26; run date not published | 41.29% | 2026-09-26 | |
| 52 | TM | Inkling Thinking Machines model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.23 pp; $0.127025/test; source snapshot 2026-09-26; run date not published | 41.19% | 2026-09-26 |
| 53 | D | DeepSeek V4.1 Flash DeepSeek model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 2.042 pp; $0.012450/test; source snapshot 2026-09-26; run date not published | 41.17% | 2026-09-26 |
| 54 | X | MiMo V2.6 Flash Xiaomi model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.032 pp; $0.002865/test; source snapshot 2026-09-26; run date not published | 41.06% | 2026-09-26 |
| 55 | GPT 5.4 Nano OpenAI model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 2.256 pp; $0.000844/test; source snapshot 2026-09-26; run date not published | 41.03% | 2026-09-26 | |
| 56 | Z | GLM 5.2 zAI model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.166 pp; $0.045010/test; source snapshot 2026-09-26; run date not published | 40.77% | 2026-09-26 |
| 57 | A | Qwen 3.8 Max Alibaba model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 2.029 pp; $0.120827/test; source snapshot 2026-09-26; run date not published | 40.67% | 2026-09-26 |
| 58 | Claude Sonnet 4.5 (Nonthinking) Anthropic model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.995 pp; $0.042403/test; source snapshot 2026-09-26; run date not published | 40.57% | 2026-09-26 | |
| 59 | Gemini 2.5 Flash Preview (9/25) (Nonthinking) Google model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.003692/test; source snapshot 2026-09-26; run date not published | 40.54% | 2026-09-26 | |
| 60 | D | DeepSeek V4 DeepSeek model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.122 pp; $0.060710/test; source snapshot 2026-09-26; run date not published | 40.45% | 2026-09-26 |
| 61 | Gemini 2.5 Flash (7/17) (Thinking) Google model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.952 pp; $0.003661/test; source snapshot 2026-09-26; run date not published | 40.36% | 2026-09-26 | |
| 62 | Gemini 2.5 Flash Preview (9/25) (Thinking) Google model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.003653/test; source snapshot 2026-09-26; run date not published | 40.33% | 2026-09-26 | |
| 63 | Kimi K2.6 Moonshot AI model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 2.041 pp; $0.041295/test; source snapshot 2026-09-26; run date not published | 40.14% | 2026-09-26 | |
| 64 | Kimi K2.5 Moonshot AI model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 2.119 pp; $0.017275/test; source snapshot 2026-09-26; run date not published | 39.32% | 2026-09-26 | |
| 65 | A | Qwen 3.7 Max Alibaba model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 2.196 pp; $0.042362/test; source snapshot 2026-09-26; run date not published | 38.75% | 2026-09-26 |
| 66 | Nemotron 3 Ultra NVIDIA model ID nvidia/nemotron-3-ultra-550b-a55b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.001 pp; source snapshot 2026-09-26; run date not published | 38.62% | 2026-09-26 | |
| 67 | Gemini 2.5 Flash (7/17) (Nonthinking) Google model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.003698/test; source snapshot 2026-09-26; run date not published | 38.42% | 2026-09-26 | |
| 68 | S | Grok 4 SpaceXAI model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.206 pp; $0.034103/test; source snapshot 2026-09-26; run date not published | 38.08% | 2026-09-26 |
| 69 | S | Grok 4.3 SpaceXAI model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.081 pp; $0.022202/test; source snapshot 2026-09-26; run date not published | 38.07% | 2026-09-26 |
| 70 | TM | Inkling Small Thinking Machines model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.206 pp; $0.016656/test; source snapshot 2026-09-26; run date not published | 37.89% | 2026-09-26 |
| 71 | S | Grok 4 Fast (Reasoning) SpaceXAI model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.941 pp; $0.002143/test; source snapshot 2026-09-26; run date not published | 37.38% | 2026-09-26 |
| 72 | A | Qwen 3.6 Plus Alibaba model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 2.017 pp; $0.015673/test; source snapshot 2026-09-26; run date not published | 36.89% | 2026-09-26 |
| 73 | Llama 4 Maverick Meta model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.994 pp; $0.002888/test; source snapshot 2026-09-26; run date not published | 36.51% | 2026-09-26 | |
| 74 | Claude Sonnet 4 (Thinking) Anthropic model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 1.939 pp; $0.069896/test; source snapshot 2026-09-26; run date not published | 34.96% | 2026-09-26 | |
| 75 | MiniMax-M2.7 MiniMax model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.985 pp; $0.007424/test; source snapshot 2026-09-26; run date not published | 34.44% | 2026-09-26 | |
| 76 | Gemini 2.5 Flash Lite (9/25) (Thinking) Google model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.736 pp; $0.001182/test; source snapshot 2026-09-26; run date not published | 34.19% | 2026-09-26 | |
| 77 | MiniMax-M2.1 MiniMax model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.943 pp; source snapshot 2026-09-26; run date not published | 34.08% | 2026-09-26 | |
| 78 | Claude Sonnet 4 (Nonthinking) Anthropic model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.906 pp; $0.039460/test; source snapshot 2026-09-26; run date not published | 33.94% | 2026-09-26 | |
| 79 | o4 Mini OpenAI model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.021 pp; $0.017605/test; source snapshot 2026-09-26; run date not published | 33.79% | 2026-09-26 | |
| 80 | Mistral Medium 3.5 Mistral model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.148 pp; $0.053370/test; source snapshot 2026-09-26; run date not published | 33.75% | 2026-09-26 | |
| 81 | A | Qwen 3.5 Flash Alibaba model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 1.787 pp; $0.003934/test; source snapshot 2026-09-26; run date not published | 33.00% | 2026-09-26 |
| 82 | Z | GLM 4.7 zAI model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.996 pp; $0.006710/test; source snapshot 2026-09-26; run date not published | 32.77% | 2026-09-26 |
| 83 | Claude Haiku 4.5 (Thinking) Anthropic model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.020099/test; source snapshot 2026-09-26; run date not published | 32.68% | 2026-09-26 | |
| 84 | X | MiMo V2.5 Pro Xiaomi model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.006719/test; source snapshot 2026-09-26; run date not published | 32.48% | 2026-09-26 |
| 85 | AG | Ling 3.0 Flash Ant Group model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.908 pp; $0.001645/test; source snapshot 2026-09-26; run date not published | 32.27% | 2026-09-26 |
| 86 | S | Grok 4.20 (Reasoning) SpaceXAI model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.036189/test; source snapshot 2026-09-26; run date not published | 32.16% | 2026-09-26 |
| 87 | X | MiMo V2.5 Xiaomi model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.025 pp; $0.002162/test; source snapshot 2026-09-26; run date not published | 31.89% | 2026-09-26 |
| 88 | A | Qwen 3 VL Plus Alibaba model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.845 pp; $0.002519/test; source snapshot 2026-09-26; run date not published | 31.65% | 2026-09-26 |
| 89 | A | Qwen 3 Max Thinking Alibaba model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.014776/test; source snapshot 2026-09-26; run date not published | 31.37% | 2026-09-26 |
| 90 | I | Mercury 2.5 Inception model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.953 pp; $0.004535/test; source snapshot 2026-09-26; run date not published | 31.33% | 2026-09-26 |
| 91 | GPT 5 Nano OpenAI model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.948 pp; $0.001729/test; source snapshot 2026-09-26; run date not published | 30.44% | 2026-09-26 | |
| 92 | S | Grok 4 Fast (Non-Reasoning) SpaceXAI model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.974 pp; $0.002149/test; source snapshot 2026-09-26; run date not published | 30.04% | 2026-09-26 |
| 93 | AG | Ling 3.0 Flash Fin Ant Group model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 1.943 pp; $0.001354/test; source snapshot 2026-09-26; run date not published | 29.30% | 2026-09-26 |
| 94 | A | Qwen 3.8 27B Alibaba model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.971 pp; $0.050145/test; source snapshot 2026-09-26; run date not published | 28.70% | 2026-09-26 |
| 95 | S | Grok 4.1 Fast Non-Reasoning SpaceXAI model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.921 pp; $0.002193/test; source snapshot 2026-09-26; run date not published | 28.35% | 2026-09-26 |
| 96 | S | Grok 4.1 Fast (Reasoning) SpaceXAI model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.992 pp; $0.002108/test; source snapshot 2026-09-26; run date not published | 28.08% | 2026-09-26 |
| 97 | Gemini 2.5 Flash Lite (Nonthinking) Google model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.843 pp; $0.001342/test; source snapshot 2026-09-26; run date not published | 27.11% | 2026-09-26 | |
| 98 | Gemini 2.5 Flash Lite (9/25) (Nonthinking) Google model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.911 pp; $0.001440/test; source snapshot 2026-09-26; run date not published | 27.08% | 2026-09-26 | |
| 99 | Llama 4 Scout Meta model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not published | 23.31% | 2026-09-26 | |
| 100 | P | Laguna M.1 Poolside model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 1.693 pp; $0.002973/test; source snapshot 2026-09-26; run date not published | 23.11% | 2026-09-26 |
| 101 | P | Laguna XS.2 Poolside model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 1.703 pp; $0.001445/test; source snapshot 2026-09-26; run date not published | 21.25% | 2026-09-26 |
| 102 | C | Command A+ Cohere model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 1.835 pp; $0.055695/test; source snapshot 2026-09-26; run date not published | 19.72% | 2026-09-26 |
Scores preserve their source precision, with any scale conversion documented (independently run). Vals AI runs this benchmark. Full Overall data contains 102 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Vals AI (dataset with Protege) |
|---|---|
| category | documentation and coding benchmarks |
| released | 2026-02 |
| size | 2,755 patient records |
| scale | percentage accuracy 0-100, higher better |
| result basis | independently run |
| source | Vals AI MedCode leaderboard |
| last frontier result | 2026-09-26 |
What is MedCode (Vals AI)?
MedCode (Vals AI) is a documentation and coding benchmark from Vals AI, released 2026-02: 2,755 patient records, scored on a percentage accuracy 0-100 scale. ICD-10-CM diagnosis coding for entire hospital stays: models assign primary and secondary codes from discharge summaries plus progress/consult notes; ground truth double-annotated by certified professional coders.
Which model leads MedCode (Vals AI)?
Claude Opus 5 (Anthropic) has the highest indexed numerical score on MedCode (Vals AI) at 63.57% (evaluation setups may differ), per Vals AI MedCode leaderboard, as of 2026-09-26.
Where do the MedCode (Vals AI) numbers come from?
From Vals AI MedCode leaderboard (independently run). Vals AI runs this benchmark. Full Overall data contains 102 scored model configurations in the 2026-09-26 snapshot, including older and reasoning variants. Scores use the published percent scale; numeric values retain source precision and displayed values round to two decimals. Snapshot update dates are not model measurement dates.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.