Health Evals

Anthropic logoClaude Sonnet 4 (Thinking): healthcare benchmark results

Anthropic · 3 boards · snapshot reviewed September 28, 2026

The index currently holds 5 results for Claude Sonnet 4 (Thinking): 34.96% on MedCode (Vals AI) (74 of 102 indexed rows), 33.94% on MedCode (Vals AI) [Claude Sonnet 4 (Nonthinking)] (78 of 102 indexed rows), 72.41% on MedScribe (Vals AI) [Claude Sonnet 4 (Nonthinking)] (83 of 104 indexed rows), 69.35% on MedScribe (Vals AI) (91 of 104 indexed rows), 48.1% on WHBench [Claude Sonnet 4] (15 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
MedCode (Vals AI)
model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 1.939 pp; $0.069896/test; source snapshot 2026-09-26; run date not published
34.96%74 of 1022026-09-26
MedCode (Vals AI)
Claude Sonnet 4 (Nonthinking)
model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.906 pp; $0.039460/test; source snapshot 2026-09-26; run date not published
33.94%78 of 1022026-09-26
MedScribe (Vals AI)
Claude Sonnet 4 (Nonthinking)
model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.038973/test; source snapshot 2026-09-26; run date not published
72.41%83 of 1042026-09-26
MedScribe (Vals AI)
model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 2.212 pp; $0.053443/test; source snapshot 2026-09-26; run date not published
69.35%91 of 1042026-09-26
WHBench
Claude Sonnet 4
95% bootstrap CI 45.5–50.6; 3 runs; temperature 0; zero-shot, closed-book
48.1%15 of 222026-03

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Sonnet 4 (Nonthinking)" and "Claude Sonnet 4".

Which healthcare benchmarks is Claude Sonnet 4 (Thinking) scored on?

As of September 28, 2026, Claude Sonnet 4 (Thinking) has indexed results on 3 tracked benchmarks: MedCode (Vals AI), MedScribe (Vals AI), WHBench.

Which results are indexed for Claude Sonnet 4 (Thinking)?

Claude Sonnet 4 (Thinking) stands at 34.96% on MedCode (Vals AI) (74 of 102 indexed rows), 33.94% on MedCode (Vals AI) [Claude Sonnet 4 (Nonthinking)] (78 of 102 indexed rows), 72.41% on MedScribe (Vals AI) [Claude Sonnet 4 (Nonthinking)] (83 of 104 indexed rows), 69.35% on MedScribe (Vals AI) (91 of 104 indexed rows), 48.1% on WHBench [Claude Sonnet 4] (15 of 22 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.