Claude Sonnet 5: healthcare benchmark results
Anthropic · 7 boards · snapshot reviewed September 28, 2026
The index currently holds 7 results for Claude Sonnet 5: 0.578 on HealthBench Professional (12 of 29 indexed rows), 58.7% on HealthBench (8 of 26 indexed rows), 34.6 on Health Optimization Bench (11 of 16 indexed rows), 56.6% on MAST (Medical AI Superintelligence Test) (7 of 8 indexed rows), 47.54% on MedCode (Vals AI) (29 of 102 indexed rows), 76.05% on MedScribe (Vals AI) (73 of 104 indexed rows), 37.4% on PhysicianBench [Claude Sonnet 5 (max)] (8 of 21 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%). | 0.578 | 12 of 29 | 2026-06 |
| HealthBench length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 card | 58.7% | 8 of 26 | 2026-06 |
| Health Optimization Bench Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 34.6 | 11 of 16 | 2026-09 |
| MAST (Medical AI Superintelligence Test) | 56.6% | 7 of 8 source board: 11 | 2026-08 |
| MedCode (Vals AI) model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not published | 47.54% | 29 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not published | 76.05% | 73 of 104 | 2026-09-26 |
| PhysicianBench Claude Sonnet 5 (max) Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 37.4% | 8 of 21 | 2026-09-28 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Sonnet 5 (max)".
Which healthcare benchmarks is Claude Sonnet 5 scored on?
As of September 28, 2026, Claude Sonnet 5 has indexed results on 7 tracked benchmarks: HealthBench Professional, HealthBench, Health Optimization Bench, MAST (Medical AI Superintelligence Test), MedCode (Vals AI), MedScribe (Vals AI), PhysicianBench.
Which results are indexed for Claude Sonnet 5?
Claude Sonnet 5 stands at 0.578 on HealthBench Professional (12 of 29 indexed rows), 58.7% on HealthBench (8 of 26 indexed rows), 34.6 on Health Optimization Bench (11 of 16 indexed rows), 56.6% on MAST (Medical AI Superintelligence Test) (7 of 8 indexed rows), 47.54% on MedCode (Vals AI) (29 of 102 indexed rows), 76.05% on MedScribe (Vals AI) (73 of 104 indexed rows), 37.4% on PhysicianBench [Claude Sonnet 5 (max)] (8 of 21 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.