Health Evals

Anthropic logoClaude Sonnet 5: healthcare benchmark results

Anthropic · 7 boards · snapshot reviewed September 28, 2026

The index currently holds 7 results for Claude Sonnet 5: 0.578 on HealthBench Professional (12 of 29 indexed rows), 58.7% on HealthBench (8 of 26 indexed rows), 34.6 on Health Optimization Bench (11 of 16 indexed rows), 56.6% on MAST (Medical AI Superintelligence Test) (7 of 8 indexed rows), 47.54% on MedCode (Vals AI) (29 of 102 indexed rows), 76.05% on MedScribe (Vals AI) (73 of 104 indexed rows), 37.4% on PhysicianBench [Claude Sonnet 5 (max)] (8 of 21 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
0.57812 of 292026-06
HealthBench
length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 card
58.7%8 of 262026-06
Health Optimization Bench
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
34.611 of 162026-09
MAST (Medical AI Superintelligence Test)56.6%7 of 8
source board: 11
2026-08
MedCode (Vals AI)
model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not published
47.54%29 of 1022026-09-26
MedScribe (Vals AI)
model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not published
76.05%73 of 1042026-09-26
PhysicianBench
Claude Sonnet 5 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
37.4%8 of 212026-09-28

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Sonnet 5 (max)".

Which healthcare benchmarks is Claude Sonnet 5 scored on?

As of September 28, 2026, Claude Sonnet 5 has indexed results on 7 tracked benchmarks: HealthBench Professional, HealthBench, Health Optimization Bench, MAST (Medical AI Superintelligence Test), MedCode (Vals AI), MedScribe (Vals AI), PhysicianBench.

Which results are indexed for Claude Sonnet 5?

Claude Sonnet 5 stands at 0.578 on HealthBench Professional (12 of 29 indexed rows), 58.7% on HealthBench (8 of 26 indexed rows), 34.6 on Health Optimization Bench (11 of 16 indexed rows), 56.6% on MAST (Medical AI Superintelligence Test) (7 of 8 indexed rows), 47.54% on MedCode (Vals AI) (29 of 102 indexed rows), 76.05% on MedScribe (Vals AI) (73 of 104 indexed rows), 37.4% on PhysicianBench [Claude Sonnet 5 (max)] (8 of 21 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.