Claude Sonnet 5.5: healthcare benchmark results
Anthropic · 5 boards · snapshot reviewed September 28, 2026
The index currently holds 9 results for Claude Sonnet 5.5: 0.692 on HealthBench Professional (2 of 29 indexed rows), 65.4 on HealthBench (1 of 26 indexed rows), 52.92% on MedCode (Vals AI) (11 of 102 indexed rows), 91.10% on MedScribe (Vals AI) (3 of 104 indexed rows), 63.2% on PhysicianBench [Claude Sonnet 5.5 (max)] (2 of 21 indexed rows), 56.4% on PhysicianBench [Claude Sonnet 5.5 (xhigh)] (5 of 21 indexed rows), 47.6% on PhysicianBench [Claude Sonnet 5.5 (high)] (6 of 21 indexed rows), 30.0% on PhysicianBench [Claude Sonnet 5.5 (medium)] (10 of 21 indexed rows), 27.2% on PhysicianBench [Claude Sonnet 5.5 (low)] (13 of 21 indexed rows). It has the highest indexed score on HealthBench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%. | 0.692 | 2 of 29 | 2026-09 |
| HealthBench Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 69.4%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%. | 65.4 | 1 of 26 | 2026-09 |
| MedCode (Vals AI) model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.119 pp; $0.391592/test; source snapshot 2026-09-26; run date not published | 52.92% | 11 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.96 pp; $0.508604/test; source snapshot 2026-09-26; run date not published | 91.10% | 3 of 104 | 2026-09-26 |
| PhysicianBench Claude Sonnet 5.5 (max) Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 63.2% | 2 of 21 | 2026-09-28 |
| PhysicianBench Claude Sonnet 5.5 (xhigh) Anthropic-run pass@1 on 100 tasks; xhigh effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 56.4% | 5 of 21 | 2026-09-28 |
| PhysicianBench Claude Sonnet 5.5 (high) Anthropic-run pass@1 on 100 tasks; high effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 47.6% | 6 of 21 | 2026-09-28 |
| PhysicianBench Claude Sonnet 5.5 (medium) Anthropic-run pass@1 on 100 tasks; medium effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 30.0% | 10 of 21 | 2026-09-28 |
| PhysicianBench Claude Sonnet 5.5 (low) Anthropic-run pass@1 on 100 tasks; low effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 27.2% | 13 of 21 | 2026-09-28 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Sonnet 5.5 (max)" and "Claude Sonnet 5.5 (xhigh)" and "Claude Sonnet 5.5 (high)" and "Claude Sonnet 5.5 (medium)" and "Claude Sonnet 5.5 (low)".
Which healthcare benchmarks is Claude Sonnet 5.5 scored on?
As of September 28, 2026, Claude Sonnet 5.5 has indexed results on 5 tracked benchmarks: HealthBench Professional, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), PhysicianBench.
Which results are indexed for Claude Sonnet 5.5?
Claude Sonnet 5.5 stands at 0.692 on HealthBench Professional (2 of 29 indexed rows), 65.4 on HealthBench (1 of 26 indexed rows), 52.92% on MedCode (Vals AI) (11 of 102 indexed rows), 91.10% on MedScribe (Vals AI) (3 of 104 indexed rows), 63.2% on PhysicianBench [Claude Sonnet 5.5 (max)] (2 of 21 indexed rows), 56.4% on PhysicianBench [Claude Sonnet 5.5 (xhigh)] (5 of 21 indexed rows), 47.6% on PhysicianBench [Claude Sonnet 5.5 (high)] (6 of 21 indexed rows), 30.0% on PhysicianBench [Claude Sonnet 5.5 (medium)] (10 of 21 indexed rows), 27.2% on PhysicianBench [Claude Sonnet 5.5 (low)] (13 of 21 indexed rows). It has the highest indexed score on HealthBench.
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.