Health Evals

Anthropic logoClaude Opus 5.5: healthcare benchmark results

Anthropic · 6 boards · snapshot reviewed September 28, 2026

The index currently holds 6 results for Claude Opus 5.5: 0.656 on HealthBench Professional (3 of 29 indexed rows), 60.6 on HealthBench (4 of 26 indexed rows), 49.80% on MedCode (Vals AI) (16 of 102 indexed rows), 91.43% on MedScribe (Vals AI) (1 of 104 indexed rows), 61 on Artificial Analysis Healthcare & Medical Index [Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)] (1 of 25 indexed rows), 68.4% on PhysicianBench [Claude Opus 5.5 (max)] (1 of 21 indexed rows). It has the highest indexed score on MedScribe (Vals AI) and Artificial Analysis Healthcare & Medical Index and PhysicianBench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. Five-trial average; refusal fallback to Claude Opus 5.
0.6563 of 292026-09
HealthBench
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 68.1%. Five-trial average; refusal fallback to Claude Opus 5.
60.64 of 262026-09
MedCode (Vals AI)
model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.273 pp; $0.658237/test; source snapshot 2026-09-26; run date not published
49.80%16 of 1022026-09-26
MedScribe (Vals AI)
model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.932 pp; $1.154156/test; source snapshot 2026-09-26; run date not published
91.43%1 of 1042026-09-26
Artificial Analysis Healthcare & Medical Index
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
611 of 25
source board: 77
2026-09
PhysicianBench
Claude Opus 5.5 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
68.4%1 of 212026-09-28

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)" and "Claude Opus 5.5 (max)".

Which healthcare benchmarks is Claude Opus 5.5 scored on?

As of September 28, 2026, Claude Opus 5.5 has indexed results on 6 tracked benchmarks: HealthBench Professional, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index, PhysicianBench.

Which results are indexed for Claude Opus 5.5?

Claude Opus 5.5 stands at 0.656 on HealthBench Professional (3 of 29 indexed rows), 60.6 on HealthBench (4 of 26 indexed rows), 49.80% on MedCode (Vals AI) (16 of 102 indexed rows), 91.43% on MedScribe (Vals AI) (1 of 104 indexed rows), 61 on Artificial Analysis Healthcare & Medical Index [Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)] (1 of 25 indexed rows), 68.4% on PhysicianBench [Claude Opus 5.5 (max)] (1 of 21 indexed rows). It has the highest indexed score on MedScribe (Vals AI) and Artificial Analysis Healthcare & Medical Index and PhysicianBench.

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.