Health Evals

Anthropic logoClaude Opus 4.8: healthcare benchmark results

Anthropic · 5 boards · snapshot reviewed September 28, 2026

The index currently holds 5 results for Claude Opus 4.8: 0.574 on HealthBench Professional (14 of 29 indexed rows), 59.3 on HealthBench (7 of 26 indexed rows), 53.22% on MedCode (Vals AI) (9 of 102 indexed rows), 85.75% on MedScribe (Vals AI) (22 of 104 indexed rows), 71.7 on MedXpertQA (MM) (9 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
Length-adjusted; Anthropic June evaluation, adaptive max effort, Claude Opus 4.8 grader, five-trial average, no tools or custom system prompt. Earlier May report was 55.8% using Claude Sonnet 4.6 as grader; the grader changed.
0.57414 of 292026-06
HealthBench
length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itself
59.37 of 262026-06
MedCode (Vals AI)
model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.165 pp; $0.350925/test; source snapshot 2026-09-26; run date not published
53.22%9 of 1022026-09-26
MedScribe (Vals AI)
model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.928 pp; $0.259121/test; source snapshot 2026-09-26; run date not published
85.75%22 of 1042026-09-26
MedXpertQA (MM)
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
71.79 of 22not reported

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page.

Which healthcare benchmarks is Claude Opus 4.8 scored on?

As of September 28, 2026, Claude Opus 4.8 has indexed results on 5 tracked benchmarks: HealthBench Professional, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM).

Which results are indexed for Claude Opus 4.8?

Claude Opus 4.8 stands at 0.574 on HealthBench Professional (14 of 29 indexed rows), 59.3 on HealthBench (7 of 26 indexed rows), 53.22% on MedCode (Vals AI) (9 of 102 indexed rows), 85.75% on MedScribe (Vals AI) (22 of 104 indexed rows), 71.7 on MedXpertQA (MM) (9 of 22 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.