Health Evals

Anthropic logoClaude Opus 4.5 (Thinking): healthcare benchmark results

Anthropic · 3 boards · snapshot reviewed September 28, 2026

The index currently holds 5 results for Claude Opus 4.5 (Thinking): 49.16% on MedCode (Vals AI) (21 of 102 indexed rows), 45.17% on MedCode (Vals AI) [Claude Opus 4.5 (Nonthinking)] (34 of 102 indexed rows), 85.32% on MedScribe (Vals AI) (25 of 104 indexed rows), 83.25% on MedScribe (Vals AI) [Claude Opus 4.5 (Nonthinking)] (43 of 104 indexed rows), 63.6% on MedXpertQA (MM) [Claude Opus 4.5] (16 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
MedCode (Vals AI)
model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 2.012 pp; $0.095846/test; source snapshot 2026-09-26; run date not published
49.16%21 of 1022026-09-26
MedCode (Vals AI)
Claude Opus 4.5 (Nonthinking)
model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.006826/test; source snapshot 2026-09-26; run date not published
45.17%34 of 1022026-09-26
MedScribe (Vals AI)
model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.896 pp; $0.410224/test; source snapshot 2026-09-26; run date not published
85.32%25 of 1042026-09-26
MedScribe (Vals AI)
Claude Opus 4.5 (Nonthinking)
model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.926 pp; $0.281674/test; source snapshot 2026-09-26; run date not published
83.25%43 of 1042026-09-26
MedXpertQA (MM)
Claude Opus 4.5
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
63.6%16 of 22

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Opus 4.5 (Nonthinking)" and "Claude Opus 4.5".

Which healthcare benchmarks is Claude Opus 4.5 (Thinking) scored on?

As of September 28, 2026, Claude Opus 4.5 (Thinking) has indexed results on 3 tracked benchmarks: MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM).

Which results are indexed for Claude Opus 4.5 (Thinking)?

Claude Opus 4.5 (Thinking) stands at 49.16% on MedCode (Vals AI) (21 of 102 indexed rows), 45.17% on MedCode (Vals AI) [Claude Opus 4.5 (Nonthinking)] (34 of 102 indexed rows), 85.32% on MedScribe (Vals AI) (25 of 104 indexed rows), 83.25% on MedScribe (Vals AI) [Claude Opus 4.5 (Nonthinking)] (43 of 104 indexed rows), 63.6% on MedXpertQA (MM) [Claude Opus 4.5] (16 of 22 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.