Health Evals

Anthropic logoClaude Opus 4.6: healthcare benchmark results

Anthropic · 7 boards · snapshot reviewed September 28, 2026

The index currently holds 10 results for Claude Opus 4.6: 0.456 on MedHELM [Claude 4.6 Opus] (8 of 10 indexed rows), 49.13% on MedCode (Vals AI) [Claude Opus 4.6 (Thinking)] (22 of 102 indexed rows), 48.24% on MedCode (Vals AI) [Claude Opus 4.6 (Nonthinking)] (26 of 102 indexed rows), 86.74% on MedScribe (Vals AI) [Claude Opus 4.6 (Nonthinking)] (18 of 104 indexed rows), 86.13% on MedScribe (Vals AI) [Claude Opus 4.6 (Thinking)] (20 of 104 indexed rows), 64.8% on MedXpertQA (MM) (15 of 22 indexed rows), 31.7 ± 2.3 on PhysicianBench (9 of 21 indexed rows), 72.1% on WHBench (1 of 22 indexed rows), 36.3% on HealthAdminBench [Claude Opus 4.6 (computer-use agent)] (1 of 7 indexed rows), 14.8% on HealthAdminBench [Claude Opus 4.6 (standardized harness)] (4 of 7 indexed rows). It has the highest indexed score on WHBench and HealthAdminBench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
MedHELM
Claude 4.6 Opus
0.4568 of 10
source board: 11
2026-05
MedCode (Vals AI)
Claude Opus 4.6 (Thinking)
model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.085 pp; $0.244127/test; source snapshot 2026-09-26; run date not published
49.13%22 of 1022026-09-26
MedCode (Vals AI)
Claude Opus 4.6 (Nonthinking)
model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.05 pp; $0.006180/test; source snapshot 2026-09-26; run date not published
48.24%26 of 1022026-09-26
MedScribe (Vals AI)
Claude Opus 4.6 (Nonthinking)
model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.942 pp; $0.115121/test; source snapshot 2026-09-26; run date not published
86.74%18 of 1042026-09-26
MedScribe (Vals AI)
Claude Opus 4.6 (Thinking)
model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.944 pp; $0.224735/test; source snapshot 2026-09-26; run date not published
86.13%20 of 1042026-09-26
MedXpertQA (MM)
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
64.8%15 of 222026-04
PhysicianBench
Pass^3 18.0
31.7 ± 2.39 of 212026-05
WHBench
95% CI 69.6-74.4; evaluations run March 2026
72.1%1 of 222026-03
HealthAdminBench
Claude Opus 4.6 (computer-use agent)
screenshot-only, task description + portal guidance; native CUA harness; subtask rate 78.4%
36.3%1 of 72026-04
HealthAdminBench
Claude Opus 4.6 (standardized harness)
screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
14.8%4 of 72026-04

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude 4.6 Opus" and "Claude Opus 4.6 (Thinking)" and "Claude Opus 4.6 (Nonthinking)" and "Claude Opus 4.6 (computer-use agent)" and "Claude Opus 4.6 (standardized harness)".

Which healthcare benchmarks is Claude Opus 4.6 scored on?

As of September 28, 2026, Claude Opus 4.6 has indexed results on 7 tracked benchmarks: MedHELM, MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM), PhysicianBench, WHBench, HealthAdminBench.

Which results are indexed for Claude Opus 4.6?

Claude Opus 4.6 stands at 0.456 on MedHELM [Claude 4.6 Opus] (8 of 10 indexed rows), 49.13% on MedCode (Vals AI) [Claude Opus 4.6 (Thinking)] (22 of 102 indexed rows), 48.24% on MedCode (Vals AI) [Claude Opus 4.6 (Nonthinking)] (26 of 102 indexed rows), 86.74% on MedScribe (Vals AI) [Claude Opus 4.6 (Nonthinking)] (18 of 104 indexed rows), 86.13% on MedScribe (Vals AI) [Claude Opus 4.6 (Thinking)] (20 of 104 indexed rows), 64.8% on MedXpertQA (MM) (15 of 22 indexed rows), 31.7 ± 2.3 on PhysicianBench (9 of 21 indexed rows), 72.1% on WHBench (1 of 22 indexed rows), 36.3% on HealthAdminBench [Claude Opus 4.6 (computer-use agent)] (1 of 7 indexed rows), 14.8% on HealthAdminBench [Claude Opus 4.6 (standardized harness)] (4 of 7 indexed rows). It has the highest indexed score on WHBench and HealthAdminBench.

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.