Health Evals

Anthropic logoAnthropic: healthcare benchmark results

31 models · 104 results · snapshot reviewed September 28, 2026

Everything the index holds for Anthropic, gathered in one place: Claude Opus 5, Claude Fable 5, Claude Fable 5.1, Claude Sonnet 5, Claude Opus 4.6, Claude Opus 5.5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 4.6, Claude Opus 4.5 (Thinking), Claude Sonnet 4 (Thinking), Claude Opus 4.1 (Thinking), Claude Sonnet 4.5 (Thinking), Claude Haiku 4.5 (Thinking), Claude 3.7 Sonnet, Claude Code (Opus 5), Claude Code (Opus 4.8), Claude Code (Opus 4.7), Claude Code (Opus 4.6), Claude Code (Sonnet 4.6), claude-code + claude-opus-5, claude-code + claude-opus-4-8, claude-code + claude-opus-4-6, claude-code + claude-sonnet-4-6, claude-code + claude-opus-4-7, claude-code + claude-fable-5, claude-code + claude-sonnet-5, openclaw + claude-opus-4-7, claude-code + claude-haiku-4-5, Claude Opus 4. The lab has the highest indexed score on MedCode (Vals AI), Health Optimization Bench, WHBench, HealthAdminBench, MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index, PhysicianBench, HealthBench, HealthAgentBench. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Claude Opus 5
length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
HealthBench Professional0.59810 of 29
Claude Opus 5
Anthropic evaluation; length-adjusted 57.8%, raw 67.1%; adaptive max effort, Opus 4.8 grader, five-trial average, no tools or customized system prompt. Previously this catalog displayed the raw value.
HealthBench57.810 of 26
Claude Opus 5
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 66.2–72.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench69.32 of 16
Claude Opus 5MAST (Medical AI Superintelligence Test)57.1%6 of 8
source board: 11
Claude Opus 5
v2 run on ARISE; 19 models on the board
First, Do NOHARM (v2)74.6%6 of 17
source board: 19
Claude Opus 5
model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.156845/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)63.57%1 of 102
Claude Opus 5
model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.236275/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)90.98%4 of 104
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index533 of 25
source board: 77
Claude Opus 5 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench57.6%4 of 21
Claude Fable 5
Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 68.9%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
HealthBench Professional0.6335 of 29
Claude Fable 5
Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 61.2%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
HealthBench60.45 of 26
Claude Fable 5
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 68.0–73.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench70.91 of 16
Claude Fable 5First, Do NOHARM (v2)65.0%11 of 17
source board: 19
Claude Fable 5
model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.203 pp; $0.591071/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)56.07%3 of 102
Claude Fable 5
model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.945 pp; $0.583239/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)88.52%10 of 104
Claude Fable 5
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
MedXpertQA (MM)80.04 of 22
Claude Fable 5.1
length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
HealthBench Professional0.6216 of 29
Claude Fable 5.1
length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%).
HealthBench60%6 of 26
Claude Fable 5.1
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 42.2–52.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Vendor safeguards declined 97/257 tasks, scored with no credit; mean over answered tasks is 75.9.
Health Optimization Bench47.38 of 16
Claude Fable 5.1
model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.165 pp; $1.116862/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)53.51%7 of 102
Claude Fable 5.1
model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.953 pp; $0.963500/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)91.29%2 of 104
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index582 of 25
source board: 77
Claude Fable 5.1 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench61.0%3 of 21
Claude Sonnet 5
length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
HealthBench Professional0.57812 of 29
Claude Sonnet 5
length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 card
HealthBench58.7%8 of 26
Claude Sonnet 5
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench34.611 of 16
Claude Sonnet 5MAST (Medical AI Superintelligence Test)56.6%7 of 8
source board: 11
Claude Sonnet 5
model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)47.54%29 of 102
Claude Sonnet 5
model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)76.05%73 of 104
Claude Sonnet 5 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench37.4%8 of 21
Claude 4.6 OpusMedHELM0.4568 of 10
source board: 11
Claude Opus 4.6 (Thinking)
model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.085 pp; $0.244127/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.13%22 of 102
Claude Opus 4.6 (Nonthinking)
model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.05 pp; $0.006180/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)48.24%26 of 102
Claude Opus 4.6 (Nonthinking)
model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.942 pp; $0.115121/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)86.74%18 of 104
Claude Opus 4.6 (Thinking)
model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.944 pp; $0.224735/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)86.13%20 of 104
Claude Opus 4.6
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
MedXpertQA (MM)64.8%15 of 22
Claude Opus 4.6
Pass^3 18.0
PhysicianBench31.7 ± 2.39 of 21
Claude Opus 4.6
95% CI 69.6-74.4; evaluations run March 2026
WHBench72.1%1 of 22
Claude Opus 4.6 (computer-use agent)
screenshot-only, task description + portal guidance; native CUA harness; subtask rate 78.4%
HealthAdminBench36.3%1 of 7
Claude Opus 4.6 (standardized harness)
screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
HealthAdminBench14.8%4 of 7
Claude Opus 5.5
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. Five-trial average; refusal fallback to Claude Opus 5.
HealthBench Professional0.6563 of 29
Claude Opus 5.5
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 68.1%. Five-trial average; refusal fallback to Claude Opus 5.
HealthBench60.64 of 26
Claude Opus 5.5
model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.273 pp; $0.658237/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.80%16 of 102
Claude Opus 5.5
model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.932 pp; $1.154156/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)91.43%1 of 104
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index611 of 25
source board: 77
Claude Opus 5.5 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench68.4%1 of 21
Claude Sonnet 5.5
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.
HealthBench Professional0.6922 of 29
Claude Sonnet 5.5
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 69.4%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.
HealthBench65.41 of 26
Claude Sonnet 5.5
model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.119 pp; $0.391592/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)52.92%11 of 102
Claude Sonnet 5.5
model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.96 pp; $0.508604/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)91.10%3 of 104
Claude Sonnet 5.5 (max)
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench63.2%2 of 21
Claude Sonnet 5.5 (xhigh)
Anthropic-run pass@1 on 100 tasks; xhigh effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench56.4%5 of 21
Claude Sonnet 5.5 (high)
Anthropic-run pass@1 on 100 tasks; high effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench47.6%6 of 21
Claude Sonnet 5.5 (medium)
Anthropic-run pass@1 on 100 tasks; medium effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench30.0%10 of 21
Claude Sonnet 5.5 (low)
Anthropic-run pass@1 on 100 tasks; low effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
PhysicianBench27.2%13 of 21
Claude Opus 4.8
Length-adjusted; Anthropic June evaluation, adaptive max effort, Claude Opus 4.8 grader, five-trial average, no tools or custom system prompt. Earlier May report was 55.8% using Claude Sonnet 4.6 as grader; the grader changed.
HealthBench Professional0.57414 of 29
Claude Opus 4.8
length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itself
HealthBench59.37 of 26
Claude Opus 4.8
model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.165 pp; $0.350925/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)53.22%9 of 102
Claude Opus 4.8
model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.928 pp; $0.259121/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.75%22 of 104
Claude Opus 4.8
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
MedXpertQA (MM)71.79 of 22
Claude Opus 4.7
length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
HealthBench Professional0.51919 of 29
Claude Opus 4.7
model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.205 pp; $0.226314/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)54.86%6 of 102
Claude Opus 4.7
model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.977 pp; $0.177841/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)82.95%45 of 104
Claude Opus 4.7
Pass^3 18.0
PhysicianBench29.3 ± 2.511 of 21
Claude Sonnet 4.6
length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
HealthBench Professional0.44225 of 29
Claude Sonnet 4.6PhysicianBench23.0 ± 2.614 of 21
Claude Sonnet 4.6
validation configuration
EHR-Complex0.3613 of 18
Claude Sonnet 4.6
95% CI 64.5-69.6
WHBench67.1%2 of 22
Claude Opus 4.5 (Thinking)
model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 2.012 pp; $0.095846/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.16%21 of 102
Claude Opus 4.5 (Nonthinking)
model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.006826/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)45.17%34 of 102
Claude Opus 4.5 (Thinking)
model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.896 pp; $0.410224/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.32%25 of 104
Claude Opus 4.5 (Nonthinking)
model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.926 pp; $0.281674/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)83.25%43 of 104
Claude Opus 4.5
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
MedXpertQA (MM)63.6%16 of 22
Claude Sonnet 4 (Thinking)
model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 1.939 pp; $0.069896/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)34.96%74 of 102
Claude Sonnet 4 (Nonthinking)
model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.906 pp; $0.039460/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)33.94%78 of 102
Claude Sonnet 4 (Nonthinking)
model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.038973/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)72.41%83 of 104
Claude Sonnet 4 (Thinking)
model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 2.212 pp; $0.053443/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)69.35%91 of 104
Claude Sonnet 4
95% bootstrap CI 45.5–50.6; 3 runs; temperature 0; zero-shot, closed-book
WHBench48.1%15 of 22
Claude Opus 4.1 (Thinking)
model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 2.067 pp; $0.269254/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)47.23%31 of 102
Claude Opus 4.1 (Nonthinking)
model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.206270/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.37%50 of 102
Claude Opus 4.1 (Thinking)
model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 1.965 pp; $0.263427/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)73.90%78 of 104
Claude Opus 4.1 (Nonthinking)
model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 2.021 pp; $0.187162/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)71.75%87 of 104
Claude Sonnet 4.5 (Thinking)
model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.101495/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)44.13%38 of 102
Claude Sonnet 4.5 (Nonthinking)
model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.995 pp; $0.042403/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.57%58 of 102
Claude Sonnet 4.5 (Nonthinking)
model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.054649/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.52%30 of 104
Claude Sonnet 4.5 (Thinking)
model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.873 pp; $0.082281/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.10%35 of 104
Claude Haiku 4.5 (Thinking)
model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.020099/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)32.68%83 of 102
Claude Haiku 4.5 (Thinking)
model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.042375/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.23%28 of 104
Claude 3.7 SonnetMedHELM0.459 of 10
source board: 11
Claude Code (Opus 5)
$3.3/task; harness+model evaluated jointly
HealthAgentBench55%1 of 12
Claude Code (Opus 4.8)
$4.0/task
HealthAgentBench32%6 of 12
Claude Code (Opus 4.7)
$4.8/task
HealthAgentBench27%8 of 12
Claude Code (Opus 4.6)
$4.1/task
HealthAgentBench19%10 of 12
Claude Code (Sonnet 4.6)
$2.9/task
HealthAgentBench17%11 of 12
claude-code + claude-opus-5
All Domains pass@1; PA 20.0%, UM 32.0%, CM 60.0%; submitted 2026-07-24; run date not published
CHI-Bench37.3%3 of 44
source board: 45
claude-code + claude-opus-4-8
All Domains pass@1; PA 32.0%, UM 28.0%, CM 40.0%; submitted 2026-05-28; run date not published
CHI-Bench33.3%4 of 44
source board: 45
claude-code + claude-opus-4-6
All Domains pass@1; PA 20.0%, UM 36.0%, CM 28.0%; submitted 2026-05-01; run date not published
CHI-Bench28.0%5 of 44
source board: 45
claude-code + claude-sonnet-4-6
All Domains pass@1; PA 24.0%, UM 34.7%, CM 20.0%; submitted 2026-05-01; run date not published
CHI-Bench26.2%6 of 44
source board: 45
claude-code + claude-opus-4-7
All Domains pass@1; PA 24.0%, UM 17.3%, CM 32.0%; submitted 2026-05-01; run date not published
CHI-Bench24.4%9 of 44
source board: 45
claude-code + claude-fable-5
All Domains pass@1; PA 24.0%, UM 24.0%, CM 24.0%; submitted 2026-07-22; run date not published
CHI-Bench24.0%10 of 44
source board: 45
claude-code + claude-sonnet-5
All Domains pass@1; PA 24.0%, UM 24.0%, CM 12.0%; submitted 2026-07-06; run date not published
CHI-Bench20.0%13 of 44
source board: 45
openclaw + claude-opus-4-7
All Domains pass@1; PA 18.7%, UM 13.3%, CM 20.0%; submitted 2026-05-01; run date not published
CHI-Bench17.3%17 of 44
source board: 45
claude-code + claude-haiku-4-5
All Domains pass@1; PA 0.0%, UM 14.7%, CM 4.0%; submitted 2026-05-01; run date not published
CHI-Bench6.2%37 of 44
source board: 45
Claude Opus 4
95% bootstrap CI 46.4–51.7; 3 runs; temperature 0; zero-shot, closed-book
WHBench49.1%14 of 22

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Anthropic appear on?

As of September 28, 2026, Anthropic models hold 104 indexed results across 16 tracked benchmarks, through Claude Opus 5, Claude Fable 5, Claude Fable 5.1, Claude Sonnet 5, Claude Opus 4.6, Claude Opus 5.5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 4.6, Claude Opus 4.5 (Thinking), Claude Sonnet 4 (Thinking), Claude Opus 4.1 (Thinking), Claude Sonnet 4.5 (Thinking), Claude Haiku 4.5 (Thinking), Claude 3.7 Sonnet, Claude Code (Opus 5), Claude Code (Opus 4.8), Claude Code (Opus 4.7), Claude Code (Opus 4.6), Claude Code (Sonnet 4.6), claude-code + claude-opus-5, claude-code + claude-opus-4-8, claude-code + claude-opus-4-6, claude-code + claude-sonnet-4-6, claude-code + claude-opus-4-7, claude-code + claude-fable-5, claude-code + claude-sonnet-5, openclaw + claude-opus-4-7, claude-code + claude-haiku-4-5, Claude Opus 4.

Where does Anthropic have the highest indexed score?

Anthropic models have the highest indexed score on MedCode (Vals AI) (Claude Opus 5, 63.57%); Health Optimization Bench (Claude Fable 5, 70.9); WHBench (Claude Opus 4.6, 72.1%); HealthAdminBench (Claude Opus 4.6, 36.3%); MedScribe (Vals AI) (Claude Opus 5.5, 91.43%); Artificial Analysis Healthcare & Medical Index (Claude Opus 5.5, 61); PhysicianBench (Claude Opus 5.5, 68.4%); HealthBench (Claude Sonnet 5.5, 65.4); HealthAgentBench (Claude Code (Opus 5), 55%).

Other labs with pages: OpenAI, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.