Health Evals

Anthropic logoClaude Fable 5: healthcare benchmark results

Anthropic · 7 boards · snapshot reviewed September 28, 2026

The index currently holds 7 results for Claude Fable 5: 0.633 on HealthBench Professional (5 of 29 indexed rows), 60.4 on HealthBench (5 of 26 indexed rows), 70.9 on Health Optimization Bench (1 of 16 indexed rows), 65.0% on First, Do NOHARM (v2) (11 of 17 indexed rows), 56.07% on MedCode (Vals AI) (3 of 102 indexed rows), 88.52% on MedScribe (Vals AI) (10 of 104 indexed rows), 80.0 on MedXpertQA (MM) (4 of 22 indexed rows). It has the highest indexed score on Health Optimization Bench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 68.9%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
0.6335 of 292026-09
HealthBench
Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 61.2%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
60.45 of 262026-09
Health Optimization Bench
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 68.0–73.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
70.91 of 162026-09
First, Do NOHARM (v2)65.0%11 of 17
source board: 19
2026-08
MedCode (Vals AI)
model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.203 pp; $0.591071/test; source snapshot 2026-09-26; run date not published
56.07%3 of 1022026-09-26
MedScribe (Vals AI)
model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.945 pp; $0.583239/test; source snapshot 2026-09-26; run date not published
88.52%10 of 1042026-09-26
MedXpertQA (MM)
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
80.04 of 22not reported

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page.

Which healthcare benchmarks is Claude Fable 5 scored on?

As of September 28, 2026, Claude Fable 5 has indexed results on 7 tracked benchmarks: HealthBench Professional, HealthBench, Health Optimization Bench, First, Do NOHARM (v2), MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM).

Which results are indexed for Claude Fable 5?

Claude Fable 5 stands at 0.633 on HealthBench Professional (5 of 29 indexed rows), 60.4 on HealthBench (5 of 26 indexed rows), 70.9 on Health Optimization Bench (1 of 16 indexed rows), 65.0% on First, Do NOHARM (v2) (11 of 17 indexed rows), 56.07% on MedCode (Vals AI) (3 of 102 indexed rows), 88.52% on MedScribe (Vals AI) (10 of 104 indexed rows), 80.0 on MedXpertQA (MM) (4 of 22 indexed rows). It has the highest indexed score on Health Optimization Bench.

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.