Claude Fable 5.1: healthcare benchmark results
Anthropic · 7 boards · snapshot reviewed September 28, 2026
The index currently holds 7 results for Claude Fable 5.1: 0.621 on HealthBench Professional (6 of 29 indexed rows), 60% on HealthBench (6 of 26 indexed rows), 47.3 on Health Optimization Bench (8 of 16 indexed rows), 53.51% on MedCode (Vals AI) (7 of 102 indexed rows), 91.29% on MedScribe (Vals AI) (2 of 104 indexed rows), 58 on Artificial Analysis Healthcare & Medical Index [Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)] (2 of 25 indexed rows), 61.0% on PhysicianBench [Claude Fable 5.1 (max)] (3 of 21 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%). | 0.621 | 6 of 29 | 2026-09 |
| HealthBench length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%). | 60% | 6 of 26 | 2026-09 |
| Health Optimization Bench Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 42.2–52.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Vendor safeguards declined 97/257 tasks, scored with no credit; mean over answered tasks is 75.9. | 47.3 | 8 of 16 | 2026-09 |
| MedCode (Vals AI) model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.165 pp; $1.116862/test; source snapshot 2026-09-26; run date not published | 53.51% | 7 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.953 pp; $0.963500/test; source snapshot 2026-09-26; run date not published | 91.29% | 2 of 104 | 2026-09-26 |
| Artificial Analysis Healthcare & Medical Index Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 58 | 2 of 25 source board: 77 | 2026-09 |
| PhysicianBench Claude Fable 5.1 (max) Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 61.0% | 3 of 21 | 2026-09-28 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)" and "Claude Fable 5.1 (max)".
Which healthcare benchmarks is Claude Fable 5.1 scored on?
As of September 28, 2026, Claude Fable 5.1 has indexed results on 7 tracked benchmarks: HealthBench Professional, HealthBench, Health Optimization Bench, MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index, PhysicianBench.
Which results are indexed for Claude Fable 5.1?
Claude Fable 5.1 stands at 0.621 on HealthBench Professional (6 of 29 indexed rows), 60% on HealthBench (6 of 26 indexed rows), 47.3 on Health Optimization Bench (8 of 16 indexed rows), 53.51% on MedCode (Vals AI) (7 of 102 indexed rows), 91.29% on MedScribe (Vals AI) (2 of 104 indexed rows), 58 on Artificial Analysis Healthcare & Medical Index [Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)] (2 of 25 indexed rows), 61.0% on PhysicianBench [Claude Fable 5.1 (max)] (3 of 21 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.