HealthBench: published results
OpenAI · 5,000 conversations · index updated September 28, 2026
Claude Sonnet 5.5 has the highest indexed numerical score on HealthBench, 65.4 as of 2026-09, per Claude Sonnet 5.5 System Card. 5,000 realistic multi-turn health conversations graded against physician-written rubrics (48,562 criteria) covering accuracy, completeness, context awareness, communication, and instruction following. OpenAI now also reports a length-adjusted variant that penalizes verbosity.
Published results
Showing top 20 of 26 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 Anthropic Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 69.4%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%. | 65.4 | 2026-09 | |
| 2 | B | Baichuan-M3 Baichuan self-run in Baichuan-M3 paper (arXiv 2602.06570) | 65.1 | 2026-02 |
| 3 | GPT-5.2-High OpenAI raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card. | 63.3 | 2026-02 | |
| 4 | Claude Opus 5.5 Anthropic Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 68.1%. Five-trial average; refusal fallback to Claude Opus 5. | 60.6 | 2026-09 | |
| 5 | Claude Fable 5 Anthropic Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 61.2%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5. | 60.4 | 2026-09 | |
| 6 | Claude Fable 5.1 Anthropic length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%). | 60% | 2026-09 | |
| 7 | Claude Opus 4.8 Anthropic length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itself | 59.3 | 2026-06 | |
| 8 | Claude Sonnet 5 Anthropic length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 card | 58.7% | 2026-06 | |
| 9 | GPT-6 Astra OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 56.9%, mean answer 1,760 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 58.3 | 2026-09 | |
| 10 | Claude Opus 5 Anthropic Anthropic evaluation; length-adjusted 57.8%, raw 67.1%; adaptive max effort, Opus 4.8 grader, five-trial average, no tools or customized system prompt. Previously this catalog displayed the raw value. | 57.8 | 2026-07 | |
| 11 | GPT-5 OpenAI OpenAI length-adjusted score, maximum reasoning effort; raw 63.1%, mean answer 2,904 characters; GPT-5.6 card Table 6. | 57.7 | 2026-06 | |
| 12 | GPT OSS 120B OpenAI reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9) | 57.6 | 2025-08 | |
| 13 | GPT-5.6 Sol OpenAI length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09 | 57.0 | 2026-06 | |
| 14 | GPT-5.6 Terra OpenAI length-adjusted (58.7 unadjusted), max reasoning effort | 57.0 | 2026-06 | |
| 15 | GPT-5.2 OpenAI OpenAI length-adjusted score, maximum reasoning effort; raw 60.7%, mean answer 2,645 characters; GPT-5.6 card Table 6. | 56.8 | 2026-06 | |
| 16 | GPT-5.5 OpenAI length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system card | 56.5 | 2026-04 | |
| 17 | GPT-5.6 Luna OpenAI length-adjusted (55.4 unadjusted), max reasoning effort | 55.8 | 2026-06 | |
| 18 | GPT-5.6 Sol (August) OpenAI ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06 | 55.0 | 2026-08 | |
| 19 | GPT-6 Luna OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 50%, mean answer 1,255 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 54.5 | 2026-09 | |
| 20 | GPT-5.3 Chat OpenAI raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted) | 54.1% | 2026-03 | |
| 21 | GPT-5.4 OpenAI OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6. | 54.0 | 2026-06 | |
| 22 | GPT-5.6 Luna (August) OpenAI ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars) | 53.3 | 2026-08 | |
| 23 | GPT-6 Sol OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 47.1%, mean answer 977 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 53.2 | 2026-09 | |
| 24 | GPT-5.5 Instant OpenAI length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11 | 51.4 | 2026-05 | |
| 25 | GPT-5.1 OpenAI OpenAI length-adjusted score, maximum reasoning effort; raw 64.2%, mean answer 4,222 characters; GPT-5.6 card Table 6. | 50.9 | 2026-06 | |
| 26 | GPT OSS 20B OpenAI reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8) | 42.5 | 2025-08 | |
Scores preserve their source precision, with any scale conversion documented (mixed sources). Rows mix documented length-adjusted and raw scores and are not a controlled cross-model comparison. Raw Baichuan, gpt-oss and GPT-5.3 launch results remain labeled in their row settings. New Claude rows use Anthropic’s Opus 4.8 grader; OpenAI reports its own evaluations. The September 22 Astra correction is applied. Use HealthBench Professional for a newer clinician task set. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | OpenAI |
|---|---|
| category | rubric-graded benchmarks |
| released | 2025-05 |
| size | 5,000 conversations |
| scale | 0-100 rubric-point percentage (some sites display 0-1), higher better; length-adjusted and unadjusted variants |
| result basis | mixed sources |
| source | Published model evaluation reports |
| official page | openai.com/index/healthbench |
| last frontier result | 2026-09 |
What is HealthBench?
HealthBench is a rubric-graded benchmark from OpenAI, released 2025-05: 5,000 conversations, scored on a 0-100 rubric-point percentage (some sites display 0-1) scale. 5,000 realistic multi-turn health conversations graded against physician-written rubrics (48,562 criteria) covering accuracy, completeness, context awareness, communication, and instruction following. OpenAI now also reports a length-adjusted variant that penalizes verbosity.
Which model leads HealthBench?
Claude Sonnet 5.5 (Anthropic) has the highest indexed numerical score on HealthBench at 65.4 (evaluation setups may differ), per Claude Sonnet 5.5 System Card, as of 2026-09.
Where do the HealthBench numbers come from?
From Published model evaluation reports (mixed sources). Rows mix documented length-adjusted and raw scores and are not a controlled cross-model comparison. Raw Baichuan, gpt-oss and GPT-5.3 launch results remain labeled in their row settings. New Claude rows use Anthropic’s Opus 4.8 grader; OpenAI reports its own evaluations. The September 22 Astra correction is applied. Use HealthBench Professional for a newer clinician task set.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.