HealthBench Professional: published results
OpenAI · 525 physician-authored tasks · index updated September 28, 2026
GPT-6 Astra (Anthropic run) has the highest indexed numerical score on HealthBench Professional, 0.703 as of 2026-09, per Claude Sonnet 5.5 System Card. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.
This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchprofessional.com. Every published row in the index appears below.
Published results
Showing top 20 of 29 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-6 Astra (Anthropic run) OpenAI Anthropic reproduction of GPT-6 Astra through the public API; max effort; no system prompt; Claude Opus 4.8 grader; length-adjusted 70.3%, raw 74.0%. Different grader/protocol from OpenAI’s own 64.7% report. | 0.703 | 2026-09 | |
| 2 | Claude Sonnet 5.5 Anthropic Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%. | 0.692 | 2026-09 | |
| 3 | Claude Opus 5.5 Anthropic Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. Five-trial average; refusal fallback to Claude Opus 5. | 0.656 | 2026-09 | |
| 4 | GPT-6 Astra OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 68.2%, mean answer 3,185 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.647 | 2026-09 | |
| 5 | Claude Fable 5 Anthropic Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 68.9%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5. | 0.633 | 2026-09 | |
| 6 | Claude Fable 5.1 Anthropic length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%). | 0.621 | 2026-09 | |
| 7 | GPT-6 Sol OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 59.5%, mean answer 1,573 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.608 | 2026-09 | |
| 8 | GPT-6 Luna OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.608 | 2026-09 | |
| 9 | GPT-5.6 Sol OpenAI length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492. | 0.605 | 2026-06 | |
| 10 | Claude Opus 5 Anthropic length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%). | 0.598 | 2026-07 | |
| 11 | Muse Spark 1.1 Meta length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44) | 0.593 | 2026-07 | |
| 12 | Claude Sonnet 5 Anthropic length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%). | 0.578 | 2026-06 | |
| 13 | GPT-5.6 Terra OpenAI length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA. | 0.577 | 2026-06 | |
| 14 | Claude Opus 4.8 Anthropic Length-adjusted; Anthropic June evaluation, adaptive max effort, Claude Opus 4.8 grader, five-trial average, no tools or custom system prompt. Earlier May report was 55.8% using Claude Sonnet 4.6 as grader; the grader changed. | 0.574 | 2026-06 | |
| 15 | Grok 4.7 xAI SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed. | 0.567 | 2026-09 | |
| 16 | GPT-5.6 Luna OpenAI length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493. | 0.557 | 2026-06 | |
| 17 | Muse Spark Meta length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44 | 0.541 | 2026-07 | |
| 18 | GPT-5.6 Sol (August) OpenAI ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars) | 0.540 | 2026-08 | |
| 19 | Claude Opus 4.7 Anthropic length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card | 0.519 | 2026-05 | |
| 20 | GPT-5.5 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars) | 0.518 | 2026-06 | |
| 21 | Grok 4.6 xAI SpaceXAI vendor report at high effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed. | 0.485 | 2026-09 | |
| 22 | GPT-5.4 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars) | 0.481 | 2026-06 | |
| 23 | GPT-5 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars) | 0.462 | 2026-06 | |
| 24 | GPT-5.2 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars) | 0.459 | 2026-06 | |
| 25 | Claude Sonnet 4.6 Anthropic length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader) | 0.442 | 2026-06 | |
| 26 | GPT-5.6 Luna (August) OpenAI ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars) | 0.441 | 2026-08 | |
| 27 | GPT-5.1 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars) | 0.396 | 2026-06 | |
| 28 | GPT-5.5 Instant OpenAI length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT. | 0.384 | 2026-05 | |
| 29 | MAI-Thinking-1 Microsoft length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision. | 0.350 | 2026-08 | |
Scores preserve their source precision, with any scale conversion documented (mixed sources). All displayed values use a 0–1 scale (source percentages divided by 100). Most are length-adjusted, but grader models, safeguards and evaluator protocols differ. Anthropic’s Astra reproduction (Opus 4.8 grader) is a separate row from OpenAI’s own run. Grok’s release table does not fully document its grading protocol. A score-source check is not an independent rerun. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | OpenAI |
|---|---|
| category | rubric-graded benchmarks |
| released | 2026-04-22 |
| size | 525 physician-authored tasks |
| scale | 0 to 1, higher is better |
| result basis | mixed sources |
| source | Published model evaluation reports |
| official page | arxiv.org/abs/2604.27470 |
| last frontier result | 2026-09 |
What is HealthBench Professional?
HealthBench Professional is a rubric-graded benchmark from OpenAI, released 2026-04-22: 525 physician-authored tasks, scored on a 0 to 1 scale. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.
Which model leads HealthBench Professional?
GPT-6 Astra (Anthropic run) (OpenAI) has the highest indexed numerical score on HealthBench Professional at 0.703 (evaluation setups may differ), per Claude Sonnet 5.5 System Card, as of 2026-09.
Where do the HealthBench Professional numbers come from?
From Published model evaluation reports (mixed sources). All displayed values use a 0–1 scale (source percentages divided by 100). Most are length-adjusted, but grader models, safeguards and evaluator protocols differ. Anthropic’s Astra reproduction (Opus 4.8 grader) is a separate row from OpenAI’s own run. Grok’s release table does not fully document its grading protocol. A score-source check is not an independent rerun.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.