Health Evals

HealthBench Professional: published results

OpenAI · 525 physician-authored tasks · index updated September 28, 2026

GPT-6 Astra (Anthropic run) has the highest indexed numerical score on HealthBench Professional, 0.703 as of 2026-09, per Claude Sonnet 5.5 System Card. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.

This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchprofessional.com. Every published row in the index appears below.

Published results

#modelscoreas of
1OpenAI logoGPT-6 Astra (Anthropic run) OpenAI
Anthropic reproduction of GPT-6 Astra through the public API; max effort; no system prompt; Claude Opus 4.8 grader; length-adjusted 70.3%, raw 74.0%. Different grader/protocol from OpenAI’s own 64.7% report.
0.7032026-09
2Anthropic logoClaude Sonnet 5.5 Anthropic
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.
0.6922026-09
3Anthropic logoClaude Opus 5.5 Anthropic
Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. Five-trial average; refusal fallback to Claude Opus 5.
0.6562026-09
4OpenAI logoGPT-6 Astra OpenAI
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 68.2%, mean answer 3,185 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
0.6472026-09
5Anthropic logoClaude Fable 5 Anthropic
Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 68.9%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
0.6332026-09
6Anthropic logoClaude Fable 5.1 Anthropic
length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
0.6212026-09
7OpenAI logoGPT-6 Sol OpenAI
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 59.5%, mean answer 1,573 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
0.6082026-09
8OpenAI logoGPT-6 Luna OpenAI
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
0.6082026-09
9OpenAI logoGPT-5.6 Sol OpenAI
length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
0.6052026-06
10Anthropic logoClaude Opus 5 Anthropic
length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
0.5982026-07
11Meta logoMuse Spark 1.1 Meta
length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
0.5932026-07
12Anthropic logoClaude Sonnet 5 Anthropic
length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
0.5782026-06
13OpenAI logoGPT-5.6 Terra OpenAI
length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
0.5772026-06
14Anthropic logoClaude Opus 4.8 Anthropic
Length-adjusted; Anthropic June evaluation, adaptive max effort, Claude Opus 4.8 grader, five-trial average, no tools or custom system prompt. Earlier May report was 55.8% using Claude Sonnet 4.6 as grader; the grader changed.
0.5742026-06
15xAI logoGrok 4.7 xAI
SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.
0.5672026-09
16OpenAI logoGPT-5.6 Luna OpenAI
length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
0.5572026-06
17Meta logoMuse Spark Meta
length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
0.5412026-07
18OpenAI logoGPT-5.6 Sol (August) OpenAI
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
0.5402026-08
19Anthropic logoClaude Opus 4.7 Anthropic
length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
0.5192026-05
20OpenAI logoGPT-5.5 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
0.5182026-06
21xAI logoGrok 4.6 xAI
SpaceXAI vendor report at high effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.
0.4852026-09
22OpenAI logoGPT-5.4 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
0.4812026-06
23OpenAI logoGPT-5 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
0.4622026-06
24OpenAI logoGPT-5.2 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
0.4592026-06
25Anthropic logoClaude Sonnet 4.6 Anthropic
length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
0.4422026-06
26OpenAI logoGPT-5.6 Luna (August) OpenAI
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
0.4412026-08
27OpenAI logoGPT-5.1 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
0.3962026-06
28OpenAI logoGPT-5.5 Instant OpenAI
length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
0.3842026-05
29Microsoft logoMAI-Thinking-1 Microsoft
length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.
0.3502026-08

Scores preserve their source precision, with any scale conversion documented (mixed sources). All displayed values use a 0–1 scale (source percentages divided by 100). Most are length-adjusted, but grader models, safeguards and evaluator protocols differ. Anthropic’s Astra reproduction (Opus 4.8 grader) is a separate row from OpenAI’s own run. Grok’s release table does not fully document its grading protocol. A score-source check is not an independent rerun. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherOpenAI
categoryrubric-graded benchmarks
released2026-04-22
size525 physician-authored tasks
scale0 to 1, higher is better
result basismixed sources
sourcePublished model evaluation reports
official pagearxiv.org/abs/2604.27470
last frontier result2026-09

What is HealthBench Professional?

HealthBench Professional is a rubric-graded benchmark from OpenAI, released 2026-04-22: 525 physician-authored tasks, scored on a 0 to 1 scale. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.

Which model leads HealthBench Professional?

GPT-6 Astra (Anthropic run) (OpenAI) has the highest indexed numerical score on HealthBench Professional at 0.703 (evaluation setups may differ), per Claude Sonnet 5.5 System Card, as of 2026-09.

Where do the HealthBench Professional numbers come from?

From Published model evaluation reports (mixed sources). All displayed values use a 0–1 scale (source percentages divided by 100). Most are length-adjusted, but grader models, safeguards and evaluator protocols differ. Anthropic’s Astra reproduction (Opus 4.8 grader) is a separate row from OpenAI’s own run. Grok’s release table does not fully document its grading protocol. A score-source check is not an independent rerun.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.