HealthBench Hard: published results
OpenAI · 1,000 conversations · index updated September 28, 2026
Baichuan-M3 has the highest indexed numerical score on HealthBench Hard, 0.444 as of 2026-02, per Baichuan-M3 Technical Report. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.
This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchhard.ai. Every published row in the index appears below.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | B | Baichuan-M3 Baichuan Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol. | 0.444 | 2026-02 |
| 2 | Muse Spark Meta raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table. | 0.428 | 2026-04 | |
| 3 | GPT-5.2-High (Baichuan run) OpenAI Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol. | 0.420 | 2026-02 | |
| 4 | GPT-6 Astra OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 34.2%, mean answer 1,697 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.366 | 2026-09 | |
| 5 | GPT-5 OpenAI length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking. | 0.347 | 2026-06 | |
| 6 | GPT-5.2 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars) | 0.343 | 2026-06 | |
| 7 | GPT-5.6 Sol OpenAI length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490. | 0.331 | 2026-06 | |
| 8 | GPT-5.6 Terra OpenAI length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA. | 0.327 | 2026-06 | |
| 9 | GPT-5.6 Luna OpenAI length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491. | 0.320 | 2026-06 | |
| 10 | GPT-5.5 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars) | 0.315 | 2026-06 | |
| 11 | GPT-5.6 Sol (August) OpenAI ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars) | 0.314 | 2026-08 | |
| 12 | GPT-6 Luna OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.314 | 2026-09 | |
| 13 | GPT-6 Sol OpenAI OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 22.1%, mean answer 974 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.301 | 2026-09 | |
| 14 | GPT OSS 120B OpenAI raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board. | 0.300 | 2025-08 | |
| 15 | GPT-5.4 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars) | 0.291 | 2026-06 | |
| 16 | GPT-5.6 Luna (August) OpenAI ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars) | 0.287 | 2026-08 | |
| 17 | GPT-5.3 Chat OpenAI raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model. | 0.259 | 2026-03 | |
| 18 | GPT-5.1 OpenAI length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars) | 0.254 | 2026-06 | |
| 19 | GPT-5.5 Instant OpenAI length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT. | 0.229 | 2026-05 | |
| 20 | GPT OSS 20B OpenAI raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board. | 0.108 | 2025-08 | |
Scores preserve their source precision, with any scale conversion documented (mixed sources). The 0–1 display divides source percentages by 100. OpenAI’s newer rows are length-adjusted; Meta, Baichuan, gpt-oss and GPT-5.3 launch rows report raw scores. These protocols are not directly interchangeable. The September 22 Astra correction and GPT-6 Sol/Luna appendix are included. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | OpenAI |
|---|---|
| category | rubric-graded benchmarks |
| released | 2025-05-12 |
| size | 1,000 conversations |
| scale | 0 to 1, higher is better |
| result basis | mixed sources |
| source | healthbenchhard.ai |
| paper | arxiv.org/abs/2505.08775 |
| last frontier result | 2026-09 |
What is HealthBench Hard?
HealthBench Hard is a rubric-graded benchmark from OpenAI, released 2025-05-12: 1,000 conversations, scored on a 0 to 1 scale. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.
Which model leads HealthBench Hard?
Baichuan-M3 (Baichuan) has the highest indexed numerical score on HealthBench Hard at 0.444 (evaluation setups may differ), per Baichuan-M3 Technical Report, as of 2026-02.
Where do the HealthBench Hard numbers come from?
From healthbenchhard.ai (mixed sources). The 0–1 display divides source percentages by 100. OpenAI’s newer rows are length-adjusted; Meta, Baichuan, gpt-oss and GPT-5.3 launch rows report raw scores. These protocols are not directly interchangeable. The September 22 Astra correction and GPT-6 Sol/Luna appendix are included.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.