Health Evals

HealthBench Hard: published results

OpenAI · 1,000 conversations · index updated September 28, 2026

Baichuan-M3 has the highest indexed numerical score on HealthBench Hard, 0.444 as of 2026-02, per Baichuan-M3 Technical Report. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.

This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchhard.ai. Every published row in the index appears below.

Published results

#modelscoreas of
1BBaichuan-M3 Baichuan
Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.
0.4442026-02
2Meta logoMuse Spark Meta
raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.
0.4282026-04
3OpenAI logoGPT-5.2-High (Baichuan run) OpenAI
Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.
0.4202026-02
4OpenAI logoGPT-6 Astra OpenAI
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 34.2%, mean answer 1,697 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
0.3662026-09
5OpenAI logoGPT-5 OpenAI
length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
0.3472026-06
6OpenAI logoGPT-5.2 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)
0.3432026-06
7OpenAI logoGPT-5.6 Sol OpenAI
length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.
0.3312026-06
8OpenAI logoGPT-5.6 Terra OpenAI
length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
0.3272026-06
9OpenAI logoGPT-5.6 Luna OpenAI
length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
0.3202026-06
10OpenAI logoGPT-5.5 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)
0.3152026-06
11OpenAI logoGPT-5.6 Sol (August) OpenAI
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)
0.3142026-08
12OpenAI logoGPT-6 Luna OpenAI
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
0.3142026-09
13OpenAI logoGPT-6 Sol OpenAI
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 22.1%, mean answer 974 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
0.3012026-09
14OpenAI logoGPT OSS 120B OpenAI
raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
0.3002025-08
15OpenAI logoGPT-5.4 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
0.2912026-06
16OpenAI logoGPT-5.6 Luna (August) OpenAI
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)
0.2872026-08
17OpenAI logoGPT-5.3 Chat OpenAI
raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
0.2592026-03
18OpenAI logoGPT-5.1 OpenAI
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)
0.2542026-06
19OpenAI logoGPT-5.5 Instant OpenAI
length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
0.2292026-05
20OpenAI logoGPT OSS 20B OpenAI
raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
0.1082025-08

Scores preserve their source precision, with any scale conversion documented (mixed sources). The 0–1 display divides source percentages by 100. OpenAI’s newer rows are length-adjusted; Meta, Baichuan, gpt-oss and GPT-5.3 launch rows report raw scores. These protocols are not directly interchangeable. The September 22 Astra correction and GPT-6 Sol/Luna appendix are included. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherOpenAI
categoryrubric-graded benchmarks
released2025-05-12
size1,000 conversations
scale0 to 1, higher is better
result basismixed sources
sourcehealthbenchhard.ai
paperarxiv.org/abs/2505.08775
last frontier result2026-09

What is HealthBench Hard?

HealthBench Hard is a rubric-graded benchmark from OpenAI, released 2025-05-12: 1,000 conversations, scored on a 0 to 1 scale. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.

Which model leads HealthBench Hard?

Baichuan-M3 (Baichuan) has the highest indexed numerical score on HealthBench Hard at 0.444 (evaluation setups may differ), per Baichuan-M3 Technical Report, as of 2026-02.

Where do the HealthBench Hard numbers come from?

From healthbenchhard.ai (mixed sources). The 0–1 display divides source percentages by 100. OpenAI’s newer rows are length-adjusted; Meta, Baichuan, gpt-oss and GPT-5.3 launch rows report raw scores. These protocols are not directly interchangeable. The September 22 Astra correction and GPT-6 Sol/Luna appendix are included.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.