Health Evals

OpenAI logoGPT-5: healthcare benchmark results

OpenAI · 6 boards · snapshot reviewed September 28, 2026

The index currently holds 6 results for GPT-5: 0.462 on HealthBench Professional (23 of 29 indexed rows), 0.347 on HealthBench Hard (5 of 20 indexed rows), 57.7 on HealthBench (11 of 26 indexed rows), 68.6% on First, Do NOHARM (v2) (10 of 17 indexed rows), 49.63% on MedCode (Vals AI) [GPT 5] (18 of 102 indexed rows), 83.65% on MedScribe (Vals AI) [GPT 5] (40 of 104 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
0.46223 of 292026-06
HealthBench Hard
length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
0.3475 of 202026-06
HealthBench
OpenAI length-adjusted score, maximum reasoning effort; raw 63.1%, mean answer 2,904 characters; GPT-5.6 card Table 6.
57.711 of 262026-06
First, Do NOHARM (v2)
from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
68.6%10 of 17
source board: 19
2026-08
MedCode (Vals AI)
GPT 5
model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not published
49.63%18 of 1022026-09-26
MedScribe (Vals AI)
GPT 5
model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not published
83.65%40 of 1042026-09-26

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT 5".

Which healthcare benchmarks is GPT-5 scored on?

As of September 28, 2026, GPT-5 has indexed results on 6 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, First, Do NOHARM (v2), MedCode (Vals AI), MedScribe (Vals AI).

Which results are indexed for GPT-5?

GPT-5 stands at 0.462 on HealthBench Professional (23 of 29 indexed rows), 0.347 on HealthBench Hard (5 of 20 indexed rows), 57.7 on HealthBench (11 of 26 indexed rows), 68.6% on First, Do NOHARM (v2) (10 of 17 indexed rows), 49.63% on MedCode (Vals AI) [GPT 5] (18 of 102 indexed rows), 83.65% on MedScribe (Vals AI) [GPT 5] (40 of 104 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.