Health Evals

OpenAI logoGPT-5.4: healthcare benchmark results

OpenAI · 9 boards · snapshot reviewed September 28, 2026

The index currently holds 11 results for GPT-5.4: 0.481 on HealthBench Professional (22 of 29 indexed rows), 0.291 on HealthBench Hard (15 of 20 indexed rows), 54.0 on HealthBench (21 of 26 indexed rows), 0.538 on MedHELM [GPT-5.4 (2026-03-05)] (5 of 10 indexed rows), 77.1% on MedXpertQA (MM) (6 of 22 indexed rows), 27.7 ± 1.5 on PhysicianBench (12 of 21 indexed rows), 0.65 on EHR-Complex [GPT-5.4 (high reasoning)] (1 of 18 indexed rows), 0.58 on EHR-Complex [GPT-5.4 (low reasoning)] (6 of 18 indexed rows), 66.8% on WHBench (3 of 22 indexed rows), 26.7% on HealthAdminBench [GPT-5.4 (computer-use agent)] (2 of 7 indexed rows), 5.9% on HealthAdminBench [GPT-5.4 (standardized harness)] (7 of 7 indexed rows). It has the highest indexed score on EHR-Complex. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
0.48122 of 292026-06
HealthBench Hard
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
0.29115 of 202026-06
HealthBench
OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6.
54.021 of 262026-06
MedHELM
GPT-5.4 (2026-03-05)
0.5385 of 10
source board: 11
2026-05
MedXpertQA (MM)
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
77.1%6 of 222026-04
PhysicianBench27.7 ± 1.512 of 212026-05
EHR-Complex
GPT-5.4 (high reasoning)
average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
0.651 of 182026-06
EHR-Complex
GPT-5.4 (low reasoning)
validation configuration
0.586 of 182026-06
WHBench
95% CI 64.5-69.2
66.8%3 of 222026-03
HealthAdminBench
GPT-5.4 (computer-use agent)
screenshot-only, task description + portal guidance; subtask rate 82.8%
26.7%2 of 72026-04
HealthAdminBench
GPT-5.4 (standardized harness)
screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
5.9%7 of 72026-04

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.4 (2026-03-05)" and "GPT-5.4 (high reasoning)" and "GPT-5.4 (low reasoning)" and "GPT-5.4 (computer-use agent)" and "GPT-5.4 (standardized harness)".

Which healthcare benchmarks is GPT-5.4 scored on?

As of September 28, 2026, GPT-5.4 has indexed results on 9 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedHELM, MedXpertQA (MM), PhysicianBench, EHR-Complex, WHBench, HealthAdminBench.

Which results are indexed for GPT-5.4?

GPT-5.4 stands at 0.481 on HealthBench Professional (22 of 29 indexed rows), 0.291 on HealthBench Hard (15 of 20 indexed rows), 54.0 on HealthBench (21 of 26 indexed rows), 0.538 on MedHELM [GPT-5.4 (2026-03-05)] (5 of 10 indexed rows), 77.1% on MedXpertQA (MM) (6 of 22 indexed rows), 27.7 ± 1.5 on PhysicianBench (12 of 21 indexed rows), 0.65 on EHR-Complex [GPT-5.4 (high reasoning)] (1 of 18 indexed rows), 0.58 on EHR-Complex [GPT-5.4 (low reasoning)] (6 of 18 indexed rows), 66.8% on WHBench (3 of 22 indexed rows), 26.7% on HealthAdminBench [GPT-5.4 (computer-use agent)] (2 of 7 indexed rows), 5.9% on HealthAdminBench [GPT-5.4 (standardized harness)] (7 of 7 indexed rows). It has the highest indexed score on EHR-Complex.

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.