GPT-5.4: healthcare benchmark results
OpenAI · 9 boards · snapshot reviewed September 28, 2026
The index currently holds 11 results for GPT-5.4: 0.481 on HealthBench Professional (22 of 29 indexed rows), 0.291 on HealthBench Hard (15 of 20 indexed rows), 54.0 on HealthBench (21 of 26 indexed rows), 0.538 on MedHELM [GPT-5.4 (2026-03-05)] (5 of 10 indexed rows), 77.1% on MedXpertQA (MM) (6 of 22 indexed rows), 27.7 ± 1.5 on PhysicianBench (12 of 21 indexed rows), 0.65 on EHR-Complex [GPT-5.4 (high reasoning)] (1 of 18 indexed rows), 0.58 on EHR-Complex [GPT-5.4 (low reasoning)] (6 of 18 indexed rows), 66.8% on WHBench (3 of 22 indexed rows), 26.7% on HealthAdminBench [GPT-5.4 (computer-use agent)] (2 of 7 indexed rows), 5.9% on HealthAdminBench [GPT-5.4 (standardized harness)] (7 of 7 indexed rows). It has the highest indexed score on EHR-Complex. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars) | 0.481 | 22 of 29 | 2026-06 |
| HealthBench Hard length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars) | 0.291 | 15 of 20 | 2026-06 |
| HealthBench OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6. | 54.0 | 21 of 26 | 2026-06 |
| MedHELM GPT-5.4 (2026-03-05) | 0.538 | 5 of 10 source board: 11 | 2026-05 |
| MedXpertQA (MM) Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 77.1% | 6 of 22 | 2026-04 |
| PhysicianBench via paper, arxiv.org | 27.7 ± 1.5 | 12 of 21 | 2026-05 |
| EHR-Complex GPT-5.4 (high reasoning) average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table via paper, arxiv.org | 0.65 | 1 of 18 | 2026-06 |
| EHR-Complex GPT-5.4 (low reasoning) validation configuration via paper, arxiv.org | 0.58 | 6 of 18 | 2026-06 |
| WHBench 95% CI 64.5-69.2 via paper, arxiv.org | 66.8% | 3 of 22 | 2026-03 |
| HealthAdminBench GPT-5.4 (computer-use agent) screenshot-only, task description + portal guidance; subtask rate 82.8% via paper, arxiv.org | 26.7% | 2 of 7 | 2026-04 |
| HealthAdminBench GPT-5.4 (standardized harness) screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA via paper, arxiv.org | 5.9% | 7 of 7 | 2026-04 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.4 (2026-03-05)" and "GPT-5.4 (high reasoning)" and "GPT-5.4 (low reasoning)" and "GPT-5.4 (computer-use agent)" and "GPT-5.4 (standardized harness)".
Which healthcare benchmarks is GPT-5.4 scored on?
As of September 28, 2026, GPT-5.4 has indexed results on 9 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedHELM, MedXpertQA (MM), PhysicianBench, EHR-Complex, WHBench, HealthAdminBench.
Which results are indexed for GPT-5.4?
GPT-5.4 stands at 0.481 on HealthBench Professional (22 of 29 indexed rows), 0.291 on HealthBench Hard (15 of 20 indexed rows), 54.0 on HealthBench (21 of 26 indexed rows), 0.538 on MedHELM [GPT-5.4 (2026-03-05)] (5 of 10 indexed rows), 77.1% on MedXpertQA (MM) (6 of 22 indexed rows), 27.7 ± 1.5 on PhysicianBench (12 of 21 indexed rows), 0.65 on EHR-Complex [GPT-5.4 (high reasoning)] (1 of 18 indexed rows), 0.58 on EHR-Complex [GPT-5.4 (low reasoning)] (6 of 18 indexed rows), 66.8% on WHBench (3 of 22 indexed rows), 26.7% on HealthAdminBench [GPT-5.4 (computer-use agent)] (2 of 7 indexed rows), 5.9% on HealthAdminBench [GPT-5.4 (standardized harness)] (7 of 7 indexed rows). It has the highest indexed score on EHR-Complex.
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.