GPT-4o: healthcare benchmark results
OpenAI · 2 boards · snapshot reviewed September 28, 2026
The index currently holds 2 results for GPT-4o: 0.31 on EHR-Complex (15 of 18 indexed rows), 44.6% on WHBench (16 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| EHR-Complex Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.31 | 15 of 18 | 2026-06 |
| WHBench via paper, arxiv.org | 44.6% | 16 of 22 | 2026-03 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page.
Which healthcare benchmarks is GPT-4o scored on?
As of September 28, 2026, GPT-4o has indexed results on 2 tracked benchmarks: EHR-Complex, WHBench.
Which results are indexed for GPT-4o?
GPT-4o stands at 0.31 on EHR-Complex (15 of 18 indexed rows), 44.6% on WHBench (16 of 22 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.