Codex (GPT 5.4 Mini): healthcare benchmark results
OpenAI · 2 boards · snapshot reviewed September 28, 2026
The index currently holds 2 results for Codex (GPT 5.4 Mini): 16% on HealthAgentBench (12 of 12 indexed rows), 8.4% on CHI-Bench [codex + gpt-5.4-mini] (34 of 44 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthAgentBench $0.6/task | 16% | 12 of 12 | 2026-07 |
| CHI-Bench codex + gpt-5.4-mini All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not published | 8.4% | 34 of 44 source board: 45 | 2026-05-01 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "codex + gpt-5.4-mini".
Which healthcare benchmarks is Codex (GPT 5.4 Mini) scored on?
As of September 28, 2026, Codex (GPT 5.4 Mini) has indexed results on 2 tracked benchmarks: HealthAgentBench, CHI-Bench.
Which results are indexed for Codex (GPT 5.4 Mini)?
Codex (GPT 5.4 Mini) stands at 16% on HealthAgentBench (12 of 12 indexed rows), 8.4% on CHI-Bench [codex + gpt-5.4-mini] (34 of 44 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.