Codex (GPT 5.5): healthcare benchmark results
OpenAI · 2 boards · snapshot reviewed September 28, 2026
The index currently holds 2 results for Codex (GPT 5.5): 42% on HealthAgentBench (3 of 12 indexed rows), 20.9% on CHI-Bench [codex + gpt-5.5] (12 of 44 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthAgentBench $2.8/task | 42% | 3 of 12 | 2026-07 |
| CHI-Bench codex + gpt-5.5 All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not published | 20.9% | 12 of 44 source board: 45 | 2026-05-01 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "codex + gpt-5.5".
Which healthcare benchmarks is Codex (GPT 5.5) scored on?
As of September 28, 2026, Codex (GPT 5.5) has indexed results on 2 tracked benchmarks: HealthAgentBench, CHI-Bench.
Which results are indexed for Codex (GPT 5.5)?
Codex (GPT 5.5) stands at 42% on HealthAgentBench (3 of 12 indexed rows), 20.9% on CHI-Bench [codex + gpt-5.5] (12 of 44 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.