Codex (GPT-5.6-sol): healthcare benchmark results
OpenAI · 2 boards · snapshot reviewed September 28, 2026
The index currently holds 2 results for Codex (GPT-5.6-sol): 45% on HealthAgentBench (2 of 12 indexed rows), 25.3% on CHI-Bench [codex + gpt-5.6-sol] (7 of 44 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthAgentBench $5.2/task | 45% | 2 of 12 | 2026-07 |
| CHI-Bench codex + gpt-5.6-sol All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not published | 25.3% | 7 of 44 source board: 45 | 2026-07-24 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "codex + gpt-5.6-sol".
Which healthcare benchmarks is Codex (GPT-5.6-sol) scored on?
As of September 28, 2026, Codex (GPT-5.6-sol) has indexed results on 2 tracked benchmarks: HealthAgentBench, CHI-Bench.
Which results are indexed for Codex (GPT-5.6-sol)?
Codex (GPT-5.6-sol) stands at 45% on HealthAgentBench (2 of 12 indexed rows), 25.3% on CHI-Bench [codex + gpt-5.6-sol] (7 of 44 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.