GPT-5.2: healthcare benchmark results
OpenAI · 6 boards · snapshot reviewed September 28, 2026
The index currently holds 8 results for GPT-5.2: 0.459 on HealthBench Professional (24 of 29 indexed rows), 0.420 on HealthBench Hard [GPT-5.2-High (Baichuan run)] (3 of 20 indexed rows), 0.343 on HealthBench Hard (6 of 20 indexed rows), 63.3 on HealthBench [GPT-5.2-High] (3 of 26 indexed rows), 56.8 on HealthBench (15 of 26 indexed rows), 49.75% on MedCode (Vals AI) [GPT 5.2] (17 of 102 indexed rows), 84.39% on MedScribe (Vals AI) [GPT 5.2] (33 of 104 indexed rows), 73.3 on MedXpertQA (MM) (8 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars) | 0.459 | 24 of 29 | 2026-06 |
| HealthBench Hard GPT-5.2-High (Baichuan run) Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol. | 0.420 | 3 of 20 | 2026-02 |
| HealthBench Hard length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars) | 0.343 | 6 of 20 | 2026-06 |
| HealthBench GPT-5.2-High raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card. | 63.3 | 3 of 26 | 2026-02 |
| HealthBench OpenAI length-adjusted score, maximum reasoning effort; raw 60.7%, mean answer 2,645 characters; GPT-5.6 card Table 6. | 56.8 | 15 of 26 | 2026-06 |
| MedCode (Vals AI) GPT 5.2 model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not published | 49.75% | 17 of 102 | 2026-09-26 |
| MedScribe (Vals AI) GPT 5.2 model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not published | 84.39% | 33 of 104 | 2026-09-26 |
| MedXpertQA (MM) Qwen-run comparison in the Qwen3.5-397B-A17B model card | 73.3 | 8 of 22 | not reported |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.2-High (Baichuan run)" and "GPT-5.2-High" and "GPT 5.2".
Which healthcare benchmarks is GPT-5.2 scored on?
As of September 28, 2026, GPT-5.2 has indexed results on 6 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM).
Which results are indexed for GPT-5.2?
GPT-5.2 stands at 0.459 on HealthBench Professional (24 of 29 indexed rows), 0.420 on HealthBench Hard [GPT-5.2-High (Baichuan run)] (3 of 20 indexed rows), 0.343 on HealthBench Hard (6 of 20 indexed rows), 63.3 on HealthBench [GPT-5.2-High] (3 of 26 indexed rows), 56.8 on HealthBench (15 of 26 indexed rows), 49.75% on MedCode (Vals AI) [GPT 5.2] (17 of 102 indexed rows), 84.39% on MedScribe (Vals AI) [GPT 5.2] (33 of 104 indexed rows), 73.3 on MedXpertQA (MM) (8 of 22 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.