GPT-5.6 Sol: healthcare benchmark results
OpenAI · 10 boards · snapshot reviewed September 28, 2026
The index currently holds 14 results for GPT-5.6 Sol: 0.605 on HealthBench Professional (9 of 29 indexed rows), 0.540 on HealthBench Professional [GPT-5.6 Sol (August)] (18 of 29 indexed rows), 0.331 on HealthBench Hard (7 of 20 indexed rows), 0.314 on HealthBench Hard [GPT-5.6 Sol (August)] (11 of 20 indexed rows), 57.0 on HealthBench (13 of 26 indexed rows), 55.0 on HealthBench [GPT-5.6 Sol (August)] (18 of 26 indexed rows), 66.6 on Health Optimization Bench [GPT-5.6 Sol (max)] (4 of 16 indexed rows), 64.6 on Health Optimization Bench [GPT-5.6 Sol (high)] (5 of 16 indexed rows), 60.2% on MAST (Medical AI Superintelligence Test) (1 of 8 indexed rows), 70.1% on First, Do NOHARM (v2) (8 of 17 indexed rows), 43.97% on MedCode (Vals AI) (39 of 102 indexed rows), 85.23% on MedScribe (Vals AI) (27 of 104 indexed rows), 81.5 on MedXpertQA (MM) (1 of 22 indexed rows), 45 on Artificial Analysis Healthcare & Medical Index [GPT-5.6 Sol (max)] (8 of 25 indexed rows). It has the highest indexed score on MAST (Medical AI Superintelligence Test) and MedXpertQA (MM). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492. | 0.605 | 9 of 29 | 2026-06 |
| HealthBench Professional GPT-5.6 Sol (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars) | 0.540 | 18 of 29 | 2026-08 |
| HealthBench Hard length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490. | 0.331 | 7 of 20 | 2026-06 |
| HealthBench Hard GPT-5.6 Sol (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars) | 0.314 | 11 of 20 | 2026-08 |
| HealthBench length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09 | 57.0 | 13 of 26 | 2026-06 |
| HealthBench GPT-5.6 Sol (August) ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06 | 55.0 | 18 of 26 | 2026-08 |
| Health Optimization Bench GPT-5.6 Sol (max) Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort. | 66.6 | 4 of 16 | 2026-09 |
| Health Optimization Bench GPT-5.6 Sol (high) Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort. | 64.6 | 5 of 16 | 2026-09 |
| MAST (Medical AI Superintelligence Test) MAST in preview; 'exact scores may change' | 60.2% | 1 of 8 source board: 11 | 2026-08 |
| First, Do NOHARM (v2) | 70.1% | 8 of 17 source board: 19 | 2026-08 |
| MedCode (Vals AI) model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not published | 43.97% | 39 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not published | 85.23% | 27 of 104 | 2026-09-26 |
| MedXpertQA (MM) Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 81.5 | 1 of 22 | not reported |
| Artificial Analysis Healthcare & Medical Index GPT-5.6 Sol (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 45 | 8 of 25 source board: 77 | 2026-09 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.6 Sol (August)" and "GPT-5.6 Sol (max)" and "GPT-5.6 Sol (high)".
Which healthcare benchmarks is GPT-5.6 Sol scored on?
As of September 28, 2026, GPT-5.6 Sol has indexed results on 10 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, Health Optimization Bench, MAST (Medical AI Superintelligence Test), First, Do NOHARM (v2), MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM), Artificial Analysis Healthcare & Medical Index.
Which results are indexed for GPT-5.6 Sol?
GPT-5.6 Sol stands at 0.605 on HealthBench Professional (9 of 29 indexed rows), 0.540 on HealthBench Professional [GPT-5.6 Sol (August)] (18 of 29 indexed rows), 0.331 on HealthBench Hard (7 of 20 indexed rows), 0.314 on HealthBench Hard [GPT-5.6 Sol (August)] (11 of 20 indexed rows), 57.0 on HealthBench (13 of 26 indexed rows), 55.0 on HealthBench [GPT-5.6 Sol (August)] (18 of 26 indexed rows), 66.6 on Health Optimization Bench [GPT-5.6 Sol (max)] (4 of 16 indexed rows), 64.6 on Health Optimization Bench [GPT-5.6 Sol (high)] (5 of 16 indexed rows), 60.2% on MAST (Medical AI Superintelligence Test) (1 of 8 indexed rows), 70.1% on First, Do NOHARM (v2) (8 of 17 indexed rows), 43.97% on MedCode (Vals AI) (39 of 102 indexed rows), 85.23% on MedScribe (Vals AI) (27 of 104 indexed rows), 81.5 on MedXpertQA (MM) (1 of 22 indexed rows), 45 on Artificial Analysis Healthcare & Medical Index [GPT-5.6 Sol (max)] (8 of 25 indexed rows). It has the highest indexed score on MAST (Medical AI Superintelligence Test) and MedXpertQA (MM).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.