GPT-5.6 Luna: healthcare benchmark results
OpenAI · 6 boards · snapshot reviewed September 28, 2026
The index currently holds 9 results for GPT-5.6 Luna: 0.557 on HealthBench Professional (16 of 29 indexed rows), 0.441 on HealthBench Professional [GPT-5.6 Luna (August)] (26 of 29 indexed rows), 0.320 on HealthBench Hard (9 of 20 indexed rows), 0.287 on HealthBench Hard [GPT-5.6 Luna (August)] (16 of 20 indexed rows), 55.8 on HealthBench (17 of 26 indexed rows), 53.3 on HealthBench [GPT-5.6 Luna (August)] (22 of 26 indexed rows), 42.39% on MedCode (Vals AI) (47 of 102 indexed rows), 84.39% on MedScribe (Vals AI) (32 of 104 indexed rows), 36 on Artificial Analysis Healthcare & Medical Index [GPT-5.6 Luna (max)] (18 of 25 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493. | 0.557 | 16 of 29 | 2026-06 |
| HealthBench Professional GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars) | 0.441 | 26 of 29 | 2026-08 |
| HealthBench Hard length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491. | 0.320 | 9 of 20 | 2026-06 |
| HealthBench Hard GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars) | 0.287 | 16 of 20 | 2026-08 |
| HealthBench length-adjusted (55.4 unadjusted), max reasoning effort | 55.8 | 17 of 26 | 2026-06 |
| HealthBench GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars) | 53.3 | 22 of 26 | 2026-08 |
| MedCode (Vals AI) model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published | 42.39% | 47 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published | 84.39% | 32 of 104 | 2026-09-26 |
| Artificial Analysis Healthcare & Medical Index GPT-5.6 Luna (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 36 | 18 of 25 source board: 77 | 2026-09 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.6 Luna (August)" and "GPT-5.6 Luna (max)".
Which healthcare benchmarks is GPT-5.6 Luna scored on?
As of September 28, 2026, GPT-5.6 Luna has indexed results on 6 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index.
Which results are indexed for GPT-5.6 Luna?
GPT-5.6 Luna stands at 0.557 on HealthBench Professional (16 of 29 indexed rows), 0.441 on HealthBench Professional [GPT-5.6 Luna (August)] (26 of 29 indexed rows), 0.320 on HealthBench Hard (9 of 20 indexed rows), 0.287 on HealthBench Hard [GPT-5.6 Luna (August)] (16 of 20 indexed rows), 55.8 on HealthBench (17 of 26 indexed rows), 53.3 on HealthBench [GPT-5.6 Luna (August)] (22 of 26 indexed rows), 42.39% on MedCode (Vals AI) (47 of 102 indexed rows), 84.39% on MedScribe (Vals AI) (32 of 104 indexed rows), 36 on Artificial Analysis Healthcare & Medical Index [GPT-5.6 Luna (max)] (18 of 25 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.