GPT-6 Luna: healthcare benchmark results
OpenAI · 6 boards · snapshot reviewed September 28, 2026
The index currently holds 6 results for GPT-6 Luna: 0.608 on HealthBench Professional (8 of 29 indexed rows), 0.314 on HealthBench Hard (12 of 20 indexed rows), 54.5 on HealthBench (19 of 26 indexed rows), 44.69% on MedCode (Vals AI) (37 of 102 indexed rows), 83.71% on MedScribe (Vals AI) (39 of 104 indexed rows), 37 on Artificial Analysis Healthcare & Medical Index [GPT-6 Luna (max)] (17 of 25 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.608 | 8 of 29 | 2026-09 |
| HealthBench Hard OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 0.314 | 12 of 20 | 2026-09 |
| HealthBench OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 50%, mean answer 1,255 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | 54.5 | 19 of 26 | 2026-09 |
| MedCode (Vals AI) model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not published | 44.69% | 37 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not published | 83.71% | 39 of 104 | 2026-09-26 |
| Artificial Analysis Healthcare & Medical Index GPT-6 Luna (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 37 | 17 of 25 source board: 77 | 2026-09 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-6 Luna (max)".
Which healthcare benchmarks is GPT-6 Luna scored on?
As of September 28, 2026, GPT-6 Luna has indexed results on 6 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index.
Which results are indexed for GPT-6 Luna?
GPT-6 Luna stands at 0.608 on HealthBench Professional (8 of 29 indexed rows), 0.314 on HealthBench Hard (12 of 20 indexed rows), 54.5 on HealthBench (19 of 26 indexed rows), 44.69% on MedCode (Vals AI) (37 of 102 indexed rows), 83.71% on MedScribe (Vals AI) (39 of 104 indexed rows), 37 on Artificial Analysis Healthcare & Medical Index [GPT-6 Luna (max)] (17 of 25 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.