Health Evals

OpenAI logoGPT-5.6 Luna: healthcare benchmark results

OpenAI · 6 boards · snapshot reviewed September 28, 2026

The index currently holds 9 results for GPT-5.6 Luna: 0.557 on HealthBench Professional (16 of 29 indexed rows), 0.441 on HealthBench Professional [GPT-5.6 Luna (August)] (26 of 29 indexed rows), 0.320 on HealthBench Hard (9 of 20 indexed rows), 0.287 on HealthBench Hard [GPT-5.6 Luna (August)] (16 of 20 indexed rows), 55.8 on HealthBench (17 of 26 indexed rows), 53.3 on HealthBench [GPT-5.6 Luna (August)] (22 of 26 indexed rows), 42.39% on MedCode (Vals AI) (47 of 102 indexed rows), 84.39% on MedScribe (Vals AI) (32 of 104 indexed rows), 36 on Artificial Analysis Healthcare & Medical Index [GPT-5.6 Luna (max)] (18 of 25 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Professional
length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
0.55716 of 292026-06
HealthBench Professional
GPT-5.6 Luna (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
0.44126 of 292026-08
HealthBench Hard
length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
0.3209 of 202026-06
HealthBench Hard
GPT-5.6 Luna (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)
0.28716 of 202026-08
HealthBench
length-adjusted (55.4 unadjusted), max reasoning effort
55.817 of 262026-06
HealthBench
GPT-5.6 Luna (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars)
53.322 of 262026-08
MedCode (Vals AI)
model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published
42.39%47 of 1022026-09-26
MedScribe (Vals AI)
model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published
84.39%32 of 1042026-09-26
Artificial Analysis Healthcare & Medical Index
GPT-5.6 Luna (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
3618 of 25
source board: 77
2026-09

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.6 Luna (August)" and "GPT-5.6 Luna (max)".

Which healthcare benchmarks is GPT-5.6 Luna scored on?

As of September 28, 2026, GPT-5.6 Luna has indexed results on 6 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index.

Which results are indexed for GPT-5.6 Luna?

GPT-5.6 Luna stands at 0.557 on HealthBench Professional (16 of 29 indexed rows), 0.441 on HealthBench Professional [GPT-5.6 Luna (August)] (26 of 29 indexed rows), 0.320 on HealthBench Hard (9 of 20 indexed rows), 0.287 on HealthBench Hard [GPT-5.6 Luna (August)] (16 of 20 indexed rows), 55.8 on HealthBench (17 of 26 indexed rows), 53.3 on HealthBench [GPT-5.6 Luna (August)] (22 of 26 indexed rows), 42.39% on MedCode (Vals AI) (47 of 102 indexed rows), 84.39% on MedScribe (Vals AI) (32 of 104 indexed rows), 36 on Artificial Analysis Healthcare & Medical Index [GPT-5.6 Luna (max)] (18 of 25 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.