OpenAI: healthcare benchmark results
33 models · 123 results · snapshot reviewed September 28, 2026
Everything the index holds for OpenAI, gathered in one place: GPT-5.6 Sol, GPT-5.4, GPT-5.5, GPT-6 Astra (Anthropic run), GPT-6 Sol, GPT-6 Luna, GPT-5.6 Luna, GPT-5, GPT-5.2, GPT-5.6 Terra, GPT-5.1, GPT-5.5 Instant, GPT OSS 120B, GPT-5.3 Chat, GPT OSS 20B, Codex (GPT-5.6-sol), Codex (GPT 5.5), Codex (GPT 5.4), Codex (GPT 5.4 Mini), o3, GPT 5 Mini, GPT 5.4 (xhigh), GPT 5.4 Nano, o4 Mini, GPT 5 Nano, GPT-4.1, GPT-4o, GPT-5.4 mini, Codex (GPT 5.3), codex + gpt-5.6-terra, codex + gpt-5.6-luna, GPT-4.1 mini, OpenAI o3. The lab has the highest indexed score on MAST (Medical AI Superintelligence Test), MedXpertQA (MM), EHR-Complex, HealthBench Professional. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| GPT-5.6 Sol length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492. | HealthBench Professional | 0.605 | 9 of 29 |
| GPT-5.6 Sol (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars) | HealthBench Professional | 0.540 | 18 of 29 |
| GPT-5.6 Sol length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490. | HealthBench Hard | 0.331 | 7 of 20 |
| GPT-5.6 Sol (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars) | HealthBench Hard | 0.314 | 11 of 20 |
| GPT-5.6 Sol length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09 | HealthBench | 57.0 | 13 of 26 |
| GPT-5.6 Sol (August) ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06 | HealthBench | 55.0 | 18 of 26 |
| GPT-5.6 Sol (max) Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort. | Health Optimization Bench | 66.6 | 4 of 16 |
| GPT-5.6 Sol (high) Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort. | Health Optimization Bench | 64.6 | 5 of 16 |
| GPT-5.6 Sol MAST in preview; 'exact scores may change' | MAST (Medical AI Superintelligence Test) | 60.2% | 1 of 8 source board: 11 |
| GPT-5.6 Sol | First, Do NOHARM (v2) | 70.1% | 8 of 17 source board: 19 |
| GPT-5.6 Sol model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 43.97% | 39 of 102 |
| GPT-5.6 Sol model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 85.23% | 27 of 104 |
| GPT-5.6 Sol Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | MedXpertQA (MM) | 81.5 | 1 of 22 |
| GPT-5.6 Sol (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 45 | 8 of 25 source board: 77 |
| GPT-5.4 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars) | HealthBench Professional | 0.481 | 22 of 29 |
| GPT-5.4 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars) | HealthBench Hard | 0.291 | 15 of 20 |
| GPT-5.4 OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6. | HealthBench | 54.0 | 21 of 26 |
| GPT-5.4 (2026-03-05) | MedHELM | 0.538 | 5 of 10 source board: 11 |
| GPT-5.4 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | MedXpertQA (MM) | 77.1% | 6 of 22 |
| GPT-5.4 | PhysicianBench | 27.7 ± 1.5 | 12 of 21 |
| GPT-5.4 (high reasoning) average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table | EHR-Complex | 0.65 | 1 of 18 |
| GPT-5.4 (low reasoning) validation configuration | EHR-Complex | 0.58 | 6 of 18 |
| GPT-5.4 95% CI 64.5-69.2 | WHBench | 66.8% | 3 of 22 |
| GPT-5.4 (computer-use agent) screenshot-only, task description + portal guidance; subtask rate 82.8% | HealthAdminBench | 26.7% | 2 of 7 |
| GPT-5.4 (standardized harness) screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA | HealthAdminBench | 5.9% | 7 of 7 |
| GPT-5.5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars) | HealthBench Professional | 0.518 | 20 of 29 |
| GPT-5.5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars) | HealthBench Hard | 0.315 | 10 of 20 |
| GPT-5.5 length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system card | HealthBench | 56.5 | 16 of 26 |
| GPT-5.5 | First, Do NOHARM (v2) | 70.0% | 9 of 17 source board: 19 |
| GPT 5.5 model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.188 pp; $0.160759/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 49.10% | 23 of 102 |
| GPT 5.5 model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.142988/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 86.87% | 17 of 104 |
| GPT-5.5 pass@1; Pass^3 28.0 | PhysicianBench | 46.3 ± 1.2 | 7 of 21 |
| GPT-6 Astra (Anthropic run) Anthropic reproduction of GPT-6 Astra through the public API; max effort; no system prompt; Claude Opus 4.8 grader; length-adjusted 70.3%, raw 74.0%. Different grader/protocol from OpenAI’s own 64.7% report. | HealthBench Professional | 0.703 | 1 of 29 |
| GPT-6 Astra OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 68.2%, mean answer 3,185 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench Professional | 0.647 | 4 of 29 |
| GPT-6 Astra OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 34.2%, mean answer 1,697 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench Hard | 0.366 | 4 of 20 |
| GPT-6 Astra OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 56.9%, mean answer 1,760 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench | 58.3 | 9 of 26 |
| GPT-6 Astra model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 2.131 pp; $0.451358/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 48.49% | 25 of 102 |
| GPT-6 Astra model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 1.938 pp; $0.581991/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 87.91% | 14 of 104 |
| GPT-6 Astra (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 52 | 4 of 25 source board: 77 |
| GPT-6 Sol OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 59.5%, mean answer 1,573 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench Professional | 0.608 | 7 of 29 |
| GPT-6 Sol OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 22.1%, mean answer 974 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench Hard | 0.301 | 13 of 20 |
| GPT-6 Sol OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 47.1%, mean answer 977 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench | 53.2 | 23 of 26 |
| GPT-6 Sol model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 2.119 pp; $0.085827/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 47.07% | 32 of 102 |
| GPT-6 Sol model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 1.942 pp; $0.083138/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 82.03% | 48 of 104 |
| GPT-6 Sol (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 43 | 11 of 25 source board: 77 |
| GPT-6 Luna OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench Professional | 0.608 | 8 of 29 |
| GPT-6 Luna OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench Hard | 0.314 | 12 of 20 |
| GPT-6 Luna OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 50%, mean answer 1,255 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration. | HealthBench | 54.5 | 19 of 26 |
| GPT-6 Luna model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 44.69% | 37 of 102 |
| GPT-6 Luna model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 83.71% | 39 of 104 |
| GPT-6 Luna (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 37 | 17 of 25 source board: 77 |
| GPT-5.6 Luna length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493. | HealthBench Professional | 0.557 | 16 of 29 |
| GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars) | HealthBench Professional | 0.441 | 26 of 29 |
| GPT-5.6 Luna length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491. | HealthBench Hard | 0.320 | 9 of 20 |
| GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars) | HealthBench Hard | 0.287 | 16 of 20 |
| GPT-5.6 Luna length-adjusted (55.4 unadjusted), max reasoning effort | HealthBench | 55.8 | 17 of 26 |
| GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars) | HealthBench | 53.3 | 22 of 26 |
| GPT-5.6 Luna model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 42.39% | 47 of 102 |
| GPT-5.6 Luna model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 84.39% | 32 of 104 |
| GPT-5.6 Luna (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 36 | 18 of 25 source board: 77 |
| GPT-5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars) | HealthBench Professional | 0.462 | 23 of 29 |
| GPT-5 length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking. | HealthBench Hard | 0.347 | 5 of 20 |
| GPT-5 OpenAI length-adjusted score, maximum reasoning effort; raw 63.1%, mean answer 2,904 characters; GPT-5.6 card Table 6. | HealthBench | 57.7 | 11 of 26 |
| GPT-5 from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking | First, Do NOHARM (v2) | 68.6% | 10 of 17 source board: 19 |
| GPT 5 model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 49.63% | 18 of 102 |
| GPT 5 model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 83.65% | 40 of 104 |
| GPT-5.2 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars) | HealthBench Professional | 0.459 | 24 of 29 |
| GPT-5.2-High (Baichuan run) Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol. | HealthBench Hard | 0.420 | 3 of 20 |
| GPT-5.2 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars) | HealthBench Hard | 0.343 | 6 of 20 |
| GPT-5.2-High raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card. | HealthBench | 63.3 | 3 of 26 |
| GPT-5.2 OpenAI length-adjusted score, maximum reasoning effort; raw 60.7%, mean answer 2,645 characters; GPT-5.6 card Table 6. | HealthBench | 56.8 | 15 of 26 |
| GPT 5.2 model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 49.75% | 17 of 102 |
| GPT 5.2 model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 84.39% | 33 of 104 |
| GPT-5.2 Qwen-run comparison in the Qwen3.5-397B-A17B model card | MedXpertQA (MM) | 73.3 | 8 of 22 |
| GPT-5.6 Terra length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA. | HealthBench Professional | 0.577 | 13 of 29 |
| GPT-5.6 Terra length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA. | HealthBench Hard | 0.327 | 8 of 20 |
| GPT-5.6 Terra length-adjusted (58.7 unadjusted), max reasoning effort | HealthBench | 57.0 | 14 of 26 |
| GPT-5.6 Terra model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.173 pp; $0.046558/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 43.41% | 41 of 102 |
| GPT-5.6 Terra model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.948 pp; $0.062588/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 82.87% | 47 of 104 |
| GPT-5.1 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars) | HealthBench Professional | 0.396 | 27 of 29 |
| GPT-5.1 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars) | HealthBench Hard | 0.254 | 18 of 20 |
| GPT-5.1 OpenAI length-adjusted score, maximum reasoning effort; raw 64.2%, mean answer 4,222 characters; GPT-5.6 card Table 6. | HealthBench | 50.9 | 25 of 26 |
| GPT 5.1 model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 2.151 pp; $0.014371/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 52.73% | 12 of 102 |
| GPT 5.1 model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 1.942 pp; $0.096508/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 88.09% | 12 of 104 |
| GPT-5.5 Instant length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT. | HealthBench Professional | 0.384 | 28 of 29 |
| GPT-5.5 Instant length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT. | HealthBench Hard | 0.229 | 19 of 20 |
| GPT-5.5 Instant length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11 | HealthBench | 51.4 | 24 of 26 |
| GPT OSS 120B raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board. | HealthBench Hard | 0.300 | 14 of 20 |
| GPT OSS 120B reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9) | HealthBench | 57.6 | 12 of 26 |
| GPT-5.3 Chat raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model. | HealthBench Hard | 0.259 | 17 of 20 |
| GPT-5.3 Chat raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted) | HealthBench | 54.1% | 20 of 26 |
| GPT OSS 20B raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board. | HealthBench Hard | 0.108 | 20 of 20 |
| GPT OSS 20B reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8) | HealthBench | 42.5 | 26 of 26 |
| Codex (GPT-5.6-sol) $5.2/task | HealthAgentBench | 45% | 2 of 12 |
| codex + gpt-5.6-sol All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not published | CHI-Bench | 25.3% | 7 of 44 source board: 45 |
| Codex (GPT 5.5) $2.8/task | HealthAgentBench | 42% | 3 of 12 |
| codex + gpt-5.5 All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not published | CHI-Bench | 20.9% | 12 of 44 source board: 45 |
| Codex (GPT 5.4) $1.3/task | HealthAgentBench | 28% | 7 of 12 |
| codex + gpt-5.4 All Domains pass@1; PA 24.0%, UM 17.3%, CM 6.7%; submitted 2026-05-01; run date not published | CHI-Bench | 16.0% | 20 of 44 source board: 45 |
| Codex (GPT 5.4 Mini) $0.6/task | HealthAgentBench | 16% | 12 of 12 |
| codex + gpt-5.4-mini All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not published | CHI-Bench | 8.4% | 34 of 44 source board: 45 |
| o3 model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.161 pp; $0.029818/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 47.29% | 30 of 102 |
| o3 model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.871 pp; $0.040334/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 76.65% | 69 of 104 |
| GPT 5 Mini model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.045 pp; $0.005560/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 43.05% | 44 of 102 |
| GPT 5 Mini model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.924 pp; $0.033478/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 80.58% | 52 of 104 |
| GPT 5.4 (xhigh) model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.148 pp; $0.212108/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 41.29% | 51 of 102 |
| GPT 5.4 (xhigh) model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 3.316 pp; $0.639282/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 77.55% | 64 of 104 |
| GPT 5.4 Nano model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 2.256 pp; $0.000844/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 41.03% | 55 of 102 |
| GPT 5.4 Nano model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.001800/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 77.09% | 67 of 104 |
| o4 Mini model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.021 pp; $0.017605/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 33.79% | 79 of 102 |
| o4 Mini model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.957 pp; $0.040605/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 69.14% | 92 of 104 |
| GPT 5 Nano model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.948 pp; $0.001729/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 30.44% | 91 of 102 |
| GPT 5 Nano model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.006961/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 72.86% | 80 of 104 |
| GPT-4.1 Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.47 | 11 of 18 |
| GPT-4.1 | WHBench | 51.8% | 11 of 22 |
| GPT-4o Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.31 | 15 of 18 |
| GPT-4o | WHBench | 44.6% | 16 of 22 |
| GPT-5.4 mini | MedHELM | 0.552 | 4 of 10 source board: 11 |
| Codex (GPT 5.3) $1.0/task; harness and model evaluated jointly | HealthAgentBench | 22% | 9 of 12 |
| codex + gpt-5.6-terra All Domains pass@1; PA 12.0%, UM 20.0%, CM 8.0%; submitted 2026-07-24; run date not published | CHI-Bench | 13.3% | 26 of 44 source board: 45 |
| codex + gpt-5.6-luna All Domains pass@1; PA 20.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published | CHI-Bench | 13.3% | 27 of 44 source board: 45 |
| GPT-4.1 mini Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.49 | 10 of 18 |
| OpenAI o3 95% bootstrap CI 61.3–65.9; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 63.6% | 5 of 22 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does OpenAI appear on?
As of September 28, 2026, OpenAI models hold 123 indexed results across 17 tracked benchmarks, through GPT-5.6 Sol, GPT-5.4, GPT-5.5, GPT-6 Astra (Anthropic run), GPT-6 Sol, GPT-6 Luna, GPT-5.6 Luna, GPT-5, GPT-5.2, GPT-5.6 Terra, GPT-5.1, GPT-5.5 Instant, GPT OSS 120B, GPT-5.3 Chat, GPT OSS 20B, Codex (GPT-5.6-sol), Codex (GPT 5.5), Codex (GPT 5.4), Codex (GPT 5.4 Mini), o3, GPT 5 Mini, GPT 5.4 (xhigh), GPT 5.4 Nano, o4 Mini, GPT 5 Nano, GPT-4.1, GPT-4o, GPT-5.4 mini, Codex (GPT 5.3), codex + gpt-5.6-terra, codex + gpt-5.6-luna, GPT-4.1 mini, OpenAI o3.
Where does OpenAI have the highest indexed score?
OpenAI models have the highest indexed score on MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol, 60.2%); MedXpertQA (MM) (GPT-5.6 Sol, 81.5); EHR-Complex (GPT-5.4, 0.65); HealthBench Professional (GPT-6 Astra (Anthropic run), 0.703).
Other labs with pages: Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.