Health Evals

OpenAI logoOpenAI: healthcare benchmark results

33 models · 123 results · snapshot reviewed September 28, 2026

Everything the index holds for OpenAI, gathered in one place: GPT-5.6 Sol, GPT-5.4, GPT-5.5, GPT-6 Astra (Anthropic run), GPT-6 Sol, GPT-6 Luna, GPT-5.6 Luna, GPT-5, GPT-5.2, GPT-5.6 Terra, GPT-5.1, GPT-5.5 Instant, GPT OSS 120B, GPT-5.3 Chat, GPT OSS 20B, Codex (GPT-5.6-sol), Codex (GPT 5.5), Codex (GPT 5.4), Codex (GPT 5.4 Mini), o3, GPT 5 Mini, GPT 5.4 (xhigh), GPT 5.4 Nano, o4 Mini, GPT 5 Nano, GPT-4.1, GPT-4o, GPT-5.4 mini, Codex (GPT 5.3), codex + gpt-5.6-terra, codex + gpt-5.6-luna, GPT-4.1 mini, OpenAI o3. The lab has the highest indexed score on MAST (Medical AI Superintelligence Test), MedXpertQA (MM), EHR-Complex, HealthBench Professional. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
GPT-5.6 Sol
length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
HealthBench Professional0.6059 of 29
GPT-5.6 Sol (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
HealthBench Professional0.54018 of 29
GPT-5.6 Sol
length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.
HealthBench Hard0.3317 of 20
GPT-5.6 Sol (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)
HealthBench Hard0.31411 of 20
GPT-5.6 Sol
length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09
HealthBench57.013 of 26
GPT-5.6 Sol (August)
ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06
HealthBench55.018 of 26
GPT-5.6 Sol (max)
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort.
Health Optimization Bench66.64 of 16
GPT-5.6 Sol (high)
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort.
Health Optimization Bench64.65 of 16
GPT-5.6 Sol
MAST in preview; 'exact scores may change'
MAST (Medical AI Superintelligence Test)60.2%1 of 8
source board: 11
GPT-5.6 SolFirst, Do NOHARM (v2)70.1%8 of 17
source board: 19
GPT-5.6 Sol
model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)43.97%39 of 102
GPT-5.6 Sol
model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.23%27 of 104
GPT-5.6 Sol
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
MedXpertQA (MM)81.51 of 22
GPT-5.6 Sol (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index458 of 25
source board: 77
GPT-5.4
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
HealthBench Professional0.48122 of 29
GPT-5.4
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
HealthBench Hard0.29115 of 20
GPT-5.4
OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6.
HealthBench54.021 of 26
GPT-5.4 (2026-03-05)MedHELM0.5385 of 10
source board: 11
GPT-5.4
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
MedXpertQA (MM)77.1%6 of 22
GPT-5.4PhysicianBench27.7 ± 1.512 of 21
GPT-5.4 (high reasoning)
average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
EHR-Complex0.651 of 18
GPT-5.4 (low reasoning)
validation configuration
EHR-Complex0.586 of 18
GPT-5.4
95% CI 64.5-69.2
WHBench66.8%3 of 22
GPT-5.4 (computer-use agent)
screenshot-only, task description + portal guidance; subtask rate 82.8%
HealthAdminBench26.7%2 of 7
GPT-5.4 (standardized harness)
screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
HealthAdminBench5.9%7 of 7
GPT-5.5
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
HealthBench Professional0.51820 of 29
GPT-5.5
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)
HealthBench Hard0.31510 of 20
GPT-5.5
length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system card
HealthBench56.516 of 26
GPT-5.5First, Do NOHARM (v2)70.0%9 of 17
source board: 19
GPT 5.5
model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.188 pp; $0.160759/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.10%23 of 102
GPT 5.5
model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.142988/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)86.87%17 of 104
GPT-5.5
pass@1; Pass^3 28.0
PhysicianBench46.3 ± 1.27 of 21
GPT-6 Astra (Anthropic run)
Anthropic reproduction of GPT-6 Astra through the public API; max effort; no system prompt; Claude Opus 4.8 grader; length-adjusted 70.3%, raw 74.0%. Different grader/protocol from OpenAI’s own 64.7% report.
HealthBench Professional0.7031 of 29
GPT-6 Astra
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 68.2%, mean answer 3,185 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench Professional0.6474 of 29
GPT-6 Astra
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 34.2%, mean answer 1,697 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench Hard0.3664 of 20
GPT-6 Astra
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 56.9%, mean answer 1,760 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench58.39 of 26
GPT-6 Astra
model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 2.131 pp; $0.451358/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)48.49%25 of 102
GPT-6 Astra
model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 1.938 pp; $0.581991/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)87.91%14 of 104
GPT-6 Astra (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index524 of 25
source board: 77
GPT-6 Sol
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 59.5%, mean answer 1,573 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench Professional0.6087 of 29
GPT-6 Sol
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 22.1%, mean answer 974 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench Hard0.30113 of 20
GPT-6 Sol
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 47.1%, mean answer 977 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench53.223 of 26
GPT-6 Sol
model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 2.119 pp; $0.085827/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)47.07%32 of 102
GPT-6 Sol
model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 1.942 pp; $0.083138/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)82.03%48 of 104
GPT-6 Sol (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index4311 of 25
source board: 77
GPT-6 Luna
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench Professional0.6088 of 29
GPT-6 Luna
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench Hard0.31412 of 20
GPT-6 Luna
OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 50%, mean answer 1,255 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
HealthBench54.519 of 26
GPT-6 Luna
model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)44.69%37 of 102
GPT-6 Luna
model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)83.71%39 of 104
GPT-6 Luna (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index3717 of 25
source board: 77
GPT-5.6 Luna
length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
HealthBench Professional0.55716 of 29
GPT-5.6 Luna (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
HealthBench Professional0.44126 of 29
GPT-5.6 Luna
length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
HealthBench Hard0.3209 of 20
GPT-5.6 Luna (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)
HealthBench Hard0.28716 of 20
GPT-5.6 Luna
length-adjusted (55.4 unadjusted), max reasoning effort
HealthBench55.817 of 26
GPT-5.6 Luna (August)
ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars)
HealthBench53.322 of 26
GPT-5.6 Luna
model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)42.39%47 of 102
GPT-5.6 Luna
model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.39%32 of 104
GPT-5.6 Luna (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index3618 of 25
source board: 77
GPT-5
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
HealthBench Professional0.46223 of 29
GPT-5
length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
HealthBench Hard0.3475 of 20
GPT-5
OpenAI length-adjusted score, maximum reasoning effort; raw 63.1%, mean answer 2,904 characters; GPT-5.6 card Table 6.
HealthBench57.711 of 26
GPT-5
from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
First, Do NOHARM (v2)68.6%10 of 17
source board: 19
GPT 5
model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.63%18 of 102
GPT 5
model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)83.65%40 of 104
GPT-5.2
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
HealthBench Professional0.45924 of 29
GPT-5.2-High (Baichuan run)
Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.
HealthBench Hard0.4203 of 20
GPT-5.2
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)
HealthBench Hard0.3436 of 20
GPT-5.2-High
raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card.
HealthBench63.33 of 26
GPT-5.2
OpenAI length-adjusted score, maximum reasoning effort; raw 60.7%, mean answer 2,645 characters; GPT-5.6 card Table 6.
HealthBench56.815 of 26
GPT 5.2
model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.75%17 of 102
GPT 5.2
model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.39%33 of 104
GPT-5.2
Qwen-run comparison in the Qwen3.5-397B-A17B model card
MedXpertQA (MM)73.38 of 22
GPT-5.6 Terra
length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
HealthBench Professional0.57713 of 29
GPT-5.6 Terra
length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
HealthBench Hard0.3278 of 20
GPT-5.6 Terra
length-adjusted (58.7 unadjusted), max reasoning effort
HealthBench57.014 of 26
GPT-5.6 Terra
model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.173 pp; $0.046558/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)43.41%41 of 102
GPT-5.6 Terra
model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.948 pp; $0.062588/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)82.87%47 of 104
GPT-5.1
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
HealthBench Professional0.39627 of 29
GPT-5.1
length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)
HealthBench Hard0.25418 of 20
GPT-5.1
OpenAI length-adjusted score, maximum reasoning effort; raw 64.2%, mean answer 4,222 characters; GPT-5.6 card Table 6.
HealthBench50.925 of 26
GPT 5.1
model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 2.151 pp; $0.014371/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)52.73%12 of 102
GPT 5.1
model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 1.942 pp; $0.096508/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)88.09%12 of 104
GPT-5.5 Instant
length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
HealthBench Professional0.38428 of 29
GPT-5.5 Instant
length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
HealthBench Hard0.22919 of 20
GPT-5.5 Instant
length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11
HealthBench51.424 of 26
GPT OSS 120B
raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
HealthBench Hard0.30014 of 20
GPT OSS 120B
reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9)
HealthBench57.612 of 26
GPT-5.3 Chat
raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
HealthBench Hard0.25917 of 20
GPT-5.3 Chat
raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted)
HealthBench54.1%20 of 26
GPT OSS 20B
raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
HealthBench Hard0.10820 of 20
GPT OSS 20B
reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8)
HealthBench42.526 of 26
Codex (GPT-5.6-sol)
$5.2/task
HealthAgentBench45%2 of 12
codex + gpt-5.6-sol
All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not published
CHI-Bench25.3%7 of 44
source board: 45
Codex (GPT 5.5)
$2.8/task
HealthAgentBench42%3 of 12
codex + gpt-5.5
All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not published
CHI-Bench20.9%12 of 44
source board: 45
Codex (GPT 5.4)
$1.3/task
HealthAgentBench28%7 of 12
codex + gpt-5.4
All Domains pass@1; PA 24.0%, UM 17.3%, CM 6.7%; submitted 2026-05-01; run date not published
CHI-Bench16.0%20 of 44
source board: 45
Codex (GPT 5.4 Mini)
$0.6/task
HealthAgentBench16%12 of 12
codex + gpt-5.4-mini
All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not published
CHI-Bench8.4%34 of 44
source board: 45
o3
model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.161 pp; $0.029818/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)47.29%30 of 102
o3
model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.871 pp; $0.040334/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)76.65%69 of 104
GPT 5 Mini
model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.045 pp; $0.005560/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)43.05%44 of 102
GPT 5 Mini
model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.924 pp; $0.033478/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)80.58%52 of 104
GPT 5.4 (xhigh)
model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.148 pp; $0.212108/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.29%51 of 102
GPT 5.4 (xhigh)
model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 3.316 pp; $0.639282/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)77.55%64 of 104
GPT 5.4 Nano
model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 2.256 pp; $0.000844/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.03%55 of 102
GPT 5.4 Nano
model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.001800/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)77.09%67 of 104
o4 Mini
model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.021 pp; $0.017605/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)33.79%79 of 102
o4 Mini
model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.957 pp; $0.040605/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)69.14%92 of 104
GPT 5 Nano
model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.948 pp; $0.001729/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)30.44%91 of 102
GPT 5 Nano
model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.006961/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)72.86%80 of 104
GPT-4.1
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.4711 of 18
GPT-4.1WHBench51.8%11 of 22
GPT-4o
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.3115 of 18
GPT-4oWHBench44.6%16 of 22
GPT-5.4 miniMedHELM0.5524 of 10
source board: 11
Codex (GPT 5.3)
$1.0/task; harness and model evaluated jointly
HealthAgentBench22%9 of 12
codex + gpt-5.6-terra
All Domains pass@1; PA 12.0%, UM 20.0%, CM 8.0%; submitted 2026-07-24; run date not published
CHI-Bench13.3%26 of 44
source board: 45
codex + gpt-5.6-luna
All Domains pass@1; PA 20.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published
CHI-Bench13.3%27 of 44
source board: 45
GPT-4.1 mini
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.4910 of 18
OpenAI o3
95% bootstrap CI 61.3–65.9; 3 runs; temperature 0; zero-shot, closed-book
WHBench63.6%5 of 22

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does OpenAI appear on?

As of September 28, 2026, OpenAI models hold 123 indexed results across 17 tracked benchmarks, through GPT-5.6 Sol, GPT-5.4, GPT-5.5, GPT-6 Astra (Anthropic run), GPT-6 Sol, GPT-6 Luna, GPT-5.6 Luna, GPT-5, GPT-5.2, GPT-5.6 Terra, GPT-5.1, GPT-5.5 Instant, GPT OSS 120B, GPT-5.3 Chat, GPT OSS 20B, Codex (GPT-5.6-sol), Codex (GPT 5.5), Codex (GPT 5.4), Codex (GPT 5.4 Mini), o3, GPT 5 Mini, GPT 5.4 (xhigh), GPT 5.4 Nano, o4 Mini, GPT 5 Nano, GPT-4.1, GPT-4o, GPT-5.4 mini, Codex (GPT 5.3), codex + gpt-5.6-terra, codex + gpt-5.6-luna, GPT-4.1 mini, OpenAI o3.

Where does OpenAI have the highest indexed score?

OpenAI models have the highest indexed score on MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol, 60.2%); MedXpertQA (MM) (GPT-5.6 Sol, 81.5); EHR-Complex (GPT-5.4, 0.65); HealthBench Professional (GPT-6 Astra (Anthropic run), 0.703).

Other labs with pages: Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.