Health Evals

xAI logoxAI: healthcare benchmark results

4 models · 13 results · snapshot reviewed September 28, 2026

Everything the index holds for xAI, gathered in one place: Grok 4.7, Grok 4.6, Grok 4.3, Grok 4.20. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Grok 4.7
SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.
HealthBench Professional0.56715 of 29
Grok 4.7
model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.55%19 of 102
Grok 4.7
model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)89.38%6 of 104
Grok 4.7 (xhigh)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index476 of 25
source board: 77
Grok 4.6
SpaceXAI vendor report at high effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.
HealthBench Professional0.48521 of 29
Grok 4.6
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.8–69.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench66.83 of 16
Grok 4.6
model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.256 pp; $0.050335/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)44.71%36 of 102
Grok 4.6
model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.956 pp; $0.037190/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)86.53%19 of 104
Grok 4.3MAST (Medical AI Superintelligence Test)53.7%8 of 8
source board: 11
Grok 4.3
model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.081 pp; $0.022202/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)38.07%69 of 102
Grok 4.3
model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.019 pp; $0.015293/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)74.40%77 of 104
Grok 4.20
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
MedXpertQA (MM)65.8%13 of 22
Grok-4.20PhysicianBench5.3 ± 3.221 of 21

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does xAI appear on?

As of September 28, 2026, xAI models hold 13 indexed results across 8 tracked benchmarks, through Grok 4.7, Grok 4.6, Grok 4.3, Grok 4.20.

Where does xAI have the highest indexed score?

xAI does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.