xAI: healthcare benchmark results
4 models · 13 results · snapshot reviewed September 28, 2026
Everything the index holds for xAI, gathered in one place: Grok 4.7, Grok 4.6, Grok 4.3, Grok 4.20. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Grok 4.7 SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed. | HealthBench Professional | 0.567 | 15 of 29 |
| Grok 4.7 model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 49.55% | 19 of 102 |
| Grok 4.7 model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 89.38% | 6 of 104 |
| Grok 4.7 (xhigh) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 47 | 6 of 25 source board: 77 |
| Grok 4.6 SpaceXAI vendor report at high effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed. | HealthBench Professional | 0.485 | 21 of 29 |
| Grok 4.6 Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.8–69.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | Health Optimization Bench | 66.8 | 3 of 16 |
| Grok 4.6 model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.256 pp; $0.050335/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 44.71% | 36 of 102 |
| Grok 4.6 model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.956 pp; $0.037190/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 86.53% | 19 of 104 |
| Grok 4.3 | MAST (Medical AI Superintelligence Test) | 53.7% | 8 of 8 source board: 11 |
| Grok 4.3 model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.081 pp; $0.022202/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 38.07% | 69 of 102 |
| Grok 4.3 model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.019 pp; $0.015293/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 74.40% | 77 of 104 |
| Grok 4.20 Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | MedXpertQA (MM) | 65.8% | 13 of 22 |
| Grok-4.20 | PhysicianBench | 5.3 ± 3.2 | 21 of 21 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does xAI appear on?
As of September 28, 2026, xAI models hold 13 indexed results across 8 tracked benchmarks, through Grok 4.7, Grok 4.6, Grok 4.3, Grok 4.20.
Where does xAI have the highest indexed score?
xAI does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.