Grok 4.7: healthcare benchmark results
xAI · 4 boards · snapshot reviewed September 28, 2026
The index currently holds 4 results for Grok 4.7: 0.567 on HealthBench Professional (15 of 29 indexed rows), 49.55% on MedCode (Vals AI) (19 of 102 indexed rows), 89.38% on MedScribe (Vals AI) (6 of 104 indexed rows), 47 on Artificial Analysis Healthcare & Medical Index [Grok 4.7 (xhigh)] (6 of 25 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed. | 0.567 | 15 of 29 | 2026-09 |
| MedCode (Vals AI) model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not published | 49.55% | 19 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not published | 89.38% | 6 of 104 | 2026-09-26 |
| Artificial Analysis Healthcare & Medical Index Grok 4.7 (xhigh) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 47 | 6 of 25 source board: 77 | 2026-09 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Grok 4.7 (xhigh)".
Which healthcare benchmarks is Grok 4.7 scored on?
As of September 28, 2026, Grok 4.7 has indexed results on 4 tracked benchmarks: HealthBench Professional, MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index.
Which results are indexed for Grok 4.7?
Grok 4.7 stands at 0.567 on HealthBench Professional (15 of 29 indexed rows), 49.55% on MedCode (Vals AI) (19 of 102 indexed rows), 89.38% on MedScribe (Vals AI) (6 of 104 indexed rows), 47 on Artificial Analysis Healthcare & Medical Index [Grok 4.7 (xhigh)] (6 of 25 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.