Grok 4.20: healthcare benchmark results
xAI · 2 boards · snapshot reviewed September 28, 2026
The index currently holds 2 results for Grok 4.20: 65.8% on MedXpertQA (MM) (13 of 22 indexed rows), 5.3 ± 3.2 on PhysicianBench [Grok-4.20] (21 of 21 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| MedXpertQA (MM) Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 65.8% | 13 of 22 | 2026-04 |
| PhysicianBench Grok-4.20 via paper, arxiv.org | 5.3 ± 3.2 | 21 of 21 | 2026-05 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Grok-4.20".
Which healthcare benchmarks is Grok 4.20 scored on?
As of September 28, 2026, Grok 4.20 has indexed results on 2 tracked benchmarks: MedXpertQA (MM), PhysicianBench.
Which results are indexed for Grok 4.20?
Grok 4.20 stands at 65.8% on MedXpertQA (MM) (13 of 22 indexed rows), 5.3 ± 3.2 on PhysicianBench [Grok-4.20] (21 of 21 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.