Gemini 3.1 Pro: healthcare benchmark results
Google · 9 boards · snapshot reviewed September 28, 2026
The index currently holds 9 results for Gemini 3.1 Pro: 58.9% on MAST (Medical AI Superintelligence Test) (4 of 8 indexed rows), 0.652 on MedHELM [Gemini 3.1 Pro (Preview)] (1 of 10 indexed rows), 62.6% on First, Do NOHARM (v2) (12 of 17 indexed rows), 59.06% on MedCode (Vals AI) [Gemini 3.1 Pro Preview (02/26)] (2 of 102 indexed rows), 76.11% on MedScribe (Vals AI) [Gemini 3.1 Pro Preview (02/26)] (72 of 104 indexed rows), 81.3% on MedXpertQA (MM) (2 of 22 indexed rows), 6.0 ± 1.0 on PhysicianBench [Gemini Pro 3.1] (20 of 21 indexed rows), 0.63 on EHR-Complex (2 of 18 indexed rows), 11.9% on HealthAdminBench (6 of 7 indexed rows). It has the highest indexed score on MedHELM. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| MAST (Medical AI Superintelligence Test) | 58.9% | 4 of 8 source board: 11 | 2026-08 |
| MedHELM Gemini 3.1 Pro (Preview) | 0.652 | 1 of 10 source board: 11 | 2026-05 |
| First, Do NOHARM (v2) | 62.6% | 12 of 17 source board: 19 | 2026-08 |
| MedCode (Vals AI) Gemini 3.1 Pro Preview (02/26) model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not published | 59.06% | 2 of 102 | 2026-09-26 |
| MedScribe (Vals AI) Gemini 3.1 Pro Preview (02/26) model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not published | 76.11% | 72 of 104 | 2026-09-26 |
| MedXpertQA (MM) Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 81.3% | 2 of 22 | 2026-04 |
| PhysicianBench Gemini Pro 3.1 via paper, arxiv.org | 6.0 ± 1.0 | 20 of 21 | 2026-05 |
| EHR-Complex validation configuration via paper, arxiv.org | 0.63 | 2 of 18 | 2026-06 |
| HealthAdminBench screenshot-only, task description + portal guidance via paper, arxiv.org | 11.9% | 6 of 7 | 2026-04 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Gemini 3.1 Pro (Preview)" and "Gemini 3.1 Pro Preview (02/26)" and "Gemini Pro 3.1".
Which healthcare benchmarks is Gemini 3.1 Pro scored on?
As of September 28, 2026, Gemini 3.1 Pro has indexed results on 9 tracked benchmarks: MAST (Medical AI Superintelligence Test), MedHELM, First, Do NOHARM (v2), MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM), PhysicianBench, EHR-Complex, HealthAdminBench.
Which results are indexed for Gemini 3.1 Pro?
Gemini 3.1 Pro stands at 58.9% on MAST (Medical AI Superintelligence Test) (4 of 8 indexed rows), 0.652 on MedHELM [Gemini 3.1 Pro (Preview)] (1 of 10 indexed rows), 62.6% on First, Do NOHARM (v2) (12 of 17 indexed rows), 59.06% on MedCode (Vals AI) [Gemini 3.1 Pro Preview (02/26)] (2 of 102 indexed rows), 76.11% on MedScribe (Vals AI) [Gemini 3.1 Pro Preview (02/26)] (72 of 104 indexed rows), 81.3% on MedXpertQA (MM) (2 of 22 indexed rows), 6.0 ± 1.0 on PhysicianBench [Gemini Pro 3.1] (20 of 21 indexed rows), 0.63 on EHR-Complex (2 of 18 indexed rows), 11.9% on HealthAdminBench (6 of 7 indexed rows). It has the highest indexed score on MedHELM.
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.