Google: healthcare benchmark results
25 models · 61 results · snapshot reviewed September 28, 2026
Everything the index holds for Google, gathered in one place: Gemini 3.1 Pro, Gemini 2.5 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3 Flash, Gemini 3.8 Flash, Gemini 3.5 Flash Lite, Gemini 3.7 Flash, Gemini 3 Pro (11/25), Gemini 3.1 Flash Lite Preview, Gemini 2.5 Flash Preview (9/25) (Nonthinking), Gemini 2.5 Flash (7/17) (Thinking), Gemini 2.5 Flash Lite (9/25) (Thinking), Gemini 2.5 Flash Lite (Nonthinking), Gemini 3.6, Gemini 2.0 Flash, gemini-cli + gemini-3-flash, gemini-cli + gemini-3.1-pro, Gemini 3 Pro, Gemma 4 31B, Gemma 4 26B A4B, Gemma 4 12B, Gemma 4 E4B, Gemma 4 E2B, Gemini 2.5 Flash. The lab has the highest indexed score on MedHELM. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Gemini 3.1 Pro | MAST (Medical AI Superintelligence Test) | 58.9% | 4 of 8 source board: 11 |
| Gemini 3.1 Pro (Preview) | MedHELM | 0.652 | 1 of 10 source board: 11 |
| Gemini 3.1 Pro | First, Do NOHARM (v2) | 62.6% | 12 of 17 source board: 19 |
| Gemini 3.1 Pro Preview (02/26) model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 59.06% | 2 of 102 |
| Gemini 3.1 Pro Preview (02/26) model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 76.11% | 72 of 104 |
| Gemini 3.1 Pro Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | MedXpertQA (MM) | 81.3% | 2 of 22 |
| Gemini Pro 3.1 | PhysicianBench | 6.0 ± 1.0 | 20 of 21 |
| Gemini 3.1 Pro validation configuration | EHR-Complex | 0.63 | 2 of 18 |
| Gemini 3.1 Pro screenshot-only, task description + portal guidance | HealthAdminBench | 11.9% | 6 of 7 |
| Gemini 2.5 Pro | MedHELM | 0.529 | 6 of 10 source board: 11 |
| Gemini 2.5 Pro | First, Do NOHARM (v2) | 61.9% | 13 of 17 source board: 19 |
| Gemini 2.5 Pro model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.015389/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 50.59% | 15 of 102 |
| Gemini 2.5 Pro model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 1.91 pp; $0.046379/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 73.55% | 79 of 104 |
| Gemini 2.5 Pro Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.31 | 16 of 18 |
| Gemini 2.5 Pro 95% bootstrap CI 32.7–38.1; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 35.3% | 21 of 22 |
| Gemini 3.6 Flash | MAST (Medical AI Superintelligence Test) | 59.3% | 3 of 8 source board: 11 |
| Gemini 3.6 Flash model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.157 pp; $0.044216/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 53.15% | 10 of 102 |
| Gemini 3.6 Flash model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.861 pp; $0.073178/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 79.66% | 57 of 104 |
| Gemini 3.5 Flash | MedHELM | 0.642 | 2 of 10 source board: 11 |
| Gemini 3.5 Flash model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.073716/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 55.83% | 5 of 102 |
| Gemini 3.5 Flash model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.166341/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 76.57% | 70 of 104 |
| Gemini 3 Flash (12/25) model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.112 pp; $0.006187/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 55.92% | 4 of 102 |
| Gemini 3 Flash (12/25) model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.014379/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 69.92% | 90 of 104 |
| Gemini 3 Flash Preview | WHBench | 64.7% | 4 of 22 |
| Gemini 3.8 Flash model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.18 pp; $0.017700/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 48.13% | 27 of 102 |
| Gemini 3.8 Flash model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.943 pp; $0.025238/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 84.50% | 31 of 104 |
| Gemini 3.8 Flash (high) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 42 | 13 of 25 source board: 77 |
| Gemini 3.5 Flash Lite model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.951 pp; $0.008093/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 43.49% | 40 of 102 |
| Gemini 3.5 Flash Lite model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.031 pp; $0.019867/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 70.89% | 88 of 104 |
| Gemini 3.5 Flash-Lite Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 23 | 23 of 25 source board: 77 |
| Gemini 3.7 Flash model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.12 pp; $0.038331/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 53.39% | 8 of 102 |
| Gemini 3.7 Flash model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.004 pp; $0.058736/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 83.94% | 36 of 104 |
| Gemini 3 Pro (11/25) model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.073 pp; $0.028248/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 52.20% | 13 of 102 |
| Gemini 3 Pro (11/25) model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.9 pp; $0.061162/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 72.04% | 86 of 104 |
| Gemini 3.1 Flash Lite Preview model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.071 pp; $0.002029/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 47.60% | 28 of 102 |
| Gemini 3.1 Flash Lite Preview model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.823 pp; $0.002195/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 63.90% | 97 of 104 |
| Gemini 2.5 Flash Preview (9/25) (Nonthinking) model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.003692/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 40.54% | 59 of 102 |
| Gemini 2.5 Flash Preview (9/25) (Thinking) model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.003653/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 40.33% | 62 of 102 |
| Gemini 2.5 Flash Preview (9/25) (Thinking) model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.014526/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 78.50% | 60 of 104 |
| Gemini 2.5 Flash Preview (9/25) (Nonthinking) model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.014385/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 77.95% | 63 of 104 |
| Gemini 2.5 Flash (7/17) (Thinking) model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.952 pp; $0.003661/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 40.36% | 61 of 102 |
| Gemini 2.5 Flash (7/17) (Nonthinking) model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.003698/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 38.42% | 67 of 102 |
| Gemini 2.5 Flash (7/17) (Thinking) model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.908 pp; $0.014824/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 82.98% | 44 of 104 |
| Gemini 2.5 Flash (7/17) (Nonthinking) model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.909 pp; $0.014869/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 82.87% | 46 of 104 |
| Gemini 2.5 Flash Lite (9/25) (Thinking) model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.736 pp; $0.001182/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 34.19% | 76 of 102 |
| Gemini 2.5 Flash Lite (9/25) (Nonthinking) model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.911 pp; $0.001440/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 27.08% | 98 of 102 |
| Gemini 2.5 Flash Lite (9/25) (Nonthinking) model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.851 pp; $0.001332/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 75.82% | 74 of 104 |
| Gemini 2.5 Flash Lite (9/25) (Thinking) model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.002567/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 66.88% | 95 of 104 |
| Gemini 2.5 Flash Lite (Nonthinking) model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.843 pp; $0.001342/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 27.11% | 97 of 102 |
| Gemini 2.5 Flash Lite (Nonthinking) model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.982 pp; $0.001211/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 72.83% | 81 of 104 |
| Gemini 3.6 Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 36.2–43.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | Health Optimization Bench | 39.7 | 9 of 16 |
| Gemini 2.0 Flash | MedHELM | 0.342 | 10 of 10 source board: 11 |
| gemini-cli + gemini-3-flash All Domains pass@1; PA 18.7%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published | CHI-Bench | 12.5% | 28 of 44 source board: 45 |
| gemini-cli + gemini-3.1-pro All Domains pass@1; PA 14.7%, UM 6.7%, CM 0.0%; submitted 2026-05-01; run date not published | CHI-Bench | 7.1% | 36 of 44 source board: 45 |
| Gemini 3 Pro Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors | MedXpertQA (MM) | 76.0% | 7 of 22 |
| Gemma 4 31B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | MedXpertQA (MM) | 61.3% | 17 of 22 |
| Gemma 4 26B A4B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | MedXpertQA (MM) | 58.1% | 18 of 22 |
| Gemma 4 12B Google's Gemma 4 model card, Unified 12B; protocol not stated | MedXpertQA (MM) | 48.7% | 19 of 22 |
| Gemma 4 E4B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | MedXpertQA (MM) | 28.7% | 21 of 22 |
| Gemma 4 E2B Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | MedXpertQA (MM) | 23.5% | 22 of 22 |
| Gemini 2.5 Flash 95% bootstrap CI 47.0–52.0; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 49.5% | 13 of 22 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Google appear on?
As of September 28, 2026, Google models hold 61 indexed results across 13 tracked benchmarks, through Gemini 3.1 Pro, Gemini 2.5 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3 Flash, Gemini 3.8 Flash, Gemini 3.5 Flash Lite, Gemini 3.7 Flash, Gemini 3 Pro (11/25), Gemini 3.1 Flash Lite Preview, Gemini 2.5 Flash Preview (9/25) (Nonthinking), Gemini 2.5 Flash (7/17) (Thinking), Gemini 2.5 Flash Lite (9/25) (Thinking), Gemini 2.5 Flash Lite (Nonthinking), Gemini 3.6, Gemini 2.0 Flash, gemini-cli + gemini-3-flash, gemini-cli + gemini-3.1-pro, Gemini 3 Pro, Gemma 4 31B, Gemma 4 26B A4B, Gemma 4 12B, Gemma 4 E4B, Gemma 4 E2B, Gemini 2.5 Flash.
Where does Google have the highest indexed score?
Google models have the highest indexed score on MedHELM (Gemini 3.1 Pro, 0.652).
Other labs with pages: OpenAI, Anthropic, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.