Health Evals

Google logoGoogle: healthcare benchmark results

25 models · 61 results · snapshot reviewed September 28, 2026

Everything the index holds for Google, gathered in one place: Gemini 3.1 Pro, Gemini 2.5 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3 Flash, Gemini 3.8 Flash, Gemini 3.5 Flash Lite, Gemini 3.7 Flash, Gemini 3 Pro (11/25), Gemini 3.1 Flash Lite Preview, Gemini 2.5 Flash Preview (9/25) (Nonthinking), Gemini 2.5 Flash (7/17) (Thinking), Gemini 2.5 Flash Lite (9/25) (Thinking), Gemini 2.5 Flash Lite (Nonthinking), Gemini 3.6, Gemini 2.0 Flash, gemini-cli + gemini-3-flash, gemini-cli + gemini-3.1-pro, Gemini 3 Pro, Gemma 4 31B, Gemma 4 26B A4B, Gemma 4 12B, Gemma 4 E4B, Gemma 4 E2B, Gemini 2.5 Flash. The lab has the highest indexed score on MedHELM. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Gemini 3.1 ProMAST (Medical AI Superintelligence Test)58.9%4 of 8
source board: 11
Gemini 3.1 Pro (Preview)MedHELM0.6521 of 10
source board: 11
Gemini 3.1 ProFirst, Do NOHARM (v2)62.6%12 of 17
source board: 19
Gemini 3.1 Pro Preview (02/26)
model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)59.06%2 of 102
Gemini 3.1 Pro Preview (02/26)
model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)76.11%72 of 104
Gemini 3.1 Pro
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
MedXpertQA (MM)81.3%2 of 22
Gemini Pro 3.1PhysicianBench6.0 ± 1.020 of 21
Gemini 3.1 Pro
validation configuration
EHR-Complex0.632 of 18
Gemini 3.1 Pro
screenshot-only, task description + portal guidance
HealthAdminBench11.9%6 of 7
Gemini 2.5 ProMedHELM0.5296 of 10
source board: 11
Gemini 2.5 ProFirst, Do NOHARM (v2)61.9%13 of 17
source board: 19
Gemini 2.5 Pro
model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.015389/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)50.59%15 of 102
Gemini 2.5 Pro
model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 1.91 pp; $0.046379/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)73.55%79 of 104
Gemini 2.5 Pro
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.3116 of 18
Gemini 2.5 Pro
95% bootstrap CI 32.7–38.1; 3 runs; temperature 0; zero-shot, closed-book
WHBench35.3%21 of 22
Gemini 3.6 FlashMAST (Medical AI Superintelligence Test)59.3%3 of 8
source board: 11
Gemini 3.6 Flash
model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.157 pp; $0.044216/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)53.15%10 of 102
Gemini 3.6 Flash
model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.861 pp; $0.073178/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)79.66%57 of 104
Gemini 3.5 FlashMedHELM0.6422 of 10
source board: 11
Gemini 3.5 Flash
model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.073716/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)55.83%5 of 102
Gemini 3.5 Flash
model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.166341/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)76.57%70 of 104
Gemini 3 Flash (12/25)
model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.112 pp; $0.006187/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)55.92%4 of 102
Gemini 3 Flash (12/25)
model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.014379/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)69.92%90 of 104
Gemini 3 Flash PreviewWHBench64.7%4 of 22
Gemini 3.8 Flash
model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.18 pp; $0.017700/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)48.13%27 of 102
Gemini 3.8 Flash
model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.943 pp; $0.025238/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.50%31 of 104
Gemini 3.8 Flash (high)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index4213 of 25
source board: 77
Gemini 3.5 Flash Lite
model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.951 pp; $0.008093/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)43.49%40 of 102
Gemini 3.5 Flash Lite
model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.031 pp; $0.019867/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)70.89%88 of 104
Gemini 3.5 Flash-Lite
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index2323 of 25
source board: 77
Gemini 3.7 Flash
model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.12 pp; $0.038331/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)53.39%8 of 102
Gemini 3.7 Flash
model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.004 pp; $0.058736/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)83.94%36 of 104
Gemini 3 Pro (11/25)
model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.073 pp; $0.028248/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)52.20%13 of 102
Gemini 3 Pro (11/25)
model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.9 pp; $0.061162/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)72.04%86 of 104
Gemini 3.1 Flash Lite Preview
model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.071 pp; $0.002029/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)47.60%28 of 102
Gemini 3.1 Flash Lite Preview
model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.823 pp; $0.002195/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)63.90%97 of 104
Gemini 2.5 Flash Preview (9/25) (Nonthinking)
model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.003692/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.54%59 of 102
Gemini 2.5 Flash Preview (9/25) (Thinking)
model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.003653/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.33%62 of 102
Gemini 2.5 Flash Preview (9/25) (Thinking)
model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.014526/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)78.50%60 of 104
Gemini 2.5 Flash Preview (9/25) (Nonthinking)
model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.014385/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)77.95%63 of 104
Gemini 2.5 Flash (7/17) (Thinking)
model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.952 pp; $0.003661/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.36%61 of 102
Gemini 2.5 Flash (7/17) (Nonthinking)
model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.003698/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)38.42%67 of 102
Gemini 2.5 Flash (7/17) (Thinking)
model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.908 pp; $0.014824/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)82.98%44 of 104
Gemini 2.5 Flash (7/17) (Nonthinking)
model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.909 pp; $0.014869/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)82.87%46 of 104
Gemini 2.5 Flash Lite (9/25) (Thinking)
model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.736 pp; $0.001182/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)34.19%76 of 102
Gemini 2.5 Flash Lite (9/25) (Nonthinking)
model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.911 pp; $0.001440/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)27.08%98 of 102
Gemini 2.5 Flash Lite (9/25) (Nonthinking)
model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.851 pp; $0.001332/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)75.82%74 of 104
Gemini 2.5 Flash Lite (9/25) (Thinking)
model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.002567/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)66.88%95 of 104
Gemini 2.5 Flash Lite (Nonthinking)
model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.843 pp; $0.001342/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)27.11%97 of 102
Gemini 2.5 Flash Lite (Nonthinking)
model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.982 pp; $0.001211/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)72.83%81 of 104
Gemini 3.6
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 36.2–43.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench39.79 of 16
Gemini 2.0 FlashMedHELM0.34210 of 10
source board: 11
gemini-cli + gemini-3-flash
All Domains pass@1; PA 18.7%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published
CHI-Bench12.5%28 of 44
source board: 45
gemini-cli + gemini-3.1-pro
All Domains pass@1; PA 14.7%, UM 6.7%, CM 0.0%; submitted 2026-05-01; run date not published
CHI-Bench7.1%36 of 44
source board: 45
Gemini 3 Pro
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
MedXpertQA (MM)76.0%7 of 22
Gemma 4 31B
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
MedXpertQA (MM)61.3%17 of 22
Gemma 4 26B A4B
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
MedXpertQA (MM)58.1%18 of 22
Gemma 4 12B
Google's Gemma 4 model card, Unified 12B; protocol not stated
MedXpertQA (MM)48.7%19 of 22
Gemma 4 E4B
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
MedXpertQA (MM)28.7%21 of 22
Gemma 4 E2B
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
MedXpertQA (MM)23.5%22 of 22
Gemini 2.5 Flash
95% bootstrap CI 47.0–52.0; 3 runs; temperature 0; zero-shot, closed-book
WHBench49.5%13 of 22

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Google appear on?

As of September 28, 2026, Google models hold 61 indexed results across 13 tracked benchmarks, through Gemini 3.1 Pro, Gemini 2.5 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3 Flash, Gemini 3.8 Flash, Gemini 3.5 Flash Lite, Gemini 3.7 Flash, Gemini 3 Pro (11/25), Gemini 3.1 Flash Lite Preview, Gemini 2.5 Flash Preview (9/25) (Nonthinking), Gemini 2.5 Flash (7/17) (Thinking), Gemini 2.5 Flash Lite (9/25) (Thinking), Gemini 2.5 Flash Lite (Nonthinking), Gemini 3.6, Gemini 2.0 Flash, gemini-cli + gemini-3-flash, gemini-cli + gemini-3.1-pro, Gemini 3 Pro, Gemma 4 31B, Gemma 4 26B A4B, Gemma 4 12B, Gemma 4 E4B, Gemma 4 E2B, Gemini 2.5 Flash.

Where does Google have the highest indexed score?

Google models have the highest indexed score on MedHELM (Gemini 3.1 Pro, 0.652).

Other labs with pages: OpenAI, Anthropic, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.