Health Evals

MedHELM: published results

Stanford CRFM / HAI and multi-institution collaborators · 121 tasks / 31 datasets · index updated September 28, 2026

Gemini 3.1 Pro (Preview) has the highest indexed numerical score on MedHELM, 0.652 as of 2026-05, per MedHELM leaderboard (medhelm.org), v5.0.0. Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate.

Published results

Scores preserve their source precision, with any scale conversion documented (official leaderboard). Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes. The current official home-page ranking is v5.0.0, updated May 14, 2026; ten of eleven models are displayed. September 28 is a source check, not a new evaluation. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherStanford CRFM / HAI and multi-institution collaborators
categorycomposite indices
released2025-02
size121 tasks / 31 datasets
scalemean win rate 0-1, higher better
result basisofficial leaderboard
sourceMedHELM leaderboard (medhelm.org), v5.0.0
last frontier result2026-05

What is MedHELM?

MedHELM is a composite benchmark from Stanford CRFM, released 2025-02: 121 tasks / 31 datasets, scored on a mean win rate 0-1 scale. Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate.

Which model leads MedHELM?

Gemini 3.1 Pro (Preview) (Google) has the highest indexed numerical score on MedHELM at 0.652 (evaluation setups may differ), per MedHELM leaderboard (medhelm.org), v5.0.0, as of 2026-05.

Where do the MedHELM numbers come from?

From MedHELM leaderboard (medhelm.org), v5.0.0 (official leaderboard). Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes. The current official home-page ranking is v5.0.0, updated May 14, 2026; ten of eleven models are displayed. September 28 is a source check, not a new evaluation.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.