Health Evals

Composite indices

3 tracked · snapshot reviewed September 28, 2026

Composite indices combine several evaluations using their publishers' weighting and normalization rules. Read the component scores and methodology alongside the aggregate: indices with different inputs or weights are not directly comparable.

MAST (Medical AI Superintelligence Test)

ARISE AI Research Network · 6-benchmark composite

A composite of clinical benchmarks spanning diagnostic and management reasoning, safety, multimodal imaging, and agentic capability.

MedHELM

Stanford CRFM · 121 tasks

Holistic clinical evaluation across 121 tasks in a clinician-validated taxonomy, ranked by mean win rate.

Artificial Analysis Healthcare & Medical Index

Artificial Analysis · 6-evaluation composite

A healthcare-weighted composite of six evaluations, covering medical knowledge, long records, knowledge work, reasoning and tools.

Which composite indices have results in this index?

3 as of September 28, 2026: MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol: highest indexed score 60.2%); MedHELM (Gemini 3.1 Pro (Preview): highest indexed score 0.652); Artificial Analysis Healthcare & Medical Index (Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback): highest indexed score 61).

The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks.