Health Evals Index
Which AI handles health questions well?
Our own index of where AI models land across 5 health boards, each result placed 0 to 100 on its board.
- Models ranked
- 58
- Boards in the index
- 5
- Updated
Latest changes
All updates- 30 SepArtificial Analysis Healthcare & Medical IndexMistral Medium 3.5, Inkling (xhigh), MiniMax-M3 and others17 added, 8 rescored
- 30 SepHealthBenchClaude Opus 5 (length-adjusted), GPT-6 Astra, Claude Fable 5 and others10 added, 1 rescored
Health Evals Index
- Anthropic
- OpenAI
- Meta
boards with a result
89.5
- HealthBench Professional0.66087.8system card
- HealthBench62.782.1system card
- Health Optimization Bench83.898.6official leaderboard
- HealthBench Hardno result on this board
- WHBenchno result on this board
- HealthBench ProfessionalWhen a doctor brings a real work question to an AI, would another doctor sign off on the reply?
- GPT-6 Astra (Anthropic run)0.703
- Claude Sonnet 5.50.692
- Claude Fable 50.660
- Claude Opus 5.50.656
- GPT-6 Astra0.647
- HealthBenchWhen someone asks an AI about their own health, do the answers meet a physician's standard?
- Claude Opus 567.1
- Claude Sonnet 5.565.4
- Baichuan-M365.1
- GPT-5.2-High63.3
- Claude Fable 562.7
- Health Optimization BenchCan a model keep up with fast-moving evidence in preventive medicine?
- Claude Fable 5.184.9
- Claude Fable 583.8
- GPT-6 Astra (max)82.2
- Grok 4.681.3
- Claude Opus 578.3
- HealthBench HardHow do models do on the health conversations that tripped them up most?
- Muse Spark0.428
- GPT-5.2-High0.420
- GPT-6 Astra0.366
- GPT-50.347
- GPT-5.20.343
- WHBenchAre questions about women's health answered safely and to current guidelines?
- Claude Opus 4.672.1
- Claude Sonnet 4.667.1
- GPT-5.466.8
- Gemini 3 Flash Preview64.7
- OpenAI o363.6
- Artificial Analysis Healthcare & Medical IndexHow do general-purpose models rank on an index weighted toward health work?
- Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)61
- Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)58
- Claude Opus 5 (Adaptive Reasoning, Max Effort)53
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)52
- GPT-6 Astra (max)52
- Health Optimization Bench: subject suitesDoes that hold when the questions span many subjects instead of one?
- Claude Fable 570.9
- Claude Opus 569.3
- Grok 4.666.8
- GPT-5.6 Sol (max)66.6
- GPT-5.6 Sol (high)64.6
Earlier changes
All updates- 30 SepHealthBench HardGPT-6 Astra, GPT-6 Sol, GPT-6 Luna and others3 added, 1 rescored
- 30 SepHealthBench ProfessionalClaude Opus 4.8 (Opus 4.8 grader), Claude Fable 5 (September card), GPT-6 Astra and others9 added, 1 rescored
- 30 SepWHBenchLlama 4 Scout, Gemini 2.5 Pro, Llama 3.1 405B and others16 added
- 10 SepHealth Optimization BenchGemini 3.8 Flash, GPT-6 Astra (max)2 added
- 8 SepHealthBench HardGPT-6 Astra1 added
- 8 SepHealthBenchGPT-6 Astra, Claude Fable 5.12 added
Lab standings
- Anthropic Claude Fable 53 of 5 boards89.5
- OpenAI GPT-6 Astra4 of 5 boards85.3
- Meta Muse Spark3 of 5 boards77.6
- SpaceX AI Grok 4.62 of 5 boards54.5
- Baichuan Baichuan-M31 of 5 boards53.0
- Moonshot AI Kimi K31 of 5 boards52.4
- Google Gemini 3 Flash1 of 5 boards46.1
- DeepSeek DeepSeek-V3.21 of 5 boards40.8
- Mistral Mistral Large1 of 5 boards39.1
- Thinking Machines Inkling1 of 5 boards34.5
- MiniMax MiniMax M31 of 5 boards20.7
- Zhipu GLM 5.21 of 5 boards16.0
- Microsoft MAI-Thinking-12 of 5 boards13.4
- NVIDIA Llama-3.1-Nemotron-70B-Instruct1 of 5 boards6.4
Retired evaluations (11)
- OpenAI Dynamic Mental Health Evaluations: Removed from the index on 2026-09-08 at the owner's request: not treated as a real, established benchmark.
- HealthBench Consensus: near-saturated physician-consensus baseline; frontier runs stopped reporting it separately
- MedQA / MultiMedQA: exam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers
- AgentClinic: no public frontier-model results since 2025
- CRAFT-MD: no public frontier-model results since 2025
- MedAgentBench: v2 lives on inside the MAST composite; the standalone board has no current frontier rows
- SDBench / MAI-DxO: Microsoft's 2025 sequential-diagnosis study was not re-run on current models
- Open Medical-LLM Leaderboard (Hugging Face): built on saturated exam sets; no frontier submissions in 2026
- MedArena: clinician preference arena; ratings pool too thin on current frontier models to quote
- AMIE evaluations: Google DeepMind research prototypes, never opened to cross-vendor comparison
- LiveClin, PrIME-LLM, MedMCP-Calc: single studies with two or fewer current-frontier rows; tracked for a future qualifying update