Health Evals

Health Evals Index

Which AI handles health questions well?

Our own index of where AI models land across 5 health boards, each result placed 0 to 100 on its board.

Models ranked
58
Boards in the index
5
Updated

Across the site: 7 boards, 163 results, 69 models.

Latest changes

All updates

Health Evals Index

Our own calculation from 5 boards: each result placed 0 to 100 within its board, averaged per model, then scaled down when a score rests on fewer than 3 boards. How it works

  • Anthropic
  • OpenAI
  • Meta

boards with a result

Every model with a result on any of the 5 boards is ranked. The squares show how many boards each score rests on. Scores from fewer than 3 boards are scaled down: ×0.816 for 2 boards, ×0.577 for 1 board.

Claude Fable 5Rank 1 of 58, Anthropic
89.5mean, 3 of 5 boards

The mean of these 3 placements, 89.5, kept in full: it rests on 3 or more boards.

Index against price

38 of 58 ranked models have a list price. Joined points: nothing cheaper scores higher.

Blended API list price, USD per 1M tokens (3 parts input to 1 part output), log scale

Index over time

55 of 58 ranked models have a release date. Joined points set a new high when they came out.

Release date

Each board ranks its own results only. Hatched bars were reported by the model's maker.

Earlier changes

All updates

Lab standings

Each lab's best model on the Health Evals Index. Squares show how many of the 5 boards its score rests on.

  1. Anthropic Claude Fable 53 of 5 boards89.5
  2. OpenAI GPT-6 Astra4 of 5 boards85.3
  3. Meta Muse Spark3 of 5 boards77.6
  4. SpaceX AI Grok 4.62 of 5 boards54.5
  5. Baichuan Baichuan-M31 of 5 boards53.0
  6. Moonshot AI Kimi K31 of 5 boards52.4
  7. Google Gemini 3 Flash1 of 5 boards46.1
  8. DeepSeek DeepSeek-V3.21 of 5 boards40.8
  9. Mistral Mistral Large1 of 5 boards39.1
  10. Thinking Machines Inkling1 of 5 boards34.5
  11. MiniMax MiniMax M31 of 5 boards20.7
  12. Zhipu GLM 5.21 of 5 boards16.0
  13. Microsoft MAI-Thinking-12 of 5 boards13.4
  14. NVIDIA Llama-3.1-Nemotron-70B-Instruct1 of 5 boards6.4
Retired evaluations (11)
  • OpenAI Dynamic Mental Health Evaluations: Removed from the index on 2026-09-08 at the owner's request: not treated as a real, established benchmark.
  • HealthBench Consensus: near-saturated physician-consensus baseline; frontier runs stopped reporting it separately
  • MedQA / MultiMedQA: exam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers
  • AgentClinic: no public frontier-model results since 2025
  • CRAFT-MD: no public frontier-model results since 2025
  • MedAgentBench: v2 lives on inside the MAST composite; the standalone board has no current frontier rows
  • SDBench / MAI-DxO: Microsoft's 2025 sequential-diagnosis study was not re-run on current models
  • Open Medical-LLM Leaderboard (Hugging Face): built on saturated exam sets; no frontier submissions in 2026
  • MedArena: clinician preference arena; ratings pool too thin on current frontier models to quote
  • AMIE evaluations: Google DeepMind research prototypes, never opened to cross-vendor comparison
  • LiveClin, PrIME-LLM, MedMCP-Calc: single studies with two or fewer current-frontier rows; tracked for a future qualifying update