Health Evals

A public index of healthcare AI evaluations

Healthcare AI benchmarks

Explore reported model results across clinical reasoning, medical knowledge, safety, and real-world workflows. See what each benchmark measures, compare results within a board, and follow the evidence behind a score.

17 benchmarks204 models503 results49 source documents
Snapshot reviewed Review notes

The benchmarks

Compare within a board · scales differ

Filter by benchmark category

17 of 17 benchmarks shown

HealthBench Professional

OpenAI · 525 tasks

525 clinician tasks with physician-written rubrics; results retain evaluator, grader and length-adjustment details.

Partial source review · Updated through the September 28 Sonnet 5.5 card, September 22 OpenAI correction and GPT-6 appendix. All source numbers checked. Evaluators use different graders; the Anthropic Astra reproduction is separate. Grok scores are numeric-source verified but protocol details are incomplete.

HealthBench Hard

OpenAI · 1,000 conversations

The 1,000 HealthBench conversations frontier models failed most, still graded on the original physician rubrics.

Showing top 12 of 20 indexed results. View all results.

Source review · Checked primary model reports, applied September 22 Astra correction, added Sol/Luna and Baichuan-report results. Historical raw results and current adjusted results are explicitly labeled and are not a matched comparison.

HealthBench

OpenAI · 5,000 conversations

Realistic multi-turn health conversations graded on physician-written rubrics for accuracy, completeness, and communication.

Source review · Updated through September 28 Sonnet 5.5 and September 22 OpenAI correction. Fixed Fable/Mythos alias and Opus 5 raw/adjusted display. Raw older results remain labeled; evaluator protocols differ.

Health Optimization Bench

Arcophos · 257 tasks

257 evidence-grounded tasks across eight preventive medicine subjects, graded blind by independent model families.

Source review · Replaced the prior 89-task incretin ranking with the official main ranking for 257 subject-suite tasks, snapshot September 10. All sixteen scores and confidence intervals checked against the source table.

MAST (Medical AI Superintelligence Test)

ARISE AI Research Network · 6-benchmark composite

A composite of clinical benchmarks spanning diagnostic and management reasoning, safety, multimodal imaging, and agentic capability.

Source review · Official source still reports August 15, 2026; all eight displayed general-ranking scores match. This is a fresh source check of the existing snapshot, not a new model evaluation.

MedHELM

Stanford CRFM · 121 tasks

Holistic clinical evaluation across 121 tasks in a clinician-validated taxonomy, ranked by mean win rate.

Source review · Official v5.0.0 home-page scores match all ten indexed models and still state May 14, 2026. No later result is implied by the September source check.

First, Do NOHARM (v2)

Stanford/Harvard consortium · 1,100 consultation cases

How often, and how severely, model consultation recommendations contain potentially harmful errors.

Showing top 12 of 17 indexed results. View all results.

Source review · Checked the official v2 technical board and added five omitted published systems/models. Seventeen of nineteen rows are indexed, with RAG systems labeled separately in settings. Board remains in preview. Newly indexed scores show a September observation date with measurement date left unknown.

HealthAgentBench

Microsoft Research · 54 agentic tasks

Agent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.

Source review · All 12 official harness results checked; added the omitted Codex GPT 5.3 row (22%). Current leader remains Claude Code (Opus 5), 55%. Source does not publish a new run date.

CHI-Bench

actAVA · 75 operations workflows

Long-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.

  1. 54.7%
  2. 37.3%
  3. 37.3%
    Anthropic logo
    claude-codeclaude-opus-5
  4. 33.3%
    Anthropic logo
    claude-codeclaude-opus-4-8
  5. 28.0%
    Anthropic logo
    claude-codeclaude-opus-4-6
  6. 26.2%
  7. 25.3%
  8. 25.3%
    Moonshot AI logo
    openai-agentskimi-k3
  9. 24.4%
    Anthropic logo
    claude-codeclaude-opus-4-7
  10. 24.0%
    Anthropic logo
    claude-codeclaude-fable-5
  11. 22.7%
    M
    hermesMedGuard
  12. 20.9%
    OpenAI logo
    codexgpt-5.5

Showing top 12 of 44 indexed results. View all results.

Source review · All 45 submissions inspected; retained all 44 numeric all-domain results and excluded one PA-only result. No newer scores found on the official board.

MedCode (Vals AI)

Vals AI · 2,755 patient records

ICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.

Source review · All 102 scored configurations checked against the first-party embedded table; source parameters captured per row. Updated 2026-09-26; individual run dates unavailable.

MedScribe (Vals AI)

Vals AI · 100 SOAP-note cases

Quality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.

Showing top 12 of 104 indexed results. View all results.

Source review · All 104 scored configurations checked against the first-party embedded table; source parameters captured per row. Updated 2026-09-26; individual run dates unavailable.

MedXpertQA (MM)

Tsinghua University · 2,000 multimodal questions

Expert-level multiple-choice questions over clinical images across 17 specialties.

Partial source review · Meta, Qwen3.5 and Google score tables verified; all published Gemma 4 and Qwen3.5 comparison columns included. Six Qwen3.7/3.8 blog values remain historical and partial because primary pages are empty. No cross-vendor protocol equivalence implied.

Artificial Analysis Healthcare & Medical Index

Artificial Analysis · 6-evaluation composite

A healthcare-weighted composite of six evaluations, covering medical knowledge, long records, knowledge work, reasoning and tools.

Partial source review · Verified the current six-evaluation methodology and 25 numeric model rows from first-party embedded chart data. This is partial coverage of 77 available variants; individual evaluation dates are undisclosed. Older incompatible index rows were removed.

PhysicianBench

academic team · 100 clinical tasks

Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.

Separate evaluation setups. Compare results within each group; the graders and harnesses differ.

Original paper · benchmark authors’ protocol

  1. 46.3 ± 1.2
  2. 31.7 ± 2.3
  3. 29.3 ± 2.5
  4. 27.7 ± 1.5
  5. 23.0 ± 2.6
  6. 18.7 ± 2.9
  7. 17.0 ± 2.6
  8. 16.7 ± 4.0
  9. 13.7 ± 4.0
  10. 8.7 ± 1.2
  11. 6.0 ± 1.0
  12. 5.3 ± 3.2

Source review · Verified all 12 original-paper rows plus 9 vendor-reported model/effort rows from today’s Sonnet 5.5 system card, pp. 137–139. Same public task set; vendor grader/harness cohort explicitly distinguished.

EHR-Complex

academic team · 3,915-task test set

Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.

Source review · Tables 3 and 10 verified: 18 configurations, including both benchmark-specific SFT variants. Current arXiv version remains v1 dated 2026-06-22.

WHBench

academic team · 47 scenarios

Women's health scenarios graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence.

Showing top 12 of 22 indexed results. View all results.

Source review · All 22 model scores in Table 3 checked against v2, revised 2026-07-23. Historical March 2026 evaluation set retained; no live refresh claimed.

HealthAdminBench

Kinetic Systems · 135 admin tasks

Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.

Source review · All seven Figure 3(a) task-success rates verified against v1 and official project page. The paper remains dated 2026-04-10; source checks are not new evaluations.

Scores use each source’s published scale and configuration. A model’s presence here does not establish clinical readiness. Review the source and evaluation setup before drawing conclusions.

How to read this index

A clinical simulation, an exam, and a safety evaluation answer different questions. Their scores are not interchangeable. Benchmark pages explain the task, scale, and result provenance; model pages gather reported results across boards. Dates belong to the underlying sources, and a snapshot review does not mean every score was newly measured. Read our methodology for the inclusion and sourcing rules.

Browse by what is measured

Follow a model across boards

113 of the 204 indexed models appear on more than one benchmark. Browse the model index or explore results by lab:

Outside this index

This is a curated collection. The following benchmarks are not actively tracked in this snapshot.

HealthBench Consensus
near-saturated physician-consensus baseline; frontier runs stopped reporting it separately
MedQA / MultiMedQA
exam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers
AgentClinic
no public frontier-model results since 2025
CRAFT-MD
no public frontier-model results since 2025
MedAgentBench
v2 lives on inside the MAST composite; the standalone board has no current frontier rows
SDBench / MAI-DxO
Microsoft's 2025 sequential-diagnosis study was not re-run on current models
Open Medical-LLM Leaderboard (Hugging Face)
built on saturated exam sets; no frontier submissions in 2026
MedArena
clinician preference arena; ratings pool too thin on current frontier models to quote
AMIE evaluations
Google DeepMind research prototypes, never opened to cross-vendor comparison
LiveClin, PrIME-LLM, MedMCP-Calc
single studies with two or fewer current-frontier rows; tracked for a future qualifying update

Open data. Traceable sources.

Download this snapshot as JSON or CSV, licensed under CC BY 4.0. Find citation guidance on the data page, and the underlying documents on the sources page.