A public index of healthcare AI evaluations
Healthcare AI benchmarks
Explore reported model results across clinical reasoning, medical knowledge, safety, and real-world workflows. See what each benchmark measures, compare results within a board, and follow the evidence behind a score.
The benchmarks
Compare within a board · scales differ
17 of 17 benchmarks shown
HealthBench Professional
OpenAI · 525 tasks525 clinician tasks with physician-written rubrics; results retain evaluator, grader and length-adjustment details.
Partial source review · Updated through the September 28 Sonnet 5.5 card, September 22 OpenAI correction and GPT-6 appendix. All source numbers checked. Evaluators use different graders; the Anthropic Astra reproduction is separate. Grok scores are numeric-source verified but protocol details are incomplete.
HealthBench Hard
OpenAI · 1,000 conversationsThe 1,000 HealthBench conversations frontier models failed most, still graded on the original physician rubrics.
Showing top 12 of 20 indexed results. View all results.
Source review · Checked primary model reports, applied September 22 Astra correction, added Sol/Luna and Baichuan-report results. Historical raw results and current adjusted results are explicitly labeled and are not a matched comparison.
HealthBench
OpenAI · 5,000 conversationsRealistic multi-turn health conversations graded on physician-written rubrics for accuracy, completeness, and communication.
Source review · Updated through September 28 Sonnet 5.5 and September 22 OpenAI correction. Fixed Fable/Mythos alias and Opus 5 raw/adjusted display. Raw older results remain labeled; evaluator protocols differ.
Health Optimization Bench
Arcophos · 257 tasks257 evidence-grounded tasks across eight preventive medicine subjects, graded blind by independent model families.
Showing top 12 of 16 indexed results. View all results.
Source review · Replaced the prior 89-task incretin ranking with the official main ranking for 257 subject-suite tasks, snapshot September 10. All sixteen scores and confidence intervals checked against the source table.
MAST (Medical AI Superintelligence Test)
ARISE AI Research Network · 6-benchmark compositeA composite of clinical benchmarks spanning diagnostic and management reasoning, safety, multimodal imaging, and agentic capability.
Source review · Official source still reports August 15, 2026; all eight displayed general-ranking scores match. This is a fresh source check of the existing snapshot, not a new model evaluation.
MedHELM
Stanford CRFM · 121 tasksHolistic clinical evaluation across 121 tasks in a clinician-validated taxonomy, ranked by mean win rate.
Source review · Official v5.0.0 home-page scores match all ten indexed models and still state May 14, 2026. No later result is implied by the September source check.
First, Do NOHARM (v2)
Stanford/Harvard consortium · 1,100 consultation casesHow often, and how severely, model consultation recommendations contain potentially harmful errors.
Showing top 12 of 17 indexed results. View all results.
Source review · Checked the official v2 technical board and added five omitted published systems/models. Seventeen of nineteen rows are indexed, with RAG systems labeled separately in settings. Board remains in preview. Newly indexed scores show a September observation date with measurement date left unknown.
HealthAgentBench
Microsoft Research · 54 agentic tasksAgent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.
Source review · All 12 official harness results checked; added the omitted Codex GPT 5.3 row (22%). Current leader remains Claude Code (Opus 5), 55%. Source does not publish a new run date.
CHI-Bench
actAVA · 75 operations workflowsLong-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.
- eriusclaude-opus-5
- eriusclaude-opus-4-8
- claude-codeclaude-opus-5
- claude-codeclaude-opus-4-8
- claude-codeclaude-opus-4-6
- claude-codeclaude-sonnet-4-6
- codexgpt-5.6-sol
- openai-agentskimi-k3
- claude-codeclaude-opus-4-7
- claude-codeclaude-fable-5
- MhermesMedGuard
- codexgpt-5.5
Showing top 12 of 44 indexed results. View all results.
Source review · All 45 submissions inspected; retained all 44 numeric all-domain results and excluded one PA-only result. No newer scores found on the official board.
MedCode (Vals AI)
Vals AI · 2,755 patient recordsICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.
Source review · All 102 scored configurations checked against the first-party embedded table; source parameters captured per row. Updated 2026-09-26; individual run dates unavailable.
MedScribe (Vals AI)
Vals AI · 100 SOAP-note casesQuality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.
Showing top 12 of 104 indexed results. View all results.
Source review · All 104 scored configurations checked against the first-party embedded table; source parameters captured per row. Updated 2026-09-26; individual run dates unavailable.
MedXpertQA (MM)
Tsinghua University · 2,000 multimodal questionsExpert-level multiple-choice questions over clinical images across 17 specialties.
Partial source review · Meta, Qwen3.5 and Google score tables verified; all published Gemma 4 and Qwen3.5 comparison columns included. Six Qwen3.7/3.8 blog values remain historical and partial because primary pages are empty. No cross-vendor protocol equivalence implied.
Artificial Analysis Healthcare & Medical Index
Artificial Analysis · 6-evaluation compositeA healthcare-weighted composite of six evaluations, covering medical knowledge, long records, knowledge work, reasoning and tools.
Showing top 12 of 25 indexed results. View all results.
Partial source review · Verified the current six-evaluation methodology and 25 numeric model rows from first-party embedded chart data. This is partial coverage of 77 available variants; individual evaluation dates are undisclosed. Older incompatible index rows were removed.
PhysicianBench
academic team · 100 clinical tasksAgents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.
Separate evaluation setups. Compare results within each group; the graders and harnesses differ.
Anthropic evaluation · Opus 5 grader
Original paper · benchmark authors’ protocol
Source review · Verified all 12 original-paper rows plus 9 vendor-reported model/effort rows from today’s Sonnet 5.5 system card, pp. 137–139. Same public task set; vendor grader/harness cohort explicitly distinguished.
EHR-Complex
academic team · 3,915-task test setAgentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.
Source review · Tables 3 and 10 verified: 18 configurations, including both benchmark-specific SFT variants. Current arXiv version remains v1 dated 2026-06-22.
WHBench
academic team · 47 scenariosWomen's health scenarios graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence.
Showing top 12 of 22 indexed results. View all results.
Source review · All 22 model scores in Table 3 checked against v2, revised 2026-07-23. Historical March 2026 evaluation set retained; no live refresh claimed.
HealthAdminBench
Kinetic Systems · 135 admin tasksComputer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.
Source review · All seven Figure 3(a) task-success rates verified against v1 and official project page. The paper remains dated 2026-04-10; source checks are not new evaluations.
Scores use each source’s published scale and configuration. A model’s presence here does not establish clinical readiness. Review the source and evaluation setup before drawing conclusions.
How to read this index
A clinical simulation, an exam, and a safety evaluation answer different questions. Their scores are not interchangeable. Benchmark pages explain the task, scale, and result provenance; model pages gather reported results across boards. Dates belong to the underlying sources, and a snapshot review does not mean every score was newly measured. Read our methodology for the inclusion and sourcing rules.
Browse by what is measured
- Rubric-graded benchmarks5 boards
- Agentic and workflow benchmarks5 boards
- Documentation and coding benchmarks2 boards
- Safety benchmarks1 boards
- Knowledge and exam benchmarks1 boards
- Composite indices3 boards
Follow a model across boards
113 of the 204 indexed models appear on more than one benchmark. Browse the model index or explore results by lab:
Outside this index
This is a curated collection. The following benchmarks are not actively tracked in this snapshot.
- HealthBench Consensus
- near-saturated physician-consensus baseline; frontier runs stopped reporting it separately
- MedQA / MultiMedQA
- exam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers
- AgentClinic
- no public frontier-model results since 2025
- CRAFT-MD
- no public frontier-model results since 2025
- MedAgentBench
- v2 lives on inside the MAST composite; the standalone board has no current frontier rows
- SDBench / MAI-DxO
- Microsoft's 2025 sequential-diagnosis study was not re-run on current models
- Open Medical-LLM Leaderboard (Hugging Face)
- built on saturated exam sets; no frontier submissions in 2026
- MedArena
- clinician preference arena; ratings pool too thin on current frontier models to quote
- AMIE evaluations
- Google DeepMind research prototypes, never opened to cross-vendor comparison
- LiveClin, PrIME-LLM, MedMCP-Calc
- single studies with two or fewer current-frontier rows; tracked for a future qualifying update
Open data. Traceable sources.
Download this snapshot as JSON or CSV, licensed under CC BY 4.0. Find citation guidance on the data page, and the underlying documents on the sources page.