Health Evals

FAQ

What people ask about the healthcare benchmark landscape, answered from the index itself.

MedPIC-Bench tests whether models change medication-safety decisions when relevant patient information changes. Health Evals reproduces the authors' 28 published model rows, analyzes the 467-question release, and distinguishes published measurements from its September 2026 source review. We have not rerun the models.

Which evaluations are in the broader benchmark index?

17 as of September 28, 2026: HealthBench Professional, HealthBench Hard, HealthBench, Health Optimization Bench, MAST (Medical AI Superintelligence Test), MedHELM, First, Do NOHARM (v2), HealthAgentBench, CHI-Bench, MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM), Artificial Analysis Healthcare & Medical Index, PhysicianBench, EHR-Complex, WHBench, HealthAdminBench. Each has a page with published scores and citations. Some sources are historical papers; consult each board’s source review and result dates.

Which AI model is best for healthcare overall?

No single number answers that, because the benchmarks measure different work on incompatible scales. What the index can say: Claude Opus 5.5 has the highest indexed score on MedScribe (Vals AI) and Artificial Analysis Healthcare & Medical Index and PhysicianBench; GPT-5.6 Sol has the highest indexed score on MAST (Medical AI Superintelligence Test) and MedXpertQA (MM); Claude Opus 4.6 has the highest indexed score on WHBench and HealthAdminBench; Claude Opus 5 has the highest indexed score on MedCode (Vals AI); GPT-5.4 has the highest indexed score on EHR-Complex; Gemini 3.1 Pro has the highest indexed score on MedHELM; Claude Fable 5 has the highest indexed score on Health Optimization Bench; GPT-6 Astra (Anthropic run) has the highest indexed score on HealthBench Professional; Claude Sonnet 5.5 has the highest indexed score on HealthBench; Baichuan-M3 has the highest indexed score on HealthBench Hard; LiSA 2.5 has the highest indexed score on First, Do NOHARM (v2); Claude Code (Opus 5) has the highest indexed score on HealthAgentBench; erius + claude-opus-5 has the highest indexed score on CHI-Bench. The models page shows every model's full footprint.

Why not combine everything into one healthcare ranking?

Because the underlying measurements do not add: a rubric score, a hard-subset score, and an exam accuracy are different quantities, produced by different graders. Averaging them would manufacture precision that does not exist, so this site keeps one table per benchmark and one page per model instead.

Where do the scores come from?

Every row names its source: the sister leaderboard sites this index's team also runs, benchmark publishers and vendor system cards, and third-party trackers. Vendor-reported numbers are labeled as such, and scale conversions or unresolved verification gaps are documented.

How current is the index?

The latest source review attempt is dated September 28, 2026, and each result row carries its own date as well, since sources refresh on their own schedules. This is a curated snapshot, not a live feed; an older result remains dated to its publication. Changes land as dated entries on the updates page.

Is the data downloadable?

Yes, the full index ships as JSON and CSV under CC BY 4.0 from fixed paths, with citation guidance on the data page. Individual scores should be cited to the original source each row names.

The index holds the current tables, the models page the per-model views, and the methodology page the sourcing rules.