Index methodology
For the featured benchmark, read the MedPIC-Bench methodology and sources.
The broader benchmark index gathers the public results from selected healthcare benchmarks, and its discipline is provenance: each score retains its published precision; HealthBench Professional and Hard percentage results are converted to the boards’ 0–1 scales with the conversion documented, with that source named on the row and cited in full on the sources page. Where the index's maintainers run the underlying leaderboard themselves, the entry says so.
Where the numbers come from
| One record per score | The pages and downloads are generated from one versioned data snapshot. Each result records its benchmark, model, published score, configuration, measurement date, and source citation. Source review dates are separate from measurement dates. A recent review does not mean a model was recently evaluated. |
|---|---|
| Which document counts | First-party documents come first: the vendor's system card, model card or launch post for a number the vendor reports about its own model, and the maintainers' official leaderboard or paper for a board number. An independent evaluator is the primary source for its own evaluation. A secondary source or sister-site compilation is labeled and receives a partial confidence flag when the original number could not be checked. |
| Sister boards | HealthBench Professional, HealthBench Hard, and Health Optimization Bench have dedicated leaderboard sites run by the same maintainers, and their published results are included with the originating citations where available. Where such a row is an in-house run, the record says so; where it is compiled from a published document, the document is cited like any other. |
| Review flags | Every record carries a confidence level. A newly entered number, a number whose document has not been located, or a number that two documents disagree on is flagged for review, shows as source pending on the board, and appears on the sources page with the flag rather than being hidden. A higher-authority document that contradicts an existing record replaces it and the change is dated on the updates page. |
| Inclusion rule | The index covers selected public evaluations of healthcare language models, including rubric scoring, agent workflows, documentation, safety, medical knowledge, and composite indices. Coverage is not exhaustive. Older paper results remain useful and keep their original dates; a missing model is not a zero score. |
What can be compared
Within one benchmark's table, rows share a task set but not always a configuration: when a source mixes grader versions or settings, the config column says so, and rows with different configs should be read loosely. Across benchmarks, nothing is comparable at all. A 0.6 on a rubric benchmark, a 46 percent on a hard subset, and an accuracy figure on an exam-style test are three different measurements that happen to share a page, which is why this site never averages them into a single healthcare number.
Limits
An index inherits the flaws of its sources. Vendor-reported scores favor the vendor's configuration, model-graded rubrics carry grader bias, and some sources update irregularly, so a date on a row is part of the data, not decoration. When a source and this index disagree, the source is authoritative, and a correction is a dated entry on the updates page.
Refresh cadence
This is a curated snapshot, not a live feed. Updates are published after checking source documents. The latest review attempt, covering 17 benchmarks, is from September 28, 2026; review details and corrections are on the updates page.