Health Evals

The benchmark

Inside MedPIC-Bench

Medication-safety rules depend on the patient. This benchmark asks whether a model applies those conditions when selecting an answer, including when a change should remove a warning.

One question, one exact answer set

Input
An English patient vignette, a medication-safety question, and answer options.
Output
A brief rationale followed by one or more selected option letters.
Scoring
Exact set match. All correct options must be selected, with no additional incorrect options.
Setting
Zero-shot. Each question is answered separately, with no indication that another question is its paired counterpart.

Keep the denominators separate

Guideline-following · GF284 questions
Choose the medication-safety answer for a fixed patient case.
Counterfactual · CF183 questions
Answer cases where a controlled change in patient information changes which rule applies.
Pair accuracy89 linked pairs, paper-reported
Both cases in a linked comparison must be correct. Each was presented to the model separately.
Risk activation71 questions
Recognize when the patient information makes a medication warning applicable.
Risk deactivation68 questions
Withdraw a warning when the patient information removes its triggering condition.

The counterfactual set also contains 26 risk-redistribution and 18 interaction-identification questions. Activation and deactivation therefore do not add up to the full counterfactual set. Pair accuracy is a different unit: 89 linked comparisons, with both answers required. We preserve the authors’ pair score because the public data does not include explicit pair links.

What is in the release?

Explore the questions behind the scores. These counts come from the public dataset; they describe coverage, not model accuracy.

  • older adults
    35275.4%
  • pediatrics
    6614.1%
  • pregnancy
    4910.5%

Percentages use the selected question set as the denominator. View the original dataset ↗

Department results: a group-level view

The paper reports mean counterfactual accuracy for model groups in five departments with at least 20 counterfactual questions. Group composition mixes model sizes and training approaches. These are descriptive comparisons, not an isolated test of medical specialization.

Mean counterfactual accuracy by department and model group, MedPIC-Bench paper Table 3.
DepartmentCF questionsMedical-specificGeneralProprietary
Neurology4626.4%31.6%40.7%
Cardiology3527.9%33.0%42.9%
Pediatrics2642.9%47.4%54.9%
Nephrology4441.1%53.0%67.2%
Obstetrics & Gynecology2080.0%81.1%96.4%

Scores: paper Table 3. Question counts: our count of the released JSON. These scores are not individual-model results.

What the design makes visible

Controlled changes help reveal whether patient information changes the model’s decision. Separate activation and deactivation scores distinguish recognizing a warning from removing it. Exact answer-set scoring avoids an LLM judge for the selected answers.

What remains outside the test

The model receives a prepared vignette and answer options. The evaluation does not observe clinical history-taking, real prescribing, ongoing care or patient outcomes. Selected rule sources and uneven category sizes limit broader conclusions.

Read the provenance and reproducibility limits ↗