The benchmark
Inside MedPIC-Bench
Medication-safety rules depend on the patient. This benchmark asks whether a model applies those conditions when selecting an answer, including when a change should remove a warning.
One question, one exact answer set
- Input
- An English patient vignette, a medication-safety question, and answer options.
- Output
- A brief rationale followed by one or more selected option letters.
- Scoring
- Exact set match. All correct options must be selected, with no additional incorrect options.
- Setting
- Zero-shot. Each question is answered separately, with no indication that another question is its paired counterpart.
Keep the denominators separate
- Guideline-following · GF284 questions
- Choose the medication-safety answer for a fixed patient case.
- Counterfactual · CF183 questions
- Answer cases where a controlled change in patient information changes which rule applies.
- Pair accuracy89 linked pairs, paper-reported
- Both cases in a linked comparison must be correct. Each was presented to the model separately.
- Risk activation71 questions
- Recognize when the patient information makes a medication warning applicable.
- Risk deactivation68 questions
- Withdraw a warning when the patient information removes its triggering condition.
The counterfactual set also contains 26 risk-redistribution and 18 interaction-identification questions. Activation and deactivation therefore do not add up to the full counterfactual set. Pair accuracy is a different unit: 89 linked comparisons, with both answers required. We preserve the authors’ pair score because the public data does not include explicit pair links.
What is in the release?
Explore the questions behind the scores. These counts come from the public dataset; they describe coverage, not model accuracy.
- older adults35275.4%
- pediatrics6614.1%
- pregnancy4910.5%
Percentages use the selected question set as the denominator. View the original dataset ↗
Department results: a group-level view
The paper reports mean counterfactual accuracy for model groups in five departments with at least 20 counterfactual questions. Group composition mixes model sizes and training approaches. These are descriptive comparisons, not an isolated test of medical specialization.
| Department | CF questions | Medical-specific | General | Proprietary |
|---|---|---|---|---|
| Neurology | 46 | 26.4% | 31.6% | 40.7% |
| Cardiology | 35 | 27.9% | 33.0% | 42.9% |
| Pediatrics | 26 | 42.9% | 47.4% | 54.9% |
| Nephrology | 44 | 41.1% | 53.0% | 67.2% |
| Obstetrics & Gynecology | 20 | 80.0% | 81.1% | 96.4% |
Scores: paper Table 3. Question counts: our count of the released JSON. These scores are not individual-model results.
What the design makes visible
Controlled changes help reveal whether patient information changes the model’s decision. Separate activation and deactivation scores distinguish recognizing a warning from removing it. Exact answer-set scoring avoids an LLM judge for the selected answers.
What remains outside the test
The model receives a prepared vignette and answer options. The evaluation does not observe clinical history-taking, real prescribing, ongoing care or patient outcomes. Selected rule sources and uneven category sizes limit broader conclusions.