Health Evals

Source review · 28 September 2026

Methodology and sources

Health Evals presents an independent analysis of MedPIC-Bench. The model results belong to the benchmark authors. Arcophos checked their published tables, counted the public dataset, and built the explanatory views.

The two sources behind the page

  1. 01 · The paper

    Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning, Zhitian Hou and colleagues. arXiv:2608.03028v1, published 4 August 2026. Table 2 supplies all 28 model rows; Table 3 supplies the department group means. Section 4.1 describes the scoring and shared protocol.

    Read the exact PDF version ↗
  2. 02 · The released dataset

    TIM0927/MedPIC-Bench on Hugging Face, licensed CC BY 4.0. The release contains 467 questions, answer sets and six annotation dimensions. We counted the pinned JSON directly; dataset coverage is independent of the model-score tables.

    Revision: 9ef6db4f13865b14fc2e6be3f94dcfaf3a0cf983

    Download the pinned original questions ↗

What “source-checked” means

Every published model row was checked against both the paper’s HTML table and PDF text. We retained the printed percentages and model names. Dataset totals and category counts were calculated from the versioned release. We did not run these models or validate their answers independently.

The study publication date is 4 August 2026. The source review date is 28 September 2026. Actual model-run dates are not available in the inspected artifacts. These are historical study results, not a continuously refreshed frontier-model leaderboard.

Protocol and missing information

Each model answers every question independently in a zero-shot setting. A common prompt requests a brief rationale and the selected option letters. A response is correct only if its selected option set matches the answer key exactly. Linked cases are presented separately, without their connection revealed.

The paper reports open-model inference through PyTorch, Transformers and SGLang on two 80 GB A800 GPUs, with proprietary models accessed through APIs. It refers to supplementary model configurations and uncertainty estimates. Those supplemental details were not found in the public PDF, HTML or dataset release we inspected. Temperature, token limits, reasoning effort and exact API model snapshots remain unspecified here.

The released JSON has no explicit pair ID, rule ID, item-level source citation, raw model response or evaluator code. The 89-pair results therefore remain paper-reported values. We cannot reconstruct them from the released questions alone. The same limitation prevents producing new per-model scores for every coverage slice.

Interpretation rules

  • Compare accuracy with its denominator. Overall, guideline-following, counterfactual and pair accuracy are different summaries.
  • Read the GF–CF gap alongside both absolute scores. A small gap can occur when both scores are low.
  • Keep the printed gap. Rounding means it can differ by 0.1 point from subtracting the displayed GF and CF values.
  • Activation and deactivation cover 139 counterfactual questions. The remaining 44 test other reasoning operations.
  • Coverage bars count questions. Drug categories overlap; model-group results mix architectures, sizes and training approaches.
  • Do not turn benchmark errors into clinical-harm estimates. The study does not measure deployment or patient outcomes.

Attribution and reuse

Credit Zhitian Hou, Yuhang Liu, Pengkai Wang and colleagues for MedPIC-Bench, and cite their paper for model measurements. Credit Arcophos / Health Evals for this independent presentation and coverage analysis. The original dataset’s CC BY 4.0 license is recorded separately from the paper’s terms and model trademarks.

The broader healthcare collection uses its own index methodology. It remains available at the benchmark index.