# Health Evals > An independent MedPIC-Bench explainer: 28 published model results on conditional medication-safety reasoning, plus dataset coverage analysis. Reviewed 2026-09-28; the underlying study was published 2026-08-04. Models were not rerun. ## MedPIC-Bench - [Results and analysis](https://healthevals.com/): guideline-following, counterfactual, activation, deactivation and pair accuracy - [Benchmark and coverage](https://healthevals.com/benchmark): task definitions, denominators and six dataset dimensions - [Methodology and sources](https://healthevals.com/medpic-methodology): source verification and reproducibility limits - [MedPIC JSON](https://healthevals.com/data/medpic.json): exact source rows, dataset counts and provenance - [MedPIC CSV](https://healthevals.com/data/medpic.csv): 28 published model rows with sources and configuration - [Original paper](https://arxiv.org/html/2608.03028v1): arXiv:2608.03028v1, Table 2 - [Original dataset](https://huggingface.co/datasets/TIM0927/MedPIC-Bench): 467 questions, CC BY 4.0 ## Index - [The broader index](https://healthevals.com/benchmarks): all 17 benchmarks with every indexed result and its source - [Full data, JSON](https://healthevals.com/data/benchmarks.json): the index with metadata and sources, CC BY 4.0 - [Full data, CSV](https://healthevals.com/data/benchmarks.csv): one flat row per result - [Plain-text index](https://healthevals.com/llms-full.txt): everything in a single document ## Benchmarks - [HealthBench Professional](https://healthevals.com/benchmarks/healthbench-professional): GPT-6 Astra (Anthropic run) leads at 0.703, via Published model evaluation reports - [HealthBench Hard](https://healthevals.com/benchmarks/healthbench-hard): Baichuan-M3 leads at 0.444, via healthbenchhard.ai - [HealthBench](https://healthevals.com/benchmarks/healthbench): Claude Sonnet 5.5 leads at 65.4, via Published model evaluation reports - [Health Optimization Bench](https://healthevals.com/benchmarks/health-optimization-bench): Claude Fable 5 leads at 70.9, via healthoptimizationbench.com - [MAST (Medical AI Superintelligence Test)](https://healthevals.com/benchmarks/mast): GPT-5.6 Sol leads at 60.2%, via MAST: Medical AI Superintelligence Test leaderboard (General board) - [MedHELM](https://healthevals.com/benchmarks/medhelm): Gemini 3.1 Pro (Preview) leads at 0.652, via MedHELM leaderboard (medhelm.org), v5.0.0 - [First, Do NOHARM (v2)](https://healthevals.com/benchmarks/first-do-noharm): LiSA 2.5 leads at 86.2, via MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) - [HealthAgentBench](https://healthevals.com/benchmarks/healthagentbench): Claude Code (Opus 5) leads at 55%, via HealthAgentBench leaderboard - [CHI-Bench](https://healthevals.com/benchmarks/chi-bench): erius + claude-opus-5 leads at 54.7%, via CHI-Bench leaderboard (actAVA) - [MedCode (Vals AI)](https://healthevals.com/benchmarks/medcode): Claude Opus 5 leads at 63.57%, via Vals AI MedCode leaderboard - [MedScribe (Vals AI)](https://healthevals.com/benchmarks/medscribe): Claude Opus 5.5 leads at 91.43%, via Vals AI MedScribe leaderboard - [MedXpertQA (MM)](https://healthevals.com/benchmarks/medxpertqa-mm): GPT-5.6 Sol leads at 81.5, via Introducing Muse Spark: Scaling Towards Personal Superintelligence - [Artificial Analysis Healthcare & Medical Index](https://healthevals.com/benchmarks/artificial-analysis-healthcare): Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) leads at 61, via Best AI for Healthcare & Medical: LLM Leaderboard - [PhysicianBench](https://healthevals.com/benchmarks/physicianbench): Claude Opus 5.5 (max) leads at 68.4%, via PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) - [EHR-Complex](https://healthevals.com/benchmarks/ehr-complex): GPT-5.4 (high reasoning) leads at 0.65, via EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) - [WHBench](https://healthevals.com/benchmarks/whbench): Claude Opus 4.6 leads at 72.1%, via WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) - [HealthAdminBench](https://healthevals.com/benchmarks/healthadminbench): Claude Opus 4.6 (computer-use agent) leads at 36.3%, via HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) ## Categories - [Rubric-graded benchmarks](https://healthevals.com/categories/rubric) - [Agentic and workflow benchmarks](https://healthevals.com/categories/agentic) - [Documentation and coding benchmarks](https://healthevals.com/categories/documentation) - [Safety benchmarks](https://healthevals.com/categories/safety) - [Knowledge and exam benchmarks](https://healthevals.com/categories/knowledge) - [Composite indices](https://healthevals.com/categories/composite) ## Background - [Methodology](https://healthevals.com/methodology): sourcing rules, comparability limits, refresh cadence - [Sources](https://healthevals.com/sources): the document, quote and locator behind every score - [Models](https://healthevals.com/models): one page per frontier model across all boards - [FAQ](https://healthevals.com/faq): the landscape questions, answered from the data ## Optional - [Updates](https://healthevals.com/updates): dated index history - [HealthBench Professional leaderboard](https://healthbenchprofessional.com/): dedicated full board - [HealthBench Hard leaderboard](https://healthbenchhard.ai/): dedicated full board - [Health Optimization Bench](https://healthoptimizationbench.com/): dedicated full board