PhysicianBench: published results
Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component · 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task) · index updated September 28, 2026
Claude Opus 5.5 (max) has the highest indexed numerical score on PhysicianBench, 68.4% as of 2026-09-28, per Claude Sonnet 5.5 System Card. LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.
Published results
Separate evaluation setups. Compare results within each group; the graders and harnesses differ.
Anthropic evaluation · Opus 5 grader
Original paper · benchmark authors’ protocol
Result detail
sources for this boardAnthropic evaluation · Opus 5 grader
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 (max) Anthropic Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 68.4% | 2026-09-28 | |
| 2 | Claude Sonnet 5.5 (max) Anthropic Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 63.2% | 2026-09-28 | |
| 3 | Claude Fable 5.1 (max) Anthropic Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 61.0% | 2026-09-28 | |
| 4 | Claude Opus 5 (max) Anthropic Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 57.6% | 2026-09-28 | |
| 5 | Claude Sonnet 5.5 (xhigh) Anthropic Anthropic-run pass@1 on 100 tasks; xhigh effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 56.4% | 2026-09-28 | |
| 6 | Claude Sonnet 5.5 (high) Anthropic Anthropic-run pass@1 on 100 tasks; high effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 47.6% | 2026-09-28 | |
| 7 | Claude Sonnet 5 (max) Anthropic Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 37.4% | 2026-09-28 | |
| 8 | Claude Sonnet 5.5 (medium) Anthropic Anthropic-run pass@1 on 100 tasks; medium effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 30.0% | 2026-09-28 | |
| 9 | Claude Sonnet 5.5 (low) Anthropic Anthropic-run pass@1 on 100 tasks; low effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished | 27.2% | 2026-09-28 | |
Original paper · benchmark authors’ protocol
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.5 OpenAI pass@1; Pass^3 28.0 via paper, arxiv.org | 46.3 ± 1.2 | 2026-05 | |
| 2 | Claude Opus 4.6 Anthropic Pass^3 18.0 via paper, arxiv.org | 31.7 ± 2.3 | 2026-05 | |
| 3 | Claude Opus 4.7 Anthropic Pass^3 18.0 via paper, arxiv.org | 29.3 ± 2.5 | 2026-05 | |
| 4 | GPT-5.4 OpenAI via paper, arxiv.org | 27.7 ± 1.5 | 2026-05 | |
| 5 | Claude Sonnet 4.6 Anthropic via paper, arxiv.org | 23.0 ± 2.6 | 2026-05 | |
| 6 | D | DeepSeek V4-Pro DeepSeek Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported via paper, arxiv.org | 18.7 ± 2.9 | 2026-05 |
| 7 | Kimi-K2.6 Moonshot AI open source via paper, arxiv.org | 17.0 ± 2.6 | 2026-05 | |
| 8 | X | MiMo-v2.5-Pro Xiaomi Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported via paper, arxiv.org | 16.7 ± 4.0 | 2026-05 |
| 9 | A | Qwen3.6-Plus Alibaba via paper, arxiv.org | 13.7 ± 4.0 | 2026-05 |
| 10 | MiniMax M2.7 MiniMax Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported via paper, arxiv.org | 8.7 ± 1.2 | 2026-05 | |
| 11 | Gemini Pro 3.1 Google via paper, arxiv.org | 6.0 ± 1.0 | 2026-05 | |
| 12 | Grok-4.20 xAI via paper, arxiv.org | 5.3 ± 3.2 | 2026-05 | |
Scores preserve their source precision, with any scale conversion documented (mixed sources). Two evaluator cohorts are shown on the same public 100-task set: the May 2026 benchmark paper (12 models, shared FHIR loop, 3 runs), and Anthropic’s September 28 system card (9 model/effort configurations, shared vendor harness, Opus 5 rubric grader). The vendor cohort is not a replication of the paper protocol. Read configuration and reporting labels before comparing cohorts. Opus 5.5 leads Anthropic’s max-effort cohort at 68.4%; GPT-5.5 leads the paper cohort at 46.3%. Differences of 5–6 points are near the resolution of this 100-task benchmark. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-05 |
| size | 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task) |
| scale | pass@1 success rate %, higher better (3 independent runs; Pass^3 also reported) |
| result basis | mixed sources |
| source | PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) |
| official page | arxiv.org/abs/2605.02240 |
| last frontier result | 2026-09-28 |
What is PhysicianBench?
PhysicianBench is a agentic and workflow benchmark from academic team, released 2026-05: 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task), scored on a pass@1 success rate % scale. LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.
Which model leads PhysicianBench?
Claude Opus 5.5 (max) (Anthropic) has the highest indexed numerical score on PhysicianBench at 68.4% (evaluation setups may differ), per Claude Sonnet 5.5 System Card, as of 2026-09-28.
Where do the PhysicianBench numbers come from?
From PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) (mixed sources). Two evaluator cohorts are shown on the same public 100-task set: the May 2026 benchmark paper (12 models, shared FHIR loop, 3 runs), and Anthropic’s September 28 system card (9 model/effort configurations, shared vendor harness, Opus 5 rubric grader). The vendor cohort is not a replication of the paper protocol. Read configuration and reporting labels before comparing cohorts. Opus 5.5 leads Anthropic’s max-effort cohort at 68.4%; GPT-5.5 leads the paper cohort at 46.3%. Differences of 5–6 points are near the resolution of this 100-task benchmark.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.