Health Evals

PhysicianBench: published results

Academic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component · 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task) · index updated September 28, 2026

Claude Opus 5.5 (max) has the highest indexed numerical score on PhysicianBench, 68.4% as of 2026-09-28, per Claude Sonnet 5.5 System Card. LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.

Published results

Separate evaluation setups. Compare results within each group; the graders and harnesses differ.

Anthropic evaluation · Opus 5 grader

Original paper · benchmark authors’ protocol

  1. 46.3 ± 1.2
  2. 31.7 ± 2.3
  3. 29.3 ± 2.5
  4. 27.7 ± 1.5
  5. 23.0 ± 2.6
  6. 18.7 ± 2.9
  7. 17.0 ± 2.6
  8. 16.7 ± 4.0
  9. 13.7 ± 4.0
  10. 8.7 ± 1.2
  11. 6.0 ± 1.0
  12. 5.3 ± 3.2
sources

Anthropic evaluation · Opus 5 grader

#modelscoreas of
1Anthropic logoClaude Opus 5.5 (max) Anthropic
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
68.4%2026-09-28
2Anthropic logoClaude Sonnet 5.5 (max) Anthropic
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
63.2%2026-09-28
3Anthropic logoClaude Fable 5.1 (max) Anthropic
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
61.0%2026-09-28
4Anthropic logoClaude Opus 5 (max) Anthropic
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
57.6%2026-09-28
5Anthropic logoClaude Sonnet 5.5 (xhigh) Anthropic
Anthropic-run pass@1 on 100 tasks; xhigh effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
56.4%2026-09-28
6Anthropic logoClaude Sonnet 5.5 (high) Anthropic
Anthropic-run pass@1 on 100 tasks; high effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
47.6%2026-09-28
7Anthropic logoClaude Sonnet 5 (max) Anthropic
Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
37.4%2026-09-28
8Anthropic logoClaude Sonnet 5.5 (medium) Anthropic
Anthropic-run pass@1 on 100 tasks; medium effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
30.0%2026-09-28
9Anthropic logoClaude Sonnet 5.5 (low) Anthropic
Anthropic-run pass@1 on 100 tasks; low effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
27.2%2026-09-28

Original paper · benchmark authors’ protocol

#modelscoreas of
1OpenAI logoGPT-5.5 OpenAI
pass@1; Pass^3 28.0
46.3 ± 1.22026-05
2Anthropic logoClaude Opus 4.6 Anthropic
Pass^3 18.0
31.7 ± 2.32026-05
3Anthropic logoClaude Opus 4.7 Anthropic
Pass^3 18.0
29.3 ± 2.52026-05
4OpenAI logoGPT-5.4 OpenAI27.7 ± 1.52026-05
5Anthropic logoClaude Sonnet 4.6 Anthropic23.0 ± 2.62026-05
6DDeepSeek V4-Pro DeepSeek
Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
18.7 ± 2.92026-05
7Moonshot AI logoKimi-K2.6 Moonshot AI
open source
17.0 ± 2.62026-05
8XMiMo-v2.5-Pro Xiaomi
Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
16.7 ± 4.02026-05
9AQwen3.6-Plus Alibaba13.7 ± 4.02026-05
10MiniMax logoMiniMax M2.7 MiniMax
Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
8.7 ± 1.22026-05
11Google logoGemini Pro 3.1 Google6.0 ± 1.02026-05
12xAI logoGrok-4.20 xAI5.3 ± 3.22026-05

Scores preserve their source precision, with any scale conversion documented (mixed sources). Two evaluator cohorts are shown on the same public 100-task set: the May 2026 benchmark paper (12 models, shared FHIR loop, 3 runs), and Anthropic’s September 28 system card (9 model/effort configurations, shared vendor harness, Opus 5 rubric grader). The vendor cohort is not a replication of the paper protocol. Read configuration and reporting labels before comparing cohorts. Opus 5.5 leads Anthropic’s max-effort cohort at 68.4%; GPT-5.5 leads the paper cohort at 46.3%. Differences of 5–6 points are near the resolution of this 100-task benchmark. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherAcademic team (Ruoqi Liu, Imran Q. Mohiuddin et al., arXiv 2605.02240); also a MAST component
categoryagentic and workflow benchmarks
released2026-05
size100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task)
scalepass@1 success rate %, higher better (3 independent runs; Pass^3 also reported)
result basismixed sources
sourcePhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF)
official pagearxiv.org/abs/2605.02240
last frontier result2026-09-28

What is PhysicianBench?

PhysicianBench is a agentic and workflow benchmark from academic team, released 2026-05: 100 real-world clinical tasks, 21 specialties, 670 structured checkpoints (~27 tool calls per task), scored on a pass@1 success rate % scale. LLM agents on long-horizon composite physician workflows inside real EHR environments, with execution-grounded verification against actual EHR systems via standard commercial APIs.

Which model leads PhysicianBench?

Claude Opus 5.5 (max) (Anthropic) has the highest indexed numerical score on PhysicianBench at 68.4% (evaluation setups may differ), per Claude Sonnet 5.5 System Card, as of 2026-09-28.

Where do the PhysicianBench numbers come from?

From PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) (mixed sources). Two evaluator cohorts are shown on the same public 100-task set: the May 2026 benchmark paper (12 models, shared FHIR loop, 3 runs), and Anthropic’s September 28 system card (9 model/effort configurations, shared vendor harness, Opus 5 rubric grader). The vendor cohort is not a replication of the paper protocol. Read configuration and reporting labels before comparing cohorts. Opus 5.5 leads Anthropic’s max-effort cohort at 68.4%; GPT-5.5 leads the paper cohort at 46.3%. Differences of 5–6 points are near the resolution of this 100-task benchmark.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.