Agentic and workflow benchmarks
5 tracked · snapshot reviewed September 28, 2026
These evaluations ask models to complete multi-step tasks in electronic health records, terminals, or administrative tools. Results depend on the model, its agent harness, available tools, and task setup. Configuration notes distinguish the systems and workflows being evaluated.
HealthAgentBench
Microsoft Research · 54 agentic tasksAgent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.
CHI-Bench
actAVA · 75 operations workflowsLong-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.
- eriusclaude-opus-5
- eriusclaude-opus-4-8
- claude-codeclaude-opus-5
- claude-codeclaude-opus-4-8
- claude-codeclaude-opus-4-6
- claude-codeclaude-sonnet-4-6
- codexgpt-5.6-sol
- openai-agentskimi-k3
- claude-codeclaude-opus-4-7
- claude-codeclaude-fable-5
- MhermesMedGuard
- codexgpt-5.5
Showing top 12 of 44 indexed results. View all results.
PhysicianBench
academic team · 100 clinical tasksAgents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.
Separate evaluation setups. Compare results within each group; the graders and harnesses differ.
Anthropic evaluation · Opus 5 grader
Original paper · benchmark authors’ protocol
EHR-Complex
academic team · 3,915-task test setAgentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.
HealthAdminBench
Kinetic Systems · 135 admin tasksComputer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.
Which agentic and workflow benchmarks have results in this index?
5 as of September 28, 2026: HealthAgentBench (Claude Code (Opus 5): highest indexed score 55%); CHI-Bench (erius + claude-opus-5: highest indexed score 54.7%); PhysicianBench (Claude Opus 5.5 (max): highest indexed score 68.4%); EHR-Complex (GPT-5.4 (high reasoning): highest indexed score 0.65); HealthAdminBench (Claude Opus 4.6 (computer-use agent): highest indexed score 36.3%).
The other categories sit on the index: rubric-graded benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.