Health Evals

Agentic and workflow benchmarks

5 tracked · snapshot reviewed September 28, 2026

These evaluations ask models to complete multi-step tasks in electronic health records, terminals, or administrative tools. Results depend on the model, its agent harness, available tools, and task setup. Configuration notes distinguish the systems and workflows being evaluated.

HealthAgentBench

Microsoft Research · 54 agentic tasks

Agent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.

CHI-Bench

actAVA · 75 operations workflows

Long-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.

  1. 54.7%
  2. 37.3%
  3. 37.3%
    Anthropic logo
    claude-codeclaude-opus-5
  4. 33.3%
    Anthropic logo
    claude-codeclaude-opus-4-8
  5. 28.0%
    Anthropic logo
    claude-codeclaude-opus-4-6
  6. 26.2%
  7. 25.3%
  8. 25.3%
    Moonshot AI logo
    openai-agentskimi-k3
  9. 24.4%
    Anthropic logo
    claude-codeclaude-opus-4-7
  10. 24.0%
    Anthropic logo
    claude-codeclaude-fable-5
  11. 22.7%
    M
    hermesMedGuard
  12. 20.9%
    OpenAI logo
    codexgpt-5.5

Showing top 12 of 44 indexed results. View all results.

PhysicianBench

academic team · 100 clinical tasks

Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.

Separate evaluation setups. Compare results within each group; the graders and harnesses differ.

Original paper · benchmark authors’ protocol

  1. 46.3 ± 1.2
  2. 31.7 ± 2.3
  3. 29.3 ± 2.5
  4. 27.7 ± 1.5
  5. 23.0 ± 2.6
  6. 18.7 ± 2.9
  7. 17.0 ± 2.6
  8. 16.7 ± 4.0
  9. 13.7 ± 4.0
  10. 8.7 ± 1.2
  11. 6.0 ± 1.0
  12. 5.3 ± 3.2

EHR-Complex

academic team · 3,915-task test set

Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.

HealthAdminBench

Kinetic Systems · 135 admin tasks

Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.

Which agentic and workflow benchmarks have results in this index?

5 as of September 28, 2026: HealthAgentBench (Claude Code (Opus 5): highest indexed score 55%); CHI-Bench (erius + claude-opus-5: highest indexed score 54.7%); PhysicianBench (Claude Opus 5.5 (max): highest indexed score 68.4%); EHR-Complex (GPT-5.4 (high reasoning): highest indexed score 0.65); HealthAdminBench (Claude Opus 4.6 (computer-use agent): highest indexed score 36.3%).

The other categories sit on the index: rubric-graded benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.