Health Evals

HealthAgentBench: published results

Microsoft Research · 54 tasks across 7 environments; 162 trials (3 attempts per task) · index updated September 28, 2026

Claude Code (Opus 5) has the highest indexed numerical score on HealthAgentBench, 55% as of 2026-07, per HealthAgentBench leaderboard. Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.

Published results

#modelscoreas of
1Anthropic logoClaude Code (Opus 5) Anthropic
$3.3/task; harness+model evaluated jointly
55%2026-07
2OpenAI logoCodex (GPT-5.6-sol) OpenAI
$5.2/task
45%2026-07
3OpenAI logoCodex (GPT 5.5) OpenAI
$2.8/task
42%2026-07
4Microsoft/Anthropic logoCopilot (Opus 4.8) Microsoft/Anthropic
$3.1/task
36%2026-07
5Microsoft/OpenAI logoCopilot (GPT 5.5) Microsoft/OpenAI
$2.6/task
35%2026-07
6Anthropic logoClaude Code (Opus 4.8) Anthropic
$4.0/task
32%2026-07
7OpenAI logoCodex (GPT 5.4) OpenAI
$1.3/task
28%2026-07
8Anthropic logoClaude Code (Opus 4.7) Anthropic
$4.8/task
27%2026-07
9OpenAI logoCodex (GPT 5.3) OpenAI
$1.0/task; harness and model evaluated jointly
22%2026-07
10Anthropic logoClaude Code (Opus 4.6) Anthropic
$4.1/task
19%2026-07
11Anthropic logoClaude Code (Sonnet 4.6) Anthropic
$2.9/task
17%2026-07
12OpenAI logoCodex (GPT 5.4 Mini) OpenAI
$0.6/task
16%2026-07

Scores preserve their source precision, with any scale conversion documented (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherMicrosoft Research
categoryagentic and workflow benchmarks
released2026-07
size54 tasks across 7 environments; 162 trials (3 attempts per task)
scalemean task success rate, 0-100%, higher better; cost per task also reported
result basisindependently run
sourceHealthAgentBench leaderboard
paperarxiv.org/abs/2606.31179
last frontier result2026-07

What is HealthAgentBench?

HealthAgentBench is a agentic and workflow benchmark from Microsoft Research, released 2026-07: 54 tasks across 7 environments; 162 trials (3 attempts per task), scored on a mean task success rate scale. Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.

Which model leads HealthAgentBench?

Claude Code (Opus 5) (Anthropic) has the highest indexed numerical score on HealthAgentBench at 55% (evaluation setups may differ), per HealthAgentBench leaderboard, as of 2026-07.

Where do the HealthAgentBench numbers come from?

From HealthAgentBench leaderboard (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.