HealthAgentBench: published results
Microsoft Research · 54 tasks across 7 environments; 162 trials (3 attempts per task) · index updated September 28, 2026
Claude Code (Opus 5) has the highest indexed numerical score on HealthAgentBench, 55% as of 2026-07, per HealthAgentBench leaderboard. Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Code (Opus 5) Anthropic $3.3/task; harness+model evaluated jointly | 55% | 2026-07 | |
| 2 | Codex (GPT-5.6-sol) OpenAI $5.2/task | 45% | 2026-07 | |
| 3 | Codex (GPT 5.5) OpenAI $2.8/task | 42% | 2026-07 | |
| 4 | Copilot (Opus 4.8) Microsoft/Anthropic $3.1/task | 36% | 2026-07 | |
| 5 | Copilot (GPT 5.5) Microsoft/OpenAI $2.6/task | 35% | 2026-07 | |
| 6 | Claude Code (Opus 4.8) Anthropic $4.0/task | 32% | 2026-07 | |
| 7 | Codex (GPT 5.4) OpenAI $1.3/task | 28% | 2026-07 | |
| 8 | Claude Code (Opus 4.7) Anthropic $4.8/task | 27% | 2026-07 | |
| 9 | Codex (GPT 5.3) OpenAI $1.0/task; harness and model evaluated jointly | 22% | 2026-07 | |
| 10 | Claude Code (Opus 4.6) Anthropic $4.1/task | 19% | 2026-07 | |
| 11 | Claude Code (Sonnet 4.6) Anthropic $2.9/task | 17% | 2026-07 | |
| 12 | Codex (GPT 5.4 Mini) OpenAI $0.6/task | 16% | 2026-07 | |
Scores preserve their source precision, with any scale conversion documented (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Microsoft Research |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-07 |
| size | 54 tasks across 7 environments; 162 trials (3 attempts per task) |
| scale | mean task success rate, 0-100%, higher better; cost per task also reported |
| result basis | independently run |
| source | HealthAgentBench leaderboard |
| paper | arxiv.org/abs/2606.31179 |
| last frontier result | 2026-07 |
What is HealthAgentBench?
HealthAgentBench is a agentic and workflow benchmark from Microsoft Research, released 2026-07: 54 tasks across 7 environments; 162 trials (3 attempts per task), scored on a mean task success rate scale. Agentic task success in realistic terminal-based healthcare environments built from real clinical artifacts; evaluates agent harnesses (Claude Code, Codex, Copilot) end to end, not bare models.
Which model leads HealthAgentBench?
Claude Code (Opus 5) (Anthropic) has the highest indexed numerical score on HealthAgentBench at 55% (evaluation setups may differ), per HealthAgentBench leaderboard, as of 2026-07.
Where do the HealthAgentBench numbers come from?
From HealthAgentBench leaderboard (independently run). Paper: arXiv 2606.31179. The rows are agent harnesses rather than bare models, and the board is run by Microsoft Research; Copilot, Microsoft's own harness, does not top it.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.