HealthAdminBench: published results
Kinetic Systems (with Stanford Hospital domain experts) · 135 tasks / 1,698 rubric-scored subtasks · index updated September 28, 2026
Claude Opus 4.6 (computer-use agent) has the highest indexed numerical score on HealthAdminBench, 36.3% as of 2026-04, per HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1). End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 (computer-use agent) Anthropic screenshot-only, task description + portal guidance; native CUA harness; subtask rate 78.4% via paper, arxiv.org | 36.3% | 2026-04 | |
| 2 | GPT-5.4 (computer-use agent) OpenAI screenshot-only, task description + portal guidance; subtask rate 82.8% via paper, arxiv.org | 26.7% | 2026-04 | |
| 3 | Kimi K2.5 Moonshot AI screenshot-only, task description + portal guidance via paper, arxiv.org | 15.6% | 2026-04 | |
| 4 | Claude Opus 4.6 (standardized harness) Anthropic screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA via paper, arxiv.org | 14.8% | 2026-04 | |
| 5 | A | Qwen 3.5 Alibaba screenshot-only, task description + portal guidance via paper, arxiv.org | 13.3% | 2026-04 |
| 6 | Gemini 3.1 Pro Google screenshot-only, task description + portal guidance via paper, arxiv.org | 11.9% | 2026-04 | |
| 7 | GPT-5.4 (standardized harness) OpenAI screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA via paper, arxiv.org | 5.9% | 2026-04 | |
Scores preserve their source precision, with any scale conversion documented (independently run). Seven agent configurations are compared using screenshots plus task description and portal guidance. Native computer-use agents and the standardized harness have separate rows. The paper also reports accessibility-tree settings, which must not be compared directly with these screenshot-only scores. No revised paper or newer scored configuration found on the official site. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Kinetic Systems (with Stanford Hospital domain experts) |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-04 |
| size | 135 tasks / 1,698 rubric-scored subtasks |
| scale | percentage end-to-end task success 0-100, higher better |
| result basis | independently run |
| source | HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) |
| official page | healthadminbench.stanford.edu |
| paper | arxiv.org/abs/2604.09937 |
| last frontier result | 2026-04 |
What is HealthAdminBench?
HealthAdminBench is a agentic and workflow benchmark from Kinetic Systems, released 2026-04: 135 tasks / 1,698 rubric-scored subtasks, scored on a percentage end-to-end task success 0-100 scale. End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.
Which model leads HealthAdminBench?
Claude Opus 4.6 (computer-use agent) (Anthropic) has the highest indexed numerical score on HealthAdminBench at 36.3% (evaluation setups may differ), per HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1), as of 2026-04.
Where do the HealthAdminBench numbers come from?
From HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) (independently run). Seven agent configurations are compared using screenshots plus task description and portal guidance. Native computer-use agents and the standardized harness have separate rows. The paper also reports accessibility-tree settings, which must not be compared directly with these screenshot-only scores. No revised paper or newer scored configuration found on the official site.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.