Health Evals

HealthAdminBench: published results

Kinetic Systems (with Stanford Hospital domain experts) · 135 tasks / 1,698 rubric-scored subtasks · index updated September 28, 2026

Claude Opus 4.6 (computer-use agent) has the highest indexed numerical score on HealthAdminBench, 36.3% as of 2026-04, per HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1). End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.

Published results

#modelscoreas of
1Anthropic logoClaude Opus 4.6 (computer-use agent) Anthropic
screenshot-only, task description + portal guidance; native CUA harness; subtask rate 78.4%
36.3%2026-04
2OpenAI logoGPT-5.4 (computer-use agent) OpenAI
screenshot-only, task description + portal guidance; subtask rate 82.8%
26.7%2026-04
3Moonshot AI logoKimi K2.5 Moonshot AI
screenshot-only, task description + portal guidance
15.6%2026-04
4Anthropic logoClaude Opus 4.6 (standardized harness) Anthropic
screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
14.8%2026-04
5AQwen 3.5 Alibaba
screenshot-only, task description + portal guidance
13.3%2026-04
6Google logoGemini 3.1 Pro Google
screenshot-only, task description + portal guidance
11.9%2026-04
7OpenAI logoGPT-5.4 (standardized harness) OpenAI
screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
5.9%2026-04

Scores preserve their source precision, with any scale conversion documented (independently run). Seven agent configurations are compared using screenshots plus task description and portal guidance. Native computer-use agents and the standardized harness have separate rows. The paper also reports accessibility-tree settings, which must not be compared directly with these screenshot-only scores. No revised paper or newer scored configuration found on the official site. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherKinetic Systems (with Stanford Hospital domain experts)
categoryagentic and workflow benchmarks
released2026-04
size135 tasks / 1,698 rubric-scored subtasks
scalepercentage end-to-end task success 0-100, higher better
result basisindependently run
sourceHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1)
official pagehealthadminbench.stanford.edu
paperarxiv.org/abs/2604.09937
last frontier result2026-04

What is HealthAdminBench?

HealthAdminBench is a agentic and workflow benchmark from Kinetic Systems, released 2026-04: 135 tasks / 1,698 rubric-scored subtasks, scored on a percentage end-to-end task success 0-100 scale. End-to-end task success of computer-use LLM agents on healthcare administration workflows: prior authorizations, denial appeals, and DME ordering; success requires completing every subtask in a task.

Which model leads HealthAdminBench?

Claude Opus 4.6 (computer-use agent) (Anthropic) has the highest indexed numerical score on HealthAdminBench at 36.3% (evaluation setups may differ), per HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1), as of 2026-04.

Where do the HealthAdminBench numbers come from?

From HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) (independently run). Seven agent configurations are compared using screenshots plus task description and portal guidance. Native computer-use agents and the standardized harness have separate rows. The paper also reports accessibility-tree settings, which must not be compared directly with these screenshot-only scores. No revised paper or newer scored configuration found on the official site.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.