CHI-Bench: published results
actAVA.ai · 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools · index updated September 28, 2026
erius + claude-opus-5 has the highest indexed numerical score on CHI-Bench, 54.7% as of 2026-07-26, per CHI-Bench leaderboard (actAVA). Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.
Published results
- eriusclaude-opus-5
- eriusclaude-opus-4-8
- claude-codeclaude-opus-5
- claude-codeclaude-opus-4-8
- claude-codeclaude-opus-4-6
- claude-codeclaude-sonnet-4-6
- codexgpt-5.6-sol
- openai-agentskimi-k3
- claude-codeclaude-opus-4-7
- claude-codeclaude-fable-5
- MhermesMedGuard
- codexgpt-5.5
- claude-codeclaude-sonnet-5
- ZAopenai-agentsglm-5.1
- ZAhermesglm-5.1
- ZAopenai-agentsglm-5.2
- openclawclaude-opus-4-7
- ZAopenclawglm-5.1
- Ahermesqwen-3.6-max
- codexgpt-5.4
Showing top 20 of 44 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | erius + claude-opus-5 Humana (harness) / Anthropic (model) All Domains pass@1; PA 72.0%, UM 36.0%, CM 56.0%; submitted 2026-07-26; run date not published | 54.7% | 2026-07-26 | |
| 2 | erius + claude-opus-4-8 Humana (harness) / Anthropic (model) All Domains pass@1; PA 40.0%, UM 16.0%, CM 56.0%; submitted 2026-06-05; run date not published | 37.3% | 2026-06-05 | |
| 3 | claude-code + claude-opus-5 Anthropic All Domains pass@1; PA 20.0%, UM 32.0%, CM 60.0%; submitted 2026-07-24; run date not published | 37.3% | 2026-07-24 | |
| 4 | claude-code + claude-opus-4-8 Anthropic All Domains pass@1; PA 32.0%, UM 28.0%, CM 40.0%; submitted 2026-05-28; run date not published | 33.3% | 2026-05-28 | |
| 5 | claude-code + claude-opus-4-6 Anthropic All Domains pass@1; PA 20.0%, UM 36.0%, CM 28.0%; submitted 2026-05-01; run date not published | 28.0% | 2026-05-01 | |
| 6 | claude-code + claude-sonnet-4-6 Anthropic All Domains pass@1; PA 24.0%, UM 34.7%, CM 20.0%; submitted 2026-05-01; run date not published | 26.2% | 2026-05-01 | |
| 7 | codex + gpt-5.6-sol OpenAI All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not published | 25.3% | 2026-07-24 | |
| 8 | openai-agents + kimi-k3 Moonshot AI All Domains pass@1; PA 28.0%, UM 32.0%, CM 16.0%; submitted 2026-07-24; run date not published | 25.3% | 2026-07-24 | |
| 9 | claude-code + claude-opus-4-7 Anthropic All Domains pass@1; PA 24.0%, UM 17.3%, CM 32.0%; submitted 2026-05-01; run date not published | 24.4% | 2026-05-01 | |
| 10 | claude-code + claude-fable-5 Anthropic All Domains pass@1; PA 24.0%, UM 24.0%, CM 24.0%; submitted 2026-07-22; run date not published | 24.0% | 2026-07-22 | |
| 11 | M | hermes + MedGuard MedGuard All Domains pass@1; PA 4.0%, UM 4.0%, CM 60.0%; submitted 2026-07-06; run date not published | 22.7% | 2026-07-06 |
| 12 | codex + gpt-5.5 OpenAI All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not published | 20.9% | 2026-05-01 | |
| 13 | claude-code + claude-sonnet-5 Anthropic All Domains pass@1; PA 24.0%, UM 24.0%, CM 12.0%; submitted 2026-07-06; run date not published | 20.0% | 2026-07-06 | |
| 14 | ZA | openai-agents + glm-5.1 Zhipu AI All Domains pass@1; PA 18.7%, UM 33.3%, CM 4.0%; submitted 2026-05-01; run date not published | 18.7% | 2026-05-01 |
| 15 | ZA | hermes + glm-5.1 Zhipu AI All Domains pass@1; PA 10.7%, UM 34.7%, CM 10.7%; submitted 2026-05-01; run date not published | 18.7% | 2026-05-01 |
| 16 | ZA | openai-agents + glm-5.2 Zhipu AI All Domains pass@1; PA 20.0%, UM 32.0%, CM 4.0%; submitted 2026-07-06; run date not published | 18.7% | 2026-07-06 |
| 17 | openclaw + claude-opus-4-7 Anthropic All Domains pass@1; PA 18.7%, UM 13.3%, CM 20.0%; submitted 2026-05-01; run date not published | 17.3% | 2026-05-01 | |
| 18 | ZA | openclaw + glm-5.1 Zhipu AI All Domains pass@1; PA 13.3%, UM 26.7%, CM 10.7%; submitted 2026-05-01; run date not published | 16.9% | 2026-05-01 |
| 19 | A | hermes + qwen-3.6-max Alibaba All Domains pass@1; PA 9.3%, UM 26.7%, CM 13.3%; submitted 2026-05-01; run date not published | 16.4% | 2026-05-01 |
| 20 | codex + gpt-5.4 OpenAI All Domains pass@1; PA 24.0%, UM 17.3%, CM 6.7%; submitted 2026-05-01; run date not published | 16.0% | 2026-05-01 | |
| 21 | A | openai-agents + qwen-3.6-max Alibaba All Domains pass@1; PA 16.0%, UM 26.7%, CM 4.0%; submitted 2026-05-01; run date not published | 15.6% | 2026-05-01 |
| 22 | hermes + kimi-k2.6 Moonshot AI All Domains pass@1; PA 18.7%, UM 21.3%, CM 6.7%; submitted 2026-05-01; run date not published | 15.6% | 2026-05-01 | |
| 23 | openai-agents + kimi-k2.6 Moonshot AI All Domains pass@1; PA 17.3%, UM 25.3%, CM 2.7%; submitted 2026-05-01; run date not published | 15.1% | 2026-05-01 | |
| 24 | D | openai-agents + deepseek-v4-pro DeepSeek All Domains pass@1; PA 10.7%, UM 28.0%, CM 4.0%; submitted 2026-05-01; run date not published | 14.2% | 2026-05-01 |
| 25 | D | hermes + deepseek-v4-pro DeepSeek All Domains pass@1; PA 8.0%, UM 25.3%, CM 8.0%; submitted 2026-05-01; run date not published | 13.8% | 2026-05-01 |
| 26 | codex + gpt-5.6-terra OpenAI All Domains pass@1; PA 12.0%, UM 20.0%, CM 8.0%; submitted 2026-07-24; run date not published | 13.3% | 2026-07-24 | |
| 27 | codex + gpt-5.6-luna OpenAI All Domains pass@1; PA 20.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published | 13.3% | 2026-07-24 | |
| 28 | gemini-cli + gemini-3-flash Google All Domains pass@1; PA 18.7%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published | 12.5% | 2026-05-01 | |
| 29 | D | openclaw + deepseek-v4-pro DeepSeek All Domains pass@1; PA 14.7%, UM 12.0%, CM 6.7%; submitted 2026-05-01; run date not published | 11.1% | 2026-05-01 |
| 30 | ZA | deepagents + glm-5.1 Zhipu AI All Domains pass@1; PA 17.3%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published | 11.1% | 2026-05-01 |
| 31 | D | deepagents + deepseek-v4-pro DeepSeek All Domains pass@1; PA 14.7%, UM 10.7%, CM 6.7%; submitted 2026-05-01; run date not published | 10.7% | 2026-05-01 |
| 32 | openclaw + kimi-k2.6 Moonshot AI All Domains pass@1; PA 12.0%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published | 10.2% | 2026-05-01 | |
| 33 | A | deepagents + qwen-3.6-max Alibaba All Domains pass@1; PA 12.0%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published | 9.3% | 2026-05-01 |
| 34 | codex + gpt-5.4-mini OpenAI All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not published | 8.4% | 2026-05-01 | |
| 35 | TM | openai-agents + TML Inkling 256K Thinking Machines All Domains pass@1; PA 4.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published | 8.0% | 2026-07-24 |
| 36 | gemini-cli + gemini-3.1-pro Google All Domains pass@1; PA 14.7%, UM 6.7%, CM 0.0%; submitted 2026-05-01; run date not published | 7.1% | 2026-05-01 | |
| 37 | claude-code + claude-haiku-4-5 Anthropic All Domains pass@1; PA 0.0%, UM 14.7%, CM 4.0%; submitted 2026-05-01; run date not published | 6.2% | 2026-05-01 | |
| 38 | SA | openai-agents + grok-4.3 SpaceX AI All Domains pass@1; PA 0.0%, UM 16.0%, CM 1.3%; submitted 2026-05-01; run date not published | 5.8% | 2026-05-01 |
| 39 | A | openclaw + qwen-3.6-max Alibaba All Domains pass@1; PA 10.7%, UM 4.0%, CM 0.0%; submitted 2026-05-01; run date not published | 4.9% | 2026-05-01 |
| 40 | SA | hermes + grok-4.3 SpaceX AI All Domains pass@1; PA 0.0%, UM 13.3%, CM 0.0%; submitted 2026-05-01; run date not published | 4.4% | 2026-05-01 |
| 41 | deepagents + kimi-k2.6 Moonshot AI All Domains pass@1; PA 8.0%, UM 1.3%, CM 0.0%; submitted 2026-05-01; run date not published | 3.1% | 2026-05-01 | |
| 42 | SA | deepagents + grok-4.3 SpaceX AI All Domains pass@1; PA 0.0%, UM 5.3%, CM 1.3%; submitted 2026-05-01; run date not published | 2.2% | 2026-05-01 |
| 43 | SA | openclaw + grok-4.3 SpaceX AI All Domains pass@1; PA 1.3%, UM 0.0%, CM 0.0%; submitted 2026-05-01; run date not published | 0.4% | 2026-05-01 |
| 44 | openai-agents + Nemotron 3 Ultra 256K NVIDIA All Domains pass@1; PA 0.0%, UM 0.0%, CM 0.0%; submitted 2026-07-24; run date not published | 0.0% | 2026-07-24 | |
Scores preserve their source precision, with any scale conversion documented (mixed sources). Official board updated 2026-08-12: 45 submitted harness configurations, 44 with all-domain accuracy. The PA-only MedArise submission has no all-domain score and is excluded here. Row dates are submission dates; run dates are not published. Community submissions and author-run baselines share the automated workspace judge. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | actAVA.ai |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-05 |
| size | 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools |
| scale | pass@1 with binary 0/1 reward, higher better |
| result basis | mixed sources |
| source | CHI-Bench leaderboard (actAVA) |
| last frontier result | 2026-08-12 |
What is CHI-Bench?
CHI-Bench is a agentic and workflow benchmark from actAVA, released 2026-05: 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools, scored on a pass@1 with binary 0/1 reward scale. Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.
Which model leads CHI-Bench?
erius + claude-opus-5 (Humana (harness) / Anthropic (model)) has the highest indexed numerical score on CHI-Bench at 54.7% (evaluation setups may differ), per CHI-Bench leaderboard (actAVA), as of 2026-07-26.
Where do the CHI-Bench numbers come from?
From CHI-Bench leaderboard (actAVA) (mixed sources). Official board updated 2026-08-12: 45 submitted harness configurations, 44 with all-domain accuracy. The PA-only MedArise submission has no all-domain score and is excluded here. Row dates are submission dates; run dates are not published. Community submissions and author-run baselines share the automated workspace judge.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.