Health Evals

CHI-Bench: published results

actAVA.ai · 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools · index updated September 28, 2026

erius + claude-opus-5 has the highest indexed numerical score on CHI-Bench, 54.7% as of 2026-07-26, per CHI-Bench leaderboard (actAVA). Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.

Published results

  1. 54.7%
  2. 37.3%
  3. 37.3%
    Anthropic logo
    claude-codeclaude-opus-5
  4. 33.3%
    Anthropic logo
    claude-codeclaude-opus-4-8
  5. 28.0%
    Anthropic logo
    claude-codeclaude-opus-4-6
  6. 26.2%
  7. 25.3%
  8. 25.3%
    Moonshot AI logo
    openai-agentskimi-k3
  9. 24.4%
    Anthropic logo
    claude-codeclaude-opus-4-7
  10. 24.0%
    Anthropic logo
    claude-codeclaude-fable-5
  11. 22.7%
    M
    hermesMedGuard
  12. 20.9%
    OpenAI logo
    codexgpt-5.5
  13. 20.0%
    Anthropic logo
    claude-codeclaude-sonnet-5
  14. 18.7%
    ZA
    openai-agentsglm-5.1
  15. 18.7%
    ZA
    hermesglm-5.1
  16. 18.7%
    ZA
    openai-agentsglm-5.2
  17. 17.3%
  18. 16.9%
    ZA
    openclawglm-5.1
  19. 16.4%
  20. 16.0%
    OpenAI logo
    codexgpt-5.4

Showing top 20 of 44 indexed results. View all results.

sources
#modelscoreas of
1Humana (harness) / Anthropic (model) logoerius + claude-opus-5 Humana (harness) / Anthropic (model)
All Domains pass@1; PA 72.0%, UM 36.0%, CM 56.0%; submitted 2026-07-26; run date not published
54.7%2026-07-26
2Humana (harness) / Anthropic (model) logoerius + claude-opus-4-8 Humana (harness) / Anthropic (model)
All Domains pass@1; PA 40.0%, UM 16.0%, CM 56.0%; submitted 2026-06-05; run date not published
37.3%2026-06-05
3Anthropic logoclaude-code + claude-opus-5 Anthropic
All Domains pass@1; PA 20.0%, UM 32.0%, CM 60.0%; submitted 2026-07-24; run date not published
37.3%2026-07-24
4Anthropic logoclaude-code + claude-opus-4-8 Anthropic
All Domains pass@1; PA 32.0%, UM 28.0%, CM 40.0%; submitted 2026-05-28; run date not published
33.3%2026-05-28
5Anthropic logoclaude-code + claude-opus-4-6 Anthropic
All Domains pass@1; PA 20.0%, UM 36.0%, CM 28.0%; submitted 2026-05-01; run date not published
28.0%2026-05-01
6Anthropic logoclaude-code + claude-sonnet-4-6 Anthropic
All Domains pass@1; PA 24.0%, UM 34.7%, CM 20.0%; submitted 2026-05-01; run date not published
26.2%2026-05-01
7OpenAI logocodex + gpt-5.6-sol OpenAI
All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not published
25.3%2026-07-24
8Moonshot AI logoopenai-agents + kimi-k3 Moonshot AI
All Domains pass@1; PA 28.0%, UM 32.0%, CM 16.0%; submitted 2026-07-24; run date not published
25.3%2026-07-24
9Anthropic logoclaude-code + claude-opus-4-7 Anthropic
All Domains pass@1; PA 24.0%, UM 17.3%, CM 32.0%; submitted 2026-05-01; run date not published
24.4%2026-05-01
10Anthropic logoclaude-code + claude-fable-5 Anthropic
All Domains pass@1; PA 24.0%, UM 24.0%, CM 24.0%; submitted 2026-07-22; run date not published
24.0%2026-07-22
11Mhermes + MedGuard MedGuard
All Domains pass@1; PA 4.0%, UM 4.0%, CM 60.0%; submitted 2026-07-06; run date not published
22.7%2026-07-06
12OpenAI logocodex + gpt-5.5 OpenAI
All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not published
20.9%2026-05-01
13Anthropic logoclaude-code + claude-sonnet-5 Anthropic
All Domains pass@1; PA 24.0%, UM 24.0%, CM 12.0%; submitted 2026-07-06; run date not published
20.0%2026-07-06
14ZAopenai-agents + glm-5.1 Zhipu AI
All Domains pass@1; PA 18.7%, UM 33.3%, CM 4.0%; submitted 2026-05-01; run date not published
18.7%2026-05-01
15ZAhermes + glm-5.1 Zhipu AI
All Domains pass@1; PA 10.7%, UM 34.7%, CM 10.7%; submitted 2026-05-01; run date not published
18.7%2026-05-01
16ZAopenai-agents + glm-5.2 Zhipu AI
All Domains pass@1; PA 20.0%, UM 32.0%, CM 4.0%; submitted 2026-07-06; run date not published
18.7%2026-07-06
17Anthropic logoopenclaw + claude-opus-4-7 Anthropic
All Domains pass@1; PA 18.7%, UM 13.3%, CM 20.0%; submitted 2026-05-01; run date not published
17.3%2026-05-01
18ZAopenclaw + glm-5.1 Zhipu AI
All Domains pass@1; PA 13.3%, UM 26.7%, CM 10.7%; submitted 2026-05-01; run date not published
16.9%2026-05-01
19Ahermes + qwen-3.6-max Alibaba
All Domains pass@1; PA 9.3%, UM 26.7%, CM 13.3%; submitted 2026-05-01; run date not published
16.4%2026-05-01
20OpenAI logocodex + gpt-5.4 OpenAI
All Domains pass@1; PA 24.0%, UM 17.3%, CM 6.7%; submitted 2026-05-01; run date not published
16.0%2026-05-01
21Aopenai-agents + qwen-3.6-max Alibaba
All Domains pass@1; PA 16.0%, UM 26.7%, CM 4.0%; submitted 2026-05-01; run date not published
15.6%2026-05-01
22Moonshot AI logohermes + kimi-k2.6 Moonshot AI
All Domains pass@1; PA 18.7%, UM 21.3%, CM 6.7%; submitted 2026-05-01; run date not published
15.6%2026-05-01
23Moonshot AI logoopenai-agents + kimi-k2.6 Moonshot AI
All Domains pass@1; PA 17.3%, UM 25.3%, CM 2.7%; submitted 2026-05-01; run date not published
15.1%2026-05-01
24Dopenai-agents + deepseek-v4-pro DeepSeek
All Domains pass@1; PA 10.7%, UM 28.0%, CM 4.0%; submitted 2026-05-01; run date not published
14.2%2026-05-01
25Dhermes + deepseek-v4-pro DeepSeek
All Domains pass@1; PA 8.0%, UM 25.3%, CM 8.0%; submitted 2026-05-01; run date not published
13.8%2026-05-01
26OpenAI logocodex + gpt-5.6-terra OpenAI
All Domains pass@1; PA 12.0%, UM 20.0%, CM 8.0%; submitted 2026-07-24; run date not published
13.3%2026-07-24
27OpenAI logocodex + gpt-5.6-luna OpenAI
All Domains pass@1; PA 20.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published
13.3%2026-07-24
28Google logogemini-cli + gemini-3-flash Google
All Domains pass@1; PA 18.7%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published
12.5%2026-05-01
29Dopenclaw + deepseek-v4-pro DeepSeek
All Domains pass@1; PA 14.7%, UM 12.0%, CM 6.7%; submitted 2026-05-01; run date not published
11.1%2026-05-01
30ZAdeepagents + glm-5.1 Zhipu AI
All Domains pass@1; PA 17.3%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published
11.1%2026-05-01
31Ddeepagents + deepseek-v4-pro DeepSeek
All Domains pass@1; PA 14.7%, UM 10.7%, CM 6.7%; submitted 2026-05-01; run date not published
10.7%2026-05-01
32Moonshot AI logoopenclaw + kimi-k2.6 Moonshot AI
All Domains pass@1; PA 12.0%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published
10.2%2026-05-01
33Adeepagents + qwen-3.6-max Alibaba
All Domains pass@1; PA 12.0%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published
9.3%2026-05-01
34OpenAI logocodex + gpt-5.4-mini OpenAI
All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not published
8.4%2026-05-01
35TMopenai-agents + TML Inkling 256K Thinking Machines
All Domains pass@1; PA 4.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published
8.0%2026-07-24
36Google logogemini-cli + gemini-3.1-pro Google
All Domains pass@1; PA 14.7%, UM 6.7%, CM 0.0%; submitted 2026-05-01; run date not published
7.1%2026-05-01
37Anthropic logoclaude-code + claude-haiku-4-5 Anthropic
All Domains pass@1; PA 0.0%, UM 14.7%, CM 4.0%; submitted 2026-05-01; run date not published
6.2%2026-05-01
38SAopenai-agents + grok-4.3 SpaceX AI
All Domains pass@1; PA 0.0%, UM 16.0%, CM 1.3%; submitted 2026-05-01; run date not published
5.8%2026-05-01
39Aopenclaw + qwen-3.6-max Alibaba
All Domains pass@1; PA 10.7%, UM 4.0%, CM 0.0%; submitted 2026-05-01; run date not published
4.9%2026-05-01
40SAhermes + grok-4.3 SpaceX AI
All Domains pass@1; PA 0.0%, UM 13.3%, CM 0.0%; submitted 2026-05-01; run date not published
4.4%2026-05-01
41Moonshot AI logodeepagents + kimi-k2.6 Moonshot AI
All Domains pass@1; PA 8.0%, UM 1.3%, CM 0.0%; submitted 2026-05-01; run date not published
3.1%2026-05-01
42SAdeepagents + grok-4.3 SpaceX AI
All Domains pass@1; PA 0.0%, UM 5.3%, CM 1.3%; submitted 2026-05-01; run date not published
2.2%2026-05-01
43SAopenclaw + grok-4.3 SpaceX AI
All Domains pass@1; PA 1.3%, UM 0.0%, CM 0.0%; submitted 2026-05-01; run date not published
0.4%2026-05-01
44NVIDIA logoopenai-agents + Nemotron 3 Ultra 256K NVIDIA
All Domains pass@1; PA 0.0%, UM 0.0%, CM 0.0%; submitted 2026-07-24; run date not published
0.0%2026-07-24

Scores preserve their source precision, with any scale conversion documented (mixed sources). Official board updated 2026-08-12: 45 submitted harness configurations, 44 with all-domain accuracy. The PA-only MedArise submission has no all-domain score and is excluded here. Row dates are submission dates; run dates are not published. Community submissions and author-run baselines share the automated workspace judge. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisheractAVA.ai
categoryagentic and workflow benchmarks
released2026-05
size75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools
scalepass@1 with binary 0/1 reward, higher better
result basismixed sources
sourceCHI-Bench leaderboard (actAVA)
last frontier result2026-08-12

What is CHI-Bench?

CHI-Bench is a agentic and workflow benchmark from actAVA, released 2026-05: 75 workflows (25 per domain), 21 healthcare applications, 200+ MCP tools, scored on a pass@1 with binary 0/1 reward scale. Long-horizon US healthcare operations workflows for agents (prior authorization, utilization management, care management), 60-80 step tasks across 4-6 stages, judged by deterministic unit tests plus an LLM judge for evidence grounding, consent, and cross-stage consistency.

Which model leads CHI-Bench?

erius + claude-opus-5 (Humana (harness) / Anthropic (model)) has the highest indexed numerical score on CHI-Bench at 54.7% (evaluation setups may differ), per CHI-Bench leaderboard (actAVA), as of 2026-07-26.

Where do the CHI-Bench numbers come from?

From CHI-Bench leaderboard (actAVA) (mixed sources). Official board updated 2026-08-12: 45 submitted harness configurations, 44 with all-domain accuracy. The PA-only MedArise submission has no all-domain score and is excluded here. Row dates are submission dates; run dates are not published. Community submissions and author-run baselines share the automated workspace judge.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.