Health Evals

Moonshot AI logoMoonshot AI: healthcare benchmark results

8 models · 20 results · snapshot reviewed September 28, 2026

Everything the index holds for Moonshot AI, gathered in one place: Kimi K3, Kimi K2.5, Kimi K2.6, openai-agents + kimi-k3, hermes + kimi-k2.6, openai-agents + kimi-k2.6, openclaw + kimi-k2.6, deepagents + kimi-k2.6. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Kimi K3
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench59.96 of 16
Kimi K3MAST (Medical AI Superintelligence Test)60.1%2 of 8
source board: 11
Kimi K3First, Do NOHARM (v2)74.0%7 of 17
source board: 19
Kimi K3
model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)48.88%24 of 102
Kimi K3
model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)87.96%13 of 104
Kimi K3 (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index459 of 25
source board: 77
Kimi K2.5
model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 2.119 pp; $0.017275/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)39.32%64 of 102
Kimi K2.5
model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 1.986 pp; $0.024890/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)76.44%71 of 104
Kimi K2.5
Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)
MedXpertQA (MM)65.314 of 22
Kimi-K2.5
headline 12-model evaluation, top open-weight
EHR-Complex0.623 of 18
Kimi K2.5
screenshot-only, task description + portal guidance
HealthAdminBench15.6%3 of 7
Kimi K2.6First, Do NOHARM (v2)59.1%15 of 17
source board: 19
Kimi K2.6
model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 2.041 pp; $0.041295/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.14%63 of 102
Kimi K2.6
model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 1.792 pp; $0.055962/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)78.15%62 of 104
Kimi-K2.6
open source
PhysicianBench17.0 ± 2.616 of 21
openai-agents + kimi-k3
All Domains pass@1; PA 28.0%, UM 32.0%, CM 16.0%; submitted 2026-07-24; run date not published
CHI-Bench25.3%8 of 44
source board: 45
hermes + kimi-k2.6
All Domains pass@1; PA 18.7%, UM 21.3%, CM 6.7%; submitted 2026-05-01; run date not published
CHI-Bench15.6%22 of 44
source board: 45
openai-agents + kimi-k2.6
All Domains pass@1; PA 17.3%, UM 25.3%, CM 2.7%; submitted 2026-05-01; run date not published
CHI-Bench15.1%23 of 44
source board: 45
openclaw + kimi-k2.6
All Domains pass@1; PA 12.0%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published
CHI-Bench10.2%32 of 44
source board: 45
deepagents + kimi-k2.6
All Domains pass@1; PA 8.0%, UM 1.3%, CM 0.0%; submitted 2026-05-01; run date not published
CHI-Bench3.1%41 of 44
source board: 45

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Moonshot AI appear on?

As of September 28, 2026, Moonshot AI models hold 20 indexed results across 11 tracked benchmarks, through Kimi K3, Kimi K2.5, Kimi K2.6, openai-agents + kimi-k3, hermes + kimi-k2.6, openai-agents + kimi-k2.6, openclaw + kimi-k2.6, deepagents + kimi-k2.6.

Where does Moonshot AI have the highest indexed score?

Moonshot AI does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.