Moonshot AI: healthcare benchmark results
8 models · 20 results · snapshot reviewed September 28, 2026
Everything the index holds for Moonshot AI, gathered in one place: Kimi K3, Kimi K2.5, Kimi K2.6, openai-agents + kimi-k3, hermes + kimi-k2.6, openai-agents + kimi-k2.6, openclaw + kimi-k2.6, deepagents + kimi-k2.6. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Kimi K3 Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | Health Optimization Bench | 59.9 | 6 of 16 |
| Kimi K3 | MAST (Medical AI Superintelligence Test) | 60.1% | 2 of 8 source board: 11 |
| Kimi K3 | First, Do NOHARM (v2) | 74.0% | 7 of 17 source board: 19 |
| Kimi K3 model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 48.88% | 24 of 102 |
| Kimi K3 model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 87.96% | 13 of 104 |
| Kimi K3 (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 45 | 9 of 25 source board: 77 |
| Kimi K2.5 model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 2.119 pp; $0.017275/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 39.32% | 64 of 102 |
| Kimi K2.5 model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 1.986 pp; $0.024890/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 76.44% | 71 of 104 |
| Kimi K2.5 Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column) | MedXpertQA (MM) | 65.3 | 14 of 22 |
| Kimi-K2.5 headline 12-model evaluation, top open-weight | EHR-Complex | 0.62 | 3 of 18 |
| Kimi K2.5 screenshot-only, task description + portal guidance | HealthAdminBench | 15.6% | 3 of 7 |
| Kimi K2.6 | First, Do NOHARM (v2) | 59.1% | 15 of 17 source board: 19 |
| Kimi K2.6 model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 2.041 pp; $0.041295/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 40.14% | 63 of 102 |
| Kimi K2.6 model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 1.792 pp; $0.055962/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 78.15% | 62 of 104 |
| Kimi-K2.6 open source | PhysicianBench | 17.0 ± 2.6 | 16 of 21 |
| openai-agents + kimi-k3 All Domains pass@1; PA 28.0%, UM 32.0%, CM 16.0%; submitted 2026-07-24; run date not published | CHI-Bench | 25.3% | 8 of 44 source board: 45 |
| hermes + kimi-k2.6 All Domains pass@1; PA 18.7%, UM 21.3%, CM 6.7%; submitted 2026-05-01; run date not published | CHI-Bench | 15.6% | 22 of 44 source board: 45 |
| openai-agents + kimi-k2.6 All Domains pass@1; PA 17.3%, UM 25.3%, CM 2.7%; submitted 2026-05-01; run date not published | CHI-Bench | 15.1% | 23 of 44 source board: 45 |
| openclaw + kimi-k2.6 All Domains pass@1; PA 12.0%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published | CHI-Bench | 10.2% | 32 of 44 source board: 45 |
| deepagents + kimi-k2.6 All Domains pass@1; PA 8.0%, UM 1.3%, CM 0.0%; submitted 2026-05-01; run date not published | CHI-Bench | 3.1% | 41 of 44 source board: 45 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Moonshot AI appear on?
As of September 28, 2026, Moonshot AI models hold 20 indexed results across 11 tracked benchmarks, through Kimi K3, Kimi K2.5, Kimi K2.6, openai-agents + kimi-k3, hermes + kimi-k2.6, openai-agents + kimi-k2.6, openclaw + kimi-k2.6, deepagents + kimi-k2.6.
Where does Moonshot AI have the highest indexed score?
Moonshot AI does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.