Health Evals

Moonshot AI logoKimi K3: healthcare benchmark results

Moonshot AI · 6 boards · snapshot reviewed September 28, 2026

The index currently holds 6 results for Kimi K3: 59.9 on Health Optimization Bench (6 of 16 indexed rows), 60.1% on MAST (Medical AI Superintelligence Test) (2 of 8 indexed rows), 74.0% on First, Do NOHARM (v2) (7 of 17 indexed rows), 48.88% on MedCode (Vals AI) (24 of 102 indexed rows), 87.96% on MedScribe (Vals AI) (13 of 104 indexed rows), 45 on Artificial Analysis Healthcare & Medical Index [Kimi K3 (max)] (9 of 25 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
Health Optimization Bench
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
59.96 of 162026-09
MAST (Medical AI Superintelligence Test)60.1%2 of 8
source board: 11
2026-08
First, Do NOHARM (v2)74.0%7 of 17
source board: 19
2026-08
MedCode (Vals AI)
model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not published
48.88%24 of 1022026-09-26
MedScribe (Vals AI)
model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not published
87.96%13 of 1042026-09-26
Artificial Analysis Healthcare & Medical Index
Kimi K3 (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
459 of 25
source board: 77
2026-09

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Kimi K3 (max)".

Which healthcare benchmarks is Kimi K3 scored on?

As of September 28, 2026, Kimi K3 has indexed results on 6 tracked benchmarks: Health Optimization Bench, MAST (Medical AI Superintelligence Test), First, Do NOHARM (v2), MedCode (Vals AI), MedScribe (Vals AI), Artificial Analysis Healthcare & Medical Index.

Which results are indexed for Kimi K3?

Kimi K3 stands at 59.9 on Health Optimization Bench (6 of 16 indexed rows), 60.1% on MAST (Medical AI Superintelligence Test) (2 of 8 indexed rows), 74.0% on First, Do NOHARM (v2) (7 of 17 indexed rows), 48.88% on MedCode (Vals AI) (24 of 102 indexed rows), 87.96% on MedScribe (Vals AI) (13 of 104 indexed rows), 45 on Artificial Analysis Healthcare & Medical Index [Kimi K3 (max)] (9 of 25 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.