MAI-Thinking-1: healthcare benchmark results
Microsoft · 2 boards · snapshot reviewed September 28, 2026
The index currently holds 2 results for MAI-Thinking-1: 0.350 on HealthBench Professional (29 of 29 indexed rows), 17.5 on Health Optimization Bench [MAI Thinking] (14 of 16 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision. | 0.350 | 29 of 29 | 2026-08 |
| Health Optimization Bench MAI Thinking Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 17.5 | 14 of 16 | 2026-09 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "MAI Thinking".
Which healthcare benchmarks is MAI-Thinking-1 scored on?
As of September 28, 2026, MAI-Thinking-1 has indexed results on 2 tracked benchmarks: HealthBench Professional, Health Optimization Bench.
Which results are indexed for MAI-Thinking-1?
MAI-Thinking-1 stands at 0.350 on HealthBench Professional (29 of 29 indexed rows), 17.5 on Health Optimization Bench [MAI Thinking] (14 of 16 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.