Microsoft: healthcare benchmark results
1 model · 2 results · snapshot reviewed September 28, 2026
Everything the index holds for Microsoft, gathered in one place: MAI-Thinking-1. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| MAI-Thinking-1 length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision. | HealthBench Professional | 0.350 | 29 of 29 |
| MAI Thinking Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | Health Optimization Bench | 17.5 | 14 of 16 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Microsoft appear on?
As of September 28, 2026, Microsoft models hold 2 indexed results across 2 tracked benchmarks, through MAI-Thinking-1.
Where does Microsoft have the highest indexed score?
Microsoft does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Baichuan, Tencent, Inception, Cohere. The full field is on the index.