Health Evals

Microsoft logoMicrosoft: healthcare benchmark results

1 model · 2 results · snapshot reviewed September 28, 2026

Everything the index holds for Microsoft, gathered in one place: MAI-Thinking-1. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
MAI-Thinking-1
length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.
HealthBench Professional0.35029 of 29
MAI Thinking
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench17.514 of 16

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Microsoft appear on?

As of September 28, 2026, Microsoft models hold 2 indexed results across 2 tracked benchmarks, through MAI-Thinking-1.

Where does Microsoft have the highest indexed score?

Microsoft does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Baichuan, Tencent, Inception, Cohere. The full field is on the index.