MAST (Medical AI Superintelligence Test): published results
ARISE AI Research Network (multi-institutional) · composite of 6 component benchmarks; 11 models · index updated September 28, 2026
GPT-5.6 Sol has the highest indexed numerical score on MAST (Medical AI Superintelligence Test), 60.2% as of 2026-08, per MAST: Medical AI Superintelligence Test leaderboard (General board). Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol OpenAI MAST in preview; 'exact scores may change' | 60.2% | 2026-08 | |
| 2 | Kimi K3 Moonshot AI | 60.1% | 2026-08 | |
| 3 | Gemini 3.6 Flash Google | 59.3% | 2026-08 | |
| 4 | Gemini 3.1 Pro Google | 58.9% | 2026-08 | |
| 5 | A | Qwen3.5 397B A17B Alibaba | 57.9% | 2026-08 |
| 6 | Claude Opus 5 Anthropic | 57.1% | 2026-08 | |
| 7 | Claude Sonnet 5 Anthropic | 56.6% | 2026-08 | |
| 8 | Grok 4.3 xAI | 53.7% | 2026-08 | |
Scores preserve their source precision, with any scale conversion documented (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release. The official overview still states August 15, 2026. Eight displayed general-ranking rows are indexed here, from eleven models on that view. MAST is in preview and scores may change. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | ARISE AI Research Network (multi-institutional) |
|---|---|
| category | composite indices |
| released | 2026-08 |
| size | composite of 6 component benchmarks; 11 models |
| scale | percentage composite, higher better |
| result basis | independently run |
| source | MAST: Medical AI Superintelligence Test leaderboard (General board) |
| last frontier result | 2026-08 |
What is MAST (Medical AI Superintelligence Test)?
MAST (Medical AI Superintelligence Test) is a composite benchmark from ARISE AI Research Network, released 2026-08: composite of 6 component benchmarks; 11 models, scored on a percentage composite scale. Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.
Which model leads MAST (Medical AI Superintelligence Test)?
GPT-5.6 Sol (OpenAI) has the highest indexed numerical score on MAST (Medical AI Superintelligence Test) at 60.2% (evaluation setups may differ), per MAST: Medical AI Superintelligence Test leaderboard (General board), as of 2026-08.
Where do the MAST (Medical AI Superintelligence Test) numbers come from?
From MAST: Medical AI Superintelligence Test leaderboard (General board) (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release. The official overview still states August 15, 2026. Eight displayed general-ranking rows are indexed here, from eleven models on that view. MAST is in preview and scores may change.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.