Health Evals

MAST (Medical AI Superintelligence Test): published results

ARISE AI Research Network (multi-institutional) · composite of 6 component benchmarks; 11 models · index updated September 28, 2026

GPT-5.6 Sol has the highest indexed numerical score on MAST (Medical AI Superintelligence Test), 60.2% as of 2026-08, per MAST: Medical AI Superintelligence Test leaderboard (General board). Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.

Published results

#modelscoreas of
1OpenAI logoGPT-5.6 Sol OpenAI
MAST in preview; 'exact scores may change'
60.2%2026-08
2Moonshot AI logoKimi K3 Moonshot AI60.1%2026-08
3Google logoGemini 3.6 Flash Google59.3%2026-08
4Google logoGemini 3.1 Pro Google58.9%2026-08
5AQwen3.5 397B A17B Alibaba57.9%2026-08
6Anthropic logoClaude Opus 5 Anthropic57.1%2026-08
7Anthropic logoClaude Sonnet 5 Anthropic56.6%2026-08
8xAI logoGrok 4.3 xAI53.7%2026-08

Scores preserve their source precision, with any scale conversion documented (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release. The official overview still states August 15, 2026. Eight displayed general-ranking rows are indexed here, from eleven models on that view. MAST is in preview and scores may change. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherARISE AI Research Network (multi-institutional)
categorycomposite indices
released2026-08
sizecomposite of 6 component benchmarks; 11 models
scalepercentage composite, higher better
result basisindependently run
sourceMAST: Medical AI Superintelligence Test leaderboard (General board)
last frontier result2026-08

What is MAST (Medical AI Superintelligence Test)?

MAST (Medical AI Superintelligence Test) is a composite benchmark from ARISE AI Research Network, released 2026-08: composite of 6 component benchmarks; 11 models, scored on a percentage composite scale. Composite score across curated clinical benchmarks spanning diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. Components: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, CPC-Bench.

Which model leads MAST (Medical AI Superintelligence Test)?

GPT-5.6 Sol (OpenAI) has the highest indexed numerical score on MAST (Medical AI Superintelligence Test) at 60.2% (evaluation setups may differ), per MAST: Medical AI Superintelligence Test leaderboard (General board), as of 2026-08.

Where do the MAST (Medical AI Superintelligence Test) numbers come from?

From MAST: Medical AI Superintelligence Test leaderboard (General board) (independently run). The board is marked as a preview and was last updated August 15, 2026; component-level breakdowns are published only for First, Do NOHARM v2. Scores may move before the full release. The official overview still states August 15, 2026. Eight displayed general-ranking rows are indexed here, from eleven models on that view. MAST is in preview and scores may change.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.