Health Evals

Artificial Analysis Healthcare & Medical Index: published results

Artificial Analysis · Six-evaluation composite; 25 default-selected models indexed · index updated September 28, 2026

Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) has the highest indexed numerical score on Artificial Analysis Healthcare & Medical Index, 61 as of 2026-09, per Artificial Analysis Healthcare & Medical Index. Artificial Analysis independently evaluates a weighted healthcare composite: knowledge 30%, agentic knowledge work 25%, long-context reasoning 15%, non-hallucination 10%, reasoning 10%, and agentic tool use 10%. Components include AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity’s Last Exam, and AutomationBench-AA.

Published results

#modelscoreas of
1Anthropic logoClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
612026-09
2Anthropic logoClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
582026-09
3Anthropic logoClaude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
532026-09
4OpenAI logoGPT-6 Astra (max) OpenAI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
522026-09
5Meta logoMuse Spark 1.3 (max) Meta
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
502026-09
6SGrok 4.7 (xhigh) SpaceXAI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
472026-09
7ZAGLM-5.3 (max) Z AI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
472026-09
8OpenAI logoGPT-5.6 Sol (max) OpenAI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
452026-09
9KKimi K3 (max) Kimi
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
452026-09
10ZAGLM 5.3 Flash Z AI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
452026-09
11OpenAI logoGPT-6 Sol (max) OpenAI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
432026-09
12XMiMo-V2.6-Pro Xiaomi
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
422026-09
13Google logoGemini 3.8 Flash (high) Google
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
422026-09
14AQwen3.8 Max (0902) Alibaba
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
412026-09
15SStep 5 Preview StepFun
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
412026-09
16DDeepSeek V4.1 Flash (Reasoning, Max Effort) DeepSeek
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
412026-09
17OpenAI logoGPT-6 Luna (max) OpenAI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
372026-09
18OpenAI logoGPT-5.6 Luna (max) OpenAI
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
362026-09
19AQwen3.8 27B (xhigh) Alibaba
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
342026-09
20MiniMax logoMiniMax-M3 MiniMax
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
302026-09
21TMInkling (xhigh) Thinking Machines
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
252026-09
22NVIDIA logoNemotron 3 Ultra 550B A55B (Reasoning) NVIDIA
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
232026-09
23Google logoGemini 3.5 Flash-Lite Google
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
232026-09
24Meta logoMuse Glimmer (high) Meta
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
182026-09
25Mistral logoMistral Medium 3.5 Mistral
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
142026-09

Scores preserve their source precision, with any scale conversion documented (independently run). Current component versions and weights were checked September 28, 2026. This snapshot indexes the 25 default-selected models with numeric data embedded in the official page; the page offers 77 model variants overall. Historical rows absent from that snapshot are omitted rather than mixed with the prior five-evaluation composition. Scores are rounded whole index points; underlying values are recorded in source locators. Individual model evaluation dates are not published. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherArtificial Analysis
categorycomposite indices
released2026-08
sizeSix-evaluation composite; 25 default-selected models indexed
scaleindex score, higher better
result basisindependently run
sourceBest AI for Healthcare & Medical: LLM Leaderboard
last frontier result2026-09

What is Artificial Analysis Healthcare & Medical Index?

Artificial Analysis Healthcare & Medical Index is a composite benchmark from Artificial Analysis, released 2026-08: Six-evaluation composite; 25 default-selected models indexed, scored on a index score scale. Artificial Analysis independently evaluates a weighted healthcare composite: knowledge 30%, agentic knowledge work 25%, long-context reasoning 15%, non-hallucination 10%, reasoning 10%, and agentic tool use 10%. Components include AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity’s Last Exam, and AutomationBench-AA.

Which model leads Artificial Analysis Healthcare & Medical Index?

Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) (Anthropic) has the highest indexed numerical score on Artificial Analysis Healthcare & Medical Index at 61 (evaluation setups may differ), per Artificial Analysis Healthcare & Medical Index, as of 2026-09.

Where do the Artificial Analysis Healthcare & Medical Index numbers come from?

From Best AI for Healthcare & Medical: LLM Leaderboard (independently run). Current component versions and weights were checked September 28, 2026. This snapshot indexes the 25 default-selected models with numeric data embedded in the official page; the page offers 77 model variants overall. Historical rows absent from that snapshot are omitted rather than mixed with the prior five-evaluation composition. Scores are rounded whole index points; underlying values are recorded in source locators. Individual model evaluation dates are not published.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.