Artificial Analysis Healthcare & Medical Index: published results
Artificial Analysis · Six-evaluation composite; 25 default-selected models indexed · index updated September 28, 2026
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) has the highest indexed numerical score on Artificial Analysis Healthcare & Medical Index, 61 as of 2026-09, per Artificial Analysis Healthcare & Medical Index. Artificial Analysis independently evaluates a weighted healthcare composite: knowledge 30%, agentic knowledge work 25%, long-context reasoning 15%, non-hallucination 10%, reasoning 10%, and agentic tool use 10%. Components include AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity’s Last Exam, and AutomationBench-AA.
Published results
Showing top 20 of 25 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 61 | 2026-09 | |
| 2 | Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 58 | 2026-09 | |
| 3 | Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 53 | 2026-09 | |
| 4 | GPT-6 Astra (max) OpenAI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 52 | 2026-09 | |
| 5 | Muse Spark 1.3 (max) Meta Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 50 | 2026-09 | |
| 6 | S | Grok 4.7 (xhigh) SpaceXAI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 47 | 2026-09 |
| 7 | ZA | GLM-5.3 (max) Z AI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 47 | 2026-09 |
| 8 | GPT-5.6 Sol (max) OpenAI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 45 | 2026-09 | |
| 9 | K | Kimi K3 (max) Kimi Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 45 | 2026-09 |
| 10 | ZA | GLM 5.3 Flash Z AI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 45 | 2026-09 |
| 11 | GPT-6 Sol (max) OpenAI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 43 | 2026-09 | |
| 12 | X | MiMo-V2.6-Pro Xiaomi Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 42 | 2026-09 |
| 13 | Gemini 3.8 Flash (high) Google Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 42 | 2026-09 | |
| 14 | A | Qwen3.8 Max (0902) Alibaba Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 41 | 2026-09 |
| 15 | S | Step 5 Preview StepFun Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 41 | 2026-09 |
| 16 | D | DeepSeek V4.1 Flash (Reasoning, Max Effort) DeepSeek Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 41 | 2026-09 |
| 17 | GPT-6 Luna (max) OpenAI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 37 | 2026-09 | |
| 18 | GPT-5.6 Luna (max) OpenAI Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 36 | 2026-09 | |
| 19 | A | Qwen3.8 27B (xhigh) Alibaba Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 34 | 2026-09 |
| 20 | MiniMax-M3 MiniMax Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 30 | 2026-09 | |
| 21 | TM | Inkling (xhigh) Thinking Machines Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 25 | 2026-09 |
| 22 | Nemotron 3 Ultra 550B A55B (Reasoning) NVIDIA Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 23 | 2026-09 | |
| 23 | Gemini 3.5 Flash-Lite Google Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 23 | 2026-09 | |
| 24 | Muse Glimmer (high) Meta Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 18 | 2026-09 | |
| 25 | Mistral Medium 3.5 Mistral Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | 14 | 2026-09 | |
Scores preserve their source precision, with any scale conversion documented (independently run). Current component versions and weights were checked September 28, 2026. This snapshot indexes the 25 default-selected models with numeric data embedded in the official page; the page offers 77 model variants overall. Historical rows absent from that snapshot are omitted rather than mixed with the prior five-evaluation composition. Scores are rounded whole index points; underlying values are recorded in source locators. Individual model evaluation dates are not published. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Artificial Analysis |
|---|---|
| category | composite indices |
| released | 2026-08 |
| size | Six-evaluation composite; 25 default-selected models indexed |
| scale | index score, higher better |
| result basis | independently run |
| source | Best AI for Healthcare & Medical: LLM Leaderboard |
| last frontier result | 2026-09 |
What is Artificial Analysis Healthcare & Medical Index?
Artificial Analysis Healthcare & Medical Index is a composite benchmark from Artificial Analysis, released 2026-08: Six-evaluation composite; 25 default-selected models indexed, scored on a index score scale. Artificial Analysis independently evaluates a weighted healthcare composite: knowledge 30%, agentic knowledge work 25%, long-context reasoning 15%, non-hallucination 10%, reasoning 10%, and agentic tool use 10%. Components include AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity’s Last Exam, and AutomationBench-AA.
Which model leads Artificial Analysis Healthcare & Medical Index?
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) (Anthropic) has the highest indexed numerical score on Artificial Analysis Healthcare & Medical Index at 61 (evaluation setups may differ), per Artificial Analysis Healthcare & Medical Index, as of 2026-09.
Where do the Artificial Analysis Healthcare & Medical Index numbers come from?
From Best AI for Healthcare & Medical: LLM Leaderboard (independently run). Current component versions and weights were checked September 28, 2026. This snapshot indexes the 25 default-selected models with numeric data embedded in the official page; the page offers 77 model variants overall. Historical rows absent from that snapshot are omitted rather than mixed with the prior five-evaluation composition. Scores are rounded whole index points; underlying values are recorded in source locators. Individual model evaluation dates are not published.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.