Health Evals

Meta logoMeta: healthcare benchmark results

9 models · 22 results · snapshot reviewed September 28, 2026

Everything the index holds for Meta, gathered in one place: Muse Spark, Muse Spark 1.1, Llama 4 Maverick, Llama 4 Scout, Muse Spark 1.2, Muse Spark 1.3 (max), Muse Glimmer (high), Llama 3.3 70B, Llama 3.1 405B. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Muse Spark
length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
HealthBench Professional0.54117 of 29
Muse Spark
raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.
HealthBench Hard0.4282 of 20
Muse Spark
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench57.27 of 16
Muse Spark (2026-04-08)MedHELM0.6213 of 10
source board: 11
Muse Spark
model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)51.31%14 of 102
Muse Spark
model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.90%21 of 104
Muse Spark
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
MedXpertQA (MM)78.4%5 of 22
Muse Spark 1.1
length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
HealthBench Professional0.59311 of 29
Muse Spark 1.1First, Do NOHARM (v2)79.7%4 of 17
source board: 19
Muse Spark 1.1
model ID meta/muse_spark_1_1; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.95 pp; $0.034628/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)88.89%8 of 104
Llama 4 Maverick
model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.994 pp; $0.002888/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)36.51%73 of 102
Llama 4 Maverick
model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.871 pp; $0.002460/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)54.22%102 of 104
Llama 4 Maverick
95% bootstrap CI 39.6–44.6; 3 runs; temperature 0; zero-shot, closed-book
WHBench42.1%17 of 22
Llama 4 Scout
model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)23.31%99 of 102
Llama 4 Scout
model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)50.59%103 of 104
Llama 4 Scout
95% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-book
WHBench35.2%22 of 22
Muse Spark 1.2
model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.187 pp; $0.039023/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)49.35%20 of 102
Muse Spark 1.2
model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.037779/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)90.06%5 of 104
Muse Spark 1.3 (max)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index505 of 25
source board: 77
Muse Glimmer (high)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index1824 of 25
source board: 77
Llama 3.3 70B
95% bootstrap CI 35.2–40.5; 3 runs; temperature 0; zero-shot, closed-book
WHBench37.8%19 of 22
Llama 3.1 405B
95% bootstrap CI 33.9–38.3; 3 runs; temperature 0; zero-shot, closed-book
WHBench36.1%20 of 22

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Meta appear on?

As of September 28, 2026, Meta models hold 22 indexed results across 10 tracked benchmarks, through Muse Spark, Muse Spark 1.1, Llama 4 Maverick, Llama 4 Scout, Muse Spark 1.2, Muse Spark 1.3 (max), Muse Glimmer (high), Llama 3.3 70B, Llama 3.1 405B.

Where does Meta have the highest indexed score?

Meta does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.