Meta: healthcare benchmark results
9 models · 22 results · snapshot reviewed September 28, 2026
Everything the index holds for Meta, gathered in one place: Muse Spark, Muse Spark 1.1, Llama 4 Maverick, Llama 4 Scout, Muse Spark 1.2, Muse Spark 1.3 (max), Muse Glimmer (high), Llama 3.3 70B, Llama 3.1 405B. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Muse Spark length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44 | HealthBench Professional | 0.541 | 17 of 29 |
| Muse Spark raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table. | HealthBench Hard | 0.428 | 2 of 20 |
| Muse Spark Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | Health Optimization Bench | 57.2 | 7 of 16 |
| Muse Spark (2026-04-08) | MedHELM | 0.621 | 3 of 10 source board: 11 |
| Muse Spark model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 51.31% | 14 of 102 |
| Muse Spark model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 85.90% | 21 of 104 |
| Muse Spark Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | MedXpertQA (MM) | 78.4% | 5 of 22 |
| Muse Spark 1.1 length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44) | HealthBench Professional | 0.593 | 11 of 29 |
| Muse Spark 1.1 | First, Do NOHARM (v2) | 79.7% | 4 of 17 source board: 19 |
| Muse Spark 1.1 model ID meta/muse_spark_1_1; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.95 pp; $0.034628/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 88.89% | 8 of 104 |
| Llama 4 Maverick model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.994 pp; $0.002888/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 36.51% | 73 of 102 |
| Llama 4 Maverick model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.871 pp; $0.002460/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 54.22% | 102 of 104 |
| Llama 4 Maverick 95% bootstrap CI 39.6–44.6; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 42.1% | 17 of 22 |
| Llama 4 Scout model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 23.31% | 99 of 102 |
| Llama 4 Scout model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 50.59% | 103 of 104 |
| Llama 4 Scout 95% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 35.2% | 22 of 22 |
| Muse Spark 1.2 model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.187 pp; $0.039023/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 49.35% | 20 of 102 |
| Muse Spark 1.2 model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.037779/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 90.06% | 5 of 104 |
| Muse Spark 1.3 (max) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 50 | 5 of 25 source board: 77 |
| Muse Glimmer (high) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 18 | 24 of 25 source board: 77 |
| Llama 3.3 70B 95% bootstrap CI 35.2–40.5; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 37.8% | 19 of 22 |
| Llama 3.1 405B 95% bootstrap CI 33.9–38.3; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 36.1% | 20 of 22 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Meta appear on?
As of September 28, 2026, Meta models hold 22 indexed results across 10 tracked benchmarks, through Muse Spark, Muse Spark 1.1, Llama 4 Maverick, Llama 4 Scout, Muse Spark 1.2, Muse Spark 1.3 (max), Muse Glimmer (high), Llama 3.3 70B, Llama 3.1 405B.
Where does Meta have the highest indexed score?
Meta does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.