Muse Spark: healthcare benchmark results
Meta · 7 boards · snapshot reviewed September 28, 2026
The index currently holds 7 results for Muse Spark: 0.541 on HealthBench Professional (17 of 29 indexed rows), 0.428 on HealthBench Hard (2 of 20 indexed rows), 57.2 on Health Optimization Bench (7 of 16 indexed rows), 0.621 on MedHELM [Muse Spark (2026-04-08)] (3 of 10 indexed rows), 51.31% on MedCode (Vals AI) (14 of 102 indexed rows), 85.90% on MedScribe (Vals AI) (21 of 104 indexed rows), 78.4% on MedXpertQA (MM) (5 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44 | 0.541 | 17 of 29 | 2026-07 |
| HealthBench Hard raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table. | 0.428 | 2 of 20 | 2026-04 |
| Health Optimization Bench Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 57.2 | 7 of 16 | 2026-09 |
| MedHELM Muse Spark (2026-04-08) | 0.621 | 3 of 10 source board: 11 | 2026-05 |
| MedCode (Vals AI) model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not published | 51.31% | 14 of 102 | 2026-09-26 |
| MedScribe (Vals AI) model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not published | 85.90% | 21 of 104 | 2026-09-26 |
| MedXpertQA (MM) Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 78.4% | 5 of 22 | 2026-04 |
Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Muse Spark (2026-04-08)".
Which healthcare benchmarks is Muse Spark scored on?
As of September 28, 2026, Muse Spark has indexed results on 7 tracked benchmarks: HealthBench Professional, HealthBench Hard, Health Optimization Bench, MedHELM, MedCode (Vals AI), MedScribe (Vals AI), MedXpertQA (MM).
Which results are indexed for Muse Spark?
Muse Spark stands at 0.541 on HealthBench Professional (17 of 29 indexed rows), 0.428 on HealthBench Hard (2 of 20 indexed rows), 57.2 on Health Optimization Bench (7 of 16 indexed rows), 0.621 on MedHELM [Muse Spark (2026-04-08)] (3 of 10 indexed rows), 51.31% on MedCode (Vals AI) (14 of 102 indexed rows), 85.90% on MedScribe (Vals AI) (21 of 104 indexed rows), 78.4% on MedXpertQA (MM) (5 of 22 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.