Health Evals

Meta logoLlama 4 Scout: healthcare benchmark results

Meta · 3 boards · snapshot reviewed September 28, 2026

The index currently holds 3 results for Llama 4 Scout: 23.31% on MedCode (Vals AI) (99 of 102 indexed rows), 50.59% on MedScribe (Vals AI) (103 of 104 indexed rows), 35.2% on WHBench (22 of 22 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
MedCode (Vals AI)
model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not published
23.31%99 of 1022026-09-26
MedScribe (Vals AI)
model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not published
50.59%103 of 1042026-09-26
WHBench
95% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-book
35.2%22 of 222026-03

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page.

Which healthcare benchmarks is Llama 4 Scout scored on?

As of September 28, 2026, Llama 4 Scout has indexed results on 3 tracked benchmarks: MedCode (Vals AI), MedScribe (Vals AI), WHBench.

Which results are indexed for Llama 4 Scout?

Llama 4 Scout stands at 23.31% on MedCode (Vals AI) (99 of 102 indexed rows), 50.59% on MedScribe (Vals AI) (103 of 104 indexed rows), 35.2% on WHBench (22 of 22 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.