Health Evals

WHBench: published results

Independent researchers (Maurya, Govindgari, Kumar) · 47 scenarios / 3,100 scored responses across 22 models · index updated September 28, 2026

Claude Opus 4.6 has the highest indexed numerical score on WHBench, 72.1% as of 2026-03, per WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2). Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.

Published results

#modelscoreas of
1Anthropic logoClaude Opus 4.6 Anthropic
95% CI 69.6-74.4; evaluations run March 2026
72.1%2026-03
2Anthropic logoClaude Sonnet 4.6 Anthropic
95% CI 64.5-69.6
67.1%2026-03
3OpenAI logoGPT-5.4 OpenAI
95% CI 64.5-69.2
66.8%2026-03
4Google logoGemini 3 Flash Preview Google64.7%2026-03
5OpenAI logoOpenAI o3 OpenAI
95% bootstrap CI 61.3–65.9; 3 runs; temperature 0; zero-shot, closed-book
63.6%2026-03
6DDeepSeek V3.2 DeepSeek
95% bootstrap CI 58.6–63.9; 3 runs; temperature 0; zero-shot, closed-book
61.3%2026-03
7SAGrok 3 SpaceX AI
95% bootstrap CI 58.0–63.4; 3 runs; temperature 0; zero-shot, closed-book
60.7%2026-03
8MAMistral Large Mistral AI
95% bootstrap CI 57.4–63.0; 3 runs; temperature 0; zero-shot, closed-book
60.2%2026-03
9SAGrok 4 SpaceX AI
95% bootstrap CI 54.9–60.8; 3 runs; temperature 0; zero-shot, closed-book
57.9%2026-03
10DDeepSeek-R1 DeepSeek
95% bootstrap CI 50.5–55.3; 3 runs; temperature 0; zero-shot, closed-book
52.9%2026-03
11OpenAI logoGPT-4.1 OpenAI51.8%2026-03
12SAGrok 3 Mini SpaceX AI
95% bootstrap CI 47.5–52.5; 3 runs; temperature 0; zero-shot, closed-book
50.0%2026-03
13Google logoGemini 2.5 Flash Google
95% bootstrap CI 47.0–52.0; 3 runs; temperature 0; zero-shot, closed-book
49.5%2026-03
14Anthropic logoClaude Opus 4 Anthropic
95% bootstrap CI 46.4–51.7; 3 runs; temperature 0; zero-shot, closed-book
49.1%2026-03
15Anthropic logoClaude Sonnet 4 Anthropic
95% bootstrap CI 45.5–50.6; 3 runs; temperature 0; zero-shot, closed-book
48.1%2026-03
16OpenAI logoGPT-4o OpenAI44.6%2026-03
17Meta logoLlama 4 Maverick Meta
95% bootstrap CI 39.6–44.6; 3 runs; temperature 0; zero-shot, closed-book
42.1%2026-03
18NVIDIA logoNemotron 70B NVIDIA
95% bootstrap CI 37.3–41.3; 3 runs; temperature 0; zero-shot, closed-book
39.3%2026-03
19Meta logoLlama 3.3 70B Meta
95% bootstrap CI 35.2–40.5; 3 runs; temperature 0; zero-shot, closed-book
37.8%2026-03
20Meta logoLlama 3.1 405B Meta
95% bootstrap CI 33.9–38.3; 3 runs; temperature 0; zero-shot, closed-book
36.1%2026-03
21Google logoGemini 2.5 Pro Google
95% bootstrap CI 32.7–38.1; 3 runs; temperature 0; zero-shot, closed-book
35.3%2026-03
22Meta logoLlama 4 Scout Meta
95% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-book
35.2%2026-03

Scores preserve their source precision, with any scale conversion documented (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherIndependent researchers (Maurya, Govindgari, Kumar)
categoryrubric-graded benchmarks
released2026-04
size47 scenarios / 3,100 scored responses across 22 models
scalemean normalized percentage 0-100, higher better
result basisindependently run
sourceWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2)
official pagearxiv.org/abs/2604.00024
last frontier result2026-03

What is WHBench?

WHBench is a rubric-graded benchmark from academic team, released 2026-04: 47 scenarios / 3,100 scored responses across 22 models, scored on a mean normalized percentage 0-100 scale. Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.

Which model leads WHBench?

Claude Opus 4.6 (Anthropic) has the highest indexed numerical score on WHBench at 72.1% (evaluation setups may differ), per WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2), as of 2026-03.

Where do the WHBench numbers come from?

From WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.