WHBench: published results
Independent researchers (Maurya, Govindgari, Kumar) · 47 scenarios / 3,100 scored responses across 22 models · index updated September 28, 2026
Claude Opus 4.6 has the highest indexed numerical score on WHBench, 72.1% as of 2026-03, per WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2). Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.
Published results
Showing top 20 of 22 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 Anthropic 95% CI 69.6-74.4; evaluations run March 2026 via paper, arxiv.org | 72.1% | 2026-03 | |
| 2 | Claude Sonnet 4.6 Anthropic 95% CI 64.5-69.6 via paper, arxiv.org | 67.1% | 2026-03 | |
| 3 | GPT-5.4 OpenAI 95% CI 64.5-69.2 via paper, arxiv.org | 66.8% | 2026-03 | |
| 4 | Gemini 3 Flash Preview Google via paper, arxiv.org | 64.7% | 2026-03 | |
| 5 | OpenAI o3 OpenAI 95% bootstrap CI 61.3–65.9; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 63.6% | 2026-03 | |
| 6 | D | DeepSeek V3.2 DeepSeek 95% bootstrap CI 58.6–63.9; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 61.3% | 2026-03 |
| 7 | SA | Grok 3 SpaceX AI 95% bootstrap CI 58.0–63.4; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 60.7% | 2026-03 |
| 8 | MA | Mistral Large Mistral AI 95% bootstrap CI 57.4–63.0; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 60.2% | 2026-03 |
| 9 | SA | Grok 4 SpaceX AI 95% bootstrap CI 54.9–60.8; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 57.9% | 2026-03 |
| 10 | D | DeepSeek-R1 DeepSeek 95% bootstrap CI 50.5–55.3; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 52.9% | 2026-03 |
| 11 | GPT-4.1 OpenAI via paper, arxiv.org | 51.8% | 2026-03 | |
| 12 | SA | Grok 3 Mini SpaceX AI 95% bootstrap CI 47.5–52.5; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 50.0% | 2026-03 |
| 13 | Gemini 2.5 Flash Google 95% bootstrap CI 47.0–52.0; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 49.5% | 2026-03 | |
| 14 | Claude Opus 4 Anthropic 95% bootstrap CI 46.4–51.7; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 49.1% | 2026-03 | |
| 15 | Claude Sonnet 4 Anthropic 95% bootstrap CI 45.5–50.6; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 48.1% | 2026-03 | |
| 16 | GPT-4o OpenAI via paper, arxiv.org | 44.6% | 2026-03 | |
| 17 | Llama 4 Maverick Meta 95% bootstrap CI 39.6–44.6; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 42.1% | 2026-03 | |
| 18 | Nemotron 70B NVIDIA 95% bootstrap CI 37.3–41.3; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 39.3% | 2026-03 | |
| 19 | Llama 3.3 70B Meta 95% bootstrap CI 35.2–40.5; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 37.8% | 2026-03 | |
| 20 | Llama 3.1 405B Meta 95% bootstrap CI 33.9–38.3; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 36.1% | 2026-03 | |
| 21 | Gemini 2.5 Pro Google 95% bootstrap CI 32.7–38.1; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 35.3% | 2026-03 | |
| 22 | Llama 4 Scout Meta 95% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-book via paper, arxiv.org | 35.2% | 2026-03 | |
Scores preserve their source precision, with any scale conversion documented (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Independent researchers (Maurya, Govindgari, Kumar) |
|---|---|
| category | rubric-graded benchmarks |
| released | 2026-04 |
| size | 47 scenarios / 3,100 scored responses across 22 models |
| scale | mean normalized percentage 0-100, higher better |
| result basis | independently run |
| source | WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) |
| official page | arxiv.org/abs/2604.00024 |
| last frontier result | 2026-03 |
What is WHBench?
WHBench is a rubric-graded benchmark from academic team, released 2026-04: 47 scenarios / 3,100 scored responses across 22 models, scored on a mean normalized percentage 0-100 scale. Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.
Which model leads WHBench?
Claude Opus 4.6 (Anthropic) has the highest indexed numerical score on WHBench at 72.1% (evaluation setups may differ), per WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2), as of 2026-03.
Where do the WHBench numbers come from?
From WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.