Health Evals

Rubric-graded benchmarks

5 tracked · snapshot reviewed September 28, 2026

These benchmarks evaluate free-text medical answers against criteria written by physicians or domain experts. Rubrics can assess accuracy, completeness, communication, and safety. Each board has its own questions, grading process, and score scale; consult the evaluation setup before comparing results.

HealthBench Professional

OpenAI · 525 tasks

525 clinician tasks with physician-written rubrics; results retain evaluator, grader and length-adjustment details.

HealthBench Hard

OpenAI · 1,000 conversations

The 1,000 HealthBench conversations frontier models failed most, still graded on the original physician rubrics.

Showing top 12 of 20 indexed results. View all results.

HealthBench

OpenAI · 5,000 conversations

Realistic multi-turn health conversations graded on physician-written rubrics for accuracy, completeness, and communication.

Health Optimization Bench

Arcophos · 257 tasks

257 evidence-grounded tasks across eight preventive medicine subjects, graded blind by independent model families.

WHBench

academic team · 47 scenarios

Women's health scenarios graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence.

Showing top 12 of 22 indexed results. View all results.

Which rubric-graded benchmarks have results in this index?

5 as of September 28, 2026: HealthBench Professional (GPT-6 Astra (Anthropic run): highest indexed score 0.703); HealthBench Hard (Baichuan-M3: highest indexed score 0.444); HealthBench (Claude Sonnet 5.5: highest indexed score 65.4); Health Optimization Bench (Claude Fable 5: highest indexed score 70.9); WHBench (Claude Opus 4.6: highest indexed score 72.1%).

The other categories sit on the index: agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.