Rubric-graded benchmarks
5 tracked · snapshot reviewed September 28, 2026
These benchmarks evaluate free-text medical answers against criteria written by physicians or domain experts. Rubrics can assess accuracy, completeness, communication, and safety. Each board has its own questions, grading process, and score scale; consult the evaluation setup before comparing results.
HealthBench Professional
OpenAI · 525 tasks525 clinician tasks with physician-written rubrics; results retain evaluator, grader and length-adjustment details.
HealthBench Hard
OpenAI · 1,000 conversationsThe 1,000 HealthBench conversations frontier models failed most, still graded on the original physician rubrics.
Showing top 12 of 20 indexed results. View all results.
HealthBench
OpenAI · 5,000 conversationsRealistic multi-turn health conversations graded on physician-written rubrics for accuracy, completeness, and communication.
Health Optimization Bench
Arcophos · 257 tasks257 evidence-grounded tasks across eight preventive medicine subjects, graded blind by independent model families.
Showing top 12 of 16 indexed results. View all results.
WHBench
academic team · 47 scenariosWomen's health scenarios graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence.
Showing top 12 of 22 indexed results. View all results.
Which rubric-graded benchmarks have results in this index?
5 as of September 28, 2026: HealthBench Professional (GPT-6 Astra (Anthropic run): highest indexed score 0.703); HealthBench Hard (Baichuan-M3: highest indexed score 0.444); HealthBench (Claude Sonnet 5.5: highest indexed score 65.4); Health Optimization Bench (Claude Fable 5: highest indexed score 70.9); WHBench (Claude Opus 4.6: highest indexed score 72.1%).
The other categories sit on the index: agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.