Knowledge and exam benchmarks
1 tracked · snapshot reviewed September 28, 2026
These evaluations test medical knowledge and reasoning through questions with reference answers. Some include clinical images and specialist-level material. Performance on a question set does not directly measure a model's ability to deliver care in practice.
MedXpertQA (MM)
Tsinghua University · 2,000 multimodal questionsExpert-level multiple-choice questions over clinical images across 17 specialties.
Which knowledge and exam benchmarks have results in this index?
1 as of September 28, 2026: MedXpertQA (MM) (GPT-5.6 Sol: highest indexed score 81.5).
The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, composite indices.