Safety benchmarks
1 tracked · snapshot reviewed September 28, 2026
These benchmarks assess potential harm in medical responses or recommendations. The definition of harm, severity weighting, and evaluation cases determine what a score means. A benchmark result alone does not establish safety for clinical use.
First, Do NOHARM (v2)
Stanford/Harvard consortium · 1,100 consultation casesHow often, and how severely, model consultation recommendations contain potentially harmful errors.
Showing top 12 of 17 indexed results. View all results.
Which safety benchmarks have results in this index?
1 as of September 28, 2026: First, Do NOHARM (v2) (LiSA 2.5: highest indexed score 86.2).
The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, knowledge and exam benchmarks, composite indices.