Health Evals

First, Do NOHARM (v2): published results

Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE · 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options · index updated September 28, 2026

LiSA 2.5 has the highest indexed numerical score on First, Do NOHARM (v2), 86.2 as of 2026-09, per ARISE MAST technical leaderboard. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.

Published results

#modelscoreas of
1ALiSA 2.5 AMBOSS
First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
86.22026-09
2DDoximity Ask 6.1 Doximity
First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
84.52026-09
3OOpenEvidence OpenEvidence
First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
80.02026-09
4Meta logoMuse Spark 1.1 Meta79.7%2026-08
5GHGlass 5.6 Max Glass Health
First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
79.72026-09
6Anthropic logoClaude Opus 5 Anthropic
v2 run on ARISE; 19 models on the board
74.6%2026-08
7Moonshot AI logoKimi K3 Moonshot AI74.0%2026-08
8OpenAI logoGPT-5.6 Sol OpenAI70.1%2026-08
9OpenAI logoGPT-5.5 OpenAI70.0%2026-08
10OpenAI logoGPT-5 OpenAI
from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
68.6%2026-08
11Anthropic logoClaude Fable 5 Anthropic65.0%2026-08
12Google logoGemini 3.1 Pro Google62.6%2026-08
13Google logoGemini 2.5 Pro Google61.9%2026-08
14AQwen3.5 397B A17B Alibaba61.1%2026-08
15Moonshot AI logoKimi K2.6 Moonshot AI59.1%2026-08
16ZGLM 5.1 Z.ai
First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. Open-weight model as listed by the evaluator. Observed September 28, 2026; the individual run date is not published.
57.92026-09
17DDeepSeek R1 DeepSeek55.8%2026-08

Scores preserve their source precision, with any scale conversion documented (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases. Official technical leaderboard checked September 28, 2026; retained August dates on existing rows; newly indexed rows have no claimed measurement date. Includes labeled RAG clinical systems and base models, which have different tool access. The board is in preview. Seventeen of nineteen overall rows are indexed. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherStanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE
categorysafety benchmarks
released2025-12
size1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options
scalepercentage safety score, higher better
result basisofficial leaderboard
sourceMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results)
paperarxiv.org/abs/2512.01241
last frontier result2026-08

What is First, Do NOHARM (v2)?

First, Do NOHARM (v2) is a safety benchmark from Stanford/Harvard consortium, released 2025-12: 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options, scored on a percentage safety score scale. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.

Which model leads First, Do NOHARM (v2)?

LiSA 2.5 (AMBOSS) has the highest indexed numerical score on First, Do NOHARM (v2) at 86.2 (evaluation setups may differ), per ARISE MAST technical leaderboard, as of 2026-09.

Where do the First, Do NOHARM (v2) numbers come from?

From MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases. Official technical leaderboard checked September 28, 2026; retained August dates on existing rows; newly indexed rows have no claimed measurement date. Includes labeled RAG clinical systems and base models, which have different tool access. The board is in preview. Seventeen of nineteen overall rows are indexed.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.