First, Do NOHARM (v2): published results
Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE · 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options · index updated September 28, 2026
LiSA 2.5 has the highest indexed numerical score on First, Do NOHARM (v2), 86.2 as of 2026-09, per ARISE MAST technical leaderboard. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | A | LiSA 2.5 AMBOSS First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published. | 86.2 | 2026-09 |
| 2 | D | Doximity Ask 6.1 Doximity First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published. | 84.5 | 2026-09 |
| 3 | O | OpenEvidence OpenEvidence First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published. | 80.0 | 2026-09 |
| 4 | Muse Spark 1.1 Meta | 79.7% | 2026-08 | |
| 5 | GH | Glass 5.6 Max Glass Health First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published. | 79.7 | 2026-09 |
| 6 | Claude Opus 5 Anthropic v2 run on ARISE; 19 models on the board | 74.6% | 2026-08 | |
| 7 | Kimi K3 Moonshot AI | 74.0% | 2026-08 | |
| 8 | GPT-5.6 Sol OpenAI | 70.1% | 2026-08 | |
| 9 | GPT-5.5 OpenAI | 70.0% | 2026-08 | |
| 10 | GPT-5 OpenAI from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking | 68.6% | 2026-08 | |
| 11 | Claude Fable 5 Anthropic | 65.0% | 2026-08 | |
| 12 | Gemini 3.1 Pro Google | 62.6% | 2026-08 | |
| 13 | Gemini 2.5 Pro Google | 61.9% | 2026-08 | |
| 14 | A | Qwen3.5 397B A17B Alibaba | 61.1% | 2026-08 |
| 15 | Kimi K2.6 Moonshot AI | 59.1% | 2026-08 | |
| 16 | Z | GLM 5.1 Z.ai First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. Open-weight model as listed by the evaluator. Observed September 28, 2026; the individual run date is not published. | 57.9 | 2026-09 |
| 17 | D | DeepSeek R1 DeepSeek | 55.8% | 2026-08 |
Scores preserve their source precision, with any scale conversion documented (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases. Official technical leaderboard checked September 28, 2026; retained August dates on existing rows; newly indexed rows have no claimed measurement date. Includes labeled RAG clinical systems and base models, which have different tool access. The board is in preview. Seventeen of nineteen overall rows are indexed. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE |
|---|---|
| category | safety benchmarks |
| released | 2025-12 |
| size | 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options |
| scale | percentage safety score, higher better |
| result basis | official leaderboard |
| source | MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) |
| paper | arxiv.org/abs/2512.01241 |
| last frontier result | 2026-08 |
What is First, Do NOHARM (v2)?
First, Do NOHARM (v2) is a safety benchmark from Stanford/Harvard consortium, released 2025-12: 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options, scored on a percentage safety score scale. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.
Which model leads First, Do NOHARM (v2)?
LiSA 2.5 (AMBOSS) has the highest indexed numerical score on First, Do NOHARM (v2) at 86.2 (evaluation setups may differ), per ARISE MAST technical leaderboard, as of 2026-09.
Where do the First, Do NOHARM (v2) numbers come from?
From MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases. Official technical leaderboard checked September 28, 2026; retained August dates on existing rows; newly indexed rows have no claimed measurement date. Includes labeled RAG clinical systems and base models, which have different tool access. The board is in preview. Seventeen of nineteen overall rows are indexed.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.