Health Evals

ZZ.ai: healthcare benchmark results

1 model · 3 results · snapshot reviewed September 28, 2026

Everything the index holds for Z.ai, gathered in one place: GLM 5.1. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
GLM 5.1
First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. Open-weight model as listed by the evaluator. Observed September 28, 2026; the individual run date is not published.
First, Do NOHARM (v2)57.916 of 17
source board: 19
GLM 5.1
model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.024196/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.60%48 of 102
GLM 5.1
model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.064 pp; $0.023717/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)72.27%84 of 104

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Z.ai appear on?

As of September 28, 2026, Z.ai models hold 3 indexed results across 3 tracked benchmarks, through GLM 5.1.

Where does Z.ai have the highest indexed score?

Z.ai does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.