Health Evals

CCohere: healthcare benchmark results

1 model · 2 results · snapshot reviewed September 28, 2026

Everything the index holds for Cohere, gathered in one place: Command A+. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Command A+
model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 1.835 pp; $0.055695/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)19.72%102 of 102
Command A+
model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 3.646 pp; $0.140316/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)55.68%100 of 104

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Cohere appear on?

As of September 28, 2026, Cohere models hold 2 indexed results across 2 tracked benchmarks, through Command A+.

Where does Cohere have the highest indexed score?

Cohere does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception. The full field is on the index.