BBaichuan: healthcare benchmark results
1 model · 2 results · snapshot reviewed September 28, 2026
Everything the index holds for Baichuan, gathered in one place: Baichuan-M3. The lab has the highest indexed score on HealthBench Hard. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Baichuan-M3 Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol. | HealthBench Hard | 0.444 | 1 of 20 |
| Baichuan-M3 self-run in Baichuan-M3 paper (arXiv 2602.06570) | HealthBench | 65.1 | 2 of 26 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Baichuan appear on?
As of September 28, 2026, Baichuan models hold 2 indexed results across 2 tracked benchmarks, through Baichuan-M3.
Where does Baichuan have the highest indexed score?
Baichuan models have the highest indexed score on HealthBench Hard (Baichuan-M3, 0.444).
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Tencent, Inception, Cohere. The full field is on the index.