Health Evals

IInception: healthcare benchmark results

1 model · 2 results · snapshot reviewed September 28, 2026

Everything the index holds for Inception, gathered in one place: Mercury 2.5. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Mercury 2.5
model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.953 pp; $0.004535/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)31.33%90 of 102
Mercury 2.5
model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.095 pp; $0.004476/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)55.09%101 of 104

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Inception appear on?

As of September 28, 2026, Inception models hold 2 indexed results across 2 tracked benchmarks, through Mercury 2.5.

Where does Inception have the highest indexed score?

Inception does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Cohere. The full field is on the index.