DDeepSeek: healthcare benchmark results
13 models · 20 results · snapshot reviewed September 28, 2026
Everything the index holds for DeepSeek, gathered in one place: DeepSeek R1, DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4, openai-agents + deepseek-v4-pro, hermes + deepseek-v4-pro, openclaw + deepseek-v4-pro, deepagents + deepseek-v4-pro, DeepSeek V4-Pro, DeepSeek-V3.2-Exp, DeepSeek-V3.1, DeepSeek V3.2. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| DeepSeek R1 | MedHELM | 0.485 | 7 of 10 source board: 11 |
| DeepSeek R1 | First, Do NOHARM (v2) | 55.8% | 17 of 17 source board: 19 |
| DeepSeek-R1 95% bootstrap CI 50.5–55.3; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 52.9% | 10 of 22 |
| DeepSeek V4.1 Flash model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 2.042 pp; $0.012450/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 41.17% | 53 of 102 |
| DeepSeek V4.1 Flash model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 1.918 pp; $0.015397/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 85.50% | 23 of 104 |
| DeepSeek V4.1 Flash (Reasoning, Max Effort) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 41 | 16 of 25 source board: 77 |
| DeepSeek V4 Pro 0813 model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.16 pp; $0.061060/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 42.47% | 46 of 102 |
| DeepSeek V4 Pro 0813 model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.004 pp; $0.041127/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 80.17% | 54 of 104 |
| DeepSeek V4 Flash 0731 model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 2.15 pp; $0.019659/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 41.41% | 49 of 102 |
| DeepSeek V4 Flash 0731 model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 1.973 pp; $0.014247/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 80.36% | 53 of 104 |
| DeepSeek V4 model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.122 pp; $0.060710/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 40.45% | 60 of 102 |
| DeepSeek V4 model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.002 pp; $0.053954/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 75.14% | 76 of 104 |
| openai-agents + deepseek-v4-pro All Domains pass@1; PA 10.7%, UM 28.0%, CM 4.0%; submitted 2026-05-01; run date not published | CHI-Bench | 14.2% | 24 of 44 source board: 45 |
| hermes + deepseek-v4-pro All Domains pass@1; PA 8.0%, UM 25.3%, CM 8.0%; submitted 2026-05-01; run date not published | CHI-Bench | 13.8% | 25 of 44 source board: 45 |
| openclaw + deepseek-v4-pro All Domains pass@1; PA 14.7%, UM 12.0%, CM 6.7%; submitted 2026-05-01; run date not published | CHI-Bench | 11.1% | 29 of 44 source board: 45 |
| deepagents + deepseek-v4-pro All Domains pass@1; PA 14.7%, UM 10.7%, CM 6.7%; submitted 2026-05-01; run date not published | CHI-Bench | 10.7% | 31 of 44 source board: 45 |
| DeepSeek V4-Pro Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported | PhysicianBench | 18.7 ± 2.9 | 15 of 21 |
| DeepSeek-V3.2-Exp Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.59 | 5 of 18 |
| DeepSeek-V3.1 Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.56 | 7 of 18 |
| DeepSeek V3.2 95% bootstrap CI 58.6–63.9; 3 runs; temperature 0; zero-shot, closed-book | WHBench | 61.3% | 6 of 22 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does DeepSeek appear on?
As of September 28, 2026, DeepSeek models hold 20 indexed results across 9 tracked benchmarks, through DeepSeek R1, DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4, openai-agents + deepseek-v4-pro, hermes + deepseek-v4-pro, openclaw + deepseek-v4-pro, deepagents + deepseek-v4-pro, DeepSeek V4-Pro, DeepSeek-V3.2-Exp, DeepSeek-V3.1, DeepSeek V3.2.
Where does DeepSeek have the highest indexed score?
DeepSeek does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.