Health Evals

DDeepSeek: healthcare benchmark results

13 models · 20 results · snapshot reviewed September 28, 2026

Everything the index holds for DeepSeek, gathered in one place: DeepSeek R1, DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4, openai-agents + deepseek-v4-pro, hermes + deepseek-v4-pro, openclaw + deepseek-v4-pro, deepagents + deepseek-v4-pro, DeepSeek V4-Pro, DeepSeek-V3.2-Exp, DeepSeek-V3.1, DeepSeek V3.2. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
DeepSeek R1MedHELM0.4857 of 10
source board: 11
DeepSeek R1First, Do NOHARM (v2)55.8%17 of 17
source board: 19
DeepSeek-R1
95% bootstrap CI 50.5–55.3; 3 runs; temperature 0; zero-shot, closed-book
WHBench52.9%10 of 22
DeepSeek V4.1 Flash
model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 2.042 pp; $0.012450/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.17%53 of 102
DeepSeek V4.1 Flash
model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 1.918 pp; $0.015397/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.50%23 of 104
DeepSeek V4.1 Flash (Reasoning, Max Effort)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index4116 of 25
source board: 77
DeepSeek V4 Pro 0813
model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.16 pp; $0.061060/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)42.47%46 of 102
DeepSeek V4 Pro 0813
model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.004 pp; $0.041127/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)80.17%54 of 104
DeepSeek V4 Flash 0731
model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 2.15 pp; $0.019659/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.41%49 of 102
DeepSeek V4 Flash 0731
model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 1.973 pp; $0.014247/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)80.36%53 of 104
DeepSeek V4
model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.122 pp; $0.060710/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.45%60 of 102
DeepSeek V4
model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.002 pp; $0.053954/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)75.14%76 of 104
openai-agents + deepseek-v4-pro
All Domains pass@1; PA 10.7%, UM 28.0%, CM 4.0%; submitted 2026-05-01; run date not published
CHI-Bench14.2%24 of 44
source board: 45
hermes + deepseek-v4-pro
All Domains pass@1; PA 8.0%, UM 25.3%, CM 8.0%; submitted 2026-05-01; run date not published
CHI-Bench13.8%25 of 44
source board: 45
openclaw + deepseek-v4-pro
All Domains pass@1; PA 14.7%, UM 12.0%, CM 6.7%; submitted 2026-05-01; run date not published
CHI-Bench11.1%29 of 44
source board: 45
deepagents + deepseek-v4-pro
All Domains pass@1; PA 14.7%, UM 10.7%, CM 6.7%; submitted 2026-05-01; run date not published
CHI-Bench10.7%31 of 44
source board: 45
DeepSeek V4-Pro
Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
PhysicianBench18.7 ± 2.915 of 21
DeepSeek-V3.2-Exp
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.595 of 18
DeepSeek-V3.1
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.567 of 18
DeepSeek V3.2
95% bootstrap CI 58.6–63.9; 3 runs; temperature 0; zero-shot, closed-book
WHBench61.3%6 of 22

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does DeepSeek appear on?

As of September 28, 2026, DeepSeek models hold 20 indexed results across 9 tracked benchmarks, through DeepSeek R1, DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4, openai-agents + deepseek-v4-pro, hermes + deepseek-v4-pro, openclaw + deepseek-v4-pro, deepagents + deepseek-v4-pro, DeepSeek V4-Pro, DeepSeek-V3.2-Exp, DeepSeek-V3.1, DeepSeek V3.2.

Where does DeepSeek have the highest indexed score?

DeepSeek does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.