Health Evals

AAlibaba: healthcare benchmark results

22 models · 36 results · snapshot reviewed September 28, 2026

Everything the index holds for Alibaba, gathered in one place: Qwen3.5 397B A17B, Qwen 3.8 Max, Qwen 3.6 Plus, Qwen 3.7 Max, Qwen 3.5 Flash, Qwen 3 VL Plus, Qwen 3 Max Thinking, Qwen 3.8 27B, hermes + qwen-3.6-max, openai-agents + qwen-3.6-max, deepagents + qwen-3.6-max, openclaw + qwen-3.6-max, Qwen3.7 Plus, Qwen3-VL-235B-A22B, Qwen3.8 27B (xhigh), Qwen3-32B-SFT, Qwen3-235B, Qwen3-14B-SFT, Qwen3-32B, Qwen3-14B, Qwen3-4B, Qwen 3.5. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Qwen3.5 397B A17BMAST (Medical AI Superintelligence Test)57.9%5 of 8
source board: 11
Qwen3.5 397B A17BFirst, Do NOHARM (v2)61.1%14 of 17
source board: 19
Qwen3.5 397B A17B
self-reported in the Qwen3.5-397B-A17B model card
MedXpertQA (MM)70.011 of 22
Qwen3.5-397B
headline evaluation
EHR-Complex0.624 of 18
Qwen 3.8 Max
model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 2.029 pp; $0.120827/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)40.67%57 of 102
Qwen 3.8 Max
model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 1.999 pp; $0.089616/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.95%29 of 104
Qwen3.8 Max
Alibaba's own Qwen3.8 launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
MedXpertQA (MM)80.4%3 of 22
Qwen3.8 Max (0902)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index4114 of 25
source board: 77
Qwen 3.6 Plus
model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 2.017 pp; $0.015673/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)36.89%72 of 102
Qwen 3.6 Plus
model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 1.917 pp; $0.029294/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)76.96%68 of 104
Qwen3.6 Plus
Qwen-run comparison in the Qwen3.7-Plus launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
MedXpertQA (MM)68.712 of 22
Qwen3.6-PlusPhysicianBench13.7 ± 4.018 of 21
Qwen 3.7 Max
model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 2.196 pp; $0.042362/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)38.75%65 of 102
Qwen 3.7 Max
model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 1.907 pp; $0.069070/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)79.40%58 of 104
Qwen 3.5 Flash
model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 1.787 pp; $0.003934/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)33.00%81 of 102
Qwen 3.5 Flash
model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 2.09 pp; $0.004425/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)70.62%89 of 104
Qwen 3 VL Plus
model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.845 pp; $0.002519/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)31.65%88 of 102
Qwen 3 VL Plus
model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.020220/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)77.13%66 of 104
Qwen 3 Max Thinking
model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.014776/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)31.37%89 of 102
Qwen 3 Max Thinking
model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.905 pp; $0.085327/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)72.71%82 of 104
Qwen 3.8 27B
model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.971 pp; $0.050145/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)28.70%94 of 102
Qwen 3.8 27B
model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.977 pp; $0.046732/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)83.85%37 of 104
hermes + qwen-3.6-max
All Domains pass@1; PA 9.3%, UM 26.7%, CM 13.3%; submitted 2026-05-01; run date not published
CHI-Bench16.4%19 of 44
source board: 45
openai-agents + qwen-3.6-max
All Domains pass@1; PA 16.0%, UM 26.7%, CM 4.0%; submitted 2026-05-01; run date not published
CHI-Bench15.6%21 of 44
source board: 45
deepagents + qwen-3.6-max
All Domains pass@1; PA 12.0%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published
CHI-Bench9.3%33 of 44
source board: 45
openclaw + qwen-3.6-max
All Domains pass@1; PA 10.7%, UM 4.0%, CM 0.0%; submitted 2026-05-01; run date not published
CHI-Bench4.9%39 of 44
source board: 45
Qwen3.7 Plus
Alibaba's own Qwen3.7 Plus launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
MedXpertQA (MM)71.0%10 of 22
Qwen3-VL-235B-A22B
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
MedXpertQA (MM)47.6%20 of 22
Qwen3.8 27B (xhigh)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index3419 of 25
source board: 77
Qwen3-32B-SFT
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
EHR-Complex0.558 of 18
Qwen3-235B
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.539 of 18
Qwen3-14B-SFT
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
EHR-Complex0.4512 of 18
Qwen3-32B
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.3614 of 18
Qwen3-14B
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.3017 of 18
Qwen3-4B
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
EHR-Complex0.1618 of 18
Qwen 3.5
screenshot-only, task description + portal guidance
HealthAdminBench13.3%5 of 7

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Alibaba appear on?

As of September 28, 2026, Alibaba models hold 36 indexed results across 10 tracked benchmarks, through Qwen3.5 397B A17B, Qwen 3.8 Max, Qwen 3.6 Plus, Qwen 3.7 Max, Qwen 3.5 Flash, Qwen 3 VL Plus, Qwen 3 Max Thinking, Qwen 3.8 27B, hermes + qwen-3.6-max, openai-agents + qwen-3.6-max, deepagents + qwen-3.6-max, openclaw + qwen-3.6-max, Qwen3.7 Plus, Qwen3-VL-235B-A22B, Qwen3.8 27B (xhigh), Qwen3-32B-SFT, Qwen3-235B, Qwen3-14B-SFT, Qwen3-32B, Qwen3-14B, Qwen3-4B, Qwen 3.5.

Where does Alibaba have the highest indexed score?

Alibaba does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.