AAlibaba: healthcare benchmark results
22 models · 36 results · snapshot reviewed September 28, 2026
Everything the index holds for Alibaba, gathered in one place: Qwen3.5 397B A17B, Qwen 3.8 Max, Qwen 3.6 Plus, Qwen 3.7 Max, Qwen 3.5 Flash, Qwen 3 VL Plus, Qwen 3 Max Thinking, Qwen 3.8 27B, hermes + qwen-3.6-max, openai-agents + qwen-3.6-max, deepagents + qwen-3.6-max, openclaw + qwen-3.6-max, Qwen3.7 Plus, Qwen3-VL-235B-A22B, Qwen3.8 27B (xhigh), Qwen3-32B-SFT, Qwen3-235B, Qwen3-14B-SFT, Qwen3-32B, Qwen3-14B, Qwen3-4B, Qwen 3.5. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Qwen3.5 397B A17B | MAST (Medical AI Superintelligence Test) | 57.9% | 5 of 8 source board: 11 |
| Qwen3.5 397B A17B | First, Do NOHARM (v2) | 61.1% | 14 of 17 source board: 19 |
| Qwen3.5 397B A17B self-reported in the Qwen3.5-397B-A17B model card | MedXpertQA (MM) | 70.0 | 11 of 22 |
| Qwen3.5-397B headline evaluation | EHR-Complex | 0.62 | 4 of 18 |
| Qwen 3.8 Max model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 2.029 pp; $0.120827/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 40.67% | 57 of 102 |
| Qwen 3.8 Max model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 1.999 pp; $0.089616/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 84.95% | 29 of 104 |
| Qwen3.8 Max Alibaba's own Qwen3.8 launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | MedXpertQA (MM) | 80.4% | 3 of 22 |
| Qwen3.8 Max (0902) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 41 | 14 of 25 source board: 77 |
| Qwen 3.6 Plus model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 2.017 pp; $0.015673/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 36.89% | 72 of 102 |
| Qwen 3.6 Plus model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 1.917 pp; $0.029294/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 76.96% | 68 of 104 |
| Qwen3.6 Plus Qwen-run comparison in the Qwen3.7-Plus launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | MedXpertQA (MM) | 68.7 | 12 of 22 |
| Qwen3.6-Plus | PhysicianBench | 13.7 ± 4.0 | 18 of 21 |
| Qwen 3.7 Max model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 2.196 pp; $0.042362/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 38.75% | 65 of 102 |
| Qwen 3.7 Max model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 1.907 pp; $0.069070/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 79.40% | 58 of 104 |
| Qwen 3.5 Flash model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 1.787 pp; $0.003934/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 33.00% | 81 of 102 |
| Qwen 3.5 Flash model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 2.09 pp; $0.004425/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 70.62% | 89 of 104 |
| Qwen 3 VL Plus model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.845 pp; $0.002519/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 31.65% | 88 of 102 |
| Qwen 3 VL Plus model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.020220/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 77.13% | 66 of 104 |
| Qwen 3 Max Thinking model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.014776/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 31.37% | 89 of 102 |
| Qwen 3 Max Thinking model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.905 pp; $0.085327/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 72.71% | 82 of 104 |
| Qwen 3.8 27B model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.971 pp; $0.050145/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 28.70% | 94 of 102 |
| Qwen 3.8 27B model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.977 pp; $0.046732/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 83.85% | 37 of 104 |
| hermes + qwen-3.6-max All Domains pass@1; PA 9.3%, UM 26.7%, CM 13.3%; submitted 2026-05-01; run date not published | CHI-Bench | 16.4% | 19 of 44 source board: 45 |
| openai-agents + qwen-3.6-max All Domains pass@1; PA 16.0%, UM 26.7%, CM 4.0%; submitted 2026-05-01; run date not published | CHI-Bench | 15.6% | 21 of 44 source board: 45 |
| deepagents + qwen-3.6-max All Domains pass@1; PA 12.0%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published | CHI-Bench | 9.3% | 33 of 44 source board: 45 |
| openclaw + qwen-3.6-max All Domains pass@1; PA 10.7%, UM 4.0%, CM 0.0%; submitted 2026-05-01; run date not published | CHI-Bench | 4.9% | 39 of 44 source board: 45 |
| Qwen3.7 Plus Alibaba's own Qwen3.7 Plus launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | MedXpertQA (MM) | 71.0% | 10 of 22 |
| Qwen3-VL-235B-A22B Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors | MedXpertQA (MM) | 47.6% | 20 of 22 |
| Qwen3.8 27B (xhigh) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 34 | 19 of 25 source board: 77 |
| Qwen3-32B-SFT Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set | EHR-Complex | 0.55 | 8 of 18 |
| Qwen3-235B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.53 | 9 of 18 |
| Qwen3-14B-SFT Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set | EHR-Complex | 0.45 | 12 of 18 |
| Qwen3-32B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.36 | 14 of 18 |
| Qwen3-14B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.30 | 17 of 18 |
| Qwen3-4B Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns | EHR-Complex | 0.16 | 18 of 18 |
| Qwen 3.5 screenshot-only, task description + portal guidance | HealthAdminBench | 13.3% | 5 of 7 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Alibaba appear on?
As of September 28, 2026, Alibaba models hold 36 indexed results across 10 tracked benchmarks, through Qwen3.5 397B A17B, Qwen 3.8 Max, Qwen 3.6 Plus, Qwen 3.7 Max, Qwen 3.5 Flash, Qwen 3 VL Plus, Qwen 3 Max Thinking, Qwen 3.8 27B, hermes + qwen-3.6-max, openai-agents + qwen-3.6-max, deepagents + qwen-3.6-max, openclaw + qwen-3.6-max, Qwen3.7 Plus, Qwen3-VL-235B-A22B, Qwen3.8 27B (xhigh), Qwen3-32B-SFT, Qwen3-235B, Qwen3-14B-SFT, Qwen3-32B, Qwen3-14B, Qwen3-4B, Qwen 3.5.
Where does Alibaba have the highest indexed score?
Alibaba does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.