Health Evals

AQwen3.5 397B A17B: healthcare benchmark results

Alibaba · 4 boards · snapshot reviewed September 28, 2026

The index currently holds 4 results for Qwen3.5 397B A17B: 57.9% on MAST (Medical AI Superintelligence Test) (5 of 8 indexed rows), 61.1% on First, Do NOHARM (v2) (14 of 17 indexed rows), 70.0 on MedXpertQA (MM) (11 of 22 indexed rows), 0.62 on EHR-Complex [Qwen3.5-397B] (4 of 18 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
MAST (Medical AI Superintelligence Test)57.9%5 of 8
source board: 11
2026-08
First, Do NOHARM (v2)61.1%14 of 17
source board: 19
2026-08
MedXpertQA (MM)
self-reported in the Qwen3.5-397B-A17B model card
70.011 of 22not reported
EHR-Complex
Qwen3.5-397B
headline evaluation
0.624 of 182026-06

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "Qwen3.5-397B".

Which healthcare benchmarks is Qwen3.5 397B A17B scored on?

As of September 28, 2026, Qwen3.5 397B A17B has indexed results on 4 tracked benchmarks: MAST (Medical AI Superintelligence Test), First, Do NOHARM (v2), MedXpertQA (MM), EHR-Complex.

Which results are indexed for Qwen3.5 397B A17B?

Qwen3.5 397B A17B stands at 57.9% on MAST (Medical AI Superintelligence Test) (5 of 8 indexed rows), 61.1% on First, Do NOHARM (v2) (14 of 17 indexed rows), 70.0 on MedXpertQA (MM) (11 of 22 indexed rows), 0.62 on EHR-Complex [Qwen3.5-397B] (4 of 18 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.