Health Evals

MedXpertQA (MM): published results

TsinghuaC3I (Tsinghua University) · 2,000 multimodal questions (MM subset) · index updated September 28, 2026

GPT-5.6 Sol has the highest indexed numerical score on MedXpertQA (MM), 81.5 as of , per Qwen3.8-Max: A New Bar for Coding and Cowork. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.

Published results

#modelscoreas of
1OpenAI logoGPT-5.6 Sol OpenAI
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
81.5not reported
2Google logoGemini 3.1 Pro Google
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
81.3%2026-04
3AQwen3.8 Max Alibaba
Alibaba's own Qwen3.8 launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
80.4%2026-08
4Anthropic logoClaude Fable 5 Anthropic
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
80.0not reported
5Meta logoMuse Spark Meta
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
78.4%2026-04
6OpenAI logoGPT-5.4 OpenAI
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
77.1%2026-04
7Google logoGemini 3 Pro Google
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
76.0%
8OpenAI logoGPT-5.2 OpenAI
Qwen-run comparison in the Qwen3.5-397B-A17B model card
73.3not reported
9Anthropic logoClaude Opus 4.8 Anthropic
Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
71.7not reported
10AQwen3.7 Plus Alibaba
Alibaba's own Qwen3.7 Plus launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
71.0%2026-05
11AQwen3.5 397B A17B Alibaba
self-reported in the Qwen3.5-397B-A17B model card
70.0not reported
12AQwen3.6 Plus Alibaba
Qwen-run comparison in the Qwen3.7-Plus launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
68.7not reported
13xAI logoGrok 4.20 xAI
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
65.8%2026-04
14Moonshot AI logoKimi K2.5 Moonshot AI
Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)
65.3not reported
15Anthropic logoClaude Opus 4.6 Anthropic
Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
64.8%2026-04
16Anthropic logoClaude Opus 4.5 Anthropic
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
63.6%
17Google logoGemma 4 31B Google
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
61.3%
18Google logoGemma 4 26B A4B Google
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
58.1%
19Google logoGemma 4 12B Google
Google's Gemma 4 model card, Unified 12B; protocol not stated
48.7%2026-04
20AQwen3-VL-235B-A22B Alibaba
Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
47.6%
21Google logoGemma 4 E4B Google
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
28.7%
22Google logoGemma 4 E2B Google
Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
23.5%

Scores preserve their source precision, with any scale conversion documented (mixed sources). Assembled table, not a single board. Rows come from several vendors' own launch documents: Meta's Muse Spark evaluation (whose values live only in a rendered table image and cover Meta's own model plus four competitors it re-ran or quoted), Alibaba's Qwen3.5, 3.7 and 3.8 posts (Alibaba's own runs of its models and of competitors), and Google's Gemma 4 model card. Protocols are not known to match across rows; each row's source line says which document it came from and who ran it. benchlm.ai presents a subset as one mirrored view while citing only Meta's methodology. Checked 2026-09-28: Meta image table visually verified from its methodology PDF; Qwen3.5 and Google model-card tables verified. Six Qwen3.7/3.8 blog rows retain their historical retrieval dates and partial confidence because those primary pages currently return no article content. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherTsinghuaC3I (Tsinghua University)
categoryknowledge and exam benchmarks
released2025-01
size2,000 multimodal questions (MM subset)
scalepercentage accuracy 0-100, higher better
result basismixed sources
sourceIntroducing Muse Spark: Scaling Towards Personal Superintelligence
official pagegithub.com/TsinghuaC3I/MedXpertQA
last frontier result2026-08

What is MedXpertQA (MM)?

MedXpertQA (MM) is a knowledge and exam benchmark from Tsinghua University, released 2025-01: 2,000 multimodal questions (MM subset), scored on a percentage accuracy 0-100 scale. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.

Which model leads MedXpertQA (MM)?

GPT-5.6 Sol (OpenAI) has the highest indexed numerical score on MedXpertQA (MM) at 81.5 (evaluation setups may differ), per Qwen3.8-Max: A New Bar for Coding and Cowork, as of null.

Where do the MedXpertQA (MM) numbers come from?

From Introducing Muse Spark: Scaling Towards Personal Superintelligence (mixed sources). Assembled table, not a single board. Rows come from several vendors' own launch documents: Meta's Muse Spark evaluation (whose values live only in a rendered table image and cover Meta's own model plus four competitors it re-ran or quoted), Alibaba's Qwen3.5, 3.7 and 3.8 posts (Alibaba's own runs of its models and of competitors), and Google's Gemma 4 model card. Protocols are not known to match across rows; each row's source line says which document it came from and who ran it. benchlm.ai presents a subset as one mirrored view while citing only Meta's methodology. Checked 2026-09-28: Meta image table visually verified from its methodology PDF; Qwen3.5 and Google model-card tables verified. Six Qwen3.7/3.8 blog rows retain their historical retrieval dates and partial confidence because those primary pages currently return no article content.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.