MedXpertQA (MM): published results
TsinghuaC3I (Tsinghua University) · 2,000 multimodal questions (MM subset) · index updated September 28, 2026
GPT-5.6 Sol has the highest indexed numerical score on MedXpertQA (MM), 81.5 as of , per Qwen3.8-Max: A New Bar for Coding and Cowork. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.
Published results
Showing top 20 of 22 indexed results. View all results.
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol OpenAI Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 81.5 | not reported | |
| 2 | Gemini 3.1 Pro Google Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 81.3% | 2026-04 | |
| 3 | A | Qwen3.8 Max Alibaba Alibaba's own Qwen3.8 launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 80.4% | 2026-08 |
| 4 | Claude Fable 5 Anthropic Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 80.0 | not reported | |
| 5 | Muse Spark Meta Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 78.4% | 2026-04 | |
| 6 | GPT-5.4 OpenAI Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 77.1% | 2026-04 | |
| 7 | Gemini 3 Pro Google Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors | 76.0% | ||
| 8 | GPT-5.2 OpenAI Qwen-run comparison in the Qwen3.5-397B-A17B model card | 73.3 | not reported | |
| 9 | Claude Opus 4.8 Anthropic Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 71.7 | not reported | |
| 10 | A | Qwen3.7 Plus Alibaba Alibaba's own Qwen3.7 Plus launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 71.0% | 2026-05 |
| 11 | A | Qwen3.5 397B A17B Alibaba self-reported in the Qwen3.5-397B-A17B model card | 70.0 | not reported |
| 12 | A | Qwen3.6 Plus Alibaba Qwen-run comparison in the Qwen3.7-Plus launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28. | 68.7 | not reported |
| 13 | Grok 4.20 xAI Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 65.8% | 2026-04 | |
| 14 | Kimi K2.5 Moonshot AI Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column) | 65.3 | not reported | |
| 15 | Claude Opus 4.6 Anthropic Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction) | 64.8% | 2026-04 | |
| 16 | Claude Opus 4.5 Anthropic Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors | 63.6% | ||
| 17 | Gemma 4 31B Google Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | 61.3% | ||
| 18 | Gemma 4 26B A4B Google Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | 58.1% | ||
| 19 | Gemma 4 12B Google Google's Gemma 4 model card, Unified 12B; protocol not stated | 48.7% | 2026-04 | |
| 20 | A | Qwen3-VL-235B-A22B Alibaba Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors | 47.6% | |
| 21 | Gemma 4 E4B Google Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | 28.7% | ||
| 22 | Gemma 4 E2B Google Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families | 23.5% | ||
Scores preserve their source precision, with any scale conversion documented (mixed sources). Assembled table, not a single board. Rows come from several vendors' own launch documents: Meta's Muse Spark evaluation (whose values live only in a rendered table image and cover Meta's own model plus four competitors it re-ran or quoted), Alibaba's Qwen3.5, 3.7 and 3.8 posts (Alibaba's own runs of its models and of competitors), and Google's Gemma 4 model card. Protocols are not known to match across rows; each row's source line says which document it came from and who ran it. benchlm.ai presents a subset as one mirrored view while citing only Meta's methodology. Checked 2026-09-28: Meta image table visually verified from its methodology PDF; Qwen3.5 and Google model-card tables verified. Six Qwen3.7/3.8 blog rows retain their historical retrieval dates and partial confidence because those primary pages currently return no article content. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | TsinghuaC3I (Tsinghua University) |
|---|---|
| category | knowledge and exam benchmarks |
| released | 2025-01 |
| size | 2,000 multimodal questions (MM subset) |
| scale | percentage accuracy 0-100, higher better |
| result basis | mixed sources |
| source | Introducing Muse Spark: Scaling Towards Personal Superintelligence |
| official page | github.com/TsinghuaC3I/MedXpertQA |
| last frontier result | 2026-08 |
What is MedXpertQA (MM)?
MedXpertQA (MM) is a knowledge and exam benchmark from Tsinghua University, released 2025-01: 2,000 multimodal questions (MM subset), scored on a percentage accuracy 0-100 scale. Expert-level multimodal medical multiple-choice QA covering clinical images (X-ray, histology, dermatology, charts) across 17 specialties; MM subset of the 4,460-question MedXpertQA benchmark.
Which model leads MedXpertQA (MM)?
GPT-5.6 Sol (OpenAI) has the highest indexed numerical score on MedXpertQA (MM) at 81.5 (evaluation setups may differ), per Qwen3.8-Max: A New Bar for Coding and Cowork, as of null.
Where do the MedXpertQA (MM) numbers come from?
From Introducing Muse Spark: Scaling Towards Personal Superintelligence (mixed sources). Assembled table, not a single board. Rows come from several vendors' own launch documents: Meta's Muse Spark evaluation (whose values live only in a rendered table image and cover Meta's own model plus four competitors it re-ran or quoted), Alibaba's Qwen3.5, 3.7 and 3.8 posts (Alibaba's own runs of its models and of competitors), and Google's Gemma 4 model card. Protocols are not known to match across rows; each row's source line says which document it came from and who ran it. benchlm.ai presents a subset as one mirrored view while citing only Meta's methodology. Checked 2026-09-28: Meta image table visually verified from its methodology PDF; Qwen3.5 and Google model-card tables verified. Six Qwen3.7/3.8 blog rows retain their historical retrieval dates and partial confidence because those primary pages currently return no article content.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.