Health Evals

EHR-Complex: published results

Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) · ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records · index updated September 28, 2026

GPT-5.4 (high reasoning) has the highest indexed numerical score on EHR-Complex, 0.65 as of 2026-06, per EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF). Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.

Published results

#modelscoreas of
1OpenAI logoGPT-5.4 (high reasoning) OpenAI
average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
0.652026-06
2Google logoGemini 3.1 Pro Google
validation configuration
0.632026-06
3Moonshot AI logoKimi-K2.5 Moonshot AI
headline 12-model evaluation, top open-weight
0.622026-06
4AQwen3.5-397B Alibaba
headline evaluation
0.622026-06
5DDeepSeek-V3.2-Exp DeepSeek
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.592026-06
6OpenAI logoGPT-5.4 (low reasoning) OpenAI
validation configuration
0.582026-06
7DDeepSeek-V3.1 DeepSeek
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.562026-06
8AQwen3-32B-SFT Alibaba
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
0.552026-06
9AQwen3-235B Alibaba
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.532026-06
10OpenAI logoGPT-4.1 mini OpenAI
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.492026-06
11OpenAI logoGPT-4.1 OpenAI
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.472026-06
12AQwen3-14B-SFT Alibaba
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
0.452026-06
13Anthropic logoClaude Sonnet 4.6 Anthropic
validation configuration
0.362026-06
14AQwen3-32B Alibaba
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.362026-06
15OpenAI logoGPT-4o OpenAI
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.312026-06
16Google logoGemini 2.5 Pro Google
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.312026-06
17AQwen3-14B Alibaba
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.302026-06
18AQwen3-4B Alibaba
Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
0.162026-06

Scores preserve their source precision, with any scale conversion documented (independently run). The paper reports 12 base models, 2 benchmark-specific SFT variants, and 4 commercial configurations used for human validation. All 18 are shown with configuration labels. The macro-average gives equal weight to 12 intent/scope columns; it is not the micro-average over all test cases. Rechecked against the June 2026 paper; no new runs implied. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherAcademic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301)
categoryagentic and workflow benchmarks
released2026-06
size~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records
scaleexact-match accuracy, 0-1, higher better
result basisindependently run
sourceEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF)
official pagearxiv.org/abs/2606.23301
last frontier result2026-06

What is EHR-Complex?

EHR-Complex is a agentic and workflow benchmark from academic team, released 2026-06: ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records, scored on a exact-match accuracy scale. Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.

Which model leads EHR-Complex?

GPT-5.4 (high reasoning) (OpenAI) has the highest indexed numerical score on EHR-Complex at 0.65 (evaluation setups may differ), per EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF), as of 2026-06.

Where do the EHR-Complex numbers come from?

From EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) (independently run). The paper reports 12 base models, 2 benchmark-specific SFT variants, and 4 commercial configurations used for human validation. All 18 are shown with configuration labels. The macro-average gives equal weight to 12 intent/scope columns; it is not the micro-average over all test cases. Rechecked against the June 2026 paper; no new runs implied.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.