EHR-Complex: published results
Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) · ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records · index updated September 28, 2026
GPT-5.4 (high reasoning) has the highest indexed numerical score on EHR-Complex, 0.65 as of 2026-06, per EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF). Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.4 (high reasoning) OpenAI average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table via paper, arxiv.org | 0.65 | 2026-06 | |
| 2 | Gemini 3.1 Pro Google validation configuration via paper, arxiv.org | 0.63 | 2026-06 | |
| 3 | Kimi-K2.5 Moonshot AI headline 12-model evaluation, top open-weight via paper, arxiv.org | 0.62 | 2026-06 | |
| 4 | A | Qwen3.5-397B Alibaba headline evaluation via paper, arxiv.org | 0.62 | 2026-06 |
| 5 | D | DeepSeek-V3.2-Exp DeepSeek Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.59 | 2026-06 |
| 6 | GPT-5.4 (low reasoning) OpenAI validation configuration via paper, arxiv.org | 0.58 | 2026-06 | |
| 7 | D | DeepSeek-V3.1 DeepSeek Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.56 | 2026-06 |
| 8 | A | Qwen3-32B-SFT Alibaba Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set via paper, arxiv.org | 0.55 | 2026-06 |
| 9 | A | Qwen3-235B Alibaba Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.53 | 2026-06 |
| 10 | GPT-4.1 mini OpenAI Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.49 | 2026-06 | |
| 11 | GPT-4.1 OpenAI Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.47 | 2026-06 | |
| 12 | A | Qwen3-14B-SFT Alibaba Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set via paper, arxiv.org | 0.45 | 2026-06 |
| 13 | Claude Sonnet 4.6 Anthropic validation configuration via paper, arxiv.org | 0.36 | 2026-06 | |
| 14 | A | Qwen3-32B Alibaba Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.36 | 2026-06 |
| 15 | GPT-4o OpenAI Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.31 | 2026-06 | |
| 16 | Gemini 2.5 Pro Google Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.31 | 2026-06 | |
| 17 | A | Qwen3-14B Alibaba Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.30 | 2026-06 |
| 18 | A | Qwen3-4B Alibaba Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns via paper, arxiv.org | 0.16 | 2026-06 |
Scores preserve their source precision, with any scale conversion documented (independently run). The paper reports 12 base models, 2 benchmark-specific SFT variants, and 4 commercial configurations used for human validation. All 18 are shown with configuration labels. The macro-average gives equal weight to 12 intent/scope columns; it is not the micro-average over all test cases. Rechecked against the June 2026 paper; no new runs implied. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Academic team (Qiao et al., Ant Group-affiliated; arXiv 2606.23301) |
|---|---|
| category | agentic and workflow benchmarks |
| released | 2026-06 |
| size | ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records |
| scale | exact-match accuracy, 0-1, higher better |
| result basis | independently run |
| source | EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) |
| official page | arxiv.org/abs/2606.23301 |
| last frontier result | 2026-06 |
What is EHR-Complex?
EHR-Complex is a agentic and workflow benchmark from academic team, released 2026-06: ~52,000 tasks (3,915-task test set) over 365K patients, 31 tables, 500M+ records, scored on a exact-match accuracy scale. Agentic clinical reasoning over MIMIC-IV EHR databases via SQL and Python across six clinical intents, at patient and population level with temporal evidence paths.
Which model leads EHR-Complex?
GPT-5.4 (high reasoning) (OpenAI) has the highest indexed numerical score on EHR-Complex at 0.65 (evaluation setups may differ), per EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF), as of 2026-06.
Where do the EHR-Complex numbers come from?
From EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) (independently run). The paper reports 12 base models, 2 benchmark-specific SFT variants, and 4 commercial configurations used for human validation. All 18 are shown with configuration labels. The macro-average gives equal weight to 12 intent/scope columns; it is not the micro-average over all test cases. Rechecked against the June 2026 paper; no new runs implied.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.