Sources
For the featured medication-safety benchmark, see MedPIC-Bench sources and verification. The collection below documents the wider healthcare benchmark index.
503 of 503 rows documented · 49 documents, 46 first-party · snapshot reviewed September 28, 2026
This page records the published citations associated with each indexed score, including review status and any remaining verification gaps. A row records where its score was read: the publication itself, the passage or table it sits in when that has been captured, who published it, when the index retrieved it, and how far the reading has been checked. First-party documents rank highest, meaning a vendor's system card for a vendor-reported number or the maintainers' own board for a leaderboard number; independent evaluators are primary sources for their own runs. Sister-site compilations and secondary citations are labeled where the original result remains unchecked. Rows whose document has not been located yet say so instead of disappearing.
Better citations do not make the boards comparable. Each benchmark below keeps its own scale, grader and task set, so a score on one section says nothing about a score on the next, and no number on this page should be lined up against a number from another section. The comparability rules are on the methodology page.
Reading the confidence field
- verified: the number was read from the document at the locator given
- partial: a citation is identified, but the original result or exact passage is not fully checked
- unverified: no document has been checked for this number yet
- disputed: documents on file disagree about this number
HealthBench Professional
29 rows · 0 to 1Board source: Published model evaluation reports · official page: arxiv.org · full board: healthbenchprofessional.com
- GPT-6 Astra (Anthropic run) OpenAI0.703as of 2026-09Anthropic reproduction of GPT-6 Astra through the public API; max effort; no system prompt; Claude Opus 4.8 grader; length-adjusted 70.3%, raw 74.0%. Different grader/protocol from OpenAI’s own 64.7% report.Claude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- Section 8.15, pp. 137–138; text above Figure 8.15.B: max-effort Astra 70.3 adjusted and 74.0 raw
- reported by
- third-party run
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5.5 Anthropic0.692as of 2026-09Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.Claude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- Section 8.15.2; pp. 137–139, Figure 8.15.B (max effort)
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5.5 Anthropic0.656as of 2026-09Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. Five-trial average; refusal fallback to Claude Opus 5.Claude Opus 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- Section 8.15.2; pp. 213–214, Figures 8.15.1.A and 8.15.2.A
- reported by
- vendor-reported
- published
- 2026-09-22
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Astra OpenAI0.647as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 68.2%, mean answer 3,185 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench Professional length-adjusted row, GPT-6 Astra column; value 64.7 (68.2, 3185)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5 Anthropic0.633as of 2026-09Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 68.9%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.Claude Fable 5.1 and Claude Mythos 5.1 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 199, Figure 8.17.2.A; Claude Fable 5, adjusted bar 63.3%; also p. 167 Table 8.1.A
- reported by
- vendor-reported
- published
- 2026-09-01
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 Anthropic0.621as of 2026-09length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).Claude Fable 5.1 and Claude Mythos 5.1 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)
- reported by
- vendor-reported
- published
- 2026-09-01
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Sol OpenAI0.608as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 59.5%, mean answer 1,573 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench Professional length-adjusted row, GPT-6 Sol column; value 60.8 (59.5, 1573)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Luna OpenAI0.608as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench Professional length-adjusted row, GPT-6 Luna column; value 60.8 (61.2, 2119)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol OpenAI0.605as of 2026-06length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 60.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL) - Claude Opus 5 Anthropic0.598as of 2026-07length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).System Card: Claude Opus 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
- reported by
- vendor-reported
- published
- 2026-07-24
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 59.8%, p. 167, Table 8.1.A, column 'Claude Opus 5'; raw 73.4% also reprinted on p. 199, sec. 8.17.2) - Muse Spark 1.1 Meta0.593as of 2026-07length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)Muse Spark 1.1 Evaluation Report · model card · first-party
- publisher
- Meta
- locator
- p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5 Anthropic0.578as of 2026-06length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).System Card: Claude Sonnet 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-28
- confidence
- verified
also reported in System Card: Claude Opus 5 (system card, 57.8%, p. 189, section 8.15.2, Figure 8.15.2.A) - GPT-5.6 Terra OpenAI0.577as of 2026-06length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA) - Claude Opus 4.8 Anthropic0.574as of 2026-06Length-adjusted; Anthropic June evaluation, adaptive max effort, Claude Opus 4.8 grader, five-trial average, no tools or custom system prompt. Earlier May report was 55.8% using Claude Sonnet 4.6 as grader; the grader changed.Claude Sonnet 5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 139, Figure 8.12.2.A; Opus 4.8 bar 57.4%
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.7 xAI0.567as of 2026-09SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.Introducing Grok 4.7 · launch post · first-party
- publisher
- SpaceXAI
- locator
- Model Improvements table, Clinical reasoning / HealthBench Professional row, Grok 4.7 xhigh column
- reported by
- vendor-reported
- published
- 2026-09-21
- retrieved
- 2026-09-28
- confidence
- partial (partial review)
- GPT-5.6 Luna OpenAI0.557as of 2026-06length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 55.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA) - Muse Spark Meta0.541as of 2026-07length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44Muse Spark 1.1 Evaluation Report · model card · first-party
- publisher
- Meta
- locator
- p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol (August) OpenAI0.540as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.7 Anthropic0.519as of 2026-05length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 cardSystem Card: Claude Opus 4.8 · system card · first-party
- publisher
- Anthropic
- locator
- p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
- reported by
- vendor-reported
- published
- 2026-05-28
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.5 OpenAI0.518as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 51.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5) - Grok 4.6 xAI0.485as of 2026-09SpaceXAI vendor report at high effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.Introducing Grok 4.7 · launch post · first-party
- publisher
- SpaceXAI
- locator
- Model Improvements table, Clinical reasoning / HealthBench Professional row, Grok 4.6 high column
- reported by
- vendor-reported
- published
- 2026-09-21
- retrieved
- 2026-09-28
- confidence
- partial (partial review)
- GPT-5.4 OpenAI0.481as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 48.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4) - GPT-5 OpenAI0.462as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 46.2, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5) - GPT-5.2 OpenAI0.459as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 45.9, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2) - Claude Sonnet 4.6 Anthropic0.442as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)System Card: Claude Sonnet 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-28
- confidence
- verified
also reported in System Card: Claude Opus 4.8 (system card, 41.7%, conflicting, p. 228, section 8.14.1) - GPT-5.6 Luna (August) OpenAI0.441as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.1 OpenAI0.396as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 39.6, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1) - GPT-5.5 Instant OpenAI0.384as of 2026-05length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.GPT-5.5 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 38.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant) - MAI-Thinking-1 Microsoft0.350as of 2026-08length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.MAI-Thinking-1: Building a Hill-Climbing Machine · model card · first-party
- publisher
- Microsoft AI
- locator
- p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.
- reported by
- vendor-reported
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
HealthBench Hard
20 rows · 0 to 1Board source: healthbenchhard.ai · paper: arxiv.org · full board: healthbenchhard.ai
- Baichuan-M3 Baichuan0.444as of 2026-02Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.Baichuan-M3 Technical Report · paper · first-party
- publisher
- Baichuan
- locator
- p. 23 (PDF page index 22), section 4.2.1 and Figure 7; HealthBench Hard
- reported by
- vendor-reported
- published
- 2026-02-06
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark Meta0.428as of 2026-04raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.
- reported by
- vendor-reported
- published
- 2026-04-08
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Muse Spark Eval Methodology (model card, 42.8, p. 5 results table image, row HealthBench Hard; protocol p. 2) - GPT-5.2-High (Baichuan run) OpenAI0.420as of 2026-02Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.Baichuan-M3 Technical Report · paper · first-party
- publisher
- Baichuan
- locator
- p. 23 (PDF page index 22), section 4.2.1 and Figure 7; HealthBench Hard
- reported by
- third-party run
- published
- 2026-02-06
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Astra OpenAI0.366as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 34.2%, mean answer 1,697 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench Hard length-adjusted row, GPT-6 Astra column; value 36.6 (34.2, 1697)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5 OpenAI0.347as of 2026-06length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 34.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5); GPT-5.5 System Card (system card, 34.7, Section 5 Health, Table 7, column GPT-5); GPT-5 System Card (system card, 46.2, conflicting, p. 18, section 3.10 Health, Figure 6 (HealthBench Hard, raw score %)) - GPT-5.2 OpenAI0.343as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 34.3, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2) - GPT-5.6 Sol OpenAI0.331as of 2026-06length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 33.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL) - GPT-5.6 Terra OpenAI0.327as of 2026-06length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 32.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA) - GPT-5.6 Luna OpenAI0.320as of 2026-06length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 32.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA) - GPT-5.5 OpenAI0.315as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 31.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5) - GPT-5.6 Sol (August) OpenAI0.314as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Luna OpenAI0.314as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench Hard length-adjusted row, GPT-6 Luna column; value 31.4 (25.4, 1241)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Sol OpenAI0.301as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 22.1%, mean answer 974 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench Hard length-adjusted row, GPT-6 Sol column; value 30.1 (22.1, 974)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT OSS 120B OpenAI0.300as of 2025-08raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.4 OpenAI0.291as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 29.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4) - GPT-5.6 Luna (August) OpenAI0.287as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.3 Chat OpenAI0.259as of 2026-03raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.GPT-5.3 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 3: HealthBench, row Hard, column GPT-5.3-INSTANT
- reported by
- vendor-reported
- published
- 2026-03-02
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.5 Instant System Card (system card, 20.2, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT); GPT-5.6 - August Updates (system card addendum) (system card, 20.2, conflicting, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.3 Instant) - GPT-5.1 OpenAI0.254as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 25.4, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1) - GPT-5.5 Instant OpenAI0.229as of 2026-05length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.GPT-5.5 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 22.9, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant) - GPT OSS 20B OpenAI0.108as of 2025-08raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-28
- confidence
- verified
HealthBench
26 rows · 0-100 rubric-point percentage (some sites display 0-1)Board source: Published model evaluation reports · official page: openai.com
- Claude Sonnet 5.5 Anthropic65.4as of 2026-09Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 69.4%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.Claude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- Section 8.15.1; pp. 137–139, Figure 8.15.B (max effort)
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Baichuan-M3 Baichuan65.1as of 2026-02self-run in Baichuan-M3 paper (arXiv 2602.06570)Baichuan-M3 Technical Report · paper · first-party
- publisher
- Baichuan
- locator
- p. 23, section 4.2.1 HealthBench-Main; Table on p. 25 (Model / HealthBench Score)
- reported by
- vendor-reported
- published
- 2026-02-06
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Baichuan-M3 Technical Report (paper, 65.1, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (Baichuan-M3-235B, HealthBench Score 65.1)) - GPT-5.2-High OpenAI63.3as of 2026-02raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card.Baichuan-M3 Technical Report · paper · first-party
- publisher
- Baichuan
- locator
- p. 25, HealthBench-Hallu table (Model / HealthBench Score column); also p. 23 prose
- reported by
- independent run
- published
- 2026-02-06
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Baichuan-M3 Technical Report (paper, 63.3, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (GPT-5.2-High, HealthBench Score 63.3)); GPT-5.6 System Card (system card, 56.8, conflicting, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2) - Claude Opus 5.5 Anthropic60.6as of 2026-09Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 68.1%. Five-trial average; refusal fallback to Claude Opus 5.Claude Opus 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- Section 8.15.1; pp. 213–214, Figures 8.15.1.A and 8.15.2.A
- reported by
- vendor-reported
- published
- 2026-09-22
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5 Anthropic60.4as of 2026-09Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 61.2%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.Claude Fable 5.1 and Claude Mythos 5.1 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 198, Figure 8.17.1.A; Claude Fable 5, adjusted bar 60.4%
- reported by
- vendor-reported
- published
- 2026-09-01
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 Anthropic60%as of 2026-09length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%).Claude Fable 5.1 and Claude Mythos 5.1 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 198, sec. 8.17.1 / Figure 8.17.1.A
- reported by
- vendor-reported
- published
- 2026-09-01
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.8 Anthropic59.3as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itselfClaude Fable 5 and Claude Mythos 5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 252, Table 8.1.A, row HealthBench, column Opus 4.8
- reported by
- vendor-reported
- published
- 2026-06-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in System Card: Claude Opus 5 (system card, 59.3%, p. 188, section 8.15.1, Figure 8.15.1.A) - Claude Sonnet 5 Anthropic58.7%as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 cardSystem Card: Claude Sonnet 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 138, section 8.12.1 HealthBench results, Figure 8.12.1.A bar label (no prose or table number)
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-28
- confidence
- verified
also reported in System Card: Claude Opus 5 (system card, 58.7%, p. 188, section 8.15.1, Figure 8.15.1.A) - GPT-6 Astra OpenAI58.3as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 56.9%, mean answer 1,760 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench length-adjusted row, GPT-6 Astra column; value 58.3 (56.9, 1760)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 Anthropic57.8as of 2026-07Anthropic evaluation; length-adjusted 57.8%, raw 67.1%; adaptive max effort, Opus 4.8 grader, five-trial average, no tools or customized system prompt. Previously this catalog displayed the raw value.System Card: Claude Opus 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 188, section 8.15.1 and Figure 8.15.1.A; adjusted 57.8%, raw 67.1%
- reported by
- vendor-reported
- published
- 2026-07-24
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5 OpenAI57.7as of 2026-06OpenAI length-adjusted score, maximum reasoning effort; raw 63.1%, mean answer 2,904 characters; GPT-5.6 card Table 6.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5 column
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
- GPT OSS 120B OpenAI57.6as of 2025-08reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9)gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-120b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol OpenAI57.0as of 2026-06length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL) - GPT-5.6 Terra OpenAI57.0as of 2026-06length-adjusted (58.7 unadjusted), max reasoning effortGPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA) - GPT-5.2 OpenAI56.8as of 2026-06OpenAI length-adjusted score, maximum reasoning effort; raw 60.7%, mean answer 2,645 characters; GPT-5.6 card Table 6.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5.2 column
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.5 OpenAI56.5as of 2026-04length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system cardGPT-5.5 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5 Health, Table 7 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
- reported by
- vendor-reported
- published
- 2026-04-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 56.5); GPT-5.6 System Card (system card, 56.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5) - GPT-5.6 Luna OpenAI55.8as of 2026-06length-adjusted (55.4 unadjusted), max reasoning effortGPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 55.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA) - GPT-5.6 Sol (August) OpenAI55.0as of 2026-08ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 Preview System Card (system card, 55.0) - GPT-6 Luna OpenAI54.5as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 50%, mean answer 1,255 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench length-adjusted row, GPT-6 Luna column; value 54.5 (50, 1255)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.3 Chat OpenAI54.1%as of 2026-03raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted)GPT-5.3 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 3: HealthBench, row HealthBench, column GPT-5.3-INSTANT
- reported by
- vendor-reported
- published
- 2026-03-02
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.5 Instant System Card (system card, 49.6, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT) - GPT-5.4 OpenAI54.0as of 2026-06OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5.4 column
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Luna (August) OpenAI53.3as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Sol OpenAI53.2as of 2026-09OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 47.1%, mean answer 977 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.GPT-6 Astra System Card — September 22 revision · system card · first-party
- publisher
- OpenAI
- locator
- Section 11.4.1, Table 29; HealthBench length-adjusted row, GPT-6 Sol column; value 53.2 (47.1, 977)
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.5 Instant OpenAI51.4as of 2026-05length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11GPT-5.5 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-28
- confidence
- verified
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 51.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant) - GPT-5.1 OpenAI50.9as of 2026-06OpenAI length-adjusted score, maximum reasoning effort; raw 64.2%, mean answer 4,222 characters; GPT-5.6 card Table 6.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5.1 column
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-28
- confidence
- verified
- GPT OSS 20B OpenAI42.5as of 2025-08reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8)gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-20b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-28
- confidence
- verified
Health Optimization Bench
16 rows · 0-100 rubric creditBoard source: healthoptimizationbench.com · official page: healthoptimizationbench.com · full board: healthoptimizationbench.com
- Claude Fable 5 Anthropic70.9as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 68.0–73.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Claude Fable 5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.709, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 Anthropic69.3as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 66.2–72.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Claude Opus 5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.693, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.6 xAI66.8as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.8–69.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Grok 4.6; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.668, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol (max) OpenAI66.6as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; GPT-5.6 Sol (max); tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.666, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol (high) OpenAI64.6as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; GPT-5.6 Sol (high); tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.646, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K3 Moonshot AI59.9as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Kimi K3; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.599, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark Meta57.2as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Muse Spark; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.572, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 Anthropic47.3as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 42.2–52.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Vendor safeguards declined 97/257 tasks, scored with no credit; mean over answered tasks is 75.9.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Claude Fable 5.1; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.473, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.6 Google39.7as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 36.2–43.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Gemini 3.6; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.397, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Inkling Thinking Machines35.6as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 32.3–39.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Inkling; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.356, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5 Anthropic34.6as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Claude Sonnet 5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.346, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.2 Zhipu20.7as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 18.1–23.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; GLM 5.2; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.207, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax M3 MiniMax18.3as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.7–20.9. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; MiniMax M3; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.183, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- MAI Thinking Microsoft AI17.5as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; MAI Thinking; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.175, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Mistral Medium 3.5 Mistral9.2as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 7.4–11.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Mistral Medium 3.5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.092, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
- Nemotron 3.5 Lightning NVIDIA4.9as of 2026-09Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 3.6–6.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.Health Optimization Bench — subject suites results · Arcophos run · first-party
- publisher
- Arcophos
- locator
- Subject suites table; Nemotron 3.5 Lightning; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.049, n=257
- reported by
- Arcophos run
- published
- 2026-09-10
- retrieved
- 2026-09-28
- confidence
- verified
MAST (Medical AI Superintelligence Test)
8 rows · percentage compositeBoard source: MAST: Medical AI Superintelligence Test leaderboard (General board)
- GPT-5.6 Sol OpenAI60.2%as of 2026-08MAST in preview; 'exact scores may change'MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K3 Moonshot AI60.1%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.6 Flash Google59.3%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.1 Pro Google58.9%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3.5 397B A17B Alibaba57.9%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 Anthropic57.1%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5 Anthropic56.6%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.3 xAI53.7%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
MedHELM
10 rows · mean win rate 0-1Board source: MedHELM leaderboard (medhelm.org), v5.0.0
- Gemini 3.1 Pro (Preview) Google0.652as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6520833333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Gemini 3.5 Flash Google0.642as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Muse Spark (2026-04-08) Meta0.621as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6208333333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - GPT-5.4 mini OpenAI0.552as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5520833333333334, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - GPT-5.4 (2026-03-05) OpenAI0.538as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5375, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Gemini 2.5 Pro Google0.529as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5291666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - DeepSeek R1 DeepSeek0.485as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.48541666666666666, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Claude 4.6 Opus Anthropic0.456as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45625, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Claude 3.7 Sonnet Anthropic0.45as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Gemini 2.0 Flash Google0.342as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-28
- confidence
- verified
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.3416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
First, Do NOHARM (v2)
17 rows · percentage safety scoreBoard source: MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · paper: arxiv.org
- LiSA 2.5 AMBOSS86.2as of 2026-09First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.ARISE MAST technical leaderboard · official leaderboard · first-party
- publisher
- ARISE
- locator
- First Do NOHARM v2 overall leaderboard; LiSA 2.5; score 86.2%
- reported by
- benchmark-owner run
- retrieved
- 2026-09-28
- confidence
- verified
- Doximity Ask 6.1 Doximity84.5as of 2026-09First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.ARISE MAST technical leaderboard · official leaderboard · first-party
- publisher
- ARISE
- locator
- First Do NOHARM v2 overall leaderboard; Doximity Ask 6.1; score 84.5%
- reported by
- benchmark-owner run
- retrieved
- 2026-09-28
- confidence
- verified
- OpenEvidence OpenEvidence80.0as of 2026-09First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.ARISE MAST technical leaderboard · official leaderboard · first-party
- publisher
- ARISE
- locator
- First Do NOHARM v2 overall leaderboard; OpenEvidence; score 80%
- reported by
- benchmark-owner run
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark 1.1 Meta79.7%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Glass 5.6 Max Glass Health79.7as of 2026-09First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.ARISE MAST technical leaderboard · official leaderboard · first-party
- publisher
- ARISE
- locator
- First Do NOHARM v2 overall leaderboard; Glass 5.6 Max; score 79.7%
- reported by
- benchmark-owner run
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 Anthropic74.6%as of 2026-08v2 run on ARISE; 19 models on the boardMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K3 Moonshot AI74.0%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol OpenAI70.1%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.5 OpenAI70.0%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5 OpenAI68.6%as of 2026-08from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships rankingMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, Model Leaderboard (Top 10 shown), row 10, SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5 Anthropic65.0%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.1 Pro Google62.6%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Pro Google61.9%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3.5 397B A17B Alibaba61.1%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K2.6 Moonshot AI59.1%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.1 Z.ai57.9as of 2026-09First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. Open-weight model as listed by the evaluator. Observed September 28, 2026; the individual run date is not published.ARISE MAST technical leaderboard · official leaderboard · first-party
- publisher
- ARISE
- locator
- First Do NOHARM v2 overall leaderboard; GLM 5.1; score 57.9%
- reported by
- benchmark-owner run
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek R1 DeepSeek55.8%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-28
- confidence
- verified
HealthAgentBench
12 rows · mean task success rateBoard source: HealthAgentBench leaderboard · paper: arxiv.org
- Claude Code (Opus 5) Anthropic55%as of 2026-07$3.3/task; harness+model evaluated jointlyHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 1, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 55%, Detailed results table, rank 1) - Codex (GPT-5.6-sol) OpenAI45%as of 2026-07$5.2/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 2, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 45%, Detailed results table, rank 2) - Codex (GPT 5.5) OpenAI42%as of 2026-07$2.8/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 3, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 42%, Detailed results table, rank 3); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 42%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Copilot (Opus 4.8) Microsoft/Anthropic36%as of 2026-07$3.1/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 4, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 36%, Detailed results table, rank 4); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 36%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Copilot (GPT 5.5) Microsoft/OpenAI35%as of 2026-07$2.6/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 5, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 35%, Detailed results table, rank 5); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 35%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Opus 4.8) Anthropic32%as of 2026-07$4.0/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 6, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 32%, Detailed results table, rank 6); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 32%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Codex (GPT 5.4) OpenAI28%as of 2026-07$1.3/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 7, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 28%, Detailed results table, rank 7); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 28%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Opus 4.7) Anthropic27%as of 2026-07$4.8/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 8, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 27%, Detailed results table, rank 8); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 27%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Codex (GPT 5.3) OpenAI22%as of 2026-07$1.0/task; harness and model evaluated jointlyHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Homepage leaderboard, rank 9, success rate and cost/task
- reported by
- benchmark publisher
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Code (Opus 4.6) Anthropic19%as of 2026-07$4.1/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 10, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 19%, Detailed results table, rank 10); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 19%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Sonnet 4.6) Anthropic17%as of 2026-07$2.9/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 11, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 17%, Detailed results table, rank 11); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 17%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Codex (GPT 5.4 Mini) OpenAI16%as of 2026-07$0.6/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 12, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-28
- confidence
- verified
also reported in HealthAgentBench detailed results (official leaderboard, 16%, Detailed results table, rank 12); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 16%, p. 11, Figure 4 (pooled task success rate, ten agents))
CHI-Bench
44 rows · pass@1 with binary 0/1 rewardBoard source: CHI-Bench leaderboard (actAVA)
- erius + claude-opus-5 Humana (harness) / Anthropic (model)54.7%as of 2026-07-26All Domains pass@1; PA 72.0%, UM 36.0%, CM 56.0%; submitted 2026-07-26; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 1, Accuracy column; submission date 2026-07-26
- reported by
- independent run
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- erius + claude-opus-4-8 Humana (harness) / Anthropic (model)37.3%as of 2026-06-05All Domains pass@1; PA 40.0%, UM 16.0%, CM 56.0%; submitted 2026-06-05; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 2, Accuracy column; submission date 2026-06-05
- reported by
- independent run
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-opus-5 Anthropic37.3%as of 2026-07-24All Domains pass@1; PA 20.0%, UM 32.0%, CM 60.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 3, Accuracy column; submission date 2026-07-24
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-opus-4-8 Anthropic33.3%as of 2026-05-28All Domains pass@1; PA 32.0%, UM 28.0%, CM 40.0%; submitted 2026-05-28; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 4, Accuracy column; submission date 2026-05-28
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-opus-4-6 Anthropic28.0%as of 2026-05-01All Domains pass@1; PA 20.0%, UM 36.0%, CM 28.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 5, Accuracy column; submission date 2026-05-01
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-sonnet-4-6 Anthropic26.2%as of 2026-05-01All Domains pass@1; PA 24.0%, UM 34.7%, CM 20.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 6, Accuracy column; submission date 2026-05-01
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- codex + gpt-5.6-sol OpenAI25.3%as of 2026-07-24All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 7, Accuracy column; submission date 2026-07-24
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + kimi-k3 Moonshot AI25.3%as of 2026-07-24All Domains pass@1; PA 28.0%, UM 32.0%, CM 16.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 8, Accuracy column; submission date 2026-07-24
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-opus-4-7 Anthropic24.4%as of 2026-05-01All Domains pass@1; PA 24.0%, UM 17.3%, CM 32.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 9, Accuracy column; submission date 2026-05-01
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-fable-5 Anthropic24.0%as of 2026-07-22All Domains pass@1; PA 24.0%, UM 24.0%, CM 24.0%; submitted 2026-07-22; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 10, Accuracy column; submission date 2026-07-22
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- hermes + MedGuard MedGuard22.7%as of 2026-07-06All Domains pass@1; PA 4.0%, UM 4.0%, CM 60.0%; submitted 2026-07-06; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 11, Accuracy column; submission date 2026-07-06
- reported by
- community submission
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- codex + gpt-5.5 OpenAI20.9%as of 2026-05-01All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 12, Accuracy column; submission date 2026-05-01
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-sonnet-5 Anthropic20.0%as of 2026-07-06All Domains pass@1; PA 24.0%, UM 24.0%, CM 12.0%; submitted 2026-07-06; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 13, Accuracy column; submission date 2026-07-06
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + glm-5.1 Zhipu AI18.7%as of 2026-05-01All Domains pass@1; PA 18.7%, UM 33.3%, CM 4.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 14, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- hermes + glm-5.1 Zhipu AI18.7%as of 2026-05-01All Domains pass@1; PA 10.7%, UM 34.7%, CM 10.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 15, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + glm-5.2 Zhipu AI18.7%as of 2026-07-06All Domains pass@1; PA 20.0%, UM 32.0%, CM 4.0%; submitted 2026-07-06; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 16, Accuracy column; submission date 2026-07-06
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openclaw + claude-opus-4-7 Anthropic17.3%as of 2026-05-01All Domains pass@1; PA 18.7%, UM 13.3%, CM 20.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 17, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openclaw + glm-5.1 Zhipu AI16.9%as of 2026-05-01All Domains pass@1; PA 13.3%, UM 26.7%, CM 10.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 18, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- hermes + qwen-3.6-max Alibaba16.4%as of 2026-05-01All Domains pass@1; PA 9.3%, UM 26.7%, CM 13.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 19, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- codex + gpt-5.4 OpenAI16.0%as of 2026-05-01All Domains pass@1; PA 24.0%, UM 17.3%, CM 6.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 20, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + qwen-3.6-max Alibaba15.6%as of 2026-05-01All Domains pass@1; PA 16.0%, UM 26.7%, CM 4.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 21, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- hermes + kimi-k2.6 Moonshot AI15.6%as of 2026-05-01All Domains pass@1; PA 18.7%, UM 21.3%, CM 6.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 22, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + kimi-k2.6 Moonshot AI15.1%as of 2026-05-01All Domains pass@1; PA 17.3%, UM 25.3%, CM 2.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 23, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + deepseek-v4-pro DeepSeek14.2%as of 2026-05-01All Domains pass@1; PA 10.7%, UM 28.0%, CM 4.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 24, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- hermes + deepseek-v4-pro DeepSeek13.8%as of 2026-05-01All Domains pass@1; PA 8.0%, UM 25.3%, CM 8.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 25, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- codex + gpt-5.6-terra OpenAI13.3%as of 2026-07-24All Domains pass@1; PA 12.0%, UM 20.0%, CM 8.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 26, Accuracy column; submission date 2026-07-24
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- codex + gpt-5.6-luna OpenAI13.3%as of 2026-07-24All Domains pass@1; PA 20.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 27, Accuracy column; submission date 2026-07-24
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- gemini-cli + gemini-3-flash Google12.5%as of 2026-05-01All Domains pass@1; PA 18.7%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 28, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openclaw + deepseek-v4-pro DeepSeek11.1%as of 2026-05-01All Domains pass@1; PA 14.7%, UM 12.0%, CM 6.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 29, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- deepagents + glm-5.1 Zhipu AI11.1%as of 2026-05-01All Domains pass@1; PA 17.3%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 30, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- deepagents + deepseek-v4-pro DeepSeek10.7%as of 2026-05-01All Domains pass@1; PA 14.7%, UM 10.7%, CM 6.7%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 31, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openclaw + kimi-k2.6 Moonshot AI10.2%as of 2026-05-01All Domains pass@1; PA 12.0%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 32, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- deepagents + qwen-3.6-max Alibaba9.3%as of 2026-05-01All Domains pass@1; PA 12.0%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 33, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- codex + gpt-5.4-mini OpenAI8.4%as of 2026-05-01All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 34, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + TML Inkling 256K Thinking Machines8.0%as of 2026-07-24All Domains pass@1; PA 4.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 35, Accuracy column; submission date 2026-07-24
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- gemini-cli + gemini-3.1-pro Google7.1%as of 2026-05-01All Domains pass@1; PA 14.7%, UM 6.7%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 36, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- claude-code + claude-haiku-4-5 Anthropic6.2%as of 2026-05-01All Domains pass@1; PA 0.0%, UM 14.7%, CM 4.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 37, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + grok-4.3 SpaceX AI5.8%as of 2026-05-01All Domains pass@1; PA 0.0%, UM 16.0%, CM 1.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 38, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openclaw + qwen-3.6-max Alibaba4.9%as of 2026-05-01All Domains pass@1; PA 10.7%, UM 4.0%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 39, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- hermes + grok-4.3 SpaceX AI4.4%as of 2026-05-01All Domains pass@1; PA 0.0%, UM 13.3%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 40, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- deepagents + kimi-k2.6 Moonshot AI3.1%as of 2026-05-01All Domains pass@1; PA 8.0%, UM 1.3%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 41, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- deepagents + grok-4.3 SpaceX AI2.2%as of 2026-05-01All Domains pass@1; PA 0.0%, UM 5.3%, CM 1.3%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 42, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openclaw + grok-4.3 SpaceX AI0.4%as of 2026-05-01All Domains pass@1; PA 1.3%, UM 0.0%, CM 0.0%; submitted 2026-05-01; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 43, Accuracy column; submission date 2026-05-01
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
- openai-agents + Nemotron 3 Ultra 256K NVIDIA0.0%as of 2026-07-24All Domains pass@1; PA 0.0%, UM 0.0%, CM 0.0%; submitted 2026-07-24; run date not publishedCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- All Domains leaderboard, rank 44, Accuracy column; submission date 2026-07-24
- reported by
- benchmark publisher
- published
- 2026-08-12
- retrieved
- 2026-09-28
- confidence
- verified
MedCode (Vals AI)
102 rows · percentage accuracy 0-100Board source: Vals AI MedCode leaderboard
- Claude Opus 5 Anthropic63.57%as of 2026-09-26model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.156845/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5"].accuracy; Overall leaderboard rank 1
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.1 Pro Preview (02/26) Google59.06%as of 2026-09-26model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-pro-preview"].accuracy; Overall leaderboard rank 2
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5 Anthropic56.07%as of 2026-09-26model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.203 pp; $0.591071/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5"].accuracy; Overall leaderboard rank 3
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3 Flash (12/25) Google55.92%as of 2026-09-26model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.112 pp; $0.006187/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3-flash-preview"].accuracy; Overall leaderboard rank 4
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.5 Flash Google55.83%as of 2026-09-26model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.073716/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash"].accuracy; Overall leaderboard rank 5
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.7 Anthropic54.86%as of 2026-09-26model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.205 pp; $0.226314/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-7"].accuracy; Overall leaderboard rank 6
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 Anthropic53.51%as of 2026-09-26model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.165 pp; $1.116862/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5-1"].accuracy; Overall leaderboard rank 7
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.7 Flash Google53.39%as of 2026-09-26model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.12 pp; $0.038331/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.7-flash"].accuracy; Overall leaderboard rank 8
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.8 Anthropic53.22%as of 2026-09-26model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.165 pp; $0.350925/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-8"].accuracy; Overall leaderboard rank 9
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.6 Flash Google53.15%as of 2026-09-26model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.157 pp; $0.044216/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.6-flash"].accuracy; Overall leaderboard rank 10
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5.5 Anthropic52.92%as of 2026-09-26model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.119 pp; $0.391592/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5-5"].accuracy; Overall leaderboard rank 11
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.1 OpenAI52.73%as of 2026-09-26model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 2.151 pp; $0.014371/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.1-2025-11-13"].accuracy; Overall leaderboard rank 12
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3 Pro (11/25) Google52.20%as of 2026-09-26model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.073 pp; $0.028248/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3-pro-preview"].accuracy; Overall leaderboard rank 13
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark Meta51.31%as of 2026-09-26model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["meta/muse_spark"].accuracy; Overall leaderboard rank 14
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Pro Google50.59%as of 2026-09-26model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.015389/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-pro"].accuracy; Overall leaderboard rank 15
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5.5 Anthropic49.80%as of 2026-09-26model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.273 pp; $0.658237/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5-5"].accuracy; Overall leaderboard rank 16
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.2 OpenAI49.75%as of 2026-09-26model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.2-2025-12-11"].accuracy; Overall leaderboard rank 17
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5 OpenAI49.63%as of 2026-09-26model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5-2025-08-07"].accuracy; Overall leaderboard rank 18
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.7 SpaceXAI49.55%as of 2026-09-26model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.7"].accuracy; Overall leaderboard rank 19
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark 1.2 Meta49.35%as of 2026-09-26model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.187 pp; $0.039023/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["meta/muse_spark_1_2"].accuracy; Overall leaderboard rank 20
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.5 (Thinking) Anthropic49.16%as of 2026-09-26model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 2.012 pp; $0.095846/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].accuracy; Overall leaderboard rank 21
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.6 (Thinking) Anthropic49.13%as of 2026-09-26model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.085 pp; $0.244127/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6-thinking"].accuracy; Overall leaderboard rank 22
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.5 OpenAI49.10%as of 2026-09-26model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.188 pp; $0.160759/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.5"].accuracy; Overall leaderboard rank 23
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K3 Moonshot AI48.88%as of 2026-09-26model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["kimi/kimi-k3"].accuracy; Overall leaderboard rank 24
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Astra OpenAI48.49%as of 2026-09-26model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 2.131 pp; $0.451358/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-6-astra"].accuracy; Overall leaderboard rank 25
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.6 (Nonthinking) Anthropic48.24%as of 2026-09-26model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.05 pp; $0.006180/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6"].accuracy; Overall leaderboard rank 26
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.8 Flash Google48.13%as of 2026-09-26model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.18 pp; $0.017700/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.8-flash"].accuracy; Overall leaderboard rank 27
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.1 Flash Lite Preview Google47.60%as of 2026-09-26model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.071 pp; $0.002029/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-flash-lite-preview"].accuracy; Overall leaderboard rank 28
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5 Anthropic47.54%as of 2026-09-26model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5"].accuracy; Overall leaderboard rank 29
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- o3 OpenAI47.29%as of 2026-09-26model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.161 pp; $0.029818/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/o3-2025-04-16"].accuracy; Overall leaderboard rank 30
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.1 (Thinking) Anthropic47.23%as of 2026-09-26model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 2.067 pp; $0.269254/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].accuracy; Overall leaderboard rank 31
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Sol OpenAI47.07%as of 2026-09-26model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 2.119 pp; $0.085827/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-6-sol"].accuracy; Overall leaderboard rank 32
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M3 MiniMax46.29%as of 2026-09-26model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.104 pp; $0.012125/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M3"].accuracy; Overall leaderboard rank 33
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.5 (Nonthinking) Anthropic45.17%as of 2026-09-26model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.006826/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101"].accuracy; Overall leaderboard rank 34
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.6 Pro Xiaomi44.97%as of 2026-09-26model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.097 pp; $0.009528/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-pro"].accuracy; Overall leaderboard rank 35
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.6 SpaceXAI44.71%as of 2026-09-26model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.256 pp; $0.050335/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.6"].accuracy; Overall leaderboard rank 36
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Luna OpenAI44.69%as of 2026-09-26model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-6-luna"].accuracy; Overall leaderboard rank 37
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4.5 (Thinking) Anthropic44.13%as of 2026-09-26model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.101495/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].accuracy; Overall leaderboard rank 38
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol OpenAI43.97%as of 2026-09-26model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-sol"].accuracy; Overall leaderboard rank 39
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.5 Flash Lite Google43.49%as of 2026-09-26model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.951 pp; $0.008093/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash-lite"].accuracy; Overall leaderboard rank 40
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Terra OpenAI43.41%as of 2026-09-26model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.173 pp; $0.046558/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-terra"].accuracy; Overall leaderboard rank 41
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.5 SpaceXAI43.29%as of 2026-09-26model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.313 pp; $0.048461/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.5"].accuracy; Overall leaderboard rank 42
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Hy4 Preview Tencent43.25%as of 2026-09-26model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.134 pp; $0.062053/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["tencent/hy4-preview"].accuracy; Overall leaderboard rank 43
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5 Mini OpenAI43.05%as of 2026-09-26model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.045 pp; $0.005560/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5-mini-2025-08-07"].accuracy; Overall leaderboard rank 44
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.3 zAI42.86%as of 2026-09-26model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.111 pp; $0.081905/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.3"].accuracy; Overall leaderboard rank 45
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4 Pro 0813 DeepSeek42.47%as of 2026-09-26model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.16 pp; $0.061060/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro-0813"].accuracy; Overall leaderboard rank 46
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Luna OpenAI42.39%as of 2026-09-26model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-luna"].accuracy; Overall leaderboard rank 47
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.1 zAI41.60%as of 2026-09-26model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.024196/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.1"].accuracy; Overall leaderboard rank 48
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4 Flash 0731 DeepSeek41.41%as of 2026-09-26model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 2.15 pp; $0.019659/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-flash-0731"].accuracy; Overall leaderboard rank 49
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.1 (Nonthinking) Anthropic41.37%as of 2026-09-26model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.206270/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805"].accuracy; Overall leaderboard rank 50
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.4 (xhigh) OpenAI41.29%as of 2026-09-26model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.148 pp; $0.212108/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-2026-03-05"].accuracy; Overall leaderboard rank 51
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Inkling Thinking Machines41.19%as of 2026-09-26model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.23 pp; $0.127025/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling"].accuracy; Overall leaderboard rank 52
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4.1 Flash DeepSeek41.17%as of 2026-09-26model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 2.042 pp; $0.012450/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4.1-flash"].accuracy; Overall leaderboard rank 53
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.6 Flash Xiaomi41.06%as of 2026-09-26model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.032 pp; $0.002865/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-flash"].accuracy; Overall leaderboard rank 54
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.4 Nano OpenAI41.03%as of 2026-09-26model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 2.256 pp; $0.000844/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-nano-2026-03-17"].accuracy; Overall leaderboard rank 55
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.2 zAI40.77%as of 2026-09-26model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.166 pp; $0.045010/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.2"].accuracy; Overall leaderboard rank 56
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.8 Max Alibaba40.67%as of 2026-09-26model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 2.029 pp; $0.120827/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-max"].accuracy; Overall leaderboard rank 57
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4.5 (Nonthinking) Anthropic40.57%as of 2026-09-26model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.995 pp; $0.042403/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929"].accuracy; Overall leaderboard rank 58
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Preview (9/25) (Nonthinking) Google40.54%as of 2026-09-26model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.003692/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025"].accuracy; Overall leaderboard rank 59
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4 DeepSeek40.45%as of 2026-09-26model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.122 pp; $0.060710/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro"].accuracy; Overall leaderboard rank 60
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash (7/17) (Thinking) Google40.36%as of 2026-09-26model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.952 pp; $0.003661/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-thinking"].accuracy; Overall leaderboard rank 61
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Preview (9/25) (Thinking) Google40.33%as of 2026-09-26model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.003653/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 62
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K2.6 Moonshot AI40.14%as of 2026-09-26model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 2.041 pp; $0.041295/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.6"].accuracy; Overall leaderboard rank 63
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K2.5 Moonshot AI39.32%as of 2026-09-26model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 2.119 pp; $0.017275/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.5-thinking"].accuracy; Overall leaderboard rank 64
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.7 Max Alibaba38.75%as of 2026-09-26model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 2.196 pp; $0.042362/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.7-max"].accuracy; Overall leaderboard rank 65
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Nemotron 3 Ultra NVIDIA38.62%as of 2026-09-26model ID nvidia/nemotron-3-ultra-550b-a55b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.001 pp; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["nvidia/nemotron-3-ultra-550b-a55b"].accuracy; Overall leaderboard rank 66
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash (7/17) (Nonthinking) Google38.42%as of 2026-09-26model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.003698/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash"].accuracy; Overall leaderboard rank 67
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 SpaceXAI38.08%as of 2026-09-26model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.206 pp; $0.034103/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-0709"].accuracy; Overall leaderboard rank 68
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.3 SpaceXAI38.07%as of 2026-09-26model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.081 pp; $0.022202/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.3"].accuracy; Overall leaderboard rank 69
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Inkling Small Thinking Machines37.89%as of 2026-09-26model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.206 pp; $0.016656/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling-small"].accuracy; Overall leaderboard rank 70
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 Fast (Reasoning) SpaceXAI37.38%as of 2026-09-26model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.941 pp; $0.002143/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-reasoning"].accuracy; Overall leaderboard rank 71
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.6 Plus Alibaba36.89%as of 2026-09-26model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 2.017 pp; $0.015673/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.6-plus"].accuracy; Overall leaderboard rank 72
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 4 Maverick Meta36.51%as of 2026-09-26model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.994 pp; $0.002888/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["fireworks/llama4-maverick-instruct-basic"].accuracy; Overall leaderboard rank 73
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4 (Thinking) Anthropic34.96%as of 2026-09-26model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 1.939 pp; $0.069896/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].accuracy; Overall leaderboard rank 74
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M2.7 MiniMax34.44%as of 2026-09-26model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.985 pp; $0.007424/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.7"].accuracy; Overall leaderboard rank 75
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Lite (9/25) (Thinking) Google34.19%as of 2026-09-26model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.736 pp; $0.001182/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 76
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M2.1 MiniMax34.08%as of 2026-09-26model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.943 pp; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.1"].accuracy; Overall leaderboard rank 77
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4 (Nonthinking) Anthropic33.94%as of 2026-09-26model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.906 pp; $0.039460/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514"].accuracy; Overall leaderboard rank 78
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- o4 Mini OpenAI33.79%as of 2026-09-26model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.021 pp; $0.017605/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/o4-mini-2025-04-16"].accuracy; Overall leaderboard rank 79
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Mistral Medium 3.5 Mistral33.75%as of 2026-09-26model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.148 pp; $0.053370/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["mistralai/mistral-medium-3.5"].accuracy; Overall leaderboard rank 80
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.5 Flash Alibaba33.00%as of 2026-09-26model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 1.787 pp; $0.003934/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.5-flash"].accuracy; Overall leaderboard rank 81
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 4.7 zAI32.77%as of 2026-09-26model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.996 pp; $0.006710/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-4.7"].accuracy; Overall leaderboard rank 82
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Haiku 4.5 (Thinking) Anthropic32.68%as of 2026-09-26model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.020099/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].accuracy; Overall leaderboard rank 83
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.5 Pro Xiaomi32.48%as of 2026-09-26model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.006719/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5-pro"].accuracy; Overall leaderboard rank 84
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Ling 3.0 Flash Ant Group32.27%as of 2026-09-26model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.908 pp; $0.001645/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-2607"].accuracy; Overall leaderboard rank 85
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.20 (Reasoning) SpaceXAI32.16%as of 2026-09-26model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.036189/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.20-0309-reasoning"].accuracy; Overall leaderboard rank 86
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.5 Xiaomi31.89%as of 2026-09-26model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.025 pp; $0.002162/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5"].accuracy; Overall leaderboard rank 87
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3 VL Plus Alibaba31.65%as of 2026-09-26model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.845 pp; $0.002519/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].accuracy; Overall leaderboard rank 88
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3 Max Thinking Alibaba31.37%as of 2026-09-26model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.014776/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-max-2026-01-23"].accuracy; Overall leaderboard rank 89
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Mercury 2.5 Inception31.33%as of 2026-09-26model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.953 pp; $0.004535/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["inception/mercury-2.5"].accuracy; Overall leaderboard rank 90
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5 Nano OpenAI30.44%as of 2026-09-26model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.948 pp; $0.001729/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5-nano-2025-08-07"].accuracy; Overall leaderboard rank 91
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 Fast (Non-Reasoning) SpaceXAI30.04%as of 2026-09-26model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.974 pp; $0.002149/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-non-reasoning"].accuracy; Overall leaderboard rank 92
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Ling 3.0 Flash Fin Ant Group29.30%as of 2026-09-26model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 1.943 pp; $0.001354/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-af-rc3"].accuracy; Overall leaderboard rank 93
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.8 27B Alibaba28.70%as of 2026-09-26model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.971 pp; $0.050145/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-27b"].accuracy; Overall leaderboard rank 94
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.1 Fast Non-Reasoning SpaceXAI28.35%as of 2026-09-26model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.921 pp; $0.002193/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-non-reasoning"].accuracy; Overall leaderboard rank 95
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.1 Fast (Reasoning) SpaceXAI28.08%as of 2026-09-26model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.992 pp; $0.002108/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-reasoning"].accuracy; Overall leaderboard rank 96
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Lite (Nonthinking) Google27.11%as of 2026-09-26model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.843 pp; $0.001342/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite"].accuracy; Overall leaderboard rank 97
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Lite (9/25) (Nonthinking) Google27.08%as of 2026-09-26model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.911 pp; $0.001440/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].accuracy; Overall leaderboard rank 98
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 4 Scout Meta23.31%as of 2026-09-26model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].accuracy; Overall leaderboard rank 99
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Laguna M.1 Poolside23.11%as of 2026-09-26model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 1.693 pp; $0.002973/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["poolside/laguna-m.1"].accuracy; Overall leaderboard rank 100
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Laguna XS.2 Poolside21.25%as of 2026-09-26model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 1.703 pp; $0.001445/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["poolside/laguna-xs.2"].accuracy; Overall leaderboard rank 101
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Command A+ Cohere19.72%as of 2026-09-26model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 1.835 pp; $0.055695/test; source snapshot 2026-09-26; run date not publishedVals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["cohere/command-a-plus-05-2026"].accuracy; Overall leaderboard rank 102
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
MedScribe (Vals AI)
104 rows · percentage accuracy 0-100Board source: Vals AI MedScribe leaderboard
- Claude Opus 5.5 Anthropic91.43%as of 2026-09-26model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.932 pp; $1.154156/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5-5"].accuracy; Overall leaderboard rank 1
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 Anthropic91.29%as of 2026-09-26model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.953 pp; $0.963500/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5-1"].accuracy; Overall leaderboard rank 2
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5.5 Anthropic91.10%as of 2026-09-26model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.96 pp; $0.508604/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5-5"].accuracy; Overall leaderboard rank 3
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 Anthropic90.98%as of 2026-09-26model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.236275/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5"].accuracy; Overall leaderboard rank 4
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark 1.2 Meta90.06%as of 2026-09-26model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.037779/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["meta/muse_spark_1_2"].accuracy; Overall leaderboard rank 5
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.7 SpaceXAI89.38%as of 2026-09-26model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.7"].accuracy; Overall leaderboard rank 6
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.3 Flash zAI88.94%as of 2026-09-26model ID zai/glm-5.3-flash; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.002235/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.3-flash"].accuracy; Overall leaderboard rank 7
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark 1.1 Meta88.89%as of 2026-09-26model ID meta/muse_spark_1_1; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.95 pp; $0.034628/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["meta/muse_spark_1_1"].accuracy; Overall leaderboard rank 8
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.3 zAI88.81%as of 2026-09-26model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.999 pp; $0.057236/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.3"].accuracy; Overall leaderboard rank 9
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5 Anthropic88.52%as of 2026-09-26model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.945 pp; $0.583239/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5"].accuracy; Overall leaderboard rank 10
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.6 Pro Xiaomi88.31%as of 2026-09-26model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.938 pp; $0.009880/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-pro"].accuracy; Overall leaderboard rank 11
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.1 OpenAI88.09%as of 2026-09-26model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 1.942 pp; $0.096508/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.1-2025-11-13"].accuracy; Overall leaderboard rank 12
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K3 Moonshot AI87.96%as of 2026-09-26model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["kimi/kimi-k3"].accuracy; Overall leaderboard rank 13
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Astra OpenAI87.91%as of 2026-09-26model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 1.938 pp; $0.581991/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-6-astra"].accuracy; Overall leaderboard rank 14
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M3 MiniMax87.25%as of 2026-09-26model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.957 pp; $0.013748/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M3"].accuracy; Overall leaderboard rank 15
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.5 SpaceXAI86.88%as of 2026-09-26model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.944 pp; $0.033208/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.5"].accuracy; Overall leaderboard rank 16
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.5 OpenAI86.87%as of 2026-09-26model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.142988/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.5"].accuracy; Overall leaderboard rank 17
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.6 (Nonthinking) Anthropic86.74%as of 2026-09-26model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.942 pp; $0.115121/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6"].accuracy; Overall leaderboard rank 18
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.6 xAI86.53%as of 2026-09-26model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.956 pp; $0.037190/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.6"].accuracy; Overall leaderboard rank 19
- reported by
- official leaderboard
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.6 (Thinking) Anthropic86.13%as of 2026-09-26model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.944 pp; $0.224735/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6-thinking"].accuracy; Overall leaderboard rank 20
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark Meta85.90%as of 2026-09-26model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["meta/muse_spark"].accuracy; Overall leaderboard rank 21
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.8 Anthropic85.75%as of 2026-09-26model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.928 pp; $0.259121/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-8"].accuracy; Overall leaderboard rank 22
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4.1 Flash DeepSeek85.50%as of 2026-09-26model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 1.918 pp; $0.015397/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4.1-flash"].accuracy; Overall leaderboard rank 23
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Inkling Thinking Machines85.41%as of 2026-09-26model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.844 pp; $0.165561/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling"].accuracy; Overall leaderboard rank 24
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.5 (Thinking) Anthropic85.32%as of 2026-09-26model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.896 pp; $0.410224/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].accuracy; Overall leaderboard rank 25
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.6 Flash Xiaomi85.28%as of 2026-09-26model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.986 pp; $0.002184/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-flash"].accuracy; Overall leaderboard rank 26
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol OpenAI85.23%as of 2026-09-26model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-sol"].accuracy; Overall leaderboard rank 27
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Haiku 4.5 (Thinking) Anthropic85.23%as of 2026-09-26model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.042375/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].accuracy; Overall leaderboard rank 28
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.8 Max Alibaba84.95%as of 2026-09-26model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 1.999 pp; $0.089616/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-max"].accuracy; Overall leaderboard rank 29
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4.5 (Nonthinking) Anthropic84.52%as of 2026-09-26model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.054649/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929"].accuracy; Overall leaderboard rank 30
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.8 Flash Google84.50%as of 2026-09-26model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.943 pp; $0.025238/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.8-flash"].accuracy; Overall leaderboard rank 31
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Luna OpenAI84.39%as of 2026-09-26model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-luna"].accuracy; Overall leaderboard rank 32
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.2 OpenAI84.39%as of 2026-09-26model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.2-2025-12-11"].accuracy; Overall leaderboard rank 33
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Inkling Small Thinking Machines84.11%as of 2026-09-26model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.87 pp; $0.019018/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling-small"].accuracy; Overall leaderboard rank 34
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4.5 (Thinking) Anthropic84.10%as of 2026-09-26model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.873 pp; $0.082281/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].accuracy; Overall leaderboard rank 35
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.7 Flash Google83.94%as of 2026-09-26model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.004 pp; $0.058736/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.7-flash"].accuracy; Overall leaderboard rank 36
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.8 27B Alibaba83.85%as of 2026-09-26model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.977 pp; $0.046732/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-27b"].accuracy; Overall leaderboard rank 37
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.5 Pro Xiaomi83.73%as of 2026-09-26model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.063 pp; $0.006030/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5-pro"].accuracy; Overall leaderboard rank 38
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Luna OpenAI83.71%as of 2026-09-26model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-6-luna"].accuracy; Overall leaderboard rank 39
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5 OpenAI83.65%as of 2026-09-26model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5-2025-08-07"].accuracy; Overall leaderboard rank 40
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Hy4 Preview Tencent83.60%as of 2026-09-26model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.065 pp; $0.053243/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["tencent/hy4-preview"].accuracy; Overall leaderboard rank 41
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.2 zAI83.53%as of 2026-09-26model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.002 pp; $0.044912/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.2"].accuracy; Overall leaderboard rank 42
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.5 (Nonthinking) Anthropic83.25%as of 2026-09-26model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.926 pp; $0.281674/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101"].accuracy; Overall leaderboard rank 43
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash (7/17) (Thinking) Google82.98%as of 2026-09-26model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.908 pp; $0.014824/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-thinking"].accuracy; Overall leaderboard rank 44
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.7 Anthropic82.95%as of 2026-09-26model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.977 pp; $0.177841/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-7"].accuracy; Overall leaderboard rank 45
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash (7/17) (Nonthinking) Google82.87%as of 2026-09-26model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.909 pp; $0.014869/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash"].accuracy; Overall leaderboard rank 46
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Terra OpenAI82.87%as of 2026-09-26model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.948 pp; $0.062588/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-terra"].accuracy; Overall leaderboard rank 47
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Sol OpenAI82.03%as of 2026-09-26model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 1.942 pp; $0.083138/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-6-sol"].accuracy; Overall leaderboard rank 48
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 Fast (Reasoning) SpaceXAI81.63%as of 2026-09-26model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.137 pp; $0.002535/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-reasoning"].accuracy; Overall leaderboard rank 49
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Ling 3.0 Flash Ant Group80.90%as of 2026-09-26model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.039 pp; $0.001358/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-2607"].accuracy; Overall leaderboard rank 50
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M2.1 MiniMax80.78%as of 2026-09-26model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.831 pp; $0.005087/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.1"].accuracy; Overall leaderboard rank 51
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5 Mini OpenAI80.58%as of 2026-09-26model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.924 pp; $0.033478/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5-mini-2025-08-07"].accuracy; Overall leaderboard rank 52
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4 Flash 0731 DeepSeek80.36%as of 2026-09-26model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 1.973 pp; $0.014247/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-flash-0731"].accuracy; Overall leaderboard rank 53
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4 Pro 0813 DeepSeek80.17%as of 2026-09-26model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.004 pp; $0.041127/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro-0813"].accuracy; Overall leaderboard rank 54
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M2.7 MiniMax79.87%as of 2026-09-26model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.86 pp; $0.005124/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.7"].accuracy; Overall leaderboard rank 55
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 Fast (Non-Reasoning) SpaceXAI79.72%as of 2026-09-26model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.871 pp; $0.002056/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-non-reasoning"].accuracy; Overall leaderboard rank 56
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.6 Flash Google79.66%as of 2026-09-26model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.861 pp; $0.073178/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.6-flash"].accuracy; Overall leaderboard rank 57
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.7 Max Alibaba79.40%as of 2026-09-26model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 1.907 pp; $0.069070/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.7-max"].accuracy; Overall leaderboard rank 58
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.1 Fast (Reasoning) SpaceXAI78.73%as of 2026-09-26model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.866 pp; $0.002387/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-reasoning"].accuracy; Overall leaderboard rank 59
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Preview (9/25) (Thinking) Google78.50%as of 2026-09-26model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.014526/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 60
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 SpaceXAI78.15%as of 2026-09-26model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.084 pp; $0.063955/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-0709"].accuracy; Overall leaderboard rank 61
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K2.6 Moonshot AI78.15%as of 2026-09-26model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 1.792 pp; $0.055962/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.6"].accuracy; Overall leaderboard rank 62
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Preview (9/25) (Nonthinking) Google77.95%as of 2026-09-26model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.014385/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025"].accuracy; Overall leaderboard rank 63
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.4 (xhigh) OpenAI77.55%as of 2026-09-26model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 3.316 pp; $0.639282/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-2026-03-05"].accuracy; Overall leaderboard rank 64
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.1 Fast Non-Reasoning SpaceXAI77.46%as of 2026-09-26model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.04 pp; $0.001782/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-non-reasoning"].accuracy; Overall leaderboard rank 65
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3 VL Plus Alibaba77.13%as of 2026-09-26model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.020220/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].accuracy; Overall leaderboard rank 66
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5.4 Nano OpenAI77.09%as of 2026-09-26model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.001800/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-nano-2026-03-17"].accuracy; Overall leaderboard rank 67
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.6 Plus Alibaba76.96%as of 2026-09-26model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 1.917 pp; $0.029294/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.6-plus"].accuracy; Overall leaderboard rank 68
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- o3 OpenAI76.65%as of 2026-09-26model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.871 pp; $0.040334/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/o3-2025-04-16"].accuracy; Overall leaderboard rank 69
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.5 Flash Google76.57%as of 2026-09-26model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.166341/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash"].accuracy; Overall leaderboard rank 70
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K2.5 Moonshot AI76.44%as of 2026-09-26model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 1.986 pp; $0.024890/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.5-thinking"].accuracy; Overall leaderboard rank 71
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.1 Pro Preview (02/26) Google76.11%as of 2026-09-26model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-pro-preview"].accuracy; Overall leaderboard rank 72
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5 Anthropic76.05%as of 2026-09-26model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5"].accuracy; Overall leaderboard rank 73
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Lite (9/25) (Nonthinking) Google75.82%as of 2026-09-26model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.851 pp; $0.001332/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].accuracy; Overall leaderboard rank 74
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Ling 3.0 Flash Fin Ant Group75.59%as of 2026-09-26model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 2.026 pp; $0.001618/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-af-rc3"].accuracy; Overall leaderboard rank 75
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4 DeepSeek75.14%as of 2026-09-26model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.002 pp; $0.053954/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro"].accuracy; Overall leaderboard rank 76
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.3 SpaceXAI74.40%as of 2026-09-26model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.019 pp; $0.015293/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.3"].accuracy; Overall leaderboard rank 77
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.1 (Thinking) Anthropic73.90%as of 2026-09-26model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 1.965 pp; $0.263427/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].accuracy; Overall leaderboard rank 78
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Pro Google73.55%as of 2026-09-26model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 1.91 pp; $0.046379/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-pro"].accuracy; Overall leaderboard rank 79
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GPT 5 Nano OpenAI72.86%as of 2026-09-26model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.006961/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/gpt-5-nano-2025-08-07"].accuracy; Overall leaderboard rank 80
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Lite (Nonthinking) Google72.83%as of 2026-09-26model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.982 pp; $0.001211/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite"].accuracy; Overall leaderboard rank 81
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3 Max Thinking Alibaba72.71%as of 2026-09-26model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.905 pp; $0.085327/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-max-2026-01-23"].accuracy; Overall leaderboard rank 82
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4 (Nonthinking) Anthropic72.41%as of 2026-09-26model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.038973/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514"].accuracy; Overall leaderboard rank 83
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.1 zAI72.27%as of 2026-09-26model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.064 pp; $0.023717/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-5.1"].accuracy; Overall leaderboard rank 84
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo V2.5 Xiaomi72.15%as of 2026-09-26model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.851 pp; $0.001351/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5"].accuracy; Overall leaderboard rank 85
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3 Pro (11/25) Google72.04%as of 2026-09-26model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.9 pp; $0.061162/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3-pro-preview"].accuracy; Overall leaderboard rank 86
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.1 (Nonthinking) Anthropic71.75%as of 2026-09-26model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 2.021 pp; $0.187162/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805"].accuracy; Overall leaderboard rank 87
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.5 Flash Lite Google70.89%as of 2026-09-26model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.031 pp; $0.019867/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash-lite"].accuracy; Overall leaderboard rank 88
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.5 Flash Alibaba70.62%as of 2026-09-26model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 2.09 pp; $0.004425/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.5-flash"].accuracy; Overall leaderboard rank 89
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3 Flash (12/25) Google69.92%as of 2026-09-26model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.014379/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3-flash-preview"].accuracy; Overall leaderboard rank 90
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4 (Thinking) Anthropic69.35%as of 2026-09-26model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 2.212 pp; $0.053443/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].accuracy; Overall leaderboard rank 91
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- o4 Mini OpenAI69.14%as of 2026-09-26model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.957 pp; $0.040605/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["openai/o4-mini-2025-04-16"].accuracy; Overall leaderboard rank 92
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 4.7 zAI68.63%as of 2026-09-26model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.123 pp; $0.019082/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["zai/glm-4.7"].accuracy; Overall leaderboard rank 93
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Mistral Medium 3.5 Mistral67.73%as of 2026-09-26model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.011 pp; $0.157641/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["mistralai/mistral-medium-3.5"].accuracy; Overall leaderboard rank 94
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Lite (9/25) (Thinking) Google66.88%as of 2026-09-26model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.002567/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 95
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Laguna M.1 Poolside65.91%as of 2026-09-26model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 2.007 pp; $0.002202/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["poolside/laguna-m.1"].accuracy; Overall leaderboard rank 96
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.1 Flash Lite Preview Google63.90%as of 2026-09-26model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.823 pp; $0.002195/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-flash-lite-preview"].accuracy; Overall leaderboard rank 97
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.20 (Reasoning) SpaceXAI63.41%as of 2026-09-26model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.095 pp; $0.031303/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["grok/grok-4.20-0309-reasoning"].accuracy; Overall leaderboard rank 98
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Laguna XS.2 Poolside61.43%as of 2026-09-26model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 2.349 pp; $0.001126/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["poolside/laguna-xs.2"].accuracy; Overall leaderboard rank 99
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Command A+ Cohere55.68%as of 2026-09-26model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 3.646 pp; $0.140316/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["cohere/command-a-plus-05-2026"].accuracy; Overall leaderboard rank 100
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Mercury 2.5 Inception55.09%as of 2026-09-26model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.095 pp; $0.004476/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["inception/mercury-2.5"].accuracy; Overall leaderboard rank 101
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 4 Maverick Meta54.22%as of 2026-09-26model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.871 pp; $0.002460/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["fireworks/llama4-maverick-instruct-basic"].accuracy; Overall leaderboard rank 102
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 4 Scout Meta50.59%as of 2026-09-26model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].accuracy; Overall leaderboard rank 103
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
- Nemotron 3.5 Lightning NVIDIA4.27%as of 2026-09-26model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 0.492 pp; $0.006241/test; source snapshot 2026-09-26; run date not publishedVals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Embedded BenchmarkView data, tasks.overall["fireworks/nemotron-lightning-3p5-30b-a3b"].accuracy; Overall leaderboard rank 104
- reported by
- benchmark publisher
- published
- 2026-09-26
- retrieved
- 2026-09-28
- confidence
- verified
MedXpertQA (MM)
22 rows · percentage accuracy 0-100Board source: Introducing Muse Spark: Scaling Towards Personal Superintelligence · official page: github.com
- GPT-5.6 Sol OpenAI81.5date not reportedQwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
- reported by
- independent run
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- partial (partial review)
- Gemini 3.1 Pro Google81.3%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Muse Spark Eval Methodology (model card, 81.3, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 81.3%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 1; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Qwen3.8 Max Alibaba80.4%as of 2026-08Alibaba's own Qwen3.8 launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.8-Max; header row: | | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max |
- reported by
- vendor-reported
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- partial (partial review)
also reported in MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 80.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 2; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Claude Fable 5 Anthropic80.0date not reportedQwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
- reported by
- independent run
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- partial (partial review)
- Muse Spark Meta78.4%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- vendor-reported
- published
- 2026-04-08
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Muse Spark Eval Methodology (model card, 78.4, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 78.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 3; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - GPT-5.4 OpenAI77.1%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Muse Spark Eval Methodology (model card, 77.1, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 77.1%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 4; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Gemini 3 Pro Google76.0%date not reportedQwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Alibaba / Qwen
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- independent run
- published
- 2026-02-16
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.2 OpenAI73.3date not reportedQwen-run comparison in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Alibaba / Qwen
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- independent run
- published
- 2026-02-16
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.8 Anthropic71.7date not reportedQwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
- reported by
- independent run
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- partial (partial review)
- Qwen3.7 Plus Alibaba71.0%as of 2026-05Alibaba's own Qwen3.7 Plus launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.Qwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
- publisher
- Alibaba
- locator
- Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.7-Plus; header row: | | GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus |
- reported by
- vendor-reported
- published
- 2026-05-31
- retrieved
- 2026-09-07
- confidence
- partial (partial review)
also reported in Qwen3.8-Max: A New Bar for Coding and Cowork (launch post, 71.0, Qwen3.8 launch post, Multimodal Benchmarks table, row MedXpertQA-MM, column Qwen3.7-Plus (same value carried forward)); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 71.0%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 5; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Qwen3.5 397B A17B Alibaba70.0date not reportedself-reported in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Alibaba / Qwen
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- vendor-reported
- published
- 2026-02-16
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3.6 Plus Alibaba68.7date not reportedQwen-run comparison in the Qwen3.7-Plus launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.Qwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
- publisher
- Alibaba
- locator
- Multimodal Benchmarks table, row MedXpertQA-MM; columns GPT-5.4 (xhigh) / Opus-4.6 Max / Gemini-3.1 Pro / Qwen3.6-Plus / Qwen3.7-Plus
- reported by
- vendor-reported
- published
- 2026-05-31
- retrieved
- 2026-09-07
- confidence
- partial (partial review)
- Grok 4.20 xAI65.8%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Muse Spark Eval Methodology (model card, 65.8, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 65.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 6; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Kimi K2.5 Moonshot AI65.3date not reportedQwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)Qwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Alibaba / Qwen
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- independent run
- published
- 2026-02-16
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.6 Anthropic64.8%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Muse Spark Eval Methodology (model card, 64.8, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 64.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 7; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Claude Opus 4.5 Anthropic63.6%date not reportedQwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Alibaba / Qwen
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- independent run
- published
- 2026-02-16
- retrieved
- 2026-09-28
- confidence
- verified
- Gemma 4 31B Google61.3%date not reportedGoogle model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card · model card · first-party
- publisher
- locator
- Evaluation Results table, MedXPertQA MM row, Gemma 4 31B column
- reported by
- vendor-reported
- published
- 2026-04-02
- retrieved
- 2026-09-28
- confidence
- verified
- Gemma 4 26B A4B Google58.1%date not reportedGoogle model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card · model card · first-party
- publisher
- locator
- Evaluation Results table, MedXPertQA MM row, Gemma 4 26B A4B column
- reported by
- vendor-reported
- published
- 2026-04-02
- retrieved
- 2026-09-28
- confidence
- verified
- Gemma 4 12B Google48.7%as of 2026-04Google's Gemma 4 model card, Unified 12B; protocol not statedGemma 4 model card · model card · first-party
- publisher
- locator
- Benchmark Results table, Vision section, row MedXPertQA MM, column Gemma 4 12B Unified; header row: | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |
- reported by
- vendor-reported
- published
- 2026-04-02
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Gemma 4 Technical Report (paper, 48.7, Gemma 4 Technical Report p. 6, Table 6 (vision benchmarks, thinking), row MedXPertQA MM, column 12B); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 48.7%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 8; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Qwen3-VL-235B-A22B Alibaba47.6%date not reportedQwen3.5 model card comparison; source-specific evaluation, not harmonized across vendorsQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Alibaba / Qwen
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- vendor-reported
- published
- 2026-02-16
- retrieved
- 2026-09-28
- confidence
- verified
- Gemma 4 E4B Google28.7%date not reportedGoogle model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card · model card · first-party
- publisher
- locator
- Evaluation Results table, MedXPertQA MM row, Gemma 4 E4B column
- reported by
- vendor-reported
- published
- 2026-04-02
- retrieved
- 2026-09-28
- confidence
- verified
- Gemma 4 E2B Google23.5%date not reportedGoogle model card; MedXPertQA MM row; vendor-reported; protocol differs from other source familiesGemma 4 model card · model card · first-party
- publisher
- locator
- Evaluation Results table, MedXPertQA MM row, Gemma 4 E2B column
- reported by
- vendor-reported
- published
- 2026-04-02
- retrieved
- 2026-09-28
- confidence
- verified
Artificial Analysis Healthcare & Medical Index
25 rows · index scoreBoard source: Best AI for Healthcare & Medical: LLM Leaderboard
- Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic61as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=claude-opus-5-5, weightedIndex=60.5359874288733; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic58as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=claude-fable-5-1, weightedIndex=58.0046680298741; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic53as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=claude-opus-5, weightedIndex=53.3389883833501; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Astra (max) OpenAI52as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-6-astra, weightedIndex=51.6668965494132; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Spark 1.3 (max) Meta50as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=muse-spark-1-3, weightedIndex=49.6912519301107; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4.7 (xhigh) SpaceXAI47as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=grok-4-7, weightedIndex=47.1903068758764; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GLM-5.3 (max) Z AI47as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=glm-5-3, weightedIndex=46.6708125611712; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Sol (max) OpenAI45as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-5-6-sol, weightedIndex=45.4043575131731; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi K3 (max) Kimi45as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=kimi-k3, weightedIndex=45.0434943845881; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GLM 5.3 Flash Z AI45as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=glm-5-3-flash, weightedIndex=44.8426192521491; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Sol (max) OpenAI43as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-6-sol, weightedIndex=43.4880669048048; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- MiMo-V2.6-Pro Xiaomi42as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=mimo-v2-6-pro, weightedIndex=41.8204296860305; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.8 Flash (high) Google42as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gemini-3-8-flash, weightedIndex=41.7830626872104; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3.8 Max (0902) Alibaba41as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=qwen3-8-max, weightedIndex=41.4279151085327; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Step 5 Preview StepFun41as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=step-5, weightedIndex=40.8171769208551; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V4.1 Flash (Reasoning, Max Effort) DeepSeek41as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=deepseek-v4-1-flash, weightedIndex=40.6388737813166; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-6 Luna (max) OpenAI37as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-6-luna, weightedIndex=36.7252603853009; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.6 Luna (max) OpenAI36as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-5-6-luna, weightedIndex=35.740444759121; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3.8 27B (xhigh) Alibaba34as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=qwen3-8-27b, weightedIndex=34.4310443115205; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- MiniMax-M3 MiniMax30as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=minimax-m3, weightedIndex=29.7042505129902; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Inkling (xhigh) Thinking Machines25as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=inkling, weightedIndex=25.3933378261902; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Nemotron 3 Ultra 550B A55B (Reasoning) NVIDIA23as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=nvidia-nemotron-3-ultra-550b-a55b, weightedIndex=23.0046725383136; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 3.5 Flash-Lite Google23as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=gemini-3-5-flash-lite, weightedIndex=23.074121756108; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Muse Glimmer (high) Meta18as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=muse-glimmer, weightedIndex=17.5980321302316; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
- Mistral Medium 3.5 Mistral14as of 2026-09Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- Embedded chart data: capability=healthcareAndMedical, initialModels slug=mistral-medium-3-5, weightedIndex=13.5556785369912; round to nearest whole point
- reported by
- third-party run
- retrieved
- 2026-09-28
- confidence
- verified
PhysicianBench
21 rows · pass@1 success rate %Board source: PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · official page: arxiv.org
- Claude Opus 5.5 (max) Anthropic68.4%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5.5 (max) Anthropic63.2%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Fable 5.1 (max) Anthropic61.0%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 5 (max) Anthropic57.6%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5.5 (xhigh) Anthropic56.4%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; xhigh effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 5.5 (high) Anthropic47.6%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; high effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.5 OpenAI46.3 ± 1.2as of 2026-05pass@1; Pass^3 28.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 46.3 ± 1.2, arXiv abs page (landing page for the PDF)) - Claude Sonnet 5 (max) Anthropic37.4%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.6 Anthropic31.7 ± 2.3as of 2026-05Pass^3 18.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 31.7 ± 2.3, arXiv abs page (landing page for the PDF)) - Claude Sonnet 5.5 (medium) Anthropic30.0%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; medium effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4.7 Anthropic29.3 ± 2.5as of 2026-05Pass^3 18.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 29.3 ± 2.5, arXiv abs page (landing page for the PDF)) - GPT-5.4 OpenAI27.7 ± 1.5as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 27.7 ± 1.5, arXiv abs page (landing page for the PDF)) - Claude Sonnet 5.5 (low) Anthropic27.2%as of 2026-09-28Anthropic-run pass@1 on 100 tasks; low effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublishedClaude Sonnet 5.5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
- reported by
- vendor-reported
- published
- 2026-09-28
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4.6 Anthropic23.0 ± 2.6as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 23.0 ± 2.6, arXiv abs page (landing page for the PDF)) - DeepSeek V4-Pro DeepSeek18.7 ± 2.9as of 2026-05Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supportedPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
- Kimi-K2.6 Moonshot AI17.0 ± 2.6as of 2026-05open sourcePhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 17.0 ± 2.6, arXiv abs page (landing page for the PDF)) - MiMo-v2.5-Pro Xiaomi16.7 ± 4.0as of 2026-05Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supportedPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3.6-Plus Alibaba13.7 ± 4.0as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 13.7 ± 4.0, arXiv abs page (landing page for the PDF)) - MiniMax M2.7 MiniMax8.7 ± 1.2as of 2026-05Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high reasoning when supportedPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini Pro 3.1 Google6.0 ± 1.0as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 6.0 ± 1.0, arXiv abs page (landing page for the PDF)) - Grok-4.20 xAI5.3 ± 3.2as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper · first-party
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (Proprietary Models block)
- reported by
- benchmark publisher
- published
- 2026-05-04
- retrieved
- 2026-09-28
- confidence
- verified
EHR-Complex
18 rows · exact-match accuracyBoard source: EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · official page: arxiv.org
- GPT-5.4 (high reasoning) OpenAI0.65as of 2026-06average over 12 intent columns; run as human-validation configuration, not in the headline 12-model tableEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.65, arXiv abs page (landing page for the PDF)) - Gemini 3.1 Pro Google0.63as of 2026-06validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.63, arXiv abs page (landing page for the PDF)) - Kimi-K2.5 Moonshot AI0.62as of 2026-06headline 12-model evaluation, top open-weightEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF)) - Qwen3.5-397B Alibaba0.62as of 2026-06headline evaluationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF)) - DeepSeek-V3.2-Exp DeepSeek0.59as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-5.4 (low reasoning) OpenAI0.58as of 2026-06validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.58, arXiv abs page (landing page for the PDF)) - DeepSeek-V3.1 DeepSeek0.56as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3-32B-SFT Alibaba0.55as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training setEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3-235B Alibaba0.53as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-4.1 mini OpenAI0.49as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-4.1 OpenAI0.47as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3-14B-SFT Alibaba0.45as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training setEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4.6 Anthropic0.36as of 2026-06validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.36, arXiv abs page (landing page for the PDF)) - Qwen3-32B Alibaba0.36as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-4o OpenAI0.31as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Pro Google0.31as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3-14B Alibaba0.30as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen3-4B Alibaba0.16as of 2026-06Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turnsEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper · first-party
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3, Avg. column
- reported by
- benchmark publisher
- published
- 2026-06-22
- retrieved
- 2026-09-28
- confidence
- verified
WHBench
22 rows · mean normalized percentage 0-100Board source: WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · official page: arxiv.org
- Claude Opus 4.6 Anthropic72.1%as of 2026-0395% CI 69.6-74.4; evaluations run March 2026WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 72.1%, same paper, abstract page) - Claude Sonnet 4.6 Anthropic67.1%as of 2026-0395% CI 64.5-69.6WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 67.1%, same paper, abstract page) - GPT-5.4 OpenAI66.8%as of 2026-0395% CI 64.5-69.2WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 66.8%, same paper, abstract page) - Gemini 3 Flash Preview Google64.7%as of 2026-03WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 64.7%, same paper, abstract page) - OpenAI o3 OpenAI63.6%as of 2026-0395% bootstrap CI 61.3–65.9; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek V3.2 DeepSeek61.3%as of 2026-0395% bootstrap CI 58.6–63.9; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 3 SpaceX AI60.7%as of 2026-0395% bootstrap CI 58.0–63.4; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Mistral Large Mistral AI60.2%as of 2026-0395% bootstrap CI 57.4–63.0; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Grok 4 SpaceX AI57.9%as of 2026-0395% bootstrap CI 54.9–60.8; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- DeepSeek-R1 DeepSeek52.9%as of 2026-0395% bootstrap CI 50.5–55.3; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-4.1 OpenAI51.8%as of 2026-03WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 51.8%, same paper, abstract page) - Grok 3 Mini SpaceX AI50.0%as of 2026-0395% bootstrap CI 47.5–52.5; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Flash Google49.5%as of 2026-0395% bootstrap CI 47.0–52.0; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Opus 4 Anthropic49.1%as of 2026-0395% bootstrap CI 46.4–51.7; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Claude Sonnet 4 Anthropic48.1%as of 2026-0395% bootstrap CI 45.5–50.6; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- GPT-4o OpenAI44.6%as of 2026-03WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 44.6%, same paper, abstract page) - Llama 4 Maverick Meta42.1%as of 2026-0395% bootstrap CI 39.6–44.6; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Nemotron 70B NVIDIA39.3%as of 2026-0395% bootstrap CI 37.3–41.3; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 3.3 70B Meta37.8%as of 2026-0395% bootstrap CI 35.2–40.5; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 3.1 405B Meta36.1%as of 2026-0395% bootstrap CI 33.9–38.3; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Gemini 2.5 Pro Google35.3%as of 2026-0395% bootstrap CI 32.7–38.1; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
- Llama 4 Scout Meta35.2%as of 2026-0395% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-bookWHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · paper · first-party
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- benchmark publisher
- published
- 2026-07-23
- retrieved
- 2026-09-28
- confidence
- verified
HealthAdminBench
7 rows · percentage end-to-end task success 0-100Board source: HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · official page: healthadminbench.stanford.edu · paper: arxiv.org
- Claude Opus 4.6 (computer-use agent) Anthropic36.3%as of 2026-04screenshot-only, task description + portal guidance; native CUA harness; subtask rate 78.4%HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 36.3%, Section 'LLMs struggle with long-horizon tasks') - GPT-5.4 (computer-use agent) OpenAI26.7%as of 2026-04screenshot-only, task description + portal guidance; subtask rate 82.8%HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 26.7%, Section 'LLMs struggle with long-horizon tasks') - Kimi K2.5 Moonshot AI15.6%as of 2026-04screenshot-only, task description + portal guidanceHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 15.6%, Section 'LLMs struggle with long-horizon tasks') - Claude Opus 4.6 (standardized harness) Anthropic14.8%as of 2026-04screenshot-only, task description + portal guidance; authors' standardized harness, no native CUAHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
- Qwen 3.5 Alibaba13.3%as of 2026-04screenshot-only, task description + portal guidanceHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 13.3%, Section 'LLMs struggle with long-horizon tasks') - Gemini 3.1 Pro Google11.9%as of 2026-04screenshot-only, task description + portal guidanceHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 11.9%, Section 'LLMs struggle with long-horizon tasks') - GPT-5.4 (standardized harness) OpenAI5.9%as of 2026-04screenshot-only, task description + portal guidance; authors' standardized harness, no native CUAHealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · paper · first-party
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- benchmark publisher
- published
- 2026-04-10
- retrieved
- 2026-09-28
- confidence
- verified
Documents on file
The distinct documents the sections above draw on, with the number of index rows each one backs.
Corrections go through the same route as everything else here: a better document replaces a weaker one, the row's confidence moves, and the change is dated on the updates page. The full record, sources included, is downloadable from the data page.