Health Evals

Sources

For the featured medication-safety benchmark, see MedPIC-Bench sources and verification. The collection below documents the wider healthcare benchmark index.

503 of 503 rows documented · 49 documents, 46 first-party · snapshot reviewed September 28, 2026

This page records the published citations associated with each indexed score, including review status and any remaining verification gaps. A row records where its score was read: the publication itself, the passage or table it sits in when that has been captured, who published it, when the index retrieved it, and how far the reading has been checked. First-party documents rank highest, meaning a vendor's system card for a vendor-reported number or the maintainers' own board for a leaderboard number; independent evaluators are primary sources for their own runs. Sister-site compilations and secondary citations are labeled where the original result remains unchecked. Rows whose document has not been located yet say so instead of disappearing.

Better citations do not make the boards comparable. Each benchmark below keeps its own scale, grader and task set, so a score on one section says nothing about a score on the next, and no number on this page should be lined up against a number from another section. The comparability rules are on the methodology page.

Reading the confidence field

HealthBench Professional

29 rows · 0 to 1

Board source: Published model evaluation reports · official page: arxiv.org · full board: healthbenchprofessional.com

  1. GPT-6 Astra (Anthropic run) OpenAI0.703as of 2026-09
    Anthropic reproduction of GPT-6 Astra through the public API; max effort; no system prompt; Claude Opus 4.8 grader; length-adjusted 70.3%, raw 74.0%. Different grader/protocol from OpenAI’s own 64.7% report.
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    Section 8.15, pp. 137–138; text above Figure 8.15.B: max-effort Astra 70.3 adjusted and 74.0 raw
    reported by
    third-party run
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  2. Claude Sonnet 5.5 Anthropic0.692as of 2026-09
    Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    Section 8.15.2; pp. 137–139, Figure 8.15.B (max effort)
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  3. Claude Opus 5.5 Anthropic0.656as of 2026-09
    Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 77.1%. Five-trial average; refusal fallback to Claude Opus 5.
    Claude Opus 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    Section 8.15.2; pp. 213–214, Figures 8.15.1.A and 8.15.2.A
    reported by
    vendor-reported
    published
    2026-09-22
    retrieved
    2026-09-28
    confidence
    verified
  4. GPT-6 Astra OpenAI0.647as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 68.2%, mean answer 3,185 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench Professional length-adjusted row, GPT-6 Astra column; value 64.7 (68.2, 3185)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  5. Claude Fable 5 Anthropic0.633as of 2026-09
    Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 68.9%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
    publisher
    Anthropic
    locator
    p. 199, Figure 8.17.2.A; Claude Fable 5, adjusted bar 63.3%; also p. 167 Table 8.1.A
    reported by
    vendor-reported
    published
    2026-09-01
    retrieved
    2026-09-28
    confidence
    verified
  6. Claude Fable 5.1 Anthropic0.621as of 2026-09
    length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
    publisher
    Anthropic
    locator
    p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)
    reported by
    vendor-reported
    published
    2026-09-01
    retrieved
    2026-09-28
    confidence
    verified
  7. GPT-6 Sol OpenAI0.608as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 59.5%, mean answer 1,573 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench Professional length-adjusted row, GPT-6 Sol column; value 60.8 (59.5, 1573)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  8. GPT-6 Luna OpenAI0.608as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 61.2%, mean answer 2,119 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench Professional length-adjusted row, GPT-6 Luna column; value 60.8 (61.2, 2119)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  9. GPT-5.6 Sol OpenAI0.605as of 2026-06
    length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 60.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL)
  10. Claude Opus 5 Anthropic0.598as of 2026-07
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    System Card: Claude Opus 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
    reported by
    vendor-reported
    published
    2026-07-24
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 59.8%, p. 167, Table 8.1.A, column 'Claude Opus 5'; raw 73.4% also reprinted on p. 199, sec. 8.17.2)
  11. Muse Spark 1.1 Meta0.593as of 2026-07
    length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
    Muse Spark 1.1 Evaluation Report · model card · first-party
    publisher
    Meta
    locator
    p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
  12. Claude Sonnet 5 Anthropic0.578as of 2026-06
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    System Card: Claude Sonnet 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-28
    confidence
    verified
    also reported in System Card: Claude Opus 5 (system card, 57.8%, p. 189, section 8.15.2, Figure 8.15.2.A)
  13. GPT-5.6 Terra OpenAI0.577as of 2026-06
    length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA)
  14. Claude Opus 4.8 Anthropic0.574as of 2026-06
    Length-adjusted; Anthropic June evaluation, adaptive max effort, Claude Opus 4.8 grader, five-trial average, no tools or custom system prompt. Earlier May report was 55.8% using Claude Sonnet 4.6 as grader; the grader changed.
    Claude Sonnet 5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    p. 139, Figure 8.12.2.A; Opus 4.8 bar 57.4%
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-28
    confidence
    verified
  15. Grok 4.7 xAI0.567as of 2026-09
    SpaceXAI vendor report at xhigh effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.
    Introducing Grok 4.7 · launch post · first-party
    publisher
    SpaceXAI
    locator
    Model Improvements table, Clinical reasoning / HealthBench Professional row, Grok 4.7 xhigh column
    reported by
    vendor-reported
    published
    2026-09-21
    retrieved
    2026-09-28
    confidence
    partial (partial review)
  16. GPT-5.6 Luna OpenAI0.557as of 2026-06
    length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 55.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA)
  17. Muse Spark Meta0.541as of 2026-07
    length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    Muse Spark 1.1 Evaluation Report · model card · first-party
    publisher
    Meta
    locator
    p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
  18. GPT-5.6 Sol (August) OpenAI0.540as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-28
    confidence
    verified
  19. Claude Opus 4.7 Anthropic0.519as of 2026-05
    length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
    System Card: Claude Opus 4.8 · system card · first-party
    publisher
    Anthropic
    locator
    p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
    reported by
    vendor-reported
    published
    2026-05-28
    retrieved
    2026-09-28
    confidence
    verified
  20. GPT-5.5 OpenAI0.518as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 51.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5)
  21. Grok 4.6 xAI0.485as of 2026-09
    SpaceXAI vendor report at high effort. HealthBench Professional release-table score; the page does not specify grader or explicitly label length adjustment, so exact protocol comparability is unconfirmed.
    Introducing Grok 4.7 · launch post · first-party
    publisher
    SpaceXAI
    locator
    Model Improvements table, Clinical reasoning / HealthBench Professional row, Grok 4.6 high column
    reported by
    vendor-reported
    published
    2026-09-21
    retrieved
    2026-09-28
    confidence
    partial (partial review)
  22. GPT-5.4 OpenAI0.481as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 48.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4)
  23. GPT-5 OpenAI0.462as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 46.2, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5)
  24. GPT-5.2 OpenAI0.459as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 45.9, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2)
  25. Claude Sonnet 4.6 Anthropic0.442as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
    System Card: Claude Sonnet 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-28
    confidence
    verified
    also reported in System Card: Claude Opus 4.8 (system card, 41.7%, conflicting, p. 228, section 8.14.1)
  26. GPT-5.6 Luna (August) OpenAI0.441as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-28
    confidence
    verified
  27. GPT-5.1 OpenAI0.396as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 39.6, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1)
  28. GPT-5.5 Instant OpenAI0.384as of 2026-05
    length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    GPT-5.5 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 38.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant)
  29. MAI-Thinking-1 Microsoft0.350as of 2026-08
    length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.
    publisher
    Microsoft AI
    locator
    p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.
    reported by
    vendor-reported
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified

HealthBench Hard

20 rows · 0 to 1

Board source: healthbenchhard.ai · paper: arxiv.org · full board: healthbenchhard.ai

  1. Baichuan-M3 Baichuan0.444as of 2026-02
    Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.
    Baichuan-M3 Technical Report · paper · first-party
    publisher
    Baichuan
    locator
    p. 23 (PDF page index 22), section 4.2.1 and Figure 7; HealthBench Hard
    reported by
    vendor-reported
    published
    2026-02-06
    retrieved
    2026-09-28
    confidence
    verified
  2. Muse Spark Meta0.428as of 2026-04
    raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.
    reported by
    vendor-reported
    published
    2026-04-08
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Muse Spark Eval Methodology (model card, 42.8, p. 5 results table image, row HealthBench Hard; protocol p. 2)
  3. GPT-5.2-High (Baichuan run) OpenAI0.420as of 2026-02
    Raw/unadjusted HealthBench Hard score as evaluated by Baichuan in its M3 report; distinct from OpenAI’s length-adjusted protocol.
    Baichuan-M3 Technical Report · paper · first-party
    publisher
    Baichuan
    locator
    p. 23 (PDF page index 22), section 4.2.1 and Figure 7; HealthBench Hard
    reported by
    third-party run
    published
    2026-02-06
    retrieved
    2026-09-28
    confidence
    verified
  4. GPT-6 Astra OpenAI0.366as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 34.2%, mean answer 1,697 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench Hard length-adjusted row, GPT-6 Astra column; value 36.6 (34.2, 1697)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  5. GPT-5 OpenAI0.347as of 2026-06
    length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 34.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5); GPT-5.5 System Card (system card, 34.7, Section 5 Health, Table 7, column GPT-5); GPT-5 System Card (system card, 46.2, conflicting, p. 18, section 3.10 Health, Figure 6 (HealthBench Hard, raw score %))
  6. GPT-5.2 OpenAI0.343as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 34.3, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2)
  7. GPT-5.6 Sol OpenAI0.331as of 2026-06
    length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 33.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL)
  8. GPT-5.6 Terra OpenAI0.327as of 2026-06
    length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 32.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA)
  9. GPT-5.6 Luna OpenAI0.320as of 2026-06
    length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 32.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA)
  10. GPT-5.5 OpenAI0.315as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 31.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5)
  11. GPT-5.6 Sol (August) OpenAI0.314as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-28
    confidence
    verified
  12. GPT-6 Luna OpenAI0.314as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 25.4%, mean answer 1,241 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench Hard length-adjusted row, GPT-6 Luna column; value 31.4 (25.4, 1241)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  13. GPT-6 Sol OpenAI0.301as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 22.1%, mean answer 974 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench Hard length-adjusted row, GPT-6 Sol column; value 30.1 (22.1, 974)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  14. GPT OSS 120B OpenAI0.300as of 2025-08
    raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-28
    confidence
    verified
  15. GPT-5.4 OpenAI0.291as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 29.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4)
  16. GPT-5.6 Luna (August) OpenAI0.287as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-28
    confidence
    verified
  17. GPT-5.3 Chat OpenAI0.259as of 2026-03
    raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
    GPT-5.3 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 3: HealthBench, row Hard, column GPT-5.3-INSTANT
    reported by
    vendor-reported
    published
    2026-03-02
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.5 Instant System Card (system card, 20.2, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT); GPT-5.6 - August Updates (system card addendum) (system card, 20.2, conflicting, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.3 Instant)
  18. GPT-5.1 OpenAI0.254as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 25.4, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1)
  19. GPT-5.5 Instant OpenAI0.229as of 2026-05
    length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    GPT-5.5 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 22.9, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant)
  20. GPT OSS 20B OpenAI0.108as of 2025-08
    raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-28
    confidence
    verified

HealthBench

26 rows · 0-100 rubric-point percentage (some sites display 0-1)

Board source: Published model evaluation reports · official page: openai.com

  1. Claude Sonnet 5.5 Anthropic65.4as of 2026-09
    Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 69.4%. The paper reports raw and adjusted scores separately; the max-effort HealthBench chart label is 65.4%.
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    Section 8.15.1; pp. 137–139, Figure 8.15.B (max effort)
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  2. Baichuan-M3 Baichuan65.1as of 2026-02
    self-run in Baichuan-M3 paper (arXiv 2602.06570)
    Baichuan-M3 Technical Report · paper · first-party
    publisher
    Baichuan
    locator
    p. 23, section 4.2.1 HealthBench-Main; Table on p. 25 (Model / HealthBench Score)
    reported by
    vendor-reported
    published
    2026-02-06
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Baichuan-M3 Technical Report (paper, 65.1, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (Baichuan-M3-235B, HealthBench Score 65.1))
  3. GPT-5.2-High OpenAI63.3as of 2026-02
    raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card.
    Baichuan-M3 Technical Report · paper · first-party
    publisher
    Baichuan
    locator
    p. 25, HealthBench-Hallu table (Model / HealthBench Score column); also p. 23 prose
    reported by
    independent run
    published
    2026-02-06
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Baichuan-M3 Technical Report (paper, 63.3, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (GPT-5.2-High, HealthBench Score 63.3)); GPT-5.6 System Card (system card, 56.8, conflicting, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2)
  4. Claude Opus 5.5 Anthropic60.6as of 2026-09
    Anthropic evaluation; length-adjusted; adaptive thinking at max effort; Claude Opus 4.8 grader; no tools or custom system prompt; safety classifiers enabled. Raw score 68.1%. Five-trial average; refusal fallback to Claude Opus 5.
    Claude Opus 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    Section 8.15.1; pp. 213–214, Figures 8.15.1.A and 8.15.2.A
    reported by
    vendor-reported
    published
    2026-09-22
    retrieved
    2026-09-28
    confidence
    verified
  5. Claude Fable 5 Anthropic60.4as of 2026-09
    Anthropic evaluation of Claude Fable 5, length-adjusted; adaptive max effort, Opus 4.8 grader, five trials, no tools or custom system prompt; raw 61.2%. Explicit Fable 5 figure in the September 1 card; earlier catalog value came from a Mythos 5 column and is not used for Fable 5.
    publisher
    Anthropic
    locator
    p. 198, Figure 8.17.1.A; Claude Fable 5, adjusted bar 60.4%
    reported by
    vendor-reported
    published
    2026-09-01
    retrieved
    2026-09-28
    confidence
    verified
  6. Claude Fable 5.1 Anthropic60%as of 2026-09
    length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%).
    publisher
    Anthropic
    locator
    p. 198, sec. 8.17.1 / Figure 8.17.1.A
    reported by
    vendor-reported
    published
    2026-09-01
    retrieved
    2026-09-28
    confidence
    verified
  7. Claude Opus 4.8 Anthropic59.3as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itself
    publisher
    Anthropic
    locator
    p. 252, Table 8.1.A, row HealthBench, column Opus 4.8
    reported by
    vendor-reported
    published
    2026-06-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in System Card: Claude Opus 5 (system card, 59.3%, p. 188, section 8.15.1, Figure 8.15.1.A)
  8. Claude Sonnet 5 Anthropic58.7%as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 card
    System Card: Claude Sonnet 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 138, section 8.12.1 HealthBench results, Figure 8.12.1.A bar label (no prose or table number)
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-28
    confidence
    verified
    also reported in System Card: Claude Opus 5 (system card, 58.7%, p. 188, section 8.15.1, Figure 8.15.1.A)
  9. GPT-6 Astra OpenAI58.3as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 56.9%, mean answer 1,760 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench length-adjusted row, GPT-6 Astra column; value 58.3 (56.9, 1760)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  10. Claude Opus 5 Anthropic57.8as of 2026-07
    Anthropic evaluation; length-adjusted 57.8%, raw 67.1%; adaptive max effort, Opus 4.8 grader, five-trial average, no tools or customized system prompt. Previously this catalog displayed the raw value.
    System Card: Claude Opus 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 188, section 8.15.1 and Figure 8.15.1.A; adjusted 57.8%, raw 67.1%
    reported by
    vendor-reported
    published
    2026-07-24
    retrieved
    2026-09-28
    confidence
    verified
  11. GPT-5 OpenAI57.7as of 2026-06
    OpenAI length-adjusted score, maximum reasoning effort; raw 63.1%, mean answer 2,904 characters; GPT-5.6 card Table 6.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5 column
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
  12. GPT OSS 120B OpenAI57.6as of 2025-08
    reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9)
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-120b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-28
    confidence
    verified
  13. GPT-5.6 Sol OpenAI57.0as of 2026-06
    length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL)
  14. GPT-5.6 Terra OpenAI57.0as of 2026-06
    length-adjusted (58.7 unadjusted), max reasoning effort
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA)
  15. GPT-5.2 OpenAI56.8as of 2026-06
    OpenAI length-adjusted score, maximum reasoning effort; raw 60.7%, mean answer 2,645 characters; GPT-5.6 card Table 6.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5.2 column
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
  16. GPT-5.5 OpenAI56.5as of 2026-04
    length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system card
    GPT-5.5 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5 Health, Table 7 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    reported by
    vendor-reported
    published
    2026-04-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 56.5); GPT-5.6 System Card (system card, 56.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5)
  17. GPT-5.6 Luna OpenAI55.8as of 2026-06
    length-adjusted (55.4 unadjusted), max reasoning effort
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 55.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA)
  18. GPT-5.6 Sol (August) OpenAI55.0as of 2026-08
    ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 Preview System Card (system card, 55.0)
  19. GPT-6 Luna OpenAI54.5as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 50%, mean answer 1,255 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench length-adjusted row, GPT-6 Luna column; value 54.5 (50, 1255)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  20. GPT-5.3 Chat OpenAI54.1%as of 2026-03
    raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted)
    GPT-5.3 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 3: HealthBench, row HealthBench, column GPT-5.3-INSTANT
    reported by
    vendor-reported
    published
    2026-03-02
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.5 Instant System Card (system card, 49.6, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT)
  21. GPT-5.4 OpenAI54.0as of 2026-06
    OpenAI length-adjusted score, maximum reasoning effort; raw 55.7%, mean answer 2,275 characters; GPT-5.6 card Table 6.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5.4 column
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
  22. GPT-5.6 Luna (August) OpenAI53.3as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-28
    confidence
    verified
  23. GPT-6 Sol OpenAI53.2as of 2026-09
    OpenAI evaluation; length-adjusted at maximum reasoning effort; raw 47.1%, mean answer 977 characters. September 22, 2026 system-card revision; the Astra values correct a prior evaluation misconfiguration.
    publisher
    OpenAI
    locator
    Section 11.4.1, Table 29; HealthBench length-adjusted row, GPT-6 Sol column; value 53.2 (47.1, 977)
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-28
    confidence
    verified
  24. GPT-5.5 Instant OpenAI51.4as of 2026-05
    length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11
    GPT-5.5 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-28
    confidence
    verified
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 51.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant)
  25. GPT-5.1 OpenAI50.9as of 2026-06
    OpenAI length-adjusted score, maximum reasoning effort; raw 64.2%, mean answer 4,222 characters; GPT-5.6 card Table 6.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1, Table 6, HealthBench length-adjusted row, GPT-5.1 column
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-28
    confidence
    verified
  26. GPT OSS 20B OpenAI42.5as of 2025-08
    reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8)
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-20b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-28
    confidence
    verified

Health Optimization Bench

16 rows · 0-100 rubric credit

Board source: healthoptimizationbench.com · official page: healthoptimizationbench.com · full board: healthoptimizationbench.com

  1. Claude Fable 5 Anthropic70.9as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 68.0–73.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Claude Fable 5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.709, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  2. Claude Opus 5 Anthropic69.3as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 66.2–72.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Claude Opus 5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.693, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  3. Grok 4.6 xAI66.8as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.8–69.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Grok 4.6; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.668, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  4. GPT-5.6 Sol (max) OpenAI66.6as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort.
    publisher
    Arcophos
    locator
    Subject suites table; GPT-5.6 Sol (max); tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.666, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  5. GPT-5.6 Sol (high) OpenAI64.6as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort.
    publisher
    Arcophos
    locator
    Subject suites table; GPT-5.6 Sol (high); tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.646, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  6. Kimi K3 Moonshot AI59.9as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Kimi K3; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.599, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  7. Muse Spark Meta57.2as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Muse Spark; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.572, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  8. Claude Fable 5.1 Anthropic47.3as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 42.2–52.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Vendor safeguards declined 97/257 tasks, scored with no credit; mean over answered tasks is 75.9.
    publisher
    Arcophos
    locator
    Subject suites table; Claude Fable 5.1; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.473, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  9. Gemini 3.6 Google39.7as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 36.2–43.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Gemini 3.6; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.397, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  10. Inkling Thinking Machines35.6as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 32.3–39.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Inkling; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.356, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  11. Claude Sonnet 5 Anthropic34.6as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Claude Sonnet 5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.346, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  12. GLM 5.2 Zhipu20.7as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 18.1–23.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; GLM 5.2; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.207, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  13. MiniMax M3 MiniMax18.3as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.7–20.9. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; MiniMax M3; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.183, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  14. MAI Thinking Microsoft AI17.5as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; MAI Thinking; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.175, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  15. Mistral Medium 3.5 Mistral9.2as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 7.4–11.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Mistral Medium 3.5; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.092, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified
  16. Nemotron 3.5 Lightning NVIDIA4.9as of 2026-09
    Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 3.6–6.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
    publisher
    Arcophos
    locator
    Subject suites table; Nemotron 3.5 Lightning; tasksets/mb2ev/analysis.json snapshot 2026-09-10; mean 0.049, n=257
    reported by
    Arcophos run
    published
    2026-09-10
    retrieved
    2026-09-28
    confidence
    verified

MAST (Medical AI Superintelligence Test)

8 rows · percentage composite

Board source: MAST: Medical AI Superintelligence Test leaderboard (General board)

  1. GPT-5.6 Sol OpenAI60.2%as of 2026-08
    MAST in preview; 'exact scores may change'
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  2. Kimi K3 Moonshot AI60.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  3. Gemini 3.6 Flash Google59.3%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  4. Gemini 3.1 Pro Google58.9%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  5. Qwen3.5 397B A17B Alibaba57.9%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  6. Claude Opus 5 Anthropic57.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  7. Claude Sonnet 5 Anthropic56.6%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  8. Grok 4.3 xAI53.7%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified

MedHELM

10 rows · mean win rate 0-1

Board source: MedHELM leaderboard (medhelm.org), v5.0.0

  1. Gemini 3.1 Pro (Preview) Google0.652as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6520833333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  2. Gemini 3.5 Flash Google0.642as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  3. Muse Spark (2026-04-08) Meta0.621as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6208333333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  4. GPT-5.4 mini OpenAI0.552as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5520833333333334, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  5. GPT-5.4 (2026-03-05) OpenAI0.538as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5375, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  6. Gemini 2.5 Pro Google0.529as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5291666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  7. DeepSeek R1 DeepSeek0.485as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.48541666666666666, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  8. Claude 4.6 Opus Anthropic0.456as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45625, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  9. Claude 3.7 Sonnet Anthropic0.45as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  10. Gemini 2.0 Flash Google0.342as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-28
    confidence
    verified
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.3416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')

First, Do NOHARM (v2)

17 rows · percentage safety score

Board source: MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · paper: arxiv.org

  1. LiSA 2.5 AMBOSS86.2as of 2026-09
    First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
    ARISE MAST technical leaderboard · official leaderboard · first-party
    publisher
    ARISE
    locator
    First Do NOHARM v2 overall leaderboard; LiSA 2.5; score 86.2%
    reported by
    benchmark-owner run
    retrieved
    2026-09-28
    confidence
    verified
  2. Doximity Ask 6.1 Doximity84.5as of 2026-09
    First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
    ARISE MAST technical leaderboard · official leaderboard · first-party
    publisher
    ARISE
    locator
    First Do NOHARM v2 overall leaderboard; Doximity Ask 6.1; score 84.5%
    reported by
    benchmark-owner run
    retrieved
    2026-09-28
    confidence
    verified
  3. OpenEvidence OpenEvidence80.0as of 2026-09
    First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
    ARISE MAST technical leaderboard · official leaderboard · first-party
    publisher
    ARISE
    locator
    First Do NOHARM v2 overall leaderboard; OpenEvidence; score 80%
    reported by
    benchmark-owner run
    retrieved
    2026-09-28
    confidence
    verified
  4. Muse Spark 1.1 Meta79.7%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  5. Glass 5.6 Max Glass Health79.7as of 2026-09
    First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. RAG clinical system, not a base-model-only result. Observed September 28, 2026; the individual run date is not published.
    ARISE MAST technical leaderboard · official leaderboard · first-party
    publisher
    ARISE
    locator
    First Do NOHARM v2 overall leaderboard; Glass 5.6 Max; score 79.7%
    reported by
    benchmark-owner run
    retrieved
    2026-09-28
    confidence
    verified
  6. Claude Opus 5 Anthropic74.6%as of 2026-08
    v2 run on ARISE; 19 models on the board
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  7. Kimi K3 Moonshot AI74.0%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  8. GPT-5.6 Sol OpenAI70.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  9. GPT-5.5 OpenAI70.0%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  10. GPT-5 OpenAI68.6%as of 2026-08
    from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, Model Leaderboard (Top 10 shown), row 10, SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  11. Claude Fable 5 Anthropic65.0%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  12. Gemini 3.1 Pro Google62.6%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  13. Gemini 2.5 Pro Google61.9%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  14. Qwen3.5 397B A17B Alibaba61.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  15. Kimi K2.6 Moonshot AI59.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified
  16. GLM 5.1 Z.ai57.9as of 2026-09
    First Do NOHARM v2 overall safety score on ARISE’s technical leaderboard; preview benchmark. Open-weight model as listed by the evaluator. Observed September 28, 2026; the individual run date is not published.
    ARISE MAST technical leaderboard · official leaderboard · first-party
    publisher
    ARISE
    locator
    First Do NOHARM v2 overall leaderboard; GLM 5.1; score 57.9%
    reported by
    benchmark-owner run
    retrieved
    2026-09-28
    confidence
    verified
  17. DeepSeek R1 DeepSeek55.8%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-28
    confidence
    verified

HealthAgentBench

12 rows · mean task success rate

Board source: HealthAgentBench leaderboard · paper: arxiv.org

  1. Claude Code (Opus 5) Anthropic55%as of 2026-07
    $3.3/task; harness+model evaluated jointly
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 1, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 55%, Detailed results table, rank 1)
  2. Codex (GPT-5.6-sol) OpenAI45%as of 2026-07
    $5.2/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 2, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 45%, Detailed results table, rank 2)
  3. Codex (GPT 5.5) OpenAI42%as of 2026-07
    $2.8/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 3, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 42%, Detailed results table, rank 3); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 42%, p. 11, Figure 4 (pooled task success rate, ten agents))
  4. Copilot (Opus 4.8) Microsoft/Anthropic36%as of 2026-07
    $3.1/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 4, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 36%, Detailed results table, rank 4); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 36%, p. 11, Figure 4 (pooled task success rate, ten agents))
  5. Copilot (GPT 5.5) Microsoft/OpenAI35%as of 2026-07
    $2.6/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 5, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 35%, Detailed results table, rank 5); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 35%, p. 11, Figure 4 (pooled task success rate, ten agents))
  6. Claude Code (Opus 4.8) Anthropic32%as of 2026-07
    $4.0/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 6, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 32%, Detailed results table, rank 6); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 32%, p. 11, Figure 4 (pooled task success rate, ten agents))
  7. Codex (GPT 5.4) OpenAI28%as of 2026-07
    $1.3/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 7, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 28%, Detailed results table, rank 7); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 28%, p. 11, Figure 4 (pooled task success rate, ten agents))
  8. Claude Code (Opus 4.7) Anthropic27%as of 2026-07
    $4.8/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 8, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 27%, Detailed results table, rank 8); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 27%, p. 11, Figure 4 (pooled task success rate, ten agents))
  9. Codex (GPT 5.3) OpenAI22%as of 2026-07
    $1.0/task; harness and model evaluated jointly
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Homepage leaderboard, rank 9, success rate and cost/task
    reported by
    benchmark publisher
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
  10. Claude Code (Opus 4.6) Anthropic19%as of 2026-07
    $4.1/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 10, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 19%, Detailed results table, rank 10); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 19%, p. 11, Figure 4 (pooled task success rate, ten agents))
  11. Claude Code (Sonnet 4.6) Anthropic17%as of 2026-07
    $2.9/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 11, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 17%, Detailed results table, rank 11); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 17%, p. 11, Figure 4 (pooled task success rate, ten agents))
  12. Codex (GPT 5.4 Mini) OpenAI16%as of 2026-07
    $0.6/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 12, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-28
    confidence
    verified
    also reported in HealthAgentBench detailed results (official leaderboard, 16%, Detailed results table, rank 12); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 16%, p. 11, Figure 4 (pooled task success rate, ten agents))

CHI-Bench

44 rows · pass@1 with binary 0/1 reward

Board source: CHI-Bench leaderboard (actAVA)

  1. erius + claude-opus-5 Humana (harness) / Anthropic (model)54.7%as of 2026-07-26
    All Domains pass@1; PA 72.0%, UM 36.0%, CM 56.0%; submitted 2026-07-26; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 1, Accuracy column; submission date 2026-07-26
    reported by
    independent run
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  2. erius + claude-opus-4-8 Humana (harness) / Anthropic (model)37.3%as of 2026-06-05
    All Domains pass@1; PA 40.0%, UM 16.0%, CM 56.0%; submitted 2026-06-05; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 2, Accuracy column; submission date 2026-06-05
    reported by
    independent run
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  3. claude-code + claude-opus-5 Anthropic37.3%as of 2026-07-24
    All Domains pass@1; PA 20.0%, UM 32.0%, CM 60.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 3, Accuracy column; submission date 2026-07-24
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  4. claude-code + claude-opus-4-8 Anthropic33.3%as of 2026-05-28
    All Domains pass@1; PA 32.0%, UM 28.0%, CM 40.0%; submitted 2026-05-28; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 4, Accuracy column; submission date 2026-05-28
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  5. claude-code + claude-opus-4-6 Anthropic28.0%as of 2026-05-01
    All Domains pass@1; PA 20.0%, UM 36.0%, CM 28.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 5, Accuracy column; submission date 2026-05-01
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  6. claude-code + claude-sonnet-4-6 Anthropic26.2%as of 2026-05-01
    All Domains pass@1; PA 24.0%, UM 34.7%, CM 20.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 6, Accuracy column; submission date 2026-05-01
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  7. codex + gpt-5.6-sol OpenAI25.3%as of 2026-07-24
    All Domains pass@1; PA 36.0%, UM 28.0%, CM 12.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 7, Accuracy column; submission date 2026-07-24
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  8. openai-agents + kimi-k3 Moonshot AI25.3%as of 2026-07-24
    All Domains pass@1; PA 28.0%, UM 32.0%, CM 16.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 8, Accuracy column; submission date 2026-07-24
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  9. claude-code + claude-opus-4-7 Anthropic24.4%as of 2026-05-01
    All Domains pass@1; PA 24.0%, UM 17.3%, CM 32.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 9, Accuracy column; submission date 2026-05-01
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  10. claude-code + claude-fable-5 Anthropic24.0%as of 2026-07-22
    All Domains pass@1; PA 24.0%, UM 24.0%, CM 24.0%; submitted 2026-07-22; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 10, Accuracy column; submission date 2026-07-22
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  11. hermes + MedGuard MedGuard22.7%as of 2026-07-06
    All Domains pass@1; PA 4.0%, UM 4.0%, CM 60.0%; submitted 2026-07-06; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 11, Accuracy column; submission date 2026-07-06
    reported by
    community submission
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  12. codex + gpt-5.5 OpenAI20.9%as of 2026-05-01
    All Domains pass@1; PA 29.3%, UM 32.0%, CM 1.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 12, Accuracy column; submission date 2026-05-01
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  13. claude-code + claude-sonnet-5 Anthropic20.0%as of 2026-07-06
    All Domains pass@1; PA 24.0%, UM 24.0%, CM 12.0%; submitted 2026-07-06; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 13, Accuracy column; submission date 2026-07-06
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  14. openai-agents + glm-5.1 Zhipu AI18.7%as of 2026-05-01
    All Domains pass@1; PA 18.7%, UM 33.3%, CM 4.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 14, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  15. hermes + glm-5.1 Zhipu AI18.7%as of 2026-05-01
    All Domains pass@1; PA 10.7%, UM 34.7%, CM 10.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 15, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  16. openai-agents + glm-5.2 Zhipu AI18.7%as of 2026-07-06
    All Domains pass@1; PA 20.0%, UM 32.0%, CM 4.0%; submitted 2026-07-06; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 16, Accuracy column; submission date 2026-07-06
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  17. openclaw + claude-opus-4-7 Anthropic17.3%as of 2026-05-01
    All Domains pass@1; PA 18.7%, UM 13.3%, CM 20.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 17, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  18. openclaw + glm-5.1 Zhipu AI16.9%as of 2026-05-01
    All Domains pass@1; PA 13.3%, UM 26.7%, CM 10.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 18, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  19. hermes + qwen-3.6-max Alibaba16.4%as of 2026-05-01
    All Domains pass@1; PA 9.3%, UM 26.7%, CM 13.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 19, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  20. codex + gpt-5.4 OpenAI16.0%as of 2026-05-01
    All Domains pass@1; PA 24.0%, UM 17.3%, CM 6.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 20, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  21. openai-agents + qwen-3.6-max Alibaba15.6%as of 2026-05-01
    All Domains pass@1; PA 16.0%, UM 26.7%, CM 4.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 21, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  22. hermes + kimi-k2.6 Moonshot AI15.6%as of 2026-05-01
    All Domains pass@1; PA 18.7%, UM 21.3%, CM 6.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 22, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  23. openai-agents + kimi-k2.6 Moonshot AI15.1%as of 2026-05-01
    All Domains pass@1; PA 17.3%, UM 25.3%, CM 2.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 23, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  24. openai-agents + deepseek-v4-pro DeepSeek14.2%as of 2026-05-01
    All Domains pass@1; PA 10.7%, UM 28.0%, CM 4.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 24, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  25. hermes + deepseek-v4-pro DeepSeek13.8%as of 2026-05-01
    All Domains pass@1; PA 8.0%, UM 25.3%, CM 8.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 25, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  26. codex + gpt-5.6-terra OpenAI13.3%as of 2026-07-24
    All Domains pass@1; PA 12.0%, UM 20.0%, CM 8.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 26, Accuracy column; submission date 2026-07-24
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  27. codex + gpt-5.6-luna OpenAI13.3%as of 2026-07-24
    All Domains pass@1; PA 20.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 27, Accuracy column; submission date 2026-07-24
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  28. gemini-cli + gemini-3-flash Google12.5%as of 2026-05-01
    All Domains pass@1; PA 18.7%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 28, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  29. openclaw + deepseek-v4-pro DeepSeek11.1%as of 2026-05-01
    All Domains pass@1; PA 14.7%, UM 12.0%, CM 6.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 29, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  30. deepagents + glm-5.1 Zhipu AI11.1%as of 2026-05-01
    All Domains pass@1; PA 17.3%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 30, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  31. deepagents + deepseek-v4-pro DeepSeek10.7%as of 2026-05-01
    All Domains pass@1; PA 14.7%, UM 10.7%, CM 6.7%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 31, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  32. openclaw + kimi-k2.6 Moonshot AI10.2%as of 2026-05-01
    All Domains pass@1; PA 12.0%, UM 18.7%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 32, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  33. deepagents + qwen-3.6-max Alibaba9.3%as of 2026-05-01
    All Domains pass@1; PA 12.0%, UM 10.7%, CM 5.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 33, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  34. codex + gpt-5.4-mini OpenAI8.4%as of 2026-05-01
    All Domains pass@1; PA 10.7%, UM 13.3%, CM 1.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 34, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  35. openai-agents + TML Inkling 256K Thinking Machines8.0%as of 2026-07-24
    All Domains pass@1; PA 4.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 35, Accuracy column; submission date 2026-07-24
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  36. gemini-cli + gemini-3.1-pro Google7.1%as of 2026-05-01
    All Domains pass@1; PA 14.7%, UM 6.7%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 36, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  37. claude-code + claude-haiku-4-5 Anthropic6.2%as of 2026-05-01
    All Domains pass@1; PA 0.0%, UM 14.7%, CM 4.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 37, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  38. openai-agents + grok-4.3 SpaceX AI5.8%as of 2026-05-01
    All Domains pass@1; PA 0.0%, UM 16.0%, CM 1.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 38, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  39. openclaw + qwen-3.6-max Alibaba4.9%as of 2026-05-01
    All Domains pass@1; PA 10.7%, UM 4.0%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 39, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  40. hermes + grok-4.3 SpaceX AI4.4%as of 2026-05-01
    All Domains pass@1; PA 0.0%, UM 13.3%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 40, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  41. deepagents + kimi-k2.6 Moonshot AI3.1%as of 2026-05-01
    All Domains pass@1; PA 8.0%, UM 1.3%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 41, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  42. deepagents + grok-4.3 SpaceX AI2.2%as of 2026-05-01
    All Domains pass@1; PA 0.0%, UM 5.3%, CM 1.3%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 42, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  43. openclaw + grok-4.3 SpaceX AI0.4%as of 2026-05-01
    All Domains pass@1; PA 1.3%, UM 0.0%, CM 0.0%; submitted 2026-05-01; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 43, Accuracy column; submission date 2026-05-01
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified
  44. openai-agents + Nemotron 3 Ultra 256K NVIDIA0.0%as of 2026-07-24
    All Domains pass@1; PA 0.0%, UM 0.0%, CM 0.0%; submitted 2026-07-24; run date not published
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    All Domains leaderboard, rank 44, Accuracy column; submission date 2026-07-24
    reported by
    benchmark publisher
    published
    2026-08-12
    retrieved
    2026-09-28
    confidence
    verified

MedCode (Vals AI)

102 rows · percentage accuracy 0-100

Board source: Vals AI MedCode leaderboard

  1. Claude Opus 5 Anthropic63.57%as of 2026-09-26
    model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.156845/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5"].accuracy; Overall leaderboard rank 1
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  2. Gemini 3.1 Pro Preview (02/26) Google59.06%as of 2026-09-26
    model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.996 pp; $0.024714/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-pro-preview"].accuracy; Overall leaderboard rank 2
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  3. Claude Fable 5 Anthropic56.07%as of 2026-09-26
    model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.203 pp; $0.591071/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5"].accuracy; Overall leaderboard rank 3
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  4. Gemini 3 Flash (12/25) Google55.92%as of 2026-09-26
    model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.112 pp; $0.006187/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3-flash-preview"].accuracy; Overall leaderboard rank 4
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  5. Gemini 3.5 Flash Google55.83%as of 2026-09-26
    model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.073716/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash"].accuracy; Overall leaderboard rank 5
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  6. Claude Opus 4.7 Anthropic54.86%as of 2026-09-26
    model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.205 pp; $0.226314/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-7"].accuracy; Overall leaderboard rank 6
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  7. Claude Fable 5.1 Anthropic53.51%as of 2026-09-26
    model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.165 pp; $1.116862/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5-1"].accuracy; Overall leaderboard rank 7
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  8. Gemini 3.7 Flash Google53.39%as of 2026-09-26
    model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.12 pp; $0.038331/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.7-flash"].accuracy; Overall leaderboard rank 8
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  9. Claude Opus 4.8 Anthropic53.22%as of 2026-09-26
    model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.165 pp; $0.350925/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-8"].accuracy; Overall leaderboard rank 9
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  10. Gemini 3.6 Flash Google53.15%as of 2026-09-26
    model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.157 pp; $0.044216/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.6-flash"].accuracy; Overall leaderboard rank 10
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  11. Claude Sonnet 5.5 Anthropic52.92%as of 2026-09-26
    model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.119 pp; $0.391592/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5-5"].accuracy; Overall leaderboard rank 11
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  12. GPT 5.1 OpenAI52.73%as of 2026-09-26
    model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 2.151 pp; $0.014371/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.1-2025-11-13"].accuracy; Overall leaderboard rank 12
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  13. Gemini 3 Pro (11/25) Google52.20%as of 2026-09-26
    model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.073 pp; $0.028248/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3-pro-preview"].accuracy; Overall leaderboard rank 13
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  14. Muse Spark Meta51.31%as of 2026-09-26
    model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 2.244 pp; $0.005341/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["meta/muse_spark"].accuracy; Overall leaderboard rank 14
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  15. Gemini 2.5 Pro Google50.59%as of 2026-09-26
    model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 2.113 pp; $0.015389/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-pro"].accuracy; Overall leaderboard rank 15
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  16. Claude Opus 5.5 Anthropic49.80%as of 2026-09-26
    model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 2.273 pp; $0.658237/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5-5"].accuracy; Overall leaderboard rank 16
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  17. GPT 5.2 OpenAI49.75%as of 2026-09-26
    model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.262 pp; $0.018852/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.2-2025-12-11"].accuracy; Overall leaderboard rank 17
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  18. GPT 5 OpenAI49.63%as of 2026-09-26
    model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.098 pp; $0.045858/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5-2025-08-07"].accuracy; Overall leaderboard rank 18
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  19. Grok 4.7 SpaceXAI49.55%as of 2026-09-26
    model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 2.171 pp; $0.105497/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.7"].accuracy; Overall leaderboard rank 19
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  20. Muse Spark 1.2 Meta49.35%as of 2026-09-26
    model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.187 pp; $0.039023/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["meta/muse_spark_1_2"].accuracy; Overall leaderboard rank 20
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  21. Claude Opus 4.5 (Thinking) Anthropic49.16%as of 2026-09-26
    model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 2.012 pp; $0.095846/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].accuracy; Overall leaderboard rank 21
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  22. Claude Opus 4.6 (Thinking) Anthropic49.13%as of 2026-09-26
    model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.085 pp; $0.244127/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6-thinking"].accuracy; Overall leaderboard rank 22
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  23. GPT 5.5 OpenAI49.10%as of 2026-09-26
    model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 2.188 pp; $0.160759/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.5"].accuracy; Overall leaderboard rank 23
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  24. Kimi K3 Moonshot AI48.88%as of 2026-09-26
    model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 2.193 pp; $0.076379/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["kimi/kimi-k3"].accuracy; Overall leaderboard rank 24
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  25. GPT-6 Astra OpenAI48.49%as of 2026-09-26
    model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 2.131 pp; $0.451358/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-6-astra"].accuracy; Overall leaderboard rank 25
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  26. Claude Opus 4.6 (Nonthinking) Anthropic48.24%as of 2026-09-26
    model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.05 pp; $0.006180/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6"].accuracy; Overall leaderboard rank 26
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  27. Gemini 3.8 Flash Google48.13%as of 2026-09-26
    model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.18 pp; $0.017700/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.8-flash"].accuracy; Overall leaderboard rank 27
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  28. Gemini 3.1 Flash Lite Preview Google47.60%as of 2026-09-26
    model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.071 pp; $0.002029/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-flash-lite-preview"].accuracy; Overall leaderboard rank 28
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  29. Claude Sonnet 5 Anthropic47.54%as of 2026-09-26
    model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 2.274 pp; $0.278799/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5"].accuracy; Overall leaderboard rank 29
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  30. o3 OpenAI47.29%as of 2026-09-26
    model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.161 pp; $0.029818/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/o3-2025-04-16"].accuracy; Overall leaderboard rank 30
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  31. Claude Opus 4.1 (Thinking) Anthropic47.23%as of 2026-09-26
    model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 2.067 pp; $0.269254/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].accuracy; Overall leaderboard rank 31
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  32. GPT-6 Sol OpenAI47.07%as of 2026-09-26
    model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 2.119 pp; $0.085827/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-6-sol"].accuracy; Overall leaderboard rank 32
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  33. MiniMax-M3 MiniMax46.29%as of 2026-09-26
    model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.104 pp; $0.012125/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M3"].accuracy; Overall leaderboard rank 33
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  34. Claude Opus 4.5 (Nonthinking) Anthropic45.17%as of 2026-09-26
    model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.006826/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101"].accuracy; Overall leaderboard rank 34
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  35. MiMo V2.6 Pro Xiaomi44.97%as of 2026-09-26
    model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.097 pp; $0.009528/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-pro"].accuracy; Overall leaderboard rank 35
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  36. Grok 4.6 SpaceXAI44.71%as of 2026-09-26
    model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.256 pp; $0.050335/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.6"].accuracy; Overall leaderboard rank 36
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  37. GPT-6 Luna OpenAI44.69%as of 2026-09-26
    model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 2.303 pp; $0.007323/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-6-luna"].accuracy; Overall leaderboard rank 37
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  38. Claude Sonnet 4.5 (Thinking) Anthropic44.13%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.101495/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].accuracy; Overall leaderboard rank 38
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  39. GPT-5.6 Sol OpenAI43.97%as of 2026-09-26
    model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 2.258 pp; $0.280517/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-sol"].accuracy; Overall leaderboard rank 39
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  40. Gemini 3.5 Flash Lite Google43.49%as of 2026-09-26
    model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.951 pp; $0.008093/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash-lite"].accuracy; Overall leaderboard rank 40
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  41. GPT-5.6 Terra OpenAI43.41%as of 2026-09-26
    model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.173 pp; $0.046558/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-terra"].accuracy; Overall leaderboard rank 41
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  42. Grok 4.5 SpaceXAI43.29%as of 2026-09-26
    model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.313 pp; $0.048461/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.5"].accuracy; Overall leaderboard rank 42
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  43. Hy4 Preview Tencent43.25%as of 2026-09-26
    model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.134 pp; $0.062053/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["tencent/hy4-preview"].accuracy; Overall leaderboard rank 43
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  44. GPT 5 Mini OpenAI43.05%as of 2026-09-26
    model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 2.045 pp; $0.005560/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5-mini-2025-08-07"].accuracy; Overall leaderboard rank 44
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  45. GLM 5.3 zAI42.86%as of 2026-09-26
    model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.111 pp; $0.081905/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.3"].accuracy; Overall leaderboard rank 45
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  46. DeepSeek V4 Pro 0813 DeepSeek42.47%as of 2026-09-26
    model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.16 pp; $0.061060/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro-0813"].accuracy; Overall leaderboard rank 46
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  47. GPT-5.6 Luna OpenAI42.39%as of 2026-09-26
    model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.27 pp; $0.015970/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-luna"].accuracy; Overall leaderboard rank 47
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  48. GLM 5.1 zAI41.60%as of 2026-09-26
    model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.024196/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.1"].accuracy; Overall leaderboard rank 48
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  49. DeepSeek V4 Flash 0731 DeepSeek41.41%as of 2026-09-26
    model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 2.15 pp; $0.019659/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-flash-0731"].accuracy; Overall leaderboard rank 49
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  50. Claude Opus 4.1 (Nonthinking) Anthropic41.37%as of 2026-09-26
    model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.206270/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805"].accuracy; Overall leaderboard rank 50
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  51. GPT 5.4 (xhigh) OpenAI41.29%as of 2026-09-26
    model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 2.148 pp; $0.212108/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-2026-03-05"].accuracy; Overall leaderboard rank 51
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  52. Inkling Thinking Machines41.19%as of 2026-09-26
    model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.23 pp; $0.127025/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling"].accuracy; Overall leaderboard rank 52
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  53. DeepSeek V4.1 Flash DeepSeek41.17%as of 2026-09-26
    model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 2.042 pp; $0.012450/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4.1-flash"].accuracy; Overall leaderboard rank 53
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  54. MiMo V2.6 Flash Xiaomi41.06%as of 2026-09-26
    model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 2.032 pp; $0.002865/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-flash"].accuracy; Overall leaderboard rank 54
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  55. GPT 5.4 Nano OpenAI41.03%as of 2026-09-26
    model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 2.256 pp; $0.000844/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-nano-2026-03-17"].accuracy; Overall leaderboard rank 55
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  56. GLM 5.2 zAI40.77%as of 2026-09-26
    model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.166 pp; $0.045010/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.2"].accuracy; Overall leaderboard rank 56
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  57. Qwen 3.8 Max Alibaba40.67%as of 2026-09-26
    model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 2.029 pp; $0.120827/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-max"].accuracy; Overall leaderboard rank 57
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  58. Claude Sonnet 4.5 (Nonthinking) Anthropic40.57%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.995 pp; $0.042403/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929"].accuracy; Overall leaderboard rank 58
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  59. model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.003692/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025"].accuracy; Overall leaderboard rank 59
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  60. DeepSeek V4 DeepSeek40.45%as of 2026-09-26
    model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.122 pp; $0.060710/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro"].accuracy; Overall leaderboard rank 60
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  61. Gemini 2.5 Flash (7/17) (Thinking) Google40.36%as of 2026-09-26
    model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.952 pp; $0.003661/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-thinking"].accuracy; Overall leaderboard rank 61
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  62. Gemini 2.5 Flash Preview (9/25) (Thinking) Google40.33%as of 2026-09-26
    model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.003653/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 62
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  63. Kimi K2.6 Moonshot AI40.14%as of 2026-09-26
    model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 2.041 pp; $0.041295/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.6"].accuracy; Overall leaderboard rank 63
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  64. Kimi K2.5 Moonshot AI39.32%as of 2026-09-26
    model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 2.119 pp; $0.017275/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.5-thinking"].accuracy; Overall leaderboard rank 64
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  65. Qwen 3.7 Max Alibaba38.75%as of 2026-09-26
    model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 2.196 pp; $0.042362/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.7-max"].accuracy; Overall leaderboard rank 65
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  66. Nemotron 3 Ultra NVIDIA38.62%as of 2026-09-26
    model ID nvidia/nemotron-3-ultra-550b-a55b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.001 pp; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["nvidia/nemotron-3-ultra-550b-a55b"].accuracy; Overall leaderboard rank 66
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  67. Gemini 2.5 Flash (7/17) (Nonthinking) Google38.42%as of 2026-09-26
    model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.003698/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash"].accuracy; Overall leaderboard rank 67
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  68. Grok 4 SpaceXAI38.08%as of 2026-09-26
    model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.206 pp; $0.034103/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-0709"].accuracy; Overall leaderboard rank 68
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  69. Grok 4.3 SpaceXAI38.07%as of 2026-09-26
    model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.081 pp; $0.022202/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.3"].accuracy; Overall leaderboard rank 69
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  70. Inkling Small Thinking Machines37.89%as of 2026-09-26
    model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.206 pp; $0.016656/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling-small"].accuracy; Overall leaderboard rank 70
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  71. Grok 4 Fast (Reasoning) SpaceXAI37.38%as of 2026-09-26
    model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.941 pp; $0.002143/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-reasoning"].accuracy; Overall leaderboard rank 71
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  72. Qwen 3.6 Plus Alibaba36.89%as of 2026-09-26
    model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 2.017 pp; $0.015673/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.6-plus"].accuracy; Overall leaderboard rank 72
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  73. Llama 4 Maverick Meta36.51%as of 2026-09-26
    model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.994 pp; $0.002888/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["fireworks/llama4-maverick-instruct-basic"].accuracy; Overall leaderboard rank 73
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  74. Claude Sonnet 4 (Thinking) Anthropic34.96%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 1.939 pp; $0.069896/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].accuracy; Overall leaderboard rank 74
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  75. MiniMax-M2.7 MiniMax34.44%as of 2026-09-26
    model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.985 pp; $0.007424/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.7"].accuracy; Overall leaderboard rank 75
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  76. Gemini 2.5 Flash Lite (9/25) (Thinking) Google34.19%as of 2026-09-26
    model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.736 pp; $0.001182/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 76
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  77. MiniMax-M2.1 MiniMax34.08%as of 2026-09-26
    model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.943 pp; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.1"].accuracy; Overall leaderboard rank 77
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  78. Claude Sonnet 4 (Nonthinking) Anthropic33.94%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.906 pp; $0.039460/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514"].accuracy; Overall leaderboard rank 78
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  79. o4 Mini OpenAI33.79%as of 2026-09-26
    model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 2.021 pp; $0.017605/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/o4-mini-2025-04-16"].accuracy; Overall leaderboard rank 79
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  80. Mistral Medium 3.5 Mistral33.75%as of 2026-09-26
    model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.148 pp; $0.053370/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["mistralai/mistral-medium-3.5"].accuracy; Overall leaderboard rank 80
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  81. Qwen 3.5 Flash Alibaba33.00%as of 2026-09-26
    model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 1.787 pp; $0.003934/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.5-flash"].accuracy; Overall leaderboard rank 81
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  82. GLM 4.7 zAI32.77%as of 2026-09-26
    model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.996 pp; $0.006710/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-4.7"].accuracy; Overall leaderboard rank 82
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  83. Claude Haiku 4.5 (Thinking) Anthropic32.68%as of 2026-09-26
    model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.998 pp; $0.020099/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].accuracy; Overall leaderboard rank 83
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  84. MiMo V2.5 Pro Xiaomi32.48%as of 2026-09-26
    model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.006719/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5-pro"].accuracy; Overall leaderboard rank 84
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  85. Ling 3.0 Flash Ant Group32.27%as of 2026-09-26
    model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.908 pp; $0.001645/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-2607"].accuracy; Overall leaderboard rank 85
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  86. Grok 4.20 (Reasoning) SpaceXAI32.16%as of 2026-09-26
    model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.036189/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.20-0309-reasoning"].accuracy; Overall leaderboard rank 86
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  87. MiMo V2.5 Xiaomi31.89%as of 2026-09-26
    model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.025 pp; $0.002162/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5"].accuracy; Overall leaderboard rank 87
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  88. Qwen 3 VL Plus Alibaba31.65%as of 2026-09-26
    model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.845 pp; $0.002519/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].accuracy; Overall leaderboard rank 88
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  89. Qwen 3 Max Thinking Alibaba31.37%as of 2026-09-26
    model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.888 pp; $0.014776/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-max-2026-01-23"].accuracy; Overall leaderboard rank 89
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  90. Mercury 2.5 Inception31.33%as of 2026-09-26
    model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.953 pp; $0.004535/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["inception/mercury-2.5"].accuracy; Overall leaderboard rank 90
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  91. GPT 5 Nano OpenAI30.44%as of 2026-09-26
    model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.948 pp; $0.001729/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5-nano-2025-08-07"].accuracy; Overall leaderboard rank 91
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  92. Grok 4 Fast (Non-Reasoning) SpaceXAI30.04%as of 2026-09-26
    model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.974 pp; $0.002149/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-non-reasoning"].accuracy; Overall leaderboard rank 92
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  93. Ling 3.0 Flash Fin Ant Group29.30%as of 2026-09-26
    model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 1.943 pp; $0.001354/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-af-rc3"].accuracy; Overall leaderboard rank 93
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  94. Qwen 3.8 27B Alibaba28.70%as of 2026-09-26
    model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.971 pp; $0.050145/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-27b"].accuracy; Overall leaderboard rank 94
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  95. Grok 4.1 Fast Non-Reasoning SpaceXAI28.35%as of 2026-09-26
    model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.921 pp; $0.002193/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-non-reasoning"].accuracy; Overall leaderboard rank 95
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  96. Grok 4.1 Fast (Reasoning) SpaceXAI28.08%as of 2026-09-26
    model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.992 pp; $0.002108/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-reasoning"].accuracy; Overall leaderboard rank 96
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  97. Gemini 2.5 Flash Lite (Nonthinking) Google27.11%as of 2026-09-26
    model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.843 pp; $0.001342/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite"].accuracy; Overall leaderboard rank 97
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  98. Gemini 2.5 Flash Lite (9/25) (Nonthinking) Google27.08%as of 2026-09-26
    model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.911 pp; $0.001440/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].accuracy; Overall leaderboard rank 98
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  99. Llama 4 Scout Meta23.31%as of 2026-09-26
    model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.749 pp; $0.002176/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].accuracy; Overall leaderboard rank 99
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  100. Laguna M.1 Poolside23.11%as of 2026-09-26
    model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 1.693 pp; $0.002973/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["poolside/laguna-m.1"].accuracy; Overall leaderboard rank 100
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  101. Laguna XS.2 Poolside21.25%as of 2026-09-26
    model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 1.703 pp; $0.001445/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["poolside/laguna-xs.2"].accuracy; Overall leaderboard rank 101
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  102. Command A+ Cohere19.72%as of 2026-09-26
    model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 1.835 pp; $0.055695/test; source snapshot 2026-09-26; run date not published
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["cohere/command-a-plus-05-2026"].accuracy; Overall leaderboard rank 102
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified

MedScribe (Vals AI)

104 rows · percentage accuracy 0-100

Board source: Vals AI MedScribe leaderboard

  1. Claude Opus 5.5 Anthropic91.43%as of 2026-09-26
    model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.932 pp; $1.154156/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5-5"].accuracy; Overall leaderboard rank 1
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  2. Claude Fable 5.1 Anthropic91.29%as of 2026-09-26
    model ID anthropic/claude-fable-5-1; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.953 pp; $0.963500/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5-1"].accuracy; Overall leaderboard rank 2
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  3. Claude Sonnet 5.5 Anthropic91.10%as of 2026-09-26
    model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000; standard error 1.96 pp; $0.508604/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5-5"].accuracy; Overall leaderboard rank 3
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  4. Claude Opus 5 Anthropic90.98%as of 2026-09-26
    model ID anthropic/claude-opus-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.236275/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-5"].accuracy; Overall leaderboard rank 4
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  5. Muse Spark 1.2 Meta90.06%as of 2026-09-26
    model ID meta/muse_spark_1_2; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.958 pp; $0.037779/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["meta/muse_spark_1_2"].accuracy; Overall leaderboard rank 5
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  6. Grok 4.7 SpaceXAI89.38%as of 2026-09-26
    model ID grok/grok-4.7; reasoning_effort=xhigh; temperature=1; top_p=0.95; standard error 1.886 pp; $0.074505/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.7"].accuracy; Overall leaderboard rank 6
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  7. GLM 5.3 Flash zAI88.94%as of 2026-09-26
    model ID zai/glm-5.3-flash; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.907 pp; $0.002235/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.3-flash"].accuracy; Overall leaderboard rank 7
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  8. Muse Spark 1.1 Meta88.89%as of 2026-09-26
    model ID meta/muse_spark_1_1; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.95 pp; $0.034628/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["meta/muse_spark_1_1"].accuracy; Overall leaderboard rank 8
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  9. GLM 5.3 zAI88.81%as of 2026-09-26
    model ID zai/glm-5.3; reasoning_effort=max; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.999 pp; $0.057236/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.3"].accuracy; Overall leaderboard rank 9
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  10. Claude Fable 5 Anthropic88.52%as of 2026-09-26
    model ID anthropic/claude-fable-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.945 pp; $0.583239/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-fable-5"].accuracy; Overall leaderboard rank 10
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  11. MiMo V2.6 Pro Xiaomi88.31%as of 2026-09-26
    model ID xiaomi/mimo-v2.6-pro; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.938 pp; $0.009880/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-pro"].accuracy; Overall leaderboard rank 11
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  12. GPT 5.1 OpenAI88.09%as of 2026-09-26
    model ID openai/gpt-5.1-2025-11-13; reasoning_effort=high; max_output_tokens=30000; standard error 1.942 pp; $0.096508/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.1-2025-11-13"].accuracy; Overall leaderboard rank 12
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  13. Kimi K3 Moonshot AI87.96%as of 2026-09-26
    model ID kimi/kimi-k3; temperature=1; max_output_tokens=30000; standard error 1.891 pp; $0.118005/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["kimi/kimi-k3"].accuracy; Overall leaderboard rank 13
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  14. GPT-6 Astra OpenAI87.91%as of 2026-09-26
    model ID openai/gpt-6-astra; reasoning_effort=max; max_output_tokens=128000; standard error 1.938 pp; $0.581991/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-6-astra"].accuracy; Overall leaderboard rank 14
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  15. MiniMax-M3 MiniMax87.25%as of 2026-09-26
    model ID minimax/MiniMax-M3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.957 pp; $0.013748/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M3"].accuracy; Overall leaderboard rank 15
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  16. Grok 4.5 SpaceXAI86.88%as of 2026-09-26
    model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.944 pp; $0.033208/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.5"].accuracy; Overall leaderboard rank 16
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  17. GPT 5.5 OpenAI86.87%as of 2026-09-26
    model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000; standard error 1.932 pp; $0.142988/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.5"].accuracy; Overall leaderboard rank 17
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  18. Claude Opus 4.6 (Nonthinking) Anthropic86.74%as of 2026-09-26
    model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.942 pp; $0.115121/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6"].accuracy; Overall leaderboard rank 18
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  19. Grok 4.6 xAI86.53%as of 2026-09-26
    model ID grok/grok-4.6; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.956 pp; $0.037190/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.6"].accuracy; Overall leaderboard rank 19
    reported by
    official leaderboard
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  20. Claude Opus 4.6 (Thinking) Anthropic86.13%as of 2026-09-26
    model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.944 pp; $0.224735/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-6-thinking"].accuracy; Overall leaderboard rank 20
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  21. Muse Spark Meta85.90%as of 2026-09-26
    model ID meta/muse_spark; temperature=1; max_output_tokens=30000; standard error 1.847 pp; $0.007681/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["meta/muse_spark"].accuracy; Overall leaderboard rank 21
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  22. Claude Opus 4.8 Anthropic85.75%as of 2026-09-26
    model ID anthropic/claude-opus-4-8; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.928 pp; $0.259121/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-8"].accuracy; Overall leaderboard rank 22
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  23. DeepSeek V4.1 Flash DeepSeek85.50%as of 2026-09-26
    model ID deepseek/deepseek-v4.1-flash; reasoning_effort=high; max_output_tokens=384000; standard error 1.918 pp; $0.015397/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4.1-flash"].accuracy; Overall leaderboard rank 23
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  24. Inkling Thinking Machines85.41%as of 2026-09-26
    model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.844 pp; $0.165561/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling"].accuracy; Overall leaderboard rank 24
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  25. Claude Opus 4.5 (Thinking) Anthropic85.32%as of 2026-09-26
    model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.896 pp; $0.410224/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].accuracy; Overall leaderboard rank 25
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  26. MiMo V2.6 Flash Xiaomi85.28%as of 2026-09-26
    model ID xiaomi/mimo-v2.6-flash; temperature=1; top_p=0.95; max_output_tokens=128000; standard error 1.986 pp; $0.002184/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.6-flash"].accuracy; Overall leaderboard rank 26
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  27. GPT-5.6 Sol OpenAI85.23%as of 2026-09-26
    model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000; standard error 1.973 pp; $0.276691/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-sol"].accuracy; Overall leaderboard rank 27
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  28. Claude Haiku 4.5 (Thinking) Anthropic85.23%as of 2026-09-26
    model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.042375/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-haiku-4-5-20251001-thinking"].accuracy; Overall leaderboard rank 28
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  29. Qwen 3.8 Max Alibaba84.95%as of 2026-09-26
    model ID alibaba/qwen3.8-max; temperature=1; max_output_tokens=30000; standard error 1.999 pp; $0.089616/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-max"].accuracy; Overall leaderboard rank 29
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  30. Claude Sonnet 4.5 (Nonthinking) Anthropic84.52%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.054649/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929"].accuracy; Overall leaderboard rank 30
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  31. Gemini 3.8 Flash Google84.50%as of 2026-09-26
    model ID google/gemini-3.8-flash; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 1.943 pp; $0.025238/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.8-flash"].accuracy; Overall leaderboard rank 31
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  32. GPT-5.6 Luna OpenAI84.39%as of 2026-09-26
    model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000; standard error 2.585 pp; $0.022813/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-luna"].accuracy; Overall leaderboard rank 32
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  33. GPT 5.2 OpenAI84.39%as of 2026-09-26
    model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.856 pp; $0.115422/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.2-2025-12-11"].accuracy; Overall leaderboard rank 33
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  34. Inkling Small Thinking Machines84.11%as of 2026-09-26
    model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.87 pp; $0.019018/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["thinkingmachines/inkling-small"].accuracy; Overall leaderboard rank 34
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  35. Claude Sonnet 4.5 (Thinking) Anthropic84.10%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000; standard error 1.873 pp; $0.082281/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-5-20250929-thinking"].accuracy; Overall leaderboard rank 35
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  36. Gemini 3.7 Flash Google83.94%as of 2026-09-26
    model ID google/gemini-3.7-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.004 pp; $0.058736/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.7-flash"].accuracy; Overall leaderboard rank 36
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  37. Qwen 3.8 27B Alibaba83.85%as of 2026-09-26
    model ID alibaba/qwen3.8-27b; reasoning_effort=xhigh; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.977 pp; $0.046732/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.8-27b"].accuracy; Overall leaderboard rank 37
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  38. MiMo V2.5 Pro Xiaomi83.73%as of 2026-09-26
    model ID xiaomi/mimo-v2.5-pro; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.063 pp; $0.006030/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5-pro"].accuracy; Overall leaderboard rank 38
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  39. GPT-6 Luna OpenAI83.71%as of 2026-09-26
    model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000; standard error 1.949 pp; $0.009562/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-6-luna"].accuracy; Overall leaderboard rank 39
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  40. GPT 5 OpenAI83.65%as of 2026-09-26
    model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.936 pp; $0.101500/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5-2025-08-07"].accuracy; Overall leaderboard rank 40
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  41. Hy4 Preview Tencent83.60%as of 2026-09-26
    model ID tencent/hy4-preview; temperature=1; top_p=1; max_output_tokens=64000; standard error 2.065 pp; $0.053243/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["tencent/hy4-preview"].accuracy; Overall leaderboard rank 41
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  42. GLM 5.2 zAI83.53%as of 2026-09-26
    model ID zai/glm-5.2; temperature=1; max_output_tokens=30000; standard error 2.002 pp; $0.044912/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.2"].accuracy; Overall leaderboard rank 42
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  43. Claude Opus 4.5 (Nonthinking) Anthropic83.25%as of 2026-09-26
    model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000; standard error 1.926 pp; $0.281674/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-5-20251101"].accuracy; Overall leaderboard rank 43
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  44. Gemini 2.5 Flash (7/17) (Thinking) Google82.98%as of 2026-09-26
    model ID google/gemini-2.5-flash-thinking; temperature=1; max_output_tokens=30000; standard error 1.908 pp; $0.014824/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-thinking"].accuracy; Overall leaderboard rank 44
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  45. Claude Opus 4.7 Anthropic82.95%as of 2026-09-26
    model ID anthropic/claude-opus-4-7; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 1.977 pp; $0.177841/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-7"].accuracy; Overall leaderboard rank 45
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  46. Gemini 2.5 Flash (7/17) (Nonthinking) Google82.87%as of 2026-09-26
    model ID google/gemini-2.5-flash; temperature=1; max_output_tokens=30000; standard error 1.909 pp; $0.014869/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash"].accuracy; Overall leaderboard rank 46
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  47. GPT-5.6 Terra OpenAI82.87%as of 2026-09-26
    model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000; standard error 1.948 pp; $0.062588/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.6-terra"].accuracy; Overall leaderboard rank 47
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  48. GPT-6 Sol OpenAI82.03%as of 2026-09-26
    model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000; standard error 1.942 pp; $0.083138/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-6-sol"].accuracy; Overall leaderboard rank 48
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  49. Grok 4 Fast (Reasoning) SpaceXAI81.63%as of 2026-09-26
    model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.137 pp; $0.002535/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-reasoning"].accuracy; Overall leaderboard rank 49
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  50. Ling 3.0 Flash Ant Group80.90%as of 2026-09-26
    model ID ant/ling-3.0-flash-2607; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.039 pp; $0.001358/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-2607"].accuracy; Overall leaderboard rank 50
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  51. MiniMax-M2.1 MiniMax80.78%as of 2026-09-26
    model ID minimax/MiniMax-M2.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.831 pp; $0.005087/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.1"].accuracy; Overall leaderboard rank 51
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  52. GPT 5 Mini OpenAI80.58%as of 2026-09-26
    model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.924 pp; $0.033478/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5-mini-2025-08-07"].accuracy; Overall leaderboard rank 52
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  53. DeepSeek V4 Flash 0731 DeepSeek80.36%as of 2026-09-26
    model ID deepseek/deepseek-v4-flash-0731; reasoning_effort=high; max_output_tokens=30000; standard error 1.973 pp; $0.014247/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-flash-0731"].accuracy; Overall leaderboard rank 53
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  54. DeepSeek V4 Pro 0813 DeepSeek80.17%as of 2026-09-26
    model ID deepseek/deepseek-v4-pro-0813; reasoning_effort=max; max_output_tokens=30000; standard error 2.004 pp; $0.041127/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro-0813"].accuracy; Overall leaderboard rank 54
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  55. MiniMax-M2.7 MiniMax79.87%as of 2026-09-26
    model ID minimax/MiniMax-M2.7; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.86 pp; $0.005124/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["minimax/MiniMax-M2.7"].accuracy; Overall leaderboard rank 55
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  56. Grok 4 Fast (Non-Reasoning) SpaceXAI79.72%as of 2026-09-26
    model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.871 pp; $0.002056/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-fast-non-reasoning"].accuracy; Overall leaderboard rank 56
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  57. Gemini 3.6 Flash Google79.66%as of 2026-09-26
    model ID google/gemini-3.6-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.861 pp; $0.073178/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.6-flash"].accuracy; Overall leaderboard rank 57
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  58. Qwen 3.7 Max Alibaba79.40%as of 2026-09-26
    model ID alibaba/qwen3.7-max; temperature=1; max_output_tokens=30000; standard error 1.907 pp; $0.069070/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.7-max"].accuracy; Overall leaderboard rank 58
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  59. Grok 4.1 Fast (Reasoning) SpaceXAI78.73%as of 2026-09-26
    model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.866 pp; $0.002387/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-reasoning"].accuracy; Overall leaderboard rank 59
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  60. Gemini 2.5 Flash Preview (9/25) (Thinking) Google78.50%as of 2026-09-26
    model ID google/gemini-2.5-flash-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.993 pp; $0.014526/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 60
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  61. Grok 4 SpaceXAI78.15%as of 2026-09-26
    model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.084 pp; $0.063955/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-0709"].accuracy; Overall leaderboard rank 61
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  62. Kimi K2.6 Moonshot AI78.15%as of 2026-09-26
    model ID kimi/kimi-k2.6; top_p=0.95; max_output_tokens=30000; standard error 1.792 pp; $0.055962/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.6"].accuracy; Overall leaderboard rank 62
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  63. model ID google/gemini-2.5-flash-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.014385/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-preview-09-2025"].accuracy; Overall leaderboard rank 63
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  64. GPT 5.4 (xhigh) OpenAI77.55%as of 2026-09-26
    model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000; standard error 3.316 pp; $0.639282/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-2026-03-05"].accuracy; Overall leaderboard rank 64
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  65. Grok 4.1 Fast Non-Reasoning SpaceXAI77.46%as of 2026-09-26
    model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.04 pp; $0.001782/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4-1-fast-non-reasoning"].accuracy; Overall leaderboard rank 65
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  66. Qwen 3 VL Plus Alibaba77.13%as of 2026-09-26
    model ID alibaba/qwen3-vl-plus-2025-09-23; temperature=1; max_output_tokens=30000; standard error 1.916 pp; $0.020220/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-vl-plus-2025-09-23"].accuracy; Overall leaderboard rank 66
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  67. GPT 5.4 Nano OpenAI77.09%as of 2026-09-26
    model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.001800/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5.4-nano-2026-03-17"].accuracy; Overall leaderboard rank 67
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  68. Qwen 3.6 Plus Alibaba76.96%as of 2026-09-26
    model ID alibaba/qwen3.6-plus; temperature=1; max_output_tokens=30000; standard error 1.917 pp; $0.029294/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.6-plus"].accuracy; Overall leaderboard rank 68
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  69. o3 OpenAI76.65%as of 2026-09-26
    model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.871 pp; $0.040334/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/o3-2025-04-16"].accuracy; Overall leaderboard rank 69
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  70. Gemini 3.5 Flash Google76.57%as of 2026-09-26
    model ID google/gemini-3.5-flash; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.166341/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash"].accuracy; Overall leaderboard rank 70
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  71. Kimi K2.5 Moonshot AI76.44%as of 2026-09-26
    model ID kimi/kimi-k2.5-thinking; top_p=0.95; max_output_tokens=30000; standard error 1.986 pp; $0.024890/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["kimi/kimi-k2.5-thinking"].accuracy; Overall leaderboard rank 71
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  72. Gemini 3.1 Pro Preview (02/26) Google76.11%as of 2026-09-26
    model ID google/gemini-3.1-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.915 pp; $0.097954/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-pro-preview"].accuracy; Overall leaderboard rank 72
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  73. Claude Sonnet 5 Anthropic76.05%as of 2026-09-26
    model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000; standard error 3.05 pp; $0.433684/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-5"].accuracy; Overall leaderboard rank 73
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  74. Gemini 2.5 Flash Lite (9/25) (Nonthinking) Google75.82%as of 2026-09-26
    model ID google/gemini-2.5-flash-lite-preview-09-2025; temperature=1; max_output_tokens=30000; standard error 1.851 pp; $0.001332/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025"].accuracy; Overall leaderboard rank 74
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  75. Ling 3.0 Flash Fin Ant Group75.59%as of 2026-09-26
    model ID ant/ling-3.0-flash-af-rc3; temperature=1; top_p=0.95; max_output_tokens=131072; standard error 2.026 pp; $0.001618/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["ant/ling-3.0-flash-af-rc3"].accuracy; Overall leaderboard rank 75
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  76. DeepSeek V4 DeepSeek75.14%as of 2026-09-26
    model ID deepseek/deepseek-v4-pro; reasoning_effort=max; max_output_tokens=128000; standard error 2.002 pp; $0.053954/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["deepseek/deepseek-v4-pro"].accuracy; Overall leaderboard rank 76
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  77. Grok 4.3 SpaceXAI74.40%as of 2026-09-26
    model ID grok/grok-4.3; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.019 pp; $0.015293/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.3"].accuracy; Overall leaderboard rank 77
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  78. Claude Opus 4.1 (Thinking) Anthropic73.90%as of 2026-09-26
    model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000; standard error 1.965 pp; $0.263427/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].accuracy; Overall leaderboard rank 78
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  79. Gemini 2.5 Pro Google73.55%as of 2026-09-26
    model ID google/gemini-2.5-pro; temperature=1; max_output_tokens=30000; standard error 1.91 pp; $0.046379/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-pro"].accuracy; Overall leaderboard rank 79
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  80. GPT 5 Nano OpenAI72.86%as of 2026-09-26
    model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000; standard error 1.891 pp; $0.006961/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/gpt-5-nano-2025-08-07"].accuracy; Overall leaderboard rank 80
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  81. Gemini 2.5 Flash Lite (Nonthinking) Google72.83%as of 2026-09-26
    model ID google/gemini-2.5-flash-lite; temperature=1; max_output_tokens=30000; standard error 1.982 pp; $0.001211/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite"].accuracy; Overall leaderboard rank 81
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  82. Qwen 3 Max Thinking Alibaba72.71%as of 2026-09-26
    model ID alibaba/qwen3-max-2026-01-23; temperature=1; max_output_tokens=30000; standard error 1.905 pp; $0.085327/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3-max-2026-01-23"].accuracy; Overall leaderboard rank 82
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  83. Claude Sonnet 4 (Nonthinking) Anthropic72.41%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000; standard error 1.929 pp; $0.038973/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514"].accuracy; Overall leaderboard rank 83
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  84. GLM 5.1 zAI72.27%as of 2026-09-26
    model ID zai/glm-5.1; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.064 pp; $0.023717/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-5.1"].accuracy; Overall leaderboard rank 84
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  85. MiMo V2.5 Xiaomi72.15%as of 2026-09-26
    model ID xiaomi/mimo-v2.5; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.851 pp; $0.001351/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["xiaomi/mimo-v2.5"].accuracy; Overall leaderboard rank 85
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  86. Gemini 3 Pro (11/25) Google72.04%as of 2026-09-26
    model ID google/gemini-3-pro-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.9 pp; $0.061162/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3-pro-preview"].accuracy; Overall leaderboard rank 86
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  87. Claude Opus 4.1 (Nonthinking) Anthropic71.75%as of 2026-09-26
    model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000; standard error 2.021 pp; $0.187162/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-opus-4-1-20250805"].accuracy; Overall leaderboard rank 87
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  88. Gemini 3.5 Flash Lite Google70.89%as of 2026-09-26
    model ID google/gemini-3.5-flash-lite; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 2.031 pp; $0.019867/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.5-flash-lite"].accuracy; Overall leaderboard rank 88
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  89. Qwen 3.5 Flash Alibaba70.62%as of 2026-09-26
    model ID alibaba/qwen3.5-flash; temperature=1; max_output_tokens=30000; standard error 2.09 pp; $0.004425/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["alibaba/qwen3.5-flash"].accuracy; Overall leaderboard rank 89
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  90. Gemini 3 Flash (12/25) Google69.92%as of 2026-09-26
    model ID google/gemini-3-flash-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.899 pp; $0.014379/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3-flash-preview"].accuracy; Overall leaderboard rank 90
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  91. Claude Sonnet 4 (Thinking) Anthropic69.35%as of 2026-09-26
    model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000; standard error 2.212 pp; $0.053443/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].accuracy; Overall leaderboard rank 91
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  92. o4 Mini OpenAI69.14%as of 2026-09-26
    model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000; standard error 1.957 pp; $0.040605/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["openai/o4-mini-2025-04-16"].accuracy; Overall leaderboard rank 92
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  93. GLM 4.7 zAI68.63%as of 2026-09-26
    model ID zai/glm-4.7; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.123 pp; $0.019082/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["zai/glm-4.7"].accuracy; Overall leaderboard rank 93
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  94. Mistral Medium 3.5 Mistral67.73%as of 2026-09-26
    model ID mistralai/mistral-medium-3.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.011 pp; $0.157641/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["mistralai/mistral-medium-3.5"].accuracy; Overall leaderboard rank 94
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  95. Gemini 2.5 Flash Lite (9/25) (Thinking) Google66.88%as of 2026-09-26
    model ID google/gemini-2.5-flash-lite-preview-09-2025-thinking; temperature=1; max_output_tokens=30000; standard error 1.923 pp; $0.002567/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-2.5-flash-lite-preview-09-2025-thinking"].accuracy; Overall leaderboard rank 95
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  96. Laguna M.1 Poolside65.91%as of 2026-09-26
    model ID poolside/laguna-m.1; temperature=1; max_output_tokens=30000; standard error 2.007 pp; $0.002202/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["poolside/laguna-m.1"].accuracy; Overall leaderboard rank 96
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  97. Gemini 3.1 Flash Lite Preview Google63.90%as of 2026-09-26
    model ID google/gemini-3.1-flash-lite-preview; reasoning_effort=high; temperature=1; max_output_tokens=30000; standard error 1.823 pp; $0.002195/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["google/gemini-3.1-flash-lite-preview"].accuracy; Overall leaderboard rank 97
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  98. Grok 4.20 (Reasoning) SpaceXAI63.41%as of 2026-09-26
    model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.095 pp; $0.031303/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["grok/grok-4.20-0309-reasoning"].accuracy; Overall leaderboard rank 98
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  99. Laguna XS.2 Poolside61.43%as of 2026-09-26
    model ID poolside/laguna-xs.2; temperature=1; max_output_tokens=30000; standard error 2.349 pp; $0.001126/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["poolside/laguna-xs.2"].accuracy; Overall leaderboard rank 99
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  100. Command A+ Cohere55.68%as of 2026-09-26
    model ID cohere/command-a-plus-05-2026; temperature=1; top_p=0.95; max_output_tokens=64000; standard error 3.646 pp; $0.140316/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["cohere/command-a-plus-05-2026"].accuracy; Overall leaderboard rank 100
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  101. Mercury 2.5 Inception55.09%as of 2026-09-26
    model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536; standard error 2.095 pp; $0.004476/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["inception/mercury-2.5"].accuracy; Overall leaderboard rank 101
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  102. Llama 4 Maverick Meta54.22%as of 2026-09-26
    model ID fireworks/llama4-maverick-instruct-basic; temperature=1; max_output_tokens=30000; standard error 1.871 pp; $0.002460/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["fireworks/llama4-maverick-instruct-basic"].accuracy; Overall leaderboard rank 102
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  103. Llama 4 Scout Meta50.59%as of 2026-09-26
    model ID together/meta-llama/Llama-4-Scout-17B-16E-Instruct; temperature=1; max_output_tokens=30000; standard error 1.901 pp; $0.001700/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["together/meta-llama/Llama-4-Scout-17B-16E-Instruct"].accuracy; Overall leaderboard rank 103
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified
  104. Nemotron 3.5 Lightning NVIDIA4.27%as of 2026-09-26
    model ID fireworks/nemotron-lightning-3p5-30b-a3b; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 0.492 pp; $0.006241/test; source snapshot 2026-09-26; run date not published
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Embedded BenchmarkView data, tasks.overall["fireworks/nemotron-lightning-3p5-30b-a3b"].accuracy; Overall leaderboard rank 104
    reported by
    benchmark publisher
    published
    2026-09-26
    retrieved
    2026-09-28
    confidence
    verified

MedXpertQA (MM)

22 rows · percentage accuracy 0-100

Board source: Introducing Muse Spark: Scaling Towards Personal Superintelligence · official page: github.com

  1. GPT-5.6 Sol OpenAI81.5date not reported
    Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    reported by
    independent run
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    partial (partial review)
  2. Gemini 3.1 Pro Google81.3%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Muse Spark Eval Methodology (model card, 81.3, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 81.3%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 1; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  3. Qwen3.8 Max Alibaba80.4%as of 2026-08
    Alibaba's own Qwen3.8 launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.8-Max; header row: | | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max |
    reported by
    vendor-reported
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    partial (partial review)
    also reported in MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 80.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 2; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  4. Claude Fable 5 Anthropic80.0date not reported
    Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    reported by
    independent run
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    partial (partial review)
  5. Muse Spark Meta78.4%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    vendor-reported
    published
    2026-04-08
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Muse Spark Eval Methodology (model card, 78.4, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 78.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 3; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  6. GPT-5.4 OpenAI77.1%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Muse Spark Eval Methodology (model card, 77.1, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 77.1%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 4; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  7. Gemini 3 Pro Google76.0%date not reported
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Alibaba / Qwen
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    independent run
    published
    2026-02-16
    retrieved
    2026-09-28
    confidence
    verified
  8. GPT-5.2 OpenAI73.3date not reported
    Qwen-run comparison in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Alibaba / Qwen
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    independent run
    published
    2026-02-16
    retrieved
    2026-09-28
    confidence
    verified
  9. Claude Opus 4.8 Anthropic71.7date not reported
    Qwen-run comparison in the Qwen3.8-Max launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    reported by
    independent run
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    partial (partial review)
  10. Qwen3.7 Plus Alibaba71.0%as of 2026-05
    Alibaba's own Qwen3.7 Plus launch table; protocol not stated; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
    Qwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
    publisher
    Alibaba
    locator
    Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.7-Plus; header row: | | GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus |
    reported by
    vendor-reported
    published
    2026-05-31
    retrieved
    2026-09-07
    confidence
    partial (partial review)
    also reported in Qwen3.8-Max: A New Bar for Coding and Cowork (launch post, 71.0, Qwen3.8 launch post, Multimodal Benchmarks table, row MedXpertQA-MM, column Qwen3.7-Plus (same value carried forward)); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 71.0%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 5; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  11. Qwen3.5 397B A17B Alibaba70.0date not reported
    self-reported in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Alibaba / Qwen
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    vendor-reported
    published
    2026-02-16
    retrieved
    2026-09-28
    confidence
    verified
  12. Qwen3.6 Plus Alibaba68.7date not reported
    Qwen-run comparison in the Qwen3.7-Plus launch post; Primary blog currently renders an empty shell; retained historical value, not reverified on 2026-09-28.
    Qwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
    publisher
    Alibaba
    locator
    Multimodal Benchmarks table, row MedXpertQA-MM; columns GPT-5.4 (xhigh) / Opus-4.6 Max / Gemini-3.1 Pro / Qwen3.6-Plus / Qwen3.7-Plus
    reported by
    vendor-reported
    published
    2026-05-31
    retrieved
    2026-09-07
    confidence
    partial (partial review)
  13. Grok 4.20 xAI65.8%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Muse Spark Eval Methodology (model card, 65.8, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 65.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 6; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  14. Kimi K2.5 Moonshot AI65.3date not reported
    Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Alibaba / Qwen
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    independent run
    published
    2026-02-16
    retrieved
    2026-09-28
    confidence
    verified
  15. Claude Opus 4.6 Anthropic64.8%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Muse Spark Eval Methodology (model card, 64.8, p. 5 benchmark image, Health section, MedXpertQA (MM); visually checked 2026-09-28); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 64.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 7; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  16. Claude Opus 4.5 Anthropic63.6%date not reported
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Alibaba / Qwen
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    independent run
    published
    2026-02-16
    retrieved
    2026-09-28
    confidence
    verified
  17. Gemma 4 31B Google61.3%date not reported
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card · model card · first-party
    publisher
    Google
    locator
    Evaluation Results table, MedXPertQA MM row, Gemma 4 31B column
    reported by
    vendor-reported
    published
    2026-04-02
    retrieved
    2026-09-28
    confidence
    verified
  18. Gemma 4 26B A4B Google58.1%date not reported
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card · model card · first-party
    publisher
    Google
    locator
    Evaluation Results table, MedXPertQA MM row, Gemma 4 26B A4B column
    reported by
    vendor-reported
    published
    2026-04-02
    retrieved
    2026-09-28
    confidence
    verified
  19. Gemma 4 12B Google48.7%as of 2026-04
    Google's Gemma 4 model card, Unified 12B; protocol not stated
    Gemma 4 model card · model card · first-party
    publisher
    Google
    locator
    Benchmark Results table, Vision section, row MedXPertQA MM, column Gemma 4 12B Unified; header row: | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |
    reported by
    vendor-reported
    published
    2026-04-02
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Gemma 4 Technical Report (paper, 48.7, Gemma 4 Technical Report p. 6, Table 6 (vision benchmarks, thinking), row MedXPertQA MM, column 12B); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 48.7%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 8; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  20. Qwen3-VL-235B-A22B Alibaba47.6%date not reported
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Alibaba / Qwen
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    vendor-reported
    published
    2026-02-16
    retrieved
    2026-09-28
    confidence
    verified
  21. Gemma 4 E4B Google28.7%date not reported
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card · model card · first-party
    publisher
    Google
    locator
    Evaluation Results table, MedXPertQA MM row, Gemma 4 E4B column
    reported by
    vendor-reported
    published
    2026-04-02
    retrieved
    2026-09-28
    confidence
    verified
  22. Gemma 4 E2B Google23.5%date not reported
    Google model card; MedXPertQA MM row; vendor-reported; protocol differs from other source families
    Gemma 4 model card · model card · first-party
    publisher
    Google
    locator
    Evaluation Results table, MedXPertQA MM row, Gemma 4 E2B column
    reported by
    vendor-reported
    published
    2026-04-02
    retrieved
    2026-09-28
    confidence
    verified

Board source: Best AI for Healthcare & Medical: LLM Leaderboard

  1. Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=claude-opus-5-5, weightedIndex=60.5359874288733; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  2. Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=claude-fable-5-1, weightedIndex=58.0046680298741; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  3. Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=claude-opus-5, weightedIndex=53.3389883833501; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  4. GPT-6 Astra (max) OpenAI52as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-6-astra, weightedIndex=51.6668965494132; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  5. Muse Spark 1.3 (max) Meta50as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=muse-spark-1-3, weightedIndex=49.6912519301107; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  6. Grok 4.7 (xhigh) SpaceXAI47as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=grok-4-7, weightedIndex=47.1903068758764; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  7. GLM-5.3 (max) Z AI47as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=glm-5-3, weightedIndex=46.6708125611712; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  8. GPT-5.6 Sol (max) OpenAI45as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-5-6-sol, weightedIndex=45.4043575131731; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  9. Kimi K3 (max) Kimi45as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=kimi-k3, weightedIndex=45.0434943845881; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  10. GLM 5.3 Flash Z AI45as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=glm-5-3-flash, weightedIndex=44.8426192521491; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  11. GPT-6 Sol (max) OpenAI43as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-6-sol, weightedIndex=43.4880669048048; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  12. MiMo-V2.6-Pro Xiaomi42as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=mimo-v2-6-pro, weightedIndex=41.8204296860305; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  13. Gemini 3.8 Flash (high) Google42as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gemini-3-8-flash, weightedIndex=41.7830626872104; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  14. Qwen3.8 Max (0902) Alibaba41as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=qwen3-8-max, weightedIndex=41.4279151085327; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  15. Step 5 Preview StepFun41as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=step-5, weightedIndex=40.8171769208551; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  16. Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=deepseek-v4-1-flash, weightedIndex=40.6388737813166; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  17. GPT-6 Luna (max) OpenAI37as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-6-luna, weightedIndex=36.7252603853009; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  18. GPT-5.6 Luna (max) OpenAI36as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gpt-5-6-luna, weightedIndex=35.740444759121; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  19. Qwen3.8 27B (xhigh) Alibaba34as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=qwen3-8-27b, weightedIndex=34.4310443115205; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  20. MiniMax-M3 MiniMax30as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=minimax-m3, weightedIndex=29.7042505129902; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  21. Inkling (xhigh) Thinking Machines25as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=inkling, weightedIndex=25.3933378261902; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  22. Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=nvidia-nemotron-3-ultra-550b-a55b, weightedIndex=23.0046725383136; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  23. Gemini 3.5 Flash-Lite Google23as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=gemini-3-5-flash-lite, weightedIndex=23.074121756108; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  24. Muse Glimmer (high) Meta18as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=muse-glimmer, weightedIndex=17.5980321302316; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified
  25. Mistral Medium 3.5 Mistral14as of 2026-09
    Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
    Artificial Analysis Healthcare & Medical Index · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    Embedded chart data: capability=healthcareAndMedical, initialModels slug=mistral-medium-3-5, weightedIndex=13.5556785369912; round to nearest whole point
    reported by
    third-party run
    retrieved
    2026-09-28
    confidence
    verified

PhysicianBench

21 rows · pass@1 success rate %

Board source: PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · official page: arxiv.org

  1. Claude Opus 5.5 (max) Anthropic68.4%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  2. Claude Sonnet 5.5 (max) Anthropic63.2%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  3. Claude Fable 5.1 (max) Anthropic61.0%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  4. Claude Opus 5 (max) Anthropic57.6%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  5. Claude Sonnet 5.5 (xhigh) Anthropic56.4%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; xhigh effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  6. Claude Sonnet 5.5 (high) Anthropic47.6%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; high effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  7. GPT-5.5 OpenAI46.3 ± 1.2as of 2026-05
    pass@1; Pass^3 28.0
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 46.3 ± 1.2, arXiv abs page (landing page for the PDF))
  8. Claude Sonnet 5 (max) Anthropic37.4%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; max effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  9. Claude Opus 4.6 Anthropic31.7 ± 2.3as of 2026-05
    Pass^3 18.0
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 31.7 ± 2.3, arXiv abs page (landing page for the PDF))
  10. Claude Sonnet 5.5 (medium) Anthropic30.0%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; medium effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  11. Claude Opus 4.7 Anthropic29.3 ± 2.5as of 2026-05
    Pass^3 18.0
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 29.3 ± 2.5, arXiv abs page (landing page for the PDF))
  12. GPT-5.4 OpenAI27.7 ± 1.5as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 27.7 ± 1.5, arXiv abs page (landing page for the PDF))
  13. Claude Sonnet 5.5 (low) Anthropic27.2%as of 2026-09-28
    Anthropic-run pass@1 on 100 tasks; low effort; shared Anthropic harness; Opus 5 rubric grader; safety classifiers enabled; differs from benchmark paper protocol; run date unpublished
    Claude Sonnet 5.5 System Card · system card · first-party
    publisher
    Anthropic
    locator
    pp. 137–139, Section 8.15.3, Figure 8.15.B, PhysicianBench pass@1
    reported by
    vendor-reported
    published
    2026-09-28
    retrieved
    2026-09-28
    confidence
    verified
  14. Claude Sonnet 4.6 Anthropic23.0 ± 2.6as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 23.0 ± 2.6, arXiv abs page (landing page for the PDF))
  15. DeepSeek V4-Pro DeepSeek18.7 ± 2.9as of 2026-05
    Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
  16. Kimi-K2.6 Moonshot AI17.0 ± 2.6as of 2026-05
    open source
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 17.0 ± 2.6, arXiv abs page (landing page for the PDF))
  17. MiMo-v2.5-Pro Xiaomi16.7 ± 4.0as of 2026-05
    Pass@1 over 3 runs; Pass^3 6.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
  18. Qwen3.6-Plus Alibaba13.7 ± 4.0as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 13.7 ± 4.0, arXiv abs page (landing page for the PDF))
  19. MiniMax M2.7 MiniMax8.7 ± 1.2as of 2026-05
    Pass@1 over 3 runs; Pass^3 1.0; shared FHIR tool harness; up to 100 turns; high reasoning when supported
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
  20. Gemini Pro 3.1 Google6.0 ± 1.0as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 6.0 ± 1.0, arXiv abs page (landing page for the PDF))
  21. Grok-4.20 xAI5.3 ± 3.2as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (Proprietary Models block)
    reported by
    benchmark publisher
    published
    2026-05-04
    retrieved
    2026-09-28
    confidence
    verified

EHR-Complex

18 rows · exact-match accuracy

Board source: EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · official page: arxiv.org

  1. GPT-5.4 (high reasoning) OpenAI0.65as of 2026-06
    average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.65, arXiv abs page (landing page for the PDF))
  2. Gemini 3.1 Pro Google0.63as of 2026-06
    validation configuration
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.63, arXiv abs page (landing page for the PDF))
  3. Kimi-K2.5 Moonshot AI0.62as of 2026-06
    headline 12-model evaluation, top open-weight
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF))
  4. Qwen3.5-397B Alibaba0.62as of 2026-06
    headline evaluation
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF))
  5. DeepSeek-V3.2-Exp DeepSeek0.59as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  6. GPT-5.4 (low reasoning) OpenAI0.58as of 2026-06
    validation configuration
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.58, arXiv abs page (landing page for the PDF))
  7. DeepSeek-V3.1 DeepSeek0.56as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  8. Qwen3-32B-SFT Alibaba0.55as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  9. Qwen3-235B Alibaba0.53as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  10. GPT-4.1 mini OpenAI0.49as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  11. GPT-4.1 OpenAI0.47as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  12. Qwen3-14B-SFT Alibaba0.45as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns; fine-tuned on trajectories from EHR-Complex training set
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  13. Claude Sonnet 4.6 Anthropic0.36as of 2026-06
    validation configuration
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.36, arXiv abs page (landing page for the PDF))
  14. Qwen3-32B Alibaba0.36as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  15. GPT-4o OpenAI0.31as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  16. Gemini 2.5 Pro Google0.31as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  17. Qwen3-14B Alibaba0.30as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified
  18. Qwen3-4B Alibaba0.16as of 2026-06
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3, Avg. column
    reported by
    benchmark publisher
    published
    2026-06-22
    retrieved
    2026-09-28
    confidence
    verified

WHBench

22 rows · mean normalized percentage 0-100

Board source: WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · official page: arxiv.org

  1. Claude Opus 4.6 Anthropic72.1%as of 2026-03
    95% CI 69.6-74.4; evaluations run March 2026
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 72.1%, same paper, abstract page)
  2. Claude Sonnet 4.6 Anthropic67.1%as of 2026-03
    95% CI 64.5-69.6
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 67.1%, same paper, abstract page)
  3. GPT-5.4 OpenAI66.8%as of 2026-03
    95% CI 64.5-69.2
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 66.8%, same paper, abstract page)
  4. Gemini 3 Flash Preview Google64.7%as of 2026-03
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 64.7%, same paper, abstract page)
  5. OpenAI o3 OpenAI63.6%as of 2026-03
    95% bootstrap CI 61.3–65.9; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  6. DeepSeek V3.2 DeepSeek61.3%as of 2026-03
    95% bootstrap CI 58.6–63.9; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  7. Grok 3 SpaceX AI60.7%as of 2026-03
    95% bootstrap CI 58.0–63.4; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  8. Mistral Large Mistral AI60.2%as of 2026-03
    95% bootstrap CI 57.4–63.0; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  9. Grok 4 SpaceX AI57.9%as of 2026-03
    95% bootstrap CI 54.9–60.8; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  10. DeepSeek-R1 DeepSeek52.9%as of 2026-03
    95% bootstrap CI 50.5–55.3; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  11. GPT-4.1 OpenAI51.8%as of 2026-03
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 51.8%, same paper, abstract page)
  12. Grok 3 Mini SpaceX AI50.0%as of 2026-03
    95% bootstrap CI 47.5–52.5; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  13. Gemini 2.5 Flash Google49.5%as of 2026-03
    95% bootstrap CI 47.0–52.0; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  14. Claude Opus 4 Anthropic49.1%as of 2026-03
    95% bootstrap CI 46.4–51.7; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  15. Claude Sonnet 4 Anthropic48.1%as of 2026-03
    95% bootstrap CI 45.5–50.6; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  16. GPT-4o OpenAI44.6%as of 2026-03
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 44.6%, same paper, abstract page)
  17. Llama 4 Maverick Meta42.1%as of 2026-03
    95% bootstrap CI 39.6–44.6; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  18. Nemotron 70B NVIDIA39.3%as of 2026-03
    95% bootstrap CI 37.3–41.3; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  19. Llama 3.3 70B Meta37.8%as of 2026-03
    95% bootstrap CI 35.2–40.5; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  20. Llama 3.1 405B Meta36.1%as of 2026-03
    95% bootstrap CI 33.9–38.3; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  21. Gemini 2.5 Pro Google35.3%as of 2026-03
    95% bootstrap CI 32.7–38.1; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified
  22. Llama 4 Scout Meta35.2%as of 2026-03
    95% bootstrap CI 33.2–37.3; 3 runs; temperature 0; zero-shot, closed-book
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    benchmark publisher
    published
    2026-07-23
    retrieved
    2026-09-28
    confidence
    verified

HealthAdminBench

7 rows · percentage end-to-end task success 0-100

Board source: HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · official page: healthadminbench.stanford.edu · paper: arxiv.org

  1. Claude Opus 4.6 (computer-use agent) Anthropic36.3%as of 2026-04
    screenshot-only, task description + portal guidance; native CUA harness; subtask rate 78.4%
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 36.3%, Section 'LLMs struggle with long-horizon tasks')
  2. GPT-5.4 (computer-use agent) OpenAI26.7%as of 2026-04
    screenshot-only, task description + portal guidance; subtask rate 82.8%
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 26.7%, Section 'LLMs struggle with long-horizon tasks')
  3. Kimi K2.5 Moonshot AI15.6%as of 2026-04
    screenshot-only, task description + portal guidance
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 15.6%, Section 'LLMs struggle with long-horizon tasks')
  4. Claude Opus 4.6 (standardized harness) Anthropic14.8%as of 2026-04
    screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified
  5. Qwen 3.5 Alibaba13.3%as of 2026-04
    screenshot-only, task description + portal guidance
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 13.3%, Section 'LLMs struggle with long-horizon tasks')
  6. Gemini 3.1 Pro Google11.9%as of 2026-04
    screenshot-only, task description + portal guidance
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 11.9%, Section 'LLMs struggle with long-horizon tasks')
  7. GPT-5.4 (standardized harness) OpenAI5.9%as of 2026-04
    screenshot-only, task description + portal guidance; authors' standardized harness, no native CUA
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    benchmark publisher
    published
    2026-04-10
    retrieved
    2026-09-28
    confidence
    verified

Documents on file

The distinct documents the sections above draw on, with the number of index rows each one backs.

documentkindrows
Vals AI MedScribe leaderboard
vals.ai
official leaderboard · first-party104
Vals AI MedCode leaderboard
vals.ai
official leaderboard · first-party102
CHI-Bench leaderboard (actAVA)
actava.ai
official leaderboard · first-party44
GPT-5.6 System Card
deploymentsafety.openai.com
system card · first-party25
Artificial Analysis Healthcare & Medical Index
artificialanalysis.ai
official leaderboard · first-party25
WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2)
arxiv.org
paper · first-party22
GPT-5.6 Preview System Card
deploymentsafety.openai.com
system card · first-party21
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF)
arxiv.org
paper · first-party18
Health Optimization Bench — subject suites results
healthoptimizationbench.com
Arcophos run · first-party16
MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results)
arise-ai.org
official leaderboard · first-party12
HealthAgentBench leaderboard
microsoft.github.io
official leaderboard · first-party12
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF)
arxiv.org
paper · first-party12
Claude Sonnet 5.5 System Card
anthropic.com
system card · first-party12
HealthAgentBench detailed results
microsoft.github.io
official leaderboard · first-party11
MedHELM leaderboard (medhelm.org), v5.0.0
medhelm.org
official leaderboard · first-party10
GPT-5.6 - August Updates (system card addendum)
cdn.openai.com
system card · first-party10
MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate)
leaderboard.medhelm.org
official leaderboard · first-party10
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1)
arxiv.org
paper · first-party9
GPT-6 Astra System Card — September 22 revision
deploymentsafety.openai.com
system card · first-party9
MAST: Medical AI Superintelligence Test leaderboard (General board)
arise-ai.org
official leaderboard · first-party8
MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai
benchlm.ai
mirror8
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240)
arxiv.org
paper · first-party8
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1)
arxiv.org
paper · first-party7
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301)
arxiv.org
paper · first-party6
WHBench (arXiv 2604.00024 abstract page)
arxiv.org
paper · first-party6
Introducing Muse Spark: Scaling Towards Personal Superintelligence
ai.meta.com
launch post · first-party6
Muse Spark Eval Methodology
ai.meta.com
model card · first-party6
Qwen/Qwen3.5-397B-A17B model card
huggingface.co
model card · first-party6
Baichuan-M3 Technical Report
arxiv.org
paper · first-party6
Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog)
kineticsystems.ai
vendor post5
GPT-5.5 Instant System Card
deploymentsafety.openai.com
system card · first-party5
System Card: Claude Opus 5
anthropic.com
system card · first-party5
Qwen3.8-Max: A New Bar for Coding and Cowork
qwen.ai
launch post · first-party5
Gemma 4 model card
ai.google.dev
model card · first-party5
Claude Fable 5.1 and Claude Mythos 5.1 System Card
anthropic.com
system card · first-party5
ARISE MAST technical leaderboard
arise-ai.org
official leaderboard · first-party5
gpt-oss-120b & gpt-oss-20b Model Card
deploymentsafety.openai.com
model card · first-party4
System Card: Claude Sonnet 5
anthropic.com
system card · first-party4
GPT-5.3 Instant System Card
deploymentsafety.openai.com
system card · first-party2
GPT-5.5 System Card
deploymentsafety.openai.com
system card · first-party2
System Card: Claude Opus 4.8
anthropic.com
system card · first-party2
Muse Spark 1.1 Evaluation Report
research.meta.ai
model card · first-party2
Qwen3.7-Plus: Multimodal Agent Intelligence
qwen.ai
launch post · first-party2
Claude Opus 5.5 System Card
anthropic.com
system card · first-party2
Introducing Grok 4.7
x.ai
launch post · first-party2
Claude Fable 5 and Claude Mythos 5 System Card
anthropic.com
system card · first-party1
MAI-Thinking-1: Building a Hill-Climbing Machine
microsoft.ai
model card · first-party1
GPT-5 System Card
cdn.openai.com
system card · first-party1
Gemma 4 Technical Report
arxiv.org
paper1

Corrections go through the same route as everything else here: a better document replaces a weaker one, the row's confidence moves, and the change is dated on the updates page. The full record, sources included, is downloadable from the data page.