Health AI Benchmarks

GPT-6.1 Sol

3 results on 3 boards

Made by
OpenAI
Released
29 September 2026
Licence
proprietary
Context
1.05M
Input, per 1M tokens
$2
Output, per 1M tokens
$10

Health Evals Index

44.6rank 9 of 593 of 5 boards

This site's own calculation from 3 of 5 benchmarks, its weighted average times a multiplier: 51.5 × 0.866 = 44.6. How it works

  • HealthBench Professional0.64264.2weight 0.743system cardReported by the maker.OpenAI; length-adjusted; 64.2 (67.2, 3038) (adjusted, raw, mean response characters); addendum Table 7 refers to the GPT-6 Astra card for method
  • HealthBench58.558.5weight 0.823system cardReported by the maker.OpenAI; length-adjusted; 58.5 (56.7, 1701) (adjusted, raw, mean response characters); addendum Table 7 refers to the GPT-6 Astra card for method
  • Health Optimization Benchno result on this benchmark
  • HealthBench Hard0.36236.2system cardReported by the maker.OpenAI; length-adjusted; 36.2 (33.4, 1646) (adjusted, raw, mean response characters); addendum Table 7 refers to the GPT-6 Astra card for method
  • WHBenchno result on this benchmark

51.5 × 0.866 = 44.6: the weighted average of these 3 scores out of 100, times 0.866 because this score rests on 3 of 5 benchmarks. From 4 benchmarks up a score keeps its full average. Each benchmark counts by its weight, which is lower the nearer its leader is to 100.

Every result

Each against its own board only. The dot strip shows the whole board; this model is the larger red dot.

  • HealthBench Professional

    OpenAI; length-adjusted; 64.2 (67.2, 3038) (adjusted, raw, mean response characters); addendum Table 7 refers to the GPT-6 Astra card for method

    0.642

    6 of 32, leader 0.703

    system card Sep 2026

  • HealthBench

    OpenAI; length-adjusted; 58.5 (56.7, 1701) (adjusted, raw, mean response characters); addendum Table 7 refers to the GPT-6 Astra card for method

    58.5

    11 of 29, leader 67.1

    system card Sep 2026

  • HealthBench Hard

    OpenAI; length-adjusted; 36.2 (33.4, 1646) (adjusted, raw, mean response characters); addendum Table 7 refers to the GPT-6 Astra card for method

    0.362

    4 of 20, leader 0.428

    system card Sep 2026

Compare GPT-6.1 Sol with other models