Health Evals

OpenAI logoGPT OSS 20B: healthcare benchmark results

OpenAI · 2 boards · snapshot reviewed September 28, 2026

The index currently holds 2 results for GPT OSS 20B: 0.108 on HealthBench Hard (20 of 20 indexed rows), 42.5 on HealthBench (26 of 26 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscoreindex positionas of
HealthBench Hard
raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
0.10820 of 202025-08
HealthBench
reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8)
42.526 of 262025-08

Position follows the rows held in this index and is not a controlled comparison. Sources may use different graders or task subsets. Positions include separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page.

Which healthcare benchmarks is GPT OSS 20B scored on?

As of September 28, 2026, GPT OSS 20B has indexed results on 2 tracked benchmarks: HealthBench Hard, HealthBench.

Which results are indexed for GPT OSS 20B?

GPT OSS 20B stands at 0.108 on HealthBench Hard (20 of 20 indexed rows), 42.5 on HealthBench (26 of 26 indexed rows).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.