Health Evals

Health Optimization Bench: published results

Arcophos · 257 tasks across eight subject suites · index updated September 28, 2026

Claude Fable 5 has the highest indexed numerical score on Health Optimization Bench, 70.9 as of 2026-09, per Health Optimization Bench — subject suites results. Questions in eight areas of preventive and optimization medicine, grounded in primary evidence and scored against task-specific rubrics. The current main ranking covers 257 released tasks, separate from the 89-task incretin therapeutics evidence suite.

This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthoptimizationbench.com. Every published row in the index appears below.

Published results

#modelscoreas of
1Anthropic logoClaude Fable 5 Anthropic
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 68.0–73.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
70.92026-09
2Anthropic logoClaude Opus 5 Anthropic
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 66.2–72.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
69.32026-09
3xAI logoGrok 4.6 xAI
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.8–69.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
66.82026-09
4OpenAI logoGPT-5.6 Sol (max) OpenAI
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort.
66.62026-09
5OpenAI logoGPT-5.6 Sol (high) OpenAI
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort.
64.62026-09
6Moonshot AI logoKimi K3 Moonshot AI
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
59.92026-09
7Meta logoMuse Spark Meta
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
57.22026-09
8Anthropic logoClaude Fable 5.1 Anthropic
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 42.2–52.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Vendor safeguards declined 97/257 tasks, scored with no credit; mean over answered tasks is 75.9.
47.32026-09
9Google logoGemini 3.6 Google
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 36.2–43.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
39.72026-09
10TMInkling Thinking Machines
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 32.3–39.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
35.62026-09
11Anthropic logoClaude Sonnet 5 Anthropic
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
34.62026-09
12Zhipu logoGLM 5.2 Zhipu
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 18.1–23.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
20.72026-09
13MiniMax logoMiniMax M3 MiniMax
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.7–20.9. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
18.32026-09
14Microsoft AI logoMAI Thinking Microsoft AI
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
17.52026-09
15Mistral logoMistral Medium 3.5 Mistral
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 7.4–11.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
9.22026-09
16NVIDIA logoNemotron 3.5 Lightning NVIDIA
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 3.6–6.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
4.92026-09

Scores preserve their source precision, with any scale conversion documented (independently run). The main ranking moved from the 89-task incretin suite to 257 subject-suite tasks. This page shows only the 257-task September 10 snapshot; the two task sets must not be pooled. Three grader families vote, with split decisions escalated to a fourth, and the authoring family excluded. Fable 5.1 safeguards declined 97 tasks; those receive no credit. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.

About the benchmark

publisherArcophos
categoryrubric-graded benchmarks
released2026-08
size257 tasks across eight subject suites
scale0-100 rubric credit, higher better
result basisindependently run
sourcehealthoptimizationbench.com
official pagehealthoptimizationbench.com
last frontier result2026-09

What is Health Optimization Bench?

Health Optimization Bench is a rubric-graded benchmark from Arcophos, released 2026-08: 257 tasks across eight subject suites, scored on a 0-100 rubric credit scale. Questions in eight areas of preventive and optimization medicine, grounded in primary evidence and scored against task-specific rubrics. The current main ranking covers 257 released tasks, separate from the 89-task incretin therapeutics evidence suite.

Which model leads Health Optimization Bench?

Claude Fable 5 (Anthropic) has the highest indexed numerical score on Health Optimization Bench at 70.9 (evaluation setups may differ), per Health Optimization Bench — subject suites results, as of 2026-09.

Where do the Health Optimization Bench numbers come from?

From healthoptimizationbench.com (independently run). The main ranking moved from the 89-task incretin suite to 257 subject-suite tasks. This page shows only the 257-task September 10 snapshot; the two task sets must not be pooled. Three grader families vote, with split decisions escalated to a fourth, and the authoring family excluded. Fable 5.1 safeguards declined 97 tasks; those receive no credit.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.