Health Optimization Bench: published results
Arcophos · 257 tasks across eight subject suites · index updated September 28, 2026
Claude Fable 5 has the highest indexed numerical score on Health Optimization Bench, 70.9 as of 2026-09, per Health Optimization Bench — subject suites results. Questions in eight areas of preventive and optimization medicine, grounded in primary evidence and scored against task-specific rubrics. The current main ranking covers 257 released tasks, separate from the 89-task incretin therapeutics evidence suite.
This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthoptimizationbench.com. Every published row in the index appears below.
Published results
Result detail
sources for this board| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 68.0–73.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 70.9 | 2026-09 | |
| 2 | Claude Opus 5 Anthropic Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 66.2–72.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 69.3 | 2026-09 | |
| 3 | Grok 4.6 xAI Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.8–69.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 66.8 | 2026-09 | |
| 4 | GPT-5.6 Sol (max) OpenAI Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 63.4–69.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Maximum reasoning effort. | 66.6 | 2026-09 | |
| 5 | GPT-5.6 Sol (high) OpenAI Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 61.5–67.6. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. High reasoning effort. | 64.6 | 2026-09 | |
| 6 | Kimi K3 Moonshot AI Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 56.5–63.3. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 59.9 | 2026-09 | |
| 7 | Muse Spark Meta Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 53.9–60.5. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 57.2 | 2026-09 | |
| 8 | Claude Fable 5.1 Anthropic Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 42.2–52.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. Vendor safeguards declined 97/257 tasks, scored with no credit; mean over answered tasks is 75.9. | 47.3 | 2026-09 | |
| 9 | Gemini 3.6 Google Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 36.2–43.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 39.7 | 2026-09 | |
| 10 | TM | Inkling Thinking Machines Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 32.3–39.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 35.6 | 2026-09 |
| 11 | Claude Sonnet 5 Anthropic Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 31.6–37.8. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 34.6 | 2026-09 | |
| 12 | GLM 5.2 Zhipu Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 18.1–23.4. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 20.7 | 2026-09 | |
| 13 | MiniMax M3 MiniMax Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.7–20.9. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 18.3 | 2026-09 | |
| 14 | MAI Thinking Microsoft AI Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 15.1–20.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 17.5 | 2026-09 | |
| 15 | Mistral Medium 3.5 Mistral Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 7.4–11.1. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 9.2 | 2026-09 | |
| 16 | Nemotron 3.5 Lightning NVIDIA Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 3.6–6.2. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | 4.9 | 2026-09 | |
Scores preserve their source precision, with any scale conversion documented (independently run). The main ranking moved from the 89-task incretin suite to 257 subject-suite tasks. This page shows only the 257-task September 10 snapshot; the two task sets must not be pooled. Three grader families vote, with split decisions escalated to a fourth, and the authoring family excluded. Fable 5.1 safeguards declined 97 tasks; those receive no credit. The line under each model names the document its score was read from; rows marked source pending are still awaiting a documented first-party source. Full citations are on the sources page.
About the benchmark
| publisher | Arcophos |
|---|---|
| category | rubric-graded benchmarks |
| released | 2026-08 |
| size | 257 tasks across eight subject suites |
| scale | 0-100 rubric credit, higher better |
| result basis | independently run |
| source | healthoptimizationbench.com |
| official page | healthoptimizationbench.com |
| last frontier result | 2026-09 |
What is Health Optimization Bench?
Health Optimization Bench is a rubric-graded benchmark from Arcophos, released 2026-08: 257 tasks across eight subject suites, scored on a 0-100 rubric credit scale. Questions in eight areas of preventive and optimization medicine, grounded in primary evidence and scored against task-specific rubrics. The current main ranking covers 257 released tasks, separate from the 89-task incretin therapeutics evidence suite.
Which model leads Health Optimization Bench?
Claude Fable 5 (Anthropic) has the highest indexed numerical score on Health Optimization Bench at 70.9 (evaluation setups may differ), per Health Optimization Bench — subject suites results, as of 2026-09.
Where do the Health Optimization Bench numbers come from?
From healthoptimizationbench.com (independently run). The main ranking moved from the 89-task incretin suite to 257 subject-suite tasks. This page shows only the 257-task September 10 snapshot; the two task sets must not be pooled. Three grader families vote, with split decisions escalated to a fourth, and the authoring family excluded. Fable 5.1 safeguards declined 97 tasks; those receive no credit.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.