Health Evals

TMThinking Machines: healthcare benchmark results

3 models · 7 results · snapshot reviewed September 28, 2026

Everything the index holds for Thinking Machines, gathered in one place: Inkling, Inkling Small, openai-agents + TML Inkling 256K. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Inkling
Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 32.3–39.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking.
Health Optimization Bench35.610 of 16
Inkling
model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.23 pp; $0.127025/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)41.19%52 of 102
Inkling
model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.844 pp; $0.165561/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)85.41%24 of 104
Inkling (xhigh)
Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points.
Artificial Analysis Healthcare & Medical Index2521 of 25
source board: 77
Inkling Small
model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.206 pp; $0.016656/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)37.89%70 of 102
Inkling Small
model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.87 pp; $0.019018/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)84.11%34 of 104
openai-agents + TML Inkling 256K
All Domains pass@1; PA 4.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published
CHI-Bench8.0%35 of 44
source board: 45

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does Thinking Machines appear on?

As of September 28, 2026, Thinking Machines models hold 7 indexed results across 5 tracked benchmarks, through Inkling, Inkling Small, openai-agents + TML Inkling 256K.

Where does Thinking Machines have the highest indexed score?

Thinking Machines does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.