TMThinking Machines: healthcare benchmark results
3 models · 7 results · snapshot reviewed September 28, 2026
Everything the index holds for Thinking Machines, gathered in one place: Inkling, Inkling Small, openai-agents + TML Inkling 256K. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
| model | benchmark | score | index position |
|---|---|---|---|
| Inkling Subject suites release set: 257 tasks across eight subjects; one answer per task, no tools; blind cross-family grading. 95% bootstrap CI 32.3–39.0. Harness snapshot September 10, 2026; not the separate 89-task incretin ranking. | Health Optimization Bench | 35.6 | 10 of 16 |
| Inkling model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.23 pp; $0.127025/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 41.19% | 52 of 102 |
| Inkling model ID thinkingmachines/inkling; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.844 pp; $0.165561/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 85.41% | 24 of 104 |
| Inkling (xhigh) Artificial Analysis independent evaluation; current six-evaluation healthcare composite, including GDPval-AA v2.1, AA-Briefcase v1.1 and AutomationBench-AA. Configuration is shown in the model name. Board snapshot checked September 28, 2026; individual run dates are not disclosed. Display rounded to whole index points. | Artificial Analysis Healthcare & Medical Index | 25 | 21 of 25 source board: 77 |
| Inkling Small model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 2.206 pp; $0.016656/test; source snapshot 2026-09-26; run date not published | MedCode (Vals AI) | 37.89% | 70 of 102 |
| Inkling Small model ID thinkingmachines/inkling-small; reasoning_effort=0.99; temperature=1; top_p=1; max_output_tokens=30000; standard error 1.87 pp; $0.019018/test; source snapshot 2026-09-26; run date not published | MedScribe (Vals AI) | 84.11% | 34 of 104 |
| openai-agents + TML Inkling 256K All Domains pass@1; PA 4.0%, UM 16.0%, CM 4.0%; submitted 2026-07-24; run date not published | CHI-Bench | 8.0% | 35 of 44 source board: 45 |
Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.
Which healthcare benchmarks does Thinking Machines appear on?
As of September 28, 2026, Thinking Machines models hold 7 indexed results across 5 tracked benchmarks, through Inkling, Inkling Small, openai-agents + TML Inkling 256K.
Where does Thinking Machines have the highest indexed score?
Thinking Machines does not have the highest indexed score on any tracked board in this snapshot.
Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, SpaceXAI, xAI, Xiaomi, MiniMax, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.