Health Evals

SSpaceXAI: healthcare benchmark results

7 models · 15 results · snapshot reviewed September 28, 2026

Everything the index holds for SpaceXAI, gathered in one place: Grok 4, Grok 4.5, Grok 4 Fast (Reasoning), Grok 4.20 (Reasoning), Grok 4 Fast (Non-Reasoning), Grok 4.1 Fast Non-Reasoning, Grok 4.1 Fast (Reasoning). Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

modelbenchmarkscoreindex position
Grok 4
model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.206 pp; $0.034103/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)38.08%68 of 102
Grok 4
model ID grok/grok-4-0709; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.084 pp; $0.063955/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)78.15%61 of 104
Grok 4
95% bootstrap CI 54.9–60.8; 3 runs; temperature 0; zero-shot, closed-book
WHBench57.9%9 of 22
Grok 4.5
model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.313 pp; $0.048461/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)43.29%42 of 102
Grok 4.5
model ID grok/grok-4.5; reasoning_effort=high; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.944 pp; $0.033208/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)86.88%16 of 104
Grok 4 Fast (Reasoning)
model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.941 pp; $0.002143/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)37.38%71 of 102
Grok 4 Fast (Reasoning)
model ID grok/grok-4-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.137 pp; $0.002535/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)81.63%49 of 104
Grok 4.20 (Reasoning)
model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.124 pp; $0.036189/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)32.16%86 of 102
Grok 4.20 (Reasoning)
model ID grok/grok-4.20-0309-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.095 pp; $0.031303/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)63.41%98 of 104
Grok 4 Fast (Non-Reasoning)
model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.974 pp; $0.002149/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)30.04%92 of 102
Grok 4 Fast (Non-Reasoning)
model ID grok/grok-4-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.871 pp; $0.002056/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)79.72%56 of 104
Grok 4.1 Fast Non-Reasoning
model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.921 pp; $0.002193/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)28.35%95 of 102
Grok 4.1 Fast Non-Reasoning
model ID grok/grok-4-1-fast-non-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 2.04 pp; $0.001782/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)77.46%65 of 104
Grok 4.1 Fast (Reasoning)
model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.992 pp; $0.002108/test; source snapshot 2026-09-26; run date not published
MedCode (Vals AI)28.08%96 of 102
Grok 4.1 Fast (Reasoning)
model ID grok/grok-4-1-fast-reasoning; temperature=1; top_p=0.95; max_output_tokens=30000; standard error 1.866 pp; $0.002387/test; source snapshot 2026-09-26; run date not published
MedScribe (Vals AI)78.73%59 of 104

Positions refer to indexed rows, including configuration variants, and are not controlled comparisons across sources or graders. Source board sizes appear separately where coverage differs.

Which healthcare benchmarks does SpaceXAI appear on?

As of September 28, 2026, SpaceXAI models hold 15 indexed results across 3 tracked benchmarks, through Grok 4, Grok 4.5, Grok 4 Fast (Reasoning), Grok 4.20 (Reasoning), Grok 4 Fast (Non-Reasoning), Grok 4.1 Fast Non-Reasoning, Grok 4.1 Fast (Reasoning).

Where does SpaceXAI have the highest indexed score?

SpaceXAI does not have the highest indexed score on any tracked board in this snapshot.

Other labs with pages: OpenAI, Anthropic, Google, Alibaba, Meta, Moonshot AI, DeepSeek, xAI, Xiaomi, MiniMax, Thinking Machines, zAI, NVIDIA, SpaceX AI, Zhipu AI, Mistral, Ant Group, Poolside, Zhipu, Z.ai, Microsoft, Baichuan, Tencent, Inception, Cohere. The full field is on the index.