Grok 4.6
X.AI·
LLMsflagshipcloud
Overview
Reasoning-first flagship succeeding Grok 4.5, with a 500K-token context window and text + image input. Same $2 / $6 per million input/output tokens as Grok 4.5, with a 75% cache-hit discount.
Capabilities and innovations
500K token context windowAgentic codingReasoning-firstText and image input
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 94.7% | independent GPQA Diamond 94.70 (± 1.13), vals.ai, reasoning effort high (accessed 2026-08-13). Same harness as the Grok 4.5 entry (92.93). Artificial Analysis measures 94.9 on the same benchmark. |
| Humanity’s Last Exam (no tools) | 42.9% | third party HLE 42.9% no-tools, measured by Artificial Analysis (high effort). Same source and setting as the Grok 4.5 entry (40.3). |
| SWE-bench Verified | 95.6% | independent SWE-bench Verified 95.60 (± 0.92), vals.ai (accessed 2026-08-13). Era-1 benchmark, near saturation — recorded for continuity, not ranked. |
| MMLU-Pro | 89.4% | independent MMLU Pro 89.40 (± 0.30), vals.ai, reasoning effort high (accessed 2026-08-13). Same harness as the Grok 4.5 entry (89.22). |
| Terminal-Bench 2.x | 78.28% | independent Terminal-Bench 2.1 78.28 (± 2.09), vals.ai, reasoning effort high (accessed 2026-08-13). Same harness as the Grok 4.5 entry (67.79). Artificial Analysis reports 88.4 on Terminal-Bench v2.1 for the same model — the vals.ai value is kept as primary so the Grok 4.5 to 4.6 delta stays within one harness. |
- SWE-bench Pro: independent evaluation pending — xAI's Grok 4.6 comparison table drops SWE-bench Pro (Grok 4.5 scored 64.7); no independent re-run published yet.
- MMMU-Pro: independent evaluation pending — Artificial Analysis lists mmmuPro as null for Grok 4.6; Grok 4.5 was measured at 80.4.
API pricing
$2 per million input tokens, $6 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-13)
Links
More from X.AI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.