Grok 4.7
X.AI·
Overview
Successor to Grok 4.6 on a new, larger base model, trained with longer reinforcement learning on multi-hour tasks. First release under the SpaceXAI name. Same $2 / $6 per million input/output tokens as Grok 4.6 (OpenRouter lists $1.60 / $4.80). 500K-token context. Red-team capabilities are invite-only for selected security partners. xAI's launch table reports only new-generation and domain benchmarks; the classic fields come from independent runs (Artificial Analysis, Vals AI).
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Humanity’s Last Exam (no tools) | 43.1% | third party HLE 43.1% no-tools, measured by Artificial Analysis at xhigh effort. Same source as the Grok 4.6 entry (42.9). |
| Terminal-Bench 2.x | 73.41% | independent Terminal-Bench 2.1 73.41% (±1.50), measured by Vals AI with the Terminus 2 harness (accessed 2026-09-23). Same harness as the Grok 4.6 entry (78.28), so the step down is a measured regression on this benchmark, not a harness change. xAI reports no 2.x figure for Grok 4.7. |
| Terminal-Bench 4.0 | 37.6% | independent Terminal-Bench 4.0 37.6% (± 3.5, 95% CI), official leaderboard maintained by the Terminal-Bench team (Stanford / Harbor / Laude Institute), xhigh effort, agent harness: Grok Build (accessed 2026-09-23). Matches xAI's launch figure; Grok 4.6 sits at 20.3 on the same board. Vals AI measures 28.28 and Artificial Analysis 25.8 in their own harnesses. |
| DeepSWE v1.1 | 71% | vendor DeepSWE v1.1 71.0 at high effort, xAI launch post. Vendor self-report over its own model. |
- GPQA Diamond: independent evaluation pending — Not reported by xAI; Artificial Analysis lists it as null and Vals AI has archived GPQA (checked 2026-09-23).
- SWE-bench Pro: not reported by the vendor — xAI reports DeepSWE v1.1 instead, recorded in deepswe_v1_1.
- MMLU-Pro: not reported by the vendor
- MMMU-Pro: independent evaluation pending — Artificial Analysis lists it as null (checked 2026-09-23).
- ARC-AGI-2: independent evaluation pending — ARC Prize results stop at Grok 4.6 (checked 2026-09-23).
- Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.
API pricing
$2 per million input tokens, $6 per million output tokens (USD, provider list price, checked 2026-09-23)
Links
More from X.AI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.