AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Grok 4.7

X.AI·

LLMsflagshipcloud

Overview

Successor to Grok 4.6 on a new, larger base model, trained with longer reinforcement learning on multi-hour tasks. First release under the SpaceXAI name. Same $2 / $6 per million input/output tokens as Grok 4.6 (OpenRouter lists $1.60 / $4.80). 500K-token context. Red-team capabilities are invite-only for selected security partners. xAI's launch table reports only new-generation and domain benchmarks; the classic fields come from independent runs (Artificial Analysis, Vals AI).

Capabilities and innovations

500K token context windowAgentic codingConfigurable reasoningText and image inputAPI: grok-4.7New, larger base modelEncrypted reasoning content in the Responses API

Benchmarks

BenchmarkScoreSource
Humanity’s Last Exam (no tools)43.1%third party
HLE 43.1% no-tools, measured by Artificial Analysis at xhigh effort. Same source as the Grok 4.6 entry (42.9).
Terminal-Bench 2.x73.41%independent
Terminal-Bench 2.1 73.41% (±1.50), measured by Vals AI with the Terminus 2 harness (accessed 2026-09-23). Same harness as the Grok 4.6 entry (78.28), so the step down is a measured regression on this benchmark, not a harness change. xAI reports no 2.x figure for Grok 4.7.
Terminal-Bench 4.037.6%independent
Terminal-Bench 4.0 37.6% (± 3.5, 95% CI), official leaderboard maintained by the Terminal-Bench team (Stanford / Harbor / Laude Institute), xhigh effort, agent harness: Grok Build (accessed 2026-09-23). Matches xAI's launch figure; Grok 4.6 sits at 20.3 on the same board. Vals AI measures 28.28 and Artificial Analysis 25.8 in their own harnesses.
DeepSWE v1.171%vendor
DeepSWE v1.1 71.0 at high effort, xAI launch post. Vendor self-report over its own model.
  • GPQA Diamond: independent evaluation pending — Not reported by xAI; Artificial Analysis lists it as null and Vals AI has archived GPQA (checked 2026-09-23).
  • SWE-bench Pro: not reported by the vendor — xAI reports DeepSWE v1.1 instead, recorded in deepswe_v1_1.
  • MMLU-Pro: not reported by the vendor
  • MMMU-Pro: independent evaluation pending — Artificial Analysis lists it as null (checked 2026-09-23).
  • ARC-AGI-2: independent evaluation pending — ARC Prize results stop at Grok 4.6 (checked 2026-09-23).
  • Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.

API pricing

$2 per million input tokens, $6 per million output tokens (USD, provider list price, checked 2026-09-23)

Links

More from X.AI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.