AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Grok 4

X.AI·

LLMsflagshipcloud

Overview

Next generation model with enhanced understanding of the physical world.

Capabilities and innovations

Physical world modelVideo reasoningRobot control

Benchmarks

BenchmarkScoreSource
GPQA Diamond88.1%unsourced
Humanity’s Last Exam (no tools)33.2%unsourced
SWE-bench Verified60.4%unsourced
MMLU90.1%unsourced
MMLU-Pro85.3%unsourced
ARC-AGI-215.9%independent
Grok 4 (Thinking) set a new commercial SOTA of 15.9% on ARC-AGI-2 (Jul 2025), independently verified by ARC Prize on a blind holdout set.
  • SWE-bench Pro: not reported by the vendor
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from X.AI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.