Grok 4.5
X.AI·
LLMsflagshipcloud
Overview
Coding- and agentic-focused flagship from xAI, trained alongside Cursor. Available in Grok Build, Cursor and the console at $2 / $6 per million input/output tokens.
Capabilities and innovations
Agentic codingReasoning-first
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 92.93% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Humanity’s Last Exam (no tools) | 40.3% | third party HLE 40,3% no-tools, measured by Artificial Analysis (high effort, the highest setting Artificial Analysis runs for Grok 4.5). |
| SWE-bench Pro | 64.7% | vendor |
| MMLU-Pro | 89.22% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 2.x | 67.79% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| MMMU-Pro | 80.4% | third party MMMU-Pro 80,4%, measured by Artificial Analysis (high effort). |
API pricing
$2 per million input tokens, $6 per million output tokens (USD, provider list price, checked 2026-07-09)
Links
More from X.AI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.