AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.1 Pro

Google·

LLMsflagshipcloud

Overview

Major reasoning upgrade with record GPQA and 2x ARC-AGI-2 improvement over Gemini 3 Pro.

Capabilities and innovations

1M context window64K outputNatively multimodalAgentic workflowsARC-AGI-2 Regime ChangeDeep Think ReasoningSVG Animation Generation

Benchmarks

BenchmarkScoreSource
GPQA Diamond94.3%third party
GPQA Diamond 94,3 — Top gelisteter Wert (nach Ausschluss des Orchestrators Sakana Fugu Ultra) auf dem benchlm GPQA-D Leaderboard.
Humanity’s Last Exam (no tools)44.4%unsourced
SWE-bench Verified80.6%unsourced
SWE-bench Pro54.2%unsourced
MMLU92.6%unsourced
MMLU-Pro90.99%unsourced
Terminal-Bench 2.x70.7%vendor
Terminal-Bench 2.1: 70,7% — von OpenAI im eigenen Vergleich gemessen (Fremdmessung), nicht von Google berichtet.
  • MMMU-Pro: not reported by the vendor

Reliability

Hallucination rate (Vectara HHEM)
10.4% (lower is better)independent
Agentic tool use (τ-bench)
76.5% (higher is better)third party

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.