Gemini 3.1 Pro
Google·
LLMsflagshipcloud
Overview
Major reasoning upgrade with record GPQA and 2x ARC-AGI-2 improvement over Gemini 3 Pro.
Capabilities and innovations
1M context window64K outputNatively multimodalAgentic workflowsARC-AGI-2 Regime ChangeDeep Think ReasoningSVG Animation Generation
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 94.3% | third party GPQA Diamond 94,3 — Top gelisteter Wert (nach Ausschluss des Orchestrators Sakana Fugu Ultra) auf dem benchlm GPQA-D Leaderboard. |
| Humanity’s Last Exam (no tools) | 44.4% | unsourced |
| SWE-bench Verified | 80.6% | unsourced |
| SWE-bench Pro | 54.2% | unsourced |
| MMLU | 92.6% | unsourced |
| MMLU-Pro | 90.99% | unsourced |
| Terminal-Bench 2.x | 70.7% | vendor Terminal-Bench 2.1: 70,7% — von OpenAI im eigenen Vergleich gemessen (Fremdmessung), nicht von Google berichtet. |
- MMMU-Pro: not reported by the vendor
Reliability
- Hallucination rate (Vectara HHEM)
- 10.4% (lower is better)independent
- Agentic tool use (τ-bench)
- 76.5% (higher is better)third party
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.