Gemini 3.6 Flash
Google·
Overview
Workhorse model in the Gemini 3 series, based on Gemini 3.5 Flash, delivering better coding, knowledge work and multimodal performance at improved token-efficiency (Google reports ~17% fewer output tokens than 3.5 Flash). 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch across the Gemini app, AI Studio and Gemini API. Vendor-reported (model card): SWE-Bench Pro 58.7, Terminal-Bench 2.1 78.0, OSWorld-Verified 83.0, MLE-Bench 63.9, CharXiv Reasoning (with tools) 89.4, GDM-MRCR v2 (128k) 91.8. Google's launch benchmark table also lists GPQA Diamond 90.4 (unchanged vs 3.5 Flash). Announced alongside 3.5 Flash-Lite and 3.5 Flash Cyber.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.43% | third party Third-party blog report of 90,4% GPQA. Not used as primary value — two independent evaluations (Vals AI 93,43%, Artificial Analysis 92,83%) sit about 3 points higher. |
| Humanity’s Last Exam (no tools) | 38.3% | third party HLE 38,3% no-tools, measured by Artificial Analysis (high effort). |
| SWE-bench Pro | 58.7% | vendor SWE-Bench Pro (Public) 58.7 — official Gemini 3.6 Flash model card, results as of July 2026. |
| MMLU-Pro | 89.28% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 2.x | 78% | vendor Terminal-Bench 2.1 (Terminus-2 harness) 78.0 — official model card. |
| MMMU-Pro | 83.2% | third party MMMU-Pro 83,2%, measured by Artificial Analysis (high effort). |
API pricing
$1.5 per million input tokens, $7.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-21)
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.