Gemini 3.6 Flash
Google·
Overview
Workhorse model in the Gemini 3 series, based on Gemini 3.5 Flash, delivering better coding, knowledge work and multimodal performance at improved token-efficiency (Google reports ~17% fewer output tokens than 3.5 Flash). 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch across the Gemini app, AI Studio and Gemini API. Vendor-reported (model card): SWE-Bench Pro 58.7, Terminal-Bench 2.1 78.0, OSWorld-Verified 83.0, MLE-Bench 63.9, CharXiv Reasoning (with tools) 89.4, GDM-MRCR v2 (128k) 91.8. Google's launch benchmark table also lists GPQA Diamond 90.4 (unchanged vs 3.5 Flash). Announced alongside 3.5 Flash-Lite and 3.5 Flash Cyber.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.43% | third party Third-party blog report of 90,4% GPQA. Not used as primary value — two independent evaluations (Vals AI 93,43%, Artificial Analysis 92,83%) sit about 3 points higher. |
| Humanity’s Last Exam (no tools) | 38.3% | third party HLE 38,3% no-tools, measured by Artificial Analysis (high effort). |
| SWE-bench Pro | 58.7% | vendor SWE-Bench Pro (Public) 58.7 — official Gemini 3.6 Flash model card, results as of July 2026. |
| MMLU-Pro | 89.28% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 2.x | 78% | vendor Terminal-Bench 2.1 (Terminus-2 harness) 78.0 — official model card. |
| DeepSWE v1.1 | 49% | vendor DeepSWE v1.1 49.0 — official Gemini 3.6 Flash model card; Google repeats the figure as the baseline in the Gemini 3.7 Flash launch (49.0 to 65.3). Datacurve's own DeepSWE board was not reachable from this environment, so the value is not cross-checked against it. |
| MMMU-Pro | 83.2% | third party MMMU-Pro 83,2%, measured by Artificial Analysis (high effort). |
API pricing
$1.5 per million input tokens, $7.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-21)
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.