AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.6 Flash

Google·

LLMsmidcloud

Overview

Workhorse model in the Gemini 3 series, based on Gemini 3.5 Flash, delivering better coding, knowledge work and multimodal performance at improved token-efficiency (Google reports ~17% fewer output tokens than 3.5 Flash). 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch across the Gemini app, AI Studio and Gemini API. Vendor-reported (model card): SWE-Bench Pro 58.7, Terminal-Bench 2.1 78.0, OSWorld-Verified 83.0, MLE-Bench 63.9, CharXiv Reasoning (with tools) 89.4, GDM-MRCR v2 (128k) 91.8. Google's launch benchmark table also lists GPQA Diamond 90.4 (unchanged vs 3.5 Flash). Announced alongside 3.5 Flash-Lite and 3.5 Flash Cyber.

Capabilities and innovations

Multimodal input (text, image, audio, video)Agentic workflows and coding1M-token contextGA at launch (2026-07-21)~17% fewer output tokens than Gemini 3.5 Flash at comparable or better quality (token-efficiency)

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.43%third party
Third-party blog report of 90,4% GPQA. Not used as primary value — two independent evaluations (Vals AI 93,43%, Artificial Analysis 92,83%) sit about 3 points higher.
Humanity’s Last Exam (no tools)38.3%third party
HLE 38,3% no-tools, measured by Artificial Analysis (high effort).
SWE-bench Pro58.7%vendor
SWE-Bench Pro (Public) 58.7 — official Gemini 3.6 Flash model card, results as of July 2026.
MMLU-Pro89.28%independent
vals.ai independent evaluation (accessed 2026-07-25).
Terminal-Bench 2.x78%vendor
Terminal-Bench 2.1 (Terminus-2 harness) 78.0 — official model card.
DeepSWE v1.149%vendor
DeepSWE v1.1 49.0 — official Gemini 3.6 Flash model card; Google repeats the figure as the baseline in the Gemini 3.7 Flash launch (49.0 to 65.3). Datacurve's own DeepSWE board was not reachable from this environment, so the value is not cross-checked against it.
MMMU-Pro83.2%third party
MMMU-Pro 83,2%, measured by Artificial Analysis (high effort).

API pricing

$1.5 per million input tokens, $7.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-21)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.