AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.6 Flash

Google·

LLMsmidcloud

Overview

Workhorse model in the Gemini 3 series, based on Gemini 3.5 Flash, delivering better coding, knowledge work and multimodal performance at improved token-efficiency (Google reports ~17% fewer output tokens than 3.5 Flash). 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch across the Gemini app, AI Studio and Gemini API. Vendor-reported (model card): SWE-Bench Pro 58.7, Terminal-Bench 2.1 78.0, OSWorld-Verified 83.0, MLE-Bench 63.9, CharXiv Reasoning (with tools) 89.4, GDM-MRCR v2 (128k) 91.8. Google's launch benchmark table also lists GPQA Diamond 90.4 (unchanged vs 3.5 Flash). Announced alongside 3.5 Flash-Lite and 3.5 Flash Cyber.

Capabilities and innovations

Multimodal input (text, image, audio, video)Agentic workflows and coding1M-token contextGA at launch (2026-07-21)~17% fewer output tokens than Gemini 3.5 Flash at comparable or better quality (token-efficiency)

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.43%third party
Third-party blog report of 90,4% GPQA. Not used as primary value — two independent evaluations (Vals AI 93,43%, Artificial Analysis 92,83%) sit about 3 points higher.
Humanity’s Last Exam (no tools)38.3%third party
HLE 38,3% no-tools, measured by Artificial Analysis (high effort).
SWE-bench Pro58.7%vendor
SWE-Bench Pro (Public) 58.7 — official Gemini 3.6 Flash model card, results as of July 2026.
MMLU-Pro89.28%independent
vals.ai independent evaluation (accessed 2026-07-25).
Terminal-Bench 2.x78%vendor
Terminal-Bench 2.1 (Terminus-2 harness) 78.0 — official model card.
MMMU-Pro83.2%third party
MMMU-Pro 83,2%, measured by Artificial Analysis (high effort).

API pricing

$1.5 per million input tokens, $7.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-21)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.