AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.5 Flash

Google·

LLMsmidcloud

Overview

Fast-tier model in the Gemini 3.5 generation, announced at Google I/O 2026. GA at launch and set as the default model in the Gemini app and AI Mode in Google Search. Architecture is not publicly disclosed. Vendor-reported benchmarks: GPQA Diamond 90.4, MMMU-Pro 81.2, SWE-bench Verified 78. Independent (Artificial Analysis, ~5 days post-launch): HLE 40.2 (exact match to vendor figure), Terminal-Bench 2.1 76.2, AA Intelligence Index v4.0: 55. Distinct from 'Gemini Omni', a separate video/world model announced at the same event.

Capabilities and innovations

Default model in Gemini app and Search AI ModeGA at launch (2026-05-19)Multimodal inputOutput speed 207.9 tok/s — AA Speed rank #2/148 at launch (Speed is the product identity behind the Flash name)

Benchmarks

BenchmarkScoreSource
GPQA Diamond90.4%vendor
GPQA Diamond 90.4.
Humanity’s Last Exam (no tools)40.2%independent
HLE 40.2 — Artificial Analysis, ~5 days post-launch. Exact match to Google's own figure (delta = 0). Confirmed as a component of Intelligence Index v4.0.
SWE-bench Verified78%vendor
SWE-bench Verified 78.
MMLU-Pro89.52%unsourced
Terminal-Bench 2.x76.2%independent
Terminal-Bench 2.1 76.2 — Artificial Analysis. Note: this is Terminal-Bench 2.1, not the older 2.0.
MMMU-Pro81.2%vendor
MMMU-Pro 81.2 (vendor). Note: llm-stats reports 83.6 in an independent run; tracked discrepancy.
  • SWE-bench Pro: not reported by the vendor
  • MMLU: not reported by the vendor

API pricing

$1.5 per million input tokens, $9 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-05-21)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.