AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.5 Flash-Lite

Google·

LLMssmallcloud

Overview

Cost-efficient, low-latency tier of the Gemini 3 series, based on Gemini 3.1 Flash-Lite and optimized for high-volume, latency-sensitive tasks (translation, classification, document processing) as well as agentic workflows. 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch and rolling out to Google Search. Vendor-reported (model card): SWE-Bench Pro 54.2, Terminal-Bench 2.1 54.0, OSWorld-Verified 74.0, MLE-Bench 39.2, CharXiv Reasoning (with tools) 76.5, GDM-MRCR v2 (128k) 72.2. Reported output speed ~350 tokens/s.

Capabilities and innovations

Multimodal input (text, image, audio, video)High-throughput, low-latency tasksAgentic workflows1M-token contextRolling out to Google SearchOutput speed ~350 tok/s at $0.30 / $2.50 per 1M tokens (cost-efficient high-volume tier)

Benchmarks

BenchmarkScoreSource
GPQA Diamond83.84%independent
vals.ai independent evaluation (accessed 2026-07-25).
Humanity’s Last Exam (no tools)17.5%third party
HLE 17,5% no-tools, measured by Artificial Analysis.
SWE-bench Pro54.2%vendor
SWE-Bench Pro (Public) 54.2 — official Gemini 3.5 Flash-Lite model card, results as of July 2026.
MMLU-Pro85.84%independent
vals.ai independent evaluation (accessed 2026-07-25).
Terminal-Bench 2.x54%vendor
Terminal-Bench 2.1 (Terminus-2 harness) 54.0 — official model card.
MMMU-Pro79%third party
MMMU-Pro 79,0%, measured by Artificial Analysis.

API pricing

$0.3 per million input tokens, $2.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-21)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.