Gemini 3.5 Flash-Lite
Google·
Overview
Cost-efficient, low-latency tier of the Gemini 3 series, based on Gemini 3.1 Flash-Lite and optimized for high-volume, latency-sensitive tasks (translation, classification, document processing) as well as agentic workflows. 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch and rolling out to Google Search. Vendor-reported (model card): SWE-Bench Pro 54.2, Terminal-Bench 2.1 54.0, OSWorld-Verified 74.0, MLE-Bench 39.2, CharXiv Reasoning (with tools) 76.5, GDM-MRCR v2 (128k) 72.2. Reported output speed ~350 tokens/s.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 83.84% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Humanity’s Last Exam (no tools) | 17.5% | third party HLE 17,5% no-tools, measured by Artificial Analysis. |
| SWE-bench Pro | 54.2% | vendor SWE-Bench Pro (Public) 54.2 — official Gemini 3.5 Flash-Lite model card, results as of July 2026. |
| MMLU-Pro | 85.84% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 2.x | 54% | vendor Terminal-Bench 2.1 (Terminus-2 harness) 54.0 — official model card. |
| MMMU-Pro | 79% | third party MMMU-Pro 79,0%, measured by Artificial Analysis. |
API pricing
$0.3 per million input tokens, $2.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-21)
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.