AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.8 Flash

Google·

LLMsmidcloud

Overview

Third Flash-tier release in six weeks and, per the model card, based on Gemini 3.7 Flash rather than a new training run. Google calls it its most intelligent Flash model yet, with the gains concentrated in software engineering and agentic knowledge work, and describes the model as spending extra reasoning steps and iterative tool calls on hard tasks instead of answering early. 1M-token context, 64K output, knowledge cutoff March 2026, three effort levels. Live at launch in the Gemini app, AI Mode, AI Studio, the Gemini API, Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported (model card, September 2026): DeepSWE v1.1 73.7, Terminal-bench 2.1 89.4, Terminal-bench 4.0 19.1 (a newer generation that belongs in neither the 2.x nor the 3.0 field), HLE-Verified 54.9, Vals Finance Agent v2 61.4, Harvey's Legal Agent Benchmark 10.0, OSWorld-2.0 59.0, CharXiv Reasoning 86.2, LABBench2 86.2, GDPval-AA v2 Elo 1545; Google's API documentation adds SWE-Bench Pro 61.6 and SWE-Atlas 51.9. Introductory API pricing of $0.75 / $3.75 per MTok until 2026-12-31, then $1.50 / $7.50. Announced together with Gemini 3.8 Flash Cyber.

Capabilities and innovations

Multimodal input (text, image, audio, video)Agentic workflows and coding1M-token context, 64K outputConfigurable effort levels (low / medium / high)Available via Gemini API, AI Studio, Antigravity and the Gemini Enterprise Agent PlatformIterates on Gemini 3.7 Flash three weeks after its release (model card: 'based on Gemini 3.7 Flash')Extra reasoning steps and iterative tool calls on complex tasks, at the cost of more output tokens at higher effort levels

Benchmarks

BenchmarkScoreSource
GPQA Diamond94.44%independent
GPQA Diamond 94.44% (±1.48, rank 4/138), measured by Vals AI (accessed 2026-09-03). Same harness as the Gemini 3.7 Flash (93.94) and 3.6 Flash (93.43) entries, so the lineage stays within one harness. Artificial Analysis measures 95.25 at high effort.
Humanity’s Last Exam (no tools)47.82%third party
HLE 47.82% no-tools, measured by Artificial Analysis at high effort — same source and setting as the Gemini 3.7 Flash (47.87) and 3.6 Flash (38.3) entries. Google's API documentation lists HLE 45.4 (3.7 Flash: 45.7); the model card's HLE-Verified 54.9 is a different item set and is not recorded here.
SWE-bench Pro61.6%vendor
SWE-Bench Pro 61.6 (Gemini 3.7 Flash: 60.4), Google's own figure from the Gemini API model documentation as quoted by launch coverage (DataCamp, eesel AI, 2026-09-02). The documentation page was not reachable from this environment and the model card table carries no SWE-Bench Pro row. No independent run is published yet.
MMLU-Pro90.22%independent
MMLU Pro 90.22% (±0.29, rank 5/138), measured by Vals AI (accessed 2026-09-03). Same harness as the Gemini 3.7 Flash entry (90.12).
Terminal-Bench 2.x81.27%independent
Terminal-Bench 2.1 81.27% (±0.38, rank 4/60), measured by Vals AI (accessed 2026-09-03). Kept as primary because it is the independent measurement and the same harness as the 3.7 Flash entry (77.53). Google's model card reports 89.4 on Terminal-bench 2.1.
MMMU-Pro85.61%third party
MMMU-Pro 85.61%, measured by Artificial Analysis at high effort — same source and setting as the Gemini 3.7 Flash entry (85.49). Vals AI measures 89.08 on its own MMMU Pro run.
  • Terminal-Bench 3.0: not reported by the vendor — The model card reports Terminal-bench 2.1 (89.4) and Terminal-bench 4.0 (19.1) but no 3.0 run. 4.0 is a separate, harder generation and is recorded in neither field.

API pricing

$0.75 per million input tokens, $3.75 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-03)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.