AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.7 Flash

Google·

LLMsmidcloud

Overview

Workhorse successor to Gemini 3.6 Flash, shipped three weeks after it. Google states the model was not trained from scratch — it replaces the predecessor via algorithmic improvements and user feedback. Same 1,048,576-token context window as 3.6 Flash, multimodal input (text, image, audio, video, files). Live through the Gemini API in AI Studio, Android Studio, Google Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported coding deltas against 3.6 Flash: DeepSWE v1.1 49.0 to 65.3, FrontierCode 1.1 Main 34.4 to 43.6. Introductory API pricing of $0.75 / $3.75 per MTok runs until 2026-12-31, after which Google lists $1.50 / $7.50.

Capabilities and innovations

Multimodal input (text, image, audio, video, files)Agentic workflows and coding1M-token contextAvailable via Gemini API, Android Studio and AntigravitySuccessor built by replacing 3.6 Flash through algorithmic improvements rather than a from-scratch training run

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.94%independent
GPQA Diamond 93.94%, measured by Vals AI (accessed 2026-08-17). Same harness as the Gemini 3.6 Flash entry (93.43), so the 3.6 to 3.7 delta stays within one harness. Artificial Analysis measures 94.55 at high effort.
Humanity’s Last Exam (no tools)47.87%third party
HLE 47.87% no-tools, measured by Artificial Analysis at high effort. Same source and setting as the Gemini 3.6 Flash entry (38.3).
SWE-bench Pro60.4%vendor
SWE-Bench Pro 60.4, Google's own figure for Gemini 3.7 Flash from the comparison table in the Gemini 3.8 Flash API documentation (3.8 Flash: 61.6), as quoted by launch coverage (DataCamp, eesel AI, 2026-09-02); the page was not reachable from this environment. Backfilled 2026-09-03 — the 3.7 Flash model card was never reachable and no independent run exists.
MMLU-Pro90.12%independent
MMLU Pro 90.12%, measured by Vals AI (accessed 2026-08-17). Same harness as the Gemini 3.6 Flash entry (89.28).
Terminal-Bench 2.x77.53%independent
Terminal-Bench 2.1 77.53%, measured by Vals AI (accessed 2026-08-17). Artificial Analysis reports 85.77 at high effort; the Vals AI value is kept as primary because it is the independent measurement and lands within half a point of the 78.0 recorded for Gemini 3.6 Flash, keeping the lineage readable.
MMMU-Pro85.49%third party
MMMU-Pro 85.49%, measured by Artificial Analysis at high effort. Same source and setting as the Gemini 3.6 Flash entry (83.2).
  • Terminal-Bench 3.0: not reported by the vendor

API pricing

$0.75 per million input tokens, $3.75 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-03)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.