AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 3.7 Flash

Google·

LLMsmidcloud

Overview

Workhorse successor to Gemini 3.6 Flash, shipped three weeks after it. Google states the model was not trained from scratch — it replaces the predecessor via algorithmic improvements and user feedback. Same 1,048,576-token context window as 3.6 Flash, multimodal input (text, image, audio, video, files). Live through the Gemini API in AI Studio, Android Studio, Google Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported coding deltas against 3.6 Flash: DeepSWE v1.1 49.0 to 65.3, FrontierCode 1.1 Main 34.4 to 43.6. Introductory API pricing of $0.75 / $3.75 per MTok runs until 2026-12-31, after which Google lists $1.50 / $7.50.

Capabilities and innovations

Multimodal input (text, image, audio, video, files)Agentic workflows and coding1M-token contextAvailable via Gemini API, Android Studio and AntigravitySuccessor built by replacing 3.6 Flash through algorithmic improvements rather than a from-scratch training run

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.94%independent
GPQA Diamond 93.94%, measured by Vals AI (accessed 2026-08-17). Same harness as the Gemini 3.6 Flash entry (93.43), so the 3.6 to 3.7 delta stays within one harness. Artificial Analysis measures 94.55 at high effort.
Humanity’s Last Exam (no tools)47.87%third party
HLE 47.87% no-tools, measured by Artificial Analysis at high effort. Same source and setting as the Gemini 3.6 Flash entry (38.3).
MMLU-Pro90.12%independent
MMLU Pro 90.12%, measured by Vals AI (accessed 2026-08-17). Same harness as the Gemini 3.6 Flash entry (89.28).
Terminal-Bench 2.x77.53%independent
Terminal-Bench 2.1 77.53%, measured by Vals AI (accessed 2026-08-17). Artificial Analysis reports 85.77 at high effort; the Vals AI value is kept as primary because it is the independent measurement and lands within half a point of the 78.0 recorded for Gemini 3.6 Flash, keeping the lineage readable.
MMMU-Pro85.49%third party
MMMU-Pro 85.49%, measured by Artificial Analysis at high effort. Same source and setting as the Gemini 3.6 Flash entry (83.2).
  • SWE-bench Pro: independent evaluation pending — Google's launch coverage reports DeepSWE v1.1 and FrontierCode 1.1 instead; the Gemini 3.7 Flash model card was not reachable from this environment, and no independent SWE-bench Pro run is published (the predecessor scored 58.7 per its model card).
  • Terminal-Bench 3.0: not reported by the vendor

API pricing

$0.375 per million input tokens, $1.875 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-17)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.