Gemini 3.7 Flash
Google·
Overview
Workhorse successor to Gemini 3.6 Flash, shipped three weeks after it. Google states the model was not trained from scratch — it replaces the predecessor via algorithmic improvements and user feedback. Same 1,048,576-token context window as 3.6 Flash, multimodal input (text, image, audio, video, files). Live through the Gemini API in AI Studio, Android Studio, Google Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported coding deltas against 3.6 Flash: DeepSWE v1.1 49.0 to 65.3, FrontierCode 1.1 Main 34.4 to 43.6. Introductory API pricing of $0.75 / $3.75 per MTok runs until 2026-12-31, after which Google lists $1.50 / $7.50.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.94% | independent GPQA Diamond 93.94%, measured by Vals AI (accessed 2026-08-17). Same harness as the Gemini 3.6 Flash entry (93.43), so the 3.6 to 3.7 delta stays within one harness. Artificial Analysis measures 94.55 at high effort. |
| Humanity’s Last Exam (no tools) | 47.87% | third party HLE 47.87% no-tools, measured by Artificial Analysis at high effort. Same source and setting as the Gemini 3.6 Flash entry (38.3). |
| MMLU-Pro | 90.12% | independent MMLU Pro 90.12%, measured by Vals AI (accessed 2026-08-17). Same harness as the Gemini 3.6 Flash entry (89.28). |
| Terminal-Bench 2.x | 77.53% | independent Terminal-Bench 2.1 77.53%, measured by Vals AI (accessed 2026-08-17). Artificial Analysis reports 85.77 at high effort; the Vals AI value is kept as primary because it is the independent measurement and lands within half a point of the 78.0 recorded for Gemini 3.6 Flash, keeping the lineage readable. |
| MMMU-Pro | 85.49% | third party MMMU-Pro 85.49%, measured by Artificial Analysis at high effort. Same source and setting as the Gemini 3.6 Flash entry (83.2). |
- SWE-bench Pro: independent evaluation pending — Google's launch coverage reports DeepSWE v1.1 and FrontierCode 1.1 instead; the Gemini 3.7 Flash model card was not reachable from this environment, and no independent SWE-bench Pro run is published (the predecessor scored 58.7 per its model card).
- Terminal-Bench 3.0: not reported by the vendor
API pricing
$0.375 per million input tokens, $1.875 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-17)
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.