Gemini 3.8 Flash
Google·
Overview
Third Flash-tier release in six weeks and, per the model card, based on Gemini 3.7 Flash rather than a new training run. Google calls it its most intelligent Flash model yet, with the gains concentrated in software engineering and agentic knowledge work, and describes the model as spending extra reasoning steps and iterative tool calls on hard tasks instead of answering early. 1M-token context, 64K output, knowledge cutoff March 2026, three effort levels. Live at launch in the Gemini app, AI Mode, AI Studio, the Gemini API, Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported (model card, September 2026): DeepSWE v1.1 73.7, Terminal-bench 2.1 89.4, Terminal-bench 4.0 19.1 (a newer generation that belongs in neither the 2.x nor the 3.0 field), HLE-Verified 54.9, Vals Finance Agent v2 61.4, Harvey's Legal Agent Benchmark 10.0, OSWorld-2.0 59.0, CharXiv Reasoning 86.2, LABBench2 86.2, GDPval-AA v2 Elo 1545; Google's API documentation adds SWE-Bench Pro 61.6 and SWE-Atlas 51.9. Introductory API pricing of $0.75 / $3.75 per MTok until 2026-12-31, then $1.50 / $7.50. Announced together with Gemini 3.8 Flash Cyber.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 94.44% | independent GPQA Diamond 94.44% (±1.48, rank 4/138), measured by Vals AI (accessed 2026-09-03). Same harness as the Gemini 3.7 Flash (93.94) and 3.6 Flash (93.43) entries, so the lineage stays within one harness. Artificial Analysis measures 95.25 at high effort. |
| Humanity’s Last Exam (no tools) | 47.82% | third party HLE 47.82% no-tools, measured by Artificial Analysis at high effort — same source and setting as the Gemini 3.7 Flash (47.87) and 3.6 Flash (38.3) entries. Google's API documentation lists HLE 45.4 (3.7 Flash: 45.7); the model card's HLE-Verified 54.9 is a different item set and is not recorded here. |
| SWE-bench Pro | 61.6% | vendor SWE-Bench Pro 61.6 (Gemini 3.7 Flash: 60.4), Google's own figure from the Gemini API model documentation as quoted by launch coverage (DataCamp, eesel AI, 2026-09-02). The documentation page was not reachable from this environment and the model card table carries no SWE-Bench Pro row. No independent run is published yet. |
| MMLU-Pro | 90.22% | independent MMLU Pro 90.22% (±0.29, rank 5/138), measured by Vals AI (accessed 2026-09-03). Same harness as the Gemini 3.7 Flash entry (90.12). |
| Terminal-Bench 2.x | 81.27% | independent Terminal-Bench 2.1 81.27% (±0.38, rank 4/60), measured by Vals AI (accessed 2026-09-03). Kept as primary because it is the independent measurement and the same harness as the 3.7 Flash entry (77.53). Google's model card reports 89.4 on Terminal-bench 2.1. |
| MMMU-Pro | 85.61% | third party MMMU-Pro 85.61%, measured by Artificial Analysis at high effort — same source and setting as the Gemini 3.7 Flash entry (85.49). Vals AI measures 89.08 on its own MMMU Pro run. |
- Terminal-Bench 3.0: not reported by the vendor — The model card reports Terminal-bench 2.1 (89.4) and Terminal-bench 4.0 (19.1) but no 3.0 run. 4.0 is a separate, harder generation and is recorded in neither field.
API pricing
$0.75 per million input tokens, $3.75 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-03)
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.