AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemini 4 Argon

Google·

LLMsflagshipcloud

Overview

First model of the Gemini 4 generation and Google's first flagship since Gemini 3.1 Pro. Phased rollout: at announcement only defenders in Google's Fairwind Program have access, with paid API customers and Google AI Ultra subscribers next and no date given. Google says the model was trained specifically for defensive cyber work, to find, validate and patch software vulnerabilities autonomously. 1M-token context and up to 1M output tokens in a single run (up from 64K). Introductory API pricing of $2 / $10 per MTok with 95% off cached input; Vals AI lists $4 / $20 on its own run. Vendor-reported (launch post, 2026-09-30): DeepSWE v1.1 77.9. Google published no GPQA, MMLU-Pro or SWE-bench Pro figure.

Capabilities and innovations

Agentic coding and long-horizon software engineeringCyber defense (vulnerability discovery, validation and patching)Multi-document knowledge work1M-token context, up to 1M output tokensMultimodal input (text, image, video, files)First Gemini 4 generation modelFairwind Program restricted access at launch1M output tokens in a single run

Benchmarks

BenchmarkScoreSource
Humanity’s Last Exam (no tools)57.09%third party
HLE 57.09% no-tools, measured by Artificial Analysis at high effort (accessed 2026-10-01), the same source and setting as the Gemini 3.8 Flash entry (47.82). Recorded as no-tools for cross-model comparability.
Terminal-Bench 4.057.58%independent
Terminal-Bench 4.0 57.58% (±2.31, rank 5/42), measured by Vals AI, reasoning effort high, temperature 1.0 (accessed 2026-10-01). Artificial Analysis measures 57.07; launch coverage quotes 57.4 from Google's own table. Not on the official tbench.ai 4.0 leaderboard yet. 4.0 is its own generation and not comparable with the 2.x or 3.0 fields.
DeepSWE v1.177.9%vendor
DeepSWE v1.1 77.9, Google's launch post (2026-09-30), labelled v1.1 and called a new state of the art by Google. Vendor self-report over its own model. The launch post and Datacurve's board were not reachable from this environment; the figure is taken from consistent launch coverage (TechCrunch, VentureBeat, DataCamp). Comparison rows for other vendors' models in Google's table are not carried over.
  • GPQA Diamond: independent evaluation pending — Not in Google's launch table. Vals AI and Artificial Analysis had run Argon on launch day but neither had published a GPQA Diamond value yet (checked 2026-10-01).
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor — Google leads with DeepSWE v1.1 instead, the pattern this catalog has seen since Jun 2026.
  • MMLU: not reported by the vendor
  • MMLU-Pro: independent evaluation pending — Not in Google's launch table; no Vals AI value yet (checked 2026-10-01).
  • Terminal-Bench 2.x: not reported by the vendor — No Terminal-Bench 2.x run reported. Vals AI and Artificial Analysis measure 4.0 only, recorded in terminal_bench_4.
  • Terminal-Bench 3.0: not reported by the vendor
  • MMMU-Pro: independent evaluation pending — Artificial Analysis lists no MMMU-Pro value for Argon yet (checked 2026-10-01).

API pricing

$2 per million input tokens, $10 per million output tokens (USD, provider list price, checked 2026-10-01)

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.