AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GPT-6 Astra

OpenAI·

LLMsflagshipcloud

Overview

OpenAI's GPT-6 generation flagship, launched 2026-09-03 in a phased rollout: participants in OpenAI's cybersecurity access program first, then ChatGPT Plus, Pro, Business and Enterprise, the API (model id gpt-6-astra), Azure and AWS Bedrock. 1.05M-token context window, 128K max output, knowledge cutoff 2026-04-30, reasoning effort selectable from none to max. First OpenAI model classified 'Critical' for cybersecurity under the Preparedness Framework: the general release ships with offensive-cyber behaviour constrained, and a fuller cyber variant is limited to vetted defenders. Priced at $10/M input and $50/M output ($1.00 cached input, $12.50 cache write); requests above 272K input tokens are billed at 2x input and 1.5x output for the whole request, and a Fast mode costs 2x. OpenAI's launch table reports Terminal-Bench 4.0, Terminal-Bench-Science, OSWorld 2.0, FrontierMath Tier 4, DeepSWE v1.1 and HLE with tools, but no SWE-bench Pro, Terminal-Bench 2.x, MMLU-Pro or no-tools HLE figure, so the tracked fields here come from independent runs (Artificial Analysis, Vals AI, ARC Prize) where one exists.

Capabilities and innovations

Frontier reasoningAgentic codingComputer useBrowser useTool use / function callingCybersecurity tasks (vulnerability discovery, exploit development)Vision inputAPI: gpt-6-astra1.05M-token context windowPreparedness Framework 'Critical' cyber classification with a gated cyber variantReasoning effort none to max, plus Fast mode at 2x priceCache reads at 10% of input price

Benchmarks

BenchmarkScoreSource
GPQA Diamond96.1%third party
GPQA Diamond 96.1% at max effort, measured by Artificial Analysis (xhigh: 96.3%), rank 1 on the AA board at intake (accessed 2026-09-04). OpenAI's launch table reports 96.0%.
Terminal-Bench 2.x87.27%independent
Terminal-Bench 2.1 87.27% (±0.38), rank 1 of 61 on the Vals AI board (updated 2026-09-03), ahead of GPT-5.6 Sol (85.77%) and Claude Fable 5.1 (85.02%) on the same harness. OpenAI publishes no 2.x figure for Astra; its Terminal-Bench 4.0 value (57.7) is a different generation and stays out of this field.
MMMU-Pro87%third party
MMMU-Pro 87% at max effort (high and xhigh: 86%), rank 1 on the Artificial Analysis board at intake (accessed 2026-09-04); the board displays whole percentages.
ARC-AGI-197.5%independent
ARC-AGI-1 public eval set 97.5% at max reasoning, ARC Prize's own run (results page, 2026-09-03). Not the semi-private set behind o3's 75.7%.
ARC-AGI-295%independent
ARC-AGI-2 public eval set 95.0% at max reasoning, ARC Prize's own run (results page, 2026-09-03); Grok 4's 15.9% was a semi-private verification. ARC Prize's semi-private testing has moved to ARC-AGI-3, where Astra scores 62.71% with the standard harness ($26,098 total) and 98.55% with the provider-adapter harness that preserves reasoning state between requests — a harness difference, not a model difference.
  • Humanity’s Last Exam (no tools): independent evaluation pending — OpenAI's launch table gives only HLE with tools (57.2; Claude Fable 5.1 65.0), which is not comparable with the no-tools convention of this field. Artificial Analysis has run the no-tools variant and reports a gain of about 6 points over GPT-5.6 Sol (47.2), but the exact figure is not in the page text at intake — fill from the AA model page once readable.
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor — OpenAI dropped SWE-bench Pro from the launch table (it had flagged ~30% of the tasks as broken at the GPT-5.6 release) and reports DeepSWE v1.1 (74.1) instead. No independent SWE-bench Pro run found at intake; check Scale's SEAL leaderboard.
  • MMLU: not reported by the vendor
  • MMLU-Pro: independent evaluation pending — Not in the launch table. Vals AI has no MMLU-Pro run for Astra at intake and now marks the benchmark saturated; Artificial Analysis dropped it from Intelligence Index v4.1.1. Together with the missing SWE-bench Pro this leaves Astra with 2 of the 4 era values, so it passes the coding-signal gate but cannot be ranked on the LLM leadership line.
  • Terminal-Bench 3.0: not reported by the vendor — OpenAI reports Terminal-Bench 4.0 (57.7), a different generation that must not be entered in the 2.x or 3.0 fields. No public Terminal-Bench 3.0 snapshot carrying Astra at intake.

API pricing

$10 per million input tokens, $50 per million output tokens (USD, provider list price, checked 2026-09-04)

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.