AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GPT-6 Sol

OpenAI·

LLMsflagshipcloud

Overview

Mid-tier GPT-6 model, released 19 days after GPT-6 Astra at half the price of GPT-5.6 Sol: $2 / $10 per million input/output tokens, cached reads at 10% of input. Generally available in the API, ChatGPT Work and Codex. 1.05M-token context, 128K max output, knowledge cutoff 2026-04-20. Launched with GPT-6 Luna ($0.10 / $0.50), the speed- and volume-tuned tier, which has no separate entry, matching the GPT-5.6 family.

Capabilities and innovations

Agentic codingComputer useTool use / function callingReasoning effort control (none to max)1.05M token context windowAPI: gpt-6-solExplicit prompt-cache breakpoints; changing effort or tools no longer breaks the cacheSol / Luna GPT-6 tiers

Benchmarks

BenchmarkScoreSource
Humanity’s Last Exam (no tools)47.9%third party
HLE 47.9% no-tools, measured by Artificial Analysis at max effort. Same source as the GPT-5.6 Sol entry (47.2).
Terminal-Bench 2.x83.15%independent
Terminal-Bench 2.1 83.15%, measured by Vals AI at max effort (accessed 2026-09-23).
Terminal-Bench 4.043.9%third party
Terminal-Bench 4.0 43.9, measured by Artificial Analysis at max effort. OpenAI publishes no 4.0 figure for Sol, and it is not yet on the official tbench.ai 4.0 board (checked 2026-09-23). 4.0 is its own generation and not comparable with the 2.x or 3.0 fields.
DeepSWE v1.168.8%vendor
DeepSWE v1.1 68.8 at max effort, OpenAI launch post. Vendor self-report over its own model.
MMMU-Pro83.3%third party
MMMU-Pro 83.3%, measured by Artificial Analysis at max effort. Same source as the GPT-5.6 Sol entry (83.4).
  • GPQA Diamond: independent evaluation pending — Not reported by OpenAI; not yet on Artificial Analysis, and Vals AI has archived GPQA (checked 2026-09-23).
  • SWE-bench Pro: not reported by the vendor — OpenAI reports DeepSWE v1.1 instead, recorded in deepswe_v1_1.
  • MMLU-Pro: not reported by the vendor
  • ARC-AGI-2: independent evaluation pending — ARC Prize lists GPT-6 Luna but not yet Sol (checked 2026-09-23).
  • Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.

API pricing

$2 per million input tokens, $10 per million output tokens (USD, provider list price, checked 2026-09-23)

Reliability

Hallucination rate (Vectara HHEM)
6.5% (lower is better)independent

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.