AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GPT-5.6 Sol

OpenAI·

LLMsflagshipcloud

Overview

Preview of OpenAI's GPT-5.6 frontier model (Sol tier), part of a Sol/Terra/Luna family in limited preview to trusted partners via Codex and the API. Sol priced at $5/M input, $30/M output tokens. Not yet generally available.

Capabilities and innovations

Frontier reasoningAgentic codingTool use / function callingComputer useCybersecurity tasksSol / Terra / Luna model familyLimited preview (Codex + API)Token-efficient agentic workflows

Benchmarks

BenchmarkScoreSource
GPQA Diamond95.2%independent
vals.ai independent evaluation (accessed 2026-07-25).
Humanity’s Last Exam (no tools)47.2%third party
HLE 47,2% no-tools, measured by Artificial Analysis (max effort; xhigh gives 44,7%).
SWE-bench Pro64.6%vendor
64.6% SWE-bench Pro. OpenAI separately reports that ~30% of SWE-bench Pro tasks are broken and advises caution on the benchmark.
MMLU-Pro89.1%independent
vals.ai independent evaluation (accessed 2026-07-25).
Terminal-Bench 2.x88.8%vendor
Terminal-Bench 2.1: 88,8% (OpenAI self-report, Sol-Tier). Sol Ultra 91,9% ist ein nicht veröffentlichtes Modell und daher nicht eingetragen.
Terminal-Bench 3.034.6%third party
Terminal-Bench 3.0 public snapshot (2026-08-28), leaderboard hosted by Snorkel AI. 3.0 runs on a different scale than 2.x — do not compare across generations.
MMMU-Pro83.4%third party
MMMU-Pro 83,4%, measured by Artificial Analysis (max effort).

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.