GPT-5.6 Sol
OpenAI·
LLMsflagshipcloud
Overview
Preview of OpenAI's GPT-5.6 frontier model (Sol tier), part of a Sol/Terra/Luna family in limited preview to trusted partners via Codex and the API. Sol priced at $5/M input, $30/M output tokens. Not yet generally available.
Capabilities and innovations
Frontier reasoningAgentic codingTool use / function callingComputer useCybersecurity tasksSol / Terra / Luna model familyLimited preview (Codex + API)Token-efficient agentic workflows
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 95.2% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Humanity’s Last Exam (no tools) | 47.2% | third party HLE 47,2% no-tools, measured by Artificial Analysis (max effort; xhigh gives 44,7%). |
| SWE-bench Pro | 64.6% | vendor 64.6% SWE-bench Pro. OpenAI separately reports that ~30% of SWE-bench Pro tasks are broken and advises caution on the benchmark. |
| MMLU-Pro | 89.1% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 2.x | 88.8% | vendor Terminal-Bench 2.1: 88,8% (OpenAI self-report, Sol-Tier). Sol Ultra 91,9% ist ein nicht veröffentlichtes Modell und daher nicht eingetragen. |
| Terminal-Bench 3.0 | 34.6% | third party Terminal-Bench 3.0 public snapshot (2026-08-28), leaderboard hosted by Snorkel AI. 3.0 runs on a different scale than 2.x — do not compare across generations. |
| MMMU-Pro | 83.4% | third party MMMU-Pro 83,4%, measured by Artificial Analysis (max effort). |
Links
More from OpenAI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.