AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GPT-5.5

OpenAI·

LLMsflagshipcloud

Overview

Agentic-focused flagship with 1M context window (400K in Codex). Thinking and Pro variants available. Codename 'Spud'.

Capabilities and innovations

1M context windowThinking modePro variant ($30/$180 per M tokens)400K context in CodexAgentic workflowsAPI pricing: $5/$30 per M tokens (base)1M Context Window (largest OpenAI flagship)Codex-Optimized 400K ContextThinking + Pro Variant Lineup

Benchmarks

BenchmarkScoreSource
SWE-bench Pro58.6%vendor
MMLU-Pro88.14%unsourced
Terminal-Bench 2.x83.4%vendor
Terminal-Bench 2.1: 83,4% (OpenAI self-report).
  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Reliability

Hallucination rate (Vectara HHEM)
9.3% (lower is better)independent

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.