AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GPT-5.3 Codex

OpenAI·

LLMscodingcloud

Overview

Agentic coding model, 25% faster, first model instrumental in creating itself.

Capabilities and innovations

Agentic coding25% fasterSelf-developingTool useSelf-DevelopmentTerminal-Bench SOTAOSWorld Leader

Benchmarks

BenchmarkScoreSource
SWE-bench Pro56.8%unsourced
  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Verified: not reported by the vendor
  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Reliability

Agentic tool use (τ-bench)
77.8% (higher is better)third party

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.