AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

o3

OpenAI·

LLMsreasoningcloud

Overview

Full reasoning model with significant improvements over o1.

Capabilities and innovations

Advanced reasoningTool useAgentic codingImproved Reasoning ChainTool IntegrationMulti-step Planning

Benchmarks

BenchmarkScoreSource
GPQA Diamond82.5%unsourced
Humanity’s Last Exam (no tools)28.4%unsourced
SWE-bench Verified69.1%unsourced
MMLU91.6%unsourced
MMLU-Pro85.59%unsourced
ARC-AGI-175.7%independent
ARC-AGI-1 semi-private eval, verified by ARC Prize at the public $10k compute limit; a high-compute (172x) configuration reached 87.5%.
  • SWE-bench Pro: benchmark did not exist at release
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.