AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

OpenAI o1

OpenAI·

LLMsreasoningcloud

Overview

Reasoning model designed to spend more time thinking before responding.

Capabilities and innovations

Chain of thoughtComplex math/codingSelf-correctionChain of Thought ReasoningTest-time Compute ScalingSelf-correction via RL

Benchmarks

BenchmarkScoreSource
GPQA Diamond78%unsourced
Humanity’s Last Exam (no tools)15.2%unsourced
SWE-bench Verified48.9%unsourced
MMLU91.8%unsourced
MMLU-Pro83.49%unsourced
  • SWE-bench Pro: benchmark did not exist at release
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.