AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Operator

OpenAI·

LLMsspecializedcloud

Overview

Agentic system capable of executing complex multi-step tasks on a computer.

Capabilities and innovations

Computer useAutonomous browsingTask executionComputer UseAutonomous AgentsMulti-step Task Execution

Benchmarks

BenchmarkScoreSource
Humanity’s Last Exam (no tools)22.1%unsourced
SWE-bench Verified52.3%unsourced
  • GPQA Diamond: not reported by the vendor
  • SWE-bench Pro: benchmark did not exist at release
  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.