AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

ChatGPT 5.1

OpenAI·

LLMsflagshipcloud

Overview

Major update with "Instant" and "Thinking" modes for adaptive reasoning.

Capabilities and innovations

Adaptive reasoningInstant/Thinking modesPersonalizationNode-based Agent Builder

Benchmarks

BenchmarkScoreSource
GPQA Diamond86.6%unsourced
Humanity’s Last Exam (no tools)35.4%unsourced
SWE-bench Verified67.8%unsourced
MMLU92.3%unsourced
  • SWE-bench Pro: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Reliability

Hallucination rate (Vectara HHEM)
12.1% (lower is better)independent
Agentic tool use (τ-bench)
74.2% (higher is better)third party

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.