AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GPT-5.4 Pro

OpenAI·

LLMsreasoningcloud

Overview

Highest-capability model optimized for quality and depth over speed.

Capabilities and innovations

Highest capabilityDecision-ready outputsExtended reasoningNative computer use1M token context (API)ARC-AGI-2 Leader (83.3%)Quality over Speed Optimization

Benchmarks

BenchmarkScoreSource
GPQA Diamond94.4%third party
GPQA Diamond 94.4 — GPT-5.4 Pro benchmark table via AI Release Tracker, which transcribes OpenAI's reported set (BrowseComp 89.3, FrontierMath, GPQA 94.4, GDPval 82). Aggregator (lowest source tier); value corroborated by multiple secondary sources but not on independent boards (benchlm/vals.ai don't list GPT-5.4 Pro). Sits above the benchlm GPQA-D leader (Gemini 3.1 Pro 94.3) but is not contradicted by it.
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor
  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Reliability

Hallucination rate (Vectara HHEM)
8.3% (lower is better)independent
Agentic tool use (τ-bench)
80.1% (higher is better)third party

Links

More from OpenAI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.