AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Opus 5.5

Anthropic·

LLMsflagshipcloud

Overview

Opus-tier successor to Claude Opus 5, released two months later and ten days after Anthropic's "Pacing the Frontier" proposal. Priced below its predecessor at $4 / $20 per million input/output tokens (cache reads $0.20), with output more than 30% faster than Opus 5. 1M-token context, 128K max output. Thinking can no longer be switched off; default effort is medium. Ships with the same cyber, bio and distillation safeguards as Fable 5.1; Life Sciences Verification Program access at launch, Cyber Verification Program access to follow. Anthropic reports no GPQA or MMLU-Pro figure.

Capabilities and innovations

Agentic codingComputer useExtended thinking (always on)Tool useVision input1M token context windowAPI: claude-opus-5-5Effort control (default medium)Fast mode ($8 / $40)

Benchmarks

BenchmarkScoreSource
Humanity’s Last Exam (no tools)61.4%third party
HLE 61.4% no-tools, measured by Artificial Analysis at max effort (with fallback). Same source as the Opus 5 entry (52.6). Anthropic reports 64.4 no-tools and 67.7 with tools in the Opus 5.5 system card (Table 8.1.A).
SWE-bench Pro89.9%vendor
SWE-bench Pro 89.9%, Anthropic system card, max effort (Opus 5: 79.2). No independent run yet.
Terminal-Bench 2.x87.64%independent
Terminal-Bench 2.1 87.64%, measured by Vals AI at high effort, rank 1, counting tasks completed via Anthropic's server-side safeguard fallback as successes (79.77 when counted as failures) (accessed 2026-09-23).
Terminal-Bench 4.066.4%vendor
Terminal-Bench 4.0 66.4, Anthropic launch table, xhigh effort. Not yet on the official tbench.ai 4.0 board (checked 2026-09-23). Vals AI measures 61.62 (max), Artificial Analysis 59.6. 4.0 is its own generation and not comparable with the 2.x or 3.0 fields.
DeepSWE v1.174.2%vendor
DeepSWE v1.1 74.2, Anthropic system card. Vendor self-report over its own model.
MMMU-Pro87.7%third party
MMMU-Pro 87.7%, measured by Artificial Analysis at max effort. Same source as the Opus 5 entry (84.7).
ARC-AGI-198.5%independent
ARC-AGI-1 98.5%, verified by ARC Prize (High effort, the best of five variants; Max and XHigh 97.5) (accessed 2026-09-23).
ARC-AGI-293.3%independent
ARC-AGI-2 93.3%, verified by ARC Prize (High effort; XHigh 92.5, Max 91.7). No ARC-AGI-3 result yet (accessed 2026-09-23).
  • GPQA Diamond: independent evaluation pending — Anthropic reports no GPQA value; Vals AI has archived GPQA as saturated and Artificial Analysis lists it as null for this model (checked 2026-09-23).
  • MMLU-Pro: independent evaluation pending — Not in the launch table or system card; no Vals AI run yet (checked 2026-09-23).
  • Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.

API pricing

$4 per million input tokens, $20 per million output tokens (USD, provider list price, checked 2026-09-23)

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.