AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Opus 5

Anthropic·

LLMsflagshipcloud

Overview

Opus-tier Claude flagship for agentic coding, long-horizon tool use and computer use, released as the new top model in the Claude apps and the API. Priced at $5 / $25 per million input/output tokens.

Capabilities and innovations

Agentic codingComputer useExtended thinkingTool useVision inputAPI: claude-opus-5Effort control

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.43%independent
vals.ai independent evaluation (accessed 2026-07-25).
Humanity’s Last Exam (no tools)52.6%third party
HLE 52,6% no-tools, measured by Artificial Analysis (Adaptive Reasoning, max effort; xhigh gives 52,5%). Anthropic reports no HLE value for Opus 5.
SWE-bench Verified96%vendor
SWE-bench Verified 96,0% (Anthropic self-report, averaged over five trials). vals.ai independent: 97,0%.
SWE-bench Pro79.2%vendor
SWE-bench Pro 79,2% (Anthropic self-reported figure per launch coverage).
MMLU-Pro91.59%independent
vals.ai independent evaluation (accessed 2026-07-25).
Terminal-Bench 2.x84.64%independent
Terminal-Bench 2.1, vals.ai independent evaluation (accessed 2026-07-25).
Terminal-Bench 3.042.7%third party
Terminal-Bench 3.0 public snapshot (2026-08-28), leaderboard hosted by Snorkel AI. 3.0 runs on a different scale than 2.x — do not compare across generations.
MMMU-Pro84.7%third party
MMMU-Pro 84,7%, measured by Artificial Analysis (Adaptive Reasoning, max effort).

API pricing

$5 per million input tokens, $25 per million output tokens (USD, provider list price, checked 2026-07-25)

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.