Claude Opus 5
Anthropic·
LLMsflagshipcloud
Overview
Opus-tier Claude flagship for agentic coding, long-horizon tool use and computer use, released as the new top model in the Claude apps and the API. Priced at $5 / $25 per million input/output tokens.
Capabilities and innovations
Agentic codingComputer useExtended thinkingTool useVision inputAPI: claude-opus-5Effort control
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.43% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Humanity’s Last Exam (no tools) | 52.6% | third party HLE 52,6% no-tools, measured by Artificial Analysis (Adaptive Reasoning, max effort; xhigh gives 52,5%). Anthropic reports no HLE value for Opus 5. |
| SWE-bench Verified | 96% | vendor SWE-bench Verified 96,0% (Anthropic self-report, averaged over five trials). vals.ai independent: 97,0%. |
| SWE-bench Pro | 79.2% | vendor SWE-bench Pro 79,2% (Anthropic self-reported figure per launch coverage). |
| MMLU-Pro | 91.59% | independent vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 2.x | 84.64% | independent Terminal-Bench 2.1, vals.ai independent evaluation (accessed 2026-07-25). |
| Terminal-Bench 3.0 | 42.7% | third party Terminal-Bench 3.0 public snapshot (2026-08-28), leaderboard hosted by Snorkel AI. 3.0 runs on a different scale than 2.x — do not compare across generations. |
| MMMU-Pro | 84.7% | third party MMMU-Pro 84,7%, measured by Artificial Analysis (Adaptive Reasoning, max effort). |
API pricing
$5 per million input tokens, $25 per million output tokens (USD, provider list price, checked 2026-07-25)
Links
More from Anthropic
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.