Claude Opus 5.5
Anthropic·
Overview
Opus-tier successor to Claude Opus 5, released two months later and ten days after Anthropic's "Pacing the Frontier" proposal. Priced below its predecessor at $4 / $20 per million input/output tokens (cache reads $0.20), with output more than 30% faster than Opus 5. 1M-token context, 128K max output. Thinking can no longer be switched off; default effort is medium. Ships with the same cyber, bio and distillation safeguards as Fable 5.1; Life Sciences Verification Program access at launch, Cyber Verification Program access to follow. Anthropic reports no GPQA or MMLU-Pro figure.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Humanity’s Last Exam (no tools) | 61.4% | third party HLE 61.4% no-tools, measured by Artificial Analysis at max effort (with fallback). Same source as the Opus 5 entry (52.6). Anthropic reports 64.4 no-tools and 67.7 with tools in the Opus 5.5 system card (Table 8.1.A). |
| SWE-bench Pro | 89.9% | vendor SWE-bench Pro 89.9%, Anthropic system card, max effort (Opus 5: 79.2). No independent run yet. |
| Terminal-Bench 2.x | 87.64% | independent Terminal-Bench 2.1 87.64%, measured by Vals AI at high effort, rank 1, counting tasks completed via Anthropic's server-side safeguard fallback as successes (79.77 when counted as failures) (accessed 2026-09-23). |
| Terminal-Bench 4.0 | 66.4% | vendor Terminal-Bench 4.0 66.4, Anthropic launch table, xhigh effort. Not yet on the official tbench.ai 4.0 board (checked 2026-09-23). Vals AI measures 61.62 (max), Artificial Analysis 59.6. 4.0 is its own generation and not comparable with the 2.x or 3.0 fields. |
| DeepSWE v1.1 | 74.2% | vendor DeepSWE v1.1 74.2, Anthropic system card. Vendor self-report over its own model. |
| MMMU-Pro | 87.7% | third party MMMU-Pro 87.7%, measured by Artificial Analysis at max effort. Same source as the Opus 5 entry (84.7). |
| ARC-AGI-1 | 98.5% | independent ARC-AGI-1 98.5%, verified by ARC Prize (High effort, the best of five variants; Max and XHigh 97.5) (accessed 2026-09-23). |
| ARC-AGI-2 | 93.3% | independent ARC-AGI-2 93.3%, verified by ARC Prize (High effort; XHigh 92.5, Max 91.7). No ARC-AGI-3 result yet (accessed 2026-09-23). |
- GPQA Diamond: independent evaluation pending — Anthropic reports no GPQA value; Vals AI has archived GPQA as saturated and Artificial Analysis lists it as null for this model (checked 2026-09-23).
- MMLU-Pro: independent evaluation pending — Not in the launch table or system card; no Vals AI run yet (checked 2026-09-23).
- Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.
API pricing
$4 per million input tokens, $20 per million output tokens (USD, provider list price, checked 2026-09-23)
Links
More from Anthropic
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.