Claude Opus 4.8
Anthropic·
Overview
Anthropic's flagship upgrade to Opus 4.7, released at the same standard price ($5/M input, $25/M output). 1M-token input context with up to 128K output tokens. Adds a Fast mode running at 2.5x output speed ($10/M input, $50/M output) alongside the standard tier. New 'dynamic workflows' tool for Claude Code, effort control in claude.ai and Cowork, and a Messages API that now accepts system entries mid-conversation. Vendor-reported benchmarks: GPQA Diamond 93.6, SWE-bench Verified 88.6, SWE-bench Pro 69.2, Terminal-Bench 2.1 74.6. Artificial Analysis Intelligence Index v4.0: 61.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.6% | vendor GPQA Diamond 93.6 — from Anthropic's launch comparison table. |
| SWE-bench Verified | 88.6% | vendor SWE-bench Verified 88.6 — Anthropic-reported (up from 87.6 on Opus 4.7). |
| SWE-bench Pro | 69.2% | vendor SWE-bench Pro 69.2 — Anthropic-reported. |
| MMLU-Pro | 89.58% | unsourced |
| Terminal-Bench 2.x | 74.6% | vendor Terminal-Bench 2.1 74.6 — Anthropic self-report. OpenAI's cross-vendor chart lists 78.9 for this model (vendor discrepancy). Version 2.1, not directly comparable to the 2.0 scores recorded for earlier models. |
- Humanity’s Last Exam (no tools): not reported by the vendor — Anthropic reports no standard HLE. Only a third-party HLE-with-tools figure (57.9% per llm-stats) exists, which is not comparable to the no-tools HLE recorded for other models here.
- MMLU: not reported by the vendor
- MMMU-Pro: not reported by the vendor
Links
More from Anthropic
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.