AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Opus 4.8

Anthropic·

LLMsflagshipcloud

Overview

Anthropic's flagship upgrade to Opus 4.7, released at the same standard price ($5/M input, $25/M output). 1M-token input context with up to 128K output tokens. Adds a Fast mode running at 2.5x output speed ($10/M input, $50/M output) alongside the standard tier. New 'dynamic workflows' tool for Claude Code, effort control in claude.ai and Cowork, and a Messages API that now accepts system entries mid-conversation. Vendor-reported benchmarks: GPQA Diamond 93.6, SWE-bench Verified 88.6, SWE-bench Pro 69.2, Terminal-Bench 2.1 74.6. Artificial Analysis Intelligence Index v4.0: 61.

Capabilities and innovations

1M-token input context windowUp to 128K output tokensVision inputComputer useAgentic workflowsDynamic Workflows (Claude Code)Effort Control (claude.ai & Cowork)Mid-Conversation System Entries (Messages API)Fast Mode (2.5x speed tier)

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.6%vendor
GPQA Diamond 93.6 — from Anthropic's launch comparison table.
SWE-bench Verified88.6%vendor
SWE-bench Verified 88.6 — Anthropic-reported (up from 87.6 on Opus 4.7).
SWE-bench Pro69.2%vendor
SWE-bench Pro 69.2 — Anthropic-reported.
MMLU-Pro89.58%unsourced
Terminal-Bench 2.x74.6%vendor
Terminal-Bench 2.1 74.6 — Anthropic self-report. OpenAI's cross-vendor chart lists 78.9 for this model (vendor discrepancy). Version 2.1, not directly comparable to the 2.0 scores recorded for earlier models.
  • Humanity’s Last Exam (no tools): not reported by the vendor — Anthropic reports no standard HLE. Only a third-party HLE-with-tools figure (57.9% per llm-stats) exists, which is not comparable to the no-tools HLE recorded for other models here.
  • MMLU: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.