AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Opus 4.6

Anthropic·

LLMsflagshipcloud

Overview

Next-generation flagship model with adaptive thinking and agent teams.

Capabilities and innovations

Adaptive thinking1M context windowAgent teams128K outputAgent TeamsAdaptive ReasoningClaude Cowork

Benchmarks

BenchmarkScoreSource
GPQA Diamond89.2%third party
GPQA Diamond: 89,2% laut BenchLM GPQA-D Leaderboard. Korrigiert einen zuvor unbelegten Wert von 94,5%.
Humanity’s Last Exam (no tools)47.3%unsourced
SWE-bench Pro56.8%unsourced
MMLU95.2%unsourced
MMLU-Pro89.11%unsourced
  • SWE-bench Verified: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Reliability

Hallucination rate (Vectara HHEM)
12.2% (lower is better)independent
Agentic tool use (τ-bench)
84.8% (higher is better)third party

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.