AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Opus 4.7

Anthropic·

LLMsflagshipcloud

Overview

Anthropic's most capable generally available model. Default in Claude Code. Dense architecture with new tokenizer, file-system-based cross-session memory, new 'xhigh' effort level, and task budgets (public beta) for agentic loops. 1M context window with no long-context premium. Sits below Mythos Preview in capability, above Opus 4.6. $5/M input, $25/M output tokens.

Capabilities and innovations

1M context windowVision input up to 3.75 MPComputer use with 1:1 pixel mappingAgentic workflowsCross-session file-system memoryxhigh Effort LevelTask Budgets (Public Beta)File-System MemoryNew Tokenizer

Benchmarks

BenchmarkScoreSource
GPQA Diamond94.2%unsourced
SWE-bench Verified87.6%unsourced
SWE-bench Pro64.3%unsourced
MMLU-Pro89.87%unsourced
Terminal-Bench 2.x69.4%unsourced
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • MMLU: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Reliability

Hallucination rate (Vectara HHEM)
12% (lower is better)independent

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.