AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude 3.5 Sonnet v2

Anthropic·

LLMsmidcloud

Overview

Upgraded Sonnet with computer use capabilities.

Capabilities and innovations

Computer useEnhanced codingTool useComputer UseAgent CapabilitiesReal-world Interaction

Benchmarks

BenchmarkScoreSource
GPQA Diamond65%unsourced
Humanity’s Last Exam (no tools)11.5%unsourced
SWE-bench Verified49%unsourced
MMLU88.3%unsourced
  • SWE-bench Pro: benchmark did not exist at release
  • MMLU-Pro: benchmark did not exist at release
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.