AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude 3.5 Sonnet

Anthropic·

LLMsmidcloud

Overview

Significant leap in coding and reasoning capabilities.

Capabilities and innovations

SOTA CodingArtifacts UIFast inferenceArtifacts UIReal-time CollaborationInteractive Content Generation

Benchmarks

BenchmarkScoreSource
GPQA Diamond59.4%unsourced
Humanity’s Last Exam (no tools)12.4%unsourced
SWE-bench Verified49%unsourced
MMLU88.7%unsourced
  • SWE-bench Pro: benchmark did not exist at release
  • MMLU-Pro: benchmark did not exist at release
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.