AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude 3.7 Sonnet

Anthropic·

LLMsmidcloud

Overview

Hybrid model with extended thinking for complex reasoning tasks.

Capabilities and innovations

Extended thinkingEnhanced codingAgentic reliabilityExtended Thinking ModeHybrid ReasoningImproved Tool Use

Benchmarks

BenchmarkScoreSource
GPQA Diamond84.8%unsourced
Humanity’s Last Exam (no tools)26.7%unsourced
SWE-bench Verified56.2%unsourced
MMLU90.7%unsourced
MMLU-Pro82.73%unsourced
  • SWE-bench Pro: benchmark did not exist at release
  • Terminal-Bench 2.x: benchmark did not exist at release
  • MMMU-Pro: benchmark did not exist at release

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.