AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Sonnet 5

Anthropic·

LLMsmidcloud

Overview

Sonnet-tier Claude model for agentic coding, computer use and tool use, released as the new default in the Claude apps and the API.

Capabilities and innovations

Agentic codingComputer useExtended thinkingTool useVision inputAPI: claude-sonnet-5OSWorld-Verified computer useIntroductory pricing through 2026-08-31

Benchmarks

BenchmarkScoreSource
GPQA Diamond88.89%independent
GPQA Diamond 88,89%, measured by Vals AI. Artificial Analysis measures 91,11% at max effort.
Humanity’s Last Exam (no tools)43.2%vendor
Humanity's Last Exam ohne Tools (43,2%); mit Tools 57,4%.
SWE-bench Pro63.2%vendor
MMLU-Pro87.55%independent
MMLU Pro 87,55%, measured by Vals AI.
Terminal-Bench 2.x80.4%vendor
Terminal-Bench 2.1: 80,4% (Anthropic self-report).
MMMU-Pro77.3%third party
MMMU-Pro 77,3%, measured by Artificial Analysis (Adaptive Reasoning, max effort).

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.