Claude Sonnet 5
Anthropic·
LLMsmidcloud
Overview
Sonnet-tier Claude model for agentic coding, computer use and tool use, released as the new default in the Claude apps and the API.
Capabilities and innovations
Agentic codingComputer useExtended thinkingTool useVision inputAPI: claude-sonnet-5OSWorld-Verified computer useIntroductory pricing through 2026-08-31
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 88.89% | independent GPQA Diamond 88,89%, measured by Vals AI. Artificial Analysis measures 91,11% at max effort. |
| Humanity’s Last Exam (no tools) | 43.2% | vendor Humanity's Last Exam ohne Tools (43,2%); mit Tools 57,4%. |
| SWE-bench Pro | 63.2% | vendor |
| MMLU-Pro | 87.55% | independent MMLU Pro 87,55%, measured by Vals AI. |
| Terminal-Bench 2.x | 80.4% | vendor Terminal-Bench 2.1: 80,4% (Anthropic self-report). |
| MMMU-Pro | 77.3% | third party MMMU-Pro 77,3%, measured by Artificial Analysis (Adaptive Reasoning, max effort). |
Links
More from Anthropic
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.