Claude Sonnet 5.5
Anthropic·
Overview
Sonnet-tier successor to Claude Sonnet 5 and the second Claude 5.5 model, released six days after Opus 5.5 at the unchanged Sonnet price of $2 / $10 per million input/output tokens (cache reads $0.20, cache writes $2.50). Output more than 30% faster than Sonnet 5. 1M-token context, 128K max output. Adaptive thinking on by default with five effort levels; Claude Code and the Claude apps default to medium, the Claude Platform to high. Cyber safeguards at the Opus 5.5 level and new reasoning-extraction classifiers; Cyber and Life Sciences Verification Program access. Anthropic reports no GPQA, MMLU-Pro or MMMU-Pro figure.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Humanity’s Last Exam (no tools) | 56.9% | vendor HLE 56.9% no-tools, Anthropic system card Table 8.1.A, max effort, 5 trials (Sonnet 5: 43.1; with tools 64.5). No Artificial Analysis HLE value yet, which is the source on the Opus 5.5 entry (checked 2026-10-01). |
| SWE-bench Pro | 81.3% | vendor SWE-bench Pro 81.3%, Anthropic system card Table 8.1.A, max effort (Sonnet 5: 63.2, Opus 5.5: 89.9). No independent run yet. |
| Terminal-Bench 4.0 | 64.14% | independent Terminal-Bench 4.0 64.14% (± 1.01), measured by Vals AI, rank 2 of 42, with Claude Sonnet 5 as server-side refusal fallback (accessed 2026-10-01). Not yet on the official tbench.ai 4.0 board. 4.0 is its own generation and not comparable with the 2.x or 3.0 fields. |
| DeepSWE v1.1 | 71% | vendor DeepSWE v1.1 71.0%, Anthropic system card section 8.3, average over five trials. Vendor self-report over its own model. |
- GPQA Diamond: independent evaluation pending — Anthropic reports no GPQA value; no Vals AI or Artificial Analysis figure yet (checked 2026-10-01).
- MMLU-Pro: independent evaluation pending — Not in the launch table or system card; no Vals AI run yet (checked 2026-10-01).
- MMMU-Pro: independent evaluation pending — Not in the launch table or system card; Artificial Analysis lists MMMU-Pro for the model but publishes no value yet (checked 2026-10-01).
- Terminal-Bench 2.x: independent evaluation pending — Vals AI shows Sonnet 5.5 on the Terminal-Bench 2.1 medium split only (86.06), no overall score yet (checked 2026-10-01). Anthropic reports no 2.x figure.
- Terminal-Bench 3.0: not reported by the vendor — Not in the launch table; 4.0 is the current release.
- ARC-AGI-1: independent evaluation pending — No ARC Prize result yet (checked 2026-10-01).
- ARC-AGI-2: independent evaluation pending — No ARC Prize result yet (checked 2026-10-01).
API pricing
$2 per million input tokens, $10 per million output tokens (USD, provider list price, checked 2026-10-01)
Links
More from Anthropic
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.