AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Fable 5

Anthropic·

LLMsflagshipcloud

Overview

Anthropic's first publicly available Mythos-class model, priced at $10/M input and $50/M output. Launched alongside the gated Claude Mythos 5 (same base model with safeguards lifted for a small set of cyberdefenders and infrastructure providers), which is not publicly available and is not tracked here. A safety system routes a minority of sessions (~5% on average) to Claude Opus 4.8. Vendor-reported benchmarks: SWE-bench Pro 80.3, HLE 59.0 (no tools; 64.5 with tools), Terminal-Bench 2.1 88.0.

Capabilities and innovations

Autonomous long-horizon tasksVision inputAgentic workflowsSoftware engineeringFirst publicly available Mythos-class modelSafety routing to Claude Opus 4.8

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.18%independent
GPQA Diamond 93,18%, measured by Vals AI. Anthropic reports no GPQA value in the launch post; Artificial Analysis measures 92,63% at max effort.
Humanity’s Last Exam (no tools)59%vendor
HLE 59.0 without tools, 64.5 with tools. No-tools figure recorded for cross-model comparability.
SWE-bench Pro80.3%vendor
SWE-bench Pro 80.3 — Anthropic launch announcement.
MMLU-Pro91.5%third party
MMLU Pro 91.5 — rank 1 on the vals.ai MMLU-Pro leaderboard. Not reported in Anthropic's launch table.
Terminal-Bench 2.x88%vendor
Terminal-Bench 2.1 88.0 — Anthropic self-report. OpenAI's cross-vendor chart lists 84.3 for this model (vendor discrepancy). Version 2.1, not directly comparable to 2.0 scores recorded for earlier models.
Terminal-Bench 3.034%third party
Terminal-Bench 3.0 public snapshot (2026-08-28), leaderboard hosted by Snorkel AI. 3.0 runs on a different scale than 2.x — do not compare across generations.
  • SWE-bench Verified: not reported by the vendor — Anthropic reports SWE-bench Pro, not Verified, for this release.
  • MMLU: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Reliability

Agentic tool use (τ-bench)
89.2% (higher is better)third party

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.