AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Claude Fable 5.1

Anthropic·

LLMsflagshipcloud

Overview

Successor to Claude Fable 5 in the Mythos-class tier, at the same $10/M input and $50/M output pricing but with cache reads cut to $0.25/M (2.5% of the input price instead of 10%). Launched alongside Claude Mythos 5.1, the same underlying model with safeguards lifted for Project Glasswing participants (tracked separately). Safety classifiers remain in place; refused requests can fall back server-side to Claude Opus 4.8 or Claude Opus 5. Adaptive thinking is always on, forced tool use is no longer supported, and all text output carries Anthropic's statistical watermark. Vendor-reported benchmarks: Terminal-Bench 4.0 55.8 (Fable 5: 42.0, Opus 5: 52.3), Terminal-Bench-Science 0.1 52.6, HLE 60.9 without tools (65.0 with tools), OSWorld 2.0 41.7 strict, CursorBench 3.2.0 73.4, AutomationBench 31.4.

Capabilities and innovations

Autonomous long-horizon tasksAgentic codingMultistep researchComputer useVision inputTool useAPI: claude-fable-5-1Cache reads at 2.5% of input priceContent provenance (text watermark, C2PA for files)Per-message effort control

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.43%independent
GPQA Diamond 93,43% (±1,86), measured by Vals AI at max effort (accessed 2026-09-01). Ties Claude Opus 5 on the same board.
Humanity’s Last Exam (no tools)60.9%vendor
HLE 60,9% without tools, 65,0% with tools — Anthropic launch announcement as reported in launch coverage (Decrypt, MarkTechPost, 2026-09-01). No-tools figure recorded for cross-model comparability. Artificial Analysis measures 59,1% at max effort with default fallback (58,7% xhigh, 55,9% high).
MMLU-Pro92.38%independent
MMLU Pro 92,38% (±0,27), rank 1 on the Vals AI MMLU-Pro leaderboard, max effort (accessed 2026-09-01). Not reported in Anthropic's launch material.
Terminal-Bench 2.x85.02%independent
Terminal-Bench 2.1 85,02% (±0,99), Vals AI at high effort, rank 2 behind GPT-5.6 Sol (85,77%), updated 2026-08-31. Anthropic itself reports Terminal-Bench 4.0 (55,8%), a different generation on a different scale — not comparable with 2.x values.
MMMU-Pro90.64%independent
MMMU Pro 90,64%, rank 1 on the Vals AI board (accessed 2026-09-01).
  • SWE-bench Verified: not reported by the vendor — Anthropic reports no SWE-bench Verified figure for this release.
  • SWE-bench Pro: independent evaluation pending — No SWE-bench Pro figure found in launch coverage (Anthropic's headline coding number for this release is Terminal-Bench 4.0). Check the launch table and Scale's SEAL leaderboard once an independent run exists.
  • MMLU: not reported by the vendor
  • Terminal-Bench 3.0: independent evaluation pending — Terminal-Bench 3.0 public snapshot (2026-08-28) predates the release; no 3.0 value for Fable 5.1 yet. Anthropic reports Terminal-Bench 4.0 (55.8), a different generation that must not be entered in the 2.x or 3.0 fields.

API pricing

$10 per million input tokens, $50 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-01)

Links

More from Anthropic

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.