Claude Mythos 5.1
Anthropic·
Overview
Same underlying model as Claude Fable 5.1, offered by invitation only to Project Glasswing participants and the Cyber and Life Sciences verification programs, with the safety classifiers and fallback routing that govern Fable 5.1 lifted. Successor to Claude Mythos 5. Shares Fable 5.1's specifications and pricing ($10/M input, $50/M output, 1M context). API ID claude-mythos-5-1 on the Claude API, Bedrock, Google Cloud and Foundry; no public access. Anthropic reports Terminal-Bench 4.0 60.9 against 55.8 for Fable 5.1 — the gap is the cost of the safeguard interventions — and HLE 60.9 without tools (65.0 with tools) for both models. Independent evaluators measure Fable 5.1, not this variant, so no independent values are carried over.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Humanity’s Last Exam (no tools) | 60.9% | vendor HLE 60,9% without tools, 65,0% with tools — Anthropic reports the same figure for Fable 5.1 and Mythos 5.1 (launch coverage, 2026-09-01). No-tools figure recorded for cross-model comparability. |
- GPQA Diamond: independent evaluation pending — Independent boards (Vals AI, Artificial Analysis) evaluate the public Fable 5.1 endpoint, not the gated Mythos 5.1 variant. Not copied over: Anthropic's own Terminal-Bench 4.0 numbers (60.9 vs 55.8) show the safeguards move scores.
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: independent evaluation pending — No SWE-bench Pro figure found in launch coverage for either 5.1 model.
- MMLU: not reported by the vendor
- MMLU-Pro: independent evaluation pending — Gated model; no independent measurement of this variant.
- Terminal-Bench 2.x: independent evaluation pending — Anthropic reports Terminal-Bench 4.0 (60.9), a different generation that must not be entered in the 2.x field. No 2.x run of this variant exists.
- Terminal-Bench 3.0: independent evaluation pending — Gated model; not on the Terminal-Bench 3.0 public snapshot.
- MMMU-Pro: independent evaluation pending — Gated model; no independent measurement of this variant.
API pricing
$10 per million input tokens, $50 per million output tokens (USD, provider list price, checked 2026-09-01)
Links
More from Anthropic
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.