Wait Which Model?
← Back to directory
Anthropic

Claude Fable 5.1

Released Sep 1, 2026 · knowledge cutoff 2026-06

Status
Frontier
Location
United States
Modality
Multimodal
Context window
1M
Max output
128K
Speed (medium effort)
47 tok/s · 6.77s to first answer token
Price ($/MTok in / out)
$10 / $50
Cost per task (medium effort)
$1.00
Open weights
No
Benchmarks7 of 10 reported

MMLU-Proretired

92.4%

SWE-bench Pro

81.2%

Terminal-Bench 2.1

91.4%

Humanity's Last Exam

60.9%

LMArena Elo

1498

GDPval-AA v2

1853

ARC-AGI-2

90%

GPQA Diamond

SWE-bench Verified

AIME (latest)retired

Strengths & weaknesses

Strengths

  • Does the task it was given and stops, instead of treating every ambiguity as more work to take on — the scope creep Fable 5 was known for is mostly gone
  • Code review is quieter and sharper: nitpick comments largely disappear while it still catches the same real issues
  • More comfortable left running unattended for hours than Fable 5, and fixes root causes rather than patching symptoms
  • Prose sounds less stereotypically Claude — denser, plainer, with far less bold-and-bullet scaffolding

Weaknesses

  • Slower in agent loops: it often issues one tool call per turn where Fable 5 batched several, and code reviews take about half again as long
  • At max effort it generates roughly twice the tokens of its peers and can sit for minutes before the first answer token, so per-task cost went up despite the cache-read cut
  • At low effort it answers from memory rather than searching, tends to rewrite whole files for small edits, and can lift passages from sources without marking them as quotes

Same model as the restricted Claude Mythos 5.1 (Project Glasswing participants only), which shares its specs and pricing and differs only in safeguards — Anthropic reports no separate Mythos score on any tracked benchmark. HLE 60.9% no tools / 65.0% with tools per Anthropic. Anthropic no longer reports SWE-bench Verified, GPQA Diamond or AIME: Terminal-Bench 2.1 91.4% is Artificial Analysis-measured at max effort (Vals AI's harness gives 85.0%), MMLU-Pro 92.4% is Vals AI's measurement at max effort, and ARC-AGI-2 90.0% is ARC Prize-verified at max effort ($4.49/task). Cache reads are $0.25/MTok (0.025x input, vs 0.1x on other Claude models); batch $5/$25. Cost per task and speed are AA's medium-effort variant — max effort runs $3.69/task at 66.2 tok/s with a 291 s first-answer latency. Anthropic's launch table also reports Terminal-Bench-Science 0.1 52.6%, Terminal-Bench 4.0 55.8% (Mythos 5.1 60.9%), GDPval-AA v2 1853, OSWorld 2.0 77.9% partial / 41.7% strict, AutomationBench 31.4% and CursorBench 3.2.0 73.4%, with Fable 5.1 scored with production safeguards on, taking zeros where they intervened. LMArena Elo 1498 is the "claude-fable-5.1-max" text-arena listing (arena.ai, 2026-09-13 update, rank 5). SWE-bench Pro 81.2% is Anthropic's system-card figure (not in the launch post), as relayed by BenchLM, CodingFleet and emergent.sh; the system-card PDF was too large to read directly this run.

For developers

API model strings

  • anthropicclaude-fable-5-1
  • aws-bedrockanthropic.claude-fable-5-1
  • google-vertexclaude-fable-5-1

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

Replaces Claude Fable 5

News