Home › Frontier Models › Claude Fable 5.1
Model · Reference

Claude Fable 5.1

The hardest work, on a cache-heavy prompt

In one line

Anthropic's current top tier, from 2026: the same $10/$50 per million as Fable 5, with cache reads at a quarter of the price — 2.5% of base input rather than 10%. On a cache-heavy prompt that is the whole difference.

Why this oneWhat it is actually for

Reach for claude-fable-5-1 when the work is demanding reasoning or long-horizon agentic — Anthropic's own guidance is to start on Opus 5 and come here when your evals on Opus 5 at higher effort still fall short. It is the slowest model in the current lineup, and that is the trade: it thinks adaptively and always on, with effort defaulting to high.

The reason to prefer it over Fable 5 is not the headline rate, which is identical. It is the cache. A pipeline that re-sends a large stable prefix — a corpus, a system prompt, a long transcript — pays $0.25 per million on the cached part here against $1.00 on Fable 5.

What it isArchitecture, lineage, training

Fable 5.1 is a closed-weight reasoning model with a 1M-token context window and 128K max output. Thinking is adaptive and always on: there is no mode that turns it off, and the manual thinking.type: enabled plus budget_tokens shape of earlier models is not accepted. You steer it with effort, which defaults to high.

It uses the tokenizer introduced with Opus 4.7, which emits roughly 30% more tokens for the same English text than the pre-4.7 models. That is a real cost factor the per-token rate alone does not show, and it applies to Fable 5 equally.

At a glanceSee it

Claude Fable 5.1 diagram

Where Fable 5.1 differs from Fable 5 — the headline rate is identical, so the cache branch is the whole economic story.

CapacityContext, output and what fits

FactValueSource
Claude API IDclaude-fable-5-1 (a pinned snapshot; the dateless id is its own snapshot, not a moving alias)Models overview
Context window1M tokensModels overview
Max output128K tokens (synchronous Messages API)Models overview
Reliable knowledge cutoffJun 2026Models overview
Training data cutoffJun 2026Models overview
Retirement commitmentNot sooner than 1 September 2027Model deprecations
Amazon Bedrock IDanthropic.claude-fable-5-1Models overview
Google Cloud IDclaude-fable-5-1Models overview
Microsoft Foundry IDclaude-fable-5-1Models overview

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
effortSteers how much the model thinksDefault highLowering it cuts output tokens, and output is 5x the input rate — this is the main cost lever on this model
cache_controlMarks a prefix cacheableAutomatic or explicit breakpointsThe difference between $0.25 and $10.00 per million on the cached span
max_tokensOutput ceilingUp to 128KA ceiling that truncates mid-answer is still billed in full
inference_geoPins inference to a regionglobal (default) or usus applies a 1.1x multiplier to every token category

SamplingShaping the output distribution

Anthropic's model pages do not document temperature or top-p ranges for this model, and this page does not restate them from an earlier generation. Steer output with effort and the prompt rather than with sampling controls.

ReasoningThinking, effort and budgets

Yes, and it cannot be switched off. Thinking is adaptive and always on: the model decides how much to think and effort steers it, defaulting to high. The older manual mode — thinking.type: "enabled" with a budget_tokens figure — is deprecated on Opus 4.6 and Sonnet 4.6 and is not accepted here. Plan for output tokens accordingly: there is no cheap non-reasoning path on this model.

ToolsFunction calling and server tools

Tool use, vision and text-plus-image input are supported across the whole current lineup, this model included. The costed extras are the server-side tools: web search is billed at $10 per 1,000 searches on top of tokens, web fetch adds no charge beyond the tokens it brings back, and code execution is free when used alongside web search or web fetch.

CostPrice, caching, batching, what drives the bill

List price from Anthropic's pricing page, per million tokens: $10.00 in, $50.00 out — identical to Fable 5. The differences are underneath.

LineFable 5.1Fable 5
Base input$10.00$10.00
Output$50.00$50.00
Cache read$0.25 (2.5% of base)$1.00 (10% of base)
5-minute cache write$12.50$12.50
1-hour cache write$20.00$20.00
Batch$5.00 / $25.00$5.00 / $25.00

The 1M context window is included at standard pricing — a 900K-token request is billed at the same per-token rate as a 9K one.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Claude APIYesclaude-fable-5-1
Amazon BedrockYesanthropic.claude-fable-5-1; Bedrock sets its own lifecycle dates
Google Cloud (Vertex)Yesclaude-fable-5-1; global, multi-region and regional endpoints
Microsoft FoundryYesclaude-fable-5-1; follows the Claude API lifecycle
Claude Platform on AWSYesBilled in Claude Consumption Units at $0.01 per CCU
Fast modeNoResearch preview, Opus 5 and Opus 4.8 only

StrengthsWhat it is good at

  • Anthropic's stated choice for demanding reasoning and long-horizon agentic work.
  • Cache reads at 2.5% of base input — a four-fold saving over Fable 5 on any prompt with a large stable prefix.
  • Full 1M context at standard pricing, with no long-context surcharge.
  • A retirement commitment no sooner than 1 September 2027, the longest in the current lineup.

LimitsWhere it falls down

  • The priciest tier on this board, and thinking cannot be turned off — there is no cheap mode.
  • Slowest comparative latency in the current lineup.
  • The post-4.7 tokenizer emits about 30% more tokens for the same English text, which the per-token rate does not show.
  • Not available in fast mode.

Against its neighboursHow it compares

Against claude-fable-5: identical headline price, identical context, identical batch rate — and a cache read that costs a quarter as much. If your workload caches, 5.1 is strictly cheaper; if it does not, the two are the same price and 5.1 is the newer snapshot with the later knowledge cutoff.

Against claude-opus-5: half the price ($5/$25), the same 1M context and 128K output, a May 2026 cutoff against June, and Anthropic recommends starting there. Come here when Opus 5 at high effort is measurably not enough.

Getting startedThe smallest call that works

Pin the id, cache your stable prefix deliberately, and measure the cache-hit share before comparing this model to anything else — it is the only number that separates it from Fable 5. Start at the default effort: high only if you have a failure that lower effort cannot clear, because output is billed at 5x input.

SourcesWhere every claim above came from

  • Anthropic pricing — platform.claude.com/docs/en/about-claude/pricing (read 18 Sep 2026)
  • Anthropic models overview — platform.claude.com/docs/en/models/overview (read 18 Sep 2026)
Checked

Price and capacity verified 2026-09-18 against https://platform.claude.com/docs/en/about-claude/pricing. first entry 2026-09-18 (primary source read today: https://platform.claude.com/docs/en/about-claude/pricing — Claude Fable 5.1 at $10 per million base input and $50 per million output, cache hits and refreshes $0.25 per million, 5m cache writes $12.50 and 1h $20, batch $5/$25. The same read confirmed Claude Fable 5 is still listed at an unchanged $10/$50 with cache hits at $1.00, so the existing row is correct and was re-stamped rather than altered.)

A living map of modern AI — kept current every morning