Anthropic's current top tier, from 2026: the same $10/$50 per million as Fable 5, with cache reads at a quarter of the price — 2.5% of base input rather than 10%. On a cache-heavy prompt that is the whole difference.
Why this oneWhat it is actually for
Reach for claude-fable-5-1 when the work is demanding reasoning or long-horizon agentic — Anthropic's own guidance is to start on Opus 5 and come here when your evals on Opus 5 at higher effort still fall short. It is the slowest model in the current lineup, and that is the trade: it thinks adaptively and always on, with effort defaulting to high.
The reason to prefer it over Fable 5 is not the headline rate, which is identical. It is the cache. A pipeline that re-sends a large stable prefix — a corpus, a system prompt, a long transcript — pays $0.25 per million on the cached part here against $1.00 on Fable 5.
What it isArchitecture, lineage, training
Fable 5.1 is a closed-weight reasoning model with a 1M-token context window and 128K max output. Thinking is adaptive and always on: there is no mode that turns it off, and the manual thinking.type: enabled plus budget_tokens shape of earlier models is not accepted. You steer it with effort, which defaults to high.
It uses the tokenizer introduced with Opus 4.7, which emits roughly 30% more tokens for the same English text than the pre-4.7 models. That is a real cost factor the per-token rate alone does not show, and it applies to Fable 5 equally.
At a glanceSee it
Where Fable 5.1 differs from Fable 5 — the headline rate is identical, so the cache branch is the whole economic story.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Claude API ID | claude-fable-5-1 (a pinned snapshot; the dateless id is its own snapshot, not a moving alias) | Models overview |
| Context window | 1M tokens | Models overview |
| Max output | 128K tokens (synchronous Messages API) | Models overview |
| Reliable knowledge cutoff | Jun 2026 | Models overview |
| Training data cutoff | Jun 2026 | Models overview |
| Retirement commitment | Not sooner than 1 September 2027 | Model deprecations |
| Amazon Bedrock ID | anthropic.claude-fable-5-1 | Models overview |
| Google Cloud ID | claude-fable-5-1 | Models overview |
| Microsoft Foundry ID | claude-fable-5-1 | Models overview |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
effort | Steers how much the model thinks | Default high | Lowering it cuts output tokens, and output is 5x the input rate — this is the main cost lever on this model |
cache_control | Marks a prefix cacheable | Automatic or explicit breakpoints | The difference between $0.25 and $10.00 per million on the cached span |
max_tokens | Output ceiling | Up to 128K | A ceiling that truncates mid-answer is still billed in full |
inference_geo | Pins inference to a region | global (default) or us | us applies a 1.1x multiplier to every token category |
SamplingShaping the output distribution
Anthropic's model pages do not document temperature or top-p ranges for this model, and this page does not restate them from an earlier generation. Steer output with effort and the prompt rather than with sampling controls.
ReasoningThinking, effort and budgets
Yes, and it cannot be switched off. Thinking is adaptive and always on: the model decides how much to think and effort steers it, defaulting to high. The older manual mode — thinking.type: "enabled" with a budget_tokens figure — is deprecated on Opus 4.6 and Sonnet 4.6 and is not accepted here. Plan for output tokens accordingly: there is no cheap non-reasoning path on this model.
ToolsFunction calling and server tools
Tool use, vision and text-plus-image input are supported across the whole current lineup, this model included. The costed extras are the server-side tools: web search is billed at $10 per 1,000 searches on top of tokens, web fetch adds no charge beyond the tokens it brings back, and code execution is free when used alongside web search or web fetch.
CostPrice, caching, batching, what drives the bill
List price from Anthropic's pricing page, per million tokens: $10.00 in, $50.00 out — identical to Fable 5. The differences are underneath.
| Line | Fable 5.1 | Fable 5 |
|---|---|---|
| Base input | $10.00 | $10.00 |
| Output | $50.00 | $50.00 |
| Cache read | $0.25 (2.5% of base) | $1.00 (10% of base) |
| 5-minute cache write | $12.50 | $12.50 |
| 1-hour cache write | $20.00 | $20.00 |
| Batch | $5.00 / $25.00 | $5.00 / $25.00 |
The 1M context window is included at standard pricing — a 900K-token request is billed at the same per-token rate as a 9K one.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Claude API | Yes | claude-fable-5-1 |
| Amazon Bedrock | Yes | anthropic.claude-fable-5-1; Bedrock sets its own lifecycle dates |
| Google Cloud (Vertex) | Yes | claude-fable-5-1; global, multi-region and regional endpoints |
| Microsoft Foundry | Yes | claude-fable-5-1; follows the Claude API lifecycle |
| Claude Platform on AWS | Yes | Billed in Claude Consumption Units at $0.01 per CCU |
| Fast mode | No | Research preview, Opus 5 and Opus 4.8 only |
StrengthsWhat it is good at
- Anthropic's stated choice for demanding reasoning and long-horizon agentic work.
- Cache reads at 2.5% of base input — a four-fold saving over Fable 5 on any prompt with a large stable prefix.
- Full 1M context at standard pricing, with no long-context surcharge.
- A retirement commitment no sooner than 1 September 2027, the longest in the current lineup.
LimitsWhere it falls down
- The priciest tier on this board, and thinking cannot be turned off — there is no cheap mode.
- Slowest comparative latency in the current lineup.
- The post-4.7 tokenizer emits about 30% more tokens for the same English text, which the per-token rate does not show.
- Not available in fast mode.
Against its neighboursHow it compares
Against claude-fable-5: identical headline price, identical context, identical batch rate — and a cache read that costs a quarter as much. If your workload caches, 5.1 is strictly cheaper; if it does not, the two are the same price and 5.1 is the newer snapshot with the later knowledge cutoff.
Against claude-opus-5: half the price ($5/$25), the same 1M context and 128K output, a May 2026 cutoff against June, and Anthropic recommends starting there. Come here when Opus 5 at high effort is measurably not enough.
Getting startedThe smallest call that works
Pin the id, cache your stable prefix deliberately, and measure the cache-hit share before comparing this model to anything else — it is the only number that separates it from Fable 5. Start at the default effort: high only if you have a failure that lower effort cannot clear, because output is billed at 5x input.
SourcesWhere every claim above came from
- Anthropic pricing —
platform.claude.com/docs/en/about-claude/pricing(read 18 Sep 2026) - Anthropic models overview —
platform.claude.com/docs/en/models/overview(read 18 Sep 2026)
Price and capacity verified 2026-09-18 against https://platform.claude.com/docs/en/about-claude/pricing. first entry 2026-09-18 (primary source read today: https://platform.claude.com/docs/en/about-claude/pricing — Claude Fable 5.1 at $10 per million base input and $50 per million output, cache hits and refreshes $0.25 per million, 5m cache writes $12.50 and 1h $20, batch $5/$25. The same read confirmed Claude Fable 5 is still listed at an unchanged $10/$50 with cache hits at $1.00, so the existing row is correct and was re-stamped rather than altered.)