Home › Frontier Models › Claude Fable 5
Model · Reference

Claude Fable 5

The hardest work

In one line

Anthropic's most capable released model: thinking is always on and cannot be disabled, and safety classifiers can decline a request with a successful HTTP 200.

Why this oneWhat it is actually for

You reach for Fable 5 when the task is genuinely at the edge of what a model can do and you are willing to pay double the Opus rate for it — long-horizon agentic runs, multi-hour autonomous work, reasoning that has to hold together across a very large context. Anthropic's own guidance is to start with Opus 5 and move up to Fable 5 only "for workloads that need the highest available capability." Two things push back. It carries safety classifiers that can decline a request outright, so your integration needs a refusal path before you ship. And it requires 30-day data retention, so a zero-data-retention organisation cannot call it at all.

What it isArchitecture, lineage, training

Fable 5 is a closed-weight model served only through APIs. Anthropic publishes no architecture at all for it: no parameter count, no statement of whether it is dense or mixture-of-experts, no layer or attention-head counts, no expert counts, no training-corpus description, no training-compute figure. There is no Hugging Face model card and no config.json, because no weights are released. Anything you read elsewhere giving Fable 5 a parameter count is not sourced from Anthropic.

What Anthropic does disclose is the serving envelope. The API model ID is claude-fable-5, a pinned snapshot rather than a moving pointer. It has a 1M-token context window "by default" and up to 128k output tokens per request. The models overview lists its reliable knowledge cutoff as January 2026 and its training data cutoff as January 2026. It became generally available on June 9, 2026.

Its lineage is visible in two mechanical details rather than in any published family tree. First, it shares a tokenizer with the Opus 4.7-and-later generation — the pricing page notes those models "use a newer tokenizer" that "produces approximately 30% more tokens for the same text." The context-window tooltip corroborates this: 1M tokens is described as roughly 555k words, against roughly 750k words per 1M tokens on Opus 4.6.

Second, it is paired with claude-mythos-5, which the docs describe as sharing "the same capabilities" and "the same specs and pricing" but without the safety classifiers, available only through the invitation-only Project Glasswing. The two are presented as the same underlying capability with different safeguards attached.

At a glanceSee it

Claude Fable 5 diagram

How a Fable 5 request flows through safety classifiers and always-on adaptive thinking to a response.

CapacityContext, output and what fits

FactValueSource
API model IDclaude-fable-5 (pinned snapshot, no date suffix)Models overview
Context window1M tokens (default and documented maximum)Models overview; Fable 5 launch page
Max output tokens128kModels overview; Fable 5 launch page
Input modalitiesText, imageModels overview ("All current Claude models support text and image input")
Output modalitiesText onlyModels overview
Reliable knowledge cutoffJanuary 2026Models overview
Training data cutoffJanuary 2026Models overview
General availabilityJune 9, 2026Fable 5 launch page
Lifecycle stateActive; tentative retirement "not sooner than June 9, 2027"Model deprecations
Weights availableNo — closed API onlyNo vendor release exists
LicenceNot applicable; commercial API terms only—
Parameter countNot disclosed—
Architecture (dense or MoE)Not disclosed—
Data retention30-day retention required; not available under zero data retentionFable 5 launch page
Minimum cacheable prefix512 tokensPrompt caching

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the modelRequired; "claude-fable-5"A wrong ID returns 404, not a fallback
messagesConversation turnsRequired; up to 100,000 messages per requestFirst message must be user
max_tokensHard ceiling on total output — thinking tokens plus response textRequired; model max 128kSet too low and a long thinking pass eats the budget, returning stop_reason: "max_tokens" with truncated or missing text. Anthropic advises a large value at high and xhigh effort
systemSystem promptString or array of text blocksSits at the front of the cache prefix; editing it invalidates everything after
output_config.effortThe primary control for intelligence, latency and cost on this modellow, medium, high, xhigh, max; default highDocs say lower settings "still perform well and often exceed xhigh performance on prior models." Changing the value between requests invalidates the prompt cache
thinkingThinking configurationOmit it, or {"type": "adaptive"}Both {"type": "enabled", "budget_tokens": N} and {"type": "disabled"} return 400. Thinking is always on and cannot be turned off
thinking.displayWhether thinking blocks carry readable text"summarized" or "omitted"; default "omitted"At the default you get thinking blocks with an empty thinking field. Billing is identical either way
temperatureSampling temperatureNon-default values rejectedAny non-default value returns 400 "on every request, regardless of whether thinking is used." Steer with prompting instead
top_pNucleus samplingNon-default values rejectedSame 400 as temperature
top_kTop-K truncationNon-default values rejectedSame 400 as temperature
stop_sequencesCustom strings that halt generationArray of stringsSupported; sets stop_reason: "stop_sequence"
tools / tool_choiceTool definitions and forcing behaviourauto (default), any, tool, noneForced tool choice works here — the restriction that blocks it applies to manual extended thinking, which this model does not have
output_config.formatConstrains output to a JSON schema{"type": "json_schema", "schema": {...}}Schema needs additionalProperties: false. No recursive schemas, no numeric or string length constraints
fallbacksServer-side retry on another model when a classifier declinesBeta, Claude API; "default" mode or a named listTurns a refusal into an answer inside one call instead of a dead response
cache_controlMarks a cache breakpoint{"type": "ephemeral", "ttl": "5m"} or "1h"; max 4 breakpointsPrefixes under 512 tokens silently do not cache and return no error
speedFast modeNot supported on this modelFast mode is documented only for Opus 5 and Opus 4.8. Use effort to trade latency here
seed, frequency_penalty, presence_penalty, logprobs, nDeterminism, repetition penalties, token probabilities, multiple completionsNot supported — no such parameters exist on the Messages APIFor repetition, prompt for it. For multiple samples, send multiple requests

Only two knobs really matter here. output_config.effort is the one Anthropic calls "the primary control for trading off intelligence, latency, and cost on Claude Fable 5" — and the counterintuitive advice is to try low and medium for routine work, because they often beat higher settings on older models. max_tokens is the second, because it caps thinking and answer together; an under-sized value on a hard task truncates the answer after the model has already spent the budget reasoning. Everything in the temperature family is gone, so output shaping happens entirely in the prompt.

SamplingShaping the output distribution

There is effectively no sampling surface on this model. temperature, top_p and top_k all return a 400 error when set to a non-default value — the thinking docs state this applies "on every request, regardless of whether thinking is used," and the deprecation page records the three parameters as deprecated from Opus 4.7 onward with the recommended replacement being to "omit and use prompting to guide model behavior." So the classic levers for making output more deterministic or more varied are simply unavailable. If you previously used temperature=0 for reproducibility, note that it never guaranteed identical outputs anyway; the API reference says "even with temperature of 0.0, the results will not be fully deterministic." The sensible starting point is therefore to send no sampling parameters at all, leave effort at its high default, and shape tone, length and variability with explicit prompt instructions. If you need several genuinely different answers, issue several requests rather than reaching for a decoding parameter that no longer exists.

ReasoningThinking, effort and budgets

Thinking is always on and cannot be turned off. The per-model configuration table lists Fable 5 as "Adaptive only," default "Always on," and rejects both "enabled" and "disabled" with a 400. You control depth through output_config.effort, not a token budget — budget_tokens does not exist on this model. Thinking tokens are billed as output tokens "even when the thinking text isn't returned to you," and they count toward max_tokens alongside the response.

The raw chain of thought is never returned. thinking.display defaults to "omitted", which gives you thinking blocks with an empty thinking field; setting "summarized" returns a readable summary produced by a different model. Billing is identical under both. In multi-turn conversations you must pass thinking blocks back exactly as received, including empty ones — modified blocks are rejected with a 400.

ToolsFunction calling and server tools

Function calling is supported with the standard tools array and tool_choice of auto, any, tool or none. Because this model has no manual extended thinking mode, the restriction that blocks forced tool choice under thinking: {"type": "enabled"} does not apply — forced tool use works. Parallel tool calls are on by default; return every tool_result in a single user message, since splitting them across messages teaches the model to stop calling tools in parallel.

Structured output uses output_config.format with a JSON schema, generally available for Claude 4.5 and later. Server-side tools listed as supported at launch include the memory tool, code execution and programmatic tool calling, plus context editing and compaction behind beta headers.

The specific failure mode to plan for is not a tool bug: it is the refusal path. A declined request returns HTTP 200 with stop_reason: "refusal", so code that reads content[0] unconditionally breaks on an empty content array.

CostPrice, caching, batching, what drives the bill

List price is $10 per million input tokens and $50 per million output tokens, matching the Fable 5 launch page and the pricing table. What actually drives the bill:

LineRate
Base input$10 / MTok
Output (includes thinking tokens)$50 / MTok
5-minute cache write$12.50 / MTok (1.25x input)
1-hour cache write$20 / MTok (2x input)
Cache hit / refresh$1 / MTok (0.1x input)
Batch API$5 in / $25 out (50% off)

Three factors matter beyond the sticker. The tokenizer is the big one: Claude 4.7-and-later models "produce approximately 30% more tokens for the same text," so a per-token comparison against an older-tokenizer model understates real cost by roughly a third. Second, the full 1M context is billed at standard rates with no long-context premium — a 900k-token request costs the same per token as a 9k one. Third, thinking is always on and always billed as output, and you cannot switch it off to save money; lowering effort is the only lever. Requests refused before any output is generated are not billed at all.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Claude API (first-party)YesModel ID claude-fable-5
Amazon BedrockYesModel ID anthropic.claude-fable-5, via the Messages-API Bedrock endpoint
Claude Platform on AWSYesAnthropic-operated; bare model ID
Google Cloud / Vertex AIYesModel ID claude-fable-5
Microsoft FoundryYesListed on the launch page as GA from June 9, 2026
Hugging Face InferenceNoClosed weights; no model card exists
Self-hostingNoWeights are not released
Zero-data-retention orgsNoRequires 30-day retention; designated a Covered Model

StrengthsWhat it is good at

  • Anthropic's own docs name it the model to use "for workloads that need the highest available capability," above its recommended default of Opus 5.
  • 1M-token context at standard per-token pricing — no long-context surcharge, and prompt caching and batch discounts apply across the full window.
  • Thinking is always on with no configuration, and depth is steered by a single effort value rather than a token budget you have to tune.
  • Documented refusal recovery: the fallbacks parameter retries on another model server-side, and fallback credit refunds the prompt-cache cost of switching.
  • Lower effort levels are documented to "often exceed xhigh performance on prior models," so the cheap settings are genuinely usable.

LimitsWhere it falls down

  • Safety classifiers can decline a request and return HTTP 200 with stop_reason: "refusal" and an empty content array — this breaks naive response parsing and has no equivalent on Opus-tier models.
  • Unavailable to zero-data-retention organisations; every request returns a 400 if the org's retention configuration does not meet the 30-day requirement.
  • Thinking cannot be disabled, so you cannot buy latency by turning reasoning off — and thinking tokens bill at the $50/MTok output rate.
  • The newer tokenizer emits roughly 30% more tokens for the same text, so effective cost is meaningfully above the headline gap with Opus 5.
  • Fast mode is not available, and the pricing page's tool-use system-prompt token table does not list Fable 5 at all, so that overhead is undocumented.

Against its neighboursHow it compares

Against Claude Opus 5, the honest framing is that Opus 5 is the default and Fable 5 is the escalation. Opus 5 costs exactly half ($5/$25 versus $10/$50), has a later reliable knowledge cutoff (May 2026 versus January 2026), supports fast mode, has a lower 512-token cache minimum shared with Fable, and lets you disable thinking at high effort or below. Fable 5 gives you the higher capability tier, and takes back the ability to turn thinking off. Against Claude Opus 4.8, Fable 5 is twice the price and, unlike 4.8, runs thinking by default rather than requiring you to opt in. The deciding question is rarely benchmark-shaped: it is whether your workload can tolerate a refusal path and 30-day retention, and whether you have measured a real quality gain at double the rate.

Getting startedThe smallest call that works

code
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-fable-5",
    "max_tokens": 4096,
    "messages": [
      {"role": "user", "content": "Explain what a write-ahead log is."}
    ]
  }'

Change two things first. Add "output_config": {"effort": "low"} and measure — on this model the cheap settings are unusually strong, and effort is your main cost lever. Then, before shipping, branch on stop_reason before reading content: a classifier refusal arrives as a successful 200 with an empty content array, and unguarded indexing into content[0] will throw.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-18 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-18 (primary source read today: https://platform.claude.com/docs/en/about-claude/pricing — Claude Fable 5 still listed at $10 per million base input and $50 per million output, UNCHANGED, with cache hits at $1.00 per million. Claude Fable 5.1 now sits above it and is its own row. Closes the 2026-09-03 id-drift reading.)

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Fable 5.1 has replaced Fable 5 as Anthropic's current top-tier release, so this page's flagship claim is stale.

    Update this page so it no longer claims Claude Fable 5 is Anthropic's most capable released model; Fable 5.1 is now the current top tier and the page should describe Fable 5 as the superseded generation.

    Anthropic · 4 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning