Anthropic's most capable released model: thinking is always on and cannot be disabled, and safety classifiers can decline a request with a successful HTTP 200.
Why this oneWhat it is actually for
You reach for Fable 5 when the task is genuinely at the edge of what a model can do and you are willing to pay double the Opus rate for it — long-horizon agentic runs, multi-hour autonomous work, reasoning that has to hold together across a very large context. Anthropic's own guidance is to start with Opus 5 and move up to Fable 5 only "for workloads that need the highest available capability." Two things push back. It carries safety classifiers that can decline a request outright, so your integration needs a refusal path before you ship. And it requires 30-day data retention, so a zero-data-retention organisation cannot call it at all.
What it isArchitecture, lineage, training
Fable 5 is a closed-weight model served only through APIs. Anthropic publishes no architecture at all for it: no parameter count, no statement of whether it is dense or mixture-of-experts, no layer or attention-head counts, no expert counts, no training-corpus description, no training-compute figure. There is no Hugging Face model card and no config.json, because no weights are released. Anything you read elsewhere giving Fable 5 a parameter count is not sourced from Anthropic.
What Anthropic does disclose is the serving envelope. The API model ID is claude-fable-5, a pinned snapshot rather than a moving pointer. It has a 1M-token context window "by default" and up to 128k output tokens per request. The models overview lists its reliable knowledge cutoff as January 2026 and its training data cutoff as January 2026. It became generally available on June 9, 2026.
Its lineage is visible in two mechanical details rather than in any published family tree. First, it shares a tokenizer with the Opus 4.7-and-later generation — the pricing page notes those models "use a newer tokenizer" that "produces approximately 30% more tokens for the same text." The context-window tooltip corroborates this: 1M tokens is described as roughly 555k words, against roughly 750k words per 1M tokens on Opus 4.6.
Second, it is paired with claude-mythos-5, which the docs describe as sharing "the same capabilities" and "the same specs and pricing" but without the safety classifiers, available only through the invitation-only Project Glasswing. The two are presented as the same underlying capability with different safeguards attached.
At a glanceSee it
How a Fable 5 request flows through safety classifiers and always-on adaptive thinking to a response.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model ID | claude-fable-5 (pinned snapshot, no date suffix) | Models overview |
| Context window | 1M tokens (default and documented maximum) | Models overview; Fable 5 launch page |
| Max output tokens | 128k | Models overview; Fable 5 launch page |
| Input modalities | Text, image | Models overview ("All current Claude models support text and image input") |
| Output modalities | Text only | Models overview |
| Reliable knowledge cutoff | January 2026 | Models overview |
| Training data cutoff | January 2026 | Models overview |
| General availability | June 9, 2026 | Fable 5 launch page |
| Lifecycle state | Active; tentative retirement "not sooner than June 9, 2027" | Model deprecations |
| Weights available | No — closed API only | No vendor release exists |
| Licence | Not applicable; commercial API terms only | — |
| Parameter count | Not disclosed | — |
| Architecture (dense or MoE) | Not disclosed | — |
| Data retention | 30-day retention required; not available under zero data retention | Fable 5 launch page |
| Minimum cacheable prefix | 512 tokens | Prompt caching |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | Required; "claude-fable-5" | A wrong ID returns 404, not a fallback |
messages | Conversation turns | Required; up to 100,000 messages per request | First message must be user |
max_tokens | Hard ceiling on total output — thinking tokens plus response text | Required; model max 128k | Set too low and a long thinking pass eats the budget, returning stop_reason: "max_tokens" with truncated or missing text. Anthropic advises a large value at high and xhigh effort |
system | System prompt | String or array of text blocks | Sits at the front of the cache prefix; editing it invalidates everything after |
output_config.effort | The primary control for intelligence, latency and cost on this model | low, medium, high, xhigh, max; default high | Docs say lower settings "still perform well and often exceed xhigh performance on prior models." Changing the value between requests invalidates the prompt cache |
thinking | Thinking configuration | Omit it, or {"type": "adaptive"} | Both {"type": "enabled", "budget_tokens": N} and {"type": "disabled"} return 400. Thinking is always on and cannot be turned off |
thinking.display | Whether thinking blocks carry readable text | "summarized" or "omitted"; default "omitted" | At the default you get thinking blocks with an empty thinking field. Billing is identical either way |
temperature | Sampling temperature | Non-default values rejected | Any non-default value returns 400 "on every request, regardless of whether thinking is used." Steer with prompting instead |
top_p | Nucleus sampling | Non-default values rejected | Same 400 as temperature |
top_k | Top-K truncation | Non-default values rejected | Same 400 as temperature |
stop_sequences | Custom strings that halt generation | Array of strings | Supported; sets stop_reason: "stop_sequence" |
tools / tool_choice | Tool definitions and forcing behaviour | auto (default), any, tool, none | Forced tool choice works here — the restriction that blocks it applies to manual extended thinking, which this model does not have |
output_config.format | Constrains output to a JSON schema | {"type": "json_schema", "schema": {...}} | Schema needs additionalProperties: false. No recursive schemas, no numeric or string length constraints |
fallbacks | Server-side retry on another model when a classifier declines | Beta, Claude API; "default" mode or a named list | Turns a refusal into an answer inside one call instead of a dead response |
cache_control | Marks a cache breakpoint | {"type": "ephemeral", "ttl": "5m"} or "1h"; max 4 breakpoints | Prefixes under 512 tokens silently do not cache and return no error |
speed | Fast mode | Not supported on this model | Fast mode is documented only for Opus 5 and Opus 4.8. Use effort to trade latency here |
seed, frequency_penalty, presence_penalty, logprobs, n | Determinism, repetition penalties, token probabilities, multiple completions | Not supported — no such parameters exist on the Messages API | For repetition, prompt for it. For multiple samples, send multiple requests |
Only two knobs really matter here. output_config.effort is the one Anthropic calls "the primary control for trading off intelligence, latency, and cost on Claude Fable 5" — and the counterintuitive advice is to try low and medium for routine work, because they often beat higher settings on older models. max_tokens is the second, because it caps thinking and answer together; an under-sized value on a hard task truncates the answer after the model has already spent the budget reasoning. Everything in the temperature family is gone, so output shaping happens entirely in the prompt.
SamplingShaping the output distribution
There is effectively no sampling surface on this model. temperature, top_p and top_k all return a 400 error when set to a non-default value — the thinking docs state this applies "on every request, regardless of whether thinking is used," and the deprecation page records the three parameters as deprecated from Opus 4.7 onward with the recommended replacement being to "omit and use prompting to guide model behavior." So the classic levers for making output more deterministic or more varied are simply unavailable. If you previously used temperature=0 for reproducibility, note that it never guaranteed identical outputs anyway; the API reference says "even with temperature of 0.0, the results will not be fully deterministic." The sensible starting point is therefore to send no sampling parameters at all, leave effort at its high default, and shape tone, length and variability with explicit prompt instructions. If you need several genuinely different answers, issue several requests rather than reaching for a decoding parameter that no longer exists.
ReasoningThinking, effort and budgets
Thinking is always on and cannot be turned off. The per-model configuration table lists Fable 5 as "Adaptive only," default "Always on," and rejects both "enabled" and "disabled" with a 400. You control depth through output_config.effort, not a token budget — budget_tokens does not exist on this model. Thinking tokens are billed as output tokens "even when the thinking text isn't returned to you," and they count toward max_tokens alongside the response.
The raw chain of thought is never returned. thinking.display defaults to "omitted", which gives you thinking blocks with an empty thinking field; setting "summarized" returns a readable summary produced by a different model. Billing is identical under both. In multi-turn conversations you must pass thinking blocks back exactly as received, including empty ones — modified blocks are rejected with a 400.
ToolsFunction calling and server tools
Function calling is supported with the standard tools array and tool_choice of auto, any, tool or none. Because this model has no manual extended thinking mode, the restriction that blocks forced tool choice under thinking: {"type": "enabled"} does not apply — forced tool use works. Parallel tool calls are on by default; return every tool_result in a single user message, since splitting them across messages teaches the model to stop calling tools in parallel.
Structured output uses output_config.format with a JSON schema, generally available for Claude 4.5 and later. Server-side tools listed as supported at launch include the memory tool, code execution and programmatic tool calling, plus context editing and compaction behind beta headers.
The specific failure mode to plan for is not a tool bug: it is the refusal path. A declined request returns HTTP 200 with stop_reason: "refusal", so code that reads content[0] unconditionally breaks on an empty content array.
CostPrice, caching, batching, what drives the bill
List price is $10 per million input tokens and $50 per million output tokens, matching the Fable 5 launch page and the pricing table. What actually drives the bill:
| Line | Rate |
|---|---|
| Base input | $10 / MTok |
| Output (includes thinking tokens) | $50 / MTok |
| 5-minute cache write | $12.50 / MTok (1.25x input) |
| 1-hour cache write | $20 / MTok (2x input) |
| Cache hit / refresh | $1 / MTok (0.1x input) |
| Batch API | $5 in / $25 out (50% off) |
Three factors matter beyond the sticker. The tokenizer is the big one: Claude 4.7-and-later models "produce approximately 30% more tokens for the same text," so a per-token comparison against an older-tokenizer model understates real cost by roughly a third. Second, the full 1M context is billed at standard rates with no long-context premium — a 900k-token request costs the same per token as a 9k one. Third, thinking is always on and always billed as output, and you cannot switch it off to save money; lowering effort is the only lever. Requests refused before any output is generated are not billed at all.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Claude API (first-party) | Yes | Model ID claude-fable-5 |
| Amazon Bedrock | Yes | Model ID anthropic.claude-fable-5, via the Messages-API Bedrock endpoint |
| Claude Platform on AWS | Yes | Anthropic-operated; bare model ID |
| Google Cloud / Vertex AI | Yes | Model ID claude-fable-5 |
| Microsoft Foundry | Yes | Listed on the launch page as GA from June 9, 2026 |
| Hugging Face Inference | No | Closed weights; no model card exists |
| Self-hosting | No | Weights are not released |
| Zero-data-retention orgs | No | Requires 30-day retention; designated a Covered Model |
StrengthsWhat it is good at
- Anthropic's own docs name it the model to use "for workloads that need the highest available capability," above its recommended default of Opus 5.
- 1M-token context at standard per-token pricing — no long-context surcharge, and prompt caching and batch discounts apply across the full window.
- Thinking is always on with no configuration, and depth is steered by a single
effortvalue rather than a token budget you have to tune. - Documented refusal recovery: the
fallbacksparameter retries on another model server-side, and fallback credit refunds the prompt-cache cost of switching. - Lower effort levels are documented to "often exceed
xhighperformance on prior models," so the cheap settings are genuinely usable.
LimitsWhere it falls down
- Safety classifiers can decline a request and return HTTP 200 with
stop_reason: "refusal"and an empty content array — this breaks naive response parsing and has no equivalent on Opus-tier models. - Unavailable to zero-data-retention organisations; every request returns a 400 if the org's retention configuration does not meet the 30-day requirement.
- Thinking cannot be disabled, so you cannot buy latency by turning reasoning off — and thinking tokens bill at the $50/MTok output rate.
- The newer tokenizer emits roughly 30% more tokens for the same text, so effective cost is meaningfully above the headline gap with Opus 5.
- Fast mode is not available, and the pricing page's tool-use system-prompt token table does not list Fable 5 at all, so that overhead is undocumented.
Against its neighboursHow it compares
Against Claude Opus 5, the honest framing is that Opus 5 is the default and Fable 5 is the escalation. Opus 5 costs exactly half ($5/$25 versus $10/$50), has a later reliable knowledge cutoff (May 2026 versus January 2026), supports fast mode, has a lower 512-token cache minimum shared with Fable, and lets you disable thinking at high effort or below. Fable 5 gives you the higher capability tier, and takes back the ability to turn thinking off. Against Claude Opus 4.8, Fable 5 is twice the price and, unlike 4.8, runs thinking by default rather than requiring you to opt in. The deciding question is rarely benchmark-shaped: it is whether your workload can tolerate a refusal path and 30-day retention, and whether you have measured a real quality gain at double the rate.
Getting startedThe smallest call that works
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-fable-5",
"max_tokens": 4096,
"messages": [
{"role": "user", "content": "Explain what a write-ahead log is."}
]
}'Change two things first. Add "output_config": {"effort": "low"} and measure — on this model the cheap settings are unusually strong, and effort is your main cost lever. Then, before shipping, branch on stop_reason before reading content: a classifier refusal arrives as a successful 200 with an empty content array, and unguarded indexing into content[0] will throw.
SourcesWhere every claim above came from
- Introducing Claude Fable 5 and Claude Mythos 5 — Anthropic
- Models overview — Anthropic
- Pricing — Anthropic
- Effort — Anthropic
- Thinking — Anthropic
- Troubleshooting thinking — Anthropic
- Prompt caching — Anthropic
- Structured outputs — Anthropic
- Model deprecations — Anthropic
- Messages API reference — Anthropic
- Could not be confirmed from a primary source as of 2026-07-25: any architecture detail whatsoever — parameter count, dense versus mixture-of-experts, expert count, layer/head counts, vocabulary size, context-extension method, or training compute. Anthropic publishes none of these for Fable 5 and releases no weights, so there is no model card or
config.jsonto read. No benchmark scores are quoted on this page because none were verified against a primary source in this session. The pricing page's per-model tool-use system-prompt token table omits Fable 5, so that overhead is undocumented. The row this page was built from listed Fable 5's reasoning mode as "Extended"; the thinking documentation contradicts that — extended thinking (thinking.type: "enabled") returns a 400 on this model, which is adaptive-only and always on. The row's "~30% more tokens" note is confirmed by the pricing page.
Price and capacity verified 2026-09-18 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-18 (primary source read today: https://platform.claude.com/docs/en/about-claude/pricing — Claude Fable 5 still listed at $10 per million base input and $50 per million output, UNCHANGED, with cache hits at $1.00 per million. Claude Fable 5.1 now sits above it and is its own row. Closes the 2026-09-03 id-drift reading.)
What changedWhat changed here
Updated this page Fable 5.1 has replaced Fable 5 as Anthropic's current top-tier release, so this page's flagship claim is stale.
Update this page so it no longer claims Claude Fable 5 is Anthropic's most capable released model; Fable 5.1 is now the current top tier and the page should describe Fable 5 as the superseded generation.
- Anthropic releases Opus 5.5 with lower prices and Fable-level performance
Anthropic released Claude Opus 5.5, which it calls "the strongest-performing model we've tested to date," at lower prices with Fable-level performance. If you're building on Claude, this is a straight capability upgrade at a lower cost basis — re-benchmark your prompts and re-check your per-token budget.
- AWS Weekly Roundup: Claude Fable 5.1 on AWS, Amazon Linux 2027 preview
Claude Fable 5.1 is now part of the AWS weekly roundup, meaning AWS customers get this generation of Claude through their existing cloud relationship. If your stack is already on AWS, you can adopt it without bolting on a separate vendor connection, which simplifies data-path and compliance work.
- Anthropic Launches Claude Fable 5.1 and Restricted Mythos 5.1 for Advanced Coding, Cybersecurity and Scientific Research
Anthropic has added Claude Fable 5.1 and a restricted Mythos 5.1 targeting advanced coding, cybersecurity, and scientific research. If you build on Claude, retest your security and coding workloads on 5.1 and treat the restricted tier as Anthropic flagging higher-risk use.
- GPT‑6 Astra
GPT-6 Astra is rolling out now to a limited set of organizations and over the coming days to Plus, Pro, Business, and Enterprise users plus the API, priced at $10 per million input and $50 per million output tokens — the same rates as Claude Fable 5 — making it OpenAI's direct frontier competitor. Note its headline ARC-AGI 3 result came via a custom Provider Adapter harness, so measure it on your
- Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work
Anthropic launched Claude Fable 5.1 (with Mythos 5.1), claiming stronger performance than Fable 5 at roughly 25 percent less typical cost and up to 45 percent less for complex agentic tasks, with looser safeguards and changed data-retention terms. If you build agents, this directly cuts your per-task token bill and the friction from overzealous refusals — worth re-benchmarking against whatever you
- Anthropic says Claude Fable 5 will reject fewer biology queries, warns of bioweapon risks | 'More work is needed to refine our safeguards' | Inshorts
Anthropic is loosening biology restrictions on Fable 5, a direct and practical change for anyone building on Claude for life-science work.
- llm-anthropic 0.26
llm-anthropic 0.26 adds claude-fable-5, claude-sonnet-5, and claude-opus-5, with server-side tools (WebSearch, WebFetch, CodeExecution, AnthropicMCP) exposed via the -T interface — so you can wire Claude into CLI/Python workflows without building each integration yourself.
- DeepSeek V4 Flash Hits 8 Trillion Daily Tokens, Tops Trending Charts at 1/105th the Cost of Claude Fable 5
DeepSeek V4 Flash is now processing 8 trillion daily tokens, topping trending charts at 1/105th the cost of Claude Fable 5. For high-volume agent or batch workloads, that price-performance gap is the reason to build against DeepSeek V4 Flash and keep cost ceilings in mind.
- Yang Zhilin Said No To Apple — Now His Kimi K3 Undercuts Claude Fable 5 On Coding
Moonshot's Kimi K3 is being pitched as undercutting Claude Fable 5 on coding — worth adding to your model-evaluation matrix before committing to a Western-only stack.
Three kinds of claim, strongest first. Signal runs every morning.