The balanced Claude: 1M context and 128k max output like the Opus tier, at $2/$10 — launched as an introductory rate, now Anthropic's standard price; the 1 September 2026 rise to $3/$15 was cancelled.
Why this oneWhat it is actually for
Sonnet 5 is the production default when you need frontier-class capacity without frontier pricing. It carries the same 1M-token context window and the same 128k max output as Opus 5 and Fable 5, and Anthropic describes it as "the best combination of speed and intelligence." You choose it over Opus 5 when your workload is high-volume enough that a 2.5x price gap matters more than the top of the quality range, and over Haiku 4.5 when 200k of context is not enough or the reasoning is genuinely hard. The $2/$10 rate was announced as introductory; Anthropic has since made it the standard price and cancelled the scheduled 1 September 2026 increase to $3/$15.
What it isArchitecture, lineage, training
Sonnet 5 is closed-weight and API-only. Anthropic publishes no architecture whatsoever — no parameter count, no dense-versus-mixture-of-experts statement, no expert, layer, attention-head, or vocabulary figures, no rope or positional-encoding details, and no training-corpus or compute description. No weights are released, so there is no Hugging Face model card and no config.json from which such numbers could be read.
The documented envelope is generous for a mid-tier model. The API model ID is claude-sonnet-5, a pinned snapshot in the dateless format. Context window is 1M tokens and max output is 128k on the synchronous Messages API, rising to 300k through the Batches API with the output-300k-2026-03-24 beta header. Reliable knowledge cutoff and training data cutoff are both January 2026.
Two mechanical details place it in the lineage. First, it uses the newer tokenizer: the context-window tooltip describes its 1M window as roughly 555k words, matching Fable 5 and Opus 5 rather than Sonnet 4.6's roughly 750k words per 1M tokens. The pricing page's note that this tokenizer "produces approximately 30% more tokens for the same text" therefore applies here, and explicitly says "Claude Sonnet 4.6 and earlier models use the previous tokenizer."
Second, its thinking configuration matches the Opus 5 generation rather than the Sonnet 4.6 one: adaptive-only, on by default, with manual budget_tokens removed. It is also the first Sonnet-tier model to support the xhigh effort level, which the effort page lists as available on only six models.
At a glanceSee it
A Sonnet 5 request, from the 1M-token context to a response of up to 128k tokens.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model ID | claude-sonnet-5 (pinned snapshot) | Models overview |
| Context window | 1M tokens | Models overview |
| Max output tokens | 128k (synchronous); 300k via Batches with output-300k-2026-03-24 | Models overview |
| Input modalities | Text, image | Models overview |
| Output modalities | Text only | Models overview |
| Reliable knowledge cutoff | January 2026 | Models overview |
| Training data cutoff | January 2026 | Models overview |
| Lifecycle state | Active; tentative retirement "not sooner than June 30, 2027" | Model deprecations |
| Weights available | No — closed API only | No vendor release exists |
| Licence | Not applicable; commercial API terms only | — |
| Parameter count | Not disclosed | — |
| Architecture (dense or MoE) | Not disclosed | — |
| Minimum cacheable prefix | 1,024 tokens | Prompt caching |
| Tool-use system prompt overhead | 354 tokens (auto/none); 474 tokens (any/tool) | Pricing |
| Comparative latency | "Fast" (versus "Moderate" for Opus 5, "Slower" for Fable 5) | Models overview |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | Required; "claude-sonnet-5" | The bare string is a pinned snapshot; do not append a date |
messages | Conversation turns | Required; up to 100,000 messages | First message must be user |
max_tokens | Hard ceiling on thinking plus response text | Required; 128k synchronous | Thinking is on by default here, so a limit tuned on Sonnet 4.6 can truncate — the symptom is stop_reason: "max_tokens" |
system | System prompt | String or array of text blocks | Front of the cache prefix; editing invalidates everything after |
thinking | Thinking configuration | Adaptive only; on by default. {"type": "adaptive"} equals omitting it | {"type": "enabled", "budget_tokens": N} returns 400 — the Sonnet 4.6 transitional escape hatch is gone. {"type": "disabled"} is accepted, at any effort level |
thinking.display | Whether thinking blocks carry text | "summarized" or "omitted"; default "omitted" | Default changed from Sonnet 4.6, which defaulted to "summarized". Streaming UIs will show a pause instead of reasoning unless you opt in |
output_config.effort | Controls thinking volume and total token spend | low, medium, high, xhigh, max; default high | First Sonnet with xhigh. Docs: medium is "comparable to Claude Sonnet 4.6 at high effort"; xhigh for the hardest coding and agentic tasks |
temperature | Sampling temperature | Non-default values rejected | Returns 400 on every request, thinking or not |
top_p / top_k | Nucleus and top-K sampling | Non-default values rejected | Same 400 as temperature |
stop_sequences | Custom halt strings | Array of strings | Supported |
tools / tool_choice | Tool definitions and forcing | auto (default), any, tool, none | Adds 354 or 474 system-prompt tokens. Forced choice works (no manual extended thinking to conflict with) |
output_config.format | JSON schema constraint on output | {"type": "json_schema", "schema": {...}} | GA on the Claude API and Google Cloud; on Bedrock, available through the Messages-API Bedrock endpoint |
cache_control | Cache breakpoint | {"type": "ephemeral", "ttl": "5m"|"1h"}; max 4 | 1,024-token minimum. Below it, nothing caches and no error is returned |
inference_geo | Pins inference geography | "global" (default) or "us" | "us" applies a 1.1x multiplier to all token categories |
speed | Fast mode | Not supported | Fast mode is documented only for Opus 5 and Opus 4.8. Use lower effort for latency here |
seed, frequency_penalty, presence_penalty, logprobs, n | Determinism, penalties, probabilities, multi-sample | Not supported | None exist on the Messages API. Repeat requests for multiple samples |
output_config.effort is the knob to spend time on, because Anthropic gives a direct cross-model anchor: Sonnet 5 at medium is "comparable to Claude Sonnet 4.6 at high effort." That makes medium a real cost saving rather than a guess. thinking.display is the second, and it bites quietly — the default flipped from "summarized" on Sonnet 4.6 to "omitted" here, so a streaming UI that used to show reasoning now shows a long pause. Third is max_tokens, which must absorb both thinking and answer now that thinking runs by default.
SamplingShaping the output distribution
Sonnet 5 accepts no meaningful sampling configuration. The thinking documentation names it alongside the Opus-tier models where "non-default temperature, top_p, or top_k values return a 400 error on every request, regardless of whether thinking is used." This is a change from Sonnet 4.6, where those parameters worked. There is no seed, so reproducibility is not offered.
What replaces them is output_config.effort plus prompting. The effort page gives Sonnet 5 an unusually concrete calibration: high is the default and suits complex reasoning, coding and agentic work; medium is a "cost-saving step-down from the default. Comparable to Claude Sonnet 4.6 at high effort"; low targets "high-volume or latency-sensitive workloads"; xhigh is reserved for "the hardest coding and agentic tasks." That mapping is the practical substitute for a temperature dial.
A sensible starting point is to send no sampling parameters, leave effort at high, and then test medium — the documented equivalence to Sonnet 4.6 at high effort means the step down is measurable rather than speculative. Hold effort constant within a cached conversation, since changing it invalidates the prefix.
ReasoningThinking, effort and budgets
Thinking is adaptive-only and on by default. The per-model configuration table lists Sonnet 5 as "Adaptive only" with default "On" — a change from Sonnet 4.6, where thinking was off until you asked for it. Omitting thinking and sending {"type": "adaptive"} are equivalent. Manual extended thinking is removed: {"type": "enabled", "budget_tokens": N} returns a 400, and unlike Sonnet 4.6 there is no deprecated-but-working escape hatch.
Unlike Opus 5, you can turn thinking off cleanly — {"type": "disabled"} is accepted with no effort-level gate. Depth is controlled entirely by output_config.effort, and Sonnet 5 is the first Sonnet-tier model to offer xhigh.
Thinking tokens bill as output tokens even when the text is not returned, and count toward max_tokens. The display default is "omitted" here, against "summarized" on Sonnet 4.6, so thinking blocks arrive with an empty thinking field unless you opt in.
ToolsFunction calling and server tools
Function calling uses the standard tools array with tool_choice of auto, any, tool or none, each accepting disable_parallel_tool_use. Forced tool choice is available because there is no manual extended thinking mode to conflict with it. Parallel calls are on by default — return all tool_result blocks in one user message, and use is_error: true for failures rather than omitting the result.
Tool definitions cost 354 system-prompt tokens at tool_choice of auto or none, and 474 at any or tool — noticeably more than Opus 5's 286/406, which is worth knowing on tool-heavy high-volume workloads where that overhead recurs on every call.
Structured outputs through output_config.format are generally available, along with strict: true on tool definitions. One platform wrinkle to watch: the structured-outputs page notes Sonnet 5 reaches Bedrock through the Messages-API Bedrock endpoint specifically, rather than the general Bedrock path used by several older models.
CostPrice, caching, batching, what drives the bill
Sonnet 5 launched with $2/$10 described as introductory pricing through 31 August 2026, with $3/$15 to follow. That increase did not happen. The pricing page, re-read on 2026-09-12, now says: "The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur."
| Line | Standard price (read 2026-09-12) |
|---|---|
| Base input | $2 / MTok |
| Output | $10 / MTok |
| 5-minute cache write | $2.50 / MTok |
| 1-hour cache write | $4 / MTok |
| Cache hit / refresh | $0.20 / MTok |
| Batch API | $1 in / $5 out |
Three factors matter. Thinking is on by default and billed at the output rate, so migrating a no-thinking Sonnet 4.6 route raises cost without a code change. The newer tokenizer emits roughly 30% more tokens for the same text than Sonnet 4.6 — measure with the token-counting endpoint against claude-sonnet-5 rather than reusing older counts. And the full 1M window is billed at standard rates with no long-context premium.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Claude API (first-party) | Yes | Model ID claude-sonnet-5 |
| Amazon Bedrock | Yes | Model ID anthropic.claude-sonnet-5, via the Messages-API Bedrock endpoint |
| Claude Platform on AWS | Yes | Anthropic-operated; bare first-party model ID |
| Google Cloud / Vertex AI | Yes | Model ID claude-sonnet-5; structured outputs GA |
| Microsoft Foundry | Yes | Listed among the platforms current models are available through |
| Fast mode | No | Documented for Opus 5 and Opus 4.8 only |
| Hugging Face Inference | No | Closed weights; no model card exists |
| Self-hosting | No | Weights are not released |
StrengthsWhat it is good at
- Same 1M-token context and same 128k max output as Opus 5 and Fable 5, at a fraction of the price — the capacity is not cut down to match the tier.
- Documented latency of "Fast," ahead of Opus 5's "Moderate" and Fable 5's "Slower," on the same context envelope.
- First Sonnet-tier model with the
xhigheffort level, one of only six models Anthropic lists as supporting it. - Anthropic gives a concrete cross-model anchor —
mediumeffort on Sonnet 5 is "comparable to Claude Sonnet 4.6 at high effort" — making the cost step-down measurable rather than guesswork. - Unlike Opus 5,
thinking: {"type": "disabled"}is accepted at every effort level with no gating.
LimitsWhere it falls down
- Thinking is on by default while Sonnet 4.6's was off, so migrated routes silently gain output-rate thinking tokens and may truncate against an unchanged
max_tokens. - The newer tokenizer emits roughly 30% more tokens for identical text than Sonnet 4.6.
- Tool-use system-prompt overhead is 354/474 tokens against Opus 5's 286/406 — worse than the more expensive model, and it recurs on every tool-enabled call.
- No fast mode, and no sampling parameters at all —
temperature,top_pandtop_know 400 where Sonnet 4.6 accepted them.
Against its neighboursHow it compares
Against Claude Opus 5, the capacity is identical — same 1M context, same 128k output — so the trade is price against capability and knowledge. Sonnet 5 runs at $2/$10 versus $5/$25, is documented as "Fast" rather than "Moderate," but has a January 2026 cutoff against Opus 5's May 2026, and no fast mode. Against Claude Haiku 4.5, Sonnet 5 gives you five times the context (1M versus 200k), double the max output (128k versus 64k), adaptive thinking with the full effort ladder rather than manual budget_tokens, and a much later knowledge cutoff — at twice the price. Against Sonnet 4.6 ($3/$15), Sonnet 5 costs less per token and adds thinking-on-by-default and xhigh effort, but the new tokenizer means the same text costs about 30% more tokens.
Getting startedThe smallest call that works
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 4096,
"messages": [
{"role": "user", "content": "Explain what a write-ahead log is."}
]
}'Change two things. Add "output_config": {"effort": "medium"} and compare — Anthropic documents that as roughly equivalent to Sonnet 4.6 at high effort, so it is the cheapest defensible setting. If you stream reasoning to users, also add "thinking": {"type": "adaptive", "display": "summarized"}; the default is "omitted" here and your UI will otherwise show a silent pause.
SourcesWhere every claim above came from
- Models overview — Anthropic
- Pricing — Anthropic (including the Sonnet 5 introductory-pricing note)
- Effort — Anthropic (Recommended effort levels for Claude Sonnet 5)
- Thinking — Anthropic
- Troubleshooting thinking — Anthropic
- Prompt caching — Anthropic
- Structured outputs — Anthropic
- Model deprecations — Anthropic
- Messages API reference — Anthropic
- Corrections to the row this page was built from, and unconfirmed items (as of 2026-07-25). Two fields in the source row are wrong and the documentation is followed instead. (1) The row gave max output as 64K until 2026-09-12, when it was corrected to 128K; the models overview latest-models table lists Claude Sonnet 5 at 128k tokens, and the Batches extended-output note names Sonnet 5 among the models reaching 300k with a beta header. (2) The row gives the reasoning mode as "Extended"; the per-model thinking table lists Sonnet 5 as adaptive only, on by default, with
thinking.type: "enabled"rejected by a 400. Not confirmed from any primary source: every architecture detail (parameter count, dense or MoE, expert/layer/head counts, vocabulary size, training compute) — Anthropic publishes none and releases no weights, so no model card orconfig.jsonexists. No benchmark scores are quoted because none were verified in this session, and the exact release date is not stated on any page read here.
Price and capacity verified 2026-09-12 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-12 (research pass, primary source read today: https://platform.claude.com/docs/en/about-claude/pricing; https://platform.claude.com/docs/en/models/sonnet-5/overview; https://platform.claude.com/docs/en/models/overview — figures confirmed: input_per_m 2.0, output_per_m 10.0, cache_hit_input_per_m 0.2, context 1M, max_output 128K. CORRECTED max_output 64K -> 128K: Vendor model page (platform.claude.com/docs/en/models/sonnet-5/overview), Capabilities table: 'Max output | 128K tokens'; header: 'Context window: 1M tokens · Max output: 128K tokens'. Models overview comparison row: 'Max output | 128K tokens | 128K tokens | 128K tokens | 64K tokens' (Fable 5.1 / Opus 5 / Sonnet 5 / Haiku 4.5).) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $2 in / $10 out per M, $0.2 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-50846c; read of platform.claude.com/docs/en/about-claude/pricing — $2/$10 as the STANDARD rate and the $0.20 cache read unchanged. ⚠︎ THIS READ IS THE ONE THAT MATTERED: the cancelled-increase note is still on the page verbatim — "The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur" — re-confirmed FOUR DAYS before the 2026-08-31 date this row used to carry in expires. Had that field survived, effective_price() would have reverted this row to $3/$15 on Tuesday and overstated Sonnet 5 by 50% on the money page. All five Anthropic rows this file carries were re-read against the model-pricing table and none moved. Claude Mythos 5 sits on that table at $10/$50 and stays omitted here because it is not GA, per the fable-5 row's own note. Cause of the hash move not determined; no row this file carries moved.)
What changedWhat changed here
- llm-anthropic 0.26
llm-anthropic 0.26 adds claude-fable-5, claude-sonnet-5, and claude-opus-5, with server-side tools (WebSearch, WebFetch, CodeExecution, AnthropicMCP) exposed via the -T interface — so you can wire Claude into CLI/Python workflows without building each integration yourself.
Three kinds of claim, strongest first. Signal runs every morning.