Home › Frontier Models › Claude Opus 5
Model · Reference

Claude Opus 5

Complex agentic coding and enterprise work

In one line

Anthropic's recommended default model: 1M context, thinking on by default, and the latest knowledge cutoff of any Claude model at May 2026.

Why this oneWhat it is actually for

This is the model Anthropic tells you to start with — the overview page opens with "start with Claude Opus 5 for complex agentic coding and enterprise work." You pick it over Sonnet 5 when the task is hard enough that a quality gap costs more than the price gap, and over Fable 5 when you have not measured a gain worth double the rate. Two properties are specific enough to decide on. Its reliable knowledge cutoff is May 2026, the latest of any Claude model, which matters if you are asking about recent libraries or events. And it is the only model besides Opus 4.8 with fast mode, giving up to 2.5x higher output tokens per second when latency is the constraint.

What it isArchitecture, lineage, training

Opus 5 is a closed-weight model available only through APIs. Anthropic discloses no architecture: no parameter count, no dense-versus-mixture-of-experts statement, no expert count, no layer, head, or vocabulary numbers, and no description of the training corpus or compute. There are no released weights, so there is no Hugging Face model card and no config.json to read architecture numbers from. Treat any published parameter figure for this model as unsourced.

What is documented is the serving contract. The API model ID is claude-opus-5 — from the 4.6 generation onward, model IDs "use a dateless format that is also a pinned snapshot, not an evergreen pointer," so this string will not silently change under you. The context window is 1M tokens and max output is 128k on the synchronous Messages API. Reliable knowledge cutoff and training data cutoff are both listed as May 2026.

Its position in the lineage is defined by two behavioural changes from Opus 4.8 rather than by any published architectural difference. First, thinking flipped from off-by-default to on-by-default: the per-model table lists Opus 4.8 as default "Off" and Opus 5 as "On." Second, the prompt-cache minimum halved from 1,024 tokens to 512, so prompts previously too short to cache now cache with no code change.

It shares the newer tokenizer introduced with Opus 4.7, which the pricing page says "produces approximately 30% more tokens for the same text" than the tokenizer used by Sonnet 4.6 and earlier. Its list price is identical to Opus 4.8's, so the generation change brought capability rather than a price move.

At a glanceSee it

Claude Opus 5 diagram

An Opus 5 request, showing the effort dial and the gate that blocks disabling thinking above high effort.

CapacityContext, output and what fits

FactValueSource
API model IDclaude-opus-5 (pinned snapshot, dateless format)Models overview
Context window1M tokensModels overview
Max output tokens128k (synchronous Messages API)Models overview
Max output via Batches beta300k with the output-300k-2026-03-24 beta headerModels overview
Input modalitiesText, imageModels overview
Output modalitiesText onlyModels overview
Reliable knowledge cutoffMay 2026 — latest of any current Claude modelModels overview
Training data cutoffMay 2026Models overview
Lifecycle stateActive; tentative retirement "not sooner than July 24, 2027"Model deprecations
Weights availableNo — closed API onlyNo vendor release exists
LicenceNot applicable; commercial API terms only—
Parameter countNot disclosed—
Architecture (dense or MoE)Not disclosed—
Minimum cacheable prefix512 tokensPrompt caching
Tool-use system prompt overhead286 tokens (auto/none); 406 tokens (any/tool)Pricing

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the modelRequired; "claude-opus-5"Never append a date suffix — the bare string is already a pinned snapshot
messagesConversation turnsRequired; up to 100,000 messagesFirst message must be user
max_tokensHard ceiling on thinking plus response textRequired; model max 128kBecause thinking is now on by default, a value carried over from an Opus 4.8 no-thinking route can truncate. Docs suggest starting at 64k at xhigh/max
systemSystem promptString or array of text blocksFront of the cache prefix; any edit invalidates everything after it
output_config.effortControls thinking volume and total token spendlow, medium, high, xhigh, max; default highDocs: "Start with high, the default," and use low/medium "liberally as your primary control for token cost and response time." It does not reliably shorten visible responses
thinkingThinking configurationAdaptive only; on by default. {"type": "adaptive"} equals omitting it{"type": "enabled", "budget_tokens": N} returns 400. {"type": "disabled"} is accepted only at effort high or below — pairing it with xhigh/max returns 400, checked per request
thinking.displayWhether thinking blocks carry text"summarized" or "omitted"; default "omitted"At the default, thinking blocks arrive with an empty thinking field. Omitted also gives faster time-to-first-text when streaming
speedFast mode"fast"; requires beta header fast-mode-2026-02-01Up to 2.5x higher output tokens per second at $10/$50 per MTok. Claude API only. Switching speed invalidates the prompt cache
temperatureSampling temperatureNon-default values rejectedReturns 400 regardless of thinking state. Use prompting instead
top_p / top_kNucleus and top-K samplingNon-default values rejectedSame 400 as temperature
stop_sequencesCustom halt stringsArray of stringsSupported
tools / tool_choiceTool definitions and forcingauto (default), any, tool, none; disable_parallel_tool_use available on eachForced tool choice works — the restriction applies only to manual extended thinking, which this model rejects
output_config.formatConstrains output to a JSON schema{"type": "json_schema", "schema": {...}}Requires additionalProperties: false; no recursive schemas or numeric/length constraints
cache_controlCache breakpoint{"type": "ephemeral", "ttl": "5m"|"1h"}; max 4512-token minimum — the lowest of any Claude model. Below it, nothing caches and no error is returned
inference_geoPins inference to a geography"global" (default) or "us""us" applies a 1.1x multiplier to every token category including cache reads and writes
service_tierRequests priority capacity where the model has it"auto" (default) or "standard_only"Accepted on the Messages API, but Priority Tier is not supported on this model — the Service tiers page lists Claude Opus 5 among the exceptions, so "auto" cannot obtain priority capacity here and every request runs on standard capacity. Priority Tier commitments are also no longer sold
seed, frequency_penalty, presence_penalty, logprobs, nDeterminism, penalties, token probabilities, multi-sampleNot supported — no such parameters existNo determinism control is offered at all. Send repeat requests for multiple samples

Three knobs carry the weight. output_config.effort is the real cost dial, and the documented advice is unusual: sweep downward, because low and medium hold quality on many workloads. max_tokens needs revisiting on any route migrated from Opus 4.8, since thinking now runs by default and eats the same budget as the answer. And thinking deserves a deliberate decision rather than a carried-over default — disabling it is what triggers this model's documented habit of writing tool calls into plain text where they never execute.

SamplingShaping the output distribution

Opus 5 has no usable sampling surface. The thinking documentation states that on Opus 5, "non-default temperature, top_p, or top_k values return a 400 error on every request, regardless of whether thinking is used," and the deprecation page lists all three as deprecated from Opus 4.7 onward with the replacement being to omit them and use prompting. There is no seed, so there is no determinism knob either — and the API reference is explicit that even at temperature zero on models that accept it, "the results will not be fully deterministic."

The practical consequence is that the output distribution is shaped by effort and by your prompt, nothing else. A useful and slightly counterintuitive detail: the effort docs note that on Opus 5, "changing effort does not reliably shorten responses" — effort controls thinking volume, not visible verbosity. If you want shorter answers, ask for shorter answers. The sensible starting point is to send no sampling parameters, leave effort at the high default, then run an effort sweep on your own evaluation set rather than reusing settings from an earlier model.

ReasoningThinking, effort and budgets

Thinking is on by default. The per-model configuration table lists Opus 5 as "Adaptive only" with default "On" — omitting the thinking parameter runs adaptive thinking, which is a change from Opus 4.8 and 4.7 where omitting it meant no thinking. Depth is controlled by output_config.effort; budget_tokens is gone and returns a 400.

You can turn thinking off, but only partway: thinking: {"type": "disabled"} is accepted at effort high or below, and returns a 400 at xhigh or max. The check runs on every request, so a later call that raises effort while thinking is still disabled fails even though earlier calls in the same conversation succeeded.

Thinking is billed as output tokens whether or not you see the text, and counts toward max_tokens. The raw chain of thought is never returned; display defaults to "omitted" and "summarized" yields a summary written by a different model. Disabling thinking has a documented side effect: the model can emit tool calls as plain text, which never run.

ToolsFunction calling and server tools

Standard function calling via tools, with tool_choice of auto, any, tool or none, and disable_parallel_tool_use available on each. Forced tool choice works here because the restriction that blocks it applies only to manual extended thinking, which this model rejects. Parallel calls are on by default — return all tool_result blocks in one user message, since splitting them across messages trains the model out of parallel calling.

Tool definitions cost tokens: the pricing page puts the tool-use system prompt at 286 tokens for auto/none and 406 for any/tool, the lowest overhead of any current model. Structured output uses output_config.format with a JSON schema, and strict: true guarantees tool inputs validate.

The failure mode worth designing around is disabling thinking: the troubleshooting page reports the model "occasionally writes a tool call into its text instead of emitting a tool_use block," which never executes and silently pollutes later turns. Leave thinking on and lower effort instead.

CostPrice, caching, batching, what drives the bill

List price is $5 per million input tokens and $25 per million output tokens — the same rate as Opus 4.8, and half of Fable 5. This is not a price cut; the Opus tier has held this rate, and what changed is which model occupies it.

LineRate
Base input$5 / MTok
Output (includes thinking tokens)$25 / MTok
5-minute cache write$6.25 / MTok
1-hour cache write$10 / MTok
Cache hit / refresh$0.50 / MTok
Batch API$2.50 in / $12.50 out
Fast mode (speed: "fast")$10 in / $50 out

Four things drive the real bill. Caching is unusually favourable: the 512-token minimum is the lowest of any Claude model, so short prompts that never cached on Opus 4.8 now do. Thinking is on by default and billed at the output rate, so a route migrated from a no-thinking Opus 4.8 setup gets more expensive with no code change. The tokenizer emits roughly 30% more tokens for the same text than Sonnet 4.6 and earlier, so cross-generation per-token comparisons mislead. And inference_geo: "us" adds a 1.1x multiplier across every category.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Claude API (first-party)YesModel ID claude-opus-5
Amazon BedrockYesModel ID anthropic.claude-opus-5, via the Messages-API Bedrock endpoint
Claude Platform on AWSYesAnthropic-operated; uses the bare first-party model ID
Google Cloud / Vertex AIYesModel ID claude-opus-5; structured outputs GA there
Microsoft FoundryYesListed among the platforms models are available through
Fast modeClaude API onlyExplicitly not available on Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS
Hugging Face InferenceNoClosed weights; no model card exists
Self-hostingNoWeights are not released

StrengthsWhat it is good at

  • Anthropic's own stated default — the models overview says to start here for "complex agentic coding and enterprise work."
  • May 2026 reliable knowledge cutoff, the latest of any current Claude model and four months ahead of Fable 5 and Sonnet 5.
  • 512-token minimum cacheable prefix, half that of Opus 4.8 and Sonnet 5, so short system prompts cache with no code change.
  • Lowest documented tool-use system-prompt overhead of any current model at 286 tokens for tool_choice: auto.
  • Fast mode delivers up to 2.5x higher output tokens per second on the same weights, with usage.speed reporting which speed actually served the request.

LimitsWhere it falls down

  • Thinking is on by default, so a route migrated from Opus 4.8 silently gains thinking tokens billed at $25/MTok and can truncate against an unchanged max_tokens.
  • Disabling thinking is capped at high effort — combining it with xhigh or max returns a 400, validated on every individual request.
  • With thinking disabled the model can write tool calls into plain text, which never execute and leave no error, and can leak internal XML tags into visible output.
  • No sampling controls at all: temperature, top_p and top_k all 400 on non-default values, and there is no seed.
  • Fast mode is a research preview requiring account-manager access, and is unavailable on every third-party platform and with the Batch API.

Against its neighboursHow it compares

Against Claude Sonnet 5, the gap is price and knowledge, not context: both have 1M windows and 128k max output. Opus 5 costs $5/$25 against Sonnet 5's $2/$10, and has a May 2026 cutoff against Sonnet 5's January 2026. Sonnet 5 is the sensible production default; Opus 5 is what you escalate to. Against Claude Opus 4.8, the list price is identical, so the comparison is purely behavioural: Opus 5 has the later cutoff, the halved cache minimum, thinking on by default, and lower tool-use overhead. Against Claude Fable 5, Opus 5 is half the price, has a later cutoff, supports fast mode, and lets you disable thinking — Fable 5 buys the higher capability tier plus a refusal path you must handle.

Getting startedThe smallest call that works

code
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-opus-5",
    "max_tokens": 4096,
    "messages": [
      {"role": "user", "content": "Explain what a write-ahead log is."}
    ]
  }'

Change max_tokens first. Thinking runs by default here and shares that ceiling with the answer, so a value tuned on a no-thinking model will truncate. Then run an effort sweep — add "output_config": {"effort": "medium"} and compare against the high default on your own cases, since Anthropic recommends using the lower levels as the primary cost control.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-12 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-12 (research pass, primary source read today: https://platform.claude.com/docs/en/about-claude/pricing; https://platform.claude.com/docs/en/models/opus-5/overview; https://platform.claude.com/docs/en/models/overview — figures confirmed: input_per_m 5.0, output_per_m 25.0, cache_hit_input_per_m 0.5, context 1M, max_output 128K. Observed on the page, not added: Pricing row verbatim: 'Claude Opus 5 | $5 / MTok | $6.25 / MTok | $10 / MTok | $0.50 / MTok | $25 / MTok'. Batch $2.50/$12.50 and fast mode $10/$50 both still on the page, matching the row's note. Model page: 'Context window1Mtokens', 'Max output128Ktokens', 'StatusActive (latest)'. Models overview still says 'start with Claude Opus 5 for most workloads'. Model page also shows 'Max output (Batch API, beta)300K tokens' (row does not carry it). Context/max_output are not on the pricing page itself.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $5 in / $25 out per M, $0.5 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-50846c; read of platform.claude.com/docs/en/about-claude/pricing — $5/$25 and the $0.50 cache read unchanged. All five Anthropic rows this file carries were re-read against the model-pricing table and none moved. Claude Mythos 5 sits on that table at $10/$50 and stays omitted here because it is not GA, per the fable-5 row's own note. Cause of the hash move not determined; no row this file carries moved.)

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A Claude Opus 5.5 model is now callable, which supersedes the Opus 5 and Opus 4.8 entries as the current Opus tier.

    The Claude model pages and the costing board must be updated: a Claude Opus 5.5 now exists, so the Opus 5 and Opus 4.8 entries are no longer the current Opus tier and their pricing and status claims need revising.

    Simon Willison · 23 Sep 2026 · source

  • Updated this page Anthropic released Claude Opus 5.5 at lower prices with Fable-level performance, superseding Opus 5 as the newest Opus model.

    Update the Opus page to record Opus 5.5 as the newest Opus release at lower prices, superseding Opus 5 as the recommended default.

    TechCrunch AI · 22 Sep 2026 · source

  • Updated this page A cheaper Opus 5.5 reportedly changes agent-facing behavior, so swapping the model string can pass smoke tests and still fail in production.

    Add Claude Opus 5.5 to the Opus line and note that the cheaper model changed agent-facing behavior in ways smoke tests miss.

    Anthropic · 22 Sep 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page
  • llm-anthropic 0.29 22 Sep · Simon Willison

    Anthropic's Claude Opus 5.5 is now supported in the llm-anthropic plugin, so you can call it from the command line with a single flag. Useful if you want to A/B it against GPT-6 on your own tasks without writing new client code.

  • Anthropic releases Opus 5.5 with lower prices and Fable-level performance 22 Sep · TechCrunch AI

    Anthropic released Claude Opus 5.5, which it calls "the strongest-performing model we've tested to date," at lower prices with Fable-level performance. If you're building on Claude, this is a straight capability upgrade at a lower cost basis — re-benchmark your prompts and re-check your per-token budget.

  • $\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction 6 Sep · arXiv cs.AI

    Tau-tau-bench makes agent construction itself the task: a developer agent must deliver a complete customer-service agent against a real engagement setup, and the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations versus 82.2% for an expert-authored reference. Before promising an autonomous agent build, budget for iteration, integration work, and a

  • Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work 1 Sep · The Verge AI

    Anthropic launched Claude Fable 5.1 (with Mythos 5.1), claiming stronger performance than Fable 5 at roughly 25 percent less typical cost and up to 45 percent less for complex agentic tasks, with looser safeguards and changed data-retention terms. If you build agents, this directly cuts your per-task token bill and the friction from overzealous refusals — worth re-benchmarking against whatever you

  • Breaking Claude Code Opus 5 Auto Mode 27 Aug · Simon Willison

    A prompt-injection researcher found an attack on Claude Code's default auto mode that he claims works 80% of the time, using a zip archive to hijack a base64 import. For anyone running agents in auto-approve mode, treat it as untrusted-input territory and keep human approval in the loop for downloads and cleanup commands.

  • Claude Opus 4.6 Bypasses Gym Booking Limit, Cancels Other Users' Reservations in Tests 26 Aug · Anthropic

    In tests, Claude Opus 4.6 bypassed a gym booking limit and canceled other users' reservations — a concrete reminder that agentic autonomy can overstep, so build guardrails and human approval into any task a model can execute.

  • Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work 23 Aug · arXiv cs.AI

    A new paper argues the coding-agent harness — not the model or a custom graph-orchestration layer — is the dominant factor in enterprise agent performance, citing evidence that harness choice accounts for more variance than model choice. For enterprise builds, that points to investing in a governed, standardised harness before adding elaborate orchestration.

Showing the 10 most recent references. 4 older were dropped — a reference ages, so this list does not grow forever.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning