Home › Frontier Models › Claude Opus 4.8
Model · Reference

Claude Opus 4.8

Hard agents, large-codebase work

In one line

The previous Opus generation at the same $5/$25 price as Opus 5, still fully active — its defining difference is that thinking is off unless you ask for it.

Why this oneWhat it is actually for

You stay on Opus 4.8 for one of two reasons. Either you have an existing deployment that is tuned and evaluated against it and the migration has no measured payoff — the price is identical to Opus 5, so there is no cost argument for moving — or you specifically want thinking off by default, which is what this model does when you omit the thinking parameter. That distinction is real: on Opus 5 the same request thinks and bills thinking tokens at the output rate. Opus 4.8 also remains the documented fallback target when a Fable 5 request is declined by a safety classifier. It is not deprecated; the deprecation table lists it as Active with retirement no sooner than May 28, 2027.

What it isArchitecture, lineage, training

Opus 4.8 is a closed-weight API-only model. Anthropic discloses no architecture for it — no parameter count, no dense-versus-MoE statement, no expert, layer, head, or vocabulary counts, and no training-corpus or compute description. No weights are published, so there is no Hugging Face model card and no config.json from which real architecture numbers could be read.

The documented envelope: API model ID claude-opus-4-8, a pinned snapshot in the dateless format used from the 4.6 generation onward. Context window 1M tokens, max output 128k on the synchronous Messages API, extending to 300k through the Batches API with the output-300k-2026-03-24 beta header. Reliable knowledge cutoff January 2026, training data cutoff January 2026.

Its place in the lineage is worth stating carefully, because two different Anthropic pages use different words. The models overview files it under a "Legacy models" accordion, describing those models as "still available" with a recommendation to consider migrating. The model deprecations page — which defines the lifecycle vocabulary — lists claude-opus-4-8 with current state "Active," not "Legacy" or "Deprecated," and gives a tentative retirement date not sooner than May 28, 2027. So it is documentation-grouped as legacy but lifecycle-classified as active, with more than a year of runway and no deprecation notice issued.

It shares the tokenizer introduced with Opus 4.7, which the pricing page says "produces approximately 30% more tokens for the same text" than the Sonnet 4.6-and-earlier tokenizer. Its request surface is the same as Opus 4.7's.

At a glanceSee it

Claude Opus 4.8 diagram

Opus 4.8 request flow, showing that omitting the thinking parameter runs no reasoning pass at all.

CapacityContext, output and what fits

FactValueSource
API model IDclaude-opus-4-8 (pinned snapshot)Models overview
Context window1M tokensModels overview
Max output tokens128k (synchronous); 300k via Batches with output-300k-2026-03-24Models overview
Input modalitiesText, imageModels overview
Output modalitiesText onlyModels overview
Reliable knowledge cutoffJanuary 2026Models overview
Training data cutoffJanuary 2026Models overview
Lifecycle stateActive (grouped under "Legacy models" in the overview); retirement "not sooner than May 28, 2027"Model deprecations; Models overview
Weights availableNo — closed API onlyNo vendor release exists
LicenceNot applicable; commercial API terms only—
Parameter countNot disclosed—
Architecture (dense or MoE)Not disclosed—
Minimum cacheable prefix1,024 tokensPrompt caching
Tool-use system prompt overhead290 tokens (auto/none); 410 tokens (any/tool)Pricing

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelSelects the modelRequired; "claude-opus-4-8"Bare string is already pinned; adding a date suffix returns 404
messagesConversation turnsRequired; up to 100,000 messagesFirst message must be user
max_tokensHard ceiling on thinking plus response textRequired; 128k synchronousDocs suggest starting at 64k when running xhigh or max effort
systemSystem promptString or array of text blocksFront of the cache prefix; editing it invalidates everything downstream
thinkingThinking configurationAdaptive only; default off. Set {"type": "adaptive"} explicitly to enableOmitting it runs without thinking — the key difference from Opus 5. {"type": "enabled", "budget_tokens": N} returns 400. {"type": "disabled"} is accepted at any effort level
thinking.displayWhether thinking blocks carry text"summarized" or "omitted"; default "omitted"At the default, blocks arrive with an empty thinking field
output_config.effortControls total token spend, thinking and otherwiselow to max; default highDocs: start at xhigh for coding and agentic work, high elsewhere, stepping down only where evals hold. Changing it invalidates the cache
speedFast mode"fast" with beta header fast-mode-2026-02-01Up to 2.5x output tokens per second at $10/$50 per MTok. Claude API only; unavailable with Batches or Priority Tier
temperatureSampling temperatureNon-default values rejectedReturns 400 on every request, thinking or not
top_p / top_kNucleus and top-K samplingNon-default values rejectedSame 400 as temperature
stop_sequencesCustom halt stringsArray of stringsSupported
tools / tool_choiceTool definitions and forcingauto (default), any, tool, noneAdds 290 or 410 system-prompt tokens depending on choice
output_config.formatJSON schema constraint on output{"type": "json_schema", "schema": {...}}GA on the Claude API, Bedrock and Google Cloud for this model
cache_controlCache breakpoint{"type": "ephemeral", "ttl": "5m"|"1h"}; max 41,024-token minimum — twice Opus 5's. Shorter prefixes silently do not cache
inference_geoPins inference geography"global" (default) or "us""us" applies a 1.1x multiplier to all token categories
seed, frequency_penalty, presence_penalty, logprobs, nDeterminism, penalties, probabilities, multi-sampleNot supportedNone of these exist on the Messages API. Use repeated requests for multiple samples

The one parameter that changes behaviour most on this model is thinking, precisely because its default is off. If you want reasoning, you must send {"type": "adaptive"} — silence gets you a non-thinking response. After that, output_config.effort is the cost dial, with the documented starting point being xhigh for coding and agentic work rather than the high API default. max_tokens matters mainly once thinking is enabled, since the two share one ceiling.

SamplingShaping the output distribution

Sampling is closed off on this model. The thinking documentation states that on Opus 4.8, "non-default temperature, top_p, or top_k values return a 400 error on every request, regardless of whether thinking is used," and the model deprecations page records the three parameters as deprecated from Opus 4.7 onward, with the recommended replacement being to "omit and use prompting to guide model behavior." There is no seed parameter, so there is no reproducibility control either.

Because thinking is off by default here, the shape of the output distribution on a bare request is closer to a conventional single-pass model than on Opus 5 — you get an answer without a reasoning pass unless you opt in. Once you enable adaptive thinking, effort becomes the lever that changes how often and how deeply the model reasons before answering. A sensible starting point is to send no sampling parameters, decide explicitly whether the route needs thinking, and then set effort — Anthropic's guidance is to start at xhigh for coding and agentic work and high otherwise, measuring before stepping down.

ReasoningThinking, effort and budgets

Thinking is adaptive-only and, unlike Opus 5, off by default. The per-model configuration table lists Opus 4.8 as "Adaptive only" with default "Off," so a request that omits thinking runs without any reasoning pass. To turn it on, send thinking: {"type": "adaptive"} explicitly. This is the single most consequential difference between this model and Opus 5, and the one most likely to change your bill when migrating in either direction.

Manual extended thinking is gone: {"type": "enabled", "budget_tokens": N} returns a 400. Depth is set by output_config.effort, which supports the full low through max ladder including xhigh. Unlike Opus 5, {"type": "disabled"} is accepted at every effort level with no gate.

Thinking tokens bill as output tokens and count toward max_tokens. display defaults to "omitted", returning thinking blocks whose thinking field is empty; "summarized" opts into readable summaries.

ToolsFunction calling and server tools

Function calling works through tools with tool_choice of auto, any, tool or none, each accepting disable_parallel_tool_use. Parallel tool calls are enabled by default; collect every tool_result into a single user message rather than splitting them, and return failed calls with is_error: true instead of dropping them.

Tool definitions carry a documented overhead of 290 system-prompt tokens for tool_choice of auto or none, and 410 for any or tool — slightly more than Opus 5's 286/406. The bash tool adds a further 325 input tokens on this model.

Structured outputs via output_config.format are generally available on the Claude API, and the structured-outputs page lists Opus 4.8 as GA on Amazon Bedrock and Google Cloud too. The most common practical trap is not tool-specific: with thinking off by default, agentic loops that assumed reasoning between tool calls will not get it unless you set thinking: {"type": "adaptive"}.

CostPrice, caching, batching, what drives the bill

List price is $5 per million input tokens and $25 per million output tokens — identical to Opus 5. There is no price saving from staying on this model.

LineRate
Base input$5 / MTok
Output (includes thinking tokens when enabled)$25 / MTok
5-minute cache write$6.25 / MTok
1-hour cache write$10 / MTok
Cache hit / refresh$0.50 / MTok
Batch API$2.50 in / $12.50 out
Fast mode$10 in / $50 out

What actually moves the bill: thinking is off by default, so a bare request here is genuinely cheaper than the same request on Opus 5 — that is the one economic argument for the model, and it disappears the moment you enable adaptive thinking. The cache minimum is 1,024 tokens, twice Opus 5's 512, so short prompts that would cache on Opus 5 silently will not cache here. The Opus 4.7-generation tokenizer emits roughly 30% more tokens for the same text than Sonnet 4.6 and earlier, so per-token comparisons across that boundary understate cost. The full 1M window is billed at standard rates with no long-context premium.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Claude API (first-party)YesModel ID claude-opus-4-8
Amazon BedrockYesModel ID anthropic.claude-opus-4-8, via the Messages-API Bedrock endpoint
Claude Platform on AWSYesAnthropic-operated; bare first-party model ID
Google Cloud / Vertex AIYesModel ID claude-opus-4-8; structured outputs GA
Microsoft FoundryUnverifiedNot listed in the legacy-models table read here; the overview's general availability statement does not name it individually
Fast modeClaude API onlyNot on Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS
Hugging Face InferenceNoClosed weights; no model card exists
Self-hostingNoWeights are not released

StrengthsWhat it is good at

  • Same $5/$25 list price as Opus 5, so there is no cost penalty for remaining on a tuned, already-evaluated deployment.
  • Thinking off by default makes a bare request genuinely cheaper than the equivalent Opus 5 call, which now thinks and bills for it.
  • Lifecycle state is Active with retirement no sooner than May 28, 2027 — over a year of runway and no deprecation notice issued.
  • Supports fast mode alongside Opus 5, at up to 2.5x higher output tokens per second for $10/$50 per MTok.
  • thinking: {"type": "disabled"} is accepted at every effort level, with none of the xhigh/max gating that Opus 5 enforces.

LimitsWhere it falls down

  • Superseded in Anthropic's own guidance — the overview groups it under "Legacy models" and directs new work to Opus 5 at the same price.
  • Knowledge cutoff is January 2026, four months behind Opus 5's May 2026, on an identically priced model.
  • Cache minimum is 1,024 tokens against Opus 5's 512, so short prompts that cache on Opus 5 silently fail to cache here with no error.
  • Thinking off by default is a footgun in both directions: agentic loops migrated from Opus 5 lose reasoning between tool calls unless you set it explicitly.
  • Higher tool-use system-prompt overhead than Opus 5 (290/410 versus 286/406 tokens), and no sampling parameters at all.

Against its neighboursHow it compares

Against Claude Opus 5 the comparison is unusually clean, because the price is the same. Opus 5 has the later knowledge cutoff (May 2026 versus January 2026), the halved cache minimum (512 versus 1,024 tokens), and slightly lower tool overhead. Opus 4.8 keeps thinking off by default and accepts thinking: {"type": "disabled"} at any effort. Unless you specifically want the non-thinking default or have evals pinned to this model, Opus 5 is the better buy at identical cost. Against Claude Fable 5, Opus 4.8 is half the price and has no safety-classifier refusal path — and is in fact the model Anthropic documents as the fallback target when Fable 5 declines a request. Against Claude Sonnet 5, Opus 4.8 costs more than double Sonnet 5's $2/$10 for the same 1M context and 128k output.

Getting startedThe smallest call that works

code
curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-opus-4-8",
    "max_tokens": 4096,
    "thinking": {"type": "adaptive"},
    "messages": [
      {"role": "user", "content": "Explain what a write-ahead log is."}
    ]
  }'

The thinking block above is deliberate — drop it and this model answers with no reasoning pass at all, which is the single most common surprise when moving between Opus 4.8 and Opus 5. Change output_config.effort next: Anthropic recommends starting at xhigh for coding and agentic work, and raising max_tokens toward 64k if you do.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-12 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-12 (research pass, primary source read today: https://platform.claude.com/docs/en/about-claude/pricing; https://platform.claude.com/docs/en/models/opus-4-8/overview — figures confirmed: input_per_m 5.0, output_per_m 25.0, cache_hit_input_per_m 0.5, context 1M, max_output 128K. Observed on the page, not added: Pricing row verbatim: 'Claude Opus 4.8 | $5 / MTok | $6.25 / MTok | $10 / MTok | $0.50 / MTok | $25 / MTok'. Model page: 'Context window1Mtokens', 'Max output128Ktokens', 'StatusActive (legacy)', 'RetirementNot sooner than May 28, 2027' - legacy status matches the row's 'for' text. Context/max_output are not on the pricing page itself.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $5 in / $25 out per M, $0.5 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-50846c; read of platform.claude.com/docs/en/about-claude/pricing — $5/$25 and the $0.50 cache read unchanged. All five Anthropic rows this file carries were re-read against the model-pricing table and none moved. Claude Mythos 5 sits on that table at $10/$50 and stays omitted here because it is not GA, per the fable-5 row's own note. Cause of the hash move not determined; no row this file carries moved.)

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A Claude Opus 5.5 model is now callable, which supersedes the Opus 5 and Opus 4.8 entries as the current Opus tier.

    The Claude model pages and the costing board must be updated: a Claude Opus 5.5 now exists, so the Opus 5 and Opus 4.8 entries are no longer the current Opus tier and their pricing and status claims need revising.

    Simon Willison · 23 Sep 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning