The previous Opus generation at the same $5/$25 price as Opus 5, still fully active — its defining difference is that thinking is off unless you ask for it.
Why this oneWhat it is actually for
You stay on Opus 4.8 for one of two reasons. Either you have an existing deployment that is tuned and evaluated against it and the migration has no measured payoff — the price is identical to Opus 5, so there is no cost argument for moving — or you specifically want thinking off by default, which is what this model does when you omit the thinking parameter. That distinction is real: on Opus 5 the same request thinks and bills thinking tokens at the output rate. Opus 4.8 also remains the documented fallback target when a Fable 5 request is declined by a safety classifier. It is not deprecated; the deprecation table lists it as Active with retirement no sooner than May 28, 2027.
What it isArchitecture, lineage, training
Opus 4.8 is a closed-weight API-only model. Anthropic discloses no architecture for it — no parameter count, no dense-versus-MoE statement, no expert, layer, head, or vocabulary counts, and no training-corpus or compute description. No weights are published, so there is no Hugging Face model card and no config.json from which real architecture numbers could be read.
The documented envelope: API model ID claude-opus-4-8, a pinned snapshot in the dateless format used from the 4.6 generation onward. Context window 1M tokens, max output 128k on the synchronous Messages API, extending to 300k through the Batches API with the output-300k-2026-03-24 beta header. Reliable knowledge cutoff January 2026, training data cutoff January 2026.
Its place in the lineage is worth stating carefully, because two different Anthropic pages use different words. The models overview files it under a "Legacy models" accordion, describing those models as "still available" with a recommendation to consider migrating. The model deprecations page — which defines the lifecycle vocabulary — lists claude-opus-4-8 with current state "Active," not "Legacy" or "Deprecated," and gives a tentative retirement date not sooner than May 28, 2027. So it is documentation-grouped as legacy but lifecycle-classified as active, with more than a year of runway and no deprecation notice issued.
It shares the tokenizer introduced with Opus 4.7, which the pricing page says "produces approximately 30% more tokens for the same text" than the Sonnet 4.6-and-earlier tokenizer. Its request surface is the same as Opus 4.7's.
At a glanceSee it
Opus 4.8 request flow, showing that omitting the thinking parameter runs no reasoning pass at all.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model ID | claude-opus-4-8 (pinned snapshot) | Models overview |
| Context window | 1M tokens | Models overview |
| Max output tokens | 128k (synchronous); 300k via Batches with output-300k-2026-03-24 | Models overview |
| Input modalities | Text, image | Models overview |
| Output modalities | Text only | Models overview |
| Reliable knowledge cutoff | January 2026 | Models overview |
| Training data cutoff | January 2026 | Models overview |
| Lifecycle state | Active (grouped under "Legacy models" in the overview); retirement "not sooner than May 28, 2027" | Model deprecations; Models overview |
| Weights available | No — closed API only | No vendor release exists |
| Licence | Not applicable; commercial API terms only | — |
| Parameter count | Not disclosed | — |
| Architecture (dense or MoE) | Not disclosed | — |
| Minimum cacheable prefix | 1,024 tokens | Prompt caching |
| Tool-use system prompt overhead | 290 tokens (auto/none); 410 tokens (any/tool) | Pricing |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | Required; "claude-opus-4-8" | Bare string is already pinned; adding a date suffix returns 404 |
messages | Conversation turns | Required; up to 100,000 messages | First message must be user |
max_tokens | Hard ceiling on thinking plus response text | Required; 128k synchronous | Docs suggest starting at 64k when running xhigh or max effort |
system | System prompt | String or array of text blocks | Front of the cache prefix; editing it invalidates everything downstream |
thinking | Thinking configuration | Adaptive only; default off. Set {"type": "adaptive"} explicitly to enable | Omitting it runs without thinking — the key difference from Opus 5. {"type": "enabled", "budget_tokens": N} returns 400. {"type": "disabled"} is accepted at any effort level |
thinking.display | Whether thinking blocks carry text | "summarized" or "omitted"; default "omitted" | At the default, blocks arrive with an empty thinking field |
output_config.effort | Controls total token spend, thinking and otherwise | low to max; default high | Docs: start at xhigh for coding and agentic work, high elsewhere, stepping down only where evals hold. Changing it invalidates the cache |
speed | Fast mode | "fast" with beta header fast-mode-2026-02-01 | Up to 2.5x output tokens per second at $10/$50 per MTok. Claude API only; unavailable with Batches or Priority Tier |
temperature | Sampling temperature | Non-default values rejected | Returns 400 on every request, thinking or not |
top_p / top_k | Nucleus and top-K sampling | Non-default values rejected | Same 400 as temperature |
stop_sequences | Custom halt strings | Array of strings | Supported |
tools / tool_choice | Tool definitions and forcing | auto (default), any, tool, none | Adds 290 or 410 system-prompt tokens depending on choice |
output_config.format | JSON schema constraint on output | {"type": "json_schema", "schema": {...}} | GA on the Claude API, Bedrock and Google Cloud for this model |
cache_control | Cache breakpoint | {"type": "ephemeral", "ttl": "5m"|"1h"}; max 4 | 1,024-token minimum — twice Opus 5's. Shorter prefixes silently do not cache |
inference_geo | Pins inference geography | "global" (default) or "us" | "us" applies a 1.1x multiplier to all token categories |
seed, frequency_penalty, presence_penalty, logprobs, n | Determinism, penalties, probabilities, multi-sample | Not supported | None of these exist on the Messages API. Use repeated requests for multiple samples |
The one parameter that changes behaviour most on this model is thinking, precisely because its default is off. If you want reasoning, you must send {"type": "adaptive"} — silence gets you a non-thinking response. After that, output_config.effort is the cost dial, with the documented starting point being xhigh for coding and agentic work rather than the high API default. max_tokens matters mainly once thinking is enabled, since the two share one ceiling.
SamplingShaping the output distribution
Sampling is closed off on this model. The thinking documentation states that on Opus 4.8, "non-default temperature, top_p, or top_k values return a 400 error on every request, regardless of whether thinking is used," and the model deprecations page records the three parameters as deprecated from Opus 4.7 onward, with the recommended replacement being to "omit and use prompting to guide model behavior." There is no seed parameter, so there is no reproducibility control either.
Because thinking is off by default here, the shape of the output distribution on a bare request is closer to a conventional single-pass model than on Opus 5 — you get an answer without a reasoning pass unless you opt in. Once you enable adaptive thinking, effort becomes the lever that changes how often and how deeply the model reasons before answering. A sensible starting point is to send no sampling parameters, decide explicitly whether the route needs thinking, and then set effort — Anthropic's guidance is to start at xhigh for coding and agentic work and high otherwise, measuring before stepping down.
ReasoningThinking, effort and budgets
Thinking is adaptive-only and, unlike Opus 5, off by default. The per-model configuration table lists Opus 4.8 as "Adaptive only" with default "Off," so a request that omits thinking runs without any reasoning pass. To turn it on, send thinking: {"type": "adaptive"} explicitly. This is the single most consequential difference between this model and Opus 5, and the one most likely to change your bill when migrating in either direction.
Manual extended thinking is gone: {"type": "enabled", "budget_tokens": N} returns a 400. Depth is set by output_config.effort, which supports the full low through max ladder including xhigh. Unlike Opus 5, {"type": "disabled"} is accepted at every effort level with no gate.
Thinking tokens bill as output tokens and count toward max_tokens. display defaults to "omitted", returning thinking blocks whose thinking field is empty; "summarized" opts into readable summaries.
ToolsFunction calling and server tools
Function calling works through tools with tool_choice of auto, any, tool or none, each accepting disable_parallel_tool_use. Parallel tool calls are enabled by default; collect every tool_result into a single user message rather than splitting them, and return failed calls with is_error: true instead of dropping them.
Tool definitions carry a documented overhead of 290 system-prompt tokens for tool_choice of auto or none, and 410 for any or tool — slightly more than Opus 5's 286/406. The bash tool adds a further 325 input tokens on this model.
Structured outputs via output_config.format are generally available on the Claude API, and the structured-outputs page lists Opus 4.8 as GA on Amazon Bedrock and Google Cloud too. The most common practical trap is not tool-specific: with thinking off by default, agentic loops that assumed reasoning between tool calls will not get it unless you set thinking: {"type": "adaptive"}.
CostPrice, caching, batching, what drives the bill
List price is $5 per million input tokens and $25 per million output tokens — identical to Opus 5. There is no price saving from staying on this model.
| Line | Rate |
|---|---|
| Base input | $5 / MTok |
| Output (includes thinking tokens when enabled) | $25 / MTok |
| 5-minute cache write | $6.25 / MTok |
| 1-hour cache write | $10 / MTok |
| Cache hit / refresh | $0.50 / MTok |
| Batch API | $2.50 in / $12.50 out |
| Fast mode | $10 in / $50 out |
What actually moves the bill: thinking is off by default, so a bare request here is genuinely cheaper than the same request on Opus 5 — that is the one economic argument for the model, and it disappears the moment you enable adaptive thinking. The cache minimum is 1,024 tokens, twice Opus 5's 512, so short prompts that would cache on Opus 5 silently will not cache here. The Opus 4.7-generation tokenizer emits roughly 30% more tokens for the same text than Sonnet 4.6 and earlier, so per-token comparisons across that boundary understate cost. The full 1M window is billed at standard rates with no long-context premium.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Claude API (first-party) | Yes | Model ID claude-opus-4-8 |
| Amazon Bedrock | Yes | Model ID anthropic.claude-opus-4-8, via the Messages-API Bedrock endpoint |
| Claude Platform on AWS | Yes | Anthropic-operated; bare first-party model ID |
| Google Cloud / Vertex AI | Yes | Model ID claude-opus-4-8; structured outputs GA |
| Microsoft Foundry | Unverified | Not listed in the legacy-models table read here; the overview's general availability statement does not name it individually |
| Fast mode | Claude API only | Not on Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS |
| Hugging Face Inference | No | Closed weights; no model card exists |
| Self-hosting | No | Weights are not released |
StrengthsWhat it is good at
- Same $5/$25 list price as Opus 5, so there is no cost penalty for remaining on a tuned, already-evaluated deployment.
- Thinking off by default makes a bare request genuinely cheaper than the equivalent Opus 5 call, which now thinks and bills for it.
- Lifecycle state is Active with retirement no sooner than May 28, 2027 — over a year of runway and no deprecation notice issued.
- Supports fast mode alongside Opus 5, at up to 2.5x higher output tokens per second for $10/$50 per MTok.
thinking: {"type": "disabled"}is accepted at every effort level, with none of thexhigh/maxgating that Opus 5 enforces.
LimitsWhere it falls down
- Superseded in Anthropic's own guidance — the overview groups it under "Legacy models" and directs new work to Opus 5 at the same price.
- Knowledge cutoff is January 2026, four months behind Opus 5's May 2026, on an identically priced model.
- Cache minimum is 1,024 tokens against Opus 5's 512, so short prompts that cache on Opus 5 silently fail to cache here with no error.
- Thinking off by default is a footgun in both directions: agentic loops migrated from Opus 5 lose reasoning between tool calls unless you set it explicitly.
- Higher tool-use system-prompt overhead than Opus 5 (290/410 versus 286/406 tokens), and no sampling parameters at all.
Against its neighboursHow it compares
Against Claude Opus 5 the comparison is unusually clean, because the price is the same. Opus 5 has the later knowledge cutoff (May 2026 versus January 2026), the halved cache minimum (512 versus 1,024 tokens), and slightly lower tool overhead. Opus 4.8 keeps thinking off by default and accepts thinking: {"type": "disabled"} at any effort. Unless you specifically want the non-thinking default or have evals pinned to this model, Opus 5 is the better buy at identical cost. Against Claude Fable 5, Opus 4.8 is half the price and has no safety-classifier refusal path — and is in fact the model Anthropic documents as the fallback target when Fable 5 declines a request. Against Claude Sonnet 5, Opus 4.8 costs more than double Sonnet 5's $2/$10 for the same 1M context and 128k output.
Getting startedThe smallest call that works
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-4-8",
"max_tokens": 4096,
"thinking": {"type": "adaptive"},
"messages": [
{"role": "user", "content": "Explain what a write-ahead log is."}
]
}'The thinking block above is deliberate — drop it and this model answers with no reasoning pass at all, which is the single most common surprise when moving between Opus 4.8 and Opus 5. Change output_config.effort next: Anthropic recommends starting at xhigh for coding and agentic work, and raising max_tokens toward 64k if you do.
SourcesWhere every claim above came from
- Models overview — Anthropic (Legacy models section)
- Model deprecations — Anthropic
- Pricing — Anthropic
- Effort — Anthropic
- Thinking — Anthropic
- Troubleshooting thinking — Anthropic
- Fast mode (research preview) — Anthropic
- Prompt caching — Anthropic
- Structured outputs — Anthropic
- Messages API reference — Anthropic
- Could not be confirmed from a primary source as of 2026-07-25: all architecture details — parameter count, dense or mixture-of-experts, expert/layer/head counts, vocabulary size, training compute. Anthropic publishes none and releases no weights, so no model card or
config.jsonexists. No benchmark numbers are quoted because none were verified in this session. The release date is not stated on any page read here. Microsoft Foundry availability for this specific model could not be confirmed and is marked Unverified above. The row this page was built from is corrected in one place: it described Opus 4.8 as "Superseded — Anthropic lists it as legacy." That is half right. The models overview does group it under a "Legacy models" accordion, but the model deprecations page — which defines the lifecycle terms — lists its current state as Active, not Legacy or Deprecated, with retirement not sooner than May 28, 2027 and no deprecation notice issued.
Price and capacity verified 2026-09-12 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-12 (research pass, primary source read today: https://platform.claude.com/docs/en/about-claude/pricing; https://platform.claude.com/docs/en/models/opus-4-8/overview — figures confirmed: input_per_m 5.0, output_per_m 25.0, cache_hit_input_per_m 0.5, context 1M, max_output 128K. Observed on the page, not added: Pricing row verbatim: 'Claude Opus 4.8 | $5 / MTok | $6.25 / MTok | $10 / MTok | $0.50 / MTok | $25 / MTok'. Model page: 'Context window1Mtokens', 'Max output128Ktokens', 'StatusActive (legacy)', 'RetirementNot sooner than May 28, 2027' - legacy status matches the row's 'for' text. Context/max_output are not on the pricing page itself.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $5 in / $25 out per M, $0.5 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-50846c; read of platform.claude.com/docs/en/about-claude/pricing — $5/$25 and the $0.50 cache read unchanged. All five Anthropic rows this file carries were re-read against the model-pricing table and none moved. Claude Mythos 5 sits on that table at $10/$50 and stays omitted here because it is not GA, per the fable-5 row's own note. Cause of the hash move not determined; no row this file carries moved.)
What changedWhat changed here
Updated this page A Claude Opus 5.5 model is now callable, which supersedes the Opus 5 and Opus 4.8 entries as the current Opus tier.
The Claude model pages and the costing board must be updated: a Claude Opus 5.5 now exists, so the Opus 5 and Opus 4.8 entries are no longer the current Opus tier and their pricing and status claims need revising.
Three kinds of claim, strongest first. Signal runs every morning.