Anthropic's cheapest and fastest current model, and the only one here still using manual extended thinking with an explicit budget_tokens value.
Why this oneWhat it is actually for
Haiku 4.5 is for volume. At $1/$5 per million tokens it is a fifth of Opus 5's rate, and the models overview describes it as "the fastest model with near-frontier intelligence." Reach for it when the unit of work is small and repetitive — routing, classification, extraction, high-turn chat — and you are running enough of them that per-call cost dominates. It is also the odd one out architecturally in the API sense: it is the only current Claude model on this board that still uses manual extended thinking, so you set a literal budget_tokens number rather than an effort level. If your codebase is standardised on output_config.effort, that call will fail here.
What it isArchitecture, lineage, training
Haiku 4.5 is closed-weight and API-only. Anthropic discloses no architecture: no parameter count, no dense-versus-mixture-of-experts statement, no expert, layer, attention-head or vocabulary figures, and no training-corpus or compute description. No weights are released, so there is no Hugging Face model card and no config.json to read real architecture numbers from.
The serving envelope is smaller than the rest of the current family, and deliberately so. The API model ID is claude-haiku-4-5-20251001, with claude-haiku-4-5 as the alias — note that this is a dated snapshot, the older naming convention, unlike the dateless pinned IDs used from the 4.6 generation onward. The context window is 200k tokens, not 1M, and max output is 64k rather than 128k. Reliable knowledge cutoff is February 2025 and training data cutoff July 2025 — by some margin the oldest of any model in this family.
Its generation shows in two places. It uses the previous tokenizer: the pricing page says the newer, roughly-30%-more-tokens tokenizer applies to "Claude 4.7 and later models," and that "Claude Sonnet 4.6 and earlier models use the previous tokenizer." The context tooltip agrees, describing 200k tokens as roughly 150k words — the same word-per-token ratio as Opus 4.6, not the denser ratio of Sonnet 5.
And its thinking mode is the legacy one. The per-model configuration table lists Haiku 4.5 as "Extended only," default off, with thinking.type: "adaptive" rejected by a 400. It also has the largest minimum cacheable prefix of any current model at 4,096 tokens.
At a glanceSee it
Haiku 4.5 uses manual budget_tokens thinking, and keeps the temperature controls newer models removed.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model ID | claude-haiku-4-5-20251001 (alias claude-haiku-4-5) | Models overview |
| Context window | 200k tokens | Models overview |
| Max output tokens | 64k | Models overview |
| Input modalities | Text, image | Models overview |
| Output modalities | Text only | Models overview |
| Reliable knowledge cutoff | February 2025 | Models overview |
| Training data cutoff | July 2025 | Models overview |
| Release date | Not stated in prose; the pinned snapshot ID encodes 2025-10-01 | Models overview (model ID) |
| Lifecycle state | Active; tentative retirement "not sooner than October 15, 2026" | Model deprecations |
| Weights available | No — closed API only | No vendor release exists |
| Licence | Not applicable; commercial API terms only | — |
| Parameter count | Not disclosed | — |
| Architecture (dense or MoE) | Not disclosed | — |
| Minimum cacheable prefix | 4,096 tokens — the largest of any current model | Prompt caching |
| Tool-use system prompt overhead | 496 tokens (auto/none); 588 tokens (any/tool) | Pricing |
| Comparative latency | "Fastest" | Models overview |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Selects the model | Required; "claude-haiku-4-5" or the dated "claude-haiku-4-5-20251001" | This model predates the dateless ID convention, so the alias resolves to the dated snapshot |
messages | Conversation turns | Required; up to 100,000 messages | First message must be user |
max_tokens | Hard ceiling on total output | Required; model max 64k | Half the Opus/Sonnet ceiling. Must exceed budget_tokens when thinking is enabled |
system | System prompt | String or array of text blocks | Front of the cache prefix |
thinking | Thinking configuration — manual mode only | {"type": "enabled", "budget_tokens": N}; default off | {"type": "adaptive"} returns a 400 with "adaptive thinking is not supported on this model". Omitting thinking runs no reasoning pass |
thinking.budget_tokens | Target thinking-token budget | Minimum 1,024; must be less than max_tokens | Below 1,024 the API rejects it. It is "a target rather than a strict cap" — max_tokens remains the hard ceiling. Changing it between requests invalidates the prompt cache |
thinking.display | Whether thinking blocks carry text | "summarized" or "omitted"; default "summarized" on this generation | Unlike the 4.7+ models, you get readable summaries without opting in |
output_config.effort | Token-spend control | Not supported on this model | The effort page's supported list omits Haiku 4.5. Use budget_tokens to control reasoning depth instead |
temperature | Sampling temperature | Supported — 0.0 to 1.0, default 1.0; incompatible with thinking | The only model here with a live temperature dial. Docs: closer to 0.0 for analytical, 1.0 for creative. Incompatible while thinking is on |
top_p | Nucleus sampling | Supported; while thinking is on, allowed only between 0.95 and 1 | "Recommended for advanced use cases only." Do not set alongside temperature |
top_k | Top-K truncation | Supported with thinking off; incompatible with thinking | Removes low-probability tail tokens |
stop_sequences | Custom halt strings | Array of strings | Supported |
tools / tool_choice | Tool definitions and forcing | auto (default), any, tool, none | With manual thinking enabled, only auto and none work — any and tool error, because forcing tool use conflicts with manual extended thinking |
output_config.format | JSON schema constraint on output | {"type": "json_schema", "schema": {...}} | Supported — structured outputs are GA for "Claude 4.5 and later models", and Haiku 4.5 is listed on Bedrock and Google Cloud too |
cache_control | Cache breakpoint | {"type": "ephemeral", "ttl": "5m"|"1h"}; max 4 | 4,096-token minimum — eight times Opus 5's. Short prompts silently do not cache and return no error |
inference_geo | Pins inference geography | Not supported — documented for Claude 4.6 and later | Requests including it on earlier models "return a 400 error" |
seed, frequency_penalty, presence_penalty, logprobs, n | Determinism, penalties, probabilities, multi-sample | Not supported | None exist on the Messages API. Repeat requests for multiple samples |
Three knobs matter, and two of them exist only here. thinking.budget_tokens is the reasoning-depth control — start near the 1,024 minimum for simple tasks and raise incrementally — and it has no effort equivalent on this model. temperature is genuinely live, unlike every other model on this page, but only while thinking is off. And cache_control deserves a look before you rely on it: the 4,096-token minimum is high enough that many routing and classification prompts fall under it and never cache, with no error to tell you.
SamplingShaping the output distribution
Haiku 4.5 is the only model in this family with a working sampling surface, and that is a direct consequence of its generation. The thinking documentation draws the line explicitly: the blanket 400 on non-default temperature, top_p and top_k applies to Fable 5, Opus 5, Opus 4.8, Opus 4.7 and Sonnet 5, while "on older models, the restriction applies only while thinking is on." Haiku 4.5 is an older model by that definition.
So with thinking off, you get the conventional dials. temperature runs 0.0 to 1.0 with a default of 1.0; the API reference advises values "closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks," while warning that even at 0.0 "the results will not be fully deterministic." top_p and top_k are both available and both flagged "recommended for advanced use cases only" — set one or the other, not both alongside temperature.
Turn thinking on and the rules tighten: temperature and top_k become incompatible, and top_p is permitted only between 0.95 and 1. A sensible start for extraction and classification is thinking off, temperature around 0.0, and nothing else set.
ReasoningThinking, effort and budgets
Haiku 4.5 uses manual extended thinking, the legacy mode, and it is the only current model on this board that does. The per-model table lists it as "Extended only," default off, with thinking.type: "adaptive" rejected by a 400 reading "adaptive thinking is not supported on this model." You enable it with thinking: {"type": "enabled", "budget_tokens": N}.
The budget rules are explicit: minimum 1,024 tokens, and it must be less than max_tokens, because thinking tokens count toward that ceiling. It is "a target rather than a strict cap" — the model may stop reasoning early, and max_tokens remains the hard limit. Anthropic suggests starting near 1,024 for simple tasks and 16,000-plus for complex ones. Thinking is billed as output tokens; usage.output_tokens_details.thinking_tokens reports how many.
Two limits are specific to this model: output_config.effort is not supported at all, and interleaved thinking is not supported — the beta header "is accepted but ignored."
ToolsFunction calling and server tools
Function calling works through the standard tools array, but there is a constraint here that does not apply to the adaptive-thinking models. With manual extended thinking enabled, tool use "only supports tool_choice: {"type": "auto"} (the default) or tool_choice: {"type": "none"}" — using any or a named tool "results in an error because these options force tool use, which is incompatible with manual extended thinking." If you need forced tool choice on this model, turn thinking off.
Tool overhead is the highest in the family: 496 system-prompt tokens at auto/none and 588 at any/tool, against Opus 5's 286/406. On a high-volume routing workload that fixed cost recurs on every call and is worth counting.
Structured outputs via output_config.format and strict: true are supported — the docs put structured outputs at GA for "Claude 4.5 and later models," and list Haiku 4.5 explicitly on Amazon Bedrock and Google Cloud. Parallel tool calls behave as elsewhere: return every tool_result in one user message.
CostPrice, caching, batching, what drives the bill
List price is $1 per million input tokens and $5 per million output tokens — the cheapest current Claude model.
| Line | Rate |
|---|---|
| Base input | $1 / MTok |
| Output (includes thinking tokens when enabled) | $5 / MTok |
| 5-minute cache write | $1.25 / MTok |
| 1-hour cache write | $2 / MTok |
| Cache hit / refresh | $0.10 / MTok |
| Batch API | $0.50 in / $2.50 out |
Two things shape the real bill in opposite directions. In your favour: this model uses the previous tokenizer. Anthropic documents the change in one direction only — Claude 4.7 and later models "produce approximately 30% more tokens for the same text," with the exact increase depending on content — so the same text tokenizes to roughly three-quarters of what it would on Sonnet 5 or the Opus tier (a ~23% reduction, the inverse of a 30% increase; Anthropic publishes no "fewer" figure of its own). That widens the effective price gap beyond the headline ratio, and it is the single most under-appreciated factor when comparing Haiku 4.5 against newer models. Measure it with the token-counting endpoint against both model IDs rather than applying a blanket multiplier.
Against you: the minimum cacheable prefix is 4,096 tokens, eight times Opus 5's 512 and four times Sonnet 5's 1,024. Many short classification and routing prompts sit below that line, so they never cache — and the API returns no error, it just processes them uncached. Verify with cache_read_input_tokens rather than assuming. Tool definitions add a further 496 to 588 input tokens per call. Anthropic's own worked example puts roughly 3,700 tokens per support-ticket conversation at about $37 per 10,000 tickets.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Claude API (first-party) | Yes | Model ID claude-haiku-4-5-20251001, alias claude-haiku-4-5 |
| Amazon Bedrock | Yes | Model ID anthropic.claude-haiku-4-5-20251001-v1:0; structured outputs GA |
| Claude Platform on AWS | Yes | Anthropic-operated; bare first-party model ID |
| Google Cloud / Vertex AI | Yes | Model ID claude-haiku-4-5@20251001 (note the @ version separator); structured outputs GA |
| Microsoft Foundry | Yes | Listed among the platforms current models are available through |
| Fast mode | No | Documented for Opus 5 and Opus 4.8 only |
inference_geo data residency | No | Documented for Claude 4.6 and later; earlier models return a 400 if the parameter is sent |
| Hugging Face Inference | No | Closed weights; no model card exists |
| Self-hosting | No | Weights are not released |
StrengthsWhat it is good at
- Cheapest current Claude model at $1/$5 per MTok, dropping to $0.50/$2.50 through the Batch API.
- Documented as the "Fastest" model in the family, with a description of "near-frontier intelligence."
- Uses the previous tokenizer, so identical text costs roughly 30% fewer tokens than on Sonnet 5 or the Opus tier — the real price gap is wider than the headline rates suggest.
- The only model here with live sampling controls:
temperature,top_pandtop_kall work when thinking is off. - Supports structured outputs and
strict: truetool schemas, and is GA for structured outputs on Bedrock and Google Cloud as well as the Claude API.
LimitsWhere it falls down
- Knowledge is badly dated — reliable cutoff February 2025, training data July 2025, roughly a year behind Sonnet 5 and fifteen months behind Opus 5.
- 200k context and 64k max output, a fifth and a half respectively of the current Opus and Sonnet tier.
output_config.effortis not supported at all, so code standardised on effort must special-case this model and usebudget_tokensinstead.- The 4,096-token cache minimum is the highest in the family; short prompts silently fail to cache with no error returned.
- With manual thinking enabled, forced tool choice (
anyor a namedtool) errors outright, and interleaved thinking is unsupported — the beta header is accepted but ignored.
Against its neighboursHow it compares
Against Claude Sonnet 5, the gap is much wider than the 2x price difference implies. Sonnet 5 gives five times the context (1M versus 200k), double the max output, adaptive thinking with the full low-to-xhigh effort ladder, and a January 2026 cutoff against Haiku's February 2025 — but partly claws the cost back through the newer tokenizer, which emits about 30% more tokens for the same text. Against Claude Opus 5 the comparison is barely a contest on capability; Haiku is the choice only when per-call cost is the binding constraint. The most honest framing is that Haiku 4.5 belongs to the previous API generation: manual budget_tokens, no effort parameter, live temperature, older tokenizer, dated snapshot ID. That makes it cheap and predictable, and it also means code written against it does not port cleanly upward.
Getting startedThe smallest call that works
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-4-5",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Classify this ticket as billing, technical, or other."}
]
}'For classification work, add "temperature": 0 first — this is one of the few current Claude models where that parameter still functions. If a task needs real reasoning, add "thinking": {"type": "enabled", "budget_tokens": 2048} and raise max_tokens above it; do not reach for output_config.effort or {"type": "adaptive"}, as neither works here.
SourcesWhere every claim above came from
- Models overview — Anthropic
- Pricing — Anthropic
- Extended thinking — Anthropic (budget rules and tuning)
- Thinking — Anthropic (sampling parameters, display defaults)
- Troubleshooting thinking — Anthropic (per-model configuration table)
- Effort — Anthropic (supported-models list)
- Prompt caching — Anthropic
- Structured outputs — Anthropic
- Model deprecations — Anthropic
- Messages API reference — Anthropic
- Could not be confirmed from a primary source as of 2026-07-25: all architecture details — parameter count, dense or mixture-of-experts, expert/layer/head counts, vocabulary size, training compute. Anthropic publishes none for Haiku 4.5 and releases no weights, so no model card or
config.jsonexists. No benchmark scores are quoted because none were verified in this session. The release date is not stated in prose on any page read here; "October 2025" is inferred only from the pinned snapshot IDclaude-haiku-4-5-20251001, and is presented as such rather than as a sourced claim. One clarification to the source row: it listed the reasoning mode as "Yes", which is true but imprecise — thinking here is manual extended thinking only, off by default, andthinking.type: "adaptive"returns a 400.
Price and capacity verified 2026-09-12 against https://platform.claude.com/docs/en/about-claude/pricing. re-read 2026-09-12 (research pass, primary source read today: https://platform.claude.com/docs/en/about-claude/pricing; https://platform.claude.com/docs/en/models/haiku-4-5/overview — figures confirmed: input_per_m 1.0, output_per_m 5.0, cache_hit_input_per_m 0.1, context 200K, max_output 64K. Observed on the page, not added: Pricing row verbatim: 'Claude Haiku 4.5 | $1 / MTok | $1.25 / MTok | $2 / MTok | $0.10 / MTok | $5 / MTok'. Model page: 'Context window200Ktokens', 'Max output64Ktokens', 'StatusActive (latest)', 'RetirementNot sooner than October 15, 2026' - that retirement floor is ~5 weeks out; not a deprecation notice, but worth watching. Context/max_output are not on the pricing page itself.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $1 in / $5 out per M, $0.1 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-27 (currency proposal src-2026-08-27-50846c; read of platform.claude.com/docs/en/about-claude/pricing — $1/$5 and the $0.10 cache read unchanged. All five Anthropic rows this file carries were re-read against the model-pricing table and none moved. Claude Mythos 5 sits on that table at $10/$50 and stays omitted here because it is not GA, per the fable-5 row's own note. Cause of the hash move not determined; no row this file carries moved.)