Home › What Happens After You Hit Enter › Reasoning token budgets and effort control
Pipeline stage · Operate

Reasoning token budgets and effort control

Telling the model how long it is allowed to think before it answers.

In one line

The budget is a ceiling, not a target: models stop thinking early on easy problems, so your average cost is set by the hard tail of your traffic.

Why you'd careThe thing you have already noticed

You enabled extended thinking, left max_tokens at the 1024 you had always used, and got back a long reasoning block and no answer at all — the stop reason said you hit the token limit. Or you shipped high reasoning effort everywhere, and the following month a summarisation endpoint that used to cost fractions of a cent per call was spending thousands of thinking tokens deciding how to summarise three paragraphs. Both are this stage. Reasoning tokens are generated tokens: they are billed as output, they run through the same decode loop at the same speed, they consume KV cache, and on most APIs they count against max_tokens alongside the answer rather than instead of it.

In and outWhat goes in, what comes out

InA request carrying either a discrete effort level or an explicit thinking budget, the assembled prompt, and a model post-trained to emit a delimited reasoning segment. On Anthropic's API this is thinking.budget_tokens; on OpenAI-compatible reasoning models it is reasoning_effort with named levels.
ProcessThe model decodes reasoning tokens inside a delimited segment while the server counts them. When the counter reaches the budget, the serving stack forces the transition out of the segment — typically by injecting the end-of-thinking delimiter — and decoding continues into the answer.
OutA reasoning segment of variable length bounded by the budget, then the answer tokens. Usage accounting reports reasoning tokens separately on some APIs and folds them into output tokens on others. The reasoning text may be returned raw, summarised, encrypted or withheld.

The answer carries forward; the reasoning usually does not. Most providers strip prior thinking blocks from the context on the next conversational turn, so the model re-derives work it has already done — a cost effect and a quality effect at once. Where a provider returns signed or encrypted thinking blocks that must be passed back verbatim to continue a tool-use loop, as Anthropic does, dropping them breaks the continuation outright.

ConceptThe idea underneath

Nothing new happens inside the network. A reasoning model is one post-trained, typically with reinforcement learning against verifiable outcomes, so that its high-probability continuation after a question is a long stretch of self-checking text rather than an immediate answer. The chain works as external scratch memory: an intermediate result written as tokens can be read back exactly by attention, whereas a single forward pass has only d_model dimensions of activation to hold it in. Depth in tokens substitutes for depth in layers, and unlike layers you can buy more of it at request time.

Test-time compute scaling has two axes. Sequential: one longer chain. Parallel: many samples plus a selection rule, which Brown et al. characterised as a coverage-versus-selection problem — sampling more nearly always finds a correct answer, but only if you can identify it. Snell et al. showed the optimal split between the two depends on problem difficulty, and that inference compute can substitute for parameters, but not uniformly: easy problems are saturated by a small budget while hard ones keep improving.

Budget control at decode time is cruder than it sounds. Muennighoff et al. demonstrated budget forcing: suppress the end-of-thinking delimiter and append a continuation word like “Wait” to buy more thinking, or inject the delimiter to cut it off. An effort dial at the low end is close to exactly this.

The accuracy-versus-tokens curve is concave, and on easy problems it is non-monotonic. Past a point the model revisits a correct early answer, finds a reason to doubt it, and changes it. That is a real, repeatable phenomenon in production, not a rounding artefact — more thinking is not monotonically better.

At a glanceSee it

Reasoning token budgets and effort control diagram

An effort level sets a ceiling on reasoning tokens before the model is forced into its answer.

The knobsHyperparameters and nuance

  • reasoning_effort(OpenAI-compatible reasoning models) — discrete levels, with which ones exist and which is the default both depending on the model: current models accept some subset of none, minimal, low, medium, high, xhigh and max. Low under-thinks genuinely multi-step problems; the top levels can multiply cost by an order of magnitude and push latency into minutes on a request that used to take two seconds.
  • output_config.effort(Anthropic) — low, medium, high, xhigh, max, defaulting to high, and the current Anthropic depth control. Unlike a thinking budget it shapes all output tokens — prose, tool calls and thinking alike — so lowering it also means fewer tool calls, not just a shorter chain. It pairs with thinking type adaptive, under which the model decides per request whether to think at all.
  • thinking.budget_tokens(Anthropic, manual mode) — an integer with a documented minimum of 1024 that must be smaller than max_tokens; the one exception is interleaved thinking, where the budget spans every thinking block in the turn and may exceed it. Check your model before reaching for this: it is deprecated on the Claude 4.6 generation and returns a 400 on Claude 4.7 and later, which use effort instead. Too small and the model's chain is cut mid-derivation; too large and you have simply raised a ceiling you rarely reach.
  • max_tokensmust cover thinking plus the visible answer, and it is the hard ceiling where a budget or an effort level is only guidance. Sizing it for the answer alone is the single most common bug when a team first enables reasoning, and it presents as an empty response rather than as an error.
  • streamwith a large budget, a non-streaming request can sit silent long enough to trip a load balancer or gateway idle timeout. Anthropic's SDKs enforce this client-side and refuse a non-streaming request when max_tokens exceeds 21,333; above roughly 32k thinking tokens per request the documented advice is to move to batch processing rather than hold a connection open.
  • temperature / top_p / top_kprovider-specific, load-bearing, and split across two eras on Anthropic. On the current models (Claude Opus 4.7 and later, Claude Sonnet 5 and the other 5-generation models) any non-default temperature, top_p or top_k returns a 400 on every request, thinking or not. On older thinking-capable models the restriction applies only while thinking is on: temperature and top_k are incompatible with thinking, while top_p is allowed between 0.95 and 1. Several other vendors' reasoning models ignore the parameter entirely. Carrying over values tuned for a non-reasoning model is at best ignored and at worst a hard error.
  • interleaved thinking / thinking-with-tools flagscontrols whether the model may reason between tool calls rather than only before the first one. On Anthropic this is the interleaved-thinking-2025-05-14 beta header, and only for manual-mode thinking on Claude 4.5 and earlier Claude 4 models; adaptive thinking on newer models interleaves automatically with no header, and the header is accepted-but-ignored there. Changes the shape of an agent loop substantially.

EffectHow this stage moves the answer

What improves is narrow and real: multi-step arithmetic, code that must satisfy several constraints simultaneously, and planning across a long tool chain. The improvement shows up as fewer confidently wrong answers rather than as better prose — the writing quality barely moves. What degrades at high budgets is more than latency. Long chains tend to produce longer, more hedged final answers, and on simple factual or formatting requests the model relitigates your instructions and sometimes talks itself into deviating from them. Under-budgeting is worse and stranger: when the reasoning segment is truncated mid-derivation and the answer is generated from an incomplete chain, the model states a conclusion drawn from work it never finished. From the outside that is indistinguishable from a hallucination, and it is the failure mode most likely to be misattributed to the model rather than to your configuration.

EvalsWhat it does to your measurements

This is the dominant confound in reasoning-model comparisons, and it is routinely omitted. An accuracy number without the token spend that produced it is not interpretable: a model that wins at 30,000 thinking tokens has not beaten one that ties at 3,000. Report accuracy-versus-tokens curves, not points. Three specific ways an eval run gets silently invalidated. First, a harness with a max_tokens cap truncates high-effort runs and scores them wrong, so “high effort performs worse” is frequently a harness artefact rather than overthinking. Second, if reasoning tokens are excluded from the reported output-token count, your cost-per-correct-answer figure is wrong by an order of magnitude. Third, pass@1 on small competition-style benchmarks has enormous variance at n=1; the honest protocol is many samples with a reported spread. An eval at default effort is comparing whatever each vendor chose as default, which is a product decision, not a capability measurement.

Failure modesWhen it goes wrong

  • Empty or truncated answer with a length stop reasonmax_tokens was sized for the answer, not for thinking plus answer, and the budget consumed the whole allowance.
  • Bill up ten to forty times after enabling reasoninga single effort setting applied across every endpoint, including classification and formatting calls that never needed a chain.
  • Gateway 504 on long requestsa non-streaming request with a large thinking budget exceeding a proxy or load-balancer idle timeout, with no partial output to keep the connection alive.
  • Accuracy drops on easy questions at high effortthe model revisits a correct early answer and revises it; more sequential compute is not monotonically better.
  • A tool loop breaks after the first turnprior thinking blocks were not passed back where the provider requires it, or a middleware layer forwarded only the visible message content and dropped them.

PapersWhere this comes from

  • Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsWei et al., 2022 (arXiv:2201.11903). Established that eliciting intermediate steps improves reasoning at a fixed model size and fixed weights. Everything on this page is downstream of that result being turned into a trained-in behaviour with a dial on it.
  • Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersSnell et al., 2024 (arXiv:2408.03314). Showed that inference compute can substitute for parameters, and that the optimal allocation is difficulty-dependent — which is precisely why a single fixed effort level across your traffic is the wrong shape.
  • s1: Simple Test-Time ScalingMuennighoff et al., 2025. Introduced budget forcing: controlling chain length at decode time by suppressing or injecting the end-of-thinking delimiter. It is the clearest published account of what an effort dial actually does mechanically.
  • Large Language Monkeys: Scaling Inference Compute with Repeated SamplingBrown et al., 2024. Characterised the parallel axis: coverage rises steeply with repeated sampling, but the gain is only realisable when you have a verifier to select with. It is the reason parallel and sequential budgets are not interchangeable.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A study of how reasoning-effort terms function as model-specific API contract clauses that affect what you pay.

    arXiv cs.AI · 18 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning