Home › What Happens After You Hit Enter › Step limits, budgets and loop detection
Pipeline stage · Operate

Step limits, budgets and loop detection

The hard caps — maximum steps, tokens, spend, wall-clock — plus the check that spots an agent going in circles.

In one line

The model has no idea what step it is on or what it has spent; every budget you believe it respects is a counter in your loop.

Why you'd careThe thing you have already noticed

You have watched an agent call the same command with the same arguments six times, each time reading the same error, each time apologising and trying again. Or you have found a run that stopped mid-task with no explanation and a stop reason your interface never rendered. Both are this stage. The model does not know it is on step 24; each turn it sees a transcript and predicts the next action, and a transcript full of failed attempts is excellent evidence that another attempt is coming. The counter that ends it lives entirely in your orchestrator, and where you place the limit decides whether your failure mode is a surprising bill or a truncated answer.

In and outWhat goes in, what comes out

InA live trajectory: the message list, plus accounting the orchestrator maintains beside it — step index, cumulative input, output and cached-token counts, wall-clock elapsed, money spent, and a rolling fingerprint of recent tool calls.
ProcessAfter each step, compare every counter against its ceiling. Hash the last call's name plus normalised arguments and compare against the previous k fingerprints. Classify the run as continue, warn or halt. A warning is text injected into context; a halt is a control-flow decision the model never sees.
OutOne of three outcomes: the next model call proceeds unchanged; a nudge message is appended stating the remaining budget or forbidding a repeated call; or the loop exits with a terminal reason, whatever partial result exists, and a usage record.

Nothing about the conversation is destroyed here, but the reason a run ended is routinely lost. A hard halt happens outside the model, so there is no final assistant turn explaining the stop; unless your interface renders the terminal reason, the user just sees an answer that stops. The budget is equally invisible in the other direction: the model cannot see any counter you do not write into the prompt.

ConceptThe idea underneath

There is no neural network in this stage. It is a watchdog — counters, ceilings and a comparison, the same pattern as a request timeout or a circuit breaker. Saying so plainly matters, because the standard mistake is to reach for prompting: telling a model to be efficient does not create a counter.

The machine-learning content sits one level away, in why loops happen at all. An autoregressive model conditions on its own previous output. A transcript in which the same command has just failed three times is strong evidence that the next thing to produce is a fourth attempt; the pattern is self-reinforcing in exactly the way token-level repetition loops are, one level up the stack. This is why raising temperature does not reliably break an agent loop. The cause is in the prefix, not in the sampler, so the fix must change the prefix: inject a message naming the repetition, or delete the failed attempts from context entirely.

The systems idea worth internalising is that the four ceilings are in different units, and only one of them is the unit you actually care about. Steps, tokens, seconds and money are not interchangeable, because cost per step is not constant: the trajectory grows, so step 40 re-sends everything step 1 sent plus 39 steps of tool output on top. A run capped at 50 steps over a 100k-token context pushes several million input tokens even though it only took 50 steps. Whether that is expensive or trivial depends almost entirely on whether your cache breakpoints held, since cached input is billed at a fraction of the base rate — check current pricing rather than assuming a ratio. Cap the currency you care about and derive the rest from it.

At a glanceSee it

Step limits, budgets and loop detection diagram

Where step, token, cost and clock ceilings sit around the agent loop, plus the repeat check.

The knobsHyperparameters and nuance

  • max_turnsOpenAI Agents SDK, passed to Runner.run and defaulting to 10. Exceeding it raises rather than returning a partial answer, so a bare max_turns with no handler converts a long task into an exception instead of a degraded but useful result.
  • recursion_limitLangGraph's per-invocation step ceiling, default 25. It counts graph super-steps rather than model calls, so a graph with several nodes per turn hits the wall far sooner than the number suggests.
  • max_iterations and early_stopping_methodLangChain's AgentExecutor. The second is the interesting one: "force" returns a fixed string when the ceiling is hit, while the generating alternative spends one more model call summarising what was found, which is nearly always the better answer. Check which values your version still supports.
  • cumulative token and cost ceilingsnot provided by default in any SDK; you sum usage.input_tokens, usage.output_tokens and the cache fields yourself. Without them a per-call max_tokens looks reassuring while the aggregate is unbounded, since it is context growth rather than output length that dominates.
  • wall-clock deadlinethe only ceiling that bounds a run whose tools hang rather than return. Set it below the patience of whatever is upstream, such as a browser or a queue's visibility timeout, or you will get duplicate runs stacked on top of the runaway one.
  • repeat threshold khow many identical fingerprints trigger the loop check. At k = 2 it fires on legitimate retries of a flaky API; at k = 6 a tight loop burns real money first. Normalise arguments before hashing, or a changing page offset hides the loop completely.

EffectHow this stage moves the answer

Where you put the ceiling decides which half of the answer the user loses. A hard halt at max steps usually lands before synthesis, so the user receives the agent's scratch work — a list of files it opened, a half-applied migration — and none of the conclusion that work was for. Budget nudges change behaviour more subtly: tell a model it has three steps left and it stops exploring and commits, producing a real answer with weaker evidence behind it. Tight budgets systematically bias toward shallow single-tool responses that read as confident precisely because the model never encountered the contradicting source. And the halt policy is itself an answer-quality decision: spending one extra model call to summarise what was established turns a truncated run into a usable partial report, which is the difference between "here is what I found and what is still open" and a blank.

EvalsWhat it does to your measurements

The numbers that matter are runaway incidents per thousand runs, p99 cost per task, truncation rate and success-under-budget. The measurement hazard is that a step ceiling is a property of the harness rather than the model, and almost nobody reports it. The same model on the same SWE-bench-style suite scores materially differently at 10 steps versus 50, so a comparison across two harnesses is not a comparison of models. Loop detection adds a second hazard: a heuristic that halts a legitimate paginated crawl records a model failure that never happened, and it stays invisible unless you log terminal reasons and read them. Report cost per solved task rather than cost per run, or you will reward agents that give up early and quietly punish the ones that finish. Publish the ceilings next to the score.

Failure modesWhen it goes wrong

  • The same call repeated until the ceilingthe failed attempts are still in context and condition the next one; resampling cannot escape a loop whose cause is the prefix.
  • A four-figure bill from one malformed tool responseunbounded retries with no cost ceiling on a trajectory that grew every step, so the last retries cost many times what the first did.
  • The run ends mid-sentence with no explanationa hard halt with no final synthesis call, and a terminal reason the interface does not surface.
  • Legitimate pagination flagged as a loopthe detector hashed the tool name and ignored the changing offset argument, or normalised it away entirely.
  • Per-step accounting looks healthy, the total does notbudgets were checked per call rather than cumulatively, missing that input tokens grow with every step.

PapersWhere this comes from

There is no research literature for this stage, and claiming otherwise would be dishonest: step ceilings, cost budgets and repeat detection are plain engineering, reinvented independently in every agent harness. What exists instead is reference implementations and their defaults — LangGraph's recursion_limit, the OpenAI Agents SDK's max_turns, and the SWE-bench agent harnesses whose published step and time limits are part of every reported score. The adjacent research is about how much compute to spend rather than how to stop spending it: Snell et al., 2024, on optimally scaling test-time compute, shows that additional inference-time steps buy accuracy with sharply diminishing returns on some task classes, which is the closest thing to a principled argument for where a ceiling belongs. Loop detection itself borrows nothing from machine learning; the useful prior art is watchdog timers and circuit breakers.

A living map of modern AI — kept current every morning