A cheap safety wrapper that caps how many steps, dollars, and seconds an agent may spend so it always stops instead of looping forever.
ConceptWhat it is
A bounded loop is a control wrapper placed around any agentic loop that enforces hard ceilings on steps, cost, and wall-clock time, with a defined behavior when a ceiling is hit. Autonomous agents choose their own next action, which means a faulty plan, a flapping tool, or a model that keeps "almost" finishing can spin indefinitely, burning tokens, hammering downstream APIs, and running up a bill with nothing to show for it.
The pattern exists to make that failure impossible rather than merely unlikely. Instead of trusting the model to know when to quit, the surrounding code counts iterations, sums token or dollar spend, watches a timer, and caps retries on transient errors. When any budget is exhausted the loop terminates deterministically and returns a partial result, an error, or a human escalation. It is a low-autonomy safeguard, not a strategy: it does not make the agent smarter, only bounded.
How it worksThe mechanics
Before the loop starts, the wrapper initializes counters for iterations, accumulated cost, and a start timestamp, plus a per-call retry budget. Each turn the agent proposes an action; the wrapper executes it, increments the step count, adds the call's tokens or cost, and checks elapsed time. If the goal check passes, it returns the result. Otherwise it tests every budget in turn: max steps, max cost, max duration, and max retries, continuing the loop only if all remain within limits. If any is breached it breaks out and returns whatever it has, optionally flagged for human review.
At a glanceSee it
What the loop flattens into one within-budget test is actually three independent meters — steps, cost, and wall-clock — ORed so any single ceiling halts the run.
Hitting a ceiling is not one behavior but a configured ladder — hard kill, graceful partial, human escalation, or a cheaper finish — each logged so the ceilings can be tuned.
When to use itWhere it fits
- Any agent that ships to production and can call tools or itself in a loop, where an unbounded run could cost money or block a request.
- Workflows built on pay-per-call tools or model usage, where a runaway loop translates directly into an unbounded bill.
- User-facing latency budgets: a chat or API endpoint that must answer within seconds and cannot wait on an open-ended agent.
- Anywhere retries on flaky tools are needed but must not stack into an effectively infinite backoff.
When NOT to use itLimits & anti-patterns
- As your only safety mechanism: it caps quantity, not correctness, and will happily return a confidently wrong but cheap answer.
- For legitimately long or open-ended research and batch jobs where a hard step cap would truncate real progress; prefer checkpointing and resumable state.
- When limits are guessed rather than measured, since caps set too low silently degrade good tasks, which is often worse than an occasional overrun.
- When a mid-action cut could leave external side effects half-applied and there is no idempotency or rollback to make an abrupt halt safe.
Trade-offsAdvantages & costs
Advantages
- Deterministic worst case: you can state the maximum cost, latency, and step count of any run before it happens.
- Cheap and simple: a handful of counters and comparisons, with no extra model calls.
- Framework-agnostic; it wraps any loop regardless of the agent architecture inside it.
- Turns catastrophic runaway failures into graceful, observable early exits you can alert on.
Trade-offs & costs
- Can cut off a task that was one step from succeeding, producing a partial or empty result.
- Choosing good limits requires measuring real workloads; static caps are a blunt instrument.
- Says nothing about output quality: a bounded agent can still be wrong, just not expensive.
- A hard cut mid-action can leave external side effects half-applied unless paired with idempotency or rollback.
ExampleIn the real world
A support agent answers billing questions by calling an orders API and a knowledge base. Most queries resolve in two or three tool calls, but a malformed account ID makes the orders lookup return an ambiguous error that the model keeps re-querying, convinced the next call will clarify. A bounded loop caps the agent at 8 steps, a 20-second wall-clock limit, roughly ten cents of model spend, and 2 retries per tool. On the bad ID the agent hits the step ceiling, breaks out, and returns "I could not verify this account, escalating to a human," instead of looping for minutes and spending dollars on one ticket. The same caps sit invisibly under thousands of healthy requests that finish well inside them.
ToolsHow to implement it
- LangChain AgentExecutorexposes
max_iterationsandmax_execution_timeto bound the reasoning loop directly. - LangGraphenforces a
recursion_limiton graph traversal and supports step timeouts and interrupts for time and human-in-the-loop bounds. - OpenAI Agents SDKcaps a run with a
max_turnsparameter that raises an error when exceeded. - Tenacityand similar retry libraries bound per-call attempts and backoff, complementing the outer step and time caps.
Cost & effortWhat it takes
This is one of the cheapest patterns to run: the guardrails themselves are plain arithmetic, incrementing counters, summing token usage, and comparing a clock, so they add negligible latency and zero extra model calls. Its entire purpose is to reduce spend by capping the worst case. Engineering effort is modest: instrument the loop with counters, thread a cost estimate through each call, and decide the exit behavior (partial result, error, or human handoff). The real work is not the code but calibration, measuring how many steps and how much time healthy tasks actually take so the ceilings sit safely above the normal case without letting a runaway go unbounded.