Home › Agents & Tool Use › Human-in-the-loop
🕹️ · Build

Human-in-the-loop

Pausing an autonomous agent to get a person's approval before a high-risk action executes.

In one line

The agent pauses before a consequential action and waits for a person to approve, edit, or reject it.

ConceptWhat it is

Human-in-the-loop (HITL) inserts an approval gate into an agent's action loop: instead of letting the agent run end to end, specific high-risk steps are marked as gated, and when the agent reaches one it stops and hands the proposed action to a person to approve, edit, or reject before anything happens.

It exists because autonomy and safety trade off directly. An agent that can send money, delete records, or publish to customers can also do real damage from a single bad decision. HITL bounds that blast radius by putting an accountable human on the small tail of actions that are irreversible or high-stakes, while everything cheap and reversible still runs automatically.

How it worksThe mechanics

The agent runs its normal reasoning loop until it selects an action; a policy classifies that action as low-risk or gated. Low-risk actions execute immediately. A gated action triggers an interrupt: the agent durably checkpoints its state, surfaces the proposed call and its reasoning to a reviewer, and blocks. The reviewer approves, edits the parameters, or rejects. On approval the agent resumes from the saved checkpoint and executes; on rejection it either revises its plan or halts, and the decision is logged for audit.

At a glanceSee it

Human-in-the-loop diagram
Human-in-the-loop diagram 1

The gate is not one yes/no switch but a policy funnel that sorts each proposed action into hard-block, human-approval, or auto-execute by reversibility, blast radius, and model confidence.

Human-in-the-loop diagram 2

What the gate does when no one answers — an SLA timer, escalation to a backup approver, and a deliberate fail-open versus fail-closed default that decides whether a timed-out action ships or cancels.

When to use itWhere it fits

  • Irreversible or costly actions: moving money, deleting data, executing trades, publishing to customers.
  • New agent deployments still earning trust before their autonomy is widened.
  • Regulated domains where an accountable human sign-off is legally or contractually required.
  • Actions whose blast radius exceeds what an automated rollback could cleanly undo.

When NOT to use itLimits & anti-patterns

  • High-frequency, low-stakes steps where a pause per action makes the agent unusably slow.
  • Fully reversible actions that a cheap automated undo already covers.
  • When no reviewer can keep pace with the agent's volume, so approvals just pile into a backlog.
  • Narrow tasks with a long, well-measured track record of reliable autonomous performance.

Trade-offsAdvantages & costs

Advantages
  • Bounds worst-case blast radius on the few actions that can actually cause harm.
  • Lets teams ship autonomy sooner by gating only the risky tail rather than everything.
  • Approvals and edits become an audit trail and a labeled dataset for later tuning.
  • Keeps an accountable human in the decision, satisfying many policy and compliance requirements.
Trade-offs & costs
  • Adds human latency, so throughput is capped by reviewer availability, not by compute.
  • Does not scale: every gated action consumes real staff time as volume grows.
  • Alert fatigue erodes review quality, and tired reviewers start rubber-stamping.
  • Mis-tuned gating either floods reviewers or lets genuinely risky actions slip through unreviewed.

ExampleIn the real world

A customer-support agent handles refunds on its own, but any refund above a set dollar amount is gated. When the agent proposes one over the limit, the run pauses, its state is checkpointed, and the proposed refund plus the agent's reasoning is posted into a review channel. A support lead approves it, edits the amount, or rejects it; only on approval does the agent resume and call the payments API. Refunds under the limit never pause, so the gate slows only the small fraction of cases where a mistake would be expensive to reverse.

ToolsHow to implement it

  • LangGraphbuilt-in interrupt and resume with checkpointed state for human-approval nodes.
  • Microsoft AutoGenUserProxyAgent human_input_mode to request human approval mid-run.
  • Temporaldurable workflows that block on a signal until a human responds, then continue.
  • Amazon Augmented AI (A2I)managed routing of low-confidence steps to human reviewers.

Cost & effortWhat it takes

The dominant cost is human wall-clock time, not tokens: a gated action can sit idle for minutes to hours waiting on a reviewer, so latency and staffing scale with the number of gated actions rather than with compute. Engineering effort is moderate — the agent must durably checkpoint and resume so a pause does not lose context, plus a review surface and routing rules to decide what gets gated. Extra token spend is negligible, limited to summarizing the proposed action for the reviewer.

A living map of modern AI — kept current every morning