Agents let an LLM plan, use tools, and act in a loop instead of just answering a single prompt.
ConceptWhat it is
An AI agent is a system where a language model plans a sequence of steps, calls external tools or APIs, observes the results, and decides what to do next, repeating until a goal is reached. It represents the most recent stage in the road to LLMs, moving models from single-turn text generators to autonomous actors.
It exists because many real tasks, like booking travel or debugging code, require multiple steps and interaction with the outside world, not just a single generated response.
How it worksThe mechanics
The model receives a goal, reasons about what action to take, calls a tool such as a search API, code interpreter, or database, reads the tool's output back into its context, and loops through plan-act-observe steps until it decides the task is complete or it hits a stopping condition.
At a glanceSee it
A single agent is really a family of control patterns — the shape of the task, not the model, decides which one fits.
The raw loop must be bounded — every step first clears a budget check and a repeat check that stop a runaway agent before it overspends or stalls.
When to use itWhere it fits
- Multi-step tasks needing tool use, like researching, coding, or booking actions.
- Workflows where the next step depends on the outcome of a previous action.
- Automating tasks that a human would otherwise do by clicking through multiple systems.
- Coding assistants that need to run and check code before answering.
When NOT to use itLimits & anti-patterns
- Single-turn question answering where one model call suffices.
- High-stakes irreversible actions without human approval in the loop.
- Tasks where a fixed, deterministic workflow is more reliable than open-ended planning.
Trade-offsAdvantages & costs
Advantages
- Automates genuinely multi-step, tool-dependent work.
- Can recover from errors by observing and re-planning.
- Extends a model's capability far beyond its training data via tools.
- Composable with other agents for complex workflows.
Trade-offs & costs
- Can loop, stall, or take unintended actions without guardrails.
- Harder to test and debug than single-turn prompting.
- Latency and cost compound across many tool calls.
- Requires careful sandboxing for safety and cost control.
ExampleIn the real world
GitHub Copilot's agent mode and Devin from Cognition both plan, write, run, and debug code autonomously across multiple files rather than just suggesting single-line completions.
ToolsHow to implement it
- LangGraphframework for building stateful, controllable agent workflows.
- AutoGPTearly open-source autonomous agent experiment.
- OpenAI function calling and Assistants APInative tool-calling support for agents.
- Model Context Protocol, MCPstandard for connecting agents to external tools.
Cost & effortWhat it takes
Cost and latency scale with the number of tool calls and reasoning steps, often several times a single LLM call; moderate to high engineering effort to build reliable guardrails and stopping conditions.