Prompt chaining breaks a complex task into a sequence of simpler prompts, each handing its output to the next.
ConceptWhat it is
Prompt chaining decomposes a single complex request into a pipeline of smaller prompts, where the output of one call becomes the input to the next — for example extract → summarize → format. It exists because a model asked to do many things at once tends to do each of them worse: instructions compete for attention, the intermediate reasoning is hidden and untestable, and one weak sub-skill drags down the whole answer.
By making each step do one job, chaining turns an opaque monolith into a sequence of stages you can inspect, test, and swap independently. Each prompt is simpler to write and easier to evaluate on its own, and you can route different steps to different models or drop validation gates between them. The cost is that you now own an orchestration problem: multiple calls to sequence, intermediate state to pass along, and errors that can propagate down the chain when an early step goes wrong.
How it worksThe mechanics
You break the task into ordered stages and write a focused prompt for each. The first prompt runs against the raw input and returns an intermediate result, often as structured JSON so the next step can consume it reliably. An orchestration layer takes that output, optionally validates or reshapes it, and injects it into the second prompt as context; the pattern repeats until the final stage emits the finished answer. Between any two links you can insert a check that retries or repairs a bad output before it flows onward, and because each call is separate you can pick the cheapest model that clears each step's bar.
At a glanceSee it
The existing pipeline is only the sequential shape — chaining is really a family of topologies (router, parallel fan-out, refinement loop) that real systems nest inside one another.
A decision test for when to chain at all — split only when the task decomposes, the intermediates need inspecting, and the extra latency is worth paying.
When to use itWhere it fits
- The task has natural sequential stages — extract, then reason, then format — that each need a different skill.
- Intermediate outputs must be inspected, validated, or logged, such as structured data checked before a final write.
- A single prompt has grown unreliable and unwieldy, with quality dropping as you pile on more instructions.
- You want to route stages to different models — a cheap one to classify, a stronger one to reason.
When NOT to use itLimits & anti-patterns
- The task is simple enough for one well-structured prompt; chaining just adds latency and moving parts.
- Latency is tight and every extra round trip hurts, since each link is another network call.
- The steps are genuinely interdependent and must be reasoned about jointly rather than in a fixed sequence.
- The control flow is dynamic and tool-driven at runtime — that is an agent's job, not a static chain.
Trade-offsAdvantages & costs
Advantages
- Each step is simpler, so prompts are easier to write, debug, and unit-test in isolation.
- Intermediate outputs are visible, enabling validation, logging, and targeted fixes at a single stage.
- You can mix models — cheap ones for easy stages, strong ones where it counts — to control cost.
- Failures localize to a specific link, making the pipeline easier to reason about and improve.
Trade-offs & costs
- More model calls mean higher aggregate latency and token cost than a single combined prompt.
- Errors propagate: a wrong output early in the chain silently corrupts everything downstream.
- You take on orchestration, state-passing, and glue code that a single prompt does not need.
- Rigid, hand-wired sequences can be brittle when inputs vary in unexpected ways.
ExampleIn the real world
A team automating support-ticket triage splits the job into four linked prompts instead of one mega-prompt. The first call classifies the ticket's intent and urgency; the second extracts structured fields such as order ID, product, and error message into JSON; the third retrieves the matching policy and drafts a reply grounded in it; the fourth rewrites that draft into the required tone and template. Because each stage is isolated, the team can unit-test the extractor against a fixed set of tickets, swap a cheaper model into the classification step, and add a validation check that rejects malformed JSON before it reaches the drafting stage — none of which is practical when all four jobs are crammed into one prompt.
ToolsHow to implement it
- LangChain (LCEL)the pipe operator composes prompt, model, and parser steps into an explicit chain.
- LangGraphgraph-based orchestration with branching, retries, and shared state between steps.
- DSPydeclarative pipelines where each step is a module whose prompts can be optimized.
- Haystacka pipeline framework for chaining retrieval and generation components into stages.
Cost & effortWhat it takes
The dominant cost is that a chain of N steps means roughly N model calls, so token spend and end-to-end latency scale with chain length rather than staying flat. Engineering effort shifts away from wrestling one giant prompt and toward orchestration: sequencing calls, passing and validating intermediate state, and adding observability so you can see which link failed. That trade is usually worth it once a task is complex enough that a single prompt is unreliable, but it is pure overhead for anything a well-written single prompt already handles.