Decompose a complex problem into a chain of simpler sub-problems and solve them in order, feeding each answer forward into the next.
ConceptWhat it is
Least-to-most prompting is a reasoning technique that splits a hard problem into a sequence of easier sub-problems and solves them in order, so that the answer to each step becomes an input to the next. It was introduced to improve compositional generalization — solving problems built from familiar steps combined in new or longer ways than the examples demonstrate — where plain chain-of-thought tends to stumble because it tries to reason through the whole thing in a single pass.
The key idea is to separate decomposition from solution. First the model (or a fixed prompt template) lists the sub-questions that build toward the final answer; then it works through them one at a time, carrying prior answers along as context. Because each step is small and grounded in explicit earlier results, the model handles tasks whose difficulty comes from their compositional structure rather than from any single hard leap.
How it worksThe mechanics
Step one, prompt the model to decompose the query into an ordered list of sub-problems, simplest first. Step two, solve the first sub-problem in isolation. Step three, append that sub-problem and its answer to the context and ask the model to solve the next one, and repeat down the list so every step can reference all prior answers. Step four, once the final sub-problem is solved its answer is the solution, or a short compose step stitches the intermediate results together. Decomposition and solving can live in one prompt or run as separate calls; the defining move is that later steps consume earlier answers rather than re-deriving them.
At a glanceSee it
The mechanism is two prompting stages — one call decomposes the problem into ordered subquestions, then a second prompt solves each in turn with every prior question and answer appended to the running context.
A routing view of when the extra stage pays off — reach for least-to-most when the test problems run longer than the demonstrated examples and steps chain on earlier answers, otherwise plain chain-of-thought is enough.
When to use itWhere it fits
- Compositional tasks where the final answer genuinely depends on intermediate results — multi-step math, multi-hop questions, or symbolic manipulation.
- Problems that are longer or deeper than your few-shot examples, where you need the model to generalize beyond the demonstrated length.
- Workflows where you want each intermediate answer to be inspectable, cacheable, or independently verifiable.
- Cases where single-pass chain-of-thought produces plausible-looking but wrong reasoning because it skips or conflates steps.
When NOT to use itLimits & anti-patterns
- Simple or single-fact queries that a direct answer or one chain-of-thought pass already handles — the extra steps add latency for no gain.
- Problems that do not decompose cleanly into an ordered sequence, where forcing a split just invents artificial sub-questions.
- Tight-latency or high-volume paths where multiple sequential calls per request are too slow or costly.
- Tasks where the hard part is one atomic leap of insight rather than a chain of easy steps — decomposition does not make the leap easier.
Trade-offsAdvantages & costs
Advantages
- Handles compositional problems and generalizes to harder, longer instances than the examples shown.
- Each step is small and grounded in explicit prior answers, which reduces skipped-reasoning errors.
- Intermediate sub-answers are visible, making failures easier to localize and debug than one opaque chain.
- Steps can be cached, validated, or swapped independently, which suits pipeline-style engineering.
Trade-offs & costs
- Multiple sequential steps mean higher latency and token cost than a single-shot prompt.
- A bad decomposition derails everything downstream — errors in early sub-problems compound.
- Choosing the right sub-problems can itself be hard and may need task-specific prompt design.
- Overkill for simple tasks, adding orchestration complexity with no accuracy benefit.
ExampleIn the real world
A billing support assistant gets the question, "If I upgrade mid-cycle, what will my prorated charge be and when does it next bill?" Instead of answering in one leap, it decomposes: (1) how many days remain in the current cycle, (2) the daily rate difference between the old and new plan, (3) the prorated charge, which is days remaining multiplied by the rate difference, and (4) the next bill date. It solves (1) to get, say, 12 days, then reuses that in (3) together with the rate difference from (2) to compute the charge, and finally reports the bill date from (4). Because step (3) explicitly consumes the answers to (1) and (2), the arithmetic stays anchored to concrete intermediate values rather than being guessed in a single pass, and a reviewer can check each figure on its own.
ToolsHow to implement it
- LlamaIndex SubQuestionQueryEngine — decomposes a query into sub-questions and answers them against underlying data before composing a final response.
- DSPy — build and optimize multi-stage pipelines where decomposition and per-step solving are declared as modules.
- LangChain — sequential chains and LCEL compose ordered steps that pass intermediate outputs forward.
- The original least-to-most prompting method (Zhou et al., Google Research) as the reference formulation of the technique.
Cost & effortWhat it takes
The main cost is extra inference: solving in sequence means more tokens and, when steps are separate calls, more round trips and higher latency than a single prompt. Engineering effort goes into designing and validating a good decomposition and wiring answers between steps, since a weak split is the dominant failure mode. It stays far cheaper than fine-tuning or training, and costs can be trimmed by caching stable sub-answers, keeping decomposition and solving in one prompt for easy cases, and reserving the full sequential treatment for genuinely compositional problems.