Self-ask makes the model pose and answer its own follow-up sub-questions until it can settle the original question.
ConceptWhat it is
Self-ask is a prompting technique where the model explicitly poses follow-up sub-questions, answers each one, and only then assembles a final answer to the original question. It targets the compositionality gap: models frequently know each fact in isolation yet fail to chain them, because a single-shot answer skips the intermediate hops and guesses.
By forcing the decomposition into a fixed, parseable shape — a Follow-up line followed by an Intermediate answer line — self-ask makes every hop visible and, crucially, lets a search engine answer any sub-question instead of the model, grounding multi-hop factual questions in retrieved evidence rather than parametric memory.
How it worksThe mechanics
A few-shot prompt demonstrates the format so the model imitates it: given a question, it decides whether a follow-up is needed, emits "Follow-up:" with the next sub-question, and receives an "Intermediate answer:" either from its own generation or from a search call keyed on that sub-question; the loop repeats, each answered hop feeding the next, until the model emits "So the final answer is:" and stops.
At a glanceSee it
The harness side — a fixed marker vocabulary lets a controller pause the decode, run a tool, splice the result back, and halt only when the final-answer phrase appears.
Why the loop pays off — one forward pass must bind the bridge entity in latent space and commits early, while writing that entity out as tokens grounds the second lookup and closes the compositionality gap.
When to use itWhere it fits
- Compositional, multi-hop factual questions where the answer depends on chaining several facts.
- Questions whose individual sub-facts are searchable even though the composed answer is not.
- Cases where you want an auditable trail of intermediate steps, not just a final claim.
- Pairing a language model with a search API so each hop is grounded in retrieved evidence.
When NOT to use itLimits & anti-patterns
- Single-hop lookups that a direct question or one search call already answers.
- Reasoning-heavy tasks with no clean factual decomposition, where free-form chain-of-thought fits better.
- Latency- or cost-sensitive paths, since every follow-up is another model or search call.
- Questions whose sub-questions cannot be answered independently of one another.
Trade-offsAdvantages & costs
Advantages
- Narrows the compositionality gap on multi-hop questions versus a single-shot answer.
- The rigid Follow-up and Intermediate-answer format is trivial to parse and to route to search.
- Exposes each reasoning hop, making a wrong step easy to spot and debug.
- Grounds factual sub-questions in live retrieval rather than model memory.
Trade-offs & costs
- Each sub-question is an extra model or search call, raising latency and cost.
- A wrong intermediate answer silently propagates into the final answer.
- The decomposition can wander, spawning irrelevant or redundant follow-ups.
- It only helps when the question genuinely decomposes, and forces structure onto ones that do not.
ExampleIn the real world
Take "Who was president of the United States when the oldest national park was founded?" A single-shot model often guesses. Under self-ask it emits "Follow-up: What is the oldest US national park?" — a search returns Yellowstone, founded 1872 — then "Follow-up: Who was US president in 1872?" — search returns Ulysses S. Grant — and finally "So the final answer is: Ulysses S. Grant", with each hop resting on a retrieved fact instead of one risky leap.
ToolsHow to implement it
- Self-Ask with Searchthe original scheme, pairing follow-up questions with a Google Search API through SerpAPI.
- LangChainships a self-ask-with-search agent that routes each follow-up to a search tool.
- DSPylets you program and compile multi-hop question decomposition instead of hand-writing the prompt.
- SerpAPI or Tavilysearch back-ends that answer each follow-up sub-question with fresh results.
Cost & effortWhat it takes
Cost scales with hop count: an n-hop question means roughly n follow-up model calls plus n search queries, so a two-to-four-hop answer runs several times a single-shot call in both latency and dollars. Engineering effort is modest — a few-shot format plus a parser loop — but the recurring search-API bill and the added latency per hop are what to budget for.