Before searching, an LLM turns a terse or conversational query into a clear, standalone one so retrieval finds the right chunks.
ConceptWhat it is
Retrieval quality hinges on the query actually matching how relevant passages are written and embedded. Real users rarely oblige: they type terse keywords, misspell terms, or ask follow-up questions that only make sense given earlier turns, like "what about the cheaper one?". Embedding search takes that raw string at face value, so a vague or context-dependent query pulls back weak or wrong chunks no matter how good the vector store is.
Query rewriting inserts one LLM step before retrieval: the model reads the user's message plus any conversation history and produces a cleaner, self-contained rewritten query that resolves pronouns, expands jargon, and fixes phrasing before it is embedded and searched. It is the classic rewrite-retrieve-read pattern, and it is especially load-bearing in conversational search, where each turn depends on the last.
How it worksThe mechanics
The user's latest message and recent chat history go to a rewriter LLM with an instruction like "rewrite this into a standalone search query." The model resolves references, turning "the second one" into the actual entity, spells out abbreviations, and drops conversational filler, emitting one or more clean queries. That rewritten text, not the original, is embedded and sent to the retriever, which fetches the top-k chunks from the vector store; those chunks plus the original user question are then handed to the answering model. Some implementations skip the rewrite when the query is already self-contained, or fan out into several rewrites for broader recall.
At a glanceSee it
The one rewrite box is really four families — each repairs a different defect in the raw query.
Multi-query fan-out turns one question into several phrasings, retrieves for each, then fuses the ranked lists — recovering chunks a single wording would miss.
When to use itWhere it fits
- Conversational or multi-turn search where follow-ups lean on earlier context, like "and its warranty?".
- Users type short, keyword-style, or sloppy queries that miss the vocabulary of your documents.
- A domain full of synonyms, acronyms, or jargon the rewriter can expand to match source text.
- Retrieval recall is your bottleneck and you have latency budget for one extra call.
When NOT to use itLimits & anti-patterns
- Single-shot, already-precise queries such as an exact error code or ID, where rewriting only adds risk and latency.
- Ultra-low-latency paths where an extra LLM round trip is unacceptable.
- When cheaper fixes should come first: better chunking, a stronger embedding model, or metadata filters.
- High-stakes literal lookups where a rewrite could silently distort the user's intent.
Trade-offsAdvantages & costs
Advantages
- Simple, model-agnostic bolt-on that sits in front of any existing retriever without touching the index.
- Large recall gains on vague, terse, or conversational queries, often the cheapest win in a RAG stack.
- Resolves cross-turn context so a single-shot retriever works in a chat setting.
- The rewrite is inspectable and loggable, which helps debugging and evals.
Trade-offs & costs
- Adds an extra LLM call, so more latency and cost on every query.
- The rewriter can drift from the user's true intent and retrieve confidently wrong chunks.
- Another prompt and model to tune, monitor, and guard against injection coming from chat history.
- Marginal or even harmful on queries that were already clean, so you may need a gate to skip it.
ExampleIn the real world
Consider a support assistant over a product help center. A user asks "How do I reset it?" one turn after discussing a home wireless router. Embedding the bare question "How do I reset it?" retrieves generic reset articles scattered across every product line. With query rewriting, the LLM reads the history and emits "How to factory reset a home wireless router," resolving the pronoun and adding the product term. Retrieval now surfaces the router's reset guide, and the answering model is grounded on the right passage instead of guessing.
ToolsHow to implement it
- LangChain, whose history-aware retriever and condense-question chains rewrite a follow-up into a standalone query.
- LlamaIndex, with query transformation modules including HyDE and sub-question rewriting.
- DSPy, which lets you optimize a rewrite module against retrieval metrics instead of hand-tuning prompts.
- RAG-Fusion, a pattern that generates multiple rewritten queries and fuses their results.
Cost & effortWhat it takes
The added cost is one LLM call per query before retrieval, so keep it on a small, fast model; the rewrite is a short, well-scoped task that does not need your top-tier generator. Budget for the extra latency, which you can soften with a lightweight model, streaming, or a cheap classifier that skips rewriting when the query is already standalone, plus a modest prompt-engineering and eval effort to confirm rewrites actually lift retrieval metrics rather than just adding a hop. Overall it is a low-effort, low-cost addition relative to its recall upside.