Instead of searching with the user's question, HyDE first has an LLM draft a plausible answer, then retrieves the real documents that most resemble that draft.
ConceptWhat it is
HyDE (Hypothetical Document Embeddings) is a query-translation trick for retrieval-augmented generation. Dense retrieval fails when a short question and a long answer are worded nothing alike: a user types "why does my container keep restarting?" while the document that holds the answer says "the liveness probe failed after the OOM-killer terminated the process." The two embeddings land far apart even though one answers the other. HyDE bridges this query-document phrasing gap by not embedding the question at all.
Instead, an LLM is asked to draft a hypothetical answer to the question. That answer is unverified and may be factually wrong, but it is written in the vocabulary, length, and declarative style of a real document. Embedding that and searching the corpus lands you near genuine passages that share its phrasing. The generated text is a throwaway: it is used only as a better-shaped search key, and the actual answer is always synthesized from the real retrieved documents.
How it worksThe mechanics
Take the user query and prompt an LLM to write a paragraph-length hypothetical answer as if it already knew the facts. Discard the query embedding and instead embed the hypothetical answer with your standard embedding model. Run a nearest-neighbor search in the vector store using that embedding, retrieving the top-k real chunks. Optionally rerank them. Finally, pass the real retrieved documents (never the hypothetical) plus the original question to the LLM to generate a grounded, cited answer. The extra LLM call sits entirely in front of retrieval; everything downstream is ordinary RAG.
At a glanceSee it
Real HyDE averages several hypothetical drafts into one search vector, so a single stray draft barely moves it — but averaging only tames random variance, not a wrong assumption shared by every draft.
HyDE is not free: a decision gate weighs vocabulary mismatch against hallucination risk, because drafting a fake answer for a fact the model does not know can invent entity names that pull retrieval off topic.
When to use itWhere it fits
- Retrieval quality is poor because questions and source documents use very different vocabulary, register, or length.
- The corpus is technical, jargon-heavy, or sparse, where lexically exact phrasing matters and near-misses are common.
- You are doing zero-shot retrieval with no labeled query-document pairs to fine-tune an embedding model on.
- Users ask terse, keyword-poor questions that a well-formed answer would express far more richly.
When NOT to use itLimits & anti-patterns
- The corpus is FAQ-style or already answer-shaped, so the question already matches documents closely and HyDE adds only latency.
- Latency or cost per query is tight, and an extra pre-retrieval LLM call is unacceptable.
- Queries touch niche facts the model knows nothing about, so its hypothetical is noise that actively misleads search.
- A strong hybrid (keyword plus dense) retriever or a fine-tuned embedding model already solves the phrasing gap more cheaply.
Trade-offsAdvantages & costs
Advantages
- Closes the query-document phrasing gap without any training data, fine-tuning, or labeled relevance judgments.
- Drops in ahead of an existing RAG pipeline; the retriever, index, and generator stay unchanged.
- Especially effective on technical and sparse corpora where a well-written answer surfaces the right terminology.
- Model-agnostic: works with any LLM for drafting and any off-the-shelf embedding model for search.
Trade-offs & costs
- Adds one extra LLM call per query, raising both latency and token spend on every request.
- A confidently wrong hypothetical can steer retrieval toward the wrong region of the corpus.
- Little or no benefit when questions already resemble documents, so the cost buys nothing.
- Harder to debug: a bad final answer may trace to the throwaway hypothetical rather than the retriever or generator.
ExampleIn the real world
An internal support assistant sits over a Kubernetes runbook corpus. A user asks "pods keep dying, what do I check?" Embedding that short phrase retrieves generic troubleshooting pages and misses the specific runbook. With HyDE, the LLM first drafts a hypothetical answer: "Pods that repeatedly restart are usually killed by the OOM-killer when memory limits are exceeded, or fail their liveness probe; inspect resource requests, limits, and probe timeouts." That paragraph is embedded and searched. Because it now carries the terms "OOM-killer," "liveness probe," and "memory limits," it lands next to the exact runbook section, which is retrieved and used to write the real, cited answer, without the fabricated draft ever reaching the user.
ToolsHow to implement it
- LangChainprovides a HypotheticalDocumentEmbedder that wraps an LLM plus a base embedding model for HyDE-style retrieval.
- LlamaIndexships a HyDEQueryTransform that rewrites the query into a hypothetical document before retrieval.
- Haystacksupports building the draft-then-embed step as a custom retrieval pipeline component.
- Any embedding model for the search key, such as OpenAI text-embedding-3, Cohere Embed, or open-source sentence-transformers.
Cost & effortWhat it takes
Complexity is medium: no training or new infrastructure, but you insert a generation step before retrieval. The dominant runtime cost is one additional LLM call per query, adding its latency and output tokens on top of the final answering call, roughly doubling the LLM round-trips. Engineering effort is a day or two using a framework's built-in HyDE transform, most of it spent tuning the hypothetical prompt and A/B testing retrieval quality against a plain dense or hybrid baseline to confirm the phrasing gap is real enough to justify the extra call.