Embed and search tiny chunks for precise matching, then hand the language model the full parent passage so it answers with complete context.
ConceptWhat it is
Parent-document retrieval (also called small-to-big) is a RAG pattern that decouples what you search from what you return. Each source document is split twice: into small child chunks used for embedding and matching, and into larger parent chunks (often a whole section or document) used for context. Only the children go into the vector index; the parents live in a separate document store keyed by id, and every child carries a reference back to its parent.
It exists to resolve the core chunking tension. Small chunks embed cleanly and match a query precisely, but a lone sentence often lacks the surrounding context a model needs to answer well. Large chunks carry context but dilute the embedding and retrieve imprecisely. Parent-document retrieval gets both at once: a precise match on the child, then the full context of its parent.
How it worksThe mechanics
At ingest, a splitter produces parent chunks, and each parent is further split into child chunks that store a reference to their parent id; the children are embedded into a vector store while the parents are written to a document store. At query time, the query is embedded and matched against the child index, each hit's parent id is resolved, duplicate parents are collapsed to one, the parent passages are fetched from the document store, and those full passages, not the matched children, are passed to the model as grounding context.
At a glanceSee it
The build-time half the runtime flow assumes — each document is split once into stored parents and again into embedded children, joined only by a shared parent id.
Choosing parent granularity is a real either—or: a tight sentence window when the answer is local, a whole section when it is spread out, tuned so the parent informs rather than floods the prompt.
When to use itWhere it fits
- Answers need the surrounding context (multi-step procedures, definitions, legal clauses) that a single sentence loses.
- Source documents are long and internally coherent, so a parent section reads as a complete, self-contained unit.
- Users ask precise questions (a specific error string, term, or figure) that match best against fine-grained child chunks.
- You already observe fragmented or context-starved snippets hurting answer quality under flat single-chunk retrieval.
When NOT to use itLimits & anti-patterns
- Documents are already short and self-contained (FAQ entries, product cards) where a child effectively equals its parent.
- Strict latency or storage budgets make a second store and an extra lookup step not worth the gain.
- The corpus is tiny enough to just pass whole documents, or so large that duplicating text across two tiers is costly.
- Precision is not the bottleneck; a plain chunk-and-retrieve baseline already answers well enough.
Trade-offsAdvantages & costs
Advantages
- Precise retrieval from small chunks plus complete context from the parent, without compromising either side.
- Fewer mid-thought, fragmented snippets reach the model, which improves answer coherence and reduces context-gap hallucination.
- Deduplication collapses several matching children into one parent, so the context window is not filled with near-duplicate passages.
- It is a retrieval-layer pattern that works with any embedding model and vector store, requiring no model change.
Trade-offs & costs
- Two-tier storage: a vector index for children plus a document store for parents, more to provision and keep in sync.
- Extra plumbing and a second lookup add latency and code complexity versus flat single-chunk retrieval.
- Parents can be large, so returning several may blow the context budget and raise per-query token cost.
- Tuning two chunk sizes, child and parent, is an added parameter surface that needs evaluation to get right.
ExampleIn the real world
An internal support assistant answers engineers' questions over a library of long operational runbooks and API references. Ingestion splits each runbook into section-sized parent chunks, then into paragraph-sized child chunks embedded in a vector store, with the parents held in a document store by id. When someone pastes a specific error string, the query matches the exact paragraph that mentions it, but the retriever returns the whole runbook section as context. The model then sees the prerequisite steps and the rollback note around that paragraph and produces a correct, complete procedure rather than an out-of-context single line.
ToolsHow to implement it
- LangChain ParentDocumentRetriever, which pairs a vector store for children with a separate docstore for parents.
- LlamaIndex hierarchical node parsing with AutoMergingRetriever and RecursiveRetriever for small-to-big retrieval.
- Vector stores such as Chroma, FAISS, or pgvector to hold the child embeddings.
- A key-value or document store (for example Redis or an in-memory docstore) to hold parents keyed by id.
Cost & effortWhat it takes
The heavy costs are storage and engineering, not compute. You persist text roughly twice, child chunks in the vector index and parents in a document store, and you maintain a second lookup path, so expect modest extra storage plus a little more retrieval latency and code. Embedding and query costs stay flat because you still embed only the small children. Query-time token cost can rise, since parents are larger than raw chunks, which is the main variable to watch and to cap with parent-size limits and deduplication.