Bad chunking guarantees bad retrieval no matter how good the embedding model is.
ConceptWhat it is
Chunking strategies are the methods used to split source documents into retrievable units before embedding: fixed-size windows, semantic chunking that splits on meaning shifts, recursive splitting that respects document structure like headers, and overlap between chunks to avoid cutting context mid-thought. It exists because a chunk that is too large dilutes relevance, and a chunk that is too small loses the context needed to answer the question.
Chunking is a first-order lever on RAG quality; no amount of reranking or agentic looping fixes context that was split badly at index time.
How it worksThe mechanics
Documents are parsed into a structured representation, split into chunks using a fixed token window, semantic similarity boundaries, or document structure like sections and headers, given a small overlap between neighbors, tagged with metadata like source and section title, and then embedded and written into the vector index.
At a glanceSee it
Pick the splitter from the document’s own shape — structure routes to recursive, drifting prose to semantic, everything else to a fixed window.
Too small loses context and too large dilutes relevance — overlap is the cheap insurance that buys the middle.
When to use itWhere it fits
- Any RAG system at the design stage, since chunking choices bound all downstream retrieval quality.
- Structured documents like manuals or contracts, where recursive splitting on headers preserves meaning.
- Long-form content where semantic chunking prevents cutting a thought in half mid-chunk.
- Corpora being re-indexed after poor retrieval results, where re-chunking is often the first fix to try.
When NOT to use itLimits & anti-patterns
- Already-short, atomic documents like FAQ entries, where further chunking adds complexity with no benefit.
- Extremely time-constrained prototypes where a naive fixed-size split is good enough to validate the idea.
- Content where semantic chunking's extra embedding calls are not worth the latency for a small corpus.
Trade-offsAdvantages & costs
Advantages
- Directly improves retrieval precision and recall more than most other RAG tuning levers.
- Semantic and recursive strategies preserve meaning better than naive fixed windows.
- Overlap reduces the risk of losing context at chunk boundaries.
- Metadata tagging enables filtering and citation back to source.
Trade-offs & costs
- Semantic chunking requires extra embedding calls, adding indexing cost.
- No single strategy works well across all document types, requiring per-corpus tuning.
- Poor chunk size choices are easy to make and hard to notice until retrieval quality is evaluated.
- Re-chunking an existing index means a full re-embedding pass.
ExampleIn the real world
A documentation platform re-indexes its help center using recursive, header-aware chunking instead of fixed 500-token windows, and measurably cuts down answers that cite the wrong section.ToolsHow to implement it
- LangChain RecursiveCharacterTextSplitterstructure-aware splitting that respects headers and paragraphs.
- LlamaIndex SemanticSplitterNodeParsersplits on embedding similarity shifts rather than fixed length.
- Unstructured.ioparses PDFs and complex layouts into clean chunkable elements first.
- Ragasmeasures context precision and recall to compare chunking strategies empirically.
Cost & effortWhat it takes
Low incremental cost for fixed-size chunking, moderate added cost for semantic chunking due to extra embedding calls; the main investment is engineering time spent tuning chunk size and overlap per corpus.