Chunking splits long documents into smaller passages so retrieval and context fitting both work.
DefinitionWhat it means
Chunking divides source documents into passages of a target size, often a few hundred tokens, with strategies ranging from fixed-length windows to sentence or section-aware splits, sometimes with overlap between neighboring chunks so a fact near a boundary is not cut in half. Each chunk becomes the unit that gets embedded, indexed, and retrieved.
Why it mattersWhy you should care
Chunk size and boundaries directly control retrieval quality: chunks too large dilute relevance and waste context window, chunks too small lose surrounding meaning and fragment a single idea across multiple retrieved pieces. Tuning chunking strategy per document type, tables, code, prose, is one of the highest-leverage and most underrated jobs in building a production RAG system.
At a glanceSee it
The splitter is not one algorithm — document shape and the need to preserve meaning pick between recursive, semantic, and fixed-window splitting.
Overlapping the sliding window by a few tokens keeps a fact that straddles a chunk seam retrievable — exactly what a zero-overlap split silently loses.
Where you see itIn the wild
- Document ingestion pipelines feeding a vector database.
- Tuning overlap and chunk size while debugging poor retrieval recall.
- Splitting PDFs, tickets, or wikis where tables and headers need special handling.