Top-k retrieval fetches the k chunks most similar to the query and hands them to the model.
DefinitionWhat it means
Top-k retrieval ranks all indexed chunks by similarity to the query embedding, or by a hybrid of dense and keyword scores, and returns the k highest-scoring ones, commonly k between three and twenty, to be inserted into the model's context. Choosing k trades off recall, missing a needed fact, against noise and cost, too much irrelevant text crowding the prompt.
Why it mattersWhy you should care
Top-k is the retrieval half of the RAG contract: even a perfect model gives a wrong answer if the right passage never made the cut. Teams tune k, add re-ranking after the initial retrieval, and monitor recall metrics because this single parameter is often the difference between a RAG system that feels reliable and one that quietly hallucinates.
At a glanceSee it
Choosing k is a tradeoff — too low drops relevant chunks, too high lets distractors dilute the prompt, so tune it on an eval set.
Production retrievers over-fetch, then rerank and de-duplicate — raw cosine top-k alone tends to return near-identical chunks.
Where you see itIn the wild
- Vector database queries specifying a k parameter alongside the embedding.
- Re-ranking layers placed after initial top-k retrieval to improve precision.
- Retrieval evaluation dashboards tracking recall at k across query sets.