Home › RAG - Retrieval-Augmented Generation › Key term › Top-k retrieval
Key term · Ground

Top-k retrieval

Pulling back only the handful of passages most likely to help.

In one line

Top-k retrieval fetches the k chunks most similar to the query and hands them to the model.

DefinitionWhat it means

Top-k retrieval ranks all indexed chunks by similarity to the query embedding, or by a hybrid of dense and keyword scores, and returns the k highest-scoring ones, commonly k between three and twenty, to be inserted into the model's context. Choosing k trades off recall, missing a needed fact, against noise and cost, too much irrelevant text crowding the prompt.

Why it mattersWhy you should care

Top-k is the retrieval half of the RAG contract: even a perfect model gives a wrong answer if the right passage never made the cut. Teams tune k, add re-ranking after the initial retrieval, and monitor recall metrics because this single parameter is often the difference between a RAG system that feels reliable and one that quietly hallucinates.

At a glanceSee it

Top-k retrieval diagram
Top-k retrieval diagram 1

Choosing k is a tradeoff — too low drops relevant chunks, too high lets distractors dilute the prompt, so tune it on an eval set.

Top-k retrieval diagram 2

Production retrievers over-fetch, then rerank and de-duplicate — raw cosine top-k alone tends to return near-identical chunks.

Where you see itIn the wild

  • Vector database queries specifying a k parameter alongside the embedding.
  • Re-ranking layers placed after initial top-k retrieval to improve precision.
  • Retrieval evaluation dashboards tracking recall at k across query sets.
A living map of modern AI — kept current every morning