Home › Ground
📚 Ground

RAG - Retrieval-Augmented Generation

Give the model your data at answer-time instead of retraining it.

OverviewWhat it is

RAG retrieves relevant documents from a knowledge base and feeds them into the prompt, so the model answers from your facts. It reduces hallucination, keeps knowledge current, and enables citations - all without changing the model's weights. It is the default way to build 'chat with our docs/product'.

At a glanceRAG - Retrieval-Augmented Generation

RAG - Retrieval-Augmented Generation diagram

The model never learns your data - you fetch the right passages at question-time and put them in the prompt.

CompareThe RAG landscape, side by side

Every retrieval pattern worth knowing, grouped by what it improves — filter by family or search. The method-by-method cards further down add a diagram for each core pattern.

Every row has a page — what it does, what it costs you, and how to tell when it is the thing biting you.

Foundational

MechanicsHow it works

Two phases. Indexing (offline): split documents into chunks, embed them, store in a vector DB. Query (live): embed the question, retrieve the top matches, optionally rerank, then pass them as context to the LLM to generate a grounded answer.

Ground levelWhat you actually build

Two phases, and the failures cluster in one of them. Instrument retrieval and generation separately or you cannot tell which half broke.

Two phases, and the failures cluster in one of them. Instrument retrieval and generation separately or you cannot tell which half broke.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

FeasibilityArchitecture & feasibility

Architecture & feasibility

  • Answer quality is capped by retrieval quality. The feasibility levers are chunking, embedding choice, hybrid search, reranking, and grounding/citation enforcement.
  • RAG is usually cheaper and more feasible than fine-tuning: no training run, no model hosting, and knowledge updates by re-indexing.
  • Failure modes cluster in retrieval (wrong or missing chunks). Instrument retrieval separately from generation so you can tell which half failed.

Method by methodThe RAG methods — each with its own diagram

RAG is a family, not one thing. Here are the main methods from simplest to most powerful, each with how it flows and where it wins or loses. Most teams start at Naïve, ship on Advanced, and reach for Agentic or GraphRAG only when the questions demand it.

Naïve RAG

Baseline

Single shot: embed the question, pull the top-k similar chunks, stuff them into the prompt, generate.

Naïve RAG diagram
Advantages
  • Simplest and fastest to build
  • Cheapest — one retrieval, one LLM call
  • Easy to debug and explain
Trade-offs
  • Misses relevant chunks when embeddings are weak
  • No reranking, so noise leaks into context
  • Exact terms (IDs, codes, names) can be missed
  • Struggles with multi-part questions
Best for: small, clean corpora and quick prototypes. Worst for: large/heterogeneous data or precise factual lookup.

Advanced RAG (hybrid + rerank)

Production default

Combine keyword (BM25) and vector search, then rerank the candidates with a cross-encoder before generating.

Advanced RAG (hybrid + rerank) diagram
Advantages
  • Large recall + precision boost
  • Catches exact terms vector-only search misses
  • Reranking pushes the best passage to the top
  • Still a predictable, single-pass pipeline
Trade-offs
  • More moving parts and latency (extra rerank call)
  • Higher cost per query
  • Needs tuning of weights and k
Best for: most production knowledge assistants — the sensible default. Worst for: ultra-low-latency paths or trivial corpora where it is overkill.

Query transformation (multi-query / HyDE)

Recall booster

Rewrite or expand the question into several queries — or a hypothetical answer (HyDE) — retrieve for each, then merge.

Query transformation (multi-query / HyDE) diagram
Advantages
  • Recovers vague or under-specified questions
  • Higher recall on hard queries
  • Cheap to bolt onto an existing pipeline
Trade-offs
  • Multiplies retrieval cost and latency
  • Can pull in off-topic chunks
  • Adds an extra LLM call to rewrite
Best for: ambiguous, natural-language questions. Worst for: cost-sensitive, high-volume paths.

Agentic RAG

Adaptive

An agent decides whether and what to retrieve, can call multiple sources or tools, and loops until it has enough — reasoning over retrieval.

Agentic RAG diagram
Advantages
  • Handles multi-hop, multi-source questions
  • Adapts retrieval to each query
  • Can verify and re-retrieve when unsure
Trade-offs
  • Slowest and most expensive
  • Hardest to make reliable — needs step limits & guardrails
  • Non-deterministic; harder to test
Best for: complex, multi-step research questions. Worst for: simple FAQ-style lookups (wasteful).

GraphRAG

Structured knowledge

Build a knowledge graph from your documents; retrieve by traversing entities and relationships, not just similarity.

GraphRAG diagram
Advantages
  • Excels at 'connect-the-dots' and global/summary questions
  • Captures relationships flat chunks lose
  • Strong for entity-heavy domains
Trade-offs
  • Expensive to build and maintain the graph
  • Complex pipeline and infra
  • Overkill for simple Q&A; slow to update
Best for: relationship-rich corpora and holistic questions. Worst for: small, simple, or fast-changing data.

In practiceWhat it means for building

Cheaper and more controllable than fine-tuning, and answers stay up to date as your data changes. The go-to pattern for knowledge assistants.

Design chunking, embeddings, hybrid search, reranking, and citation enforcement. Retrieval quality is the ceiling on answer quality.

GlossaryKey terms

CheckCheck your understanding

RAG vs fine-tuning?

RAG injects knowledge (current, private, cite-able) without retraining; fine-tuning changes behaviour (tone, format, skill). Prompt first; combine both for complex systems.

My RAG bot gives wrong answers - where do you look?

Almost always retrieval: bad chunking, weak embeddings, or low top-k means the right passage never reaches the model. Fix retrieval before the prompt.

How do you evaluate a RAG system?

Measure retrieval (context recall/precision) and generation (faithfulness, answer relevance) separately, on a golden set.

What changedWhat changed here

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning