OverviewWhat it is
RAG retrieves relevant documents from a knowledge base and feeds them into the prompt, so the model answers from your facts. It reduces hallucination, keeps knowledge current, and enables citations - all without changing the model's weights. It is the default way to build 'chat with our docs/product'.
At a glanceRAG - Retrieval-Augmented Generation
The model never learns your data - you fetch the right passages at question-time and put them in the prompt.
CompareThe RAG landscape, side by side
Every retrieval pattern worth knowing, grouped by what it improves — filter by family or search. The method-by-method cards further down add a diagram for each core pattern.
Every row has a page — what it does, what it costs you, and how to tell when it is the thing biting you.
MechanicsHow it works
Two phases. Indexing (offline): split documents into chunks, embed them, store in a vector DB. Query (live): embed the question, retrieve the top matches, optionally rerank, then pass them as context to the LLM to generate a grounded answer.
Ground levelWhat you actually build
Two phases, and the failures cluster in one of them. Instrument retrieval and generation separately or you cannot tell which half broke.
LandscapeTypes & approaches
Click a highlighted type to open its own page — concept, use case, and diagram.
FeasibilityArchitecture & feasibility
Architecture & feasibility
- Answer quality is capped by retrieval quality. The feasibility levers are chunking, embedding choice, hybrid search, reranking, and grounding/citation enforcement.
- RAG is usually cheaper and more feasible than fine-tuning: no training run, no model hosting, and knowledge updates by re-indexing.
- Failure modes cluster in retrieval (wrong or missing chunks). Instrument retrieval separately from generation so you can tell which half failed.
Method by methodThe RAG methods — each with its own diagram
RAG is a family, not one thing. Here are the main methods from simplest to most powerful, each with how it flows and where it wins or loses. Most teams start at Naïve, ship on Advanced, and reach for Agentic or GraphRAG only when the questions demand it.
Naïve RAG
BaselineSingle shot: embed the question, pull the top-k similar chunks, stuff them into the prompt, generate.
Advantages
- Simplest and fastest to build
- Cheapest — one retrieval, one LLM call
- Easy to debug and explain
Trade-offs
- Misses relevant chunks when embeddings are weak
- No reranking, so noise leaks into context
- Exact terms (IDs, codes, names) can be missed
- Struggles with multi-part questions
Advanced RAG (hybrid + rerank)
Production defaultCombine keyword (BM25) and vector search, then rerank the candidates with a cross-encoder before generating.
Advantages
- Large recall + precision boost
- Catches exact terms vector-only search misses
- Reranking pushes the best passage to the top
- Still a predictable, single-pass pipeline
Trade-offs
- More moving parts and latency (extra rerank call)
- Higher cost per query
- Needs tuning of weights and k
Query transformation (multi-query / HyDE)
Recall boosterRewrite or expand the question into several queries — or a hypothetical answer (HyDE) — retrieve for each, then merge.
Advantages
- Recovers vague or under-specified questions
- Higher recall on hard queries
- Cheap to bolt onto an existing pipeline
Trade-offs
- Multiplies retrieval cost and latency
- Can pull in off-topic chunks
- Adds an extra LLM call to rewrite
Agentic RAG
AdaptiveAn agent decides whether and what to retrieve, can call multiple sources or tools, and loops until it has enough — reasoning over retrieval.
Advantages
- Handles multi-hop, multi-source questions
- Adapts retrieval to each query
- Can verify and re-retrieve when unsure
Trade-offs
- Slowest and most expensive
- Hardest to make reliable — needs step limits & guardrails
- Non-deterministic; harder to test
GraphRAG
Structured knowledgeBuild a knowledge graph from your documents; retrieve by traversing entities and relationships, not just similarity.
Advantages
- Excels at 'connect-the-dots' and global/summary questions
- Captures relationships flat chunks lose
- Strong for entity-heavy domains
Trade-offs
- Expensive to build and maintain the graph
- Complex pipeline and infra
- Overkill for simple Q&A; slow to update
In practiceWhat it means for building
Cheaper and more controllable than fine-tuning, and answers stay up to date as your data changes. The go-to pattern for knowledge assistants.
Design chunking, embeddings, hybrid search, reranking, and citation enforcement. Retrieval quality is the ceiling on answer quality.
GlossaryKey terms
CheckCheck your understanding
RAG vs fine-tuning?
RAG injects knowledge (current, private, cite-able) without retraining; fine-tuning changes behaviour (tone, format, skill). Prompt first; combine both for complex systems.
My RAG bot gives wrong answers - where do you look?
Almost always retrieval: bad chunking, weak embeddings, or low top-k means the right passage never reaches the model. Fix retrieval before the prompt.
How do you evaluate a RAG system?
Measure retrieval (context recall/precision) and generation (faithfulness, answer relevance) separately, on a golden set.
What changedWhat changed here
As of 2026-09-25 — retrieval research in the daily brief: 12 items in the last 7 days
6 of the 12, newest first, from the Signal daily brief
- Agent Memory with Episodic Retrieval for Financial Decision-Making
- CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars
- Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit
- LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning
- Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices
- UniDataAgent: An Ontology-Grounded Agent for Enterprise Question-to-Report Automation
Three kinds of claim, strongest first. Signal runs every morning.