RAG metrics check whether an answer is actually grounded in the documents it cited.
ConceptWhat it is
RAG metrics are a family of scores purpose-built for retrieval-augmented generation systems, measuring dimensions such as faithfulness to retrieved context, answer relevance, and retrieval precision and recall.
They exist because generic quality metrics do not catch a RAG system's specific failure modes, like answering fluently but ungrounded in the retrieved documents, or retrieving irrelevant chunks that never should have been surfaced.
How it worksThe mechanics
Given a question, the retrieved context chunks, and the generated answer, each metric checks a specific relationship: faithfulness checks whether every claim in the answer is supported by the retrieved context, context relevance checks whether retrieved chunks actually relate to the question, and these scores are computed automatically, often using an LLM judge under the hood.
At a glanceSee it
The RAG triad — each metric scores one edge of the question / context / answer triangle: context relevance for retrieval, faithfulness for grounding, answer relevance for usefulness.
Split by pipeline stage: retrieval metrics (precision, recall, MRR, nDCG) tell you whether the right chunks came back; generation metrics (faithfulness, answer relevance) tell you whether the model used them well — and each points at a different fix.
When to use itWhere it fits
- Evaluating whether a RAG pipeline's answers are grounded in retrieved sources.
- Diagnosing whether failures stem from bad retrieval or bad generation.
- Continuously monitoring a production RAG system for hallucination drift.
- Comparing retrieval strategies like hybrid search versus pure vector search.
When NOT to use itLimits & anti-patterns
- Non-retrieval-augmented systems, where these metrics do not apply.
- Very early prototyping stages, where qualitative spot-checks may be faster than setting up formal metrics.
- Tasks where the answer legitimately requires synthesis beyond the retrieved context, which faithfulness scoring can unfairly penalize.
Trade-offsAdvantages & costs
Advantages
- Pinpoints whether retrieval or generation is the source of a failure.
- Directly targets hallucination, a top concern for RAG systems.
- Enables systematic comparison of retrieval strategies and chunk sizes.
- Can run continuously in production for drift detection.
Trade-offs & costs
- Metrics computed via an LLM judge inherit that judge's own biases.
- Faithfulness scoring can penalize valid inference beyond literal context.
- Requires careful chunk-level instrumentation to compute correctly.
- Adds evaluation-time cost on top of the generation itself.
ExampleIn the real world
A healthcare documentation assistant runs Ragas faithfulness and context-precision metrics nightly on a sample of production queries to catch any drift toward ungrounded or off-topic answers before it reaches clinicians.
ToolsHow to implement it
- Ragasleading open-source library for faithfulness, relevance, and precision metrics.
- TruLensinstruments and scores RAG pipelines with groundedness metrics.
- Arize Phoenixobservability platform with built-in RAG evaluation metrics.
- LlamaIndex evaluation modulebuilt-in retrieval and response evaluators.
Cost & effortWhat it takes
Adds the cost of one or more LLM-judge calls per evaluated query on top of the generation cost. Moderate engineering effort to instrument retrieval and wire metrics into a pipeline.