Faithfulness asks whether the answer is supported by what was retrieved, not just plausible.
DefinitionWhat it means
Faithfulness is an evaluation metric for retrieval-augmented generation that measures whether every claim in a generated answer is actually supported by the retrieved passages the model was given, as opposed to being a fluent but unsupported addition drawn from the model's parametric memory.
Why it mattersWhy you should care
A high-faithfulness answer can still be wrong if the retrieved sources themselves are bad, but a low-faithfulness answer is a direct signal of ungrounded hallucination, so faithfulness scoring alongside retrieval-quality metrics is standard in any RAG evaluation suite before a product goes to customers.
At a glanceSee it
Faithfulness is scored per claim — each atomic claim is tested for entailment against a source, the share that pass becomes the score, and any claim no source supports counts as a hallucination however fluent it reads.
Two distinct ways an answer betrays its sources — intrinsic errors contradict the retrieved text while extrinsic errors add facts the sources never mention — both repaired by tying each claim to a cited span.
Where you see itIn the wild
- RAG evaluation frameworks like RAGAS reporting faithfulness alongside relevance.
- Grounding checks embedded in enterprise document Q&A products.
- Eval discussions distinguishing faithfulness from factual correctness.