Home › Evals & Testing › Key term › Faithfulness
Key term · Operate

Faithfulness

Checking whether a RAG answer is actually backed by its retrieved sources.

In one line

Faithfulness asks whether the answer is supported by what was retrieved, not just plausible.

DefinitionWhat it means

Faithfulness is an evaluation metric for retrieval-augmented generation that measures whether every claim in a generated answer is actually supported by the retrieved passages the model was given, as opposed to being a fluent but unsupported addition drawn from the model's parametric memory.

Why it mattersWhy you should care

A high-faithfulness answer can still be wrong if the retrieved sources themselves are bad, but a low-faithfulness answer is a direct signal of ungrounded hallucination, so faithfulness scoring alongside retrieval-quality metrics is standard in any RAG evaluation suite before a product goes to customers.

At a glanceSee it

Faithfulness diagram
Faithfulness diagram 1

Faithfulness is scored per claim — each atomic claim is tested for entailment against a source, the share that pass becomes the score, and any claim no source supports counts as a hallucination however fluent it reads.

Faithfulness diagram 2

Two distinct ways an answer betrays its sources — intrinsic errors contradict the retrieved text while extrinsic errors add facts the sources never mention — both repaired by tying each claim to a cited span.

Where you see itIn the wild

  • RAG evaluation frameworks like RAGAS reporting faithfulness alongside relevance.
  • Grounding checks embedded in enterprise document Q&A products.
  • Eval discussions distinguishing faithfulness from factual correctness.
A living map of modern AI — kept current every morning