Home › Guardrails & Responsible AI › Hallucination / grounding checks
🛡️ · Operate

Hallucination / grounding checks

Verifying model claims against source evidence before they reach a user.

In one line

A grounding check catches confident-sounding answers that aren't actually backed by any source.

ConceptWhat it is

Hallucination is when a model states something false or unsupported with the same fluent confidence as a true statement, because next-token prediction has no built-in notion of factual grounding. Grounding checks are the verification layer that compares generated claims against retrieved source documents or a knowledge base to catch this before it reaches a user.

They exist because RAG pipelines reduce hallucination but do not eliminate it; a model can still ignore, misread, or extrapolate beyond its retrieved context.

How it worksThe mechanics

After generation, a second pass, often a smaller LLM judge or an NLI-style entailment model, checks whether each claim in the answer is entailed by the retrieved chunks; unsupported claims get flagged, cited with lower confidence, or trigger a regeneration with stricter instructions to only use provided context.

At a glanceSee it

Hallucination / grounding checks diagram
Hallucination / grounding checks diagram 1

Inside the grounding-check box — the answer is split into atomic claims, each scored against retrieved chunks into one of three NLI verdicts: entailed, contradicted, or neutral, with unsupported claims kept distinct from contradicted ones.

Hallucination / grounding checks diagram 2

Why grounding is necessary but not sufficient — it enforces faithfulness to a source, so a claim backed by a wrong or outdated source sails through the check while remaining false.

When to use itWhere it fits

  • Enterprise knowledge assistants where wrong answers erode trust in the whole system.
  • Medical, legal, or financial copilots where an unsupported claim carries real liability.
  • Any RAG system being evaluated or monitored for production readiness.
  • Summarization tools where fabricated details must be caught before publishing.

When NOT to use itLimits & anti-patterns

  • Purely creative or brainstorming use cases where factual grounding is not the point.
  • Low-stakes internal tools where added latency and cost outweigh the risk of an occasional wrong answer.

Trade-offsAdvantages & costs

Advantages
  • Directly targets the failure mode users fear most from LLMs.
  • Can auto-flag or auto-correct without a human in the loop for every request.
  • Produces measurable hallucination rates for regression testing across model versions.
Trade-offs & costs
  • Adds a full extra inference call, roughly doubling latency and cost.
  • The grounding judge can itself be wrong, flagging a mismatch that isn't there.
  • Only catches claims traceable to retrieved text, not general reasoning errors.

ExampleIn the real world

Morgan Stanley's internal GPT-4 assistant for financial advisors runs a grounding evaluation layer that checks every generated answer against the firm's own research corpus before displaying it.

ToolsHow to implement it

  • Ragasopen-source RAG evaluation with faithfulness and context-precision metrics.
  • TruLensinstrumentation for tracking groundedness scores in production.
  • Galileocommercial hallucination-detection and evaluation platform.
  • Vectara HHEMa dedicated hallucination-evaluation model for scoring factual consistency.

Cost & effortWhat it takes

Roughly doubles per-query LLM cost and adds 200ms to a few seconds of latency depending on judge model size; moderate engineering effort to wire into a pipeline.

A living map of modern AI — kept current every morning