Home › Guardrails & Responsible AI › PII redaction
🛡️ · Operate

PII redaction

Detecting and masking personal data before it enters prompts, logs, or outputs.

In one line

PII redaction keeps names, numbers, and addresses out of prompts, logs, and model outputs by default.

ConceptWhat it is

PII redaction is the automated detection and masking of personally identifiable information, such as names, phone numbers, social security numbers, and addresses, in text flowing into or out of an LLM. It exists because prompts and completions routinely get logged, cached, or sent to third-party model providers, and unredacted personal data in any of those places creates a privacy and compliance liability.

Redaction can happen before data reaches the model, after generation, or both, depending on whether the risk is input exposure or output leakage.

How it worksThe mechanics

A named-entity recognition model or regex-and-dictionary pipeline scans text for PII categories, replaces matches with placeholder tokens like NAME or a reversible pseudonym, and optionally maintains a secure mapping so redacted values can be restored for authorized downstream steps.

At a glanceSee it

PII redaction diagram
PII redaction diagram 1

Detection is never perfect — the redaction threshold trades false positives that strip context the model needs against false negatives that leak real PII to logs, cache, and providers.

PII redaction diagram 2

Masking techniques split by reversibility — irreversible methods maximize privacy, while reversible ones preserve recovery at the cost of a vault or key that becomes its own breach target.

When to use itWhere it fits

  • Healthcare and financial applications subject to HIPAA, GDPR, or similar regulation.
  • Any pipeline sending customer text to a third-party model API outside the company's trust boundary.
  • Logging and observability systems that store prompts and completions for debugging.
  • Customer support transcripts being used for fine-tuning or analytics.

When NOT to use itLimits & anti-patterns

  • Fully on-premise systems processing only synthetic or already-anonymized data.
  • Internal developer tools with no real customer data flowing through, where redaction adds friction with no benefit.

Trade-offsAdvantages & costs

Advantages
  • Directly reduces regulatory and breach-notification risk.
  • Can be applied uniformly across logging, caching, and model-call layers.
  • Reversible pseudonymization preserves utility for legitimate downstream use.
Trade-offs & costs
  • Imperfect detection means rare-format PII, like unusual ID numbers, can slip through.
  • Over-redaction can strip context the model needs to answer correctly.
  • Adds a processing step and secure key-management overhead for reversible schemes.

ExampleIn the real world

A healthcare chatbot vendor runs Presidio-based redaction on every patient message before it reaches OpenAI's API, so no protected health information ever leaves the vendor's own infrastructure unmasked.

ToolsHow to implement it

  • Microsoft Presidioopen-source PII detection and anonymization framework.
  • AWS Comprehend PIImanaged detection and redaction as part of AWS's NLP suite.
  • Google Cloud DLPentity detection with tokenization and format-preserving encryption.
  • spaCy NERcustomizable open-source entity recognition for building bespoke pipelines.

Cost & effortWhat it takes

Low per-call cost for managed APIs, roughly fractions of a cent per document; the bigger cost is engineering effort for reversible mapping and secure key storage.

A living map of modern AI — kept current every morning