Home › Security › Data privacy and consent
🔒 · Operate

Data privacy and consent

Whether you are allowed to use this data this way, which is a different question from whether you can.

In one line

Consent given for one use does not extend to a second, and production traffic becoming training data is exactly that second use.

ConceptWhat it is

Consent is permission for a specific use. Privacy is the broader question of whether a use is one the person would recognise and accept. Neither is answered by having the data, and both are decided long before the model sees it.

For an AI system the live question is almost always secondary use. A corpus gathered to deliver a service is proposed as training data, or production conversations are proposed as an eval set. Those are new purposes, and the original consent usually does not stretch to cover them.

How it worksThe mechanics

Data arrives with a purpose attached, and that purpose travels with it. In practice that means a provenance field on every corpus and every derived artefact recording where it came from and what it may be used for, checked at the point of use rather than remembered by whoever was in the room.

The second mechanism is minimisation at the boundary. Personal data that does not need to reach the model should not — detection and redaction before the prompt is assembled, not after the answer comes back. Anything that does cross the boundary is governed by the provider contract, which is why zero-retention terms matter more than they look.

At a glanceSee it

Data privacy and consent diagram

Purpose travels with the data and is checked at the point of use. Training is a separate purpose, so it needs its own basis rather than inheriting one.

When to use itWhere it fits

  • Before turning production traffic into training or eval data, which is the most common quiet breach.
  • When a corpus contains anything about identifiable people, including free-text fields nobody classified.
  • When choosing a model provider, because retention terms are the control over data you cannot keep inside.
  • When a kit or demo needs a corpus — a self-authored or properly licensed one removes the question entirely.

When NOT to use itLimits & anti-patterns

  • As a blanket refusal; most uses are fine and the point is to know which, not to stop.
  • As a redaction-only strategy, since re-identification from quasi-identifiers defeats naive masking.
  • Where the data is genuinely public and licensed for the use, in which case the analysis is short.
  • As a late review step — consent cannot be granted retroactively for processing already done.

Trade-offsAdvantages & costs

Advantages
  • Prevents the failure that is hardest to remediate, because you cannot un-train or un-disclose.
  • A provenance field costs almost nothing and answers most questions before they become investigations.
  • Minimisation at the boundary reduces both privacy exposure and token cost, which rarely align so neatly.
  • Forces the question of corpus licensing early, where it is cheap to solve.
Trade-offs & costs
  • Cuts off production traffic as an eval source unless consent was designed for at the start.
  • Redaction degrades the very context that makes answers good, so there is a real quality trade.
  • Purpose tracking is discipline rather than technology, and discipline decays without a check enforcing it.
  • Detection of personal data in free text is imperfect, so minimisation reduces exposure rather than eliminating it.

ExampleIn the real world

A support assistant is working well and the obvious next step is to fine-tune on a year of resolved tickets. The tickets are full of customer names, addresses and account details supplied to get a problem fixed. Using them to improve a model is a purpose nobody agreed to. The workable version redacts and samples with a basis, and the useful realisation is that an eval set of two hundred well-chosen cases would have served better anyway.

ToolsHow to implement it

  • Microsoft Presidioopen-source detection and redaction of personal data before the prompt is assembled.
  • A provenance field on every corpus recordsource and permitted purpose, checked at use rather than recalled from memory.
  • Zero-retention provider agreementsthe only meaningful control once data leaves your network for inference.
  • A self-authored or clearly licensed corpusfor demos and kits, this removes the analysis rather than passing it.

Cost & effortWhat it takes

Redaction adds a processing step measured in milliseconds and a quality cost that is real but usually small. The dominant cost is design: deciding purposes up front and carrying them, which is days of thought and very little code. The alternative cost is a retraining or disclosure event, which is not budgetable.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page OpenAI now confirms Zero Data Retention for eligible API customers and previews Private Safety Processing, affecting what can be promised about data handling.

    OpenAI · 19 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning