Home › Data for AI › Enrich
🗄️ · Build

Enrich

Adding the metadata that makes filtering, permissions and citation possible later.

In one line

Metadata you do not attach at ingest is metadata you cannot filter on at query time, and adding it later means re-indexing everything.

ConceptWhat it is

Enrichment attaches information to a record that was not in its text: source, owner, date, classification, access tags, document type, language, and whatever the domain needs. It is the last ingest stage and the one whose absence is felt furthest downstream.

The reason it matters more than it looks is that retrieval filters can only use fields that exist. Access control at retrieval time, tenant scoping, date-bounded questions and citation back to a source document are all downstream of a decision made here, and none of them can be retrofitted without rebuilding the index.

How it worksThe mechanics

Metadata that already exists is carried rather than recreated — file owner, permissions, modified date, source URL, the licence determination. That carrying is the part most often lost, because each transform stage can drop fields it does not understand, and nothing reports the loss.

Derived metadata is added where it earns its place: language detection, document type classification, entity extraction, a summary used for contextual retrieval. Each addition costs ingest time and index size, so the test is whether a query or a filter will actually use it, not whether it is interesting.

At a glanceSee it

Enrich diagram

Carrying existing metadata is the half that gets lost, silently, one transform at a time. Everything a filter can do later is decided here.

When to use itWhere it fits

  • Always, for anything that will be filtered, permissioned, cited or scoped by tenant.
  • Wherever access control depends on source metadata, which has to survive every earlier stage.
  • When questions are naturally bounded by time, department or document type.
  • For contextual retrieval, where a generated blurb about each chunk measurably improves recall.

When NOT to use itLimits & anti-patterns

  • Deriving metadata nothing will filter on, which costs ingest time and index size for nothing.
  • As a substitute for carrying real metadata; a classifier's guess at ownership is not ownership.
  • Expensive per-record model calls at ingest, unless the retrieval gain has been measured.
  • After indexing, which is not enrichment but a rebuild.

Trade-offsAdvantages & costs

Advantages
  • Makes retrieval-time filtering possible at all, which is the foundation of access control.
  • Enables citation back to a real source, which is most of what makes an answer trustworthy.
  • Cheap for carried metadata, since the information already exists and only has to survive.
  • Contextual blurbs are one of the few interventions with a well-measured retrieval gain.
Trade-offs & costs
  • Model-derived metadata costs a call per record, which at corpus scale is a real number.
  • Fields dropped by an intermediate stage vanish silently, with no error anywhere.
  • Index size grows with metadata, and highly selective filters can slow search rather than speed it.
  • Derived metadata can be wrong, and a wrong access tag is worse than a missing one.

ExampleIn the real world

A retrieval system needs to answer questions scoped to a single department. The corpus was built from a file share where every document sat in a departmental folder, so the information existed at ingest — and the chunking step kept only the text. Adding the field meant re-extracting, re-chunking and re-embedding the entire corpus for a value that had been in the file path all along.

ToolsHow to implement it

  • A metadata schema defined before ingestso every stage knows which fields it must pass through untouched.
  • Contextual retrieval blurbsa short generated summary per chunk, one of the few well-measured recall improvements.
  • Language and type detectioncheap classifiers that add genuinely useful filter dimensions.
  • A field-survival assertiona validation check that required metadata is still present after the last stage.

Cost & effortWhat it takes

Carried metadata is free. Derived metadata costs whatever the deriving costs, and a model call per chunk over a large corpus is the one line here that can become significant — measure the retrieval gain before buying it. The expensive mistake is omission, since the remedy is a full re-index.

A living map of modern AI — kept current every morning