Lineage is the difference between saying a system scored 82% and being able to show which corpus, which chunks and which run produced that 82%.
ConceptWhat it is
A data catalog is the inventory: what corpora exist, who owns them, what is in them, and what may legally be done with each one. Lineage is the trail across them — which source produced which chunk, which chunk was retrieved for which answer, and which corpus version a published measurement was computed against.
The reason it belongs in ops rather than in a data team's backlog is that an LLM system publishes numbers. An eval score, a cost-per-call, a regression flagged in review — each is a claim about a specific set of bytes. Without lineage the claim survives only as long as nobody asks the follow-up question, and the follow-up question is always which data was this?
How it worksThe mechanics
Every ingested document carries an identifier and a source reference from the moment it is extracted. Chunking preserves it, so each chunk knows its parent document and the parent knows its origin. The retrieval trace records the chunk ids that were actually returned, not just the answer text, which is what turns a bad answer into a question with an address.
The catalog side is a register rather than a pipeline: one row per corpus, carrying its owner, its licence, its refresh cadence and its current version. A published result cites the version. That citation is the whole mechanism — it is what lets a number from six weeks ago be recomputed rather than merely believed.
At a glanceSee it
The identifier set at extraction is what survives all the way to a published number. Break the chain at any stage and the number stops being traceable.
When to use itWhere it fits
- Whenever a measurement is published to anyone outside the team, which is the case that makes it mandatory.
- In any regulated setting, where the question is not whether an answer was good but which record it came from.
- When a corpus is refreshed on a cadence, so a change in score can be attributed to the data rather than the model.
- Before a first eval run rather than after it, because lineage cannot be reconstructed backwards.
When NOT to use itLimits & anti-patterns
- As a reason to build a catalog product; a register with the right columns beats a platform nobody fills in.
- In genuine one-off exploration against a fixed local folder, where the folder itself is the version.
- As a substitute for dataset versioning — lineage says which bytes, versioning is what keeps those bytes retrievable.
- Recording lineage nobody can query; a trail stored where no investigation will reach it is cost with no control.
Trade-offsAdvantages & costs
Advantages
- Turns a disputed number into a recomputation, which ends arguments rather than extending them.
- Makes a bad answer diagnosable — the retrieved chunks are named, so the failure is either retrieval or generation.
- Answers the licensing and privacy question directly, since the catalog already records what may be used where.
- Costs almost nothing when built at ingest and is close to impossible to add afterwards.
Trade-offs & costs
- Every stage of the pipeline has to carry the identifier, so one careless transform breaks the chain silently.
- The trace grows with traffic and needs a retention policy of its own.
- A catalog decays the moment it stops being required by the code that reads corpora.
- It records what happened without judging it; lineage tells you which chunks were retrieved, not that they were wrong.
ExampleIn the real world
A retrieval assistant is reported to have got materially worse over a month, and the obvious suspect is a model upgrade that landed in the same window. Because every answer's trace carries chunk ids and every chunk carries its parent document, the regression is localised in an afternoon: a scheduled corpus refresh had begun pulling a summary export of the source rather than the full records, so the retriever was returning correctly-ranked chunks of a thinner document. The model was never involved.
ToolsHow to implement it
- An id set at extractionthe single field the entire trail hangs from.
- OpenLineage or Marquezwhen lineage has to span pipelines that several teams own.
- A corpus registerowner, licence, cadence and current version, one row per corpus.
- The retrieval tracechunk ids on every answer, which is where most investigations actually start.
Cost & effortWhat it takes
Effort is small at ingest and large retrospectively, which is the whole argument for doing it first. Storage is dominated by the retrieval trace rather than the catalog. The cost that is easy to miss is discipline: a transform that drops the identifier breaks lineage without failing anything, so the field has to be required by the code rather than by a convention.