A database is the easiest source to read and the hardest to reproduce, because it is correct at a moment rather than fixed.
ConceptWhat it is
An operational database is usually the richest source available: structured, already governed, already permissioned. Reading a corpus out of it is a query, which makes it feel like the cheapest option on this list.
The cost arrives later. A database is designed to be current, not to be re-readable. Run the same query next week and you get different rows, so an eval number computed against it cannot be reproduced and a regression cannot be distinguished from the world having moved.
How it worksThe mechanics
The corpus is defined as a query plus a moment: a snapshot taken at a recorded time, written to immutable storage, and given a version. Everything downstream — chunking, embedding, eval — cites that version rather than the live table.
Change data capture is the incremental form. Rather than re-snapshotting, a stream of row changes updates the index, and the version becomes a log position. That keeps the index fresh, but it only preserves reproducibility if the log position is recorded alongside every published measurement.
At a glanceSee it
The snapshot is what makes the number reproducible. Evaluating against the live table measures the world moving as much as the system changing.
When to use itWhere it fits
- When the authoritative answer already lives in a table and re-deriving it from documents would be worse.
- Where row-level permissions already exist and can be carried into chunk metadata rather than reinvented.
- For structured lookups an agent can query directly, where retrieval over prose would be a downgrade.
- When freshness genuinely matters and a nightly snapshot is an acceptable definition of current.
When NOT to use itLimits & anti-patterns
- As a live source under an eval, where the shifting population makes every comparison meaningless.
- When the schema changes frequently and the ingest breaks quietly rather than loudly.
- For text that lives in a blob column and was never modelled — that is a document source wearing a table's clothes.
- When querying production risks the production workload; read from a replica or a snapshot instead.
Trade-offsAdvantages & costs
Advantages
- Structure is already there, so extraction is a query rather than a parsing problem.
- Permissions and ownership usually exist and can be mirrored rather than invented.
- Incremental refresh is well-understood, with change data capture as a mature path.
- Data quality is typically far better than anything parsed out of documents.
Trade-offs & costs
- Not reproducible without an explicit snapshot discipline nobody applies by default.
- Schema drift breaks ingestion silently, and the symptom is a thinner index rather than an error.
- Joins that a human would do mentally have to be materialised before chunking, or context arrives incomplete.
- Text in a database is often terse and context-free, so chunks retrieve badly without enrichment.
ExampleIn the real world
A support assistant retrieves over resolved tickets read live from the ticketing database. Eval scores drift down over a month and the team hunts prompt regressions. Nothing about the system changed; the ticket mix changed, because a product launch shifted what people were asking about. A pinned snapshot would have separated the two questions on the first day rather than the thirtieth.
ToolsHow to implement it
- Debezium or native CDCstreaming row changes into the index instead of re-snapshotting everything.
- Parquet on object storagea cheap, immutable, versioned home for the snapshot the eval will cite.
- dbtexpressing the corpus query as versioned, tested, reviewable code rather than a saved query.
- A read replicaso corpus builds cannot compete with the production workload.
Cost & effortWhat it takes
Storage for snapshots is the only new recurring cost and it is small next to inference. The engineering is a day for a first snapshot pipeline and considerably more for CDC. The cost people miss is the join work: assembling enough context around a terse row that the chunk retrieves usefully is usually more effort than reading the rows.