Home › Data for AI › Files and documents
🗄️ · Build

Files and documents

PDFs, Office files and wiki pages as a corpus — the shape most retrieval systems actually face.

In one line

Documents are the most common corpus and the least uniform, so most of the work is recovering the structure the format threw away.

ConceptWhat it is

Documents are where institutional knowledge actually lives — contracts, policies, reports, handbooks, slide decks. It is the dominant shape for retrieval work; of the fifteen measured kits on this site, most judge a document rather than a row or a conversation.

The difficulty is that a document format encodes appearance, not meaning. A PDF knows where a glyph sits on a page; it does not know that those glyphs are a table header. Everything downstream depends on how much of that lost structure you can recover.

How it worksThe mechanics

Files are collected with their metadata intact — path, owner, permissions, modified date, version — because that metadata is what later makes filtering and access control possible, and it is trivially lost by a naive copy. Each file is then routed by type to an extractor suited to it, rather than through one parser that handles everything badly.

Structure recovery is the substance of the work: headings become chunk boundaries, tables become something other than run-on text, and page furniture like headers and footers is dropped. A document that survives this well produces chunks a reader would recognise as coherent passages, which is the only real test.

At a glanceSee it

Files and documents diagram

Routing by type beats one parser for everything. The test of the whole pipeline is whether a chunk reads as a coherent passage.

When to use itWhere it fits

  • Whenever the knowledge exists as prose that people already read and trust.
  • Policy, legal, clinical and compliance material, where the document is the authority.
  • When the source system offers no API and the export is a folder of files.
  • As the default assumption for a first retrieval build, because this is what most corpora turn out to be.

When NOT to use itLimits & anti-patterns

  • When the same content is available structured, from a database or API, which is almost always better.
  • For scanned images without a deliberate decision to pay for OCR and to accept its error rate.
  • Where documents contradict each other and nobody will say which wins; retrieval will surface both.
  • When superseded versions live in the same folder as current ones, which is the most common corpus defect there is.

Trade-offsAdvantages & costs

Advantages
  • It is where the knowledge actually is, so coverage is high without asking anyone to rewrite anything.
  • Documents carry their own context, so chunks are more self-contained than database rows.
  • File metadata gives you owner, date and permissions almost for free.
  • Incremental updates are easy to detect by modified time or content hash.
Trade-offs & costs
  • Extraction fidelity varies enormously by format, and tables are where it fails worst.
  • Version confusion is endemic — the superseded handbook is usually in the same folder.
  • Layout artefacts survive naive extraction and pollute chunks with headers, footers and page numbers.
  • Scanned documents need OCR, which adds cost, latency and a new error mode.

ExampleIn the real world

A policy assistant is reported as unreliable. Reading a sample of failures shows the model is not hallucinating at all — tables extracted from the source PDFs arrive as run-on text, so a benefits matrix becomes an unreadable sentence and the model answers from it confidently. The defect was in the extractor, and no amount of prompt work would have reached it.

ToolsHow to implement it

  • Unstructured or Doclingtype-aware extraction that preserves headings and table structure rather than flattening everything.
  • PyMuPDFfast, precise PDF text and layout access when you need control over what is kept.
  • Tesseract or a hosted OCR serviceonly where scans are genuinely unavoidable, and with the error rate measured.
  • A content hash per filethe cheapest reliable way to know what actually changed since the last ingest.

Cost & effortWhat it takes

Extraction is cheap per document and adds up at corpus scale; OCR is an order of magnitude more expensive and much slower. The dominant cost is engineering attention on the long tail of formats. Budget for reading extracted output by eye before trusting it, because table damage is invisible to every automated check you are likely to have.

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning