Extraction is where most retrieval systems quietly lose their quality, and tables are where extraction loses most.
ConceptWhat it is
Extraction converts a source format into text a pipeline can work with. It sounds mechanical and it is the stage with the largest unmeasured effect on final answer quality, because everything downstream inherits whatever it produces.
The reason is that formats encode presentation. A PDF stores glyph positions; a spreadsheet stores cells with no statement of what is a header; a slide stores boxes. Meaning that was obvious to a human reader has to be reconstructed, and where it is not, the model receives something that looks like text and reads like noise.
How it worksThe mechanics
Files are routed by type to an extractor built for that type, and the extractor's job is to emit structure rather than a flat string: headings marked as headings, tables as rows and columns, lists as lists, and page furniture discarded. Reading order matters particularly for multi-column layouts, where naive extraction interleaves two columns into nonsense.
Because failures are silent, the stage needs its own evidence. A sample of extracted output is read by a person before the corpus is trusted, and cheap invariants run over the rest — a document that yields far less text than its page count suggests, or a table that emerged as a single line, is flagged rather than indexed.
At a glanceSee it
Extraction fails silently, so it needs its own evidence: cheap invariants over everything, and a sample read by eye before the corpus is trusted.
When to use itWhere it fits
- Always, for any non-plain-text source — this is a stage, not an option.
- With extra care wherever tables carry the answer, which is most policy and finance material.
- Before measuring anything else, because an extraction defect looks exactly like a model defect.
- When answer quality is disappointing and the prompt has already been rewritten twice.
When NOT to use itLimits & anti-patterns
- As a single parser applied to every format, which is how the long tail silently degrades.
- Without inspecting output, since no automated check reliably detects a mangled table.
- For scanned documents, without deciding explicitly to pay for OCR and to accept its errors.
- As a place to also clean and normalise; keeping extraction separate is what makes a re-run cheap.
Trade-offsAdvantages & costs
Advantages
- The highest-leverage quality work available, and usually the least glamorous.
- Structure preserved here makes chunking, filtering and citation possible downstream.
- Mature open tooling handles the common formats well.
- Fixes here improve every question at once, unlike a prompt change tuned to one.
Trade-offs & costs
- Fails silently, producing plausible-looking text that is subtly wrong.
- Quality varies enormously by format and by document, so a sample can mislead.
- Tables remain genuinely hard, and no extractor is reliable across all of them.
- The long tail of formats consumes far more engineering time than the common cases.
ExampleIn the real world
An assistant over benefits documents answers coverage questions confidently and wrongly. The source is a PDF whose eligibility matrix extracts as one continuous paragraph, so row and column relationships are gone and the model pairs the wrong plan with the wrong limit. Nothing in the prompt, the retrieval or the model was at fault, and none of the gates could see it.
ToolsHow to implement it
- Unstructured or Doclingstructure-aware extraction across many formats, emitting elements rather than a flat string.
- PyMuPDFprecise control over PDF text and layout, including reading order in multi-column pages.
- Camelot or Tabulatable-specific extraction for the cases a general extractor mangles.
- Length and density invariantscheap automated flags for documents that yielded suspiciously little text.
Cost & effortWhat it takes
Compute is modest for text formats and significant for OCR. The real budget line is engineering attention and human reading time, which is the part most often skipped and the part that finds the defects. A day spent reading extracted output early is routinely worth more than a week of prompt iteration later.