Most cleaning is invisible until it is missing, when two identical-looking strings fail to match because one carries a different Unicode code point.
ConceptWhat it is
Cleaning makes text uniform: consistent Unicode normalisation, consistent whitespace, consistent quotation marks and dashes, and the removal of artefacts that extraction could not avoid. It is unglamorous and it decides whether exact matching, deduplication and keyword search work at all.
The characteristic failure is that two strings look identical on screen and differ in bytes. An accented character can be one code point or two; a quotation mark can be one of several; a non-breaking space is not a space. Every one of those breaks an exact match while remaining invisible to the person debugging it.
How it worksThe mechanics
A single normalisation form is chosen and applied everywhere — the same form at ingest, at query time, and anywhere a comparison happens. Applying it in one place and not the others is worse than not applying it, because it manufactures mismatches between the index and the query.
Beyond normalisation, cleaning removes what extraction left behind: repeated page headers, footers, page numbers, navigation fragments, control characters. The rule that keeps this safe is conservatism — a cleaner that strips too much destroys meaning, and unlike a missing normalisation, that damage is not recoverable downstream.
At a glanceSee it
The same normalisation form has to apply at ingest and at query time. Applying it on one side only manufactures the mismatches it exists to prevent.
When to use itWhere it fits
- Always, and identically on both the indexing and the query path.
- Before any deduplication or exact-match step, which cannot work on unnormalised text.
- For any corpus mixing sources, where each source brings its own conventions.
- Wherever keyword or hybrid search is used, since lexical matching is the most sensitive to this.
When NOT to use itLimits & anti-patterns
- Aggressively, where the removed detail carries meaning — code, identifiers and legal text all suffer from over-cleaning.
- As a substitute for fixing extraction; cleaning cannot recover a table that arrived as a paragraph.
- Case-folding or stripping accents for a language where they are semantic.
- Differently on the two paths, which is the single most common way this stage is got wrong.
Trade-offsAdvantages & costs
Advantages
- Very cheap and makes matching, deduplication and lexical search behave predictably.
- Removes a whole class of bug that is otherwise extremely hard to see.
- Standard library support means almost none of it needs writing.
- Improves chunk quality by stripping repeated furniture that would otherwise dominate short chunks.
Trade-offs & costs
- Over-cleaning destroys meaning irreversibly, and the loss is discovered much later.
- Silent when wrong — the symptom is a match that does not happen, which nothing reports.
- Language-specific decisions do not generalise, so a multilingual corpus needs care per language.
- Easy to apply inconsistently across paths, which creates the exact defect it prevents.
ExampleIn the real world
A retrieval system fails to find a supplier by name, though the name is plainly in the corpus. The document contains a non-breaking space between the two words and the query contains an ordinary one. Semantic search still half-works, so the failure is intermittent rather than total, which made it look like a ranking problem for a week.
ToolsHow to implement it
- Python unicodedata.normalizeone chosen form, applied on both the index and the query path.
- ftfyrepairing text that was already mis-decoded before it reached you.
- A shared normalisation functionone importable implementation, so the two paths cannot drift apart.
- Conservative regex for furnituretargeted removal of known headers and footers rather than broad stripping.
Cost & effortWhat it takes
Negligible compute, measured in microseconds per document. Engineering effort is small and front-loaded: decide the form, put it in one function, and call it from both paths. The expensive version is discovering months later that the two paths normalise differently and re-indexing the corpus.