Home › What Happens After You Hit Enter › Unicode normalization & text hygiene
Pipeline stage · Operate

Unicode normalization & text hygiene

Before tokenizing, text gets normalised — and invisible characters, smart quotes and lookalike letters either get cleaned up or quietly change the meaning.

In one line

Normalization decides the exact bytes the tokenizer sees, so a character nobody can see can change your token count, your cache key, and the model's instructions.

Why you'd careThe thing you have already noticed

You pasted a paragraph out of a PDF and the token count came back a third higher than the same text typed by hand. You sent what you were certain was an identical prompt twice and the second one missed the cache. You watched a model follow an instruction that appears nowhere in the visible text. These are the same stage. Between decoding the bytes off the wire and handing a string to the tokenizer, something either normalises the text or does not, and the difference between a non-breaking space and a space, or a Cyrillic а and a Latin a, is invisible to you and decisive to everything downstream.

In and outWhat goes in, what comes out

InDecoded strings from every source that reaches a prompt — typed input, clipboard paste, OCR output, scraped HTML, tool results, file contents — carrying combining marks, non-breaking spaces, soft hyphens, zero-width joiners, bidirectional controls, smart quotes and homoglyphs borrowed from other scripts.
ProcessApply a normalisation form: NFC composes canonically equivalent sequences, NFKC additionally folds compatibility variants such as ligatures and full-width forms. Then scan for format-category and bidi control codepoints, strip or flag them, and optionally collapse runs of whitespace.
OutA canonical UTF-8 byte string, byte-identical for inputs that differed only in encoding, plus a list of spans that were removed or flagged as suspicious so the decision is auditable rather than silent.

Canonical normalisation is round-trippable in principle; compatibility normalisation is not. NFKC turns a superscript two into a digit two, a ligature into two letters, and full-width CJK punctuation into ASCII — meaning changes, quietly. Whitespace collapsing destroys indentation, which destroys Python and YAML. Keep the original bytes if you will ever echo the user's text back verbatim; the normalised copy is for the tokenizer, not for display.

ConceptThe idea underneath

There is no neural network in this stage. It is a canonicalisation problem of the same family as case-folding a database key, and it matters only because of what sits immediately downstream. Unicode permits more than one codepoint sequence for the same rendered text: é is either U+00E9, or U+0065 followed by combining acute U+0301. Normalisation picks one representative per equivalence class. NFC and NFD use canonical equivalence, which preserves meaning; NFKC and NFKD add compatibility equivalence, which does not.

The coupling to the model runs through bytes. A byte-level BPE tokenizer has a merge table learned on a corpus that was itself normalised in some particular way, so text matching that convention compresses into few tokens and text that does not shatters. A regular space is one byte and usually merges into the following word; a non-breaking space is the two bytes C2 A0, rarely appears in the merge table, and typically splits what would have been one token into three. A document full of them — which is exactly what a PDF extractor produces — costs a third more for identical content.

Prompt caching compounds it. Cache lookup is on the exact token id prefix, not on the text, so two strings that render identically but normalise differently are two different cache entries. One smart quote introduced by an editor is enough to turn a warm prefix cold.

Finally, the tokenizer reads what your eyes skip. Codepoints in Unicode's format category render as nothing and still encode as bytes, so a span of zero-width or bidi-override characters can carry text that is invisible in every UI and fully legible to the model. That is the mechanism behind imperceptible-perturbation attacks and behind hidden instructions planted in documents.

At a glanceSee it

Unicode normalization & text hygiene diagram

Raw text is canonicalised, scanned for invisible controls, then handed to the tokenizer as bytes.

The knobsHyperparameters and nuance

  • unicodedata.normalize(form, s)(Python), with form in NFC, NFD, NFKC, NFKD — NFC is the safe default for user text. NFKC is right for search keys and cache keys and wrong for anything you will display back, or that contains code, mathematics or CJK punctuation.
  • --normalization_rule_name(SentencePiece training; default nmt_nfkc, alternatives nfkc, identity, and nmt_nfkc_cf for case folding) — baked into the tokenizer at training time. Whatever it was, your runtime preprocessing should not contradict it.
  • --remove_extra_whitespaces(SentencePiece; defaults to true) — collapses whitespace runs before tokenizing. Left on, indentation-sensitive text is corrupted before the model ever sees it, which is precisely why code-oriented tokenizers turn it off.
  • normalizers.Sequence([...])(Hugging Face tokenizers, and the normalizer block in tokenizer.json) — the composed NFC/NFKC/Strip/Replace pipeline that actually runs at inference. Inspect it rather than assuming; families differ substantially.
  • zero-width and bidi strip policyan application-level denylist, typically Unicode general category Cf plus U+202A to U+202E and U+2066 to U+2069. Strip and you break legitimate emoji ZWJ sequences and correct mixed-direction rendering; keep and the injection channel stays open. Flag-and-log is usually the right middle.

EffectHow this stage moves the answer

Over-normalising changes answers by changing the question. NFKC folding rewrites x² as x2, so a mathematics question quietly becomes a different one; it converts full-width Japanese punctuation to ASCII, which changes tone and sometimes parsing; it merges ligatures, which is fine in prose and wrong in a filename. Whitespace collapsing is the destructive one for developers — strip leading indentation from a Python snippet and the model answers about code that would not run, or reformats the whole thing while guessing at structure. Under-normalising changes answers in the opposite direction. Text carrying hidden format-category characters can instruct the model to disregard your instructions, and the resulting answer looks like an inexplicable model failure: you will read the prompt in your logs, see nothing wrong, and conclude the model is unreliable. The observable signature is behaviour that reproduces only from the original copy-paste.

EvalsWhat it does to your measurements

This stage is almost never in an eval suite, which is precisely why it is worth measuring. It shows up first in cost and cache metrics: tokens-per-request drifts upward as more of your traffic arrives from documents rather than typed text, and cache hit rate falls with no change to your prompt templates. It shows up second in reproducibility — an eval corpus normalised on ingest and a production path that is not are two different distributions, and the offline scores are optimistic by an amount nobody has measured. The specific silent invalidation is fixture drift: eval fixtures stored in a repository get normalised by an editor, a linter or a JSON round-trip, so the exact bytes that produced a recorded failure no longer exist and the bug becomes unreproducible. Store fixtures as escaped byte sequences, and assert on token counts rather than text equality.

Failure modesWhen it goes wrong

  • Token count much higher for pasted text than for the same text typednon-breaking spaces, soft hyphens and ligatures from a PDF or word processor do not match the merge table.
  • An identical-looking prompt misses the prompt cachecache identity is the token id prefix, and a smart quote or a differently-composed accent produces different ids.
  • The model obeys an instruction nobody wrotezero-width or bidi-control characters carrying text that renders as nothing and tokenizes as content.
  • Code answers come back with the indentation guessed atwhitespace collapsing ran before tokenization and removed the structure.
  • Non-Latin text degrades after a preprocessing changeNFKC compatibility folding or homoglyph folding rewrote script-specific forms.

PapersWhere this comes from

  • Bad Characters: Imperceptible NLP AttacksBoucher et al., 2022 (arXiv:2106.09898). Showed that invisible characters, homoglyphs, reorderings and deletions can change an NLP system's output while leaving the rendered text unchanged to a human reader, establishing text hygiene as a security control rather than a tidiness measure.
  • Trojan Source: Invisible VulnerabilitiesBoucher and Anderson, 2021. Applied the same bidirectional-override trick to source code, so that a reviewer and a compiler see different programs; the identical mechanism applies to any document a model reads.
  • Unicode Standard Annex #15, Normalization FormsUnicode Consortium, maintained. The normative definition of NFC, NFD, NFKC and NFKD, including which transformations preserve meaning and which do not. A specification rather than research, and the correct authority for this stage.
  • Unicode Technical Standard #39, Unicode Security MechanismsUnicode Consortium, maintained. Defines the confusables data and restriction levels used for homoglyph detection; it is where a defensible strip-or-flag policy comes from instead of an ad-hoc denylist.
A living map of modern AI — kept current every morning