UC0524 · store-sensitivity · what a customer actually tunes at deployment

Move the slider.
Watch what stops leaving the building.

Nothing on this page was typed by hand. The right-hand column is produced by running the masker over this kit's real corpus — 440 documents across 64 stores — and the page is built from its output file. No model is involved in masking at any point: it is pattern substitution, deterministic, free, and identical on every run.

The deployment dial

Five levels, each a superset of the one below. Move it and every pair underneath re-renders from the masker's own output.

Nothing maskedNothing sent
L0 · RAW
None
The document as held. Nothing is replaced.
L1 · LIBERAL
Identifiers only
Replace the three things that are dangerous on their own — a government identifier, a payment instrument, a live secret. Names, addresses and prose all survive.
L2 · BALANCED
Personal data
Everything that identifies a person goes. The sentence around it stays, so a judgement about what the sentence says is still possible.
L3 · STRINGENT
Stringent
Also every date and every figure. Almost nothing quantitative leaves the building.
L4 · MAXIMUM
Metadata only
No document content leaves at all — only the one-word kind of each document.

Every category, real against its decoy

Each pair is rendered by the corpus generator from one skeleton — same opening, same layout, same vocabulary — so only the entity in the slot separates them. That is what makes the test below sharp: mask both and see whether they are still different strings.

Where the patterns come from

This is the answer to “how would you know what to mask without seeing my data?”. You do not need to. These are published specifications, shipped as vertical packs and switched on per deployment.

tokenpackwhere the shape comes from
[EMAIL]coreRFC 5322 local@domain
[PHONE]coreITU E.164 / national trunk form
[NAME]coreDEMO ONLY — corpus shape — production uses NER or a gazetteer, not a capital-letter rule
[ADDRESS]coreDEMO ONLY — corpus shape — production uses the postal authority's thoroughfare list
[DATE]coreISO 8601
[MONEY]coreDEMO ONLY — corpus shape
[CARD]paymentsISO/IEC 7812 PAN; production adds a Luhn check and published BIN ranges
[SECRET]secretsvendor key prefix convention
[GOV-ID]us-personUS SSA area-group-serial form
⛔ A masker recognises; it does not discover. Anything outside the enabled packs survives into the prompt untouched. Masking is risk reduction, never a guarantee — which is exactly why the in-perimeter option has to stay on the table.

Where 480 comes from — and why it is not your number

sum over stores of (categories carried by that store's schema revision)

This corpus48032 stores × 7 categories + 32 stores × 8 categories
TransfersTHE FORMULAnot the number
Grows with3 thingsmore stores in the register; more categories in the schema; a schema revision that carries categories an older one did not
the FORMULA and the method transfer. THE NUMBER DOES NOT — a customer's cell count is their own register times their own schema, and an industry pack changes the category set, not the arithmetic.

Every masked string above is this code's output. The page is generated by 2026-09-23-masking-lab.py from the JSON that 2026-09-23-masking-levels.py writes while running the masker over data/corpus/*.txt. There is no hand-written “after” column anywhere on it.

⚠︎ Separability is a ceiling, not a promise. That a real case and its decoy remain different strings after masking proves no reader is prevented from telling them apart. It does not prove a model will — the measured run scores 450 of 480, below its own ceiling. Proving the floor under masking needs a fresh paid run and is not claimed here.

The corpus is synthetic, generated from a fixed seed. Every name, identifier, card number, address and email in it is fabricated. The schema DS-CLASS-2026 is invented for this kit and is not a regulator's rule.