The business caseThe problem this solves
A do-not-contact register only means anything if somebody checks it against what actually went out. That check is a person opening two systems side by side — the register, and every sending system's contact log — for one subject at a time, and deciding two things per subject that no query can decide for them: which of the entries filed under that surname is actually this person, and which of the messages that went out is a purpose the register was supposed to stop. Everything after those two decisions is calendar arithmetic and set membership. The queue is therefore expensive for the wrong reason: it is mostly free work, gated behind two judgements. Nothing. It is a first read that puts a packet in front of a reviewer with the evidence attached. It does not add anybody to a register or remove anybody from one, does not extend, shorten, end or reinstate a suppression, does not release a channel or resume marketing, does not contact anybody, does not open or close an account, does not state that two records are the same human being, and does not make, record or report a finding of breach against any person or operator.
Audience
A compliance reviewer working a reconciliation queue, and the person who has to decide whether to buy a model for it. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual reconciliation packets
The corpus is 60 reconciliation packets, 0.12 MB (txt 60). Because the two readings had to be separable from the arithmetic, and no real register data could be used to do it. The corpus is built so that the name reading and the purpose reading each fail in a DIFFERENT way that pure code cannot recover: 16 name-variant packets (diminutives, married surnames, compound surnames, dropped diacritics, initials, and two relatives on one household account) and four purpose styles running from the vocabulary's own words to a sentence carrying a DIFFERENT purpose's vocabulary. ⚠︎ AND THE SURFACE FORM VARIES INDEPENDENTLY OF THE LABEL, WHICH IT DID NOT AT FIRST: the first cut held one stem per purpose per style — measured, 26 distinct sentences across 182 events, most repeated 17 times, a closed phrase list a keyword floor memorises rather than reads. Each (purpose, style) now carries six stems, re-measured at 82 distinct sentences most repeated 8 times, with vocabulary deliberately shared ACROSS labels — statement appears under three purposes, credit under two. data/SOURCES.md states all of it.
The corpus
- The 60 reconciliation packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — every one of the 60 packets, the register index and the whole answer key is generated. Nothing here is fetched, scraped, licensed or derived from anything that was, and there is no personal data in it anywhere: evals/check_labels.py sweeps every packet for seven real-identifier shapes (email, telephone, card-length digit runs, national-insurance and social-security patterns, postcodes and any mention of a date of birth) and reports 0.
Swap this folder for your own material and the kit is pointed at your reconciliation packets. That is the whole change — there is no database to migrate.
CONTACT SUPPRESSION RECONCILIATION PACKET XCR-0001
Alderhaven Group - register REG-CS-9 - standard CSR-2026 as at 2026-09-05
review period 2026-01-01 to 2026-09-05 - review date 2026-09-05
SUBJECT UNDER REVIEW as the contact log records them
name in the log Hana Alderby
account ref ACC-4011
properties seen PR-CORE
log note none
REGISTER ENTRIES ON FILE narrowed in code to the entries the register held under this surname
RG-2370 registered name Marcus Alderby
RG-2370 effective from 2026-02-03
RG-2370 duration indefinite
RG-2370 scope channels all
RG-2370 scope properties PR-CORE, PR-LIVE
RG-2370 entry status in force
RG-2370 register note none
RG-5426 registered name Hana Alderby
RG-5426 effective from 2026-02-08
RG-5426 duration 12 months
RG-5426 scope channels EMAIL, SMS, POST, PUSH, IN-APP
RG-5426 scope properties PR-CORE, PR-LIVE
RG-5426 entry status in force
RG-5426 register note none
RG-9407 registered name Selim Alderby
RG-9407 effective from 2026-03-04
RG-9407 duration 5 years
RG-9407 scope channels EMAIL, SMS, POST, PUSH, IN-APP
RG-9407 scope properties all
RG-9407 entry status in force
RG-9407 register note none
CONTACT EVENTS RECORDED AGAINST THIS ACCOUNT IN THE REVIEW PERIOD
EV-0101 date 2026-02-19
EV-0101 channel EMAIL
EV-0101 property PR-CORE
EV-0101 purpose as logged promotional offer, 25 off the next order, one use, expires in fourteen days
EV-0101 reference SND-1013
EV-0102 date 2026-01-31Abridged — the file continues.
The outcomeWhat a good result looks like
Every packet comes back with the two readings named, the breach set re-derived from the standard in code, the verdict, and ONE line from the packet that the verdict turns on — so a reviewer can see the evidence in a glance rather than taking the answer on somebody's word. On this corpus, 53 of 60 packets come back with all five fields correct after the free station, against 33 of 60 for a free floor that costs nothing and 26 of 60 for the raw reply.
And when it cannot
A breach the arm does not name is a contact that reached somebody who had asked not to be contacted and that nobody will now look at. The paid arm leaves 7 of the 27 breaches unnamed after the station; the free rules floor leaves 14. Both numbers are published, and this kit never averages them against false flags — CSR-2026 section 7 says the two errors are not the same size and there is no F-score anywhere in it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your register entries and your contact log agree on the name, and your campaign metadata uses your own vocabulary's words — the free rules floor, alone
on this corpus that arm answers 55.0 pct of packets completely correctly at $0.00 with no network, and it BEATS the raw paid reply's 43.3 pct. There is no measured case here for spending anything on the easy packets. - Your register was keyed years ago and your contact log is full of familiar forms, married surnames and household accounts — the paid call for the ENTRY reading, then the free station for everything after it
this is the measured gap: 95.0 pct against the floor's 88.3 pct on the entry, and on the married-surname family the floor answers NO-ENTRY outright while the paid arm reads the log note and resolves it. - Your campaign metadata describes what a message DOES rather than naming its purpose, or carries another purpose's vocabulary — the paid call for the PURPOSE reading — the widest measured gap in the kit
69 of 70 suppressed contacts found against the floor's 47 of 70, and the floor leaves 14 of the 27 breaches unnamed against the paid arm's 7. - You need this to be a control rather than a first read — unknown from this kit — it is NOT demonstrated here
7 of the 27 breaches are still unnamed after the station, the station itself is wrong on every minimum-term entry, and the whole thing is measured on an invented standard over an invented corpus with the candidate entries already narrowed.
And where nothing here is good enough:
- Your register durations are conditional — minimum terms, renewals, 'until the subject applies in writing' — NEITHER, until src/policy.py::duration_end is replaced or taught to abstain
the pure-code station is wrong on every such packet on EVERY arm, including both free floors. 3 of 3 here, and the model had all 3 right before the station overrode it — this is the one place the free half is the weak half.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-06. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/purposes.json with your own contact-purpose taxonomy and data/policy.md and data/policy.json with your own standard, then drop your packets into data/corpus/ in the same column shape. ⚠︎ DO NOT PUT REAL REGISTER DATA IN IT. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a reading that a normalised string match and a keyword table already produce. That is the case against the best-fitting scenario (“Your register entries and your contact log agree on the name, and your campaign metadata uses your own vocabulary's words”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet that is not a fixed-width column table. Values are padded and separated by runs of two or more spaces, and the generator asserts on every build that no value contains two consecutive spaces — one that did would split into another column. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | THE KEY'S TWO READINGS ARE THE GENERATOR'S STRUCTURAL TRUTH. evals/check_labels.py re-derives the verdict, the breach set and the citation independently and reports 0 disagreements — but it takes the PURPOSE CODE from the key, because recovering it from the sentence is exactly the reading the paid call is bought for. 4 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-06 — r001-exclusion-contact. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, then python3 -m src.app. Nothing to install, no key needed: the 60 packets, the answer key, both free floors' results, two paid runs' results and the adversarial run's results all ship, and every one of them replays off disk for $0.00. python3 -m evals.check_labels re-derives the whole key from the packets in about a second.




