The business caseThe problem this solves
Screening an item against a schedule of exemptions is not one yes-or-no. A schedule carries several exemptions, each with numbered conditions, and an item meets some of them. The reviewer's answer has four parts and the interesting one is the third: which exemptions qualify, which numbered condition decided each verdict, which exemptions cannot be decided AT ALL because the dossier is silent on a condition they turn on, and -- when more than one qualifies -- which is cheapest to hold. A screen with only two verdicts has to lie about the third class. It will pick the same lie every time: on this corpus that is 62 of 480 rows written off as failures. Nothing is replaced. It replaces the FIRST PASS of reading a dossier against a schedule -- the part where somebody works eight exemptions one at a time and writes down which condition ended each. It does not file, does not claim, does not decide, and produces no regulatory advice: the schedule, the registry, the substances and the lists in this kit are all invented and restate nothing.
Audience
Whoever confirms the screening record before it is filed -- a registry reviewer, a compliance desk, a consignor's own regulatory affairs team. The output is a record a person checks, with the deciding condition printed beside every verdict so the reasoning can be checked rather than believed. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual item dossiers, one screening record each
The corpus is 60 item dossiers, one screening record each, 0.07 MB (csv_row 15 · dossier_memo 15 · registry_form 15 · supplier_packet 15). There is no public corpus of exemption dossiers, and the reason is more than 'nobody collected one'. An exemption dossier is a filing by a party asking for relief: it names them, states what they hold and in what quantity, and exists to persuade a regulator. Where such filings are published at all they are published redacted, in the jurisdiction's own format, against that jurisdiction's own schedule -- and the schedule is the thing that would have to be reproduced to grade anything.
That would matter less if the PROSE were what is being measured. It is not. This key needs three facts that cannot be recovered from a document by reading it. (1) WHICH CONDITION IS DECISIVE -- the key does not say EX-03 fails, it says EX-03 fails on EX-03.c1, which needs an agreed schedule, an agreed precedence and an agreed tie-break, three decisions no external corpus carries. (2) WHICH DOSSIERS ARE DELIBERATELY SILENT -- 62 of 480 rows stand on a fact being stated NOWHERE, and proving a real document is silent about something is proving a negative over a document you did not write; here it is a fact of construction, 51 particulars withheld at generation time and every one re-checked by searching the rendered text for it. (3) WHICH EXEMPTION IS CHEAPEST TO HOLD -- the claim rule ranks by BURDEN, a property of the schedule rather than of the item.
⚠︎ AND THE GENERATOR COSTS THE MEASUREMENT MORE THAN IT COSTS THE KEY. Every narrative sentence comes from one table of exactly two phrasings per field. A regex reader built from those stems was written and run against this corpus THIS TURN and recovers 43 of 44 prose-only facts, taking the free floor to 480/480 rows and 60/60 dossiers -- past the paid arm. Separately, 37 of those 44 facts are physical_form, WHICH NO CONDITION IN THE SCHEDULE TESTS; only 7 decide anything. And the intake record is unrealistically clean -- every declared particular is already a typed field with a legal value, so the floor is handed a perfect reading and only has to apply the rulebook. Since the floor is the column this kit exists to beat, an unrealistically strong floor is the conservative error to make.
The corpus
- The 60 item dossiers, one screening record eachgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your item dossiers, one screening record each. That is the whole change — there is no database to migrate.
NORTHGATE SUBSTANCE REGISTRY (INVENTED)
ITEM SCREENING FORM -- Registry Notice 12
==============================================================
Dossier reference IT-0001
Consignor Portsoy Compounding
Delivery site Building C, Ellersby Science Park
Filed 2026-07-21
DECLARED PARTICULARS
Intended use research
Annual quantity (kg) 3.5 kg
Per-shipment quantity (kg) 22 kg
Physical form solid
Packaging ampoule
Max listed-component concentration (%) 12 %
Composition NSR-04255 (Tellurate ester GX), NSR-04701 (Pelmerol), NSR-09340 (Erbanite)
Transport mode sea
Prior registration none
Registration status pending
Contained system yes
Site licence held no
Disposal route declared no
Sample or trial only yes
End user declared yes
Notice given before dispatch yes
Resale permitted yes
NARRATIVE
No further narrative.
Signed for the consignor. This form is the consignor's own declaration.
The outcomeWhat a good result looks like
One screening record per dossier: 8 verdicts, the condition that settled each, the exemption to claim and the cheaper exemptions a single missing fact is standing in front of. Measured on r001: 479 of 480 verdicts (99.8%), 56 of 60 dossiers right on all four graded fields (93.3%), 60 of 60 claims, and silence read as satisfaction ZERO times.
And when it cannot
⚠︎ THE HEADLINE DOES NOT SURVIVE ITS OWN CORPUS. Against the strong free floor -- the same intake record and the same engine, allowed to answer NOT STATED -- the paid arm wins 9 verdict rows and loses 1 (exact McNemar p = 0.0215, verified this turn), and at the dossier level it is 7 against 4, p = 0.549. 93.3% versus 88.3% IS NOT A GAP; it is noise, and must not be quoted as one. Every one of the 9 rows and all 7 dossiers the model wins are the same planted case, prose_only_fact: 7 dossiers in 60.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A screening population whose particulars already arrive as structured intake data — a register or a form-capture system that records each declared field with its value, and OMITS a field the filer did not declare — the free rules floor in NOT-STATED mode, and no model at all
It scores 471 of 480 verdicts, 53 of 60 dossiers, 62 of 62 silences, 26 of 26 non-empty better-if-resolved sets and 438 of 442 citations, for $0.00 in 0.08 s over the whole corpus — and it BEATS the paid arm on three of those and ties it on two more (claiming nothing when nothing qualifies, 33/33; two-or-more-qualifying, 9/9). Every piece of arithmetic in the job is already free, because no arm does it: they all hand their reading to the same src/screen.py. - An intake form that does not capture everything the schedule tests, so a decisive particular sometimes appears only in a covering memo, a remarks block or an annex — the model, published RAW
This is the one thing the money buys and it is worth naming exactly: 7 of 7 onprose_only_factwhere BOTH free floors score 0 of 7. Paired row by row against the strong floor the model wins 9 and loses 1 (p = 0.0215) and every one of the 9 is that case. ⚠︎ AND PUBLISH RAW, NOT RECHECKED. The pure-code station is a net loss here — it turns three correct QUALIFIES into INSUFFICIENT_INFORMATION by re-deriving from a reading whose prose-only field came backnull.
And where nothing here is good enough:
- A corpus whose narratives are templated, generated, or drawn from a small set of house phrasings — neither — write the regex reader
Measured on THIS corpus this turn: a reader built from the generator's two-phrasings-per-field table recovers 43 of 44 prose-only facts and takes the free floor to 480/480 rows and 60/60 dossiers, past the paid arm, for nothing. If your narratives are that regular, the model is buying you a tenth of a point and a per-call bill. - A desk that wants a verdict shipped straight through, with a person only on the exceptions — neither — this kit does not produce a decision
It produces the record a reviewer confirms, with the deciding condition printed beside every verdict so the reasoning can be checked rather than believed. INSUFFICIENT_INFORMATION is a real answer here rather than a refusal, and 62 rows over 26 dossiers carry it — a desk that auto-ships would have to choose one of the two lies this kit exists to avoid.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own dossiers as .txt into data/corpus/ and add one row per dossier to data/items.json (id, format, case, hard, target, consignor, site, filed, and a declared block holding ONLY the particulars your intake form actually captured). ⚠︎ OMIT A KEY RATHER THAN NULLING IT. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per dossier for a set of comparisons a fixed rulebook engine already makes correctly for nothing. And note what the money does NOT buy even here: the two floors and the model are indistinguishable at the dossier level, p = 0.549. That is the case against the best-fitting scenario (“A screening population whose particulars already arrive as structured intake data — a register or a form-capture system that records each declared field with its value, and OMITS a field the filer did not declare”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A dossier that states a particular in a phrasing this reader has not met. The eighteen reading fields are three-valued and the whole kit stands on null meaning 'the dossier does not say'. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | COST IN MONEY. Not measured, and not estimated in the kit. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-exemption-screen. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured this turn on three scratch copies, none of which had a key. python3 -m tools.build_corpus rebuilds the whole corpus in 0.11 s; python3 -m evals.check_labels re-derives all 480 rows independently in 0.03 s and prints 'no disagreement'; each free floor scores the corpus in 0.08 s. Nothing pip-installs -- requirements.txt names nothing and the standard library is the entire dependency list. ⚡ AND THE REBUILD IS BYTE-IDENTICAL, MEASURED RATHER THAN CLAIMED: three rebuilds at PYTHONHASHSEED=1, 97531 and 424242, each diffed recursively against the committed data/ directory, clean every time. The re-run free floors also reproduce the committed result files' score blocks exactly. That check exists because a sibling kit in this batch claimed byte-identical rebuilds and did not have them.



