Home › Use Cases › Screen one item against a schedule of exemptions, condition by condition
Use caseUC0229
🧪 Use-case kit · runnable

Screen one item against a schedule of exemptions, condition by condition

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

Screening an item against a schedule of exemptions is not one yes-or-no. A schedule carries several exemptions, each with numbered conditions, and an item meets some of them. The reviewer's answer has four parts and the interesting one is the third: which exemptions qualify, which numbered condition decided each verdict, which exemptions cannot be decided AT ALL because the dossier is silent on a condition they turn on, and -- when more than one qualifies -- which is cheapest to hold. A screen with only two verdicts has to lie about the third class. It will pick the same lie every time: on this corpus that is 62 of 480 rows written off as failures. Nothing is replaced. It replaces the FIRST PASS of reading a dossier against a schedule -- the part where somebody works eight exemptions one at a time and writes down which condition ended each. It does not file, does not claim, does not decide, and produces no regulatory advice: the schedule, the registry, the substances and the lists in this kit are all invented and restate nothing.

Audience

Whoever confirms the screening record before it is filed -- a registry reviewer, a compliance desk, a consignor's own regulatory affairs team. The output is a record a person checks, with the deciding condition printed beside every verdict so the reasoning can be checked rather than believed. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual item dossiers, one screening record each

The corpus is 60 item dossiers, one screening record each, 0.07 MB (csv_row 15 · dossier_memo 15 · registry_form 15 · supplier_packet 15). There is no public corpus of exemption dossiers, and the reason is more than 'nobody collected one'. An exemption dossier is a filing by a party asking for relief: it names them, states what they hold and in what quantity, and exists to persuade a regulator. Where such filings are published at all they are published redacted, in the jurisdiction's own format, against that jurisdiction's own schedule -- and the schedule is the thing that would have to be reproduced to grade anything. That would matter less if the PROSE were what is being measured. It is not. This key needs three facts that cannot be recovered from a document by reading it. (1) WHICH CONDITION IS DECISIVE -- the key does not say EX-03 fails, it says EX-03 fails on EX-03.c1, which needs an agreed schedule, an agreed precedence and an agreed tie-break, three decisions no external corpus carries. (2) WHICH DOSSIERS ARE DELIBERATELY SILENT -- 62 of 480 rows stand on a fact being stated NOWHERE, and proving a real document is silent about something is proving a negative over a document you did not write; here it is a fact of construction, 51 particulars withheld at generation time and every one re-checked by searching the rendered text for it. (3) WHICH EXEMPTION IS CHEAPEST TO HOLD -- the claim rule ranks by BURDEN, a property of the schedule rather than of the item. ⚠︎ AND THE GENERATOR COSTS THE MEASUREMENT MORE THAN IT COSTS THE KEY. Every narrative sentence comes from one table of exactly two phrasings per field. A regex reader built from those stems was written and run against this corpus THIS TURN and recovers 43 of 44 prose-only facts, taking the free floor to 480/480 rows and 60/60 dossiers -- past the paid arm. Separately, 37 of those 44 facts are physical_form, WHICH NO CONDITION IN THE SCHEDULE TESTS; only 7 decide anything. And the intake record is unrealistically clean -- every declared particular is already a typed field with a legal value, so the floor is handed a perfect reading and only has to apply the rulebook. Since the floor is the column this kit exists to beat, an unrealistically strong floor is the conservative error to make.

The corpus

  • The 60 item dossiers, one screening record eachgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your item dossiers, one screening record each. That is the whole change — there is no database to migrate.

One item dossiers, one screening record each, as the model receives itIT-0001.txt · 1 of 60
NORTHGATE SUBSTANCE REGISTRY (INVENTED)
ITEM SCREENING FORM -- Registry Notice 12
==============================================================
  Dossier reference     IT-0001
  Consignor             Portsoy Compounding
  Delivery site         Building C, Ellersby Science Park
  Filed                 2026-07-21

DECLARED PARTICULARS
  Intended use                               research
  Annual quantity (kg)                       3.5 kg
  Per-shipment quantity (kg)                 22 kg
  Physical form                              solid
  Packaging                                  ampoule
  Max listed-component concentration (%)     12 %
  Composition                                NSR-04255 (Tellurate ester GX), NSR-04701 (Pelmerol), NSR-09340 (Erbanite)
  Transport mode                             sea
  Prior registration                         none
  Registration status                        pending
  Contained system                           yes
  Site licence held                          no
  Disposal route declared                    no
  Sample or trial only                       yes
  End user declared                          yes
  Notice given before dispatch               yes
  Resale permitted                           yes

NARRATIVE
  No further narrative.

Signed for the consignor. This form is the consignor's own declaration.

The outcomeWhat a good result looks like

One screening record per dossier: 8 verdicts, the condition that settled each, the exemption to claim and the cheaper exemptions a single missing fact is standing in front of. Measured on r001: 479 of 480 verdicts (99.8%), 56 of 60 dossiers right on all four graded fields (93.3%), 60 of 60 claims, and silence read as satisfaction ZERO times.

And when it cannot

⚠︎ THE HEADLINE DOES NOT SURVIVE ITS OWN CORPUS. Against the strong free floor -- the same intake record and the same engine, allowed to answer NOT STATED -- the paid arm wins 9 verdict rows and loses 1 (exact McNemar p = 0.0215, verified this turn), and at the dossier level it is 7 against 4, p = 0.549. 93.3% versus 88.3% IS NOT A GAP; it is noise, and must not be quoted as one. Every one of the 9 rows and all 7 dossiers the model wins are the same planted case, prose_only_fact: 7 dossiers in 60.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A screening population whose particulars already arrive as structured intake data — a register or a form-capture system that records each declared field with its value, and OMITS a field the filer did not declare — the free rules floor in NOT-STATED mode, and no model at all
    It scores 471 of 480 verdicts, 53 of 60 dossiers, 62 of 62 silences, 26 of 26 non-empty better-if-resolved sets and 438 of 442 citations, for $0.00 in 0.08 s over the whole corpus — and it BEATS the paid arm on three of those and ties it on two more (claiming nothing when nothing qualifies, 33/33; two-or-more-qualifying, 9/9). Every piece of arithmetic in the job is already free, because no arm does it: they all hand their reading to the same src/screen.py.
  • An intake form that does not capture everything the schedule tests, so a decisive particular sometimes appears only in a covering memo, a remarks block or an annex — the model, published RAW
    This is the one thing the money buys and it is worth naming exactly: 7 of 7 on prose_only_fact where BOTH free floors score 0 of 7. Paired row by row against the strong floor the model wins 9 and loses 1 (p = 0.0215) and every one of the 9 is that case. ⚠︎ AND PUBLISH RAW, NOT RECHECKED. The pure-code station is a net loss here — it turns three correct QUALIFIES into INSUFFICIENT_INFORMATION by re-deriving from a reading whose prose-only field came back null.

And where nothing here is good enough:

  • A corpus whose narratives are templated, generated, or drawn from a small set of house phrasings — neither — write the regex reader
    Measured on THIS corpus this turn: a reader built from the generator's two-phrasings-per-field table recovers 43 of 44 prose-only facts and takes the free floor to 480/480 rows and 60/60 dossiers, past the paid arm, for nothing. If your narratives are that regular, the model is buying you a tenth of a point and a per-call bill.
  • A desk that wants a verdict shipped straight through, with a person only on the exceptions — neither — this kit does not produce a decision
    It produces the record a reviewer confirms, with the deciding condition printed beside every verdict so the reasoning can be checked rather than believed. INSUFFICIENT_INFORMATION is a real answer here rather than a refusal, and 62 rows over 26 dossiers carry it — a desk that auto-ships would have to choose one of the two lies this kit exists to avoid.

At a glanceHow the whole thing runs

93%item all correct pct
51,070 msp50, end to end
$25.60per 1,000 item dossiers · Google Gemini 3 Flash

Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own dossiers as .txt into data/corpus/ and add one row per dossier to data/items.json (id, format, case, hard, target, consignor, site, filed, and a declared block holding ONLY the particulars your intake form actually captured). ⚠︎ OMIT A KEY RATHER THAN NULLING IT. Corpus lens →
When is this the wrong choice?Avoid: Paying per dossier for a set of comparisons a fixed rulebook engine already makes correctly for nothing. And note what the money does NOT buy even here: the two floors and the model are indistinguishable at the dossier level, p = 0.549. That is the case against the best-fitting scenario (“A screening population whose particulars already arrive as structured intake data — a register or a form-capture system that records each declared field with its value, and OMITS a field the filer did not declare”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A dossier that states a particular in a phrasing this reader has not met. The eighteen reading fields are three-valued and the whole kit stands on null meaning 'the dossier does not say'. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?COST IN MONEY. Not measured, and not estimated in the kit. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-31 — r001-exemption-screen. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured this turn on three scratch copies, none of which had a key. python3 -m tools.build_corpus rebuilds the whole corpus in 0.11 s; python3 -m evals.check_labels re-derives all 480 rows independently in 0.03 s and prints 'no disagreement'; each free floor scores the corpus in 0.08 s. Nothing pip-installs -- requirements.txt names nothing and the standard library is the entire dependency list. ⚡ AND THE REBUILD IS BYTE-IDENTICAL, MEASURED RATHER THAN CLAIMED: three rebuilds at PYTHONHASHSEED=1, 97531 and 424242, each diffed recursively against the committed data/ directory, clean every time. The re-run free floors also reproduce the committed result files' score blocks exactly. That check exists because a sibling kit in this batch claimed byte-identical rebuilds and did not have them.

A living map of modern AI — kept current every morning