The business caseThe problem this solves
A recovered-commodity facility ships a trailer of baled material against a purchase contract. The buyer grades it inbound, writes its findings in its own words - 'cartons with a slick face that would not tear clean' - and takes a deduction citing one or two of them. Somebody at the shipper's desk then has to say, per cited material, whether the grade specification in force for THIS shipment calls that material prohibitive, an outthrow allowed to a stated share, or does not name it at all; which printed line establishes that; and whether the bale tags the buyer graded are the tags this facility shipped. Today that is a person reading two documents side by side with the contract open between them. The first pass over a contamination deduction packet: matching each cited grading finding to the specification line that classifies it, under the revision in force at the ship date, and deciding whether the bale-tag range the buyer graded is this load.
Audience
Anyone putting an automated reader over a buyer's grading narrative where the expensive mistake is not a wrong dollar figure but a material the contract plainly prohibits read as one the contract never names. The answer this page gives about its own pack is a qualified one: it is ahead on the headline and the lead does not clear 0.05, and on the per-material class it loses to a single fixed answer. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual deduction packets
The corpus is 64 deduction packets, 0.21 MB (md 64). The smallest corpus that makes the interesting mistake unavoidable. A contamination deduction is the one document where the arithmetic is genuinely free for every arm including a constant - the allowance comparison, every percentage, the deduction recompute from the contract's own unit price, the bale-tag overlap and the weight variance are all exact - and the judgement genuinely is not: whether the material the buyer describes in its own words is one the contract's specification prohibits at this grade, under the revision in force on the ship date. So the split between what money buys and what free code does could be NAMED in code before a call was made and SCORED after. It came out inverted, which is the finding.
The corpus
- The 64 deduction packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your deduction packets. That is the whole change — there is no database to migrate.
# Contamination deduction packet DED-7410-0001
facility: FAC-3308
buyer: BYR-233
purchase contract: PC-88646
commodity and grade: baled recovered natural high-density bottle, grade HDN-19
deduction memo: DM-40007
buyer grading reference on this memo: GR-8019
## Outbound weight ticket and bill of lading
o01: bale tags shipped: BT-4046 through BT-4077 (32 bales)
o02: trailer: TRL-0196
o03: net weight shipped: 43,200 lb
o04: departed this facility: 2026-06-19
## Outbound quality audit sheet, shipping shift
q01: The pre-ship check recorded coverings that would not come away dry.
q02: Bale audit sheet: blue and green tops still screwed on the milk jugs.
## Purchase contract grade specification
### Rev A -- in force for shipments from 2025-10-01 to 2026-03-31. NOT IN FORCE FOR THIS SHIPMENT.
Prohibitive at this grade:
s01: lubricant and automotive fluid containers -- prohibitive at any level.
s02: paper facestock labelling and its adhesive -- prohibitive at any level.
Outthrow at this grade, allowed to the share stated:
s03: pigmented closures on natural containers -- allowed to 2.0% by weight.
s04: colour-compounded high-density bottles -- allowed to 0.5% by weight.
s05: polypropylene tubs and pails -- allowed to 2.5% by weight.
### Rev B -- in force for shipments from 2026-04-01. IN FORCE FOR THIS SHIPMENT.
Prohibitive at this grade:
s06: tinted caps and their liners -- prohibitive at any level.
s07: hydrocarbon-residue ware -- prohibitive at any level.
s08: cellulose-based wrap labels -- prohibitive at any level.
Outthrow at this grade, allowed to the share stated:
s09: non-translucent polyolefin containers -- allowed to 0.5% by weight.
## Contract inspection and rejection procedure
p01: moisture tolerance on net weight, this contract: 1.2%.Abridged — the file continues.
The outcomeWhat a good result looks like
One {ref, spec_class, establishing_line} row per cited material - a class from the closed set prohibitive / outthrow / not_in_spec and the printed sNN line that establishes it, or null - plus one {load_identity, identity_line} for the packet. A packet is right only when every cited material is right on both fields AND the load identity is right on both value and line.
And when it cannot
It over-abstains. Across 75 cited materials it answered not_in_spec 31 times against a corpus that holds 8, and of the 42 materials it classed wrongly, 30 were answered not_in_spec where the key says prohibitive (21) or outthrow (9). It is not reaching for the wrong line and it is not inventing a prohibition - it declines to match the grader's plain-English narrative to the trade name the specification prints, and answers the one value that means the contract is silent. Under the older of the two revisions it reads 8 packets of 32 whole against 16 of 32 under the newer one.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Deciding which cited materials a purchase contract's grade specification actually prohibits — free code first - the desk phrase list in evals/floors.py
It reads 44 of 64 packets whole and 58 of 75 material classes for $0.00, against the paid call's 24 and 33, and the loss is significant at p = 0.000535053 and p = 4.12576e-05. A word list that has read your own specification is a real thing a desk can build. - Deciding which printed line establishes a class — the paid call, with the class grader beside it
30 of 75 against the floor of record's 21 and a fixed answer's 8 - the widest room of any cell here at 54 materials, and the cell where a keyword reader is most confidently wrong because of the 13 planted attractor pairs. - Deciding which shipped load the buyer actually graded — free code, and then the paid call only for the sentences a grep misses
A printed-prefix grep already reaches 48 of 64 on the line and a fixed this_load reaches 55 of 64 on the value. The paid call takes the remaining 16 lines, which is real and narrow, and it still reads the overlap wrongly on 3 packets. - Holding a hard cap - no regulator, no asserted threshold, no environmental determination, no settlement — this pack's structural contract
0 breaches over 30 attacked calls under three framings, including ten wordings demanding a regulatory claim, a statutory limit, a dollar credit and asettlementfield outright. The cap is structural - the answer has room for a schedule reference, a class, a line and an identity and nothing else - rather than a prompt sentence, and src/recheck.py drops anything outside it. RCS is the one family in this vertical with no regulator in its rule set, and nothing in this kit is asserted as regulator-governed.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/packets at your own deduction packets and rewrite src/record.py::parse to return the same six things: the header (ship date, load, grade, buyer), both specification blocks with their revision marks and their printed sNN lines, the numbered grading findings, the deduction schedule citing them by dNN, the correspondence, and the printed commercial terms (unit price, moisture tolerance, shipped weight). Nothing measured here transfers to your packets. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for this reading on material of this shape without measuring it first. The paid call is ahead of the floor of record and the lead does not clear 0.05. That is the case against the best-fitting scenario (“Deciding which cited materials a purchase contract's grade specification actually prohibits”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A grading report whose findings are not numbered. The whole answer is keyed on the printed dNN schedule reference and the sNN specification line; without them there is nothing for a class to attach to and no way to cite the line that establishes it. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the headline is a real difference. 24 packets against 14 is discordant 16/6 and p = 0.0524788 on 64 packets. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is OpenAI-compatible endpoint; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-21 — r001-bale-deduction - 64 deduction packets, 75 cited materials, 64 billed model calls, the fast tier, reasoning off. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on this build: a clean checkout with no key configured rebuilds all 66 corpus files byte-identically under two PYTHONHASHSEEDs, re-derives every label with 0 disagreements, scores all five free arms and the whole scoring path at $0.00, recomputes all 45 paired claims, and renders the whole board on port 9621 with the one live control disabled and the reason printed beside it. The reply cache ships, and that was decided by measurement rather than convention: the run record stores a 400-character head and this kit's widest reply is 208 characters, so the head is not lossy here - but the cap claim is only checkable against the whole reply, and --resume --rescore, which otherwise re-derives every published figure for $0.00, would find nothing cached and buy 64 calls.





















