The business caseThe problem this solves
A hazardous-waste manifest leaves the generator with the shipment. The designated facility signs it on receipt and returns a signed copy, and until that copy is in hand the shipment is not evidenced as delivered. Today a shipping desk keys a status word off whatever arrives in the mail, and a compliance officer works the list that word produces. The word is wrong in exactly the two places that matter: an envelope that held the transporter's page rather than the facility's, and a block carrying a rubber stamp rather than a printed name and a date. Opening the envelope, deciding whether what is in it is the designated facility's own signed copy, checking the block for a name AND a date, comparing every line of the returned copy against the line that was signed out, reading the correspondence for an extension that was actually granted, and counting the window out on a calendar.
Audience
A generator's compliance officer working the manifest exception list, and the shipping desk that keys the return log. The decision is which manifests go on the exception report and which are simply not due yet. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual manifest tracking record
The corpus is 62 manifest tracking record, 0.12 MB (txt 62). It is generated because it has to be. A real manifest tracking record is a generator's own regulated file, it carries the names of the people who signed each block, and publishing one would publish a shipment. Generating it also buys the thing a scraped corpus cannot: the ANSWER IS DERIVED from the same structure the record is rendered from, so there is no second place the key lives and a corpus change cannot leave a stale key behind. The identifier formats are wrong on purpose — a corpus that could be mistaken for a real shipment is a corpus somebody eventually treats as one.
The corpus
- The 62 manifest tracking recordgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every generator, facility and transporter is an invented trading name; every EPA-style identifier, manifest tracking number, waste code and line id is in a format NO REAL REGISTRY USES (ZZG9000219, MTN-91002329-ZQ, QW122) and every record prints (invented) beside the company name. There is no personal data at all — evals/check_labels.py sweeps all 62 records for five families of identifier on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your manifest tracking record. That is the whole change — there is no database to migrate.
HAZARDOUS WASTE MANIFEST TRACKING RECORD MTR-0001
Generator: ZZG9000100 - Braxmoor Coatings LLC (invented)
Manifest: MTN-91000000-ZQ Shipped: 2026-03-02 Reporting date: 2026-03-29 Procedure: MRP-2026
MANIFEST AND RETURN WINDOW AS THE REGISTER HOLDS IT
generator ZZG9000100
designated facility ZZF9000400 Kestrel Ridge Treatment (invented)
transporter ZZT9000700 Redlake Haulage (invented)
shipment date 2026-03-02
reporting date 2026-03-29
return window 35 days
shipment state routine
MANIFEST LINES AS THE GENERATOR SIGNED THEM, WITH THE RETURNED COPY BESIDE
LINE CODE QUANTITY UNIT CONT CLASS STATUS COPY-CODE COPY-QTY
ML-0001 QW101 25.00 P 1 hazard ACTIVE QW101 25.00
ML-0002 QW156 66.25 G 3 hazard ACTIVE QW156 66.25
ML-0003 QW228 107.50 T 5 nonreg ACTIVE QW228 107.50
RETURN LOG AS THE MAILROOM KEYED IT
RL-01 2026-03-04 SENT carrier-copy
NOTE: our own page filed on dispatch.
RL-02 2026-03-29 RECEIVED inbound-mail
NOTE: opened and filed: the receiving facility's copy for this tracking number.
RL-03 2026-03-29 SCANNED records
NOTE: facility block completed in full, printed name and date of receipt both entered.
RETURN TRACKING AS THE SHIPPING DESK REPORTS IT
status COPY RECEIVED
signature block NAME AND DATE
copy received 2026-03-29
days open 27
CORRESPONDENCE NOTES
correspondence: quarterly manifest register reconciled against the shipping log.
correspondence: annual training record for the shipping desk refreshed this month.
The outcomeWhat a good result looks like
One manifest in, one row out: what actually came back, whether a written extension is in force, which manifest lines the returned copy contradicts, the due date, and one of five verdicts.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed and 0 stopped at the output ceiling; 61 of 62 records came back with all five graded fields right on the published column, and the one that did not is named on this page by its record id.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your returned copies are transcribed into columns beside the signed ones — free code
a quantity or a waste code keyed differently is a comparison of two printed columns, and both free arms and the call score 10 of 10 on that family - Your mailroom logs a status word and nobody opens the envelope — the call
this is the whole product.signature_statusis 62 of 62 for the call against the best free arm's 49, and the envelope-reading family is 11 of 11 against 3 - Extensions are recorded in a workflow system with a state field — free code
a state field is a lookup. R-4's whole difficulty is that a request, a grant and a withdrawal are three sentences that state the same manifest number, the same date and the same term - Your exception list is worked by one person and a false entry costs an hour — the call, with the station
the station's verdict is derived from the reading rather than taken from the reply, so a wrong verdict traces to a wrong reading and never to arithmetic. 61 of 62 verdicts on the published column
And where nothing here is good enough:
- Returned copies arrive as scans of handwriting — neither, yet
this kit is text end to end and has no capability stage. A transcription step in front of it changes what is measured and none of these numbers carries
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own tracking records in the same shape and data/manifests.json with one register row each. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call that adds nothing to a numeric compare. That is the case against the best-fitting scenario (“Your returned copies are transcribed into columns beside the signed ones”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A returned copy that arrives as a SCAN OF HANDWRITING rather than as a transcription. Every annotation in this corpus is a sentence in plain English; a real one is a carbon page somebody wrote on at a weighbridge. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the labelled line is the line a compliance officer would have quoted. The key names ONE per discrepancy — that line's own printed row — and an officer might accept the annotation printed under it carrying the same fact. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier + the station + the free corrections, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-manifest-return. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all 62 manifests, all three free floors and every committed run, and scores every one of them offline — the empty shot on this page was taken against a server started with API_KEY blanked in its environment. The corpus rebuilds byte-identically from SEED 20260911, the second implementation re-derives the answer key with 0 disagreements over 848 checks, and every grader re-scores at $0.00. The only thing a key buys is a NEW call.







