The business caseThe problem this solves
A municipality's service period closes with a few hundred scale tickets on it — a row per ticket with its gross and tare weight, its date and the facility it went to — plus the facility records that certify what each site is allowed to be credited for, the hauler submissions somebody filed against them, and a notes block where three desks have written to each other all period. Nobody disputes the tickets. What nobody has is the REPORT: which stream each ticket belongs to, which tickets are disqualified and on which of six grounds, what tonnage that makes per stream, which streams have enough backing to be reported at all, and — the part that takes the afternoon — WHICH FACILITY RECORD carries each ticket, once the notes that void a record have been read against the dates they carry. Today that is an analyst with DIV-METHOD-2026 open in one window and the pack in the other. Reading one service period's pack end to end: testing every facility record against the disqualification ladder in order — voided, weight not reconciled, facility outside the declared set, no facility record, record not certified, record outside the allocation window — then reading every note for a void and joining that note's own date to the record's, then placing each ticket on the first record that survives, and only then doing the arithmetic.
Audience
A waste and recycling operator's sustainability desk assembling one municipality's service-period diversion tonnage report, and the manager deciding whether a model call is worth buying for it. This report's answer for this corpus is YES, and every number is laid out so that answer can be checked rather than taken — including the slice where the money bought nothing at all. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual diversion packs
The corpus is 64 diversion packs, 0.20 MB (txt 64). 64 packs, 860 scale tickets, 201 facility records and 195 hauler submissions, all generated at seed 20260917. The mix is the whole experiment: 32 packs carry an INERT void (a note voiding a record that is not the one the ticket sits on), 61 carry a confirmation DECOY quoting the deciding clause verbatim, 27 carry a note dated after the period end, 12 are cap probes asking in a desk's own voice for something the method never does, and 61 are decided by a note at all — which leaves 3 that the printed columns alone decide, and those 3 are the packs where the free floor and the paid call are exactly equal.
The corpus
- The 64 diversion packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from and — section 4 — what the generator was built to make the model fail on, written BEFORE any call was bought. ⚠︎ THAT PREDICTION CAME OUT BACKWARDS AND IT IS PUBLISHED UNTIDIED. It predicted the arm would win the reading and lose the arithmetic. Half of it was a category error — the answer contract has no tonnage field, so there is no arithmetic slice to lose. The other half is INVERTED by the measurement: the note reading WON (voided_an_inert_record 0 across the 32-pack trap, decoy_read_as_void 0 over 61 packs, acted_on_a_note_dated_after_period_end 0 over 27), and the deterministic COLUMN RULEBOOK is what lost (method_change 0/4, provisional_record 0/4, record_tie 0/4, weight_mismatch 0/3, window_edge 0/4). Free code losing its own arithmetic edge cases is a result nobody predicted, in either direction.
Swap this folder for your own material and the kit is pointed at your diversion packs. That is the whole change — there is no database to migrate.
DIVERSION PACK DVP-0001
SERVICE AREA: SVA-10 PERIOD: 2026-02 PERIOD START: 2026-02-01 PERIOD END: 2026-02-28 ASSEMBLED: 2026-03-08
DECLARED METHOD: DIV-METHOD-2026 - recovery facility classes FAC-CDR, FAC-MRF, FAC-ORG, FAC-REU; declared disposal class FAC-LFL; allocation window 14 days after the weigh date, inclusive; minimum backing 3 tickets; the residue rate in force at period start governs the whole period.
BLOCK 1 - SCALE TICKETS
TICKET WEIGHED HAULER ROUTE FACILITY CLASS MATERIAL GROSS TARE NET STATUS
SCT-40001 2026-02-22 HAU-06 RTE-106 FCY-10 FAC-CDR MAT-CDB 9.87 6.02 3.85 FINAL
SCT-40002 2026-02-27 HAU-05 RTE-104 FCY-10 FAC-CDR MAT-FOD 15.00 6.95 8.05 FINAL
SCT-40003 2026-02-24 HAU-10 RTE-168 FCY-10 FAC-CDR MAT-CDB 27.24 4.00 23.24 FINAL
SCT-40004 2026-02-26 HAU-05 RTE-129 FCY-10 FAC-CDR MAT-CDB 24.98 7.90 17.08 FINAL
SCT-40005 2026-02-15 HAU-07 RTE-165 FCY-19 FAC-MRF MAT-CDB 9.74 5.02 4.72 FINAL
SCT-40006 2026-02-07 HAU-12 RTE-114 FCY-19 FAC-MRF MAT-MXD 21.15 5.61 15.54 FINAL
SCT-40007 2026-02-16 HAU-01 RTE-136 FCY-19 FAC-MRF MAT-BLK 11.38 4.19 7.19 FINAL
SCT-40008 2026-02-18 HAU-09 RTE-127 FCY-19 FAC-MRF MAT-FOD 12.42 5.39 7.03 FINAL
SCT-40009 2026-02-05 HAU-01 RTE-139 FCY-33 FAC-REU MAT-FOD 21.95 7.52 14.43 FINAL
SCT-40010 2026-02-04 HAU-07 RTE-106 FCY-33 FAC-REU MAT-BLK 13.92 4.40 9.52 FINAL
SCT-40011 2026-02-17 HAU-03 RTE-128 FCY-33 FAC-REU MAT-OCC 13.39 7.61 5.78 FINAL
SCT-40012 2026-02-28 HAU-09 RTE-158 FCY-33 FAC-REU MAT-YRD 22.89 4.51 18.38 FINALAbridged — the file continues.
The outcomeWhat a good result looks like
One diversion pack in, one assembled report out: every scale ticket placed in a diversion stream or carrying exactly one disqualification reason with the facility record that decided it, then — in code, for every arm alike — each stream's ticket count, the rate applied, the diverted and disposed tonnage, the REPORTED/HELD state at the declared minimum backing of 3 tickets, the disposal split, the excluded block and the five kinds of open gap.
And when it cannot
And what it does when it cannot. On the scored run all 64 replies parsed, none stopped at the ceiling, and EVERY shape defect is zero — no ticket missing from the answer, no ticket invented, no ticket carrying both a stream and a reason, no stream or reason outside the closed vocabulary, no excluded or disposal ticket citing a record, no multi-citation and no phantom record id. What it produces instead is a CONFIDENT WRONG READING: 30 of the 64 packs have something wrong in them, at a median self-reported confidence of 0.9 against 0.9 on the 34 it gets completely right. There is no signal in that number and a reader must not route on it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your packs record a void, a withdrawal or a decertification as a STRUCTURED COLUMN on the facility record rather than as a sentence in a notes block. — the free columns floor, and do not buy a call at all
on the 3 packs no note decides, the free floor and the paid call are BOTH 40 of 40 tickets — identical, p = 1.00. Every pack you move from the notes block to a column is a pack that stops needing a call. - Your notes block is where the voids actually live, and a note can take a facility record out of play with a date join nobody has coded. — the call
on the 61 packs a note decides, it is 748 of 820 tickets against the floor's 693, exact McNemar p = 1.70051e-06. That is the slice the money is bought for and it is the only slice where it moves. - Your hard cases are date windows, record ties and provisional records. — free Python, and keep the call away from them
free code BEATS the call on all three: provisional_record 43 vs 37, record_tie 43 vs 38, window_edge 64 vs 62. Those are the deterministic parts of the method and the call is worse at them than a rulebook. - You want the report arithmetic, not the reading. — src/method.py::rollup(), free, with any reading you like
given the key's reading the station reproduces the key's report on 64 of 64. None of the stream counts, the rates, the tonnage, the minimum-backing test or the gap detection needs a model. - You want to know whether the call can be talked into acting. — evals/injection.py, 16 live calls, about $0.005898
four wordings on four packs planted in the pack's own notes block — certify the figures, an 'operator instruction' amending the method mid-period, roll the held streams in and sign and submit, and a fabricated state mandate. 0 adjudicated breaches and 0 regressions.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own service-period packs in the same shape — header, DECLARED METHOD line, scale ticket table, facility records, hauler submissions, notes — and rebuild the key with python3 -m tools.build_corpus. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per pack for arithmetic you already have. That is the case against the best-fitting scenario (“Your packs record a void, a withdrawal or a decertification as a STRUCTURED COLUMN on the facility record rather than as a sentence in a notes block.”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A pack with no DECLARED METHOD line. The whole reading is 'inside the declared method', and with nothing declared there is no facility set, no allocation window, no minimum backing and no weight tolerance — the report cannot be computed at all, by any arm. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | ON 40 OF THE 860 TICKETS THE MONEY BOUGHT NOTHING: free code and the paid call are BOTH 40 of 40 there, p = 1.00. Those are the 3 packs no note decides. 11 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-diversion-tonnage. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, runs all three free floors live, re-scores every committed arm from its result file and serves all 8 screenshots. What it cannot do is buy a call: the 'Ask the model' control disables itself and says why. Measured from this repository with API_KEY unset, on the no-key board at PORT+100 (9590); nothing else is required beyond Python 3 — requirements.txt pins no package.







