The business caseThe problem this solves
A disposal facility bills a hauling account monthly, one charge line per load. Each line is supposed to reconcile to a scale ticket that weighed the load and to a haul record that says this account's truck delivered it — tonnage, material class, gate rate, service date. Working that by hand means matching every line to a ticket number that is sometimes blank and sometimes mis-keyed, deciding from the weighbridge's own typed remarks what was actually tipped rather than what was booked at the gate, applying the facility minimum where a load came in short, looking the rate up in the window containing the service date, and checking three surcharges — for every line, every facility, every month. On this corpus 165 of 281 lines do not reconcile. Matching every invoice line to a scale ticket by hand, reading the weighbridge's remarks to decide what was actually tipped, applying the minimum-load rule, looking the gate rate up by service date and checking each surcharge against the ticket.
Audience
A waste operator's accounts-payable reviewer working a disposal invoice, and the contracts analyst behind them. Whoever has to say which lines to query and what the net amount at issue is, before anybody sends a letter to a facility. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual disposal invoice
The corpus is 44 disposal invoice, 0.11 MB (txt 44). It is generated because it has to be. A real disposal invoice, its ticket file and a hauler's dispatch log are three parties' trading records and the remarks name drivers and sites — and the exact shapes this kit measures (a load reclassified at the face that the docket never recorded, a ticket column mis-keyed at billing, a ticket with no haul behind it) are the rare rows in a real month and the ones nobody would let out. So every byte is invented from one seed, and the SHAPES are planted at a stated frequency: 45 reclassifications, 46 decoy remarks, 37 mis-keyed columns, 19 decoy ticket numbers in a narrative, 29 short loads that hit the facility minimum and 23 contaminated tickets.
The corpus
- The 44 disposal invoicegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every facility, hauling account, truck, site, scale ticket, haul record, invoice line and gate rate is invented; there is no real person, no address, no telephone number and no national identifier anywhere. It also states the limitation that matters most: the remarks come from 32 templates, so a keyword floor scores better here than it would on real weighbridge prose.
Swap this folder for your own material and the kit is pointed at your disposal invoice. That is the whole change — there is no database to migrate.
================================================================================================
DISPOSAL INVOICE DI-0001 AGREEMENT DA-2026-W
FACILITY Riverbend Transfer Station (FAC-RIV)
ACCOUNT Ridgeway Hauling Co (ACCT-4417)
PERIOD 2026-01-01 to 2026-01-31 ISSUED 2026-01-31
================================================================================================
INVOICE LINES
LINE|TICKET|SERVICE_DATE|CLASS|TONS|RATE|GATE|ENVFEE|FUEL|CONTAM|TOTAL|DESCRIPTION
L01|880001|2026-01-15|FILL|21.431|20.00|428.62|66.44|18.00|0.00|513.06|CLEAN FILL DISPOSAL PER TICKET 880001
L02|880002|2026-01-24|SPECIAL|17.748|152.00|2697.70|55.02|113.30|0.00|2866.02|SPECIAL WASTE DISPOSAL PER TICKET 880002
L03|880003|2026-01-22|CD|22.864|78.00|1783.39|70.88|74.90|75.00|2004.17|CONSTRUCTION AND DEMOLITION DEBRIS DISPOSAL PER TICKET 880003
L04|880004|2026-01-16|SPECIAL|14.617|152.00|2221.78|45.31|93.31|0.00|2360.40|SPECIAL WASTE DISPOSAL PER TICKET 880004
L05|880005|2026-01-22|MSW|16.386|72.00|1179.79|50.80|49.55|0.00|1280.14|MUNICIPAL SOLID WASTE DISPOSAL PER TICKET 880005
L06|880006|2026-01-23|YARD|14.618|45.00|657.81|45.32|27.63|0.00|730.76|YARD WASTE DISPOSAL PER TICKET 880006
L07|880007|2026-01-24|CD|3.542|78.00|276.28|10.98|11.60|0.00|298.86|CONSTRUCTION AND DEMOLITION DEBRIS DISPOSAL PER TICKET 880007
SCALE TICKETS
TICKET|DATE|FACILITY|TRUCK|CLASS|GROSS_LB|TARE_LB|NET_TONS|FLAGS|REMARKS
880001|2026-01-15|FAC-RIV|RH-124|FILL|76062|33200|21.431|-|no issues at the gate
880002|2026-01-26|FAC-RIV|RH-230|SPECIAL|66966|31470|17.748|-|wet weather, load sheeted
880003|2026-01-22|FAC-RIV|RH-266|CD|79558|33830|22.864|CONTAM|no issues at the gate
880004|2026-01-18|FAC-RIV|RH-270|SPECIAL|64914|35680|14.617|-|wet weather, load sheetedAbridged — the file continues.
The outcomeWhat a good result looks like
One invoice in, one reviewed invoice out: every line matched to its ticket, one break code with the clause that proves it, the entitled amount and the amount at issue to the cent and signed, the queried set, the net total, and a PASS or QUERY recommendation — with the ticket row that decided each reading quoted verbatim.
And when it cannot
And what it does when it cannot. On the scored run 44 of 44 replies parsed and none stopped at the ceiling, so there is no unparsed-reply behaviour to show from this run — a reply that returned nothing would be counted WRONG and stay in the denominator, never dropped. The failure that DID happen is quieter and is named: on 6 lines the arm read the class the gate booked rather than the class the weighbridge recorded, and the pure-code engine priced that reading faithfully and published a CLASS_BREAK against a line that was correct.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your invoice lines carry a clean ticket number and your tickets carry no prose — the deterministic key join, and do not buy a call at all
A key join on ticket, date and tonnage is free, instant and exact. On THIS corpus it only reaches 201 of 281 lines because 37 columns are mis-keyed and 45 tickets were reclassified in prose; take those away and it is perfect. - Your ticket columns are sometimes mis-keyed and the real number is in the narrative — buy the call, or write the narrative regex
This is where the money is unambiguous against the join everybody runs: 280 to 244, p = < 0.000001. But the freeremarksfloor gets 279 of the same 281 lines with a nine-line regex, so the honest comparison is against THAT, not against the join. - Your weighbridge remarks are templated, short, and written by the same few people — write the word list, and do not buy a call
That is exactly this corpus, and the result is a tie: 267 lines whole each, 14 discordant each way, p = 1.000000. A lexicon of about 40 phrases in 30 lines of Python reaches the same answer for $0.00. - Your remarks are open-ended English typed by dozens of people across many sites — buy the call — but measure it first, on your own remarks
A lexicon degrades as the vocabulary widens and a reading does not, so the margin should grow. NOTHING ON THIS PAGE MEASURES THAT: this corpus's remarks come from 32 templates and the floor's score here is an upper bound. The honest answer is that the experiment is cheap — $0.038865 for 44 invoices — so run it rather than reasoning about it.
And where nothing here is good enough:
- You want a number you can put in a credit note or a payment run — neither arm, and this kit says so
Nothing here pays, approves, holds or releases anything, raises a credit or a rebill, or amends the agreement. It stops at a reviewed invoice: the lines to query, the clause behind each, and the net amount at issue. What happens next is a person's decision.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own billing extracts in the same three-table shape and data/contract.json with your own Schedule A, then run python3 -m evals.run --run-id b000-<yours>-remarks --floor remarks — no key, no spend — and read what free code already gets. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per invoice for a join you already have. That is the case against the best-fitting scenario (“Your invoice lines carry a clean ticket number and your tickets carry no prose”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A facility whose extract is shaped differently. src/docsheet.py is bound to three literal section headings and to three fixed column counts; rename a block or move a column and every figure downstream is wrong at once rather than gradually. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE PAID CALL IS WORTH ANYTHING OVER THE BEST FREE CODE. It is not, on this corpus: 267 lines whole each, 14 discordant each way, p = 1.000000. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-scale-ticket-match. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors, all six committed runs and every screenshot. python3 tools/build_corpus.py --check rebuilds every byte and diffs; python3 -m evals.check_labels re-derives the key with a second implementation over 2,380 checks. Neither needs a network. requirements.txt is deliberately empty of packages.






