The business caseThe problem this solves
A waste hauler's revenue from EXTRAS — an unscheduled lift, a container overfilled at the kerb, a contaminated load, an overweight lift — is captured at the kerb by a driver on a run sheet and invoiced weeks later by somebody who was not there. Two failures run in opposite directions at the same time and neither is visible in a column. Extras that were performed and are chargeable never reach the invoice at all, because nothing connected the driver's sentence to a billing code; and extras that were NOT chargeable get invoiced anyway, because dispatch typed the row EXTRA-LIFT and the billing system priced what it was typed. The second is worse than the first: it invoices a customer for a lift the hauler already owed it. On this corpus the hauler's own extras reconciliation is wrong on 44 of 63 periods. Opening one period's file, ticking every unscheduled lift on the run sheet against the service calendar, reading each driver note to decide whether the site asked for the lift or the hauler went back on its own account, chasing each scale ticket to the lift it belongs to, recomputing the tonnage above the contracted allowance, and then reading the invoice's extra lines back the other way to see which of them anything on the file supports.
Audience
The billing analyst working a period's extras on a commercial waste book, and the billing supervisor behind them for disputed or aged items. Whoever reads the output is deciding which candidates get drafted as charges, which invoice lines get queried before the customer does, and which periods close. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual service-account billing period file
The corpus is 63 service-account billing period file, 0.17 MB (txt 63). It is generated because it has to be. A real extras capture runs on a hauler's own route sheets, a customer's signed service agreement and an invoice register for a real account, and the exact shapes this kit is about — a lift the hauler went back for on its own account, a contamination the site cleared before the tip, an overage a driver described on a row typed as routine service — are precisely the ones that are commercially sensitive and legally load-bearing. Generating them is what makes the answer key DERIVABLE rather than typed: each period is built as a structure, the blocks are rendered from it, and XCP-2026 is applied to the same structure by src/agreement.py. It also makes the difficulty placeable, which is the honest cost of the method and is stated as such in data/SOURCES.md: the decisive evidence was deliberately put in a sentence in the notes for four families, because the first build put it in a column and free code decided everything.
The corpus
- The 63 service-account billing period filegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every account is an invented code and trading name, every event id, scale ticket, invoice line and rate is arithmetic on the file index, and there are no people in the corpus at all — which is the first place this kit's own refusal would have leaked.
Swap this folder for your own material and the kit is pointed at your service-account billing period file. That is the whole change — there is no database to migrate.
==============================================================================
EXTRA SERVICE CHARGE CAPTURE -- ONE SERVICE ACCOUNT, ONE BILLING PERIOD
==============================================================================
FILE XTR-0001
ACCOUNT 7401 Northgate Crossing Retail
SITE rear dock, 2 x 8 yd front load
AGREEMENT SA-7401-01, extras schedule A
PERIOD 2026-06-01 to 2026-06-30
CUT-OFF 2026-07-02
BILLING LAG extras invoiced within 5 business days of the event
TOLERANCE 2.50
CURRENCY USD
-- SERVICE AGREEMENT, EXTRAS SCHEDULE (what may be charged, and for what) ----
EXTRA LIFT 95.00 per lift, only where the customer requested it
CONTAINER OVERAGE 45.00 per event, lid unable to close or material outside the container
CONTAMINATION 125.00 per event, load rejected or a contaminated fraction recorded
OVERWEIGHT 45.00 per ton above the contracted allowance of 3.000 tons per lift
-- SCHEDULED SERVICE CALENDAR ------------------------------------------------
Mon, Thu front load, 2 lifts per week, 8 scheduled lifts in the period
-- ROUTE AND DRIVER EVENT LOG ------------------------------------------------
EVENT DATE TYPE CONTAINER STATUS DRIVER NOTE
EVT-3001 2026-06-03 SERVICE CTR-7401A LOGGED routine lift, nothing to report
EVT-3002 2026-06-09 SERVICE CTR-7401B LOGGED presented on time, lid closed
EVT-3003 2026-06-15 SERVICE CTR-7401A LOGGED clean lift, container level
EVT-3004 2026-06-22 SERVICE CTR-7401B LOGGED worked to the run sheet
EVT-3005 2026-06-10 EXTRA-LIFT CTR-7401A LOGGED second lift for the week, ordered by the siteAbridged — the file continues.
The outcomeWhat a good result looks like
One service account and one billing period in, one row out: which events on the route and driver log are chargeable extras under the agreement, which invoice extra lines a chargeable event supports with the invoice row quoted verbatim, what the agreement's own schedule allows recomputed in code, the period's variance, and one of five XCP-2026 verdicts. Anything it could not evidence is flagged by name. It never posts a charge, issues an invoice, raises a credit or waives a fee, and it never names a person.
And when it cannot
And what it does when it cannot. On the scored run 63 of 63 replies parsed, nothing stopped at the ceiling and no call failed. Where it is wrong it is wrong about a READING, and four of its six misses are the same reading: it returned an EMPTY supported-line set on three periods whose invoice line names no event id, and it supported a line on two periods where nothing on the file evidences it. The station cannot notice either — it re-derives faithfully from whatever reading it is handed and returns a verdict with full confidence.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your dispatch codes are reliable and your drivers write nothing — the free rules floor, and do not buy a call at all
50 of 63 periods for $0.00, and 37 of 37 on everything decidable from columns and dates: the overweight recompute, the phantom ticket, the duplicated line, the billing-lag window and the periods with no extras at all. - Your drivers DO write, and your depot codes an unscheduled lift the same way whoever the lift was for — the paid call
On the chargeable-event set it is 61 of 63 against free code's 52 — eleven periods to two, exact McNemar p = 0.022. Seven of those eleven are one family: a lift a driver's note says the hauler worked on its own account, which free code gets 0 of 7 and the hauler's own system gets wrong every time. - You mostly want to catch lines billed with nothing behind them — the free rules floor
Free code is AHEAD on the supported-line set, 58 to 57, and ahead outright on the quoted row, 63 to 59. A floor that copies the invoice row out of the file cannot mis-quote it, and a floor that supports every line it cannot disprove never returns an empty set.
And where nothing here is good enough:
- Your overage evidence is a photograph — neither, yet
The candidate row says it in terms: photo evidence lives outside any system of record. This kit reads a driver's SENTENCE; an image stage would sit in front of it, and nothing here has been measured against one. - You want the pack to draft the charge — neither
It never will. The answer contract offers no field that could express a charge, a credit or a waiver, and src/prompt.py fails the BUILD if one is added. The pack's contribution stops at a candidate with its evidence attached.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own period files in the same shape and data/accounts.json with your own rate schedules, allowances, tolerances and billing lags. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per period for arithmetic you already have. That is the case against the best-fitting scenario (“Your dispatch codes are reliable and your drivers write nothing”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A period with no driver notes at all. Every reading this kit is paid for lives in a sentence; a dispatch system that records only codes reduces this kit to its own free floor, which is 50 of 63 here and costs nothing. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-overage-capture. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run. python3 tools/build_corpus.py --check rebuilds all 63 files byte-identically (verified under PYTHONHASHSEED 0, 12345 and 999), python3 -m evals.check_labels re-derives the key from the printed files alone and reports 0 disagreements and 0 person-shaped names, and python3 -m evals.run --run-id b000-overage-capture-rules --floor rules scores all 63 periods for $0.00. There is nothing to pip install: the kit is standard library only.







