The business caseThe problem this solves
A buyer gets a month's invoice from a service contractor: sixty of them at once, five to nine lines each, every line a single run on a single date. Most of the check is arithmetic the buyer can already do — was the run in our own dispatch log, has it been billed twice, is the base the rate card's rate, is the accessorial one this contract permits, does the figure match what the log's own minutes come to. Two of the questions are not arithmetic. The service-type column is the CONTRACTOR'S CODING, not a fact about the run, and the rate card prices what was actually performed: a short repeated link between two points is a shuttle loop wherever the line says <code>Charter run</code>. And the service log records minutes stationary and never says why, while Clause 6.3 makes a wait the contractor itself caused unchargeable — so the only place that cause appears is the line's own narrative. Joining every invoice line to a service log by run reference and date, to a rate card by service type, and to an accessorial schedule by both — then reading each line's narrative to decide whether the coding is right and who caused any delay, and totalling the difference. One invoice at a time, for a month of them.
Audience
The reconciler in a buyer's finance or contract-management function who works the month's contractor invoices, and the contract manager who has to take whatever comes back to the contractor. The decision they are making is which lines to query — not whether to pay. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual contractor service invoices
The corpus is 60 contractor service invoices, 0.21 MB (txt 60). It is generated because a real one cannot be published. A contractor's invoice names a real buyer, a real supplier, real rates that are usually commercially confidential, and a dispute over it is a live commercial matter. And a corpus of real invoices would carry no answer key: the whole point here is a verdict per line that something can be right or wrong against, so the generator holds two facts the invoice never prints — what each run ACTUALLY was, and whether a recorded wait was the contractor's own — and derives the key from them by running the same engine every arm is graded against. No verdict in gold.jsonl was typed.
The corpus
- The 60 contractor service invoicesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 60 invoices, data/contracts.json and the whole answer key are generated in-process by the file that sits beside them. Every buyer, contractor, service area, contract number, run reference, rate, date, narrative and signature is invented.
Swap this folder for your own material and the kit is pointed at your contractor service invoices. That is the whole change — there is no database to migrate.
CONTRACTOR SERVICE INVOICE - RUN BILLING
INVOICE HEADER
Invoice CI-0001
Contract SVC-3100-01
Buyer Ardsley Unified District
Contractor Stonebridge Vehicle Hire
Service area Hill and valley sector
Billing period 2026-01-01 to 2026-01-31
Submitted 2026-02-04
CONTRACT TERMS
Rate card version RC-2026-A, in force for the whole of this billing period
Contracted base rate by service type
Charter run $369.00 per run
Extended route $264.00 per run
Standard route $214.00 per run
Shuttle loop $137.00 per run
Accessorials permitted, by service type
Charter run Waiting time; Fuel adjustment
Extended route Waiting time; Extra stop; Fuel adjustment
Standard route Waiting time; Extra stop
Shuttle loop none
How a permitted accessorial is computed
Waiting time $0.90 per minute beyond 10 free minutes, on the minutes the service log records
Extra stop $21.00 per stop beyond the scheduled count, on the stops the service log records
Fuel adjustment 4 pct of the contracted base rate for the service performed
SERVICE LOG (runs the buyer's dispatch system recorded in this period)
Run Date Wait min Stops over
R-4001 2026-01-02 1 0
R-4004 2026-01-05 0 0
R-4007 2026-01-08 34 0
R-4010 2026-01-11 0 0
R-4013 2026-01-14 0 0
INVOICE LINES
# Run Date Service billed Base Accessorial Units Amount Line total NarrativeAbridged — the file continues.
The outcomeWhat a good result looks like
One invoice in, and per line: one verdict from a closed set of six, the contract term it rests on, the amount at issue to the cent, and one row of the invoice copied verbatim as evidence. Then for the invoice as a whole: PASS or DISPUTE, the list of lines that do not tie, and the total at issue.
And when it cannot
And what it does when it cannot. On this corpus it got 398 of 400 line verdicts and 58 of 60 invoices completely right, and both failures point the SAME WAY: it named money that is not at issue. CI-0015 line 6 ties and it answered RATE-NOT-CONTRACTED, $85.00; CI-0050 line 1 is ACCESSORIAL-NOT-PERMITTED at $7.70 and it answered RATE-NOT-CONTRACTED at $73.70. Across the whole corpus it OVERSTATED the amount at issue by $151.00 and understated it by $0.00. That is the direction that writes a dispute letter to a contractor who billed correctly, and it is the opposite of how the free floors fail — the column floor understates by $1,229.89 and never overstates by a cent.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your contractors code the service type correctly and your disputes are about arithmetic, duplicates and runs that never happened. — the free column floor alone — evals/baseline.py
Every one of those is a lookup, an equality or one multiplication. The floor is 384 of 400 lines on this corpus and it never overstates a cent. - Your invoice narratives come from a small template pool — a dispatch system writes them, not a person. — the keyword floor, and measure it the way evals/baseline.py does
A word list over a template pool is memorisation and memorisation is free. Run memorisation_check() on your own corpus first; if your narratives repeat, the paid call is buying nothing. - Your contractors write the narrative by hand and the disputes turn on what the run actually was, or on who caused a delay. — the paid call, and keep the recheck station
That is the whole measured difference here: the defanged keyword floor gets 0 of the 7 contractor-caused-wait rows and the call gets them. The station then re-derives the term, the amount and the totals so only the READING is bought. - A wrong DISPUTE costs you more than a missed one — a contractor relationship, or a dispute process with a fee. — the free column floor, or the paid call with a human on every DISPUTE
The failure directions are opposite and measured: the column floor understates by $1,229.89 and overstates by nothing; the paid call overstated by $151.00 on this run. Two of its 60 invoices named money that was not at issue.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-05. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own invoices in the same columnar shape, data/contracts.json with your own rate cards, permitted schedules and dispatch log, and data/policy.md + data/policy.json with your own rulebook. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a reading you do not need. That is the case against the best-fitting scenario (“Your contractors code the service type correctly and your disputes are about arithmetic, duplicates and runs that never happened.”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An invoice whose line rows are not # run date service base accessorial units amount total narrative separated by runs of spaces. src/rules.py matches one regex and a row it cannot match is silently skipped — an invoice in a different layout parses to zero rows and every arm answers nothing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | One scored run. Every percentage here is a single measurement of a stochastic system, and a repeat at the same settings was not bought. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 6 models on the fast tier, reasoning left to the provider (the published run) and the free column floor — the whole lookup pass and the free keyword floor as fitted to this corpus and the free keyword floor, memorising patterns mechanically removed and the modal floor — TIES on every line, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-05 — r001-route-invoice-recon. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9312, scores all four free floors over every invoice offline, rebuilds the corpus and the key from the seed, runs the independent label gate, and replays the committed scored run at $0.00. Nothing to install — Python standard library end to end.




