Home › Use Cases › Recompute a waste hauler's fuel fee and environmental fee against their own eligible bases
Use caseUC0299
🧪 Use-case kit · runnable

Recompute a waste hauler's fuel fee and environmental fee against their own eligible bases

A small, forkable project that does one job end to end. Run twice for real over the same set, and every figure on these pages captured from those runs.

The business caseThe problem this solves

A waste hauler's monthly service invoice carries two percentage fees on top of the service and disposal charges, and they are not applied to the same money. The hauler publishes a rate per service month; the customer's contract may cap either fee lower. Somebody in accounts payable is supposed to check that each fee was applied to the right lines at the right rate, on every invoice, for every site — and what they actually do is compare the percentage on the fee line against the schedule and move on, because reconstructing the base means deciding what every charge line IS. Re-adding the eligible base of two percentage fees by hand, invoice by invoice, and deciding line by line which charges each fee is even allowed to sit on. It does not replace the call to the hauler, the credit, or the person who decides whether a divergence is worth arguing.

Audience

A billing analyst working a month of hauler invoices, and the billing supervisor who reads what they produce. The decision this report is for is narrow: is a percentage fee on this invoice worth a call to the hauler, and what exactly would you say. It is not for anyone choosing a hauler, negotiating a rate or approving a credit. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual waste hauler service invoices

The corpus is 62 waste hauler service invoices, 0.14 MB (txt 62). It is generated because it has to be. A real hauler service invoice is a customer's own accounts-payable document naming a real hauler, a real site and real money, and there is no public corpus of them. Generating it also buys the one thing a fetched corpus cannot: the answer key is DERIVED from the same structure the invoice is rendered from, so every base, every fee and every verdict is what the generator constructed rather than what somebody later decided it should be — and evals/check_labels.py re-derives all of it from the PRINTED invoice with its own rule table to prove the two agree.

The corpus

  • The 62 waste hauler service invoicesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 62 invoices, the structured charge-line feed, the flat rule list and the whole answer key are generated in-process from seed 20260903 by tools/build_corpus.py. --check rebuilds them into a temp directory and diffs byte for byte, and it was run under three different PYTHONHASHSEEDs and agreed all three times.

Swap this folder for your own material and the kit is pointed at your waste hauler service invoices. That is the whole change — there is no database to migrate.

One waste hauler service invoice, as the model receives itWSI-0001.txt · 1 of 62
WASTE SERVICE INVOICE
==============================================================================
  Hauler:        Northreach Waste Group                     Invoice:        WSI-0001
  Customer:      Belfount Distribution Centre               Service month:  2026-05
  Site:          Bay 11 - waste pad                         Agreement:      WSA-5914

CHARGE LINES
  id   code description                                                            amount
  --------------------------------------------------------------------------------------
  L01  SVC  Front-load 6 yd container C-2410, 2 lifts/week                       2,607.39
  L02  SVC  Front-load 8 yd container C-1928, 3 lifts/week                         798.89
  L03  SVC  Rear-load 20 yd container C-1783, 2 lifts/week                       3,096.49
  L04  DSP  Tonnage 18.60 T at 47.87/T - transfer station disposal                 553.89
  L05  ONE  Container relocation 2026-05-18 - one-time                             459.47
  L06  TAX  State and local sales tax at 8.75% on services and disposal            657.66
  L07  FEE  Fuel recovery fee at 6.36% of the eligible base                        413.58
  L08  FEE  Environmental fee at 5.38% of the eligible base                        379.65
  --------------------------------------------------------------------------------------
  invoice total                                                                  8,967.02

PUBLISHED FEE SCHEDULE - Northreach Waste Group, effective 2026-05
  Fuel recovery fee                6.36 % of the eligible base
  Environmental fee                5.38 % of the eligible base
  Eligible base, fuel recovery fee: recurring container service charges.
  Eligible base, environmental fee: recurring container service charges and

Abridged — the file continues.

The outcomeWhat a good result looks like

One invoice in, two fee rows out. Each row carries the eligible base as recomputed, the recomputed fee, the difference against what was billed, and one verdict from a closed set of six — with the invoice line that establishes it, quoted verbatim.

And when it cannot

And what it does when it cannot. On the adversarial arm one invoice's reply ran past the 32,000-token ceiling and came back unparseable; both of its diverging fee rows are counted WRONG and stay inside the published denominator, and the run is not re-fired. That is 1 call of the 198 this kit bought and it is the only product failure this kit has ever measured. It is on the board as a shot.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the fee recomputed and the difference valued, on invoices whose charge lines are honestly coded — the free rules floor alone — evals/baseline.py
    It is right on every fee row it reads correctly and it costs $0.00 with no network. On the 42 invoices in this corpus with no trap line it is a TIE with the paid call at 84 of 84 fee rows, to the cent.
  • Your charge-line descriptions carry more than the code column does — a delivery coded as a service, a tonnage line coded as a service, a recurring line coded as a one-time — the paid call
    That is the whole measured margin: 124 of 124 against the floor's 106, discordant 18 to 0, exact two-sided McNemar p = 7.6e-06. All 18 of the floor's errors are on the four trap forms and all 18 are FALSE DIVERGENCES.
  • You care most about never flagging a correct invoice — the paid call, and watch false_divergence
    0 of 81 on both scored runs against the floor's 18 of 81 (22.2 pct). That is the error direction that gets a recheck switched off, and it is counted apart from the other direction and never averaged with it.
  • You want a pure-code station in front of the model to catch its mistakes — src/recheck.py, but read what it measured first
    It changed 0 fields on 124 fee rows across both scored runs, because the model never disagreed with it. What it DID measure is its own limit: on the null floor, which supplies no reading, it takes 81 correct MATCHES down to 16 and manufactures 81 false divergences.

At a glanceHow the whole thing runs

100%verdict correct pct · 2 runs, no ordering
68,350 msp50, end to end
$34.42per 1,000 waste hauler service invoices · Gemini 3 Flash

Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own invoices and data/invoices.json with the matching structured feed — one row per invoice carrying lines (id, code, description, cents, and fee_of on the two fee lines), schedule (the published rate per fee, omitted where none is published), contract (a cap per fee, omitted where there is none) and billed (what each fee was charged and at what stated rate). ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Buying a model to do arithmetic. Nothing on this board suggests it is worth it. That is the case against the best-fitting scenario (“You want the fee recomputed and the difference valued, on invoices whose charge lines are honestly coded”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A scanned or photographed invoice. Every arm here reads a fixed-layout text charge table, and the amounts arrive structured in data/invoices.json rather than being read out of a column. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHERE THE MODEL BREAKS. It did not break, on 124 fee rows, twice, and a third time under a different prompt. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, r001, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-surcharge-recheck. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured rebuilds the whole corpus and answer key from seed in 0.07 s, re-derives the key independently in 0.04 s at 0 failures over 4,057 checks, scores both free arms in 0.11 s, renders the entire board on 127.0.0.1:9299 and replays both committed paid runs offline. What it cannot do without a key is fire a new call: the one control that would is disabled and says so on the page.

A living map of modern AI — kept current every morning