Home › Use Cases › Reconcile one transfer station's tonnage balance for one period, ton by ton
Use caseUC0352
🧪 Use-case kit · runnable

Reconcile one transfer station's tonnage balance for one period, ton by ton

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A transfer station weighs everything twice and counts the floor at both ends of the period, and the four numbers are supposed to close. When they do not, the site system prints an imbalance and a status word — and on 34 of these 62 periods that figure is wrong, because a load that was weighed in and turned back at the tipping floor still reads POSTED, one truck was weighed twice, or an outbound load was raised here and cancelled before the trailer left. The only record of any of it is a sentence printed under the row, in ordinary English. Working that by hand means reading every note on every ticket, converting any column that is a container count rather than a tonnage, deciding whether a moisture deduction was actually agreed or merely quoted, and then doing the arithmetic — for every station, every period. Opening one station's period file, reading the note under every scale ticket, checking each inbound memo's container arithmetic, deciding whether the moisture note in the station notes was an agreement or a quote, and then striking the balance against both floor counts by hand.

Audience

A waste operator's transfer desk working a period balance list, and the operations analyst behind it. Whoever has to say how many tons are unaccounted for and on which side, before anybody upstream turns that figure into anything else. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual balance period

The corpus is 62 balance period, 0.20 MB (txt 62). It is generated because it has to be. A real scale-ticket ledger is an operator's own trading record and its notes name haulers and drivers, and the exact shapes this kit measures — a load turned back at the tipping floor that still reads POSTED, a moisture deduction that was quoted and refused, a NET TONS column that is a container count — are the rare rows in a real period and the ones nobody would let out. So every byte is invented from one seed, and the SHAPES are planted at a stated frequency: a disagreement on 30 of 62 periods, the site system's own panel wrong on 34, 6 moisture deductions in force and 11 that were never agreed.

The corpus

  • The 62 balance periodgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every operator, station, permit reference, destination facility, hauler run, manifest and note is invented; there is no name of a person, no address, no telephone number and no national identifier anywhere, and evals/check_labels.py sweeps all 62 files for five families of identifier on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your balance period. That is the whole change — there is no database to migrate.

One balance period, as the model receives itTBR-0001.txt · 1 of 62
============================================================================================
TRANSFER STATION TONNAGE BALANCE FILE                                  TBR-0001
Operator: OPR-3100 - Ardsley Common Waste Services (invented)
Station: TS-1104   Ardsley Common Transfer Station   Period: 2026-07-06 to 2026-07-12   Procedure: TB-2026
============================================================================================

STATION AND PERIOD AS THE MASTER HOLDS IT
  tonnage unit                        ton
  permit reference               PRM-8000
  opening floor                    22.583   ton
  closing floor                    21.000   ton
  weighed throughput              270.750   ton weighed inbound
  tolerance pct                      0.25   pct of weighed throughput
  threshold tons                    5.000   or more
  threshold pct                      1.50   pct or more
  period                        full-week

DESTINATION FACILITIES AS THE MASTER HOLDS THEM
  FACILITY   NAME                                   DISPOSITION
  FAC-2201   Kestrel Ridge Landfill                 disposal
  FAC-2244   Marchbank Valley Landfill              disposal
  FAC-3310   Rushmere Materials Recovery            recovery

SCALE TICKET LEDGER AS POSTED BY THE SITE
  TICKET     DATE             NET TONS  KIND       STATUS   REF        MEMO
  TKT-0100   2026-07-06         70.000  inbound    POSTED   HR-4403    commingled recyclables, 10 ctr at 7.000 ton per ctr, hauler run HR-4403
  TKT-0101   2026-07-07         49.000  outbound   POSTED   FAC-2244   residue, 2 trl, manifest MF-9901
  TKT-0102   2026-07-08         61.750  inbound    POSTED   HR-4401    C&D, 13 ctr at 4.750 ton per ctr, hauler run HR-4401

Abridged — the file continues.

The outcomeWhat a good result looks like

One period in, one row out: which tickets this procedure treats differently from the ledger that printed them and the row quoted verbatim for each, whether a moisture deduction is in force, the outbound tonnage the admitted loads account for, the imbalance against both floor counts, which side it is on, and one verdict — with the residue-to-disposal against recovered-tonnage split reported beside it.

And when it cannot

And what it does when it cannot. On the scored run 62 of 62 replies parsed and none stopped at the ceiling, so there is no unparsed-reply behaviour to show from this run — a reply that returned nothing would be counted WRONG and stay in the denominator, never dropped. The failure that did happen is quieter and is named: on 10 periods the arm returned the site system's own figure where a posted ticket had not moved material, and the pure-code station reproduces that faithfully because the station register carries no ticket at all.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your ledger's bad rows announce themselves — every turned-back load says "rejected", every duplicate says "duplicate" — the free rules floor, and do not buy a call at all
    A keyword list reaches everything a keyword can reach for $0.00. On this corpus 10 of the removal wordings carry no keyword at all, which is the only reason the paid arm is ahead on that family (10 of 16 against 5).
  • Your moisture and rebate notes are prose, and some of them are quotes rather than agreements — buy the call
    This is the one place the money is unambiguous. On the 11 periods carrying a deduction that was costed and NOT agreed the paid arm is right 6 times and the regex 0 — because the refused wordings carry a real date, a real facility code and a real rate, so every shape test and every date test passes and only the sentence says no.
  • Your columns are sometimes in the wrong unit — a container count, a pounds printout, a pallet count — write the arithmetic, do not buy it
    That is not a reading. The memo states the count and the per-unit weight and they either multiply out to the column or they do not. Free code gets 3 of 5 here and the paid arm 1.
  • You need to know WHICH rows were wrong, not just whether the total came out right — buy the call
    The offsetting family is the whole argument: a turned-back inbound and a wrong-station outbound of the same size leave the imbalance correct and two rows wrong. No tonnage and no verdict can say so. The paid arm names both rows on 3 of 4; the free floor on 0.

And where nothing here is good enough:

  • You want a number you can put in front of a regulator — neither arm, and this kit says so
    Nothing here establishes what a permit or a host-fee report requires — no threshold, no retention period, no deadline — and the kit refuses to state one on any arm. It stops at a reconciled, explained balance and at what it could not reconcile.

At a glanceHow the whole thing runs

45%all five correct pct
1,717 msp50, end to end
$0.00per 1,000 balance period · google/gemini-3-flash

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own balance files in the same shape and data/stations.json with your own site master, then run python3 -m evals.run --run-id b000-<yours>-rules --floor rules — no key, no spend — and read what free code already gets. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per period for a keyword list you could write in an afternoon. That is the case against the best-fitting scenario (“Your ledger's bad rows announce themselves — every turned-back load says "rejected", every duplicate says "duplicate"”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A real ledger of individual truck tickets rather than hauler-run rows. These files carry 678 ticket rows across 62 periods; a metro station's real week is hundreds per period and the prompt stops being 3459 bytes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE MARGIN OVER FREE CODE IS REAL. p = 0.5716 on 28 discordant pairs. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-tonnage-balance. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors, all six committed runs and every screenshot. python3 tools/build_corpus.py --check rebuilds every byte and diffs; python3 -m evals.check_labels re-derives the key with a second implementation. Neither needs a network. requirements.txt is deliberately empty of packages.

A living map of modern AI — kept current every morning