Home › Use Cases › Reconcile a billed container rental line against the site it is actually assigned to
Use caseUC0276
🧪 Use-case kit · runnable

Reconcile a billed container rental line against the site it is actually assigned to

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A container rental line keeps running until somebody stops it. The container is picked up, swapped for another unit, or carted by the customer to a different yard; the route stops visiting; the monthly rental keeps invoicing against a site the container is not on. Nobody notices because the billing system knows the ASSIGNMENT and the route system knows the ACTIVITY, and reconciling the two is an annual manual sweep of thousands of lines. The annual manual container reconciliation sweep — pulling the rental register and the route history for each line, finding the last time anybody saw THAT container at THAT site, and adding up what was billed after it. It does not replace the coordinator's correction of the register, the recovery-trip approval, or any adjustment to a bill; all three are outside the pack by construction.

Audience

The billing analyst working a container reconciliation sweep, and the assignment coordinator who reads what they produce. Neither of them is being replaced: the pack produces a line to read, and the correction of the record stays with the coordinator. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual container rental reconciliation packets

The corpus is 64 container rental reconciliation packets, 0.12 MB (txt 64). It is generated because it has to be. A real container reconciliation packet is a join of a hauler's billing register and its route history — commercial records of a named customer at a named address, with driver names on every line. There is no public corpus of these and there could not be one. What is generated is the SHAPE of the problem: a billing schedule, an assignment of record, an activity log that names containers other than the billed one, and a rulebook that has to be applied in its published order.

The corpus

  • The 64 container rental reconciliation packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 64 packets, the container register and the whole answer key are generated in-process from one seed. There is no upstream dataset, no scrape and no licence to inherit.

Swap this folder for your own material and the kit is pointed at your container rental reconciliation packets. That is the whole change — there is no database to migrate.

One container rental reconciliation packet, as the model receives itCRP-0001.txt · 1 of 64
CONTAINER RENTAL RECONCILIATION PACKET  CRP-0001
Location is INFERRED from route and driver records. No RFID, GPS or barcode scan exists for this container.
As of 2026-08-31.  Verification window 90 days: an event dated on or after 2026-06-02 verifies presence.

BILLED LINE
  Customer        Pemberton Foods
  Site            S-1122  Elm St yard
  Container       CN-29107  (6 yd front-load)
  Billed to site  S-1122

RENTAL BILLING  (monthly container rental only; haul and disposal charges are not shown)
  period                     amount
  2026-04-01..2026-04-30      109.50
  2026-05-01..2026-05-31      109.50
  2026-06-01..2026-06-30      109.50
  2026-07-01..2026-07-31      118.50   rate change
  2026-08-01..2026-08-31      118.50

ASSIGNMENT OF RECORD  (container register)
  CN-29107  ->  S-1533  since 2026-05-18  ACTIVE
  CN-29107  ->  S-1122  2025-11-21..2026-05-17  ENDED

SERVICE ACTIVITY  (events naming a container at this site, and this container anywhere; oldest first)
  2026-04-03  S-1122  CN-29107  service completed, blocked access cleared  route 39  drv K. Mbeki
  2026-04-14  S-1122  CN-29107  service completed, lid damage noted  route 39  drv K. Mbeki
  2026-04-23  S-1122  CN-29107  serviced, full  route 39  drv K. Mbeki
  2026-04-30  S-1122  CN-29107  service completed  route 39  drv K. Mbeki
  2026-05-13  S-1122  CN-29107  service completed, blocked access cleared  route 39  drv K. Mbeki
  2026-05-17  S-1122  CN-29107  removed to Central Yard; service cancelled  route 39  drv K. Mbeki
  2026-05-20  S-1533  CN-29107  delivered and set  route 9  drv L. Whitcombe
  2026-05-27  S-1533  CN-29107  service completed  route 9  drv L. Whitcombe
  2026-06-09  S-1533  CN-29107  service completed, blocked access cleared  route 9  drv L. Whitcombe

Abridged — the file continues.

The outcomeWhat a good result looks like

One billed rental line in, six graded answers out: what the latest event naming that container attests, what the assignment of record says, the rental billed past the last attestation to the cent, the CAR-2026 position, the action, and the activity line the reading came from — quoted verbatim so a person can check it in one glance.

And when it cannot

And what it does when it cannot. On the scored run the paid call read seven haul-backs as relocations — removed answered as elsewhere — at confidence 0.90 to 0.99. Both labels are departures, so CAR-2026 fired identically and the position, the action and the figure were right on all 64; what is wrong is the label, not the decision. It cleared no departed container (0 of 36), assumed absence from silence on none of the 13 stale packets, and flagged none of the 15 clean ones.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your activity log is a fixed-width export with a closed set of event codes — the free regex floor in evals/baseline.py
    It scored 64 of 64 packets on this corpus for $0.00 and beat the paid call. If your phrasings are enumerable, enumerate them.
  • Your activity log is free text — driver notes typed at the cab — the paid call
    The floor's whole result rests on 27 enumerable phrasings. That table cannot be written against prose, and this kit's corpus cannot tell you how the call does on prose either — it is the measurement nobody here has made.
  • Anybody can write into the case-notes field that reaches the pack — the paid call plus src/recheck.py, and read the adversarial arm first
    Under a fabricated site-verification sentence over $7,970.75 of rental the call cleared 0 of 20. The regex floor is immune only because it never reads the notes.

And where nothing here is good enough:

  • You have RFID, GPS tags or scan-at-swap on your containers — neither arm — read the tag
    This whole pack is an inference layer over service activity and driver reports. If the location is directly observable the inference is the wrong design.

At a glanceHow the whole thing runs

89%packet all correct pct
24,348 msp50, end to end
$15.52per 1,000 container rental packets · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packets, data/register.json with your own assignment rows (one per packet), and data/gold.jsonl with your own key. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call that, on a corpus shaped like this one, does not beat an expression table. That is the case against the best-fitting scenario (“Your activity log is a fixed-width export with a closed set of event codes”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A free-text activity log. Every expression in the free floor is bound to fixed-width columns and a closed phrase list; a real depot's notes field defeats all of it, and that is the measurement this corpus cannot make. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A SECOND MODEL. One tier was measured, on one run. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-container-assign. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9276 and scores every graded cell twice for $0.00 — --floor rules and --floor modal need no key, and --stub proves the whole pipe end to end without one.

A living map of modern AI — kept current every morning