Home › Use Cases › Check one outbound booking against the customer routing guide that travelled with it
Use caseUC0285
🧪 Use-case kit · runnable

Check one outbound booking against the customer routing guide that travelled with it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A customer publishes a ROUTING GUIDE — which carrier gets the tender on which lane, which mode at which weight, how much notice a delivery appointment needs, when orders must be consolidated, what has to be labelled and what has to travel with the freight. A vendor's transport desk tenders against it every day, and a booking that departs from it is a chargeback the customer raises weeks later against a shipment nobody can change any more. Today that is a planner reading a PDF guide beside a booking screen. The expensive version of getting it wrong is not the carrier that appears nowhere in the guide — that one is easy. It is the booking tendered to the guide's SECONDARY carrier with no decline from the primary behind it: on every mechanical test it looks exactly like a conforming tender. A transport planner reading a routing guide beside a booking screen, one tender at a time, and the chargeback that arrives weeks later when they miss one.

Audience

A transport planner working a tender queue, and the vendor-compliance analyst who reads what they produce. The decision this report is for is narrower than it looks: not 'should we buy a model', but WHICH OF THE EIGHT CHECKS a model should touch — because a free regex arm scores seven of them perfectly. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual outbound bookings

The corpus is 62 outbound bookings, 0.11 MB (txt 62). A real routing guide and the bookings tendered against it are commercially confidential on both sides of the contract — the guide is the customer's own instruction to its vendors and the booking is the vendor's own tender file — so there was never a real corpus to reduce. What is generated instead is built to make ONE thing hard and to make the other seven easy on purpose, because that split is the finding: seven of the eight conflict classes are decidable from fixed-layout panels and a free regex gets 25 of 25 of them, and the eighth is a sentence in a notes panel typed by people.

The corpus

  • The 62 outbound bookingsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 62 bookings, the guide register and the whole answer key are generated in-process from seed 20260903 by tools/build_corpus.py, and --check proves a rebuild is byte-identical to what is committed under two different PYTHONHASHSEEDs. Nothing is fetched, scraped or licensed from anywhere, so there is no source URL to map and no third-party dedication to verify.

Swap this folder for your own material and the kit is pointed at your outbound bookings. That is the whole change — there is no database to migrate.

One outbound booking, as the model receives itBKG-0001.txt · 1 of 62
Customer routing compliance review -- Calloway Hardware Co
Booking BKG-0001        prepared 2026-09-03

BOOKING FACTS
  Customer                Calloway Hardware Co
  Booking reference       BKG-0001
  Tender reference        TND-41001
  Tendered on             2026-08-05
  Ship date               2026-08-10
  Lane tendered           NDC31-STR4203
  Mode tendered           LTL
  Carrier tendered        Northvale Freightways
  Total weight lb         5,216
  Guide version printed   RG-2026.2

ROUTING GUIDE EXTRACT   (RG-2026.2, as printed on the booking)
  LANE NDC31-STR4203 PRIMARY Northvale Freightways SECONDARY Fallowbrook Freight
  LANE NDC31-STR4620 PRIMARY Ambercrest Logistics SECONDARY Wexham Road Services
  MODE parcel up to 149 lb, LTL 150 to 9,999 lb, truckload 10,000 lb and above
  CONSOLIDATION orders to one destination on one ship date above 5 pallets move as one shipment
  APPOINTMENT minimum 24 hours notice before the delivery window opens
  LABELS Pallet placard, Destination store barcode
  DOCUMENTS Carrier bill of lading, Purchase-order reference
  SEQUENCE the tender goes to the primary carrier; the secondary is reached only after the primary has declined this tender

SHIPMENT LINES
  ORDER        DESTINATION  SHIP DATE   PALLETS  WEIGHT LB  PIECES  SHIPMENT
  ORD-74198    STR4203      2026-08-10  1        4,661      9       SHP-0001-1
  ORD-77597    STR4491      2026-08-10  1        555        12      SHP-0001-1

APPOINTMENT AND DOCUMENTS
  Appointment requested at   2026-08-10T04:01
  Delivery window opens      2026-08-12T11:00
  Pallet placard             evidenced
  Destination store barcode  evidenced
  Carrier bill of lading     evidenced
  Purchase-order reference   evidenced

CASE NOTES

Abridged — the file continues.

The outcomeWhat a good result looks like

One booking in, five graded answers out: the appointment lead time in whole hours, the conflict class, the conformance position, the action, and the line that establishes the conflict quoted verbatim and locatable at character offsets. On this corpus the paid call returns 60 of 62 conflict classes right against the free floor's 51, names the secondary-carrier trap on 12 of 12 against the floor's 5, and — after RGC-2026 is re-applied in code — 60 of 62 actions.

And when it cannot

And what it does when it cannot. ⚠︎ TWO REPLIES WERE CUT OFF AT THE 32,000-TOKEN CEILING and returned nothing at all — BKG-0001 and BKG-0053. Both are counted WRONG on all five fields and stay in every denominator on this page; neither was re-fired. ⚠︎ AND THE CITATION IS A LOSS TO A FREE REGEX, plainly: 41 of 62 bookings scored citation credit against the free floor's 51, and where the key names a line the paid call earned 16 of 37 against the floor's 30. Not one of its quotes was fabricated — 0 of 37 returned quotes were unlocatable — it quoted the OTHER line of the same argument, the guide clause where the key names the booking line and the booking line where the key names the clause. That is exactly the disagreement evals/check_labels.py said in its own output it could not settle, and it is recorded here rather than argued away.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the appointment lead time and nothing else — the free rules floor alone — evals/baseline.py::lead_of
    62 of 62 exact for $0.00 over two printed ISO timestamps, against the paid call's 60 (it answered 60 of 60 and got every one of those right). A real tender export prints both timestamps too.
  • You want the carrier, mode, consolidation, timing and evidence checks — the free rules floor, again
    seven of the eight conflict classes are decidable from fixed-layout panels and it misses 0 of 25 of them. There is nothing here for a model to win.
  • You want to catch the tender that went to the secondary with no decline behind it — the paid call, and read the conflict class on its own
    12 of 12 against the floor's 5. Every one of the floor's misses is a mechanism, not bad luck: 4 carry a decline of a DIFFERENT tender, 3 a decline dated before the tender existed, and 4 carry a genuine decline in words the phrase list does not hold.
  • You want the action to be right — the paid call WITH src/recheck.py bolted to it, never without
    60 of 62 rechecked against 53 raw and the floor's 51 — paired, p = 0.0225. The recheck fired 14 times, and every override is a booking whose guide register the prompt never carried.
  • You want the line quoted so a planner can check the flag — the free rules floor
    ⚠︎ THE PAID CALL LOSES THIS ONE: 41 of 62 credited against 51. Not one of its quotes was fabricated (0 of 37 unlocatable) — it quoted the other half of the same argument, the guide clause where the key names the booking line.

At a glanceHow the whole thing runs

32%booking all correct
70,680 msp50, end to end
$33.86per 1,000 outbound bookings · google/gemini-3-flash

Run once, for real, on 2026-09-03. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Three files. Your bookings and your guide, on your machine. Corpus lens →
When is this the wrong choice?Avoid: Paying for a datetime subtraction. That is the case against the best-fitting scenario (“You want the appointment lead time and nothing else”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A booking whose panels are shaped differently. Every parser here — the free floor, the independent key check and the board — binds to literal headings, so a renamed panel returns an empty section and a conflict nobody flags. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A SECOND RUN of the same corpus. Every figure here is one run, and two of its calls were cut off at the output ceiling — a second run would very likely truncate a different pair. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-routing-conformance. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone, run the rules floor, open the board. No install, no key, no network: requirements.txt names nothing and every free path is standard library. The floor answers and scores all 310 graded cells in wall_seconds 0.0.

A living map of modern AI — kept current every morning