Home › Use Cases › Say why a carrier's freight bill disagrees with the audit pre-rate, and quote the line
Use caseUC0484
🧪 Use-case kit · runnable

Say why a carrier's freight bill disagrees with the audit pre-rate, and quote the line

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A carrier's freight bill does not match what the shipper's audit rating engine pre-rated for the same shipment. The variance is a number, not a reason: the carrier may have billed a weight nobody certified, an amendment the rating table has not loaded, an amendment the contract record does not hold, a rate that is not in force, the wrong week's fuel percentage or an accessorial outside the schedule — or the bill is right and the pre-rate is stale. The contract record that decides it is a mix of tabulated versions and letters, and the letters propose, withdraw, correct and scope each other in whatever words their authors used. So an auditor opens all eight panels and reads. the auditor's read of eight panels per mismatched bill to decide why the carrier's bill and the pre-rate differ — it does not correct a rate, load a table, amend a contract, dispute or pay a bill, or rate a carrier.

Audience

A freight audit lead deciding whether to put a model in front of a queue of rating mismatches. This report's answer is NOT ON ACCURACY: on the cause the paid call ties a free walk with a written contract vocabulary, loses the recurrence count and the complete row to it, and wins only the bills whose deciding letter is paraphrased. What is worth taking is the framework and that one reading. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual mismatched freight bills

The corpus is 64 mismatched freight bills, 0.21 MB (json 5 · jsonl 1 · md 1 · txt 64). A real carrier's bills, contract letters and rate tables carry customer names, negotiated rates and lanes, so they cannot be shipped and a redacted set cannot be labelled. This one is BUILT as structures and the key is RMT-2026 applied to those same structures, which is the only way to have 64 labelled bills whose key is derivable rather than opinion. It was rebuilt once (v2) so that READING decides: the contract record carries four sentence operators — a proposal not countersigned, a withdrawal, a correction, a scope — each worded in a standard form and in paraphrases, and 43 of 64 bills change cause when a sentence operator is switched off. 16 bills are decided by a paraphrased letter, 15 by a standard-worded one, 16 by the tables alone; 21 have no operator threshold; 5 have no contract record attached.

The corpus

  • The 64 mismatched freight billsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your mismatched freight bills. That is the whole change — there is no database to migrate.

One mismatched freight bill, as the model receives itRMM-0001.txt · 1 of 64
FREIGHT BILL RATING MISMATCH  RMM-0001
====================================================================================================
PANEL 1 — SHIPMENT AND BILL OF LADING
  Shipper                  Fernhollow Housewares (synthetic shipper)
  Carrier                  CAR-35 (regional LTL carrier)
  Lane                     LN-558  DC-West to delivery zone 6
  Bill of lading           BOL-238737
  Ship date                2026-05-26
  Delivered                2026-05-30
  Weight tendered          9,230 lb
  Accessorials requested   none

PANEL 2 — THE FREIGHT BILL AS THE CARRIER BILLED IT
  Freight bill             FB-887516
  Bill date                2026-06-09
  Lane billed              LN-558
  Billed weight            9,230 lb
  Linehaul                 92.30 cwt at 22.15 per cwt               USD 2,044.45
  Fuel surcharge           28.5% of linehaul                        USD 582.67
  Billed total                                                      USD 2,627.12

PANEL 3 — AUDIT PRE-RATE (the shipper's rating engine)
  Rated on                 2026-06-10
  Rating table loaded      contract versions BASE, AMD-1, AMD-2
  Lane as rated            LN-558
  Weight as rated          9,230 lb (bill of lading)
  Linehaul as rated        92.30 cwt at 22.60 per cwt               USD 2,085.98
  Fuel as rated            28.5% of linehaul                        USD 594.50
  Accessorial as rated     none requested on the bill of lading     USD 0.00
  Pre-rated total                                                   USD 2,680.48
  Variance                 billed total minus pre-rated total       USD -53.36

PANEL 4 — CONTRACT RATE RECORD FOR THIS CARRIER AND LANE
  Record                   RR-558  CAR-35 on LN-558
  Version    Signed       Effective    Linehaul per cwt

Abridged — the file continues.

The outcomeWhat a good result looks like

One row per mismatched bill an auditor could work as written: the cause under RMT-2026, the line of the file that establishes it, the recurrence flag against the operator's threshold, and — derived in pure code, never asked — the queue it goes to.

And when it cannot

It abstains. A bill whose contract rate record never came through cannot be walked, and the answer is NEEDS-REVIEW whatever the carrier's remark or the analyst's note says — 5 of 64 bills, and every arm gets all 5 because the rule is structural in src/recheck.py. Where the operator supplied no recurrence threshold for a carrier it answers the literal threshold not supplied and never borrows one — 21 of 21 with 0 breaches, a number a constant emitting that literal would also reach, so it is reported and never carded alone.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • your contract letters are written in the contract's standard words, and your export prints the rate tables — the free domain floor — evals/baseline.py, $0.00
    it takes 14 of the 15 standard-wording bills and 15 of the 16 the tables decide; the paid call takes 8 and 10. A call there buys nothing and loses bills a vocabulary list already had.
  • your carriers word their proposals, withdrawals and corrections in their own words — the paid call on those bills only — and measure it against a tables-only walk as well as your vocabulary floor
    this is the one place the call reads what the floor of record cannot: 10 of the 16 paraphrase-decided bills against the floor's 2 (p = 0.0078). It is not separable from the tables-only walk or the constant through the station, which each take 8 there.
  • the recurrence flag matters, because a pattern triggers a rate-table review — free code, unconditionally
    counting PANEL 8 against the operator's threshold is 64 of 64 in src/policy.py and 47 on the paid call, which under-counted 16 bills. The literal threshold not supplied is 21 of 21 with 0 breaches on both.
  • you want what this kit is actually worth on your corpus — the FRAMEWORK — the graded refusal behaviour, the pure-code station applied to every arm alike, the derived queue and the free recurrence count, the adversarial arm and the measured cost by tariff
    the accuracy result here is a tie with one narrow slice, and it only exists because a free floor was written and measured first. On letters worded further from the floor's vocabulary the same machinery would report a larger slice, and it would be believable for the same reason.

At a glanceHow the whole thing runs

62%rechecked cause accuracy pct
1,553 msp50, end to end
$0.38per 1,000 mismatched freight bills · the fast tier, outside the peak window

Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/policy.json with your own card — the reading rules, the walk, the evidence line per cause and the recurrence thresholds your operator actually supplies — and src/taxonomy.py's causes and their ORDER with yours. The measured result does not travel. Corpus lens →
When is this the wrong choice?Avoid: Paying per bill for a date comparison and a word list. That is the case against the best-fitting scenario (“your contract letters are written in the contract's standard words, and your export prints the rate tables”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a contract letter worded in a way nobody anticipated. The free floor of record reads the usual contract words for a proposal, a withdrawal, a correction and a scope and takes 2 of 16 paraphrase-decided bills; the paid call takes 10. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether a paid call is worth making at all on this corpus — on the headline it is not. It ties the free floor of record on the cause, 40 against 42 (p = 0.8601), and loses the recurrence flag (47 v 64, p < 0.0001) and the complete row (25 v 42, p = 0.0076). 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, off-peak tariff. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-16 — r001-rating-mismatch. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — The kits repository is private, so this is a record of what was run rather than an offer: on a copy of this folder with no key configured and no .env above it, all four free floors (42, 33, 28 and 22 of 64 rechecked), evals/check_labels.py (0 disagreements), and a re-score of the committed paid run from its own cache — every score and all 52 paired tests byte-equal to the committed file — reproduced at 0 calls, $0.00 and about one second of wall clock, and the board on port 9484 served every route. Re-running the paid arm needs a provider key.

A living map of modern AI — kept current every morning