The business caseThe problem this solves
A carrier's freight bill does not match what the shipper's audit rating engine pre-rated for the same shipment. The variance is a number, not a reason: the carrier may have billed a weight nobody certified, an amendment the rating table has not loaded, an amendment the contract record does not hold, a rate that is not in force, the wrong week's fuel percentage or an accessorial outside the schedule — or the bill is right and the pre-rate is stale. The contract record that decides it is a mix of tabulated versions and letters, and the letters propose, withdraw, correct and scope each other in whatever words their authors used. So an auditor opens all eight panels and reads. the auditor's read of eight panels per mismatched bill to decide why the carrier's bill and the pre-rate differ — it does not correct a rate, load a table, amend a contract, dispute or pay a bill, or rate a carrier.
Audience
A freight audit lead deciding whether to put a model in front of a queue of rating mismatches. This report's answer is NOT ON ACCURACY: on the cause the paid call ties a free walk with a written contract vocabulary, loses the recurrence count and the complete row to it, and wins only the bills whose deciding letter is paraphrased. What is worth taking is the framework and that one reading. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual mismatched freight bills
The corpus is 64 mismatched freight bills, 0.21 MB (json 5 · jsonl 1 · md 1 · txt 64). A real carrier's bills, contract letters and rate tables carry customer names, negotiated rates and lanes, so they cannot be shipped and a redacted set cannot be labelled. This one is BUILT as structures and the key is RMT-2026 applied to those same structures, which is the only way to have 64 labelled bills whose key is derivable rather than opinion. It was rebuilt once (v2) so that READING decides: the contract record carries four sentence operators — a proposal not countersigned, a withdrawal, a correction, a scope — each worded in a standard form and in paraphrases, and 43 of 64 bills change cause when a sentence operator is switched off. 16 bills are decided by a paraphrased letter, 15 by a standard-worded one, 16 by the tables alone; 21 have no operator threshold; 5 have no contract record attached.
The corpus
- The 64 mismatched freight billsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your mismatched freight bills. That is the whole change — there is no database to migrate.
FREIGHT BILL RATING MISMATCH RMM-0001
====================================================================================================
PANEL 1 — SHIPMENT AND BILL OF LADING
Shipper Fernhollow Housewares (synthetic shipper)
Carrier CAR-35 (regional LTL carrier)
Lane LN-558 DC-West to delivery zone 6
Bill of lading BOL-238737
Ship date 2026-05-26
Delivered 2026-05-30
Weight tendered 9,230 lb
Accessorials requested none
PANEL 2 — THE FREIGHT BILL AS THE CARRIER BILLED IT
Freight bill FB-887516
Bill date 2026-06-09
Lane billed LN-558
Billed weight 9,230 lb
Linehaul 92.30 cwt at 22.15 per cwt USD 2,044.45
Fuel surcharge 28.5% of linehaul USD 582.67
Billed total USD 2,627.12
PANEL 3 — AUDIT PRE-RATE (the shipper's rating engine)
Rated on 2026-06-10
Rating table loaded contract versions BASE, AMD-1, AMD-2
Lane as rated LN-558
Weight as rated 9,230 lb (bill of lading)
Linehaul as rated 92.30 cwt at 22.60 per cwt USD 2,085.98
Fuel as rated 28.5% of linehaul USD 594.50
Accessorial as rated none requested on the bill of lading USD 0.00
Pre-rated total USD 2,680.48
Variance billed total minus pre-rated total USD -53.36
PANEL 4 — CONTRACT RATE RECORD FOR THIS CARRIER AND LANE
Record RR-558 CAR-35 on LN-558
Version Signed Effective Linehaul per cwtAbridged — the file continues.
The outcomeWhat a good result looks like
One row per mismatched bill an auditor could work as written: the cause under RMT-2026, the line of the file that establishes it, the recurrence flag against the operator's threshold, and — derived in pure code, never asked — the queue it goes to.
And when it cannot
It abstains. A bill whose contract rate record never came through cannot be walked, and the answer is NEEDS-REVIEW whatever the carrier's remark or the analyst's note says — 5 of 64 bills, and every arm gets all 5 because the rule is structural in src/recheck.py. Where the operator supplied no recurrence threshold for a carrier it answers the literal threshold not supplied and never borrows one — 21 of 21 with 0 breaches, a number a constant emitting that literal would also reach, so it is reported and never carded alone.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- your contract letters are written in the contract's standard words, and your export prints the rate tables — the free domain floor — evals/baseline.py, $0.00
it takes 14 of the 15 standard-wording bills and 15 of the 16 the tables decide; the paid call takes 8 and 10. A call there buys nothing and loses bills a vocabulary list already had. - your carriers word their proposals, withdrawals and corrections in their own words — the paid call on those bills only — and measure it against a tables-only walk as well as your vocabulary floor
this is the one place the call reads what the floor of record cannot: 10 of the 16 paraphrase-decided bills against the floor's 2 (p = 0.0078). It is not separable from the tables-only walk or the constant through the station, which each take 8 there. - the recurrence flag matters, because a pattern triggers a rate-table review — free code, unconditionally
counting PANEL 8 against the operator's threshold is 64 of 64 in src/policy.py and 47 on the paid call, which under-counted 16 bills. The literalthreshold not suppliedis 21 of 21 with 0 breaches on both. - you want what this kit is actually worth on your corpus — the FRAMEWORK — the graded refusal behaviour, the pure-code station applied to every arm alike, the derived queue and the free recurrence count, the adversarial arm and the measured cost by tariff
the accuracy result here is a tie with one narrow slice, and it only exists because a free floor was written and measured first. On letters worded further from the floor's vocabulary the same machinery would report a larger slice, and it would be believable for the same reason.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/policy.json with your own card — the reading rules, the walk, the evidence line per cause and the recurrence thresholds your operator actually supplies — and src/taxonomy.py's causes and their ORDER with yours. The measured result does not travel. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per bill for a date comparison and a word list. That is the case against the best-fitting scenario (“your contract letters are written in the contract's standard words, and your export prints the rate tables”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a contract letter worded in a way nobody anticipated. The free floor of record reads the usual contract words for a proposal, a withdrawal, a correction and a scope and takes 2 of 16 paraphrase-decided bills; the paid call takes 10. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether a paid call is worth making at all on this corpus — on the headline it is not. It ties the free floor of record on the cause, 40 against 42 (p = 0.8601), and loses the recurrence flag (47 v 64, p < 0.0001) and the complete row (25 v 42, p = 0.0076). 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, off-peak tariff. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-rating-mismatch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — The kits repository is private, so this is a record of what was run rather than an offer: on a copy of this folder with no key configured and no .env above it, all four free floors (42, 33, 28 and 22 of 64 rechecked), evals/check_labels.py (0 disagreements), and a re-score of the committed paid run from its own cache — every score and all 52 paired tests byte-equal to the committed file — reproduced at 0 calls, $0.00 and about one second of wall clock, and the board on port 9484 served every route. Re-running the paid arm needs a provider key.







