The business caseThe problem this solves
A freight bill audit programme produces a ledger of adjudicated findings every month, and finance needs one number out of it: what was recovered this period. The number is easy to produce and hard to defend. Some findings settled for less than they claimed; some credits released an accrual raised months ago; some rate corrections prevented an overcharge rather than recovering one; a carrier occasionally debits back a recovery already published. Which of those count is the programme's own recognition policy, printed in the pack -- and it changes. When the figure moves, the first question is always whether the freight moved or the rules did. The spreadsheet pass a freight cost accountant makes over an audit recovery pack before the close: reading the finding ledger line by line, applying the recognition rules to each row, chasing the notes log for the rows whose treatment is not in a column, and writing the narrative that goes to the controller with the number.
Audience
The freight cost accountant who assembles the monthly recovery figure and the controller who signs it. The decision is not 'is this finding valid' -- an auditor already settled that. It is 'is this total the one our policy produces, can I show the rows behind it, and would I get the same number if I ran it again next quarter'. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual freight audit recovery pack (one SYNTHETIC monthly close for one audit programme: period header, recognition methodology, printed finding ledger, prior-period carry, programme notes, and a withheld finance position)
The corpus is 24 freight audit recovery pack (one SYNTHETIC monthly close for one audit programme: period header, recognition methodology, printed finding ledger, prior-period carry, programme notes, and a withheld finance position), 0.11 MB (txt 24). Because the answer is COMPUTED rather than opined, which makes every axis on this page a fact rather than a judgement -- and because it lets the reproducibility question be asked sharply. A summary of prose can differ between two passes and both versions be fine. A recovered total cannot: it is one number, and two passes either produce it or they do not. That is why this row was picked over the other summarisation candidates.
The corpus
- The 24 freight audit recovery pack (one SYNTHETIC monthly close for one audit programme: period header, recognition methodology, printed finding ledger, prior-period carry, programme notes, and a withheld finance position)generated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your freight audit recovery pack (one SYNTHETIC monthly close for one audit programme: period header, recognition methodology, printed finding ledger, prior-period carry, programme notes, and a withheld finance position). That is the whole change — there is no database to migrate.
FREIGHT AUDIT RECOVERY PACK -- FAR-0001
=======================================
Recovery Period Header
----------------------
Programme : Northbound Regional Freight Audit
Shipper : Calderon Foods Distribution
Period : 2026-01 (2026-01-01 to 2026-01-31)
Close date : 2026-02-06
Settlement currency : USD
Prepared for : Freight Cost Accounting
Recognition Methodology
-----------------------
M-1 In-period settlement only
A finding counts toward this period's recovered total only when its credit, short-pay
or refund actually settled between the period start and the period end printed above.
A credit that settles an accrual raised in an earlier period belongs to that earlier
period and is excluded here. A settlement confirmed inside the period counts even
where the ledger row had not been updated at print time and its settled column is
still blank; it counts at the amount the confirmation states.
M-2 Settled amount governs
Where a finding settled for less than the amount claimed, the recovered figure is the
SETTLED amount. The claimed amount is never recognised.
M-3 Cost avoidance excluded
A rate correction applied before the invoice was paid prevented an overcharge; it did
not recover one. It is excluded from recovered savings and reported separately as
cost avoidance.
M-4 Reversals deducted
Where a carrier debits back a recovery published in an earlier period, the amount is
deducted in the period the debit posts. It appears in the ledger as a negative
settled amount.
M-5 Open and disputed findings excluded
A finding still open, in dispute or awaiting carrier response at close has recovered
nothing and is excluded.Abridged — the file continues.
The outcomeWhat a good result looks like
One recovered total, the list of findings it is built from with the amount each contributed, the category breakdown, every excluded finding with the rule id that excluded it, the cost-avoidance figure reported separately, and an explicit amended/unchanged verdict on the methodology.
And when it cannot
The total does not equal the sum of the rows the same reply names. Measured at 0.0 pct traceability -- it happened on every answered pack.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your finding ledger is a clean table and your recognition rules map onto its columns — the free floor (evals/baseline.py, columns-only)
It is measured at 100.0 pct on the 120 column-decidable findings here -- every one of them -- cites the right rule on each exclusion, reconciles by construction and returns the same answer every time, for $0.00 and under a second. - Some findings' treatment lives in a free-text notes log rather than in a column — a keyword pass first, then measure the residue before assuming a model closes it
Three keywords got the floor to 24.71 pct on the 85 prose-decidable findings. The paid arm got 34.12 pct. That gap is real and it is 9.41 points -- worth knowing before it is worth buying. - You want the model's reading of the prose AND you need the total to be right — the split architecture this run points at, which this kit does not implement: floor for the column classes, model for the prose classes, and code for the arithmetic
The three failures are independent. The arm is fine at citing a rule (100 pct over 36 overlapping exclusions), mediocre at the columns, ahead on prose, and cannot sum its own itemisation on a single pack. Only the last one is fatal, and it is the one no model needs to be doing. - You need the figure to be defensible next quarter, recomputed — any deterministic arm
Reproducibility here is a property of the code, not of the corpus. The floors return byte-identical answers by construction; the paid arm returned the same total on 0 of 24 packs.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-27. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own recovery packs into data/corpus/ in the layout data/SOURCES.md describes -- underlined section headings, a finding ledger whose rows end with status, settled date and channel, and your own Recognition Methodology section -- and nothing in src/ changes. ⚠︎ WHAT DOES NOT TRANSFER. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying a model to do addition. On the same findings the paid arm scored 83.33 pct. That is the case against the best-fitting scenario (“Your finding ledger is a clean table and your recognition rules map onto its columns”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A close pack larger than the context window. There is no chunking and adding it would be the defect, not the fix -- see the index note. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A second paid tier. One tier was run through the model seam; that is a spend decision, not a comparison, and nothing here says a deliberating tier would do better or worse. 11 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-27 — r001-audit-savings. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, then python3 -m src.app -- no install, no key, no network. The packs, the answer key and every committed run record are in the repo, and the UI's second control replays the paid arm's actual answer to any pack with the reconciliation check recomputed live.



