Home › Use Cases › Match a filed loss-and-damage claim to its shipment, receipt and invoice
Use caseUC0423
🧪 Use-case kit · runnable

Match a filed loss-and-damage claim to its shipment, receipt and invoice

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A loss-and-damage claim arrives with the references the claimant copied onto the form — a bill of lading, a purchase order, a ship date, a consignee — and the claims desk has to find which shipment it is about, whether an exception was actually taken at delivery, whether that exception still stands, what those pieces are worth on the shipment's own invoice, and whether the claim came in before the CLAIM BY date. The claims system matches on intake and prints a status, and that status is computed from columns. It disagrees with the rulebook on 28 of the 64 claims in this corpus and matched the WRONG shipment on 18, because what decides the claim is often a sentence: the claimant's office writing that the references were copied in error, a report cancelled after a recount, the same damage keyed twice, a note naming which of two deliveries the cartons came on. Opening one claim file, finding the shipment the claim describes among every shipment the account has in the period, reading every delivery receipt and OS&D exception under it, reading each note to see whether it is about this claim and whether it changes anything, pricing the excepted pieces off the invoice, and comparing the RECEIVED stamp with the CLAIM BY date.

Audience

The claims desk at a carrier or a 3PL working a period's loss-and-damage claims — and the adjuster the claim is routed to, who alone decides whether anything is paid. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual claim files

The corpus is 64 claim files, 0.16 MB (txt 64). It is generated because it has to be. A real claims book is a carrier's own commercial record: shippers, consignees and goods under contract, invoice values that are commercially confidential, and correspondence written by named people about named people — drivers, dock hands, clerks. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.

The corpus

  • The 64 claim filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every claimant is an invented account code and trading name, every shipment, bill of lading, purchase order, receipt, OS&D report and invoice is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note is from the claims desk, the terminal, customer service or the claimant's shipping office, and a delivery receipt is signed at a DOCK. evals/check_labels.py sweeps all 64 files for a person-shaped name, an honorific or a 'signed by' on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your claim files. That is the whole change — there is no database to migrate.

One claim file, as the model receives itCLM-0001.txt · 1 of 64
==============================================================================
FREIGHT LOSS AND DAMAGE CLAIM -- ONE CLAIM, THE SHIPMENTS IT COULD BE ABOUT
==============================================================================
FILE           CLM-0001
CLAIM NO       LD-26-0101
CLAIMANT       5521  Fernmoor Ceramics
RECEIVED       2026-05-13
CLAIM TYPE     DAMAGE
TOLERANCE      10.00
CURRENCY       USD
TERMINAL       terminal 12 cross-dock

-- THE CLAIM AS FILED --------------------------------------------------------
BOL GIVEN        --
PO GIVEN         55489
SHIP DATE GIVEN  2026-05-01
CONSIGNEE GIVEN  Fennick Row DC
PIECES CLAIMED   1 CTN
WEIGHT CLAIMED   42 LB
VALUE CLAIMED    185.00
NARRATIVE: One carton of ceramic tiles arrived crushed at one corner; the contents cannot be sold.

-- CANDIDATE SHIPMENTS (shipment system, this account, this period) ----------
PRO         BOL      PO     SHIPPED     DELIVERED   CONSIGNEE           PCS  UNIT  WEIGHT  CLAIM BY
PRO-441000  7789587  56766  2026-05-07  2026-05-09  Quenby Park DC        4  CTN    168 LB  2026-06-23
PRO-441001  7728156  55489  2026-05-01  2026-05-03  Fennick Row DC        6  CTN    252 LB  2026-06-17

-- DELIVERY RECEIPTS AND OS&D EXCEPTIONS -------------------------------------
RECEIPT   PRO         DELIVERED   PIECES    SIGNED AT
DR-8000   PRO-441000  2026-05-09   4 of 4    yard gate 2 dock
   no exception taken, signed clear
DR-8001   PRO-441001  2026-05-03   6 of 6    south receiving dock
   OSD-5000  DAMAGED   1 PCS  crushed at one corner, noted at unloading

-- INVOICES ------------------------------------------------------------------
INVOICE    PRO         PIECES  UNIT   UNIT VALUE       TOTAL
INV-30000  PRO-441000       4  CTN       185.00      740.00

Abridged — the file continues.

The outcomeWhat a good result looks like

One claim file in, one row out: which shipment the claim is about, which OS&D exceptions on it are in force with each report row quoted verbatim, the supported pieces, the supported value and the value gap to the cent, and one of six LDC-2026 verdicts. 62 of 64 claims come back completely right, against 42 for the best free arm and 32 for the claims system's own auto-match.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, nothing stopped at the ceiling and no call failed. The 2 claims it got wrong are named in the kit README with what it answered: on CLM-0048 the reply's sentence named both deliveries and its shipments field named neither, so the station read NOT-FOUND against a key of AMBIGUOUS; on CLM-0064 a note cancelled one of two reports and the reply dropped both. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your terminal records a cancelled or re-keyed OS&D report as a status, and your claim form is filled from the shipment record rather than typed by the claimant — the free rules floor, and do not buy a call at all
    42 of 64 claims for $0.00. The window, the clear receipt, the over-invoice claim, the split delivery, the sister shipment, the wrong-kind exception, the transposed bill of lading and the claim that matches nothing are all decidable from columns and dates, and the floor gets every one.
  • Corrections, cancellations and re-keys arrive as free text — a note from the claimant's office, a terminal memo, a line in the OS&D log — the paid call
    This is the whole product. On the 22 claims where a sentence decides the reading the paid arm is 21 and every free arm is 0.
  • You need the ambiguous claims — one PO, two deliveries, nothing saying which — found reliably, because that is the queue a named human works — the free rules floor, and read the paid arm beside it
    The rules floor got all 5 ambiguous claims and the paid arm 4 — on CLM-0048 it wrote the ambiguity into its sentence and left the shipments field empty. Two shipments on one quoted PO is a column fact, so free code is the right tool and the call is the risk.
  • You want the claims system's own auto-match audited — either paid or free — both beat it
    The auto-match is right on 32 of 64 claims and matched the wrong shipment on 18. It is published as an arm precisely so the comparison is against what is running today rather than against nothing.

And where nothing here is good enough:

  • Your claims span several shipments by design, or carry mixed invoice lines — neither, yet
    No file in this corpus does. The unit of work is one claim about one shipment and every percentage on this page is against that unit.

At a glanceHow the whole thing runs

50%all correct pct
1,585 msp50, end to end
$1.46per 1,000 claim files · the fast tier

Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own claim files in the same shape and data/claims.json with your own register, then run python3 -m evals.run --run-id b000-<yours>-rules --floor rules — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per claim for a comparison you already have. That is the case against the best-fitting scenario (“Your terminal records a cancelled or re-keyed OS&D report as a status, and your claim form is filled from the shipment record rather than typed by the claimant”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A shipment export with no CLAIM BY date per shipment. CR-3 compares the RECEIVED stamp with that printed date and this kit computes no claim window of its own; a forker must supply the dates their own contracts produce. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-12 — r001-freight-claim-match. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning