Home › Use Cases › Reconcile one chemical shipment's four documents against its order and hazard table
Use caseUC0436
🧪 Use-case kit · runnable

Reconcile one chemical shipment's four documents against its order and hazard table

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A specialty-chemicals shipment travels with four documents written by four desks — the carrier's bill of lading, the dangerous-goods declaration, the quality desk's certificate of analysis and the trade desk's commercial invoice — and before it moves they have to agree with the sales order and the shipper's own hazard table on product, lot, shipping name, identification number, class, packing group, net quantity, consignee and tariff code. The desk's checklist reads the PRINTED columns. What decides the answer is sometimes a sentence — an order amended after the export, a lot re-allocated, a document corrected or cancelled, a trading name given up — and far more often a print style: the same packing group in roman or arabic, the same quantity in kilograms or pounds. The desk's own checklist status disagrees with the procedure on 33 of the 64 sets in this corpus. Opening one shipment's document set, reading every cell of four documents against the sales order as amended on or before AS AT and the hazard table row, deciding which differences are print style and which are real, reading each note to see whether it amends the order, re-allocates the lot, corrects or cancels a document or gives up a trading name, and naming what does not agree.

Audience

The trade-compliance and dispatch desks at a chemicals shipper, working the week's document sets before collection — and the engineer deciding whether a model call is worth putting in front of them. ⚠︎ On this corpus the answer is no: free code wins. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual document sets

The corpus is 64 document sets, 0.24 MB (txt 64). It is generated because it has to be. A real shipper's document sets are commercial records carrying real customers, real products with real classifications and desk notes written by named people. None of that can be published, and a corpus that could be published would have had the two things this kit measures — the sentence that decides a reading, and the print style that does not — normalised out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.

The corpus

  • The 64 document setsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every shipper, carrier, consignee, product, lot, order and document number is invented, SDR-2026 and HZT-2026 are invented, no product-to-hazard pairing is a claim about any real substance, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — every note speaks for a desk. evals/check_labels.py sweeps all 64 files for a person-shaped name on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your document sets. That is the whole change — there is no database to migrate.

One document set, as the model receives itSHP-0001.txt · 1 of 64
================================================================================
SHIPMENT DOCUMENT SET RECONCILIATION -- ONE SHIPMENT, ONE AS-AT DATE
================================================================================
FILE              SHP-0001
SHIPPER           S-310  Orrinford Specialty Chemicals
SHIPMENT          SH-66100
SALES ORDER       SO-48100
CARRIER           K-41  Varnhollow Haulage
MODE              road, packaged goods
SHIP DATE         2026-07-01
AS AT             2026-07-03
SET COMPILED      2026-07-05

-- HAZARD TABLE EXTRACT (HZT-2026, the shipper's own classification table) -----
PRODUCT  TRADE NAME                   SHIPPING NAME                      ID      CLASS  PG   TARIFF
PX-8450  SALINDREL DISPERSE 9         PIGMENT DISPERSION, ECOTOXIC       HZ-907  9      III  5712.30
PX-8451  SALINDREL DISPERSE 9W        PIGMENT DISPERSION, AQUEOUS        HZ-908  9      III  5712.39
PX-9560  OSCUVANE PRIMER 7            ADHESIVE PRIMER, FLAMMABLE         HZ-213  3      II   5935.06

-- SALES ORDER SO-48100 (order system export 2026-06-28) -----------------------
PRODUCT                 PX-8450
LOT ALLOCATED           LT-2606-101
NET QUANTITY            1,400.0 KG
PACKAGING               7 x 200 KG steel drum
CONSIGNEE               C-2207  Brindlecombe Coatings Ltd
ALSO TRADES AS          BCL Industrial; Brindlecombe Coatings East Depot

-- DOCUMENTS IN THIS SET -------------------------------------------------------
BILL OF LADING                 BL-77100
DANGEROUS GOODS DECLARATION    DGD-30100
CERTIFICATE OF ANALYSIS        COA-24100
COMMERCIAL INVOICE             CI-58100

-- BILL OF LADING BL-77100 (carrier K-41 Varnhollow Haulage, issued 2026-07-01) --
SHIPPER                 S-310 Orrinford Specialty Chemicals

Abridged — the file continues.

The outcomeWhat a good result looks like

One document set in, one row out: every cell that disagrees, every document not in the set or not in force, every exception and one of seven SDR-2026 verdicts. 34 of 64 sets come back with all four graded fields right — against 43 for the free rules floor, which reads no note. The paid call is ahead where a sentence decides (11 of 21 against 0) and on missing documents (64 against 60), and behind on the 39 sets the columns decide (19 against 39).

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 30 sets it got wrong are named in the kit README with what it answered. The dominant cause is print style read as disagreement — 52 false cells, 16 of them packing groups printed 'PG 3' or '2' against a roman table value — and twice it let a real disagreement through as RECONCILED (SHP-0053, SHP-0064). A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your documents are compared as printed — packing groups in roman or arabic, quantities in kilograms, tonnes or pounds, legal names with and without suffixes — the free rules floor, and do not buy a call at all
    43 of 64 sets for $0.00, perfect on the 39 the columns decide. The paid call gets 19 of those 39 because it reads print style as disagreement.
  • Order amendments, lot re-allocations, corrections and cancellations arrive as free text in the notes — the rules floor first, then a call ONLY on what the notes could overturn
    On the 21 sets a sentence decides the paid arm is 11 and the rules floor 0; on missing documents 64 against 60. That half is real — and smaller than the half it loses.
  • A set whose checklist already disagrees with itself — the rules floor
    The checklist status matches the key on 31 of 64; the rules floor gets 43 sets whole, the paid arm 34.
  • You need to know no reply will certify, sign or classify — any arm — the cap is a schema, not a model property
    0 on every committed arm and 0 of 21 under attack. The answer contract has no field for it and src/refusal.py reads the prose.
  • A string-similarity matcher looks like the cheap middle ground — the rules floor instead
    Similarity at 0.85 got 0 of 64 sets whole and flagged 259 agreeing cells.

At a glanceHow the whole thing runs

53%rechecked all correct pct
1,665 msp50, end to end
$1.61per 1,000 document sets · the fast tier

Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own document sets in the same labelled-block shape, data/hazard_table.json with your own table and data/shipments.json with your own register, then run python3 -m evals.run --run-id b000-<yours>-rules --floor rules — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per shipment to be told a roman III and an arabic 3 disagree. That is the case against the best-fitting scenario (“Your documents are compared as printed — packing groups in roman or arabic, quantities in kilograms, tonnes or pounds, legal names with and without suffixes”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A document that does not travel as labelled text — a scanned bill, a declaration image, a certificate table. The parser anchors on field labels; without them there are no cells. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO WIN OVER FREE CODE COULD BE VERIFIED. The free rules floor is whole on 67.2 pct of sets against the scored run's 53.1 pct (p = 0.150), and no change that might close the gap was measured. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-12 — r001-shipment-docset. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all six free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning