The business caseThe problem this solves
A shipment arrives with two documents that were produced by two different people at two different times: a commercial invoice cut by the seller's sales system, and a packing list printed in the warehouse as the cartons were filled. An entry preparer has to put them side by side and answer, line by line, whether they describe the same goods in the same quantities at figures that add up — before anything downstream can happen. Today that is done by eye, and the part that takes the time is not the arithmetic. It is deciding that HEX BOLT ZN M8x40 DIN933 on the invoice and M8 HEX BOLT ZN 40MM in the warehouse are the same product, on the rows where the warehouse did not write the part number down. Putting an invoice and a packing list side by side and checking, line by line, that the goods match, the quantities tie out, every extension multiplies, the origins are stated, and the declared totals, package count and weights agree with the rows — then writing the conflicts down.
Audience
An entry preparer or a customs broker's clerk working a shipment's paperwork, and the person deciding whether a model belongs anywhere in that job. ⚠︎ THIS REPORT'S OWN ANSWER TO THE SECOND QUESTION IS NO, on this corpus, and that is the finding rather than a caveat. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual shipment document sets
The corpus is 60 shipment document sets, 0.11 MB (txt 60). It is generated because it has to be. A real commercial invoice and packing list carry an importer of record, a broker's licence, tax identifiers and banking details, and no such file can be published. What matters here is not the parties but the SHAPE — a fixed-column invoice, a packing list with its own reference column, and the specific places that column is left blank. The 15 families are built around the one question worth measuring: 19 of 326 invoice lines have no reference to join on, and 10 of those are two near-identical products a page apart with nothing printed anywhere to separate them.
The corpus
- The 60 shipment document setsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — every one of the 60 document sets and the whole answer key is generated. Nothing here is fetched, scraped, licensed or derived from anything that was.
Swap this folder for your own material and the kit is pointed at your shipment document sets. That is the whole change — there is no database to migrate.
SHIPMENT DOCUMENT SET - COMMERCIAL INVOICE AND PACKING LIST
SET HEADER
Shipment SH-0001
Seller Anadolu Doku AS, Bursa
Buyer Arcadia Trading LLC, Newark
Invoice number ANA-77262
Invoice date 2026-06-16
Delivery term FOB Izmir
Currency USD
COMMERCIAL INVOICE
# Part Qty Unit Unit price Extended Origin Description
1 Z-6451 20 piece $4.28 $85.60 MY Copper braid earthing strap, 25 mm2, 300 mm
2 N-5818 72 piece $0.96 $69.12 SE Rubber-lined pipe clamp, 32 mm
3 B-1394 12 roll $2.89 $34.68 TR Kraft paper tape, 48 mm x 50 m, reinforced
4 D-6920 250 length $19.60 $4,900.00 IN Aluminium extrusion, 40 x 40 slot profile, 3 m
INVOICE TOTAL $5,089.40
PACKING LIST
# Ref Qty Pkgs Gross kg Net kg Origin Marks Description
1 D-6920 250 55 1399.500 1350.000 IN SK-4 ALU PROF 40X40 3M
2 Z-6451 20 1 4.500 3.600 MY BX-1 CU BRAID EARTH 25 300
3 B-1394 12 1 5.460 4.560 TR PL-3 TAPE KRAFT 48X50 REIN
4 N-5818 72 1 6.228 5.328 SE CT-2 CLAMP RUB LINED 32
DECLARED TOTALS
Total packages 58
Gross weight kg 1415.688
Net weight kg 1363.488
SET NOTES
The packing list was produced two days after the invoice.
END OF DOCUMENT SET
The outcomeWhat a good result looks like
One shipment in, one entry worksheet out: every invoice line with the packing rows that carry its goods, the quantity, unit price, extended value and country of origin copied as printed, one ENT-2026 finding per line with the rule it cites, the six declared header fields, every document-level conflict, the exact set of conflicting lines and one status — READY or HOLD.
And when it cannot
And what it does when it cannot. On 6 of the 60 calls of the scored run the reply never contained a worksheet at all: with reasoning sent disabled, the model wrote its arithmetic out as prose in the answer channel and stopped. finish_reason is stop on all 6 of them and none reached the 8000-token ceiling — the largest ran to 1936 — so this is format adherence and not truncation. ⚑ IT WAS PARTLY TRUNCATION BEFORE, AND THAT IS MEASURED RATHER THAN ASSUMED: the first pair of runs used a 3,000-token ceiling, 14 and 8 replies failed to parse, and one hit 3,000 exactly. Raising the ceiling to 8000 cut the failures to 6 and 6 and lifted the rechecked headline from 76.7 and 86.7 pct to 90.0 and 90.0. Those calls are billed and counted as failures inside every published denominator; none was re-fired. The board shows the empty column and says why.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your packing list fills in its reference column reliably — the free rules floor, alone
307 of 326 lines here are settled by that column and nothing else, at $0.00 and with no network. There is no measured case on this corpus for spending anything. - Your reference column is often blank, and your two wordings share characters — the free rules floor, and measure it before assuming it fails
on this corpus trigram overlap pairs 10 of 10 genuinely contested rows correctly, which nobody expected when the trap was built - Your reference column is blank and the two wordings share almost no characters — unknown from this kit — the paid arm is the candidate, and it is NOT demonstrated here
this is the case the kit was built to measure and the corpus turned out not to contain it. The generator's two wordings of one product are too similar. - You need the arithmetic half only — extensions, totals, weights, package counts — src/policy.py, with no model at all
every one of those is integer arithmetic over parsed columns and is exact
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-05. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own document sets in the same columnar shape — a SET HEADER, a COMMERCIAL INVOICE panel, a PACKING LIST panel, DECLARED TOTALS and SET NOTES — then run python3 -m evals.run --run-id b001-yours --floor rules, which needs no key and no network. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO, AND IT IS THIS KIT'S HEADLINE. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call whose answer the join already produced. That is the case against the best-fitting scenario (“Your packing list fills in its reference column reliably”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A column that moves POSITION in the row. The parser reads runs of whitespace, so columns need not line up at the same character offsets — but an invoice printing origin before extended value parses to nonsense or not at all. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the paid arm would beat the free floor on descriptions that genuinely diverge. The corpus was built to contain that case and the measurement says it does not: a generic trigram matcher solved all ten contested lines. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-05 — r003-entry-worksheet. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9313 — all 60 document sets, both free floors computed live, ENT-2026, and every committed run replayed off its result file — and scores both free floors offline. The one control that would spend is disabled and says why. Nothing was installed: requirements.txt names no package.




