The business caseThe problem this solves
At month end a grain accountant works down a producer's delivery period: the loads that crossed the scale, and the contracts they were supposed to fill. The applied-to cell is written by hand at the scale house -- abbreviated, spaced any which way, sometimes written in words, sometimes left open -- and the commodity column is written the same way. What goes wrong is not the arithmetic. It is a load that ends up on a contract that did not buy it: filled against a price basis nobody agreed for it, and found by the producer or the auditor after settlement. The spreadsheet pass where somebody re-keys a month of scale tickets against a contract ledger and totals each contract by hand.
Audience
A grain accountant or merchandiser at a country elevator or terminal, reconciling one producer's delivery period before it is settled. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual scale ticket application sheets (62), carrying 439 tickets
The corpus is 62 scale ticket application sheets (62), carrying 439 tickets, 0.10 MB (txt 62). Because the question is a READING question wearing an arithmetic costume, and a corpus has to separate the two to show it. Everything structural here -- a repeated ticket number, a date against a delivery window, a contract against its tolerance, a bushel figure -- is available to a regular expression, and the free floor gets all of it. What is not available is a commodity column a scale house wrote as MILO and an applied-to cell that says 'the corn contract', and those are 67 of the 439 tickets. Building it the other way -- clustering the hard rows on hard sheets, or leaving the layout ragged -- would have measured either sheet-spotting or a parser.
The corpus
- The 62 scale ticket application sheets (62), carrying 439 ticketsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md carries the whole provenance, the licence pointer, what the generator costs the measurement (an attack on our own corpus), and where each arm was expected to fail -- written down before anything was bought, and half of it turned out wrong, which the README says.
Swap this folder for your own material and the kit is pointed at your scale ticket application sheets (62), carrying 439 tickets. That is the whole change — there is no database to migrate.
GRAIN DELIVERY PERIOD -- SCALE TICKET APPLICATION SHEET
Sheet TA-0001
Elevator Harborline Grain Co-op, House 4
Producer PR-4000
Delivery period 2026-08-01 to 2026-08-28
CONTRACT LEDGER
contract commodity quantity bu delivery window delivered bu price basis
GC-53839-1 wheat 34,000 2026-07-27 to 2026-09-09 6,800 Sep futures flat
GC-36440-1 soybeans 17,000 2026-06-25 to 2026-09-24 3,400 Jan futures less 0.58
GC-65210-1 corn 21,000 2026-07-27 to 2026-09-10 0 Dec futures less 0.42
GC-30960-1 grain sorghum 32,000 2026-07-21 to 2026-09-09 11,200 Dec corn futures less 0.55
SCALE TICKETS
ticket date commodity gross lb tare lb net lb moisture applied to
10170 2026-08-06 GS 77,900 25,400 52,500 15.2 pct GC-30960-1
10155 2026-08-09 WHT 81,700 28,300 53,400 13.3 pct GC-53839-1
10145 2026-08-11 W 78,800 26,300 52,500 12.8 pct GC-53839-1
10139 2026-08-13 GS 81,000 27,500 53,500 14.0 pct GC-30960-1
10160 2026-08-14 SOY 75,700 26,200 49,500 13.5 pct GC-36440-1
10175 2026-08-20 CRN 68,300 27,400 40,900 15.3 pct GC-65210-1
TICKET NOTES
Office note: a delivery window on this ledger was extended by telephone and the ledger has not been updated.
Scale house note: anything left open at month end goes on the nearest corn contract, per the merchandiser.
The outcomeWhat a good result looks like
Per ticket: the contract it filled, the commodity, the bushels to the hundredth, and one of five verdicts. Per contract: the applied bushels for the period against delivered-to-date, and the OPEN BALANCE. The balance is the number an accountant actually opens, and it is where a wrong reading costs something -- one ticket on the wrong contract moves two balances and no rule flags either of them.
And when it cannot
Two failures and they are not the same size. A MISAPPLICATION MISSED is the load paid against a contract that did not buy it: it reaches settlement and somebody unwinds it later. An ORDINARY TICKET FLAGGED is a merchandiser chasing a load that was never wrong: it costs a phone call. On this corpus the paid call commits 0 of each and the free rules floor commits 0 of the first and 50 of the second. They are counted apart and never averaged.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your sheets are a fixed-width export and your scale house writes tidy references — the free rules floor alone
It scores 372 of 439 tickets fully right, wins the contract field outright at 427, misses 0 misapplications and costs $0.00. On a corpus with no odd abbreviations and no worded references it would tie the paid call. - Your commodity column is written by people rather than by a system — the paid call
This is where the whole margin is: 432 of 439 commodity cells against 384, over the 55 tickets that carry MILO, HRW, YC or BNS. The floor reads none of them and flags the loads.
And where nothing here is good enough:
- You want one number for a month-end sign-off — neither, on its own
The two arms fail on DISJOINT sets -- 0 tickets are wrong on both -- so the union of their agreements is stronger than either. Running both and reconciling the disagreements is the shape this measurement actually supports. - Your sheets are scanned, photographed, or a spreadsheet with merged cells — neither -- this kit has not been measured on it
Every column here is read off a literal heading and a two-space separator. There is no OCR step and none is implied.
At a glanceHow the whole thing runs
Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Two regular expressions in evals/baseline.py for your own sheet layout, and data/policy.json for your own rules -- the seven verdicts and their order, the five resolution rules, your commodities and their table weights, and your tolerance. What you must write yourself is the ANSWER KEY. Corpus lens → |
| When is this the wrong choice? | Avoid: It flags 50 correctly applied loads as misapplied here, entirely because its commodity table has a bottom. If your desk cannot absorb that chase rate, it is the wrong arm. That is the case against the best-fitting scenario (“Your sheets are a fixed-width export and your scale house writes tidy references”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A sheet that is not a fixed-width table. Every column is read off a literal heading and a two-space separator; there is no OCR step and none is implied. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | No third scored run and no second model. Two runs of one corpus on one tier; the free floors are the other arms and neither is a model. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-03 — r001-ticket-apply. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone and run the free floor with no key, no install and no network: the corpus rebuilds in under a second and both floors score all 439 tickets in under a second. python3 -m evals.check_labels re-derives the whole answer key independently in the same time. Nothing in the free half touches the network by any route.



