Home › Use Cases › Reconcile a producer's scale tickets against the grain contracts they were applied to
Use caseUC0298
🧪 Use-case kit · runnable

Reconcile a producer's scale tickets against the grain contracts they were applied to

A small, forkable project that does one job end to end. Run twice for real over the same set, and every figure on these pages captured from those runs.

The business caseThe problem this solves

At month end a grain accountant works down a producer's delivery period: the loads that crossed the scale, and the contracts they were supposed to fill. The applied-to cell is written by hand at the scale house -- abbreviated, spaced any which way, sometimes written in words, sometimes left open -- and the commodity column is written the same way. What goes wrong is not the arithmetic. It is a load that ends up on a contract that did not buy it: filled against a price basis nobody agreed for it, and found by the producer or the auditor after settlement. The spreadsheet pass where somebody re-keys a month of scale tickets against a contract ledger and totals each contract by hand.

Audience

A grain accountant or merchandiser at a country elevator or terminal, reconciling one producer's delivery period before it is settled. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual scale ticket application sheets (62), carrying 439 tickets

The corpus is 62 scale ticket application sheets (62), carrying 439 tickets, 0.10 MB (txt 62). Because the question is a READING question wearing an arithmetic costume, and a corpus has to separate the two to show it. Everything structural here -- a repeated ticket number, a date against a delivery window, a contract against its tolerance, a bushel figure -- is available to a regular expression, and the free floor gets all of it. What is not available is a commodity column a scale house wrote as MILO and an applied-to cell that says 'the corn contract', and those are 67 of the 439 tickets. Building it the other way -- clustering the hard rows on hard sheets, or leaving the layout ragged -- would have measured either sheet-spotting or a parser.

The corpus

  • The 62 scale ticket application sheets (62), carrying 439 ticketsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md carries the whole provenance, the licence pointer, what the generator costs the measurement (an attack on our own corpus), and where each arm was expected to fail -- written down before anything was bought, and half of it turned out wrong, which the README says.

Swap this folder for your own material and the kit is pointed at your scale ticket application sheets (62), carrying 439 tickets. That is the whole change — there is no database to migrate.

One scale ticket application sheets (62), carrying 439 ticket, as the model receives itTA-0001.txt · 1 of 62
GRAIN DELIVERY PERIOD -- SCALE TICKET APPLICATION SHEET

  Sheet             TA-0001
  Elevator          Harborline Grain Co-op, House 4
  Producer          PR-4000
  Delivery period   2026-08-01 to 2026-08-28

CONTRACT LEDGER
  contract        commodity         quantity bu    delivery window               delivered bu    price basis
  GC-53839-1      wheat             34,000         2026-07-27 to 2026-09-09      6,800           Sep futures flat
  GC-36440-1      soybeans          17,000         2026-06-25 to 2026-09-24      3,400           Jan futures less 0.58
  GC-65210-1      corn              21,000         2026-07-27 to 2026-09-10      0               Dec futures less 0.42
  GC-30960-1      grain sorghum     32,000         2026-07-21 to 2026-09-09      11,200          Dec corn futures less 0.55

SCALE TICKETS
  ticket     date          commodity    gross lb     tare lb     net lb      moisture     applied to
  10170      2026-08-06    GS           77,900       25,400      52,500      15.2 pct     GC-30960-1
  10155      2026-08-09    WHT          81,700       28,300      53,400      13.3 pct     GC-53839-1
  10145      2026-08-11    W            78,800       26,300      52,500      12.8 pct     GC-53839-1
  10139      2026-08-13    GS           81,000       27,500      53,500      14.0 pct     GC-30960-1
  10160      2026-08-14    SOY          75,700       26,200      49,500      13.5 pct     GC-36440-1
  10175      2026-08-20    CRN          68,300       27,400      40,900      15.3 pct     GC-65210-1

TICKET NOTES
  Office note: a delivery window on this ledger was extended by telephone and the ledger has not been updated.
  Scale house note: anything left open at month end goes on the nearest corn contract, per the merchandiser.

The outcomeWhat a good result looks like

Per ticket: the contract it filled, the commodity, the bushels to the hundredth, and one of five verdicts. Per contract: the applied bushels for the period against delivered-to-date, and the OPEN BALANCE. The balance is the number an accountant actually opens, and it is where a wrong reading costs something -- one ticket on the wrong contract moves two balances and no rule flags either of them.

And when it cannot

Two failures and they are not the same size. A MISAPPLICATION MISSED is the load paid against a contract that did not buy it: it reaches settlement and somebody unwinds it later. An ORDINARY TICKET FLAGGED is a merchandiser chasing a load that was never wrong: it costs a phone call. On this corpus the paid call commits 0 of each and the free rules floor commits 0 of the first and 50 of the second. They are counted apart and never averaged.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your sheets are a fixed-width export and your scale house writes tidy references — the free rules floor alone
    It scores 372 of 439 tickets fully right, wins the contract field outright at 427, misses 0 misapplications and costs $0.00. On a corpus with no odd abbreviations and no worded references it would tie the paid call.
  • Your commodity column is written by people rather than by a system — the paid call
    This is where the whole margin is: 432 of 439 commodity cells against 384, over the 55 tickets that carry MILO, HRW, YC or BNS. The floor reads none of them and flags the loads.

And where nothing here is good enough:

  • You want one number for a month-end sign-off — neither, on its own
    The two arms fail on DISJOINT sets -- 0 tickets are wrong on both -- so the union of their agreements is stronger than either. Running both and reconciling the disagreements is the shape this measurement actually supports.
  • Your sheets are scanned, photographed, or a spreadsheet with merged cells — neither -- this kit has not been measured on it
    Every column here is read off a literal heading and a two-space separator. There is no OCR step and none is implied.

At a glanceHow the whole thing runs

96%ticket all correct pct · 2 runs, no ordering
77,124 msp50, end to end
$15.64per 1,000 scale ticket sheets · GPT-5.6 Luna

Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Two regular expressions in evals/baseline.py for your own sheet layout, and data/policy.json for your own rules -- the seven verdicts and their order, the five resolution rules, your commodities and their table weights, and your tolerance. What you must write yourself is the ANSWER KEY. Corpus lens →
When is this the wrong choice?Avoid: It flags 50 correctly applied loads as misapplied here, entirely because its commodity table has a bottom. If your desk cannot absorb that chase rate, it is the wrong arm. That is the case against the best-fitting scenario (“Your sheets are a fixed-width export and your scale house writes tidy references”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A sheet that is not a fixed-width table. Every column is read off a literal heading and a two-space separator; there is no OCR step and none is implied. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?No third scored run and no second model. Two runs of one corpus on one tier; the free floors are the other arms and neither is a model. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-ticket-apply. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run the free floor with no key, no install and no network: the corpus rebuilds in under a second and both floors score all 439 tickets in under a second. python3 -m evals.check_labels re-derives the whole answer key independently in the same time. Nothing in the free half touches the network by any route.

A living map of modern AI — kept current every morning