The business caseThe problem this solves
A grain elevator's position desk carries a BOOK balance for every bin — what it opened with, every scale ticket posted in and out, every transfer to and from another bin, and the shrink its own posted schedule computes on what came in. Once a period the bin is MEASURED with a gauge or a bin scan. The two do not agree, and the difference is somebody's grain: the elevator's own, a depositor's open storage, or grain against which a warehouse receipt has been issued. Reconciling one bin means reading every note printed under a ticket to find the loads that never moved, recomputing the shrink at the schedule that was actually in force, and then deciding whose bushels the gap reaches. Opening one bin's file, adding the ticket column up by hand, reading each note printed under a row, recomputing the shrink per ticket against the base moisture, and walking the ownership split to decide whose grain a shortfall reaches.
Audience
A country elevator's position desk or grain accountant working a month-end bin list, and the merchandiser who has to explain a shortfall to a depositor. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual bin reconciliation file
The corpus is 62 bin reconciliation file, 0.14 MB (txt 62). It is generated because it has to be. A real bin reconciliation is a licensed warehouseman's own trading record: it carries producer names against open-storage balances, warehouse-receipt numbers against bonded grain, and — on exactly the files that make this job interesting — a location manager writing down that a load was turned back or that a shortfall should come out of somebody's storage. Those are the rows a grain business would least want published and the only rows worth measuring. BIR-2026 is invented for the same reason and is named as invented on every surface.
The corpus
- The 62 bin reconciliation filegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every location is an invented trading name, every bin id, ticket reference, schedule code and moisture reading is arithmetic on the file index, and there is no personal data at all — evals/check_labels.py sweeps all 62 files for five families of identifier on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your bin reconciliation file. That is the whole change — there is no database to migrate.
BIN RECONCILIATION FILE BIN-0001
Location: LOC-02 - Harrow Creek Elevator (invented)
Bin: BN-10 Yellow corn 2026 Period: 2026-05-01 to 2026-05-30 Procedure: BIR-2026
==============================================================================
BIN AND PERIOD AS THE POSITION BOOK HOLDS IT
stock unit bu
test weight 56.000 lb per bu
capacity 148000.000 bu
opening book balance 68000.000 bu
measured inventory 76176.500 bu
measurement method bin-scan
tolerance pct 0.30
tolerance min 150.000 bu
OWNERSHIP AS THE POSITION BOOK HOLDS IT
company-owned 40120.000 bu
open storage 11988.000 bu
warehouse receipt 15892.000 bu
SHRINK SCHEDULE AS POSTED
SS-2200 moisture shrink 1.200 pct per point over 15.0 pct base
SS-2300 handling shrink 0.400 pct of gross inbound
TICKET LEDGER AS POSTED BY THE POSITION DESK
TKT-0101 2026-05-03 10000.000 inbound POSTED SC-6371 gross 10000.000 bu at 14.0 pct moisture
TKT-0102 2026-05-07 13000.000 inbound POSTED SC-4020 gross 13000.000 bu at 18.5 pct moisture
TKT-0103 2026-05-11 6000.000 inbound POSTED SC-8832 gross 6000.000 bu at 18.3 pct moisture
TKT-0104 2026-05-05 9000.000 outbound POSTED LO-6147 net 504000.000 lb at 56.000 lb per bu
TKT-0105 2026-05-10 9000.000 outbound POSTED LO-4601 net 504000.000 lb at 56.000 lb per bu
TRF-0106 2026-05-10 2000.000 transfer-out POSTED TR-8586 moved to BN-13 for drying
DRY-0107 2026-05-14 471.000 drying-loss LOGGED DL-1750 burned off in the continuous-flow dryerAbridged — the file continues.
The outcomeWhat a good result looks like
One bin in, one row out: which posted tickets this procedure treats differently from the ledger that printed them, with the ledger row quoted verbatim for each; the shrink schedule in force; the book balance, the difference against the measurement and its percentage; one of five verdicts; and whose grain the difference is against.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed, 0 calls failed and 0 stopped at the token ceiling — so there is no no-answer rate to publish and none is invented. What it does get wrong is published by name: the arm cited nothing on 3 of the 5 bins whose quantity column is a net weight, cited rows the key does not label on 3 more, and cited a VOID row the desk had already excluded on 4.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your ticket ledger's bad rows announce themselves — every rejected load says "rejected", every duplicate says "duplicate" — the free rules floor, and do not buy a call at all
The floor is 37 of 62 here and it reaches everything a keyword can reach for $0.00. On the pounds column it BEATS the paid arm 3 to 2. - Your location notes routinely describe a rejection, a wrong bin or a double keying OF SOMETHING ELSE — an earlier load on the same contract, another bin the same day — the paid call, and read the keep-note family's own rate
10 of 15 against 0 of 15. Every removal word a rules engine looks for is inside those sentences and the row still moved. - Your shrink schedule gets revised, and revisions get circulated to producers before they are adopted — the paid call
8 of 9 against 0 of 9 on revisions that were never adopted. A regex applies every one it can parse, and a circulated revision states its code, its factor and its date just as fully as an adopted one. - You need to know whose grain a shortfall reaches, not just how big it is — either arm, and read B-8 rather than the model
The ownership band is derived by src/policy.py from the register's own split, so it is free and it is the same for every arm once the reading is right. 57 of 62 on the paid arm, 52 on the free floor — the gap is the reading upstream, not the ladder. - Your bins are reconciled monthly across a few hundred bins — the paid call, and price it
$0.000739 a bin measured off-peak. 400 bins a month is under $0.30.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own bin files in the same shape and data/bins.json with your own register, then run the three free floors and evals/check_labels.py before buying a call. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per bin for a regex you could write in an afternoon. That is the case against the best-fitting scenario (“Your ticket ledger's bad rows announce themselves — every rejected load says "rejected", every duplicate says "duplicate"”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A ledger that is not fixed-width columns. src/ledger.py's row regex is the shape these files print; a CSV or a grain-accounting API export needs a different parser and nothing above it changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-09 — r001-bin-inventory. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run, and python3 tools/build_corpus.py --check rebuilds every byte and diffs at 0 problems. Nothing in the kit needs a credential to be read.






