The business caseThe problem this solves
A producer delivers grain against one contract across a delivery period. The scale office weighs and grades every load and issues a scale ticket; weeks later the elevator issues a settlement statement and pays a net figure. Between the two sits a chain of arithmetic nobody at the producer's end can see: pounds to bushels, a moisture shrink at the contract's own factor, a price built from a futures month and a basis, three per-bushel discounts off whichever discount schedule governs, and freight that is deducted under one set of terms and not the other. Today the producer's only check is the elevator's own settlement panel, which verifies one thing -- that the printed lines add up to the net -- and reports ACCEPTED whenever they do, on settlements the contract says are hundreds of dollars out. The line-by-line recompute of a grain settlement -- every ticket through the contract's shrink, price and discount schedule, plus freight -- that a farm accountant does by hand, or does not do at all, before deciding whether to query the elevator.
Audience
The producer's agronomist or farm accountant deciding whether to query a settlement, and the elevator's settlement clerk deciding whether a query is right. The decision is not 'how much' -- it is 'which line departs, and what in the contract says so'. This report's own answer about buying a model call for it is a clear no on the headline and a qualified no on the reading. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual settlements
The corpus is 64 settlements, 0.20 MB (txt 64). There is no public corpus of commodity settlements and there could not be one: a settlement is a priced commercial document between two named counterparties, and the price, the basis and the discount schedule one elevator gave one producer are their commercial terms. A generated corpus is the only kind this kit could publish, and it buys something a scraped one never could -- an answer key that is TRUE BY CONSTRUCTION. The generator knows the true reading of every file, writes it, parses it back, and runs the kit's own station over the parsed file with that reading; whatever comes out IS the key. A perfect score therefore means exactly one thing: the three readings were right.
The corpus
- The 64 settlementsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your settlements. That is the whole change — there is no database to migrate.
GRAIN SETTLEMENT FILE GSF-0001
GSR-2026 -- one commodity settlement, recomputed against the governing contract and the
scale tickets. Every producer, elevator, farm, contract number, ticket number,
futures month and price below is INVENTED for this kit. No real contract, schedule
or settlement is quoted, paraphrased or relied on, and there is no personal data.
1. THE CONTRACT as signed
CONTRACT HC-1065
COMMODITY SOYBEANS NO 1
PRODUCER SILVERMERE ROW CROP CO
ELEVATOR TALLGRASS RIVER TERMINAL
DELIVERY PERIOD 2026-02-17 to 2026-03-25
SETTLEMENT DATE 2026-03-28
PRICING BASIS FLAT
FUTURES MONTH MAY-2026
FUTURES PRICE 11.3069
BASIS -0.5588
FLAT PRICE 11.6232
BUSHEL WEIGHT 60
DISCOUNT AUTHORITY SCHEDULE-B
SHRINK POLICY NONE
SHRINK BASE MOISTURE 14.5
SHRINK FACTOR 1.13
FREIGHT TERMS FOB-FARM
FREIGHT RATE 0.1390
TOLERANCE AMOUNT 29.11
TOLERANCE PERCENT 0.30
2. THE DISCOUNT SCHEDULES on file
SCHEDULE-A
DS-A01 MOISTURE each 1.0 pct over 15.5 at 0.0386 per bushel
DS-A02 TEST WEIGHT each 1.0 lb under 54.8 at 0.0486 per bushel
DS-A03 DAMAGE each 1.0 pct over 1.5 at 0.0687 per bushel
SCHEDULE-B
DS-B01 MOISTURE each 1.0 pct over 15.0 at 0.0531 per bushel
DS-B02 TEST WEIGHT each 1.0 lb under 54.0 at 0.0497 per bushel
DS-B03 DAMAGE each 1.0 pct over 2.5 at 0.0999 per bushel
3. THE SCALE TICKETS as weighed and graded
LINE TICKET DATE GROSS LB TARE LB NET LB MOIST TEST DMGAbridged — the file continues.
The outcomeWhat a good result looks like
One row per settlement: the net this contract and these tickets support, which discount schedule governs, which loads belong to the settlement at all, whether a post-delivery amendment moves anything, the exact lines that depart with each one quoted verbatim, and one of six verdicts naming what must move.
And when it cannot
A re-grade. A different moisture, test weight, damage figure or weight is a number NOTHING in the file supports -- not the ticket, not the contract, not either schedule -- and it turns a recompute into a negotiation with no record behind it. Across 64 settlements and 16 adversarial calls this kit proposed 0, including on the 2 corpus files carrying a settlement manager's note asking for exactly that. The failure it DOES ship is louder: its own arithmetic. On 61 of 64 settlements the net it computed was not the net the contract produces.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You want the settlement recomputed and the departing lines named — free code -- evals/baseline.py --floor strict
it is 78.1 pct right on the whole row against the paid call's 4.7 pct, it costs nothing, it never has a bad day and it re-runs forever - Your correspondence actually discusses changing the governing discount schedule — the paid call for the READING, with the arithmetic re-done in code
it read the governing schedule right on 64 of 64, including every file where a note agreed a change and every file where one was only proposed; the keyword floor gets 59 of 64 because it cannot read a verb - Your settlements are clean and you want a check, not a finding — the elevator's own footing check, or
panel
it catches every case where the printed lines do not add up to the net, which is the only failure that needs no contract at all
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own settlements in the same six-section layout -- the headings in src/statement.py::HEADINGS are exact-match anchors -- and rewrite data/gold.jsonl with your own key. The measured result does not travel with them. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for arithmetic. The model's own net is right on 3 of 64. That is the case against the best-fitting scenario (“You want the settlement recomputed and the departing lines named”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A statement that never settled a delivered load. Every departure in this corpus is a line that is present and wrong, because a citation quotes a printed line and an absent line cannot be quoted. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | REPEATABILITY. The scored arm ran ONCE. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN), one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-settlement-recompute. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board -- all 64 settlements, both discount schedules, the ticket-to-statement alignment and every committed arm -- and scores all three free floors offline at $0.00. The empty screenshot is that state, shot against a server whose API_KEY was blanked in its own environment. python3 -m evals.check_labels re-derives all 64 answer keys in a second implementation with nothing installed.





