Home › Use Cases › Re-do the grade-discount math on a grain settlement against the schedule in force
Use caseUC0272
🧪 Use-case kit · runnable

Re-do the grade-discount math on a grain settlement against the schedule in force

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A grain elevator settles a load and takes a grade discount off the gross: so much per bushel for every tenth of a point the moisture, test weight, foreign material, total damage or heat damage sits outside the bracket. The schedule is published, it is revised during the season, and the version that governs is the one in force ON THE DELIVERY DATE — not the one printed on the statement, and not the one the settlement system happened to apply. Nobody re-does the arithmetic. The grower reads a total, the desk reads a total, and the one number neither of them has is what the schedule actually owed. Re-doing the bracket arithmetic on a settlement statement by hand against a schedule history, which is what happens today only when a grower complains.

Audience

A settlement clerk or grain accountant with a season of tickets and a schedule that was revised twice, deciding which statements to open. The answer this report gives is a qualified one: the arithmetic is a solved problem for $0.00 and the READING is not, so the decision is whether the reading is worth $0.0086 a statement. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual settlement tickets

The corpus is 62 settlement tickets, 0.09 MB (json 3 · jsonl 1 · md 2 · txt 62). Grain settlement statements are not public: they are commercial documents between an elevator and a named grower, and the interesting ones are the disputed ones. So the corpus is generated, and it is generated to make the arithmetic REAL — every ticket is built as a structure, the five factors are priced against the governing version's own brackets in integer cents with each factor floored separately, and the panels are rendered FROM those numbers. The gold row is then src/schedule.py run over the same structure, so a corpus change cannot leave a stale key behind. Five house layouts exist for one reason: three of them do not print the delivery date in ISO, and the delivery date is the field the whole schedule selection turns on.

The corpus

  • The 62 settlement ticketsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your settlement tickets. That is the whole change — there is no database to migrate.

One settlement ticket, as the model receives itGDT-0001.txt · 1 of 62
GRAIN SETTLEMENT STATEMENT                                 GDT-0001
Wrenfield Grain Co-operative, House 3
Settled under discount schedule DS-2026R2

LOAD FACTS
  Delivered (unload)            2026-07-17
  Settlement date               2026-07-30
  Statement printed             2026-07-31
  Ticket number                 GDT-0001
  Commodity                     soybeans
  Contract reference            CT-4076
  Net bushels                   856.00
  Schedule applied              DS-2026R2

GRADE FACTORS AS ASSAYED
  Moisture                      13.2 pct
  Test weight                   56.5 lb/bu
  Foreign material              4.4 pct
  Total damage                  6.6 pct
  Heat-damaged kernels          0.5 pct

DEDUCTIONS AS APPLIED
  foreign material discount      4.4 pct over 2.0 pct @ 3.0 c/bu per pt               $        61.63
  total damage discount          6.6 pct over 5.0 pct @ 2.5 c/bu per pt               $        34.24
  heat-damaged kernels discount  0.5 pct over 0.2 pct @ 25.0 c/bu per pt              $        64.20
  handling fee                   unload and elevate, flat                             $        29.43
  TOTAL DEDUCTIONS                                                                    $       189.50

SETTLEMENT NOTES
  Contract CT-4076 is priced and this delivery is applied against it in full.
  The grower elected cash settlement; no storage or deferred pricing applies to this load.

The outcomeWhat a good result looks like

Per statement: the delivery date, the discount the statement applied, what the governing version owed, the signed delta in cents, whom it favoured, one of six outcomes, one of three verdicts and the line that establishes the finding. Across a season: every ticket computed on a superseded schedule, the dollar delta, and which side it fell on.

And when it cannot

Three ways, all of them recorded rather than hidden. It answers UNVERIFIABLE when the grade factors are not on the statement — 7 of 62 tickets here — and that is a REFUSAL to conclude, not a finding. It answers a finding that is wrong: 1 of 62 on this run (GDT-0050, where it named a non-discount deduction on a statement that carries none). And it quotes a line the key does not name: 7 of 62, every one of them located in the statement and every one a line a reviewer might well accept — which is a property of the key's tie-break as much as of the arm, and both readings are published.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A season of tickets against a schedule that was revised mid-season — the paid call, with the rule engine bolted to it
    the version in force is selected from the DELIVERY DATE, and on 23 of 62 statements here the delivery date is not the first date printed. The floor gets 23 of those wrong; the paid arm gets all 62 dates right.
  • Statements from ONE elevator in ONE fixed format — the free rules floor
    a regular expression over a fixed-layout panel reads the printed figures exactly — 57 of 62 applied figures on this corpus, for $0.00, and 12 of 12 on the tickets carrying the confusable charge.
  • Deciding whether a discount was applied correctly on ONE disputed ticket — the board, not the batch
    src/app.py prices the five factors one at a time against the governing version's own brackets and prints the whole published version history beside it, with no model in the loop and no key configured.

And where nothing here is good enough:

  • Deciding whether the GRADE itself is right — nothing here
    the cap is final-grade-determination and the answer contract has no field that could express a grade. 3 tickets carry a grower letter asking for a re-assay and the correct behaviour on all of them is to not engage.

At a glanceHow the whole thing runs

98%graded cells correct pct
102,391 msp50, end to end
$41.34per 1,000 settlement tickets · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/schedule.json with your own brackets, rates, effective ranges and rule order, and data/schedule.md with the same rulebook as prose — the prompt inserts that file verbatim, so the two cannot drift. The measured margin does NOT travel. Corpus lens →
When is this the wrong choice?Avoid: Believing the margin transfers. It was measured on five layouts this kit invented. That is the case against the best-fitting scenario (“A season of tickets against a schedule that was revised mid-season”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A statement whose deductions panel is not fixed-layout — a scanned exception note, a hand-annotated correction. Every expression in the free floor is bound to the printed column positions, and nothing in this kit detects that it is reading something else. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?That the margin survives on real settlement statements. Every statement here was generated by this kit, in five layouts this kit chose. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-grade-discount. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Observed from a clean checkout with no key configured: src/app.py serves the whole board, prices all five factors on every statement, computes the free rules floor live on any statement, replays the committed run and renders the 62-row corpus table — every route but /api/check, which returns 200 saying nothing was called. tools/build_corpus.py --check rebuilt the corpus byte-identically under two different PYTHONHASHSEEDs, and evals/check_labels.py re-derived all 62 rows at 0 problems. What could NOT be reproduced without a key is the paid arm itself: python3 -m evals.run --stub proves the whole pipeline including the scorer, but the reply it scores is a stub.

A living map of modern AI — kept current every morning