The business caseThe problem this solves
A grain merchandiser's contracts and settlements desk gets an amendment file: the contract terms of record, the amendments executed against them, the grower instructions that authorise those amendments, the pricing records behind a repricing, the confirmations and the desk review, and a handful of file notes written in prose. Seven AMND-2026 rules have to be answered and two gates sit beside them — was the amendment instructed, was the instruction not post-dated, is the price basis of record, does the repricing quote the reference of record, is it inside the 90-day window, does the amended quantity reconcile, is the delivery period of record — each MET with the record ids that carry it or BREACHED for one of seventeen stated reasons. Almost all of that is column work. On 20 of the 64 files in this corpus a FILE NOTE somewhere else in the file takes a printed record out of play, and 10 of those notes are dated AFTER the file itself, which under AMND-2026 saves three rules and not the other six. That sentence is what a desk reads for and it is the only thing this kit buys. The desk minutes spent reading every printed record block against seven rules and two gates, testing each candidate row against the file's declared register scope and 90-day pricing window, and reading every file note for whether it takes a record out of play AS AT the date the note carries.
Audience
The contracts and settlements desk of a grain merchandiser checking an amendment and a repricing before a settlement runs, and the reviewer whose desk review and grower confirmation the file has to evidence. The decision it supports is "is this file clean enough to settle on" — never whether the price may be set, never whether the amendment may be made, and never whether anybody is in breach. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual contract amendment files
The corpus is 64 contract amendment files, 0.14 MB (json 4 · jsonl 1 · md 2 · txt 64). No public corpus of grain contract amendment files exists, and one built from real trade files would carry real growers, real prices and real disputes. Every file here is generated in process from SEED 20264890, so the key is DERIVED by the same rulebook the kit applies and the whole set rebuilds byte-identically under two different hash seeds. The mixture is the measurement: 14 files a file note decides, 40 carrying a decoy note that names the record which STANDS inside a withdrawal-shaped sentence, 10 where the deciding note is dated AFTER the file date, 14 source traps, 7 arithmetic files, 7 with two exceptions and 10 carrying a desk note that asks in terms for something AMND-2026 never does.
The corpus
- The 64 contract amendment filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every contract, grower, commodity, register, reference and record id is invented, generated from SEED 20264890, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL - a grower is a code, a confirmer is a role, and every file note speaks for the contracts, merchandising, settlements, grain accounting, origination or trade desk. evals/check_labels.py sweeps all 64 files for a person-shaped name, an honorific and an unattributed note on every run and reports 0. The best free-code floor is published there too: 44 of 64 whole-file, measured before any call was bought.
Swap this folder for your own material and the kit is pointed at your contract amendment files. That is the whole change — there is no database to migrate.
COMMODITY CONTRACT AMENDMENT FILE CTR-26-0001
FILE ASSEMBLED 2026-06-05 STANDARD AMND-2026
GROWER GRW-3113 COMMODITY CMD-CAN SEASON S2026
CONTRACT DATE 2026-01-26 FILE DATE 2026-05-31
DELIVERY PERIOD OF RECORD 2026-10 CONTRACTED QUANTITY 5000 BU AMENDED QUANTITY 7500 BU
DECLARED SCOPE REGISTERS CONTRACTBOOK, PRICEDESK PRICING WINDOW 2026-02-09 TO 2026-05-10
[1] CONTRACT TERMS
ID DATED TYPE REGISTER REFERENCE VALUE STATUS
TRM-70000 2026-01-26 QUANTITY CONTRACTBOOK - 5000 EXECUTED
TRM-70001 2026-01-26 DELIVERY-PERIOD CONTRACTBOOK - 2026-09 EXECUTED
TRM-70002 2026-01-26 PRICE-BASIS PRICEDESK REF-4409 - EXECUTED
TRM-70003 2026-01-26 QUALITY-SPEC CONTRACTBOOK - - EXECUTED
[2] AMENDMENTS
ID EXECUTED EFFECTIVE AFFECTS REGISTER REFERENCE VALUE STATUS
AMD-30000 2026-03-01 2026-03-04 QUANTITY CONTRACTBOOK - +2500 EXECUTED
AMD-30001 2026-03-14 2026-03-19 DELIVERY-PERIOD CONTRACTBOOK - 2026-10 EXECUTED
[3] INSTRUCTIONS ON FILE
ID RECORDED CHANNEL SUBJECT SCOPE STATUS
INS-50000 2026-02-27 PORTAL AMD-30000 QUANTITY RECORDED
INS-50001 2026-03-11 DESK-CALL AMD-30001 DELIVERY-PERIOD RECORDED
[4] PRICING RECORDS
ID PRICED KIND REGISTER REFERENCE STATUS
PRC-61000 2026-02-11 PRICE-SET PRICEDESK REF-4409 POSTED
PRC-61001 2026-03-04 REPRICE PRICEDESK REF-4409 POSTED
PRC-61002 2026-04-02 REPRICE PRICEDESK REF-4409 POSTED
[5] CONFIRMATION AND DESK REVIEW
ID DATED KIND SUBJECT ROLE STATUSAbridged — the file continues.
The outcomeWhat a good result looks like
One contract amendment file in, one determination out: the seven AMND-2026 rule rows with the record ids that support each or the reason it fails, one FILE-VERIFIED or FILE-EXCEPTIONS verdict, both gate rows — the desk review of THIS file and the confirm-with-grower gate on every repricing — a second GATE-SATISFIED or GATE-BREACH verdict, and the three lineage lists the cited records drag out with them. 63 of 64 files come back with all seven graded fields right once the free station has re-applied the rulebook, against 44 for the floor of record.
And when it cannot
And what it does when it cannot. 64 of 64 replies parsed, 0 stopped at the output ceiling and no call failed. The one file it gets wrong is named in the kit README with the sentence that decided it — CTR-26-0016, where a 2026-06-09 file note withdraws desk-review record DRV-11105 on a file dated 2026-06-07. The note is LATER than the file, which under AMND-2026 saves nothing on I1 because the desk-review gate is retrospective; the call returned an empty withdrawal map, the station found DRV-11105 printed and complete, and the file came back GATE-SATISFIED against a GATE-BREACH key. Five replies returned a row that is not merely wrong but IMPOSSIBLE — a BREACHED rule carrying a source, or a MET one carrying none — and every one is flagged and recorded rather than silently corrected. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- The record changes in your amendment files arrive as a STATUS column, not as a sentence — the printed-columns floor (b000-amendment-reprice-columns)
on the 44 files this corpus decides from the columns alone, the free rule and the paid call are level at 44 of 44, p = 1.0000. Nothing is bought. - Records are taken out of play by a desk note in prose, and the note's DATE decides what it touches — the scored arm — one call for the reading, free code for everything else
19 of the 20 files a sentence decides, against 0 for every free arm anybody could write from the domain. p = 0.000004 against the floor of record. - You wrote the generator, or your corpus has a fixed house phrasing for withdrawals — the free
tunedrule
20 of 20 on the predicted slice and 64 of 64 whole, at $0.00 — it beats the call. A regex over a phrasing you control is not an oracle, it is maintenance. - You need a number to put in front of a reviewer today — the WHOLE-FILE metric on both columns, never a single cell
the modal complete answer reaches 1 of 64, so whole-file is the one figure a fixed answer cannot reach. The gate verdict, the file verdict and any single rule's pass rate reach 49, 35 and 55 of 64 respectively on a constant.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own contract amendment files in the same shape — header, DECLARED SCOPE line, printed record blocks, file notes — and data/contracts.json with your own register, then edit data/policy.json so the seven rules, the seventeen reason codes, the per-rule ladders, the two gates and the note-effect split are YOURS. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Buying a call per file to re-read columns you already have in a database. That is the case against the best-fitting scenario (“The record changes in your amendment files arrive as a STATUS column, not as a sentence”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A file with no DECLARED SCOPE line. The register-scope and pricing-window tests are "inside the declared registers and the declared window"; with neither declared there is nothing to test against and the kit would be inventing a scope. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | THE HEADLINE IS A WHOLE-FILE FIGURE AND THREE CELLS INSIDE IT ARE ALREADY WON BY FREE CODE — none of them may be quoted alone, here or on any card, deck or reel. A fixed GATE-SATISFIED reaches 49 of 64 (76.6 pct), a fixed FILE-VERIFIED 35 of 64 (54.7 pct) and a fixed MET on C1 alone 55 of 64 (85.9 pct); WHOLE-FILE is safe as the headline only because the modal complete answer reaches 1 of 64. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-amendment-reprice. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, runs the free floors live in the browser and replays every committed arm. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run; without one the button is disabled and says so.












