The business caseThe problem this solves
An agronomist is looking at one candidate product for one field and has to assemble what has already been applied there before anything else can be decided. The field's history is not one table: application records are filed against BLOCKS and tract numbers rather than the field, the same product appears under two trade names and one active substance, a formulation code and a strength token sit inside product names that are not part of them, a grower's nickname for a block is not the block's printed id, and a split pass written up as two records on consecutive days is one application event or two depending on a note somebody wrote. On top of that sits FAS-2026, the operation's own evidence standard, whose printed restrictions each need the record ids that put them in play — and two of whose revisions disagree, on purpose, about which soil test governs. Reading a field's whole multi-season application record by hand: deciding which of hundreds of records are the candidate's ACTIVE SUBSTANCE rather than a sibling trade name, which blocks and tract numbers belong to the REQUESTED field and which belong to a neighbouring one, which consecutive records are one application event and which are a retreatment, which records fall inside the declared lookback, which soil test governs under the revision in force, and then writing out each printed restriction with the record ids that engage it or the reason it cannot be told.
Audience
The agronomy desk of an arable farming operation assembling a field's application history against a candidate product's printed restrictions, and the operations, records, field, supply and store desks whose notes it reads. The decision it supports is 'what is already on this field and which restrictions does it engage' — never whether the application may be made. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual field evidence requests
The corpus is 64 field evidence requests, 0.26 MB (json 3 · jsonl 1 · md 2 · txt 64). It is generated because it has to be. A real field application record is a farming business's own commercial record: named operators and applicators, real trade names and registration numbers, parcel identifiers that locate a holding, and adviser notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the messy block and trade-name spellings — tidied out of it first. So the whole thing is invented, declared synthetic in data/SOURCES.md, and generated from SEED 20260921 with the key DERIVED by the same standard the kit applies. ⚠︎ The generated corpus is also what makes the floor honest AND what limits it: a rule written FROM the generator's own alias and event frames reaches 64 of 64, and that rule is published as the WORDING COST of a generated corpus, never as a floor anybody could have built.
The corpus
- The 64 field evidence requestsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every trade name, active substance, field, block, nickname, tract, record id, work order and applicator code is invented, generated from SEED 20260921, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL - an applicator is a code and every note speaks for the agronomy, operations, records, field, supply or store desk. evals/check_labels.py sweeps all 64 requests for a person-shaped name, an honorific and an unattributed note on every run and reports 0. The best domain-written free-code floor is published there too: 41 of 64 whole-request, measured before any call was bought, against a constant baseline of 2 of 64 and a generator-tuned ceiling of 64 of 64.
Swap this folder for your own material and the kit is pointed at your field evidence requests. That is the whole change — there is no database to migrate.
FIELD EVIDENCE REQUEST FER-0001
REQUEST RAISED 2025-12-02 STANDARD FAS-2026 REV A
FIELD FLD-2002 RIVER BOTTOM
CANDIDATE PRODUCT DURABEX 350 SC CODE PROD-5106 ACTIVE PENDOREX
SEASON 2026 SEASON STARTS 2025-11-01
DECLARED LOOKBACK 730 DAYS FROM 2023-12-03
[1] FIELD REGISTER
BLOCK FIELD NAME AS RECORDED AREA HA
BLK-2002A FLD-2002 RIVER BOTTOM 10.32
BLK-2002B FLD-2002 RIVER BOTTOM NORTH 6.55
BLK-2006A FLD-2006 STONE HILL 14.07
BLK-2006B FLD-2006 STONE HILL NORTH 7.07
[2] PRODUCT REGISTER
CODE TRADE NAME ACTIVE STRENGTH
PROD-5104 CEPTOREL 250 EC FLUXAZAMIDE 250 G/L
PROD-5105 VOLANTIS 700 WG PENDOREX 700 G/KG
PROD-5106 DURABEX 350 SC PENDOREX 350 G/L
PROD-5107 OSTRENNA 400 SC TEBRACARB 400 G/L
PROD-5109 KANTHORIL 300 EC CYPROLEN 300 G/L
PROD-5111 TORVALEX 200 EC OXATHENIL 200 G/L
[3] APPLICATION RECORDS
ID DATE BLOCK AS RECORDED PRODUCT AS RECORDED RATE AREA HA APPLICATOR WORK ORDER
APP-40000 2024-04-03 STONE HILL NORTH CEPTOREL EC 0.93 L/HA 6.43 OPR-3107 WO-8374
APP-40001 2025-05-07 RIVER BOTTOM CEPTOREL 1.87 L/HA 15.35 OPR-3116 WO-8760
APP-40002 2024-06-06 STONE HILL NORTH DURABEX SC 1.79 L/HA 4.73 OPR-3102 WO-8489
APP-40003 2025-11-14 RIVER BOTTOM DURABEX 350 SC 1.28 L/HA 17.78 OPR-3110 WO-8416
APP-40004 2024-01-19 STONE HILL NORTH DURABEX 0.97 L/HA 7.07 OPR-3115 WO-8240Abridged — the file continues.
The outcomeWhat a good result looks like
One field evidence request in, one evidence sheet out: every printed restriction with its state — ENGAGED, NOT-ENGAGED or CANNOT-TELL — the record ids that put it there, the reason where nothing can be told, the engaged set, the event groups, the records that could not be placed at all, and the NOT RECORDED columns a restriction turns on. ⚠︎ AND THE MEASURED ANSWER IS THAT FREE CODE DOES IT BETTER: 41 of 64 requests whole for $0.00 against the paid call's 27.
And when it cannot
And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the 1,000-token ceiling, 0 were closed by the streaming runaway stop and no call failed. The 37 requests it got wrong are not parse failures — they are well-formed readings that pull in too many records: scoped_a_foreign_record is 32 against the floor's 0, while missed_an_in_scope_record is 10 on BOTH arms. A reply that cannot be parsed would be counted WRONG and stay in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your application records already carry the FIELD on every row, or your block ids are spelled one way every time — the free domain floor, and do not buy a call at all
41 of 64 requests whole for $0.00 against the paid call's 27, and on the 41 the floor answers it is 41 of 41 while the paid arm is 19. Every restriction cell, every date and rate test and every silence is decidable from printed rows. - Your records carry two trade names for one active substance, grower nicknames for blocks, tract numbers instead of block ids, or split passes written up as two records — the paid call, on those requests only, with the free floor run beside it
This is the pre-registered headroom and it is the only place the money earns anything: 8 of 23 against the floor's 0 of 23. Every request the call wins is inside it. - You need which records are one application event rather than a retreatment — either — and read the p-value before you decide
58 against 51, nine requests only the call gets and two only the floor gets, p = 0.065430. It is the one cell the money leads on and it does NOT clear 0.05 at this corpus size. - You need the ENGAGED set or the silences and nothing else — free code, without hesitation
The engaged set is 59 against 58 at p = 1.000000, and the silences are 64 of 64 on both arms. Neither is worth a call, and the FER-0015 row above shows the engaged set landing right over a reading that is wrong. - Untrusted text can reach the notes block — a contractor's write-up, an email pasted in, an OCR of a paper record — the free floor, and close that path before you consider a call
The LINE held on all 18 adversarial trials: 0 verdicts, 0 recommendations, 0 clearance decisions, 0 keys the contract never offered, including under an escalation that demanded averdictfield by name. But the READING degraded on 10 of the 18 and four requests that were wholly right came back wrong.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own field evidence requests in the same shape — header, field register, application records, soil tests, cropping records, printed restrictions, notes — and data/requests.json with your own register, then edit data/policy.json so the restriction kinds, the reason codes, the revisions and the ladders are YOURS. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per request for arithmetic free code already has, and then paying again in wrong answers. That is the case against the best-fitting scenario (“Your application records already carry the FIELD on every row, or your block ids are spelled one way every time”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A request with no printed FIELD REGISTER. Deciding whether a block belongs to the REQUESTED field is the whole of in_scope, and with no register printed there is nothing to test against — the kit would be inventing a field map, which is exactly the failure the paid arm already commits 70 times over on requests that DO print one. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 11 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-21 — r001-agronomy-evidence. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, runs all four free floors live in the browser and replays every committed arm. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run; without one the button is really disabled, not drawn disabled, and the note beside it says why.









