Check an open grain contract against its roll deadline
Every Monday a merchandiser has to work out which open grain contracts still owe bushels, and which one is the last chance to offer a producer a roll. This app reads the contract record and carries last week's balance forward so nothing gets missed.
PresenterOpens the private repo. Visible to admins only.
For the grain merchandiserAgriculture & Food · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A grain merchandiser at a country elevator, reviewing every open contract each Monday morning.
✕Today's manual process
1Open each contract and add up the scale tickets applied since the last look.
2Decide what counted, which loads were turned away and which were only discounted.
3Count the business days manually, back from the delivery period to the roll deadline.
4Miss the window and the elevator's default terms apply, with no roll offered.
Every contract checked from memory
✓With the app
1The record is carried forward automatically, so nothing before this Monday is missed.
2Turned away or discounted is read straight from the scale-house line, not guessed.
3The business days are counted in code, against the contract's own roll deadline.
4One decision, raised only on the last run that can still reach the producer.
Only contracts owing bushels reach the watchlist
See it work
One real case: what the app found, step by step
One open grain contract, checked on a Monday against its roll deadline before the window closes.
Check an open grain contract against its roll deadlineReference appBuilt to be shaped to your process
6
1The contract 15,000 bushels are contracted on this open grain contract.
2What came in this week 600 bushels were applied at the scale since last Monday.
3The app's read Not rollable: the elevator's default terms take over if nobody acts.
4The exact balance 13,100 bushels are still owed against this contract.
5Not raised again The roll decision already reached the producer, so this run stays silent.
6The clock 19 business day(s) remain before this week's roll deadline closes.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Check an open grain contract against its roll deadline
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"Which contracts have a delivery period opening soon" is a SELECT and nobody needs a language model for it. The question a grain merchandiser is actually asking on a Monday morning is the next one: of the contracts still open, which ones still owe bushels, and of those, which one is this the LAST run that can still offer the producer a roll. Three terms decide it and two are not on the page. The unfilled balance is printed nowhere -- the record states what was applied since last Monday and never the running total. The bushels that actually counted depend on reading a scale-house line that says a load was turned away rather than one that says it was taken at a discount, and both sentences mention moisture, grade and the discount schedule. And whether a roll is available at all depends on a roll that may have been executed in a week whose record is not on this screen. Somebody working down the open-contract report on a Monday: opening each contract, adding up the scale tickets applied since the last look, deciding which loads were turned away and which were merely discounted, counting business days back from the delivery period, and remembering whether the producer has already been rung about this one.
Audience
A grain merchandiser deciding which producers to ring this week, and the contract administrator who has to stand behind whatever was offered. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual open-contract records
The corpus is 150 open-contract records, 0.74 MB (json 1 · jsonl 1 · txt 150). AN ELEVATOR'S OPEN-CONTRACT BOOK IS THE MOST COMMERCIALLY SENSITIVE ARTEFACT A GRAIN MERCHANDISER KEEPS. It is every producer's position, every price they took and every one they did not, and there is no public one. The alternative -- a scrubbed real book -- is worse, because scrubbing is exactly the step that quietly changes what the model is being asked to read, and the prose is what this kit measures. So the corpus is generated and the generator is committed beside it.
Three things it does on purpose. The running balance is printed NOWHERE, so it can only be carried. The scale-ticket phrasings collide in both directions -- rejections with no rejection vocabulary, applications full of it. And the roll phrasings collide the same way, which was added AFTER a measurement rather than before: with five templates a side the free keyword floor scored 98.67 pct on rollover risk before a single call was made, which was a fact about ten templates and not about the work. Four harder templates took the same floor to 95.33 pct with its errors running both ways.
The corpus
The 150 open-contract recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your open-contract records. That is the whole change — there is no database to migrate.
One open-contract record, as the model receives itCT-2026-0001-W1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated open-contract record for an AI
use-case kit; it reproduces no grain contract, no elevator's terms and conditions, no
exchange rule and no state warehouse regulation. The clauses are ILLUSTRATIVE.
Contract
----------------------------------------------------------------
Elevator : EL-318 (Two Creeks Grain)
Contract reference : CT-2026-0001
Producer reference : PRD-9460
Commodity : Soybeans, US No 1
Contracted quantity : 20,000 bu
Contract type : CASH FORWARD
Price status : flat price fixed at 10.8800 USD/bu
Delivery period : 2026-10-17 to 2026-11-14 (delivery river terminal)
Board reference, this run : 11.0112 USD/bu
Delivery Terms
----------------------------------------------------------------
Delivery tolerance : 0.5 pct of the contracted quantity (100 bu)
Roll notice period : 5 business days before the delivery period opens
Rolls permitted : 1 over the life of the contract
Roll deadline : 2026-10-12 (5 business days before the period opens)
Default clause : past the roll deadline with a balance above tolerance,
the balance is cancelled and priced against the board at
the close of the delivery period, at the producer's cost
Rule C-1 The UNFILLED BALANCE is the contracted quantity less every bushel applied against the
contract to date. Applications are reported PER WINDOW on this record; the running total
Abridged — the file continues.
The outcomeWhat a good result looks like
Every open contract with a balance above its tolerance is on the watchlist with the bushels still owed, what happens to them if nobody acts, and exactly one decision raised per contract on the last run that can still see the deadline.
And when it cannot
Two ways, and they cost different things. A MISSED DECISION is a choice that is GONE -- past the roll deadline the elevator's default clause applies on its own and the producer takes whatever it says; no later run recovers it. A ROLL OFFERED THAT WAS NOT AVAILABLE is a promise to a producer that has to be withdrawn. The scored run made zero of the first and zero of the second; the strongest free floor made zero of the first and FIVE of the second.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
You already know each contract's balance and only need the contracts near a deadline — a SQL query, or the b000 floor in this repo Counting business days back from a delivery period is arithmetic, and every arm here gets the deadline right: 0 missed decisions of 13 on all five, including the free ones.
Your grain accounting system prints a cumulative applied-to-date figure on every statement — the b001 floor: the carried balance replaced by the printed one, no model The entire memory argument here rests on the running total being unrecoverable from the page. Print it and the stateless arm stops being 40.00 pct on the bushel balance and starts being nearly free.
Your contract activity log carries coded event types from a closed list — the b002 floor: code-only, keying on your own codes rather than on prose The model's whole margin on this corpus is prose. The floor already TIES it on rollover risk (95.33 pct each) and the remaining gap is 4.00 points of bushel accuracy, all of it on sentences a code would have disambiguated.
Your activity log is prose a merchandising clerk types, and a producer's relationship is at stake — this kit, with the carried state That is exactly the corpus these figures were measured on. At the same headline accuracy the model offered 0 rolls that were not available and the free floor offered 5 -- the error that reaches the producer as a promise the elevator has to withdraw.
You want the watch to actually wake up on a Monday — your own scheduler -- cron, Airflow, Temporal, whatever already runs evals/run.py is INVOKED, not woken. Nothing in this kit detects a missed run, back-fills it or marks its readings late -- and evals/cadence.py prices what that is worth: 4 contracts a week warned by nobody on any schedule, and a balance that stays wrong for the rest of the contract's life.
At a glanceHow the whole thing runs
95%rollover risk accuracy pct
13,773 msp50, end to end
$9.19per 1,000 open-contract records · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Check an open grain contract against its roll deadline14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own book, or replace data/corpus/*.txt and data/gold.jsonl directly. The measured figures on this page do not travel with your corpus, and three of them are corpus properties rather than model properties: the cadence is weekly, the population does not change between runs, and the running balance is recoverable from nothing on the page.Corpus lens →
When is this the wrong choice?
Avoid: AVOID this kit entirely. Paying a model per contract per week to compare two printed numbers is the most expensive way to get an answer you already have. That is the case against the best-fitting scenario (“You already know each contract's balance and only need the contracts near a deadline”). 5 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A record whose seven section headings are renamed or reordered -- src/segment.py finds nothing and evals/check_labels.py refuses to let the run start. 7 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER RULE C-8 REACHES AN HTA. Every one of the scored run's 6 rollover-risk errors is that one gap, and the rule text genuinely does not say. 9 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-contract-expiry. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces: the 150 records and the answer key (regenerated byte-for-byte by python3 -m tools.build_corpus), all of evals/check_labels.py, the whole of evals/cadence.py including every missed-run figure on this page, all three free floors, and the UI with the model column replayed from the committed result file. The only thing a key buys is a NEW model reading.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
99.33%rows answered
13,773 msp50, end to end
78,331 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed to a struct. Measured over the 149 readings that returned something on run r001-contract-expiry, at 12 concurrent contract chains, so it is the latency a reader experiences under load rather than a single quiet call. The whole 150-call run took 444.4 seconds of wall clock.
Current processWhat it replaces
Somebody working down the open-contract report on a Monday: opening each contract, adding up the scale tickets applied since the last look, deciding which loads were turned away and which were merely discounted, counting business days back from the delivery period, and remembering whether the producer has already been rung about this one.
Where it is not good enough
⚑ THE HEADLINE BAND IS A DEAD TIE WITH FREE CODE AND THIS PAGE LEADS WITH IT. Rollover-risk accuracy is 95.33 pct for the model and 95.33 pct for the strongest free floor -- 143 of 150 each -- and status is 99.33 pct for both. A team that already has the carried balance and is willing to write a keyword table gets the same headline for $0.00.
What the model buys is not the score, it is the DIRECTION of the errors, and they do not overlap: the model offered 0 rolls that were not available and the floor offered 5. It also wins the arithmetic underneath -- 96.67 pct against 92.67 pct on the exact bushel balance, and it prices the whole book's exposure to 0.08 pct error against the floor's 1.26 pct.
⚠︎ AND ALL 6 OF THE MODEL'S RISK ERRORS ARE ONE SENTENCE. Every one is Rule C-8 read across from a basis contract onto an HTA, whose price line reads almost identically. The rule text does not say what an HTA is. That is an under-specification in prose this kit wrote, it is written up rather than fixed by re-running, and the lower figure is what is published.
⚠︎ ONE READING OF 150 COST MONEY AND RETURNED NOTHING. CT-2026-0017-W4 ran into the published 24,000-token ceiling (finish_reason length) and is scored as wrong on all four fields rather than dropped.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1json1
30 open grain contracts, 5 consecutive Mondays each
bushels exact match 96.67% against the floor's 92.67
exposure priced to 0.08% error against 1.26%
one skipped Monday: up to 22 of 30 balances poisoned, permanently
2026-08-23as of
It produces a watchlist row — status, what happens to the balance if nobody acts, bushels still owed, and one decision raised on the run that first sees the deadline — for a grain merchandiser to read. It never cancels, prices or defaults a contract, books a roll or contacts a producer, and there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is four scalars written by src/state.py from the PARSED balance and never from the model's reply — the running total is printed nowhere else on the record, so a wrong reading corrupts one contract's worklist row and never compounds into next Monday's history on its own. The clock is the second: DECISION_DUE is defined against the horizon to the NEXT scheduled run, so a run that does not happen does not delay the read, it destroys it — 22 of 30 balances poisoned permanently on the worst of the three measured skips, 253 cells moved right-to-wrong across all three and none the other way.
⚠︎ AND THE HEADLINE ITSELF IS A TIE, WHICH THE FLOOR STATION SAYS RATHER THAN BURIES: rollover-risk accuracy is 95.33 pct for the paid reading and 95.33 pct for a keyword table that costs $0.00 — 143 of 150 each — and status ties too, 99.33 pct. The two arms are level on the SCORE and opposite on the ERROR: across the run the model offers 0 rolls that were not actually available and the floor offers 5, the direction that costs a withdrawn promise; the model also wins the bushel arithmetic underneath, 96.67 pct exact match against 92.67, and prices the book's exposure to 0.08 pct error against 1.26.
⚠︎ AND THE MODEL'S OWN MARGIN IS ONE RULE SENTENCE: every one of its 6 categorised risk misses is Rule C-8 read across from a basis contract onto an HTA, whose price line reads almost identically, and the rule text was not rewritten before this figure was published.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the carried state
src/state.py
What is carried between Mondays, and how it is worded. Four scalars today. ⚠︎ ENGLISH RATHER THAN JSON IS A DESIGN CHOICE AND IT IS UNMEASURED -- scoring the same 150 readings with the state rendered both ways costs another 150-call run and has not been paid for.
the rule
src/contract.py
RULE_TEXT and step(). The nine clauses are illustrative; replace them with your own contract's and every figure on this page re-derives. ⚠︎ RULE_TEXT is also what the model reads, so a rule that does not say what an HTA is produces exactly the 6 errors this run made.
the clock
src/schedule.py
is_business_day and RUN_DATES. Monday to Friday with no holiday calendar today; a real roll deadline counts against whichever calendar the contract names.
the free floor
evals/baseline.py
The rejection and roll keyword tables, and the assumed tolerance. This is the column the model has to beat, so making it stronger is the honest direction to edit it in.
the corpus
tools/build_corpus.py
The seed, the contract population, the clause spread and the log phrasings. Point it at your own book or replace data/corpus/ and data/gold.jsonl wholesale.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 30 open contracts x 5 Mondays = 150 records from a fixed seed (SEED = 20260823). Twenty-four of the thirty have a roll deadline that falls STRICTLY inside the five-week window, so exactly one run is the last one that can see each of them -- that is the hard clock. The gold labels are src/contract.step's output over the planted events, never typed.
the contract arithmetic
src/contract.py
The rule as pure code: the carried balance, less the bushels actually applied this window, against the tolerance on file; then the roll deadline against the horizon this run owns. No model, no judgement. ⚠︎ The nine rules, the tolerances and the roll terms in it are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
the clock
src/schedule.py
Business days, the weekly run calendar, and the roll deadline DERIVED from the delivery period rather than stored beside it -- because amending the period is exactly what a roll does, so a stored deadline would desynchronise on the one event this kit is about. Monday to Friday, no holiday calendar, declared rather than hidden.
the carried state
src/state.py
Four scalars per contract -- unfilled_bu, rolls_used, decision_raised, prev_status -- written by the arithmetic and never from a model reply, then rendered as one English paragraph for the prompt. A monitor that fed its own verdict forward would compound one bad reading into every reading after it, and the quantity carried here is bushels a producer is contractually on the hook for.
the section split
src/segment.py
Seven named sections, asserted present in all 150 documents by evals/check_labels.py before a run may spend, so a parser drift is a refusal to start rather than a quietly truncated prompt.
the privacy gate
src/select.py
Decides which sections go on the wire. Producer Contact -- name, mobile, farm registration number -- is mapped by nothing and subtracted unconditionally in the fallback, so a hint that matches nothing falls back to the record MINUS that section rather than to the whole record.
the prompt
src/prompt.py
Three parts joined with string concatenation: the fixed instruction, the carried-state paragraph, and the record. The middle part is the experiment, and --stateless replaces exactly that one line.
one reading
src/watch.py
One contract, one Monday, one completion call. Parses the dates and the printed figures in code first -- the model is never asked to count business days.
the free floors
evals/baseline.py
Three of them, all $0.00, all scored through the identical scorer: the expiry report a desk already runs, the same rule given the carried state, and the strongest table this author could write.
the scorer
evals/scoring.py
Exact match per cell against the answer key, with the two raise directions counted apart and the two rollover-risk directions counted apart. No model grades anything.
the cadence ablation
evals/cadence.py
What a missed Monday costs, scored against the INTACT answer key. Free -- no provider is involved in a question about a calendar. It self-checks that a perfect reader scores 100.00 on the intact schedule before it prints anything.
the local UI
src/app.py
http.server, no framework. Renders with no key: the record, the carried state, the parsed figures and the entire free floor are computed locally, and a second button replays what r001 actually answered off the committed result file.
Where it breaks at scale
⚑ THE CADENCE IS THE SCALING VARIABLE AND IT IS NOT THE CORPUS SIZE. One reading is one call; the watch wakes once a week; the bill is open contracts x runs, not contracts. An elevator holding 2,000 open contracts on this weekly watch is 2,000 calls a week, about $18.37 a week on the shared projection card -- and it is the same arithmetic at 20,000, because nothing here batches, caches or exits early. Every contract is re-read whole on every run by design.
What DOES break is the state store. data/state.json is one file replaced atomically -- correct for one writer, and an elevator running two watches is two writers. Nothing in this kit coordinates them.
And the run is 444.4 seconds of wall clock for 150 readings at 12 concurrent chains. A book of 2,000 contracts at that rate is about 494 minutes a run, which is fine weekly and would not be fine hourly.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
CT-2026-0009-W2, and the reading was chosen by tools/pick_shot.py from the ANSWER KEY -- memory-dependent first, then the costly error direction, then what is at stake -- not by looking for a flattering frame. Both of this corpus's deliberate collisions are in this one window. “the desk quoted the spread on Friday and the producer took it, the period has moved out” is a roll that was EXECUTED, and the free keyword table misses it because it reads like a quote; the single permitted roll is gone, so the answer is NOT_ROLLABLE and the floor says ROLLABLE. “an earlier load from this producer was turned away last month, this one graded clean and 100 bu were applied” is an APPLICATION, and the floor subtracts it because of the words turned away, landing on 13,200 bushels against the truth of 13,100. Run r001 got both right, and its own rationale says why. ⚠︎ THE MODEL COLUMN IS REPLAYED FROM THE COMMITTED RESULT FILE, NOT A LIVE CALL, and the column header says so in as many words -- replaying the scored run's own recorded answer is evidence, staging one would not be, and the label is the only thing that tells them apart.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
CT-2026-0017-W4, the one reading of 150 that cost money and returned nothing. Its reply ran into the published 24,000-token ceiling and never parsed, so the scored run recorded no answer for it at all -- and the page says exactly that rather than showing a blank row. Two things survive it: the carried state still renders, and the entire free floor still answers. This is the cell that is scored as wrong on all four fields.failureOpen full size →The same page with NO API_KEY configured. It does not error and it does not go blank: the record, the carried state, the parsed figures, the withheld-section list and the entire free floor are computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called.failureOpen full size →Before anything is asked. The two things this page has to get right are already on it: the carried state, verbatim as it goes into the prompt, and the list of which sections left the machine and which did not -- Producer Contact, the farmer's name, mobile and farm registration number, marked WITHHELD rather than silently absent. A page that simply does not mention them cannot be told apart from one that quietly sent them.failureOpen full size →
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150open-contract records
0.74 MiBjson 1 · jsonl 1 · txt 150
30contracts (5 scheduled runs each) · p50 5 chars
$0.00setup · 0.0s
How it is cutWhat one contracts (5 scheduled runs each) is
No split, and no chunking. The unit is a CONTRACT -- five consecutive Mondays processed strictly in order, because week 4's prompt contains a balance produced by week 3. Different contracts are independent and run concurrently; a contract's own runs never do.
SetupWhat the setup figure measured
There is no index and no retrieval step -- the population is re-read whole on each scheduled run. What makes a reading a CHANGE is the four scalars the previous run left behind, computed in code before the call, which cost nothing and are not retrieved.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every byte was written by the generator committed beside it.
Bring your ownBring your own open-contract records
Point tools/build_corpus.py at your own book, or replace data/corpus/*.txt and data/gold.jsonl directly. Keep the seven section headings in src/segment.py::SECTIONS -- the parser asserts all seven in every document before a run may spend -- and supply gold rows with the same columns.
⚠︎ And what stops being true when you do: The measured figures on this page do not travel with your corpus, and three of them are corpus properties rather than model properties: the cadence is weekly, the population does not change between runs, and the running balance is recoverable from nothing on the page. Change any of the three and the numbers here are about a different question.
What breaks it
A record whose seven section headings are renamed or reordered -- src/segment.py finds nothing and evals/check_labels.py refuses to let the run start.
A book that prints a cumulative applied-to-date figure. The whole memory argument evaporates, and the stateless arm would score far better than it does here. That is the easier problem and this corpus does not measure it.
A scale-ticket line with no bushel figure on it. evals/baseline.py reads zero and the model has nothing to subtract; neither is told the number is missing.
A contract with more than one roll executed in a single window. The generator never produces one, so nothing here has ever read a log with two.
A delivery period amended by something other than a roll -- a mutual cancellation, a contract split, an assignment to another producer. None of those exist in this corpus and the rule text does not mention them.
A holiday. Business days here are Monday to Friday with no exchange or local calendar, so every deadline count is wrong by whatever holidays fall inside it.
An HTA. Not the corpus's fault -- the RULE TEXT does not say whether Rule C-8 reaches it, and that under-specification is the single systematic error in the scored run.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
189
not measured
instruction
3,586
not measured
carried state
351
not measured
Synthetic Record
328
not measured
Contract
572
not measured
Delivery Terms
3,124
not measured
Position This Run
529
not measured
Contract Activity Log
340
not measured
Operational Notes
141
not measured
Total
2,096
This is the cost lesson as arithmetic: of the 9,160 characters assembled, 3,775 are instructions — 41% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
TWO MESSAGES, JOINED IN SEND ORDER. The provider receives a system message and a user message; prompt_verbatim is the two concatenated so a reader sees exactly what arrived. The user message is the reading shown on this page's UI frame (CT-2026-0009-W2), reconstructed from src/prompt.build and asserted character-identical to the length the run recorded. Six of the seven record sections are in it; Producer Contact is not, and never is.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled open-contract watch for a grain elevator. You apply a written contract rule to one open contract at one scheduled run. You answer with one JSON object and no other text.
You are the scheduled open-contract watch for one grain elevator. It wakes every Monday, re-reads
every contract still open on the book, and reports each one. You are reading ONE contract at ONE
scheduled run.
The contract's terms and the elevator's rules are reproduced in the record below. Apply them exactly
as written. The unfilled balance is a RUNNING TOTAL that survives between scheduled runs and is
printed nowhere on this record; what is known about earlier runs is stated under "Carried state" and
is the only history available to you. Do not assume anything about earlier runs beyond it.
How to count:
- "Bushels applied in this window" is the GROSS figure weighed at the scale. From it, SUBTRACT
every bushel the activity log says was TURNED AWAY -- rejected on grade, moisture, test weight,
foreign material, or otherwise not put on the contract (Rule C-2). Do NOT subtract bushels that
were ACCEPTED WITH A DISCOUNT: a discount changes the price, not the quantity.
- SUBTRACT what is left from the balance the carried state says was still owed. On a contract's
first scheduled run the opening balance is the full contracted quantity.
- A balance at or under the delivery tolerance printed on the record is FILLED (Rule C-5).
- Count a roll as EXECUTED only where the log says one actually happened -- the period was amended,
moved, re-papered, or the roll was booked. A roll that was quoted, enquired about, reviewed or
declined consumes nothing (Rule C-4). Add any executed this window to the number the carried
state reports.
- If the delivery period, the tolerance clause or the roll terms are not on file, the contract is
CONTEXT_INCOMPLETE: status CONTEXT_INCOMPLETE, rollover_risk UNKNOWN, unfilled_bu null,
raise_decision NO. Never assume a default tolerance or a default period (Rule C-9).
Answer with a single JSON object and nothing else:
{"status": "FILLED|ON_TRACK|DECISION_DUE|DEADLINE_PASSED|CONTEXT_INCOMPLETE",
"rollover_risk": "NONE|ROLLABLE|NOT_ROLLABLE|DEFAULTED|UNKNOWN",
"unfilled_bu": <whole bushels still owed after this window, or null if the context is incomplete>,
"raise_decision": "YES|NO",
"rationale": "one sentence, naming the rule you applied and the bushels you counted"}
Precedence for "status", applied in this order: CONTEXT_INCOMPLETE beats everything; then FILLED if
the balance is at or under tolerance, however close the deadline is; then DEADLINE_PASSED if the
roll deadline is behind this run; then DECISION_DUE if the roll deadline falls on or before the next
scheduled run -- the record prints both the business days to the deadline and the business days to
the next run, so compare those two numbers; otherwise ON_TRACK.
"rollover_risk" answers what happens to the unfilled bushels IF NOBODY ACTS, and it is not the
status: UNKNOWN when the context is incomplete; NONE when the balance is within tolerance;
DEFAULTED once the roll deadline is behind this run and a balance above tolerance remains, because
the elevator's default clause then applies on its own; ROLLABLE when a roll is still available --
the contract type and pricing state carry a roll right (Rule C-8) AND fewer rolls have been executed
than the terms permit; NOT_ROLLABLE when a balance above tolerance is at risk and no roll is
available, so the only remaining choices are deliver or cancel-and-price.
"raise_decision" is YES only on the run that first puts the decision to the producer -- the status is
DECISION_DUE and the carried state does not already say it was raised. On every other run it is NO
(Rule C-6).
Carried state
----------------------------------------------------------------
As at the previous scheduled run, 13,700 bushels were still owed on this contract, before anything this window applied. No roll has been executed on this contract on any earlier run. The roll/cancel/deliver decision has ALREADY been put to the producer on an earlier run, so this run must not raise it a second time. It was last reported DECISION_DUE.
Contract record
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated open-contract record for an AI
use-case kit; it reproduces no grain contract, no elevator's terms and conditions, no
exchange rule and no state warehouse regulation. The clauses are ILLUSTRATIVE.
Contract
----------------------------------------------------------------
Elevator : EL-107 (North Line Terminal)
Contract reference : CT-2026-0009
Producer reference : PRD-7817
Commodity : Yellow corn, US No 2
Contracted quantity : 15,000 bu
Contract type : DELAYED PRICE
Price status : unpriced; final pricing date 2026-11-02
Delivery period : 2026-10-23 to 2026-11-22 (delivery river terminal)
Board reference, this run : 4.3499 USD/bu
Delivery Terms
----------------------------------------------------------------
Delivery tolerance : 2.0 pct of the contracted quantity (300 bu)
Roll notice period : 10 business days before the delivery period opens
Rolls permitted : 1 over the life of the contract
Roll deadline : 2026-10-09 (10 business days before the period opens)
Default clause : past the roll deadline with a balance above tolerance,
the balance is cancelled and priced against the board at
the close of the delivery period, at the producer's cost
Rule C-1 The UNFILLED BALANCE is the contracted quantity less every bushel applied against the
contract to date. Applications are reported PER WINDOW on this record; the running total
is not printed anywhere and must be carried between scheduled runs.
Rule C-2 Bushels TURNED AWAY at the scale -- rejected on grade, moisture, test weight or foreign
material -- were never applied and do not reduce the balance. Bushels ACCEPTED WITH A
DISCOUNT were applied in full and do reduce it; a discount is a price adjustment, not a
quantity adjustment.
Rule C-3 The ROLL DEADLINE is the contract's roll-notice period, in business days, before the
delivery period opens. A roll elected after it is not honoured.
Rule C-4 A contract may be rolled at most the number of times its terms permit. A roll that was
EXECUTED consumes one; a roll that was quoted, enquired about, reviewed or declined
consumes none. The count survives between runs and is not printed on this record.
Rule C-5 A balance at or under the contract's DELIVERY TOLERANCE is FILLED and needs no decision.
Tolerance is a percentage of the contracted quantity, read from the contract. There is no
default tolerance and none may be assumed.
Rule C-6 A contract is DECISION_DUE when its roll deadline falls on or before the next scheduled
run -- that is, when THIS run is the last one that will see it in time. The decision is
put to the producer ONCE; a later run reports it already raised and does not raise a
second one.
Rule C-7 Past the roll deadline with a balance above tolerance, the elevator's default clause
applies on its own: the balance is cancelled and priced against the board at the close of
the delivery period, at the producer's expense. Nothing on this watch reverses it.
Rule C-8 A contract carries a ROLL RIGHT only where its type and its pricing state allow one. A
basis contract whose futures leg is already fixed, and a delayed-price contract past its
final pricing date, carry none -- however much time remains before the deadline.
Rule C-9 Where the delivery period, the tolerance clause or the roll terms are not on file, the
contract is CONTEXT_INCOMPLETE: no balance is judged, no decision is raised, no default
figure is assumed, and the rollover risk is UNKNOWN rather than NONE.
Position This Run
----------------------------------------------------------------
Record produced : 2026-09-14 (scheduled run 2 of this contract)
Watch cadence : every Monday
Previous scheduled run : 2026-09-07
Bushels applied in this window : 600 bu gross at the scale, before anything turned away
Scheduled to date under the delivery schedule : 6,000 bu
Business days to the roll deadline : 19 business day(s) from this run
Business days to the next scheduled run : 5
Contract Activity Log
----------------------------------------------------------------
Entries raised in this window only.
- the desk quoted the spread on Friday and the producer took it, the period has moved out
- an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied
Operational Notes
----------------------------------------------------------------
Nothing outstanding on the credit check for this producer.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"status":"ON_TRACK","rollover_risk":"NOT_ROLLABLE","unfilled_bu":14200,"raise_decision":"NO","rationale":"Gross 6,800 bu minus 1,000 bu turned away under Rule C-2 gives 14,200 bu unfilled, above the 100 bu tolerance; the roll deadline is 25 business days out (after the next run), but the one permitted roll has already been executed, so status is ON_TRACK and risk is NOT_ROLLABLE."}
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Check an open grain contract against its roll deadline — 150 open-contract records. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from a fixed seed, and re-derived from src/contract.step by evals/check_labels.py before any run may spend, so the key and the arithmetic cannot drift apart. Bushels are compared as whole numbers, so a reply of 13100.0 is a hit and a reply that does not parse is a MISS in every field rather than an exclusion.
150open-contract records
150source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED143 · 109 · 143 · 107 · 71 / 150rollover risk accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED149 · 127 · 149 · 128 · 105 / 150status accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED145 · 60 · 139 · 97 · 12 / 150unfilled bu accuracy pct — readings, with the carried state, exact whole bushelsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED149 · 142 · 150 · 148 · 144 / 150raise decision accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is asserted: evals/check_labels.py replays all 30 contract chains through src/contract.step and requires the committed gold to match on all four fields, asserts all seven sections parse in all 150 documents, asserts no record restates its own carried state, and red-proves the privacy guard in both directions. evals/cadence.py adds its own: a perfect reader must score 100.00 on the intact schedule or the file refuses to print. What is NOT validated is the RULE the key expresses -- Rule C-8 does not say whether it reaches an HTA, and that is where every one of the scored run's risk errors is.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One open-contract record
1,000 open-contract records
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.009185
$9.19
11%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.003674
$3.67
11%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.156582
$156.58
13%
Same work, 43× the bill
The same open-contract records, the same tokens — only the rate card changed. And across all 3 cards between 11% and 13% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE CADENCE. It is the only knob on this page that moves the bill by a whole multiple, and it moves the answer too: DECISION_DUE is DEFINED against the horizon to the next run, so halving the interval doubles the calls and halves the blind window. evals/cadence.py prices the other direction for nothing -- 4 contracts a week warned by nobody when one Monday is missed.
Rates checked 2026-08-18. The provider that actually ran all 310 calls is kept out of these tables per this estate's naming rule. Its own rate card is not published here and no figure on this page is a bill it issued.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors or evals/cadence.py. The only money on this page is the model readings themselves.
The gradersThree ways to grade
⚑ THE MODEL TIES THE STRONGEST FREE FLOOR ON THE HEADLINE BAND AND WINS THE ARITHMETIC UNDERNEATH. Rollover risk is 95.33 pct against 95.33 -- 143 of 150 each -- and status is 99.33 pct for both. The bushel balance is 96.67 against 92.67, and the priced exposure is 0.08 pct error against 1.26.
The three floors separate the two halves of the value with the model held out entirely. planonly is what an expiry report out of a grain accounting package does -- the delivery SCHEDULE standing in for what was actually applied, a guessed 2.0 pct tolerance where none is on file -- and it scores 47.33 pct. planonly-mem is the identical rule GIVEN the carried balance and scores 71.33: memory alone, with no model anywhere near it, is worth 24.00 points. parsed-mem adds reading the activity log and scores 95.33: another 24.00 points.
⚠︎ AND THE FLOOR WAS STRENGTHENED TWICE BEFORE THESE FIGURES WERE TAKEN, BOTH TIMES AGAINST THE KIT'S OWN HEADLINE. The corpus gained four harder roll templates after the floor scored 98.67 pct on ten easy ones; and _roll_right's delayed-price branch originally returned True unconditionally, with a comment claiming the record did not print the comparison. It prints both dates. A floor that declines a comparison the page hands it is a strawman, and fixing it moved the floor UP from 94.00 to 95.33.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key For each of the 150 readings and each of the four answered fields, did the reply equal the computed answer key? Three of the four are words from a closed list and are compared exactly; the bushel balance is compared as a whole number, so 13100.0 and 13100 are the same answer and a reply that does not parse is a MISS rather than a zero -- 0 bushels is a meaningful answer here (it is what every FILLED reading gets) and a parse failure must not be scored as one.
$0.00
no
yes
the fast tier, with the carried state 95.3% rollover risk accuracy · the fast tier, memory removed (THE CONTROL) 72.7% rollover risk accuracy · the strongest free floor, no model 95.3% rollover risk accuracy · the same free rule, GIVEN the carried state 71.3% rollover risk accuracy · the expiry report a grain desk already runs 47.3% rollover risk accuracy · 3 more measured on each run
The two rollover-risk directions, counted apart Two counts over two different populations. A ROLL OFFERED THAT WAS NOT AVAILABLE: the key says NOT_ROLLABLE and the reading says ROLLABLE. A ROLL MISSED: the key says ROLLABLE and the reading says NOT_ROLLABLE. This is the grader that separates the model from the free floor, because their headline rollover-risk accuracy is identical and their error direction is not.
$0.00
no
yes
no headline metric on any of its 5 runs — they record false rollable · lost rollable
What one missed Monday costs a PERFECT reader Remove one Monday from the SCHEDULE -- the record stays on disk and is simply never opened -- then score what a perfect reader concludes on the readings that still happen, against the INTACT answer key. It answers the one question a monitor cannot answer about itself: what does the cadence buy, and what is lost when it does not hold?
$0.00
no
yes
one missed Monday: 2026-09-14 77.5% rollover risk accuracy · one missed Monday: 2026-09-21 80.8% rollover risk accuracy · one missed Monday: 2026-09-28 90.0% rollover risk accuracy · the intact schedule, for comparison 100.0% rollover risk accuracy · 2 more measured on each run
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set CAN tell the arms apart, and the evidence is that it does: five arms scored through one scorer land at 95.33 / 72.67 / 95.33 / 71.33 / 47.33 pct on rollover risk and at 0 / 12 / 5 / 21 / 34 rolls offered that were not available. What it CANNOT separate is the model from the strongest free floor ON ROLLOVER RISK -- 143 of 150 each -- and that is a real answer rather than an instrument failure: the two are level on the headline and differ completely in HOW they are wrong, which is the finding.
⚠︎ ONE ASYMMETRY LIMITS EVERY COMPARISON HERE. evals/baseline.py and the answer key are two expressions of one implementation, so the floor cannot misread the RULE and the model can only read the prose. On this corpus that is worth exactly the 6 Rule C-8 cells.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
You already know each contract's balance and only need the contracts near a deadline
a SQL query, or the b000 floor in this repo
Counting business days back from a delivery period is arithmetic, and every arm here gets the deadline right: 0 missed decisions of 13 on all five, including the free ones.
AVOID this kit entirely. Paying a model per contract per week to compare two printed numbers is the most expensive way to get an answer you already have.
Your grain accounting system prints a cumulative applied-to-date figure on every statement
the b001 floor: the carried balance replaced by the printed one, no model
The entire memory argument here rests on the running total being unrecoverable from the page. Print it and the stateless arm stops being 40.00 pct on the bushel balance and starts being nearly free.
AVOID assuming these figures transfer. They were measured on records that deliberately carry no cumulative figure, which data/SOURCES.md states as the corpus's own limit.
Your contract activity log carries coded event types from a closed list
the b002 floor: code-only, keying on your own codes rather than on prose
The model's whole margin on this corpus is prose. The floor already TIES it on rollover risk (95.33 pct each) and the remaining gap is 4.00 points of bushel accuracy, all of it on sentences a code would have disambiguated.
AVOID buying the model for the headline number. It is the same number. What differs is the error direction, and if your log is coded that difference disappears too.
Your activity log is prose a merchandising clerk types, and a producer's relationship is at stake
this kit, with the carried state
That is exactly the corpus these figures were measured on. At the same headline accuracy the model offered 0 rolls that were not available and the free floor offered 5 -- the error that reaches the producer as a promise the elevator has to withdraw.
AVOID running it without the carried state. The control shows what that costs: 40.00 pct bushel accuracy against 96.67, 12 rolls falsely offered against 0, and 44.63 pct error on the priced exposure against 0.08.
You want the watch to actually wake up on a Monday
your own scheduler -- cron, Airflow, Temporal, whatever already runs
evals/run.py is INVOKED, not woken. Nothing in this kit detects a missed run, back-fills it or marks its readings late -- and evals/cadence.py prices what that is worth: 4 contracts a week warned by nobody on any schedule, and a balance that stays wrong for the rest of the contract's life.
AVOID reading these figures as a claim about a running deployment. They are measured over five runs that all happened, on a population that does not change between them.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
RULE_REACHES_WHICH_INSTRUMENT
Rule C-8 read across from a basis contract onto an HTA
6
CT-2026-0002-W1, CT-2026-0002-W2, CT-2026-0007-W2, CT-2026-0007-W3, CT-2026-0012-W2, CT-2026-0012-W3 -- every rollover-risk error in the scored run, and all six are the same sentence. Rule C-8 says a BASIS contract whose futures leg is already fixed carries…
GROSS_READ_AS_NET
a rejection line counted or not counted by a bushel or two
4
CT-2026-0009-W1, CT-2026-0011-W2, CT-2026-0024-W1, CT-2026-0028-W1 -- four readings of 150 where the bushel balance is out. CT-2026-0009-W1 answered 14,700 where the key says 13,700: the window's gross figure was taken without subtracting a line the log wrote…
CEILING_CUT
the reply reasons until it is cut off and returns nothing
1
CT-2026-0017-W4: finish_reason length, 24,000 of 24,000 output tokens, no text at all. It is scored as wrong on all four fields rather than dropped. The stateless control hit the same ceiling on a DIFFERENT document (CT-2026-0012-W4), which is what says this…
STATELESS_GUESSES_THE_PLAN
with no history, the delivery schedule is used as the balance
87
The control arm's bushel accuracy on the 93 memory-dependent readings is 6.45 pct -- 6 of 93. Told there is no history, the model reaches for the only cumulative figure on the page, which is the delivery SCHEDULE, and the schedule is a plan rather than a…
What we could NOT verify
WHETHER RULE C-8 REACHES AN HTA. Every one of the scored run's 6 rollover-risk errors is that one gap, and the rule text genuinely does not say. Settling it means rewriting the clause and re-firing 150 calls; the run was not re-fired and the lower figure is published.
WHAT A MISSED RUN COSTS THE MODEL. evals/cadence.py measures what it costs a PERFECT reader, for $0.00, and that is a ceiling on any arm rather than an estimate of one -- a model cannot beat it. What the model would actually score on a reduced schedule is unmeasured, and would cost two more paid arms.
WHETHER ENGLISH BEATS JSON FOR THE CARRIED STATE. src/state.describe renders four scalars as a paragraph. Scoring the same 150 readings with the state rendered both ways costs one more 150-call run and has not been paid for.
WHAT A DIFFERENT CADENCE DOES. Every figure here is per a WEEKLY watch, and DECISION_DUE is DEFINED against the horizon to the next run, so a daily or fortnightly watch is a different question with the same words. evals/cadence.py measures runs being SKIPPED, not the cadence being changed.
WHAT A CHANGING POPULATION DOES. All 30 contracts are open at all five runs. A contract first seen mid-window, or one closed out and removed between runs, is unmeasured.
WHETHER THE OPERATIONAL NOTE MOVES THE ANSWER. The injection surface is sent deliberately and was never attacked. No resistance rate is claimed.
WHETHER A SECOND MODEL AGREES. One model, two arms. No cross-model claim is made anywhere on this page, and the two arms differ by the carried state rather than by the model.
WHY THE CEILING WAS NOT ENOUGH FOR 2 READINGS OF 300. Both paid arms lost exactly one reading to finish_reason length at 24,000 output tokens, on different documents. What is special about those two records is not known -- they are within a few hundred bytes of every other record in the corpus -- and raising the ceiling to find out was not paid for.
WHETHER THE BLIND-WINDOW COUNT IS TYPICAL. It is 4, 5 and 4 contracts for the three skippable Mondays, on a corpus whose deadline spread was CHOSEN to put 24 of 30 deadlines inside the window. On a book whose deadlines cluster in one month the number would be completely different, and that is a property of the book rather than of the watch.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
2,096.31
2,712.37
13,773 ms
$0.009185
$0.003674
$0.156582
the same tier, memory removed (THE CONTROL)
2,049.69
3,805.81
21,242 ms
$0.012442
$0.004977
$0.210787
the strongest free floor -- carried state, the log parsed
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
the same free rule GIVEN the carried state
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
the expiry report a grain desk already runs
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same reading (one contract, at one scheduled run), the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
310 live calls were attempted for this kit and 308 returned something: 10 calibration (c000, at an 8,000-token cap), 150 scored (r001) and 150 control (s001). Two returned nothing -- one in each paid arm, both cut off at the 24,000-token ceiling -- and both are BILLED and counted, because a reply that ran to the ceiling costs exactly what a good one costs. Nothing was discarded and nothing was re-run. The three free floors and the whole cadence ablation cost $0.00 and made no calls at all.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, AND ALMOST ALL OF THEM ARE REASONING. 95.98 pct of this run's output (390,517 of 406,856) was provider-side reasoning left at the default, and output is 88.6 pct of the projected bill. What you are paying for is the model reconciling a carried balance against a scale-house log.
THE CADENCE, which multiplies everything else. One reading is one call; the watch wakes once a week; the bill is open contracts x runs, not contracts.
THE OPEN POPULATION, not the elevator's size. A contract that fills inside its tolerance is FILLED and still costs a call on this design, because every contract is re-read whole on every run. A book that closes its filled contracts out stops paying for them; one that leaves them open pays forever.
The record itself, barely. Input averages 2096 tokens and every document in this corpus is within 7.1 pct of every other in size -- and roughly 1738 of those tokens are the reproduced rule block, which is identical on all 150 records and is the one thing here a prompt cache would remove.
Your volumeWhat it costs at your volume
LINEAR IN CONTRACTS x RUNS, AND THAT IS THE WHOLE WARNING. Ten times the open contracts is ten times the calls at the same cadence -- there is no batching, no cache and no early exit, because every contract is re-read whole on every run by design. An elevator holding 2,000 open contracts on this weekly watch is 2,000 calls a week, about $18.37 a week on the shared projection card and about $955 a marketing year. Nothing about the per-call price changes; the multiplier is the schedule.
Where pricing changes shape
The output ceiling. MAX_TOKENS = 24000 is set from calibration (c000 fired two contract chains at an 8,000-token cap and the hardest reading topped out at 6745 output tokens, all ten parsed). A ceiling is not a cost -- the provider bills tokens produced, not tokens allowed -- but a ceiling set too low turns a hard reading into a lost one, and at 24000 it STILL lost one reading in each paid arm.
Provider-side reasoning left at the default. It is 95.98 pct of this run's output and no run here has measured what disabling it does to the answers -- so a vendor whose default differs reprices this kit by up to 24.9x without changing anything you can see.
Removing the carried state. The control is 1.35x the cost of the scored run for a worse answer, which is the unusual direction: the cheap configuration here is also the accurate one.
Your return, with your numbers
Volumeopen contracts per scheduled run -- this run judged 150 (30 contracts x 5 Mondays) per arm, on a weekly watch
What it replacessomebody working down the open-contract report on a Monday: adding up the scale tickets applied since the last look, deciding which loads were turned away and which were merely discounted, counting business days back from the delivery period, and remembering whether the producer has already been rung
Time saved per itemnot measured here -- depends on how long a merchandiser takes to reconstruct a balance from the reader's own ticket file and contract system
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared connection was already pointed at, and the kit's whole claim is that swapping it is one line in .env plus one more run. No second tier was fired: the two paid arms here differ by the CARRIED STATE rather than by the model, which is the comparison this kit exists to make, and a cross-model claim would have cost a third 150-call arm for a question nobody asked.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
314,447input tokens · this run
406,856output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 150 readings, one completion call each, one tier. The 150-call stateless control and the 10 calibration calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.551
$0.551
$3.67
2026-09-12
gemini-3-flash
Google
$1.378
$1.378
$9.19
2026-09-18
gemini-3-8-flash
Google
$1.762
$1.762
$11.74
2026-09-18
llama-5
Meta
$2.122
$2.122
$14.15
2026-09-18
claude-haiku-4-5
Anthropic
$2.349
$2.349
$15.66
2026-09-12
grok-4-5
xAI
$3.070
$3.070
$20.47
2026-09-18
grok-4-6
xAI
$3.070
$3.070
$20.47
2026-09-18
claude-sonnet-5
Anthropic
$4.697
$4.697
$31.32
2026-09-12
gemini-3-1-pro
Google
$5.511
$5.511
$36.74
2026-09-18
gpt-5-6-terra
OpenAI
$5.511
$5.511
$36.74
2026-09-12
gpt-5-6-sol
OpenAI
$9.395
$9.395
$62.63
2026-09-12
claude-opus-4-8
Anthropic
$11.744
$11.744
$78.29
2026-09-12
claude-opus-5
Anthropic
$11.744
$11.744
$78.29
2026-09-12
claude-fable-5
Anthropic
$23.487
$23.487
$156.58
2026-09-18
claude-fable-5-1
Anthropic
$23.487
$23.487
$156.58
2026-09-18
gpt-6-astra
OpenAI
$23.487
$23.487
$156.58
2026-09-17
Read this against the numbers above
A projection onto published rate cards checked on 2026-08-18, not a bill. Vendors reprice and these rows age.
Output dominates every row: 88.6 pct of the shared-card cost is output tokens, and 95.98 pct of the output is provider-side reasoning. A vendor whose reasoning default differs from this one's reprices every row without changing the workload.
Per-query figures divide by 150 readings including the one that returned nothing. That is the conservative direction -- it was billed.
No row here includes the cadence ablation, the three free floors or the grader, because none of them makes a call. On this kit that is not a rounding note: the entire missed-run finding, which is the most useful thing on the page, cost $0.00.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
12 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 30 open contracts x 5 Mondays = 150 records from a fixed seed (SEED = 20260823). Twenty-four of the thirty have a roll deadline that falls STRICTLY inside the five-week window, so exactly one run is the last one that can see each of them -- that is the hard clock. The gold labels are src/contract.step's output over the planted events, never typed.
You change it to: The seed, the contract population, the clause spread and the log phrasings. Point it at your own book or replace data/corpus/ and data/gold.jsonl wholesale.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
CONTRACTS = 30
RULE = "-" * 64
ELEVATORS = [("EL-041", "River Bend Elevator"), ("EL-107", "North Line Terminal"),
COMMODITIES = [("Yellow corn, US No 2", 4.42), ("Soybeans, US No 1", 10.86),
src/contract.pythe contract arithmetic — a swap seam
The rule as pure code: the carried balance, less the bushels actually applied this window, against the tolerance on file; then the roll deadline against the horizon this run owns. No model, no judgement. ⚠︎ The nine rules, the tolerances and the roll terms in it are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
You change it to: RULE_TEXT and step(). The nine clauses are illustrative; replace them with your own contract's and every figure on this page re-derives. ⚠︎ RULE_TEXT is also what the model reads, so a rule that does not say what an HTA is produces exactly the 6 errors this run made.
Business days, the weekly run calendar, and the roll deadline DERIVED from the delivery period rather than stored beside it -- because amending the period is exactly what a roll does, so a stored deadline would desynchronise on the one event this kit is about. Monday to Friday, no holiday calendar, declared rather than hidden.
You change it to: is_business_day and RUN_DATES. Monday to Friday with no holiday calendar today; a real roll deadline counts against whichever calendar the contract names.
src/schedule.py
# The clock. Business days, the weekly run calendar, and the roll deadline derived from both.
RUN_DATES = ("2026-09-07", "2026-09-14", "2026-09-21", "2026-09-28", "2026-10-05")
CADENCE_DAYS = 7
CADENCE_LABEL = "every Monday"
CADENCE_BUSINESS_DAYS = 5
def d(s):
def is_business_day(day):
def bdays_between(a, b):
def bdays_before(day, n):
def roll_deadline(delivery_start, notice_business_days):
src/state.pythe carried state — a swap seam
Four scalars per contract -- unfilled_bu, rolls_used, decision_raised, prev_status -- written by the arithmetic and never from a model reply, then rendered as one English paragraph for the prompt. A monitor that fed its own verdict forward would compound one bad reading into every reading after it, and the quantity carried here is bushels a producer is contractually on the hook for.
You change it to: What is carried between Mondays, and how it is worded. Four scalars today. ⚠︎ ENGLISH RATHER THAN JSON IS A DESIGN CHOICE AND IT IS UNMEASURED -- scoring the same 150 readings with the state rendered both ways costs another 150-call run and has not been paid for.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_contract(store, contract_id):
def describe(state):
src/segment.pythe section split
Seven named sections, asserted present in all 150 documents by evals/check_labels.py before a run may spend, so a parser drift is a refusal to start rather than a quietly truncated prompt.
src/segment.py
# Split a contract record into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Contract", "Delivery Terms", "Position This Run",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe privacy gate
Decides which sections go on the wire. Producer Contact -- name, mobile, farm registration number -- is mapped by nothing and subtracted unconditionally in the fallback, so a hint that matches nothing falls back to the record MINUS that section rather than to the whole record.
src/select.py
# Pick which sections of a record are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
CONTRACT = "Contract"
TERMS = "Delivery Terms"
POSITION = "Position This Run"
LOG = "Contract Activity Log"
PRODUCER = "Producer Contact"
NOTES = "Operational Notes"
NEVER_SENT = (PRODUCER,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts joined with string concatenation: the fixed instruction, the carried-state paragraph, and the record. The middle part is the experiment, and --stateless replaces exactly that one line.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/watch.pyone reading
One contract, one Monday, one completion call. Parses the dates and the printed figures in code first -- the model is never asked to count business days.
src/watch.py
# One contract, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled open-contract watch for a grain elevator. You apply a written "
MAX_TOKENS = 24000
FIELDS = ("status", "rollover_risk", "unfilled_bu", "raise_decision")
def documents():
def contracts():
def load_doc(doc_id):
def _num(text, pat, cast=float):
evals/baseline.pythe free floors — a swap seam
Three of them, all $0.00, all scored through the identical scorer: the expiry report a desk already runs, the same rule given the carried state, and the strongest table this author could write.
You change it to: The rejection and roll keyword tables, and the assumed tolerance. This is the column the model has to beat, so making it stronger is the honest direction to edit it in.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("planonly", "planonly-mem", "parsed-mem")
ASSUMED_TOLERANCE_PCT = 2.0
REJECT_WORDS = ("turned away", "refused", "rejected", "backed off", "not accepted",
ROLL_DONE_WORDS = ("amended to the", "executed on the desk", "re-papered", "was booked",
ROLL_NOT_DONE_WORDS = ("declined", "no election", "nothing was booked", "quote expires",
def _log_lines(text):
def _bushels_in(line):
def _rejected_bushels(text):
def _rolls_executed(text):
evals/scoring.pythe scorer
Exact match per cell against the answer key, with the two raise directions counted apart and the two rollover-risk directions counted apart. No model grades anything.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("status", "rollover_risk", "unfilled_bu", "raise_decision")
def _pct(n, d):
def _bu(v):
def score(records, golds):
evals/cadence.pythe cadence ablation
What a missed Monday costs, scored against the INTACT answer key. Free -- no provider is involved in a question about a calendar. It self-checks that a perfect reader scores 100.00 on the intact schedule before it prints anything.
evals/cadence.py
# What a MISSED RUN costs. Free -- no provider is involved in a question about a calendar.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FIELDS = ("status", "rollover_risk", "unfilled_bu", "raise_decision")
def load_gold():
def perfect_chain(gold, doc_ids):
def floor_chain(gold, doc_ids, mode):
def blind_window(gold, skip_run):
def report(gold, skip_run):
def main():
src/app.pythe local UI
http.server, no framework. Renders with no key: the record, the carried state, the parsed figures and the entire free floor are computed locally, and a second button replays what r001 actually answered off the committed result file.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9018"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-contract-expiry")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 30 open contracts x 5 Mondays = 150 records from a fixed seed (SEED = 20260823). Twenty-four of the thirty have a roll deadline that falls STRICTLY inside the five-week window, so exactly one run is the last one that can see each of them -- that is the hard clock. The gold labels are src/contract.step's output over the planted events, never typed. A swap seam.
src/contract.pyThe rule as pure code: the carried balance, less the bushels actually applied this window, against the tolerance on file; then the roll deadline against the horizon this run owns. No model, no judgement. ⚠︎ The nine rules, the tolerances and the roll terms in it are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus. A swap seam.
src/schedule.pyBusiness days, the weekly run calendar, and the roll deadline DERIVED from the delivery period rather than stored beside it -- because amending the period is exactly what a roll does, so a stored deadline would desynchronise on the one event this kit is about. Monday to Friday, no holiday calendar, declared rather than hidden. A swap seam.
src/state.pyFour scalars per contract -- unfilled_bu, rolls_used, decision_raised, prev_status -- written by the arithmetic and never from a model reply, then rendered as one English paragraph for the prompt. A monitor that fed its own verdict forward would compound one bad reading into every reading after it, and the quantity carried here is bushels a producer is contractually on the hook for. A swap seam.
src/segment.pySeven named sections, asserted present in all 150 documents by evals/check_labels.py before a run may spend, so a parser drift is a refusal to start rather than a quietly truncated prompt.
src/select.pyDecides which sections go on the wire. Producer Contact -- name, mobile, farm registration number -- is mapped by nothing and subtracted unconditionally in the fallback, so a hint that matches nothing falls back to the record MINUS that section rather than to the whole record.
src/prompt.pyThree parts joined with string concatenation: the fixed instruction, the carried-state paragraph, and the record. The middle part is the experiment, and --stateless replaces exactly that one line.
src/watch.pyOne contract, one Monday, one completion call. Parses the dates and the printed figures in code first -- the model is never asked to count business days.
evals/baseline.pyThree of them, all $0.00, all scored through the identical scorer: the expiry report a desk already runs, the same rule given the carried state, and the strongest table this author could write. A swap seam.
evals/scoring.pyExact match per cell against the answer key, with the two raise directions counted apart and the two rollover-risk directions counted apart. No model grades anything.
evals/cadence.pyWhat a missed Monday costs, scored against the INTACT answer key. Free -- no provider is involved in a question about a calendar. It self-checks that a perfect reader scores 100.00 on the intact schedule before it prints anything.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2096 input and 2712 output tokens per reading (one contract, at one scheduled run), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one contract, at one scheduled run)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one contract, at one scheduled run) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Operational Notes, which are one of three fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- prose a merchandising clerk types into the contract system -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
The experimentWe did NOT attack it — and this is the surface we left open
An indirect prompt injection needs a field somebody outside can write into, and this kit has exactly one: the Operational Notes a merchandising clerk types into the contract system. It is SENT rather than hidden, because hiding a surface does not close it. No attack was fired at this kit and none is claimed. The three boundaries below were checked in code on 2026-08-23, and only the first of them is red-proven in both directions rather than argued from an absent code path.
Boundary checked
What could go wrong
What the code guarantees
Does a producer's name, mobile and farm registration number ever leave the machine?
Every record carries a Producer Contact section -- the farmer's name, mobile number and FSA farm registration number. A farm registration number ties a person to a legal entity, a field boundary and a payment history, and not one field this kit answers asks for any of it. A selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the record MINUS that section rather than to the record. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py over all 150 documents: 0 of 150 leak with the guard, 150 of 150 leak with the naive or list(secs) AND the condition that reaches it reproduced (a grain accounting package renaming its blocks on an upgrade, after which only the personal-data section still parses). ⚠︎ WITHOUT REPRODUCING THAT CONDITION THE PROOF MEASURES ZERO AND IS THE TEST NOT FIRING -- on today's corpus every hint names a section all 150 documents carry, so the fallback is not on any live code path.
Can anything here cancel, price or default a contract, book a roll or ring a producer?
A watchlist that produces DECISION_DUE is one function call away from a system that acts on it, and the acting version is the one a buyer asks for next.
No such code path exists. The only writers in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). evals/check_labels.py greps every .py and .js file for such names and passes at 0. ⚠︎ IT ASSERTS THE ABSENCE OF NAMES IT KNOWS; a path called something else would pass it, which is stated here rather than hidden.
Can a wrong reading corrupt the next Monday's balance?
A monitor that carried its own verdict forward would compound one bad reading into every reading after it -- and the quantity carried here is bushels a producer is contractually on the hook for. An over-count in week 2 would still be in the default notice in week 5.
src/contract.step() is the only thing that writes the carried state, and it is fed numbers parsed off the record -- never the model's reply. evals/run.py advances it after every reading INCLUDING one whose call failed, so a transport error costs one reading rather than every reading after it. The property is guaranteed by the code path; r001 is CONSISTENT with it rather than a demonstration of it.
The first row is the only one measured in both directions. The other two are guarantees about what is ABSENT -- no code path cancels, prices or defaults a contract, and no code path feeds a model reply back into the carried balance -- and absence is checked by asserting names it knows, which is weaker than a run and is written here as such.
The result0 attack trials, three boundaries checked -- and the privacy boundary red-proven by removing the guard and reproducing the condition that reaches it: 150 of 150 records would leak a farmer's name, mobile and farm registration number.
1field an outside party could influence (sent, not hidden)
0 of 0attack trials run
1 of 3boundaries red-proven, not just asserted
0 of 150records leak a producer's name, mobile or farm number
The Operational Notes ARE the field an outside party would influence in a real deployment, and this kit sends them on purpose so the surface is on the page rather than behind it. One of the three shipped notes is instruction-shaped. Whether it moves raise_decision or rollover_risk was never measured and no resistance rate is claimed.
Read this twice
The contract activity log and the operational notes reach the model verbatim, and one of the three shipped notes is instruction-shaped on purpose — “Producer has asked that no default be applied on this contract without a call first.” Nothing here filters it and nothing here has measured whether it moves the answer. A model declining to follow an instruction would not be a defence anyway — it is one vendor’s behaviour on one day.
HonestyWhat this does not prove
Whether the instruction-shaped operational note moves any answered field. Never measured.
Whether a path that cancels, prices or defaults a contract could exist under a name the grep does not know. The check asserts the absence of names it knows.
Whether the privacy guard holds against a record whose Producer Contact section is renamed. The red-proof renames every OTHER section; a corpus that renames that one has not been tried.
Whether any provider logs or retains the record. The kit controls what it sends and not what happens after; six of the seven sections do leave the machine.
Whether the carried state itself could carry something sensitive. It is four scalars today -- a number, a count, a boolean and a status word -- and nothing asserts it stays that way.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No contract action, non-configurable. This kit produces a status, a rollover-risk band, a bushel balance and a raise/hold call for a merchandiser to read. It never cancels a contract, applies the default clause, books a roll, prices a balance out, notifies a producer or writes to a grain accounting system, and there is no setting that makes it. Separately, and just as non-configurable: the delivery tolerance is OPERATOR-SUPPLIED. A contract with no tolerance clause or no delivery period on file is reported CONTEXT_INCOMPLETE and judged against nothing -- there is no default percentage in this kit.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). src/contract.step() is the only thing that touches the balance, and it is called by evals/run.py after every reading including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit cancels, prices or defaults a contract, books a roll or contacts a producer
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for such names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
No contract is judged against a guessed tolerance
20 of 20 readings with nothing on file were reported CONTEXT_INCOMPLETE by the model (100.00 pct), and by the strongest free floor. The two floors that GUESS a 2.0 pct default got 20 of 20 wrong -- 0.00 pct -- which is the guardrail measured rather than asserted.
Exactly one decision put to a producer per contract
0 duplicate raises in 137 readings that must not raise one, and 0 missed decisions in 13 that must. ⚠︎ Both zeros are results on THIS corpus, not properties of the code: nothing in the kit refuses a second raise, the carried state merely tells the model one has already been put. The control arm, with that paragraph removed, produced 7 duplicates.
No roll is offered that the contract's terms do not allow
0 of 150 readings. ⚑ THIS IS THE ONE THE MODEL WINS AND THE FREE FLOOR LOSES: at the identical headline accuracy the floor offered 5 and the control offered 12. It is a result on this corpus and not a property of the code -- nothing here refuses to say ROLLABLE.
A wrong answer does not propagate into the next scheduled run
The property is guaranteed by the code path -- src/contract.step() never reads the reply -- and r001 is CONSISTENT with it rather than a demonstration of it: the reading that returned nothing at all (CT-2026-0017-W4) did not damage the balance carried into CT-2026-0017-W5, which was answered correctly.
The producer's name, mobile and farm registration never reach the provider
0 of 150 records leak the Producer Contact section, and 150 of 150 leak it when the guard is removed AND the condition that reaches the fallback is reproduced. Both directions asserted by evals/check_labels.py before any run may spend.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The balance is correct whatever the model says, which means a wrong reading is a wrong watchlist row and not a corrupted history. Those are different problems and only the first one is on this page.
IT IS NOT A SCHEDULER. evals/run.py is invoked; it is not woken. Nothing here detects a missed run, back-fills it or marks its readings late -- and evals/cadence.py measures exactly what that is worth rather than leaving it as a caveat.
IT IS NOT AN AUTHORITY ON GRAIN CONTRACTS. The nine clauses, the tolerances and the roll terms are invented for this kit and reproduce no grain contract, elevator's terms and conditions, exchange rule or state warehouse regulation.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 42 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
18 measured by the latest run24 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key
alarm
rollover_risk_accuracy_pct; status_accuracy_pct; unfilled_bu_accuracy_pct; raise_decision_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Both paid arms had exactly 1 of 150, both at the ceiling, and the two are different documents -- so 1 is a property of this ceiling on this workload rather than of either arm.
rollover-directions
The two rollover-risk directions, counted apart
alarm
false_rollable; lost_rollable — alarm on false_rollable above 0. A missed roll is visible to whoever reads the watchlist next week; an offer that was never available is visible to the producer first.
missed-run-ablation
What one missed Monday costs a PERFECT reader
alarm
blind_window_contracts; contracts_with_a_poisoned_balance; cells_moved_direction; unfilled_bu_accuracy_pct — alarm on blind_window_contracts above 0 on any schedule. Every other number here degrades; this one is a permanent loss.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
778,831
open-contract records edited — the count held, the bytes did not
split.count
30
the contracts count moved — a different set was scored
split.size_p50
5
the median size of one contract moved
split.size_p95
5
the 95th-percentile size of one contract moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cadence_days 7, context_incomplete_cells 20, contracts 30, decision_cells 13, documents 150, memory_cells 94, quiet_cells 137, readings_scored 150, stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
rollover risk -- the headline
not yet known
150 readings
A band is the spread between runs of the SAME arm at the same settings, and there has been no repeat. r001 and s001 are one run each of two DIFFERENT prompts, which is a comparison and not a band. ⚠︎ And note what this figure is level with: the strongest FREE floor scores 95.33 pct on the same 150 readings, which is the same number to two decimal places.
rolls offered that were not available
0 -- and this is the one band on the page that is a hard floor rather than a spread
150 readings; 39 of them NOT_ROLLABLE in the key
evals/scoring.py::score, r001-contract-expiry. The free floor makes 5 of these on the identical cells and the stateless control makes 12, so the grader DOES discriminate -- just not on the headline rate.
available rolls missed
6 of 46 ROLLABLE readings -- all six the same Rule C-8 gap
46 readings ROLLABLE in the key
evals/scoring.py::score, r001-contract-expiry. Every one is Rule C-8 read across from a basis contract onto an HTA. It is the conservative direction and it is a PROSE defect, not a model defect.
the decision call, both directions
0 missed and 0 duplicate -- saturated, and that is a statement about the corpus rather than about the arm
13 readings that must raise a decision, 137 that must not
A grader an arm aces has stopped discriminating -- and here EVERY arm aces the missed direction, including both free floors, because the deadline arithmetic is printed on the page. The duplicate direction still separates: the stateless control raises 7.
the bushel balance, exact
96.67 pct -- 145 of 150, and the four misses are scale-ticket lines
150 readings; 93 of them memory-dependent
evals/scoring.py::score, r001-contract-expiry. The free floor scores 92.67 pct on the same cells and the stateless control 40.00, so this grader discriminates by 4.00 points against free code and 56.67 against memorylessness -- the widest separation on the page.
context-incomplete recall -- the guardrail, scored
100.00 pct, saturated
20 readings with no tolerance clause or no delivery period on file
Both paid arms and the strongest free floor abstain on all 20. The two floors that GUESS a 2.0 pct default get all 20 wrong. That is the guardrail measured rather than asserted -- and the denominator is small, so one row is 5.0 points.
replies that returned nothing
1 of 150 in each paid arm, on different documents
150 readings per arm
finish_reason length at 24,000 of 24,000 output tokens. Two arms, two different documents, neither longer nor structurally odd. Counted as wrong in all four fields rather than dropped from the denominator.
the schedule holding
not a metric of any run -- 4 contracts a week warned by NOBODY when one Monday is missed
30 contracts, 3 skippable Mondays
evals/cadence.py, free, scored against the intact key. Across the three ablations 253 cells moved and 0 went the right way. This band cannot be improved by a better model; the perfect-reader figure is the ceiling on any arm.
status
99.33 pct -- and the free floor scores the same number to two decimal places
150 readings
evals/scoring.py::score. Status is FILLED / ON_TRACK / DECISION_DUE / DEADLINE_PASSED / CONTEXT_INCOMPLETE, and all five inputs to it are PRINTED on the record -- the tolerance, the balance the carried state supplies, and the two day counts. So this grader stopped discriminating between the model and free code, and the stateless control at 84.67 pct is the only arm it separates.
the priced exposure
0.08 pct error against the key's total
18,280,201.99 USD of unfilled position across 150 readings
The bushels each reading claims are still owed, priced at the board reference on file. It is a WHOLE-BOOK figure and it hides sign: over- and under-statements cancel, which is why the exact bushel band above is the one to read first. The free floor lands at 1.26 pct and the stateless control at 44.63.
latency
p50 13773 ms, p95 78331 ms -- no repeat, so this is one run's spread and not a tolerance
149 readings that returned something, at 12 concurrent contract chains
END-TO-END per reading, measured under load rather than on a quiet call. The p95 is 5.7x the p50 because output length varies with how much reconciling a window needs. The stateless control runs at p50 21242 ms for the same work, which is the cost of reasoning without history.
tokens, and where they went
314,447 in, 406,856 out over 150 readings -- and 95.98 pct of the output is provider-side reasoning
150 readings
Recorded per call by src/adapters and summed by evals/run.py, BILLED OR NOT -- the reading that returned nothing is in these totals because it was billed. Output exceeds input on this kit, which is unusual for a document task and is entirely the reasoning default.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-contract-expiry-planonly 2026-08-23
b001-contract-expiry-planonly-mem 2026-08-23
b002-contract-expiry-parsed-mem 2026-08-23
answered, %
100.0
100.0
100.0
context incomplete recall, %
0.0
0.0
100.0
duplicate raise rate, %
4.38
1.46
0.00
exposure usd error, %
20.85
29.95
1.26
false rollable
34
21
5
input tokens, whole run
0
0
0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
lost rollable
0
0
1
memory raise accuracy, %
95.74
100.00
100.00
memory risk accuracy, %
46.81
84.04
96.81
memory unfilled accuracy, %
12.77
72.34
91.49
missed decision, %
0.0
0.0
0.0
output tokens, whole run
0
0
0
raise decision accuracy, %
96.00
98.67
100.00
rollover risk accuracy, %
47.33
71.33
95.33
status accuracy, %
70.00
85.33
99.33
unfilled bu accuracy, %
8.00
64.67
92.67
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 3 runs. Columns here are only ever compared with each other.
Metric
c000-contract-expiry-calibration 2026-08-23
r001-contract-expiry 2026-08-23
s001-contract-expiry-stateless 2026-08-23
answered, %
100.00
99.33
99.33
context incomplete recall, %
—
100.0
100.0
duplicate raise rate, %
0.00
0.00
5.11
exposure usd error, %
0.00
0.08
44.63
false rollable
0
0
12
input tokens, whole run
21117
314447
307454
model latency p50 ms
17812.00
13773.00
21242.00
model latency p95 ms
59109.00
78331.00
105900.00
lost rollable
3
6
6
memory raise accuracy, %
100.00
100.00
92.47
memory risk accuracy, %
75.00
94.62
59.14
memory unfilled accuracy, %
100.00
98.92
6.45
missed decision, %
0.0
0.0
0.0
output tokens, whole run
29880
406856
570871
raise decision accuracy, %
100.00
99.33
94.67
rollover risk accuracy, %
70.00
95.33
72.67
status accuracy, %
100.00
99.33
84.67
unfilled bu accuracy, %
100.00
96.67
40.00
not a time series No two of these 3 runs measured the same system — they differ on context_incomplete_cells, contracts, decision_cells, documents, max_tokens, memory_cells, quiet_cells, readings_scored, stateless — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-contract-expiry-stub 2026-08-23
answered, %
100.0
context incomplete recall, %
0.0
duplicate raise rate, %
4.38
exposure usd error, %
20.85
false rollable
34
input tokens, whole run
337210
model latency p50 ms
0.00
model latency p95 ms
0.00
lost rollable
0
memory raise accuracy, %
95.74
memory risk accuracy, %
46.81
memory unfilled accuracy, %
12.77
missed decision, %
0.0
output tokens, whole run
6257
raise decision accuracy, %
96.0
rollover risk accuracy, %
47.33
status accuracy, %
70.0
unfilled bu accuracy, %
8.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 18 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
r001-contract-expiry against s001-contract-expiry-stateless. Same corpus, same model, same grader, same 150 readings; the prompts differ in exactly one block and evals/check_labels.py asserts they are identical everywhere else. ⚠︎ Each arm lost one call to the ceiling and both are counted as wrong rather than dropped, so every figure on both sides carries up to 0.67 points of that.
whether a run happens on the Monday it was scheduled for
a PERFECT reader's bushel accuracy 100.00 pct -> 45.00 pct, rollover risk 100.00 -> 77.50, 121 cells moved and 121 of 121 right to wrong, 22 of 30 contracts left with a permanently wrong balance, 173,800 bushels never seen by any run, and 4 contracts whose roll deadline nobody warns about on any schedule
measured
results/cadence-contract-expiry.json, the 2026-09-14 ablation. Free -- no provider was called. Scored against the INTACT answer key, because the truth about a contract does not change because nobody looked at it. ⚑ THIS IS A CEILING ON EVERY ARM, not an estimate of one: a model cannot beat a perfect reader, only add its own errors on top.
whether the free floor is allowed to read the printed pricing date
the strongest free floor's rollover risk 94.00 pct -> 95.33 pct, and rolls falsely offered 7 -> 5
measured
evals/baseline._roll_right, before and after. The first version returned True for every delayed-price contract with a comment claiming the record did not print the comparison; it prints both dates and watch.position_of already parses one of them. ⚠︎ THIS LEVER MOVES THE KIT'S OWN HEADLINE THE WRONG WAY AND WAS PULLED ANYWAY. A floor that declines a comparison the page hands it is a strawman, and the published tie is against the stronger version.
how many collision templates the activity log carries
the strongest free floor's rollover risk 98.67 pct -> 95.33 pct, with its errors moving from one direction to both
measured
tools/build_corpus.ROLL_EXECUTED and ROLL_NOT_EXECUTED, five templates a side and then seven. With ten easy templates a keyword table separated them almost perfectly, which was a fact about the templates and not about the work. Four harder ones -- two executions that read like refusals, two refusals that read like executions -- took the same untouched floor down 3.34 points.
whether the free rule is given the carried balance
rollover risk 47.33 pct -> 71.33 pct and the bushel balance 8.00 -> 64.67, with NO MODEL anywhere near either arm
measured
b000-contract-expiry-planonly against b001-contract-expiry-planonly-mem. Identical code but for the carried state, both $0.00, both scored through the same scorer. This is the cleanest statement on the page of what MEMORY is worth as against what a MODEL is worth.
the delivery tolerance on a contract
which contracts are on the watchlist at all -- it is the FILLED boundary, and 23 of 150 readings sit inside it
reasoning
src/contract.tolerance_bushels and Rule C-5. Not measured: no run varied the tolerance. A contract inside tolerance is off the decision list however close its deadline is, so widening it silently removes rows from the watch -- which is a design consequence read off the code, not a figure from a result file.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
rollover risk -- the headline
nothing yet. What SHOULD fire is the direction, not the rate -- see the next two rows.
rolls offered that were not available
any single one. A roll offered that the terms do not allow reaches the producer as a promise the elevator then has to withdraw, and unlike a missed roll it cannot be quietly corrected next Monday.
available rolls missed
a seventh, or any one on a contract type other than an HTA. The first would say the clause is worse than believed; the second would say this is not the failure mode anybody thought it was.
the decision call, both directions
any missed decision at all. That direction is not a late alert -- past the deadline the default clause has applied and no later run recovers the choice.
the bushel balance, exact
a miss on a window whose activity log names no bushels at all. Every miss in this run is on a line that does; one without would be a different failure.
context-incomplete recall -- the guardrail, scored
any accrual against a contract with nothing on file. It is the one thing this row's guardrail forbids outright.
replies that returned nothing
a second in one arm, or one on a document another arm also lost. Either would say the ceiling is the problem rather than the reading.
the schedule holding
a gap detected by src/schedule.gap(). Detecting is free and recovering is impossible -- the bushels applied in an unread week are printed on one record and one only.
status
a status error on a reading whose bushel balance is right. Every status miss in this run inherits a balance miss; one that did not would be a precedence failure instead.
the priced exposure
this rate moving while the exact bushel rate does not. That combination means the errors have stopped cancelling and are concentrating on the big contracts.
latency
p95 approaching the point where a full book would not finish inside its own cadence. At this p95, 2,000 contracts at 12 workers is roughly 218 minutes a run -- fine weekly.
tokens, and where they went
the reasoning share moving without the prompt changing. It would mean the vendor changed a default, which reprices every row in the Cost lens with nothing visible on any page.
NextThe three you would add first
A scheduler, and something that notices when a run did not happenThis is the gap evals/cadence.py prices: 4 contracts a week warned by NOBODY when one Monday is missed, and a balance that stays wrong for the rest of the contract's life. src/schedule.gap() detects a gap for free and cannot repair one.
A second writer lock on data/state.jsonOne file replaced atomically is correct for one writer. An elevator running two watches is two writers and nothing here coordinates them.
A rule text that says which instruments Rule C-8 reachesEvery rollover-risk error in the scored run is that one gap. It is a prose fix, not a model fix, and it is the cheapest improvement available to this kit.
A refusal to raise a second decision, in codeToday the carried state TELLS the model one has already been put and the model complies. The control shows what happens when it is not told: 7 duplicates.
A holiday calendarBusiness days here are Monday to Friday with nothing else. Every deadline count is wrong by whatever holidays fall inside it, and a roll deadline is not a soft date.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py and evals/cadence.py (both free, seconds) on any change to tools/build_corpus.py, src/contract.py, src/schedule.py, src/segment.py or src/select.py. Re-run the scored eval AND its stateless control together (paid, 300 calls) on any change to src/prompt.py or src/state.py.
What this cannot tell you
Whether a contract-action path could exist under a name the grep does not know.
Whether zero duplicate raises survives a corpus with more DECISION_DUE cells. There are 13 in this one, so the claim is about 13 readings and nothing wider.
Whether zero falsely-offered rolls survives a different roll-template mix. It is a result on this corpus, and the corpus's roll templates were rewritten once already.
Whether the instruction-shaped operational note can move raise_decision. Never attacked.
Whether the carried state stays four scalars. Nothing asserts it, and a fifth field carrying free text would put producer data back on the wire.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a BALANCE -- four scalar fields written by arithmetic. The cost is flat in history length because nothing here remembers what was said, only what was counted, and a marketing year is fifty-two runs. It would also give back the failure this design exists to avoid: a running total fed from its own model output compounds forever, and here the total is bushels a producer owes
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the schedule
evals/run.py and src/schedule.py
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
THIS is the seam where a framework genuinely earns its place, and this kit does not argue it -- it PRICES it. evals/cadence.py measures what one missed Monday costs for $0.00: 4 contracts warned by nobody, 173,800 bushels never seen, and a balance wrong for the rest of the contract's life. 30 independent chains of 5 strictly-ordered readings, with the cadence, retries, missed-run detection and late back-fill all OUTSIDE the kit. A ThreadPoolExecutor is the right size for an eval and the wrong size for a watch that has to notice it did not wake
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
four exact-match comparisons, two confusion directions counted apart, three slices of the same cells and one dollar sum is a dict comprehension, not a platform
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each contract is a chain of five readings -- five consecutive Mondays -- with no branching and exactly one edge between consecutive runs, carrying four scalars. Different contracts never touch. A framework would add an orchestrator to a for-loop that already runs 30 chains wide.
The other sideWhat a framework costs you
No scheduler. This is a watch that carries state and is still INVOKED, not woken -- and unlike most kits that say this, the price is measured here rather than implied: see the cadence ablation.
No missed-run detection that recovers anything. src/schedule.gap() notices; nothing repairs.
No concurrency model for the state store. data/state.json is one file replaced atomically -- correct for one writer, and an elevator with two watches is two writers.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
The scheduling seam is priced but not tested. Nothing here has ever run under cron, Airflow or Temporal -- what is measured is the COST of the schedule not holding, not the benefit of any particular engine holding it.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-contract-expiry on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
13,773 ms
p50 13773 ms, p95 78331 ms -- no repeat, so this is one run's spread and not a tolerance
p95 approaching the point where a full book would not finish inside its own cadence. At this p95, 2,000 contracts at 12 workers is roughly 218 minutes a run -- fine weekly.
Model, p95
78,331 ms
p50 13773 ms, p95 78331 ms -- no repeat, so this is one run's spread and not a tolerance
p95 approaching the point where a full book would not finish inside its own cadence. At this p95, 2,000 contracts at 12 workers is roughly 218 minutes a run -- fine weekly.
Input tokens
314,447
314,447 in, 406,856 out over 150 readings -- and 95.98 pct of the output is provider-side reasoning
the reasoning share moving without the prompt changing. It would mean the vendor changed a default, which reprices every row in the Cost lens with nothing visible on any page.
Output tokens
406,856
314,447 in, 406,856 out over 150 readings -- and 95.98 pct of the output is provider-side reasoning
the reasoning share moving without the prompt changing. It would mean the vendor changed a default, which reprices every row in the Cost lens with nothing visible on any page.
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-contract-expiry-calibration17,812 ms
r001-contract-expiry13,773 ms
s001-contract-expiry-stateless21,242 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-contract-expiry-planonly, b001-contract-expiry-planonly-mem, b002-contract-expiry-parsed-mem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
contract records
data/corpus/CT-2026-<n>-W<k>.txt -- 150 files, 778,831 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Producer Contact -- the farmer's name, mobile and farm registration number -- never does, by src/select.NEVER_SENT
the carried state
data/state.json in a deployment, and in-memory per run in the eval -- four scalars per contract, written by src/contract.step and never by a model
as one English paragraph in every prompt, and the UI prints the same paragraph verbatim so a reader can audit what the model was told
the answer key
data/gold.jsonl -- 150 rows, the output of src/contract.step over the planted events, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the run records
results/eval-*.json plus results/cadence-contract-expiry.json in the kit, and one small record per run in the app repo's run register
never -- they are read by the site build, not by a provider
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a WEEKLY watch: every Monday, 5 runs over the shipped window. One run owns exactly the business days to the next one, and DECISION_DUE is DEFINED against that horizon (Rule C-6) -- so the schedule is part of the ANSWER and not only part of the bill.
30 contracts x 5 Mondays = 150 calls per arm, 0 gaps and 0 out-of-order readings; evals/check_labels.py asserts all 30 sequences are complete and that the five run dates are business days one cadence apart before a run may spend. Wall clock 444.4s at 12 workers. A book of 2,000 open contracts on this cadence is 2,000 calls a week, $18.37 on the shared projection card. (r001-contract-expiry, evals/check_labels.py, evals/cadence.py, src/schedule.RUN_DATES)
⚑ WHAT A MISSED RUN COSTS IS MEASURED HERE, NOT ARGUED, AND IT COST $0.00. Skip one Monday and two things happen. Every roll deadline that falls after it and on or before the next Monday is warned about by NOBODY on any schedule -- 4, 5 and 4 contracts for the three skippable Mondays -- because the run before it correctly saw the deadline beyond its horizon and the next run to fire is already past it. And the bushels applied that week are printed on one record and one only, so the running total is understated for the REST of the contract's life: missing 2026-09-14 leaves 22 of 30 contracts permanently wrong, 173,800 bushels never seen, $1,072,213.04 of position at the board reference. A perfect reader's bushel accuracy falls from 100.00 to 45.00 pct, 121 cells move and 121 of 121 go right to wrong. ⚠︎ AND THE CALENDAR UNDER ALL OF IT IS MONDAY TO FRIDAY WITH NO HOLIDAYS. src/schedule.is_business_day is a weekday test and nothing else, and every roll deadline on this page is derived from it, so each one is wrong by whatever holidays fall inside its notice period. A roll deadline is not a soft date. Replacing that one function re-derives every figure here.
every figure on this page is per a weekly watch. A daily watch pays 5x and shrinks the blind window to one day; a fortnightly watch pays half and doubles it. Nothing here measures either -- evals/cadence.py measures runs being SKIPPED on this cadence, which is a different question from the cadence being changed.
state
FOUR SCALARS PER CONTRACT, written by arithmetic and rendered as one English paragraph: bushels still owed, rolls executed, whether the decision has been put, and what the last run reported. Not the previous record, not the previous prompt, not the previous reply.
The control removes exactly that paragraph and nothing else -- evals/check_labels.py asserts the two prompts differ on one line. It costs 22.66 points of rollover risk, 56.67 points of bushel accuracy, and on the 93 memory-dependent readings the bushel column falls to 6.45 pct. It also costs 1.35x more to run, because a model told there is no history for a balance printed nowhere reasons at length: 570,871 output tokens against 406,856. (r001-contract-expiry, s001-contract-expiry-stateless, evals/check_labels.py)
The state is written from src/contract.step and NEVER from a model reply, so a wrong reading costs one reading rather than every reading after it. That is a code-path guarantee; the run is consistent with it rather than a demonstration of it.
a deployment that persisted the model's own answer into data/state.json would invalidate every figure here, because the failure mode being measured is one this design cannot have.
model
ONE completion call per reading, one provider, one key, read from .env by src/config.py -- shared repo root first, then the kit's own, then the real environment. MAX_TOKENS is 24000 and the harness refuses to override it for any run id that does not begin with c, so a figure measured under a ceiling the page does not name cannot be mistaken for a scored one. thinking is never sent; provider-side reasoning is left at the vendor default.
150 readings, 149 answered, 1 lost to the ceiling. Input averages 2096 tokens and output 2712, of which 95.98 pct is provider-side reasoning. p50 13773 ms, p95 78331 ms at 12 concurrent chains. The ceiling itself came from c000: two chains at an 8,000-token cap, hardest reading 6745 output tokens, all 10 parsed. (r001-contract-expiry, c000-contract-expiry-calibration, src/watch.MAX_TOKENS)
⚠︎ 24000 WAS NOT ENOUGH FOR ONE READING IN EACH PAID ARM, on different documents (CT-2026-0017-W4 and CT-2026-0012-W4), and neither is longer or structurally odd -- every record in this corpus is within a few hundred bytes of every other. A reply that runs to the ceiling is BILLED and returns nothing, and both are counted as wrong rather than dropped. Why those two is not known and raising the ceiling to find out was not paid for.
every figure on this page is one tier at one ceiling. A different model, a different reasoning default or a different ceiling is a different run, and no cross-model claim is made anywhere here -- the two paid arms differ by the CARRIED STATE, not by the model.
labels
A GENERATED answer key, not a captured one. data/gold.jsonl is src/contract.step's own output over the events tools/build_corpus.py planted, so the key and the arithmetic cannot drift; evals/check_labels.py re-derives all 30 chains and requires an exact match on all four fields before any run may spend. Scoring stops at exact equality per cell -- no model grades anything, and the two error directions on the raise call and on the rollover risk are counted apart rather than averaged.
150 rows, 94 of them memory-dependent, 20 with no tolerance clause or delivery period on file, 778,831 bytes of corpus from seed 20260823. Five arms scored through one scorer land at 95.33 / 72.67 / 95.33 / 71.33 / 47.33 pct on rollover risk, which is the evidence the set can tell them apart. (data/corpus-stats.json, evals/check_labels.py, evals/scoring.py)
The key is only as good as the RULE it expresses, and one clause in it is known to be under-specified: Rule C-8 says a basis contract with a fixed futures leg carries no roll right and says nothing about an HTA. Every rollover-risk error in the scored run is that gap. The rule was not rewritten and the run was not re-fired. ⚠︎ AND THE FLOOR SHARES THE KEY'S IMPLEMENTATION, so it cannot misread the rule while the model can only read the prose -- an asymmetry that flatters the floor by exactly those cells.
a real book. This population is 30 contracts open at all 5 runs, on records that deliberately carry no cumulative applied-to-date figure -- A contract that DISAPPEARS between runs is the case nothing here handles: its state stays in the store forever and nothing prunes it. A book that prints one, or that gains and loses contracts mid-window, is a different and easier question.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a producer told a roll is available, and the contract's terms do not allow one
something counted a roll that was never executed, or missed one that was. Both live in the activity log's prose: a roll can be booked in words that read like a refusal, and declined in words that read like a booking
read the activity-log line, not the band. On CT-2026-0009-W2 the free keyword floor reads "the desk quoted the spread on Friday and the producer took it, the period has moved out" as a quote, leaves the single permitted roll unspent and reports ROLLABLE; the key says NOT_ROLLABLE. The floor did this 5 times and the scored run did it 0 times (results/eval-b002-contract-expiry-parsed-mem.json against results/eval-r001-contract-expiry.json, CT-2026-0009-W2)
a bushel balance that is out by exactly the size of one scale ticket
a load turned away at the probe was counted as applied, or a load taken at a discount was subtracted. A discount changes the price, not the quantity (Rule C-2), and both sentences mention moisture, grade and the discount schedule
compare the window's gross figure against the log lines that name bushels. The scored run got 96.67 pct of balances exactly right against the free floor's 92.67, and the gap is entirely these lines (results/eval-r001-contract-expiry.json misses, field unfilled_bu, against results/eval-b002-contract-expiry-parsed-mem.json)
a balance that has been wrong since some week, and stays wrong
a scheduled run did not happen. The bushels applied that week are printed on one record and one only, and nothing on any later page carries a cumulative figure, so the running total is understated for the rest of the contract's life
run evals/cadence.py -- it is free -- and check src/schedule.gap(). Missing one Monday leaves 22 of 30 contracts permanently wrong, 173,800 bushels never seen, and 4 contracts whose roll deadline nobody warns about on any schedule (results/cadence-contract-expiry.json)
a second call to a producer about a contract they were already rung about
the reading was made without the carried state -- nothing on a record says a decision has already been put, so a reader who cannot see last Monday raises it again
check what the carried state said before disputing the call. The stateless control raised 7 duplicates in 137 quiet readings; the stateful arm raised 0 (results/eval-s001-contract-expiry-stateless.json against results/eval-r001-contract-expiry.json)
a contract being judged against a tolerance nobody filed
something filled a missing clause with a default. That is the one thing this row's guardrail forbids -- the tolerance is operator-supplied and a guessed one produces a decision nobody can defend
look at the Delivery tolerance line. Both floors that guess 2.0 pct get all 20 of these readings wrong; both paid arms that abstain get all 20 right (results/eval-b000-contract-expiry-planonly.json against results/eval-r001-contract-expiry.json)
['What a missed run costs the MODEL (only what it costs a perfect reader, which bounds it).', 'What a different cadence does, as opposed to a run on this cadence being skipped.', 'What a changing population does -- contracts joining or leaving the book mid-window.', 'What a holiday calendar does to every deadline count.', 'Whether English beats JSON for the carried state.', 'Whether a second model or a second tier agrees with this one.', 'Whether the instruction-shaped operational note moves any answered field.', 'Why the 24,000-token ceiling was not enough for exactly one reading in each paid arm, on different documents.']
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every byte was written by the generator committed beside it. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineThe rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key
For each of the 150 readings and each of the four answered fields, did the reply equal the computed answer key? Three of the four are words from a closed list and are compared exactly; the bushel balance is compared as a whole number, so 13100.0 and 13100 are the same answer and a reply that does not parse is a MISS rather than a zero -- 0 bushels is a meaningful answer here (it is what every FILLED reading gets) and a parse failure must not be scored as one.
$0.00per 1,000 open-contract records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function all three free floors and the cadence ablation are scored through.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
CT-2026-0009-W2 -- scheduled run 2 of 5, Monday 2026-09-14, a 15,000 bu delayed-price soybean contract with a 2.0 pct tolerance, one roll permitted, and 600 bu applied gross at the scale this week
What last Monday left behind
As at the previous scheduled run, 13,700 bushels were still owed on this contract, before anything this window applied. No roll has been executed on this contract on any earlier run. The roll/cancel/deliver decision has ALREADY been put to the producer on an earlier run, so this run must not raise it a second time. It was last reported DECISION_DUE.
The activity log, this window only
the desk quoted the spread on Friday and the producer took it, the period has moved out / an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied
The answer key
ON_TRACK, rollover risk NOT_ROLLABLE, 13,100 bushels still owed, raise_decision NO
What run r001 answered
ON_TRACK, rollover risk NOT_ROLLABLE, 13,100 bushels still owed, raise_decision NO
What the free floor answered
ON_TRACK, rollover risk ROLLABLE, 13,200 bushels still owed, raise_decision NO
Why this one
It is the only reading in this corpus where BOTH of the generator's deliberate collisions fire in one window, and the free floor loses both. The roll line reads like a quote and is an execution; the ticket line reads like a rejection and is an application. Chosen by tools/pick_shot.py off the answer key -- memory-dependent first, then the costly error direction, then exposure -- rather than by hand.
Grader
Verdict
Why
The rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key
rollover risk hit, status hit, bushels hit, decision hit -- four of four
CT-2026-0009-W2, run 2 of 5 on a 15,000 bu delayed-price contract with a 2.0 pct tolerance and one roll permitted. 13,700 bushels were carried in and 600 came over the scale gross. Two log lines decide the reading and both are collisions: "the desk quoted the spread on Friday and the producer took it, the period has moved out" is a roll that was EXECUTED, spending the single permitted roll; "an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied" is an APPLICATION, not a turnaway. So 13,700 minus 600 is 13,100 owed, above the 300 bu tolerance, no roll remains, and the deadline is 19 business days out against a 5-day horizon: ON_TRACK, NOT_ROLLABLE, 13,100, no. The model answered all four correctly and its rationale names both facts.
The two rollover-risk directions, counted apart
in scope on the NOT_ROLLABLE side, and a hit
The key says NOT_ROLLABLE, so this reading sits inside the population the falsely-offered count is measured over. The model did not offer a roll. The free floor DID: its execution table keys on the desk's verbs for a completed roll -- amended, executed, re-papered, booked -- and "the period has moved out" matches none of them, so it leaves the roll unspent and reports ROLLABLE. That is the costly direction: a producer told a roll is available when the contract's single permitted roll is already gone. The floor did this 5 times across the run and the model 0.
What one missed Monday costs a PERFECT reader
not in scope
This grader scores schedules, not readings -- it removes one Monday and asks what a PERFECT reader then gets wrong, so no individual arm's answer is inside its denominator. It is shown rather than omitted because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. What it WOULD say about this contract is that CT-2026-0009's balance is one of the ones poisoned by a missed 2026-09-14 run: the 600 bushels applied in this very window are printed on this record and nowhere else.
The formulaWhat it computes
accuracy = hits / 150 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
95.3% rollover risk accuracy · 3 more measured on this row
the fast tier, memory removed (THE CONTROL)
72.7% rollover risk accuracy · 3 more measured on this row
the strongest free floor, no model
95.3% rollover risk accuracy · 3 more measured on this row
the same free rule, GIVEN the carried state
71.3% rollover risk accuracy · 3 more measured on this row
the expiry report a grain desk already runs
47.3% rollover risk accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by tools/build_corpus.py over the planted events at generation time and re-derived from src/contract.step by evals/check_labels.py before any run may spend. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known to be arguable -- Rule C-8 says a BASIS contract with a fixed futures leg carries no roll right and says nothing about an HTA, whose price line reads almost identically. Six readings turn on that gap and the key assumes the rule does not reach an HTA. The other three fields have no known ambiguity.
Watch these
rollover_risk_accuracy_pct
status_accuracy_pct
unfilled_bu_accuracy_pct
raise_decision_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Both paid arms had exactly 1 of 150, both at the ceiling, and the two are different documents -- so 1 is a property of this ceiling on this workload rather than of either arm.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because two of them are small: 13 readings must raise a decision and 20 have nothing on file, so one row moves the first by 7.7 points and the second by 5.0.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/contract.py, src/schedule.py, src/segment.py or src/select.py. Re-run the scored eval AND its stateless control together (paid, 300 calls) on any change to src/prompt.py or src/state.py -- the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. All three floors and the whole cadence ablation are free and should be re-run on any change at all.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real elevator, where whether a load was turned away or discounted is decided by a person reading a scale clerk's note. That is why this corpus is generated rather than captured.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineThe two rollover-risk directions, counted apart
Two counts over two different populations. A ROLL OFFERED THAT WAS NOT AVAILABLE: the key says NOT_ROLLABLE and the reading says ROLLABLE. A ROLL MISSED: the key says ROLLABLE and the reading says NOT_ROLLABLE. This is the grader that separates the model from the free floor, because their headline rollover-risk accuracy is identical and their error direction is not.
$0.00per 1,000 open-contract records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in the same pass as the exact-match grader.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
CT-2026-0009-W2 -- scheduled run 2 of 5, Monday 2026-09-14, a 15,000 bu delayed-price soybean contract with a 2.0 pct tolerance, one roll permitted, and 600 bu applied gross at the scale this week
What last Monday left behind
As at the previous scheduled run, 13,700 bushels were still owed on this contract, before anything this window applied. No roll has been executed on this contract on any earlier run. The roll/cancel/deliver decision has ALREADY been put to the producer on an earlier run, so this run must not raise it a second time. It was last reported DECISION_DUE.
The activity log, this window only
the desk quoted the spread on Friday and the producer took it, the period has moved out / an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied
The answer key
ON_TRACK, rollover risk NOT_ROLLABLE, 13,100 bushels still owed, raise_decision NO
What run r001 answered
ON_TRACK, rollover risk NOT_ROLLABLE, 13,100 bushels still owed, raise_decision NO
What the free floor answered
ON_TRACK, rollover risk ROLLABLE, 13,200 bushels still owed, raise_decision NO
Why this one
It is the only reading in this corpus where BOTH of the generator's deliberate collisions fire in one window, and the free floor loses both. The roll line reads like a quote and is an execution; the ticket line reads like a rejection and is an application. Chosen by tools/pick_shot.py off the answer key -- memory-dependent first, then the costly error direction, then exposure -- rather than by hand.
Grader
Verdict
Why
The rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key
rollover risk hit, status hit, bushels hit, decision hit -- four of four
CT-2026-0009-W2, run 2 of 5 on a 15,000 bu delayed-price contract with a 2.0 pct tolerance and one roll permitted. 13,700 bushels were carried in and 600 came over the scale gross. Two log lines decide the reading and both are collisions: "the desk quoted the spread on Friday and the producer took it, the period has moved out" is a roll that was EXECUTED, spending the single permitted roll; "an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied" is an APPLICATION, not a turnaway. So 13,700 minus 600 is 13,100 owed, above the 300 bu tolerance, no roll remains, and the deadline is 19 business days out against a 5-day horizon: ON_TRACK, NOT_ROLLABLE, 13,100, no. The model answered all four correctly and its rationale names both facts.
The two rollover-risk directions, counted apart
in scope on the NOT_ROLLABLE side, and a hit
The key says NOT_ROLLABLE, so this reading sits inside the population the falsely-offered count is measured over. The model did not offer a roll. The free floor DID: its execution table keys on the desk's verbs for a completed roll -- amended, executed, re-papered, booked -- and "the period has moved out" matches none of them, so it leaves the roll unspent and reports ROLLABLE. That is the costly direction: a producer told a roll is available when the contract's single permitted roll is already gone. The floor did this 5 times across the run and the model 0.
What one missed Monday costs a PERFECT reader
not in scope
This grader scores schedules, not readings -- it removes one Monday and asks what a PERFECT reader then gets wrong, so no individual arm's answer is inside its denominator. It is shown rather than omitted because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. What it WOULD say about this contract is that CT-2026-0009's balance is one of the ones poisoned by a missed 2026-09-14 run: the 600 bushels applied in this very window are printed on this record and nowhere else.
The formulaWhat it computes
Two whole-number counts off the confusion matrix. Never averaged, never combined into an F-score, and never expressed as one 'rollover accuracy' figure -- that figure is the same for both arms and hides the entire finding.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
no headline metric on this row — it records false rollable 0 · lost rollable 6
the strongest free floor
no headline metric on this row — it records false rollable 5 · lost rollable 1
the fast tier, memory removed (THE CONTROL)
no headline metric on this row — it records false rollable 12 · lost rollable 6
the free rule GIVEN the carried state
no headline metric on this row — it records false rollable 21 · lost rollable 0
the expiry report a desk already runs
no headline metric on this row — it records false rollable 34 · lost rollable 0
In operationWhat to monitor
Reference standard: data/gold.jsonl's rollover_risk column, computed by src/contract.step. Scored against the reference; not itself the reference.
These rates are UNKNOWN, on purpose
Its own TPR and TNR are not measured because it does not answer a true/false question about another grader's verdict -- it counts two directions of one field against the key, so there is nothing here for an agreement rate to be about.
Watch these
false_rollable
lost_rollable
Alarm on
false_rollable above 0. A missed roll is visible to whoever reads the watchlist next week; an offer that was never available is visible to the producer first.
How tight can the band be? 46 readings are ROLLABLE in the key and 39 are NOT_ROLLABLE, so the finest band these counts support is 2.2 and 2.6 points respectively. They are published as counts rather than rates for exactly that reason.
Cadence: Free. Re-run with every scored run, and re-run all three floors whenever tools/build_corpus.py's roll templates change -- these two counts are the most template-sensitive figures on this page.
The decisionWhen to reach for it
Use it
The two error directions cost different things and are fixed by different people, which is exactly this kit's situation.
Do not use it
A task where one direction is free. Here neither is: an offer that has to be withdrawn damages a producer relationship, and a missed roll is an option the producer never heard about.
Check an open grain contract against its roll deadline
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one missed Monday costs a PERFECT reader
Remove one Monday from the SCHEDULE -- the record stays on disk and is simply never opened -- then score what a perfect reader concludes on the readings that still happen, against the INTACT answer key. It answers the one question a monitor cannot answer about itself: what does the cadence buy, and what is lost when it does not hold?
$0.00per 1,000 open-contract records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/cadence.py, in-process, no key and no model. It refuses to print unless a perfect reader scores 100.00 on the INTACT schedule first -- a self-check added after the first version chained all 150 readings into one contract and reported 34.17 pct as the intact baseline.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
CT-2026-0009-W2 -- scheduled run 2 of 5, Monday 2026-09-14, a 15,000 bu delayed-price soybean contract with a 2.0 pct tolerance, one roll permitted, and 600 bu applied gross at the scale this week
What last Monday left behind
As at the previous scheduled run, 13,700 bushels were still owed on this contract, before anything this window applied. No roll has been executed on this contract on any earlier run. The roll/cancel/deliver decision has ALREADY been put to the producer on an earlier run, so this run must not raise it a second time. It was last reported DECISION_DUE.
The activity log, this window only
the desk quoted the spread on Friday and the producer took it, the period has moved out / an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied
The answer key
ON_TRACK, rollover risk NOT_ROLLABLE, 13,100 bushels still owed, raise_decision NO
What run r001 answered
ON_TRACK, rollover risk NOT_ROLLABLE, 13,100 bushels still owed, raise_decision NO
What the free floor answered
ON_TRACK, rollover risk ROLLABLE, 13,200 bushels still owed, raise_decision NO
Why this one
It is the only reading in this corpus where BOTH of the generator's deliberate collisions fire in one window, and the free floor loses both. The roll line reads like a quote and is an execution; the ticket line reads like a rejection and is an application. Chosen by tools/pick_shot.py off the answer key -- memory-dependent first, then the costly error direction, then exposure -- rather than by hand.
Grader
Verdict
Why
The rollover risk, the status, the bushels still owed and the raise/hold call, per reading, exact match against the computed answer key
rollover risk hit, status hit, bushels hit, decision hit -- four of four
CT-2026-0009-W2, run 2 of 5 on a 15,000 bu delayed-price contract with a 2.0 pct tolerance and one roll permitted. 13,700 bushels were carried in and 600 came over the scale gross. Two log lines decide the reading and both are collisions: "the desk quoted the spread on Friday and the producer took it, the period has moved out" is a roll that was EXECUTED, spending the single permitted roll; "an earlier load from this producer was turned away last month, this one graded clean and 100 bu on ticket 42472 were applied" is an APPLICATION, not a turnaway. So 13,700 minus 600 is 13,100 owed, above the 300 bu tolerance, no roll remains, and the deadline is 19 business days out against a 5-day horizon: ON_TRACK, NOT_ROLLABLE, 13,100, no. The model answered all four correctly and its rationale names both facts.
The two rollover-risk directions, counted apart
in scope on the NOT_ROLLABLE side, and a hit
The key says NOT_ROLLABLE, so this reading sits inside the population the falsely-offered count is measured over. The model did not offer a roll. The free floor DID: its execution table keys on the desk's verbs for a completed roll -- amended, executed, re-papered, booked -- and "the period has moved out" matches none of them, so it leaves the roll unspent and reports ROLLABLE. That is the costly direction: a producer told a roll is available when the contract's single permitted roll is already gone. The floor did this 5 times across the run and the model 0.
What one missed Monday costs a PERFECT reader
not in scope
This grader scores schedules, not readings -- it removes one Monday and asks what a PERFECT reader then gets wrong, so no individual arm's answer is inside its denominator. It is shown rather than omitted because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. What it WOULD say about this contract is that CT-2026-0009's balance is one of the ones poisoned by a missed 2026-09-14 run: the 600 bushels applied in this very window are printed on this record and nowhere else.
The formulaWhat it computes
Same four exact-match comparisons, over the 120 readings that survive the reduced schedule, against the unreduced key. Plus two counts a rate cannot express: contracts whose roll deadline falls in the blind window (never warned on ANY schedule), and contracts whose carried balance is permanently wrong afterwards.
The analysisWhat it actually did
Model
Result
one missed Monday: 2026-09-14
77.5% rollover risk accuracy · 2 more measured on this row
one missed Monday: 2026-09-21
80.8% rollover risk accuracy · 2 more measured on this row
one missed Monday: 2026-09-28
90.0% rollover risk accuracy · 2 more measured on this row
the intact schedule, for comparison
100.0% rollover risk accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: the INTACT data/gold.jsonl. The truth about a contract does not change because nobody looked at it, so the reduced-schedule arm is scored against the unreduced key -- any other choice would measure the watch's own confusion as if it were the world.
These rates are UNKNOWN, on purpose
What a missed run costs the MODEL is not measured. It would take a fourth and fifth paid arm, and the perfect-reader figure is the ceiling on any of them -- a model cannot beat it, only add its own errors on top. That bound is why the free measurement was judged sufficient and the money was not spent.
Watch these
blind_window_contracts
contracts_with_a_poisoned_balance
cells_moved_direction
unfilled_bu_accuracy_pct
Alarm on
blind_window_contracts above 0 on any schedule. Every other number here degrades; this one is a permanent loss.
How tight can the band be? 30 contracts, so one contract in the blind window is 3.3 pct of the book. The three ablations put 4, 5 and 4 there. Across all three, 253 cells moved and 0 went the right way.
Cadence: Free, and it should be re-run on any change to src/schedule.py or to the corpus's deadline spread. It is the only grader here that measures the thing the kit cannot control.
The decisionWhen to reach for it
Use it
A monitor's answer depends on WHEN it ran, and nothing inside it records that a run did not happen.
Do not use it
A stateless task, or one where the source system carries a cumulative figure that lets a later reading reconstruct what was missed. Neither is true here.
A living map of modern AI — kept current every morning