Home › Use Cases › Warehouse Receipt Obligation Tracking Against Licensed Capacity
Use caseUC0089
🧪 Use-case kit · runnable
Warehouse Receipt Obligation Tracking Against Licensed Capacity
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A grain warehouse licence is not a capacity number. It is a capacity number plus a handful of sentences about TIME -- a first day over capacity is a notice, the third is a referral to the state for licence suspension, an open notice does not close until the facility has been back inside capacity for two days running. Two facilities filing exactly the same bushels against exactly the same capacity need different actions, because one of them was over yesterday and the other was not. Nothing in the day you are holding tells you which. And the clauses themselves differ facility by facility, including one -- whether company-owned grain counts toward the storage obligation -- that decides whether a facility is over capacity at all. So somebody opens every record, every day, and goes back through the filings by hand. Somebody opening every licensed facility's daily position record, working out from the facility's OWN licence which obligation total governs, and then going back through the previous filings to count how many consecutive business days it has been over -- because the day you are holding does not say.
Audience
Grain warehouse compliance and elevator operations -- anyone who files or reviews a daily position record against a state licence, and anyone deciding whether a language model has any business near one. The honest answer here is a qualified yes with one field carved out: on the days it answered, the model got every status and every action right and lost to free code on naming the largest obligation category. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual daily position records
The corpus is 150 daily position records, 0.48 MB (txt 150). Generated rather than collected, because a real daily position record names a licence, a bonded surety and a licensed manager who is personally liable under it -- publishing thirty facilities' worth of those, redacted or not, is not something a public kit gets to do. Generating it also buys the two things this kit is measuring: 25 of the 30 facilities carry licence terms other than the majority set, which is what makes reading the clauses worth anything, and the day patterns are chosen rather than sampled so that every clause fires somewhere -- 66 of 150 days turn on a consecutive-day count, 23 on the inclusion clause, and 10 sit after the middle day of a reset trap.
The corpus
The 150 daily position recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your daily position records. That is the whole change — there is no database to migrate.
One daily position record, as the model receives itWH-0417-D1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This is a GENERATED daily position record for a grain warehouse that does not exist. The
licence number, the facility, the people and every bushel figure on this page were produced
by tools/build_corpus.py from a fixed seed. It is not a filing and reproduces no real one.
Facility
----------------------------------------------------------------
Licence no : WH-0417
Facility : Prairie Bend Grain Co-operative
Business day : 2026-08-17
Filed by : the licensed manager, at close of business
Licence Terms
----------------------------------------------------------------
Licensed capacity : 750,000 bu
Obligation basis : storage obligations INCLUDE company-owned grain held in store
Suspension referral : refer the licence for suspension after 3 consecutive
business days over licensed capacity
Notice release : release an open capacity notice after 2 consecutive
business days back within licensed capacity
The licence requires, on each business day and in this order:
1. On the FIRST business day storage obligations exceed licensed capacity, issue a
capacity notice. A notice stays open until it is released.
2. Once a notice is open, refer the licence for suspension on the day the count of
consecutive business days over capacity reaches the figure above. Refer once.
3. While a notice is open and obligations are back within capacity, hold the notice
pending release until the consecutive-day figure above is reached, then release it.
4. Otherwise take no action. One action per facility per business day.
Abridged — the file continues.
The outcomeWhat a good result looks like
Per facility per business day: the status to report after the licence's clauses have been applied, the single action the licence requires today, and the largest obligation category counted under that licence -- plus four carried counters the next run is judged against, written by code from the parsed totals and never from the model's reply.
And when it cannot
It answers with a status and an action that are wrong, in a well-formed JSON object, and nothing raises an error. There is no confidence score and no abstention: the scorer's 'absent' verdict exists for a call that never returned, and nothing downstream reads it. On this run the failure that actually happened was not a wrong answer at all -- it was nine calls the provider refused, which the harness recorded by name and the scorer counted as misses.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Your licence terms are the same at every facility and never change — b001, the coded rule -- free The whole gap between the free rule and the model on this corpus is the three varying clauses. With one set of terms, b001 IS src/licence.py and scores what the answer key scores.
Terms vary by facility, by state or by licence renewal — the model arm, r001's configuration 94.0 pct against 77.33 pct on action accuracy over all 150 days, and 95.65 pct against 4.35 pct on the days the inclusion clause inverts the verdict -- which is the failure that produces a wrong answer rather than a late one.
You only need to know which obligation category is largest — a free argmax over the printed table 98.58 pct for nothing against 97.87 pct for a model call, on the same 141 days -- and the STATELESS arm of this very kit scores 99.29 pct, better than both. Removing the carried state IMPROVES this field, which is the sharpest statement available that it does not want a model or a memory.
You cannot guarantee the job runs every business day — fix the schedule first, then this kit MEASURED TWICE, AND THE MODEL IS THE WORSE OF THE TWO. One skipped business day out of five costs the free coded rule 17.5 points of action accuracy (a000 against b001, 37 of 47 moved cells right to wrong); it costs the MODEL 49.06 points -- 100.0 pct to 50.94 pct over the 53 days both arms answered, with 43 of 43 moved cells right to wrong and NOT ONE the other way (a001 against r001). A model does not absorb a broken clock, it compounds it, and no prompt change recovers a day nobody read.
At a glanceHow the whole thing runs
94%action accuracy pct
7,774 msp50, end to end
$3.94per 1,000 daily position records · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Warehouse Receipt Obligation Tracking Against Licensed Capacity14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Drop your own records into data/corpus/ named <FACILITY>-D<n>.txt in business-day order and write data/gold.jsonl with the same keys. The moment your records carry a carry-in balance, this kit's central measurement stops transferring.Corpus lens →
When is this the wrong choice?
Avoid: Do not pay for a model to re-read a rulebook you already encoded. On a single-terms population this kit's measured advantage is zero by construction. That is the case against the best-fitting scenario (“Your licence terms are the same at every facility and never change”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A record with no Licence Terms section -- a facility operating under a blanket state licence, which is entirely ordinary. Every field's hints match sections that all 150 records carry, so today the selector's fallback is never reached; src/select.py's subtraction is unconditional precisely so that when it IS reached the licensed manager's name does not go out with it. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER THE THREE WRONG CATEGORY NAMES WOULD REPRODUCE. There is no repeat probe on this kit -- no configuration ran twice -- so 3 of 150 is one observation and not a rate. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-grain-position. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured reproduces every figure on this page that did not need a model: the pre-flight and both of its red-proofs, both free floors, the cadence ablation, and the two comparisons that read the committed results back -- five of the seven recorded runs. The corpus and the answer key are regenerated byte-for-byte by python3 -m tools.build_corpus. What it cannot reproduce without a key is r001-grain-position and the calibration.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
94.0%rows answered
7,774 msp50, end to end
23,167 msp95
2 minclone to first result
What the clock covers. END-TO-END per facility-day: one HTTP request carrying the assembled prompt, and the reply parsed to a status, an action and an obligation category. Splitting the record into sections, dropping the section no field asks for, parsing the two subtotals and the three licence clauses, and advancing the carried counters all happen outside this clock and cost no network at all. THE TAIL IS THE STORY: p95 (23,167 ms) is 3.0 times p50 (7,774 ms), and the slowest single call took 53.4 seconds. Output length is driven by how many licence clauses are in play, not by how long the record is -- every record in this corpus is within 29 bytes of every other.
Current processWhat it replaces
Somebody opening every licensed facility's daily position record, working out from the facility's OWN licence which obligation total governs, and then going back through the previous filings to count how many consecutive business days it has been over -- because the day you are holding does not say.
Where it is not good enough
⚑ NINE OF THE 150 SCORED CALLS NEVER HAPPENED, AND THAT IS THE FIRST THING ON THIS PAGE RATHER THAN A FOOTNOTE. Six were refused with 402 Insufficient Balance on the shared provider account and three lost a 429 concurrency limit that outlasted four backoffs. The scorer counts an unanswered day as wrong, which is the right default for a headline -- a monitor that does not answer has not answered -- so every figure over 150 days is depressed by exactly those nine. THE HEADLINE IS 94.0 PCT ACTION ACCURACY OVER ALL 150 LABELLED DAYS. On the 141 days it answered it is 100.0 pct, and the action confusion matrix is a clean diagonal: there is no wrong action anywhere in this run. Both scopes are published because neither answers the other's question.
⚑ THE FREE FLOORS BEAT THE MODEL ON ONE OF THE THREE FIELDS, AND SO DOES ITS OWN CONTROL. Naming the largest counted obligation category is 98.58 pct for a free argmax over the printed table, 99.29 pct for the STATELESS arm, and 97.87 pct for the full kit -- on the same 141 days. Over all 150 the model reads 92.0 pct against the floor's 98.67 pct. That field does not need a model, and the fact that REMOVING memory improves it is the sharper version of the same finding: on a pure reading task the carried state is a distraction that costs about a point and a half.
⚑ AND A MISSED RUN COSTS THE MODEL NEARLY THREE TIMES WHAT IT COSTS FREE CODE. Skip one business day out of five and the free coded rule loses 17.5 points of action accuracy; the model loses 49.06 -- 100.0 pct to 50.94 pct over the 53 days both arms answered, with 43 of 43 moved cells going right to wrong and NOT ONE going the other way. The model does not absorb a broken clock, it compounds it, and that is the single most important number on this page for anyone deciding whether to run this on a schedule they do not control.
⚠︎ THE CORPUS REMOVES A ROUTE REAL RECORDS LEAVE OPEN. A real daily position record carries a carry-in balance, from which yesterday's total -- and therefore whether yesterday was over capacity -- is recoverable by subtraction. These records do not, deliberately, so that the carried state is the only history available. That makes this a HARDER corpus than a real one for a stateless reader, and no figure here transfers to a deployment reading real filings until that has been measured.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1
30 licensed facilities, 5 consecutive business days each
It produces a position status and one action for a compliance officer to read, and refers nothing to anybody — SUSPENSION_REFERRAL is a string in a JSON reply and a column in a result file. TWO STATIONS MAKE THIS A MONITOR AND MOST KITS ONLY BUILD ONE. The second is what the last run left behind: four counters, written by src/licence.py from the totals parsed off the page and never from the model's answer, so one wrong day costs one day. Remove that block and the same prompt scores 56.74 pct instead of 100 — and 0 of 61 on the days a consecutive-day count decides the answer, not one. The THIRD station is the one that keeps getting skipped: those counters count RUNS, and they equal business days only while something outside this repository fires the job every business day. ⚑ SO IT WAS PRICED, TWICE. Remove business day 3 from the SCHEDULE — the record still exists and is simply never read — and the free coded rule falls 80.0 to 62.5 pct over the 120 days it covers, while the MODEL falls 100.0 to 50.94 pct over the 53 days both arms answered, with 43 of 159 moved cells going right to wrong and not one the other way. The cadence guarantee is worth more here than the memory it protects. The kit detects the gap afterwards for nothing (src/schedule.gap(), two dates and a weekday test, flagged 30 of those days) and repairs none of them, because a position for a day nobody read is not recoverable from a later close-of-business snapshot. ⚑ WHAT THE MODEL IS PAID FOR IS THE RULEBOOK, NOT THE ARITHMETIC. Three clauses differ facility by facility and all three are printed on the record: how many consecutive days over capacity trigger a referral, how many days back inside release a notice, and whether company-owned grain counts toward the storage obligation. Only 5 of the 30 facilities carry the majority terms. A free implementation of the same rules with those terms hardcoded scores 77.3 pct and 4.55 pct on the days the inclusion clause inverts the verdict, against 100 and 100 — and that clause does not shift a deadline, it inverts the answer. Notably the STATELESS arm still scores 86.36 pct on it, because reading a clause needs no memory: the two mechanisms are separable and each arm loses a different one.
⚠︎ NINE CALLS NEVER HAPPENED, WHICH IS WHY TWO NUMBERS ARE ON THE REPORT. Six were refused with 402 Insufficient Balance on the shared account and three lost a concurrency limit; the scorer counts an unanswered day as wrong. On the 141 that returned the action confusion matrix is a clean diagonal — no wrong action anywhere in the run — and a zero on 141 days is a claim about those 141 days, not the population.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL and MODEL in .env. Adding a vendor is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the licence clauses
src/licence.py
step() holds the four clauses and their order. Your jurisdiction's licence is a different step(); everything above it -- the counters, the cadence, the scorer, the prompt -- is unchanged.
what is never sent
src/select.py
NEVER_SENT and SECTION_HINTS. A section no field maps to does not leave the machine, and the subtraction is unconditional rather than a fallback.
what is carried between runs
src/state.py
describe() renders the counters into English. Four scalars is the whole contract; anything that grows with history turns this into a kit whose input cost climbs every run.
the cadence
src/schedule.py
WEEKEND and business_days_between. A jurisdiction with holidays, or a facility filing weekly rather than daily, changes this file and nothing else -- but see breaks_at_scale before believing that.
the corpus
tools/build_corpus.py
Point data/corpus/ at your own records, or regenerate with different patterns and licence terms. The record format is asserted by evals/check_labels.py, so a mismatch refuses to run rather than truncating a prompt.
Components
Component
File
Role
the licence
src/licence.py
the three varying clauses and the day-count arithmetic they imply; the answer key is this module replayed
the carried state
src/state.py
four counters per facility, written from the parsed totals and never from the model's reply
the cadence
src/schedule.py
which business day a run owns, and free detection of a run that did not happen
the section splitter
src/segment.py
one record into its seven named sections, pure code
the selector
src/select.py
which sections reach the provider; Facility Contacts is mapped by nothing and therefore never sent
the prompt
src/prompt.py
instruction, carried state, selected sections -- three parts, string concatenation, no template engine
the model call
src/adapters/__init__.py
one completion per facility-day, over stdlib urllib
the scorer
evals/scoring.py
exact match per cell against the computed gold, plus five slices; no model, no judge
the pre-flight
evals/check_labels.py
everything that must be true before a run may spend, including two red-proofs
Where it breaks at scale
⚑ THE SCHEDULER IS OUTSIDE THE KIT AND THAT IS THE FIRST THING THAT BREAKS. This is a monitor that carries state and is still INVOKED, not woken. Cron, Airflow or a person decides when it runs, and every consecutive-day count in src/licence.py is a count of RUNS -- it equals a count of business days only while that outside thing keeps its promise. Measured, not asserted: one skipped business day out of five moves 47 of 360 cells for a system that cannot make a mistake, 37 of them from right to wrong, and drops action accuracy from 80.0 pct to 62.5 pct (evals/cadence.py, a000 against b001). The kit detects the gap afterwards for nothing -- src/schedule.gap() flagged 30 of 120 days -- and detecting is not recovering: a position for a day nobody read cannot be reconstructed from a later record, so the kit refuses to backfill.
THE STATE STORE IS ONE FILE. data/state.json is replaced atomically, which is correct for one writer and is not a concurrency model. Two schedulers advancing the same facility would race and the loser's day would vanish -- which, in this kit, silently un-opens a capacity notice and restarts a referral clock at zero.
WALL CLOCK IS BOUNDED BY THE CHAIN, NOT BY WIDTH. Facilities are independent and ran 30 chains wide; the days inside one facility are strictly serial because day 3's prompt contains counters produced by day 2. A year of history is a queue of 250 calls deep that no amount of concurrency shortens. What does NOT grow is the input: four counters cost the same on business day 400 as on day 4.
AND THE PROVIDER IS A REAL CEILING, WHICH THIS RUN HIT. 3 of 150 calls lost a 429 concurrency limit after four exponential backoffs at 12 workers, and 6 more were refused outright when the shared account balance ran out mid-run. Neither is a property of the design and both are properties of running it, and a monitor on a clock has no operator watching to notice.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
The pre-flight every run must pass before it is allowed to spend: 150 records, 30 gap-free facility chains, the answer key re-derived from src/licence.replay, no status token or day count anywhere outside the rulebook -- and the rulebook proven to be ONE template across all 150, so exempting it from that scan is earned rather than assumed. THREE guards are RED-PROVEN in both directions: the privacy selector under a hint that matches nothing (0 leak with the guard, 150 of 150 without), the cadence detector (0 on an intact schedule, exactly 1 on a skipped business day, 0 across a weekend), and the chain resume (0 of 180 cells disagree -- the first version of that flag disagreed on 21 and this line is what caught it, before any money was spent on the ablation it enables).successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
Five arms, one grader, printed at both scopes -- and the kit losing a field to its own control. On the 141 days every arm answered the model is 100 pct on status and action; on naming the largest counted obligation category it is 97.87 pct, against 98.58 pct for a free argmax over the printed table and 99.29 pct for the SAME MODEL with the carried state removed. Removing memory improves that field. The row that matters most is the control's 0.0 pct in the clock-dependent column: 0 of 66.failureOpen full size →What a missed run costs, priced twice on the same perturbation -- business day 3 removed from the schedule, the record still on disk and simply never read. WITHOUT a model the free coded rule falls 80.0 -> 62.5 pct action accuracy over the 120 days it covers. WITH the model it falls 100.0 -> 50.94 pct over the 53 days both arms answered, and 43 of 43 moved cells go right to wrong with NOT ONE going the other way. A model compounds a broken clock rather than absorbing it. The kit detected the gap for free on 30 of those days, which is not the same as recovering it.failureOpen full size →The scored run started with no API_KEY configured: it refuses before any call is made and points at the free stub, rather than failing at the HTTP layer or spending to find out. Exit code 1, nothing billed, no result file written. This is the honest failure state, captured rather than staged.failureOpen full size →
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150daily position records
0.48 MiBtxt 150
p50 3,345chars per daily position record
$0.00setup · 0s
How it is cutWhat one daily position record is
none -- one daily position record is one unit and goes whole into one call, minus the one section no field asks for
SetupWhat the setup figure measured
There is no index and nothing was built. A monitor does not search its population, it RE-READS it: every open facility's record goes whole into one call on every scheduled run, whether or not anything about it changed. The zeros are the honest value of a step that does not exist, not a fast index.
LicenceLicence
MIT, with the rest of the kits repository. The corpus carries no third-party material.
Bring your ownBring your own daily position records
Drop your own records into data/corpus/ named <FACILITY>-D<n>.txt in business-day order and write data/gold.jsonl with the same keys. If your record format differs, src/segment.SECTIONS and src/position.reading_of are the two files to change, and evals/check_labels.py will refuse to let a run spend until the split, the sequence and the leak scan all pass -- so a format mismatch is a refusal, not a quietly truncated prompt. Your licence's clauses go in src/licence.step; everything above it is unchanged.
⚠︎ And what stops being true when you do: The moment your records carry a carry-in balance, this kit's central measurement stops transferring. Every figure here is measured on a corpus where the carried state is the ONLY route from one run to the next, and a real filing leaves a second route open by arithmetic. Nothing on this page says how much of the work the carried state is doing once that route exists.
What breaks it
A record with no Licence Terms section -- a facility operating under a blanket state licence, which is entirely ordinary. Every field's hints match sections that all 150 records carry, so today the selector's fallback is never reached; src/select.py's subtraction is unconditional precisely so that when it IS reached the licensed manager's name does not go out with it. evals/check_labels.py red-proves both directions.
A record that carries a CARRY-IN BALANCE. Real filings do, and yesterday's total -- therefore whether yesterday was over capacity -- is then recoverable by subtraction. This corpus removes it deliberately, so nothing measured here says what a stateless reader could do on a real record.
A licence clause this kit does not model: a cure period counted in CALENDAR days rather than business days, a referral that resets on a bond increase, a capacity that varies by commodity or by bin. src/licence.step holds four clauses in a fixed order and anything outside them is not read at all.
A public holiday. src/schedule.py treats Saturday and Sunday as non-business days and models no holiday calendar, so a jurisdiction's holiday reads as a missed run -- which, given what a missed run costs here, is the wrong failure to have.
More than one action due on one facility on one day. The record has one action column and src/licence.step returns one action, so a licence whose clauses can fire together needs a different return shape, not a different prompt.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
124
not measured
instruction
1,354
not measured
carried state
134
not measured
Synthetic Record
353
not measured
Facility
239
not measured
Licence Terms
1,545
not measured
Position, as of close of business
494
not measured
Bin and Receipt Movement
292
not measured
Manager Commentary
176
not measured
Total
1,038
This is the cost lesson as arithmetic: of the 4,711 characters assembled, 1,545 are rulebooks — 33% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.build() for record WH-0447-D4 carrying the counters run r001 actually recorded for it (carried_in in results/eval-r001-grain-position.json), rather than logged verbatim by the run. That is only trustworthy because the replayed assembly is deterministic and the run recorded the total prompt length it sent -- 4,816 characters for the user message against the 4,769 replayed here plus the 124-character system message, and the difference is the system message not being counted in the run's own figure. The worked example is a SUSPENSION_REFERRAL day on a facility whose licence excludes company-owned grain: the model's own rationale cites both the exclusion clause and the day count, which is the whole job in one sentence.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You apply a written grain warehouse licence to one day's position record. You answer with one JSON object and no other text.
You are reviewing one licensed grain warehouse's daily position record for one business day.
The facility's own licence terms, its thresholds and the actions the licence requires are reproduced
in the record below. Apply them exactly as written. THE TERMS ARE NOT THE SAME AT EVERY FACILITY --
read this facility's own numbers rather than what is usual.
The licence's actions depend on how many CONSECUTIVE BUSINESS DAYS a condition has held. The record
covers one day only and states nothing about any earlier day. What is known about this facility's
history is stated under "Carried state" and is the only history available to you. Do not assume
anything about earlier business days beyond it.
Answer with a single JSON object and nothing else:
{"position_status": "WITHIN|OVER|UNDER_NOTICE|UNDER_REFERRAL",
"action": "NONE|CAPACITY_NOTICE|SUSPENSION_REFERRAL|HOLD_PENDING_RELEASE|NOTICE_RELEASED",
"top_obligation": "the largest obligation category COUNTED under this licence, named exactly as written in the record",
"rationale": "one sentence, naming the licence clause or the figure you applied"}
"position_status" is the status to report for this facility at close of business today, after the
licence's actions have been applied. "action" is the single action the licence requires today, or
NONE. One action per facility per business day.
Carried state
----------------------------------------------------------------
A capacity notice is OPEN against this facility. As of the last run it had been over licensed capacity on 2 consecutive business days.
Daily position record
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This is a GENERATED daily position record for a grain warehouse that does not exist. The
licence number, the facility, the people and every bushel figure on this page were produced
by tools/build_corpus.py from a fixed seed. It is not a filing and reproduces no real one.
Facility
----------------------------------------------------------------
Licence no : WH-0447
Facility : Turkey Creek Grain Co-operative
Business day : 2026-08-20
Filed by : the licensed manager, at close of business
Licence Terms
----------------------------------------------------------------
Licensed capacity : 620,000 bu
Obligation basis : storage obligations EXCLUDE company-owned grain held in store
Suspension referral : refer the licence for suspension after 3 consecutive
business days over licensed capacity
Notice release : release an open capacity notice after 2 consecutive
business days back within licensed capacity
The licence requires, on each business day and in this order:
1. On the FIRST business day storage obligations exceed licensed capacity, issue a
capacity notice. A notice stays open until it is released.
2. Once a notice is open, refer the licence for suspension on the day the count of
consecutive business days over capacity reaches the figure above. Refer once.
3. While a notice is open and obligations are back within capacity, hold the notice
pending release until the consecutive-day figure above is reached, then release it.
4. Otherwise take no action. One action per facility per business day.
Report exactly one status, testing these in order and taking the FIRST that applies:
UNDER_REFERRAL the licence has been referred for suspension under a notice that has
not since been released
OVER storage obligations exceed licensed capacity today
UNDER_NOTICE a capacity notice is open and obligations are within capacity today
WITHIN none of the above
Position, as of close of business
----------------------------------------------------------------
Obligation category Bushels
Warehouse receipts outstanding 207,500 bu
Open storage (no receipt issued) 195,000 bu
Delayed-price contracts 292,292 bu
Company-owned grain in store 55,500 bu
Total, all lines 750,292 bu
Total excluding company-owned grain 694,792 bu
Bin and Receipt Movement
----------------------------------------------------------------
Scale tickets posted today : 21
Warehouse receipts issued today : 5
Warehouse receipts cancelled : 2
Delayed-price contracts priced : 1
(counts only; this section carries no bushel figures)
Manager Commentary
----------------------------------------------------------------
Rail cars were short again; we are holding more than we would like until the next placement.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"position_status":"WITHIN","action":"NONE","top_obligation":"Open storage (no receipt issued)","rationale":"Total storage obligations of 594,612 bu are below licensed capacity of 750,000 bu and no capacity notice is open, so no action is required."}
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Warehouse Receipt Obligation Tracking Against Licensed Capacity — 150 daily position records. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
No LLM-as-judge, and that is not a shortcut. Every field is a closed set, so correctness is a matter of equality rather than of reading, and asking a model to grade NONE == NONE would add cost, variance and a second thing to be wrong.
150daily position records
150source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED141 · 141 · 116 · 69 · 75 · 84 · 33 / 150action accuracy pct — facility-days, ALL 150 LABELLED -- the 9 the provider refused count as wrongDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED141 / 150status accuracy pct — facility-days, all 150 labelledDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED61 · 44 · 0 · 18 / 66consecutive day rule accuracy pct — facility-days whose action turns on a count the record does not stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 · 1 · 20 / 23exclusion trap accuracy pct — facility-days where the licence's inclusion clause inverts the verdictDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED9 / 10reset accuracy pct — facility-days after the middle day of a reset-trap facilityDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED61 / 67action when one is due pct — facility-days where the licence requires an action -- 6 missedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED3 · 27 / 83false action rate pct — quiet facility-days -- 3 actions raised that the licence does not ask for, all three of them unanswered callsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED138 · 149 / 150top obligation accuracy pct — facility-days -- THE FIELD A FREE RULE WINS, 98.67 pct for nothingDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 141unparsed replies — calls that returned -- every reply parsed; largest 6,405 output tokens under a 34,000 capDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED14 / 150calls refused by the provider — scheduled calls -- 6 refused for an exhausted account balance, 3 lost a concurrency limit after four backoffsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The grader is exact string equality over three closed vocabularies -- four statuses, five actions, and four obligation category names copied from the record's own table. There is nothing to validate about equality; what needed validating was the ANSWER KEY, and it is re-derived from the records themselves by evals/check_labels.py (150 of 150) rather than trusted from the file the generator wrote. No model grades anything in this kit.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One daily position record
1,000 daily position records
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.003935
$3.94
13%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.001574
$1.57
13%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.067315
$67.32
15%
Same work, 43× the bill
The same daily position records, the same tokens — only the rate card changed. And across all 3 cards between 13% and 15% of what you pay is the prompt this pipeline sends, not the answer it writes.
provider-side reasoning -- 93.35 pct of this run's output tokens, and output is 87 pct of the projected bill on the shared card. Nothing else on this page is close, and this kit has measured neither what disabling it costs in accuracy nor what a lower ceiling would have saved.
Rates checked 2026-08-18. The provider that actually ran the calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Free means free: evals/scoring.py, both floors, the cadence ablation and the two comparisons make no network call of any kind. The only paid thing on this page is the answer being graded.
The gradersThree ways to grade
⚑ b001 IS THE NUMBER TO BEAT AND PUBLISHING ONLY b000 WOULD HAVE FLATTERED THIS KIT. A memoryless threshold at 46.0 pct action accuracy is a straw man: nobody runs a compliance monitor with no memory. What a team actually runs is a scheduled job that implements the licence -- once, against one facility's terms -- and that is b001 at 77.33 pct. The model's 94.0 pct over all 150 days is measured against THAT, and the whole of the gap is the three clauses b001 cannot see: it scores 4.35 pct on the 23 days the inclusion clause inverts, against the model's 95.65 pct. ⚠︎ AND THE FLOORS ARE NOT THE MEMORY CONTROL. That is s001-grain-position-stateless, the identical prompt with the carried-state block replaced by one line; it is a separate arm and it is what the memory claim rests on.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Action, exact match whether the single action the licence requires today was named exactly
$0.00
no
yes
the fast tier, on a schedule that held 94.0% · the fast tier, memory removed (THE CONTROL) 56.0% · free coded rule, usual terms hardcoded 77.3% · free four-column threshold, no memory 46.0% · the wiring stub, no key and no model 55.3% · the fast tier, ONE MISSED RUN (60 days, not 150) 55.0%
the fast tier, on a schedule that held 94.0% · the fast tier, memory removed (THE CONTROL) 76.7% · free coded rule, usual terms hardcoded 74.7% · free four-column threshold, no memory 63.3% · the wiring stub, no key and no model 63.3%
the fast tier, on a schedule that held 92.0% · the fast tier, memory removed (THE CONTROL) 99.3% · free coded rule, usual terms hardcoded 98.7% · free four-column threshold, no memory 98.7% · the wiring stub, no key and no model 98.7%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
The three graders are not independent and this set does not pretend otherwise: a status miss and an action miss on the same facility-day are usually one mistake counted twice. What the set DOES separate, and what it was built to separate, is the two mechanisms -- and the control settles it rather than an argument. Remove the carried state and the 66 clock-dependent days fall to 0.0 pct while the 23 exclusion-trap days stay at 86.96 pct: the CLOCK half collapses completely and the READING half barely moves, because the licence clause is on the page and the day count is not. The free floors sit on the opposite diagonal -- 66.67 pct on the clock days and 4.35 pct on the traps. Four arms, two mechanisms, and each arm loses a different one. ⚠︎ It still cannot tell two MODELS apart: one tier was run.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Your licence terms are the same at every facility and never change
b001, the coded rule -- free
The whole gap between the free rule and the model on this corpus is the three varying clauses. With one set of terms, b001 IS src/licence.py and scores what the answer key scores.
Do not pay for a model to re-read a rulebook you already encoded. On a single-terms population this kit's measured advantage is zero by construction.
Terms vary by facility, by state or by licence renewal
the model arm, r001's configuration
94.0 pct against 77.33 pct on action accuracy over all 150 days, and 95.65 pct against 4.35 pct on the days the inclusion clause inverts the verdict -- which is the failure that produces a wrong answer rather than a late one.
Do not run it without the pre-flight. Every guarantee on this page -- the privacy guard, the leak-free corpus, the gap-free chains -- is asserted by evals/check_labels.py before a run spends, and none of it is checked at runtime.
You only need to know which obligation category is largest
a free argmax over the printed table
98.58 pct for nothing against 97.87 pct for a model call, on the same 141 days -- and the STATELESS arm of this very kit scores 99.29 pct, better than both. Removing the carried state IMPROVES this field, which is the sharpest statement available that it does not want a model or a memory.
Do not fold it into a headline accuracy. Averaged in, it hides the one line item a buyer could delete.
You cannot guarantee the job runs every business day
fix the schedule first, then this kit
MEASURED TWICE, AND THE MODEL IS THE WORSE OF THE TWO. One skipped business day out of five costs the free coded rule 17.5 points of action accuracy (a000 against b001, 37 of 47 moved cells right to wrong); it costs the MODEL 49.06 points -- 100.0 pct to 50.94 pct over the 53 days both arms answered, with 43 of 43 moved cells right to wrong and NOT ONE the other way (a001 against r001). A model does not absorb a broken clock, it compounds it, and no prompt change recovers a day nobody read.
Do not let src/schedule.gap()'s free detection stand in for a guarantee. It found 30 of 120 affected days and could not repair one of them.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
PROVIDER_REFUSED
the call never happened
9
WH-0486-D5 -- 402 Payment Required: {"error":{"message":"Insufficient Balance"}}. Six of these; the shared provider account ran out of credit part-way through the run. Three more are 429 Too Many Requests after four exponential backoffs at 12 workers…
WRONG_CATEGORY
the largest obligation category named wrongly
3
WH-0423-D1 -- gold Warehouse receipts outstanding, model Delayed-price contracts. All three misses of this kind are the same direction, and the free argmax gets all three right. This is the only category of WRONG ANSWER in the entire run.
CLOCK_BLIND
a day count the model could not know
66
WITH the carried state, 61 of 66. WITHOUT it, 0 of 66 -- the control does not get one of them right, and its rationales say why: it reports the position the record shows and answers NONE or a first-day notice, because a first day over capacity is the only…
STALE_CLOCK
right answer, one business day out of date
43
Skip business day 3 and 43 of the 43 cells that move go from right to WRONG, and not one moves the other way. Action accuracy over the 53 days both arms answered falls 100.0 -> 50.94 pct. The free coded rule loses 17.5 points on the same perturbation; the…
NO_WRONG_ACTION
actions answered wrongly
0
There are none, and the confusion matrix says so rather than the prose: `[{gold NONE, got NONE, n 80}, {CAPACITY_NOTICE, CAPACITY_NOTICE, 25}, {NOTICE_RELEASED, NOTICE_RELEASED, 12}, {HOLD_PENDING_RELEASE, HOLD_PENDING_RELEASE, 12}, {SUSPENSION_REFERRAL…
What we could NOT verify
WHETHER THE THREE WRONG CATEGORY NAMES WOULD REPRODUCE. There is no repeat probe on this kit -- no configuration ran twice -- so 3 of 150 is one observation and not a rate. The same caution applies to the control BEATING the scored run on that field by 1.42 points on the shared scope: two runs of two different prompts, not two runs of one.
WHETHER THE MODEL'S 100 PCT ON THE ANSWERED DAYS SURVIVES A HARDER CORPUS. Every licence clause here resolves inside five business days and every record states its own terms in one block. A clause spanning a longer window, or terms held in a separate licence file the record only references, are both ordinary and neither is in this set.
WHETHER A MISSED RUN THE KIT ANNOUNCES COSTS LESS THAN ONE IT DOES NOT. a001 is SILENT by design -- a scheduler that does not fire does not announce anything -- so the measured damage is the naive case. src/schedule.gap() can detect the gap for free and src/state.describe could carry a warning; whether telling the model would recover any of the 43 lost cells is 60 more calls and was not paid for.
WHETHER RENDERING THE CARRIED STATE AS ENGLISH BEATS RENDERING IT AS JSON. Two more 150-call runs; not paid for.
WHAT DISABLING PROVIDER-SIDE REASONING WOULD DO. 93.35 pct of the scored run's output tokens were reasoning left at the provider's default. src/adapters can send a thinking field and this kit never has, so the accuracy cost of turning it off is unknown and so is the saving.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,038.42
1,138.62
7,774 ms
$0.003935
$0.001574
$0.067315
the same tier, memory removed (THE CONTROL)
1,023.49
712.53
6,072 ms
$0.002649
$0.001060
$0.045861
the strong free floor -- the same licence arithmetic, majority terms hardcoded
0
0
0 ms
$0.000000
$0.000000
$0.000000
the memoryless free floor -- today's four-column sum against capacity
0
0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
⚑ FIVE OF THE SEVEN RECORDED RUNS ON THIS PAGE COST NOTHING. The wiring stub, both rules floors, the cadence ablation and the two comparisons that read the committed results back are pure code and make no network call of any kind, and the pre-flight with both of its red-proofs is free too. Only r001 and the calibration were paid for, and the scored run's 9 refused calls were billed for nothing because they returned no completion. ⚠︎ THE MEMORY CONTROL AND THE MODEL CADENCE ABLATION -- 210 further calls -- WERE NOT PAID FOR AT ALL, because the shared account ran out; that is a hole in the evidence and not a saving. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING, 93.35 pct of output tokens (149,872 of 160,545) and therefore most of the bill. It was left at the provider's default and this kit has not measured what turning it off does to the answers.
HOW MANY LICENCE CLAUSES ARE IN PLAY, NOT HOW LONG THE RECORD IS. Every record is between 3,329 and 3,358 bytes -- a 29-byte spread -- and output tokens ran from 347 to 6,405 on the scored run and from 1,085 to 16,780 on the calibration. A quiet day inside capacity resolves in one clause; a day with an open notice, a referral clock one short and an exclusion clause is three deep.
THE FIXED HEAD OF THE PROMPT. The system message, the instruction and the licence rulebook are 3,023 of 4,893 characters -- 62 pct -- and the rulebook is one template across all 150 records. The day's own numbers are the smaller half of what is sent.
THE NUMBER OF FACILITIES, LINEARLY, AND NOT THE LENGTH OF THEIR HISTORY. The carried state is four scalars, so business day 400 costs exactly what business day 4 costs -- the opposite curve to a conversation kit, whose input grows with the square of the turns.
Your volumeWhat it costs at your volume
Linear in FACILITY-DAYS and flat in history length. Each day is one call whose input is one record plus four carried counters, so 300 facilities cost ten times 30 and a facility on its four-hundredth business day costs what one on its fourth costs. 141 calls project to $0.5548 on the shared rate card, so ten times the population is about $5.55 a day. What does NOT scale is wall clock inside a facility: its days are strictly serial, so a long backfill is a long queue no width can shorten -- and the provider's own concurrency limit is a real ceiling, which this run hit three times at 12 workers.
Where pricing changes shape
THE OUTPUT CEILING IS A CLIFF AND A CEILING IS NOT BILLED. The provider charges tokens produced, not tokens allowed, so the calls that finish in a few hundred tokens pay nothing for the headroom the hard ones need -- but a call that hits the ceiling is billed in full and returns nothing at all. The calibration measured a 16,780-token maximum at a 24,000 cap; the published ceiling is 34,000 and the scored run's largest reply was 6,405, 19 pct of it.
THE ACCOUNT BALANCE IS A CLIFF AND IT IS NOT A GRADUAL ONE. This run met it: six calls returned 402 Insufficient Balance and every subsequent call on that key failed identically. A monitor on a clock has nobody watching, so the failure mode is a compliance report that silently stops being produced -- which is worse than a wrong one.
CONCURRENCY IS A CLIFF BEFORE MONEY IS. Three calls exhausted four exponential backoffs against a 429 at 12 workers. Widening the pool to shorten the wall clock buys refusals, not throughput.
Your return, with your numbers
Volumefacility-days per business day -- this run judged 150 (30 facilities x 5 business days) in one pass per arm; a real elevator group runs its whole population once every business day, every business day
What it replacessomebody opening each facility's daily position record, working out from that facility's own licence which obligation total governs, and going back through the previous filings to count consecutive days
Time saved per itemnot measured here -- depends on how long applying a warehouse licence by hand takes at the reader's own operation, and on how many facilities share one set of terms
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start a question whose honest answer might be 'free code already does one of the three fields better', which here it does.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
146,417input tokens · this run
160,545output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every model figure on these pages: 141 facility-days answered of 150 scheduled, one completion call each, one tier. The 10 calibration calls are accounted for separately in Cost.cost_of_evaluation_usd; the five free runs cost nothing and are projected nowhere.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.222
$0.222
$1.57
2026-09-12
gemini-3-flash
Google
$0.555
$0.555
$3.94
2026-09-18
gemini-3-8-flash
Google
$0.712
$0.712
$5.05
2026-09-18
llama-5
Meta
$0.865
$0.865
$6.14
2026-09-18
claude-haiku-4-5
Anthropic
$0.949
$0.949
$6.73
2026-09-12
grok-4-5
xAI
$1.256
$1.256
$8.91
2026-09-18
grok-4-6
xAI
$1.256
$1.256
$8.91
2026-09-18
claude-sonnet-5
Anthropic
$1.898
$1.898
$13.46
2026-09-12
gemini-3-1-pro
Google
$2.219
$2.219
$15.74
2026-09-18
gpt-5-6-terra
OpenAI
$2.219
$2.219
$15.74
2026-09-12
gpt-5-6-sol
OpenAI
$3.797
$3.797
$26.93
2026-09-12
claude-opus-4-8
Anthropic
$4.746
$4.746
$33.66
2026-09-12
claude-opus-5
Anthropic
$4.746
$4.746
$33.66
2026-09-12
claude-fable-5
Anthropic
$9.491
$9.491
$67.32
2026-09-18
claude-fable-5-1
Anthropic
$9.491
$9.491
$67.32
2026-09-18
gpt-6-astra
OpenAI
$9.491
$9.491
$67.32
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 93.35 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (149,872 of 160,545), left at the provider's default, so every row below prices a reasoning-on workload. Output is 87 pct of the projected bill on the shared card, which means most of what these rows charge for is the model thinking about four licence clauses. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
EVERY ROW PRICES 141 CALLS, NOT 150. Nine were refused by the provider and produced no tokens; a reader reproducing this run on a healthy account should expect 150 calls and about 6 pct more than the figures here.
THE CHEAPEST OPTION ON THIS PAGE IS NOT ON THIS TABLE. The free coded rule scores 77.33 pct action accuracy at $0.00 and beats the model on naming the largest obligation category. A cheaper model row is not a cheaper kit.
Output length here is driven by how many licence clauses are in play, not by record size -- 347 to 6,405 tokens on records within 29 bytes of each other. A population with different licence terms, or with longer breaches, would move every row below.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Nine modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/licence.pythe licence — a swap seam
the three varying clauses and the day-count arithmetic they imply; the answer key is this module replayed
You change it to:step() holds the four clauses and their order. Your jurisdiction's licence is a different step(); everything above it -- the counters, the cadence, the scorer, the prompt -- is unchanged.
src/licence.py
# The warehouse licence, and the day-count arithmetic it implies. Pure code, no model.
WITHIN, OVER, UNDER_NOTICE, UNDER_REFERRAL = "WITHIN", "OVER", "UNDER_NOTICE", "UNDER_REFERRAL"
NONE = "NONE"
CAPACITY_NOTICE = "CAPACITY_NOTICE"
SUSPENSION_REFERRAL = "SUSPENSION_REFERRAL"
HOLD_PENDING_RELEASE = "HOLD_PENDING_RELEASE"
NOTICE_RELEASED = "NOTICE_RELEASED"
STATUSES = (WITHIN, OVER, UNDER_NOTICE, UNDER_REFERRAL)
ACTIONS = (NONE, CAPACITY_NOTICE, SUSPENSION_REFERRAL, HOLD_PENDING_RELEASE, NOTICE_RELEASED)
CATEGORIES = ("Warehouse receipts outstanding", "Open storage (no receipt issued)",
src/state.pythe carried state — a swap seam
four counters per facility, written from the parsed totals and never from the model's reply
You change it to:describe() renders the counters into English. Four scalars is the whole contract; anything that grows with history turns this into a kit whose input cost climbs every run.
src/state.py
# The carried state — what yesterday's run left behind, and the only reason a reading is a CHANGE.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_facility(store, facility_id):
def advance(store, facility_id, total_bu, capacity_bu, referral_after_days, release_after_days):
def describe(state):
src/schedule.pythe cadence — a swap seam
which business day a run owns, and free detection of a run that did not happen
You change it to: WEEKEND and business_days_between. A jurisdiction with holidays, or a facility filing weekly rather than daily, changes this file and nothing else -- but see breaks_at_scale before believing that.
src/schedule.py
# THE CADENCE. When this monitor wakes, what one run owns, and what a missed run costs.
WEEKEND = (5, 6)
def _d(s):
def is_business_day(day):
def business_days_between(last, today):
def gap(last_run_day, today):
def due(last_run_day, today):
src/segment.pythe section splitter
one record into its seven named sections, pure code
src/segment.py
# Split a daily position record into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Facility", "Licence Terms",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe selector — a swap seam
which sections reach the provider; Facility Contacts is mapped by nothing and therefore never sent
You change it to: NEVER_SENT and SECTION_HINTS. A section no field maps to does not leave the machine, and the subtraction is unconditional rather than a fallback.
src/select.py
# Pick which sections of a position record are sent. Pure code — the last deterministic step
BANNER = "Synthetic Record"
FACILITY = "Facility"
TERMS = "Licence Terms"
POSITION = "Position, as of close of business"
MOVEMENT = "Bin and Receipt Movement"
CONTACTS = "Facility Contacts"
COMMENTARY = "Manager Commentary"
NEVER_SENT = (CONTACTS,)
SECTION_HINTS = {
src/prompt.pythe prompt
instruction, carried state, selected sections -- three parts, string concatenation, no template engine
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework — string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model call — a swap seam
one completion per facility-day, over stdlib urllib
You change it to: PROVIDER, BASE_URL and MODEL in .env. Adding a vendor is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
evals/scoring.pythe scorer
exact match per cell against the computed gold, plus five slices; no model, no judge
evals/scoring.py
# Score a run against the computed gold. Pure code, no model, no judge.
FIELDS = ("position_status", "action", "top_obligation")
def _pct(n, d):
def _norm(v):
def score(records, golds):
def compare(a, b, a_label="a", b_label="b"):
def answers_that_moved(a_cells, b_cells):
evals/check_labels.pythe pre-flight
everything that must be true before a run may spend, including two red-proofs
evals/check_labels.py
# The pre-flight. Everything that must be true before a run is allowed to spend anything.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAIL = []
def ok(label, cond, detail=""):
def main():
Start hereThe shortest path into it
src/licence.pythe three varying clauses and the day-count arithmetic they imply; the answer key is this module replayed A swap seam.
src/state.pyfour counters per facility, written from the parsed totals and never from the model's reply A swap seam.
src/schedule.pywhich business day a run owns, and free detection of a run that did not happen A swap seam.
src/segment.pyone record into its seven named sections, pure code
src/select.pywhich sections reach the provider; Facility Contacts is mapped by nothing and therefore never sent A swap seam.
src/prompt.pyinstruction, carried state, selected sections -- three parts, string concatenation, no template engine
src/adapters/__init__.pyone completion per facility-day, over stdlib urllib A swap seam.
evals/scoring.pyexact match per cell against the computed gold, plus five slices; no model, no judge
evals/check_labels.pyeverything that must be true before a run may spend, including two red-proofs
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1038 input and 1138 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Manager Commentary, which is one of five fixed sentences. src/select.py names the commentary as the field an outside party WOULD control in a real deployment and sends it deliberately, so the surface is visible rather than hidden — but on this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant.
The experimentWe did not attack it — and the boundary that matters most was proven by breaking it on purpose, twice
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus has none. So the four rows below are boundaries, not payloads. Two of them were red-proven: the privacy guard was replaced with the naive fallback every sibling kit once shipped and all 150 records leaked the licensed manager's name, then it was restored and all 150 held — and the FIRST version of that proof was worthless, because on today's corpus the fallback is never reached and swapping it changed nothing. The cadence detector was proven the same way, in three directions: 0 on an intact schedule, exactly 1 on a skipped business day, 0 across a weekend. Confirmed by assertion and by reading the recorded run, not by an attack trial, on 2026-08-23 — no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does the licensed manager's name ever leave the machine?
The Facility Contacts section names the licensed manager, their employee reference, a desk extension and a named contact at the bonded surety — the only place in the corpus whose subject is a PERSON, and the manager is personally liable under the licence. A selector that fell back to the whole record would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint matching nothing falls back to the record MINUS that section rather than to the record. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py, and the red-proof had to be built correctly to mean anything: on today's corpus every hint matches a section that exists, so _fallback is never REACHED and patching it alone proves nothing — the first version of the check did exactly that and reported 0 leaks under the naive guard. The check now breaks a hint as well, which is the scenario src/select.py's docstring names. Guard on: 0 of 150 leak. Guard removed, same broken hint: 150 of 150 leak.
Can a record tell the model what happened on an earlier business day?
If any record named an earlier day's status, action or day count, or carried a carry-in balance, a stateless reader could reconstruct the history from the page and every claim this kit makes about the carried state would be measuring nothing.
evals/check_labels.py scans all 150 records for any status token, action token, history phrase or ISO date other than the record's own, outside the rulebook, and refuses to let a run spend if it finds one. 0 leaks. The rulebook is exempt from that scan because it is where the rules are WRITTEN — and the exemption is EARNED rather than asserted: strip the three varying numbers and the one varying clause out of every Licence Terms block and all 150 are the same bytes, so a facility fact smuggled into the rules would not survive it. The movement section is separately proven to carry no bushel figures, because bushels plus today's total reconstructs yesterday's.
Can a wrong answer poison the next business day?
A monitor that fed the model's verdict back into its own next prompt would compound one bad day into every day after it — the failure mode that makes stateful systems frightening.
It cannot. src/state.advance() calls src/licence.step() on the totals PARSED from the record; the model's reply is scored and never written. ⚠︎ AND THIS RUN CANNOT DEMONSTRATE IT, WHICH THE PAGE SAYS RATHER THAN LETTING THE ZERO IMPLY OTHERWISE: there is no wrong action anywhere in the run to propagate. The three wrong answers are all in the obligation-category field, which is never carried. The property is in the code path and this run is consistent with it, not a demonstration of it.
Can a figure measured under a non-published output ceiling be mistaken for a scored one?
A calibration fired at a different cap produces numbers under a configuration the page does not name.
evals/run.py refuses --max-tokens unless the run id begins with 'c'. The one calibration is c000; the scored run carries the published MAX_TOKENS = 34,000 from src/position.py.
Each boundary above was checked by running an assertion or by reading the recorded run, not by an attack trial — there is no untrusted field on this corpus to construct a payload against. Two of the four are red-proven, meaning the guard was removed and the failure observed rather than merely asserted while passing, and one of those two had to be rebuilt because the first version proved nothing.
The result0 attack trials, and four boundaries checked — two of them red-proven by removing the guard and watching all 150 records fail.
0untrusted input fields on this corpus
0 of 0attack trials run
2 of 4boundaries red-proven, not just asserted
The Manager Commentary IS the field an outside party would control in a real deployment — a licensed manager writes it freely — and this kit sends it rather than hiding it. On this corpus it is one of five sentences chosen by a seeded random number generator, so there is nothing adversarial in it to catch. A version pointed at real filings reopens the question and should be attacked before it ships.
Read this twice
The carried counters are written from the arithmetic, never from the model's reply. That is the difference between a monitor and a system that compounds its own mistakes, and it is why a wrong answer here costs one day rather than every day after it. ⚠︎ AND THIS RUN CANNOT PROVE IT: there is no wrong action anywhere in it to propagate, and the three wrong answers are in the one field that is never carried. The guarantee is in the code path — licence.step() never reads the reply — and the run is consistent with it. What the boundary gives you in any case is narrower than correctness: it guarantees only that a wrong answer stays where it is. It says nothing about a missed RUN, which corrupts every day after it and is measured separately at 47 cells moved on one skipped day.
HonestyWhat this does not prove
Whether a real deployment's Manager Commentary — prose a licensed manager writes freely — would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at real filings.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model.
Whether the nine refused calls leaked anything. Six returned 402 and three returned 429, so no completion was produced and none was billed — but the prompt had already been sent over the wire in every case, which is worth saying rather than counting them as calls that never happened.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
The carried counters are written by code from the parsed obligation totals and the previous counters. The model's status, action and category are scored and are never written back into the state the next run is judged against.
src/state.advance() -> src/licence.step(), called by evals/run.py after every facility-day INCLUDING one whose call failed. There is no code path anywhere in the kit that writes a model reply into the state store, and no code path that notifies anybody of anything.
EvidenceDoes it hold?
What
Measured
A wrong answer does not propagate to the next business day
There is no wrong ACTION anywhere in r001-grain-position to propagate — the confusion matrix is a clean diagonal over the 141 answered days — and the 3 wrong answers are all in the obligation-category field, which is never carried. ⚠︎ So the property is guaranteed by the code path and this run is CONSISTENT with it rather than a demonstration of it.
A facility-day whose CALL failed still advances the counters
evals/run.py's exception branch calls licence.step() before continuing, because the day WAS read — the record was parsed and the arithmetic done — and only the model did not answer. This run exercised that branch 9 times, which is the first time in this series it has been exercised at all: 6 provider refusals and 3 concurrency failures, and every later day in those facilities was still scored against the right counters.
A MISSED RUN does not advance the counters, and that is the opposite decision
evals/run.py's skip branch does not call licence.step() at all, because advancing on a day nobody read would be inventing a reading. The consequence is measured rather than argued: 47 of 360 cells move on one skipped day, 37 of them from right to wrong.
Nothing in this kit notifies anybody
SUSPENSION_REFERRAL is a string in a JSON reply and a column in a result file. There is no notifier, no webhook and no writer of any kind outside results/*.json and data/state.json — confirmed by reading every call site.
The licensed manager's details never reach the provider
0 of 150 records leak the Facility Contacts section, and 150 of 150 leak it when the guard is removed under a hint that matches nothing. Both directions asserted by evals/check_labels.py before any run may spend.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The counters are correct whatever the model says, which means a wrong action is reported wrongly for that day and nothing catches it. On this run that never happened; on yours it will.
IT IS NOT A RATE LIMIT OR AN APPROVAL GATE. There is no apply path to gate: nothing here refers anything to anybody, so there is no 'undo' because there is nothing that could 'do'.
IT DOES NOTHING ABOUT THE SCHEDULE, WHICH IS THE BIGGER HOLE. The guardrail bounds the damage from a wrong ANSWER to one day. A missed RUN is unbounded in exactly the way this guardrail is not, and no code in the kit can prevent it — only detect it afterwards.
It does not make the state file safe to share. One writer, one file, replaced atomically; two schedulers on one facility would race and the loser's day would vanish.
It does not validate the answer key. The key is arithmetic, and the status precedence it implements was not written down anywhere until a calibration run exposed the ambiguity — see environment.ladder's labels row.
WatchedWhat is watched, and why that one
8runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 43 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
12 measured by the latest run31 need the model half
Metric
Owner
Role
Why this one
action-exact
Action, exact match
alarm
action_accuracy_pct; consecutive_day_rule_accuracy_pct; false_action_rate_pct; action_when_one_is_due_pct — alarm on action_when_one_is_due_pct -- a missed action leaves a facility over capacity with nobody told, which is the direction a state regulator cares about; a false action costs review time
status-exact
Position status, exact match
alarm
status_accuracy_pct; exclusion_trap_accuracy_pct — alarm on exclusion_trap_accuracy_pct -- it is the slice where a wrong reading of the licence inverts the verdict rather than shifting a deadline
top-obligation-exact
Largest counted obligation, exact match
alarm
top_obligation_accuracy_pct — alarm on top_obligation_accuracy_pct falling below the free floor's 98.67 pct, which on this run it already has
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
501,737
daily position records edited — the count held, the bytes did not
split.count
150
the daily position records count moved — a different set was scored
split.size_p50
3,345
the median size of one daily position record moved
split.size_p95
3,354
the 95th-percentile size of one daily position record moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (failures 0, stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
action accuracy
not yet known
150 facility-days
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. One run exists, and 9 of its 150 calls never happened.
the two scopes
6.0 points — the difference between counting the 9 refused calls and not
150 scheduled, 141 answered
94.0 pct over all 150 labelled days against 100.0 pct over the 141 the run answered. That is not a band and not noise: it is one arithmetic fact about a run that lost 9 calls, and both figures are published because neither answers the other's question.
the free floor's margin on naming the obligation category
the floor is ahead, and by less than one row on the shared scope
141 facility-days both arms answered
98.58 pct free against 97.87 pct for the model on the 141 days both answered — 0.71 points, and one row of 141 is 0.71 points. Over all 150 the gap is 6.67 points and is dominated by the refused calls. The finding is that the field does not need a model, not that the model is bad at it.
the cost of a missed run
17.5 points, and it is a measurement rather than a spread
120 facility-days
80.0 pct intact against 62.5 pct with business day 3 never read, both for the free coded rule over the 120 days both cover. Pure arithmetic on both sides, so there is no noise in it at all.
the denominator itself
0 — a constant of the corpus, not a result
150 facility-days
150 facility-days, 67 with an action due, 83 quiet, 66 clock-dependent, 23 exclusion traps and 10 reset days on every arm. Which days do what is decided by tools/build_corpus.py before any model sees anything, and the status-vocabulary clarification changed the records without moving a single gold label — the file is byte-identical either side of it.
input volume
not yet known
141 calls
146,417 input tokens over 141 calls WITH the carried state, against 153,524 over 150 without it -- and per call the two are within 15 tokens (1023.49 against 1038.42), because the one-line 'no history is available' sentence is about the same size as the block it replaces. The carried block is not what costs money here; the reasoning it provokes is, and output tells that story instead (1138.62 tokens a day with it, 712.53 without). Prompt assembly is pure code and deterministic, so both figures are model-independent -- but no configuration ran twice, so neither is a band.
provider-side reasoning share
not yet known
141 calls
93.35 pct of output on the scored run (149,872 of 160,545) and 99.04 pct on the calibration (86,507 of 87,349). Two points at two corpus versions, which is a reason to keep the ceiling high rather than a band.
latency
not yet known
141 calls
p50 7,774 ms and p95 23,167 ms, with the slowest single call at 53.4 seconds — the tail is 3.0 times the median and is the number to plan a nightly window against. One recorded run at one configuration.
position status
not yet known
150 facility-days
94.0 pct over all 150 labelled days and 100.0 pct over the 141 answered. No configuration ran twice, so there is no spread to state; what IS known is the control's 76.67 pct on the same field, which is a different prompt and not a band.
where a consecutive-day count decides the answer
0 -> 100 between the two arms, and that is a DIFFERENCE, not a spread
66 clock-dependent facility-days
92.42 pct with the carried state over all 150, 100.0 pct over the answered subset, and 0.0 pct without it -- the control gets 0 of 66 right. This is the slice the kit exists for, so it is banded by what removing the mechanism does rather than by run-to-run noise, of which none has been measured.
where the inclusion clause inverts the verdict
95.65 pct across three arms that all read the page, against 4.35 pct for the one that does not
23 exclusion-trap facility-days
r001 95.65 pct, the stateless control 86.96 pct, the free coded rule 4.35 pct. The READING half of this kit survives losing memory almost intact and collapses when the clause is hardcoded instead of read -- which is the cleanest evidence on this page that the two mechanisms are separable.
the reset trap
one row is 10 points, so nothing finer than that is supportable
10 facility-days
90.0 pct on 10 facility-days -- the days after the middle day of an OOWOO facility, where a rule counting days-over rather than CONSECUTIVE days-over refers a licence the terms do not. The free memoryless floor scores 20.0 pct and the control 20.0 pct on the same 10.
actions missed when one is due
not yet known
67 facility-days where the licence requires an action
91.04 pct of the 67 days an action is due, and all 6 misses are calls the provider refused rather than answers the model got wrong. The control reaches 41.79 pct on the same denominator. THIS IS THE ALARM METRIC: a missed action leaves a facility over capacity with nobody told.
false actions on quiet days
not yet known -- and 3 of 83 is not a rate
83 quiet facility-days
3.61 pct on the scored run, and all 3 are unanswered calls rather than invented actions. The stateless control raises 27 (32.53 pct) and the memoryless floor 42 (50.6 pct) -- both of which ARE invented, because a system with no history answers every quiet day as though it were a first day over capacity.
replies that did not parse
0 of 141 — and it is the calls that never RETURNED that this run lost
141 calls that returned
Every reply that arrived parsed, on both the scored run and the calibration. The 9 failures are refusals, not malformed answers, and they are counted apart because they need opposite fixes: a ceiling change repairs a truncated reply and does nothing for an exhausted account.
HistoryRun history
8 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
ablation · with the model in the path — 2 runs. Columns here are only ever compared with each other.
Metric
a000-grain-position-missedrun-rules 2026-08-23
a001-grain-position-missedrun 2026-08-23
action accuracy, %
62.5
55.0
action when one is due, %
48.94
35.00
consecutive day rule accuracy, %
41.3
50.0
exclusion trap accuracy, %
5.88
90.00
false action rate, %
28.77
35.00
input tokens, whole run
0
62315
model latency p50 ms
0.00
9041.00
model latency p95 ms
0.00
19255.00
output tokens, whole run
0
68836
reset accuracy, %
40.0
60.0
status accuracy, %
68.33
70.00
top obligation accuracy, %
98.33
98.33
not a time series No two of these 2 runs measured the same system — they differ on actions_due, clock_dependent_days, days_scored, documents, max_tokens, provider, schedule_replayed, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 5 runs. Columns here are only ever compared with each other.
Metric
b000-grain-position-fourcolumn 2026-08-23
b001-grain-position-codedrule 2026-08-23
c000-grain-position-calibration 2026-08-23
r001-grain-position 2026-08-23
s001-grain-position-stateless 2026-08-23
action accuracy, %
46.00
77.33
100.00
94.00
56.00
action when one is due, %
41.79
64.18
100.00
91.04
41.79
consecutive day rule accuracy, %
0.00
66.67
100.00
92.42
0.00
exclusion trap accuracy, %
0.00
4.35
100.00
95.65
86.96
false action rate, %
50.60
12.05
0.00
3.61
32.53
input tokens, whole run
0
0
9641
146417
153524
model latency p50 ms
0.00
0.00
89463.00
7774.00
6072.00
model latency p95 ms
0.00
0.00
147133.00
23167.00
10393.00
output tokens, whole run
0
0
87349
160545
106880
reset accuracy, %
20.0
70.0
100.0
90.0
20.0
status accuracy, %
63.33
74.67
60.00
94.00
76.67
top obligation accuracy, %
98.67
98.67
100.00
92.00
99.33
not a time series No two of these 5 runs measured the same system — they differ on actions_due, clock_dependent_days, days_scored, documents, failures, max_tokens, provider, schedule_replayed, stateless, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-grain-position-stub 2026-08-23
action accuracy, %
55.33
action when one is due, %
0.0
consecutive day rule accuracy, %
40.91
exclusion trap accuracy, %
0.0
false action rate, %
0.0
input tokens, whole run
179507
model latency p50 ms
0.00
model latency p95 ms
0.00
output tokens, whole run
4340
reset accuracy, %
70.0
status accuracy, %
63.33
top obligation accuracy, %
98.67
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 12 chips that all say so.
DeviationsWhat deviated
0 breaches across 8 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
the carried-state block in the prompt (r001 against s001)
the CLOCK half and the READING half move in opposite directions, and one field moves the wrong way. Action accuracy and the consecutive-day slice collapse without it; the inclusion-clause slice barely moves, because that clause is on the page; and naming the largest obligation category gets BETTER without it.
measured
r001 against s001 over the 141 days every arm answered: action 100.0 -> 56.74 pct, consecutive-day slice 100.0 -> 0.0 pct (0 of 66), inclusion clause 100.0 -> 86.36 pct, largest obligation 97.87 -> 99.29 pct. Output falls 1138.62 -> 712.53 tokens a day while INPUT barely moves (1038.42 -> 1023.49, under 15 tokens): the one-line no-history sentence is about the same size as the block it replaces, so the memory shows up in what the model THINKS and not in what it is sent.
one business day removed from the SCHEDULE (r001 against a001)
everything that depends on a count, in one direction only -- and it moves the model further than it moves free code, which is the opposite of what a reader would guess. Nothing raises an error on either side.
measured
a001 against r001 over the 53 days both answered: action 100.0 -> 50.94 pct, status 100.0 -> 67.92 pct, largest obligation unmoved at 98.11 pct. 43 of 159 cells moved and ALL 43 went right to wrong, none the other way. The same perturbation costs the free coded rule 17.5 points (a000 against b001, 80.0 -> 62.5 pct) against the model's 49.06.
reading the licence clauses instead of hardcoding them (r001 against b001)
the inclusion-clause slice, catastrophically, and nothing else much. It does NOT move the largest-obligation field, where the hardcoded rule is ahead.
measured
b001 is the IDENTICAL src/licence.step with the majority terms hardcoded, so any gap is the three varying clauses and nothing else. Inclusion clause 4.35 pct against the model's 95.65 over 23 days; action 77.33 against 94.0 over all 150; largest obligation 98.67 against 92.0, where the free rule WINS. Only 5 of the 30 facilities carry the majority terms.
a call the provider refuses
the published headline, and nothing about the system. Nine refusals moved action accuracy from %s pct to %s pct without a single wrong answer being given, which is why this page prints both scopes.
measured
results/eval-r001-grain-position.json failures[]: 6 x 402 Insufficient Balance, 3 x 429 after four exponential backoffs at 12 workers. 7 of the 9 landed on business days 4 and 5, which is also why the model cadence pair compares 53 days rather than 60.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
action accuracy
nothing yet
the two scopes
any run whose refused-call count is not 0
the free floor's margin on naming the obligation category
the model going ahead of the floor on this field, which would be the finding
the cost of a missed run
any change at all, on any re-run — both arms are deterministic
the denominator itself
any change at all, on any run
input volume
nothing yet
provider-side reasoning share
nothing yet
latency
nothing yet
position status
nothing yet
where a consecutive-day count decides the answer
any answer at all on this slice from an arm with no carried state
where the inclusion clause inverts the verdict
a drop toward the free rule's figure, which would mean the clause stopped being read
the reset trap
any change at all -- at 10 rows every movement is a whole row
actions missed when one is due
any missed action that is an ANSWER rather than a refusal -- this run has none
false actions on quiet days
a false action that is an answer -- one row is 1.20 points
replies that did not parse
any unparsed reply at all, which would point at the output ceiling
NextThe three you would add first
A liveness check on the schedule, outside the kitEvery figure on this page assumes the run fired. The measured cost of one missed business day is 47 cells moved and 17.5 points of action accuracy for a system that cannot make a mistake, and the kit can only detect the gap AFTER the fact. A monitor whose scheduler nobody watches is a compliance report that silently stops.
An alarm on refused calls, not just a count in a result fileThis run lost 9 of 150 calls — 6 to an exhausted account balance and 3 to a concurrency limit — and the only reason anybody knows is that a person read the output. On a clock, with nobody watching, the same failure is a day that produces no exception list at all.
A human step in front of any real referralSUSPENSION_REFERRAL is the highest-consequence output this kit has: it puts a licensed warehouse in front of a state regulator. 12 of 150 days carry one, and a notifier without a person in front of it would send them on the strength of a run that also silently lost 9 calls.
A state store with a concurrency modeldata/state.json is one file replaced atomically. That is right for one process and is not a design for a scheduler, and the moment two facilities are advanced in parallel by different workers this needs to be a real store.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-check this guardrail's evidence on any change to src/licence.py::step (it writes the counters), to src/state.py (what is carried and how it is rendered), or to src/schedule.py (what one run owns). Re-run the FREE arms on every commit -- the pre-flight, both floors and the free cadence ablation cost nothing, and together they re-prove the privacy guard, the leak-free corpus, the chain-resume and the price of a missed run. Re-run the PAID arms (r001, s001, a001) only on a change to the prompt, the model, MAX_TOKENS or the corpus -- 360 calls.
What this cannot tell you
Whether the no-propagation guarantee holds in practice. It is guaranteed by the code path -- licence.step() never reads the model's reply -- and this run cannot demonstrate it, because there is no wrong ACTION anywhere in it to propagate.
Whether a missed run the kit ANNOUNCES costs less than one it does not. a001 is silent by design; src/schedule.gap() detects the gap for free and src/state.describe could carry a warning, and whether that recovers any of the 43 lost cells was not paid for.
Whether any of these levers reproduces. No configuration ran twice, so every ripple row above is one observation of a difference between two runs, not a difference plus a band.
Whether the state file is safe under concurrency. One writer, one file, replaced atomically; two schedulers advancing one facility have never been run and would race.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all — json, os, re, random, argparse, datetime, urllib, subprocess and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the cadence
kits/UC0089-grain-position/src/schedule.py
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
THIS is the seam where a framework genuinely earns its place, and this kit is the one that can price it: 30 independent chains of 5 strictly-ordered steps, with the trigger, the retry policy, the backfill question and the liveness alarm all OUTSIDE the kit. One missed run costs 47 cells. A ThreadPoolExecutor is the right size for an eval and the wrong size for a monitor that actually runs on a clock, and the difference is measured rather than asserted.
the carried state
kits/UC0089-grain-position/src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior runs)
those carry a TRANSCRIPT; this carries four COUNTERS written by arithmetic. The whole reason the cost is flat in history length is that nothing here remembers what was said, only what was counted, and a memory layer would give back the growth this design exists to avoid
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install. What a router WOULD have bought on this run is a second key to fail over to when the first returned 402 six times
the scorer
kits/UC0089-grain-position/evals/scoring.py
eval harnesses (promptfoo, DeepEval)
three exact-match comparisons and five slices of the same cells is a dict comprehension, not a platform
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each facility is a chain of five steps — business day 1 to 5 — with no branching and exactly one edge between consecutive days, carrying four counters. Different facilities never touch. A framework would add an orchestrator to a for-loop that already runs 30 chains wide. What it would NOT be adding an orchestrator to is the thing that actually needs one, which is the clock outside the loop.
The other sideWhat a framework costs you
No scheduler, and this kit publishes what that costs: 47 cells move on one missed business day. The cadence, the retry policy, the liveness alarm and the question of what to do when a day does not arrive are all outside the kit.
No state store worth the name. data/state.json is one file replaced atomically; a framework would bring a checkpointer with a concurrency model, which this does not have.
No observability beyond what evals/run.py prints and writes to results/. There is no tracing, no dashboard and no alarm — which is how 9 refused calls become a number in a JSON file that somebody has to read.
No retry/backoff beyond src/adapters' own bounded retry, which covers a busy provider and a dropped connection and nothing else. It did not cover an exhausted account, because nothing can.
No failover between providers or keys. One BASE_URL, one API_KEY; when that account returned 402 the run simply stopped producing answers.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation. The cadence row is the exception in one direction only: what a missed run COSTS is measured, what a scheduler would have PREVENTED is not.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-grain-position on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
7,774 ms
not yet known
nothing yet
Model, p95
23,167 ms
not yet known
nothing yet
Input tokens
146,417
not yet known
nothing yet
Output tokens
160,545
not yet known
nothing yet
No movement column. Not one of the 4 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-grain-position-calibration89,463 ms
r001-grain-position7,774 ms
s001-grain-position-stateless6,072 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
2 runs not plotted. b000-grain-position-fourcolumn, b001-grain-position-codedrule recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 8 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
daily position records
data/corpus/<FACILITY>-D<n>.txt — 150 files, 501,737 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Facility Contacts never does, by src/select.NEVER_SENT
the answer key
data/gold.jsonl — 150 rows, the output of src/licence.replay over the corpus totals, never hand-authored and re-derived FROM THE RECORDS by evals/check_labels.py before any run may spend
never — evals/scoring.py is pure code, no model and no key
the carried state
data/state.json in a deployment (src/state.py, written atomically). ⚠︎ IN AN EVAL IT IS SCOPED TO THE RUN AND NEVER TOUCHES DISK — evals/run.py builds a fresh in-memory store per facility, because a run that started from the previous run's memory could not be re-run or compared with its own control
one or two English sentences per call, produced by state.describe from four scalar counters — never a prior record and never a prior reply
the schedule
src/schedule.py — a weekday calendar and two date functions. There is NO scheduler in this repository: what fires the run is cron, Airflow or a person, and that boundary is deliberate
never — two dates and a weekday test, no network
the recorded runs
results/eval-*.json — the scored run, both free floors, the cadence ablation, the calibration, the wiring stub and the two comparisons, all committed
never — they are read by the page and by evals/floors.py and evals/cadence.py, which re-score them for $0.00
the key
.env or the shared repo-root .env — never committed, 0600
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
one run per BUSINESS DAY, after close of business, once the day's scale tickets and receipt cancellations are posted. One run owns exactly one business day for the WHOLE population — every licensed facility with an open storage obligation, re-read whole, including the ones that were quiet yesterday, because the licence's clocks advance whether or not anything changed and a notice release is an event that happens on a day when nothing happened. What this run owns that the last one did not is the day's transitions: every counter advances by exactly one, a notice opens on the first day over, a referral fires on the Nth, a hold becomes a release. src/schedule.py holds the calendar and the gap detector; nothing in this repository fires the run.
MEASURED TWICE, AND THE MODEL IS THE WORSE OF THE TWO. Without a model: the same free coded rule over the same corpus, once on an intact schedule and once with business day 3 never read -- action accuracy 80.0 -> 62.5 pct over the 120 days both cover, 47 of 360 cells moved, 37 right to wrong. Both arms are deterministic, so there is no noise in that figure at all. WITH the model: 100.0 -> 50.94 pct over the 53 days both arms answered, 43 of 159 cells moved, 43 right to wrong and NOT ONE the other way. A model compounds a broken clock rather than absorbing it. Detection is free and does not help: src/schedule.gap() flagged the gap on 30 of 53 days at $0.00 and repaired none of them. (a000-grain-position-missedrun-rules against b001-grain-position-codedrule, and a001-grain-position-missedrun against r001-grain-position, via evals/cadence.py; results/eval-cadence-grain-position.json)
one missed run is not recoverable, and the kit refuses to pretend otherwise. A position for a day nobody read cannot be reconstructed from a later record — a close-of-business snapshot does not say what the total was two days ago — so src/schedule.due() returns TODAY plus a count of what was lost rather than backfilling. Beyond that: public holidays are not modelled at all, so a jurisdiction's holiday reads as a missed run, which given the figure above is the wrong failure to have; and a facility that files weekly rather than daily needs a different WEEKEND and a different business_days_between.
every consecutive-day figure on this page assumes the schedule held. The counters count RUNS. If yours does not fire every business day, the correct reading of the headline is not 94.0 pct — it is the cadence ablation's 62.5 pct, and that is for a system that cannot make a mistake.
state
four scalar counters per facility — consecutive business days over capacity, consecutive days back within, whether a capacity notice is open, whether the licence has already been referred — written by src/licence.step from the PARSED totals and never from the model's reply, and rendered by src/state.describe into one or two English sentences. That block is the entire route from one run to the next.
134 characters on the worked example, and it is worth 43.26 POINTS OF ACTION ACCURACY -- measured against a real control, not against a different system. s001-grain-position-stateless is this prompt with this block replaced by one line saying no history is available, all 150 calls answered: action 56.0 pct against 94.0 pct, and 0 OF 66 on the days a consecutive-day count decides the answer against 61 of 66. The control does not get a single one of them right, and its own rationales say why -- a first day over capacity is the only thing one record can support. ⚠︎ AND IT WINS ONE FIELD: naming the largest counted obligation category is 99.29 pct WITHOUT the carried state and 97.87 pct with it, on the 141 days every arm answered. On a pure reading task the state is a distraction. (r001-grain-position against s001-grain-position-stateless, via evals/floors.py; lenses.LLM.prompt_parts[2])
four scalars, so business day 400 costs what business day 4 costs — the OPPOSITE curve to a conversation kit, whose input grows with the square of the turns. What is NOT bounded is the store: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model. Two schedulers advancing one facility would race, and the loser's day vanishing silently un-opens a capacity notice and restarts a referral clock at zero.
a corpus whose records carry a CARRY-IN BALANCE. Real filings do, and yesterday's total — therefore whether yesterday was over capacity — is then recoverable by subtraction, so the carried state stops being the only route and nothing here says how much of the work it is still doing. This corpus removes that route deliberately and evals/check_labels.py asserts the removal on all 150 records.
model
one call per FACILITY-DAY carrying the instruction, the carried-state block and six of the record's seven sections, behind src/adapters/__init__.py, at the published MAX_TOKENS = 34,000 with provider-side reasoning left at the default.
0 of 141 returned replies failed to parse; the largest was 6,405 output tokens (19 pct of the cap). 9 of the 150 scheduled calls never returned at all — 6 refused with 402 Insufficient Balance on the shared account and 3 lost a 429 concurrency limit after four exponential backoffs at 12 workers. (r001-grain-position; results/eval-r001-grain-position.json failures[])
the 34,000-token ceiling was set at roughly twice the largest reply the calibration saw (16,780 on c000, fired deliberately at a 24,000 cap so the true maximum could be observed rather than clipped) and this run came nowhere near it. ⚠︎ The calibration was fired against the corpus as it stood BEFORE the status vocabulary was written into the rulebook, so the ceiling is inherited rather than re-derived — the gold is byte-identical either way and evals/check_labels.py re-derives it, but the prompt is five lines longer than the one that set the ceiling.
one tier was run, so nothing here compares two models. Point .env at your own model and the free scorer, both free floors and the cadence ablation re-run on your numbers for $0.00.
labels
data/gold.jsonl, 150 rows, computed by src/licence.replay over the corpus totals at generation time — and then RE-DERIVED FROM THE RECORDS by evals/check_labels.py before any run may spend, so the key is not merely whatever the generator wrote down.
83 gold NONE / 28 CAPACITY_NOTICE / 12 SUSPENSION_REFERRAL / 14 HOLD_PENDING_RELEASE / 13 NOTICE_RELEASED over 150 facility-days; 66 days turn on a consecutive-day count the record does not state, 23 on the licence's inclusion clause, and 10 sit after the middle day of a reset trap. 150 of 150 rows re-derive. (lenses.Eval.dataset, grain-position-2026-08-23-30facilities-150days; evals/check_labels.py)
the key is only as good as ITS reading of the licence, and this kit already paid for that once. The first calibration scored 100 pct on the action and 60 pct on the status, and every status miss was defensible: the record named four status values and never said which governs when two fit at once. The fix was to write the precedence into the rulebook, not to mark the reading wrong — an underspecified policy is the defect, and it is the one this kit is most likely to meet on real licences.
your own facilities: hand-label the gold, which is the real work. This key is a luxury of controlling both the generator and the licence, and a hand-labelled key has an error rate nothing here has measured.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a facility over capacity for a third day with no SUSPENSION_REFERRAL, and no error anywhere
the referral clock is behind the calendar. A scheduled run did not happen, so the counters — which count RUNS, not days — are short by however many were missed. Nothing in the kit raises on this: the arithmetic is internally consistent and simply describes a different week from the one that happened.
run src/schedule.gap() over the state's last run date and today before disputing the answer. It is two dates and a weekday test, it costs nothing, and on the cadence ablation it flagged 30 of 120 days correctly. Then check whether the day it names has a record nobody read. (results/eval-a000-grain-position-missedrun-rules.json and results/eval-cadence-grain-position.json — 47 cells moved on one skipped day, 37 of them from right to wrong)
a facility reported OVER whose four obligation lines sum past capacity but whose counted total does not
the licence's inclusion clause was not read. This facility's terms EXCLUDE company-owned grain from the storage obligation and something added all four columns anyway — which is not carelessness, it is applying the terms of a different licence.
read the Obligation basis line on the record itself, then compare against the two subtotals the record prints. Both are on the page and neither is marked as governing; the clause is what decides. (results/eval-b001-grain-position-codedrule.json — the free coded rule scores 4.35 pct on the 23 days this inverts the verdict, against the model's 95.65 pct)
a SUSPENSION_REFERRAL on a facility whose days over capacity were not consecutive
something counted days over rather than CONSECUTIVE days over. One day back inside capacity resets the referral clock, and a counter driven by actions rather than by readings never resets at all.
look at what the carried state said about days_within, not just days_over. Ten days of this corpus sit after the middle day of a reset facility and exist to catch exactly this. (results/eval-b000-grain-position-fourcolumn.json — the memoryless floor scores 20.0 pct on those 10 days; r001 scores 90.0 pct)
No machine symptom — this failure leaves no trace in any output.
the harness records every refusal by name in results/eval-*.json failures[] and the scorer counts an unanswered day as WRONG rather than skipping it — so the score falls instead of the row disappearing. That is the whole defence, and it only works if somebody reads the score: there is no alerting in this kit and no notifier of any kind.
['Whether a missed run the kit ANNOUNCES costs less than one it does not. a001 is silent by design, because a scheduler that does not fire does not announce anything, so the measured damage is the naive case. src/schedule.gap() detects the gap for free and src/state.describe could carry a warning; whether telling the model recovers any of the 43 lost cells is 60 more calls and was not paid for.', 'Whether rendering the carried counters as English beats rendering them as JSON. Two more 150-call runs; not paid for.', "What disabling provider-side reasoning does. 93.35 pct of the scored run's output tokens were reasoning at the provider's default; src/adapters can send a thinking field and this kit never has.", 'Whether any figure here reproduces. No configuration ran twice, so there is no repeat probe and no band on any metric on this page.', 'Concurrency and GPU sizing. The runs used 12 worker threads against one hosted endpoint and hit a 429 three times; nothing here measures what width the provider will actually sustain, and nothing here runs on hardware anybody owns.', "Provider-side retention. What the configured BASE_URL does with a prompt after answering it is the reader's contract with their vendor and is not observable from this repository.", 'Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two schedulers advancing one facility have never been run and would race.', 'Anything about public holidays. src/schedule.py models weekends only, so every figure here is measured on a week with none.']
The corpus licence, from the Data lens: MIT, with the rest of the kits repository. The corpus carries no third-party material. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineAction, exact match
whether the single action the licence requires today was named exactly
$0.00per 1,000 daily position records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score -- string equality after upper-casing and trimming, against a gold set computed by src/licence.replay
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The record
WH-0447-D4
What the last run left behind
A capacity notice is OPEN against this facility. As of the last run it had been over licensed capacity on 2 consecutive business days.
This facility's own clauses
refer after 3 consecutive business days over capacity; release after 2 back within; storage obligations EXCLUDE company-owned grain
Today's position
694,792 bu counted under this licence against 620,000 bu licensed capacity (the four-column sum is 750,292 bu)
Gold status
UNDER_REFERRAL
Gold action
SUSPENSION_REFERRAL
What the model reported
UNDER_REFERRAL
What the model did
SUSPENSION_REFERRAL
In the model's own words
Excluding company-owned grain, today's storage obligations are 694,792 bu against 620,000 bu licensed capacity, making this the 3rd consecutive business day over capacity and triggering the suspension referral under the licence terms.
Verdict
hit on all three fields
Grader
Verdict
Why
Action, exact match
hit
gold SUSPENSION_REFERRAL, model SUSPENSION_REFERRAL -- exact. This is the reference grader and this is the day the whole kit is about: the third consecutive business day over capacity at a facility whose licence refers after 3.
Position status, exact match
hit
gold UNDER_REFERRAL, model UNDER_REFERRAL -- exact. Reaching it needs the inclusion clause: the four-column sum is 750,292 bu, the counted total is 694,792 bu, and only one of those exceeds the 620,000 bu licence.
Largest counted obligation, exact match
hit
gold 'Delayed-price contracts', model 'Delayed-price contracts' -- exact. ⚠︎ AND THIS IS THE FIELD A FREE ARGMAX AND THE STATELESS CONTROL BOTH BEAT THE SCORED RUN ON across the corpus (98.58 and 99.29 pct against 97.87), so a hit here is one row and not a recommendation.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier, on a schedule that held
scored 94.0%
the fast tier, memory removed (THE CONTROL)
scored 56.0%
free coded rule, usual terms hardcoded
scored 77.3%
free four-column threshold, no memory
scored 46.0%
the wiring stub, no key and no model
scored 55.3%
the fast tier, ONE MISSED RUN (60 days, not 150)
scored 55.0%
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/licence.replay over the corpus totals and re-derived FROM THE RECORDS by evals/check_labels.py before any run may spend.
These rates are UNKNOWN, on purpose
It is the reference standard, so it cannot be scored against itself: a TPR or TNR here would be measuring this grader against its own verdicts and would read as evidence. There is also no binary positive class to compute one over -- action is a five-value closed vocabulary, and the two directions that DO matter (a missed action against a false one) are counted apart and published as counts rather than folded into a rate.
Watch these
action_accuracy_pct
consecutive_day_rule_accuracy_pct
false_action_rate_pct
action_when_one_is_due_pct
Alarm on
action_when_one_is_due_pct -- a missed action leaves a facility over capacity with nobody told, which is the direction a state regulator cares about; a false action costs review time
How tight can the band be? One row of the 67 days where an action is due is 1.49 points, so no band finer than that is supportable. The false-action rate's denominator is 83 quiet days, one row of which is 1.20 points.
Cadence: every scored run, and it is free -- there is no reason not to
The decisionWhen to reach for it
Use it
always -- it is the whole point of the kit and it costs nothing
Do not use it
never; there is no cheaper or more faithful way to grade a closed vocabulary
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one linePosition status, exact match
whether the status to report after the licence's clauses is right
$0.00per 1,000 daily position records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, same equality, four-value vocabulary
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The record
WH-0447-D4
What the last run left behind
A capacity notice is OPEN against this facility. As of the last run it had been over licensed capacity on 2 consecutive business days.
This facility's own clauses
refer after 3 consecutive business days over capacity; release after 2 back within; storage obligations EXCLUDE company-owned grain
Today's position
694,792 bu counted under this licence against 620,000 bu licensed capacity (the four-column sum is 750,292 bu)
Gold status
UNDER_REFERRAL
Gold action
SUSPENSION_REFERRAL
What the model reported
UNDER_REFERRAL
What the model did
SUSPENSION_REFERRAL
In the model's own words
Excluding company-owned grain, today's storage obligations are 694,792 bu against 620,000 bu licensed capacity, making this the 3rd consecutive business day over capacity and triggering the suspension referral under the licence terms.
Verdict
hit on all three fields
Grader
Verdict
Why
Action, exact match
hit
gold SUSPENSION_REFERRAL, model SUSPENSION_REFERRAL -- exact. This is the reference grader and this is the day the whole kit is about: the third consecutive business day over capacity at a facility whose licence refers after 3.
Position status, exact match
hit
gold UNDER_REFERRAL, model UNDER_REFERRAL -- exact. Reaching it needs the inclusion clause: the four-column sum is 750,292 bu, the counted total is 694,792 bu, and only one of those exceeds the 620,000 bu licence.
Largest counted obligation, exact match
hit
gold 'Delayed-price contracts', model 'Delayed-price contracts' -- exact. ⚠︎ AND THIS IS THE FIELD A FREE ARGMAX AND THE STATELESS CONTROL BOTH BEAT THE SCORED RUN ON across the corpus (98.58 and 99.29 pct against 97.87), so a hit here is one row and not a recommendation.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier, on a schedule that held
scored 94.0%
the fast tier, memory removed (THE CONTROL)
scored 76.7%
free coded rule, usual terms hardcoded
scored 74.7%
free four-column threshold, no memory
scored 63.3%
the wiring stub, no key and no model
scored 63.3%
In operationWhat to monitor
Reference standard: the same data/gold.jsonl.
These rates are UNKNOWN, on purpose
It answers a different question from the grader the others are measured against -- which state the facility is IN rather than what to DO about it -- so there is nothing here for an agreement rate to be about, and a rate against the four-value status vocabulary would be accuracy under another name. Its own TPR and TNR are not measured.
Watch these
status_accuracy_pct
exclusion_trap_accuracy_pct
Alarm on
exclusion_trap_accuracy_pct -- it is the slice where a wrong reading of the licence inverts the verdict rather than shifting a deadline
How tight can the band be? The exclusion-trap slice is 23 days, so one row is 4.35 points and no band finer than that means anything.
Cadence: every scored run, free
The decisionWhen to reach for it
Use it
always -- it is the field the licence's inclusion clause inverts, so it is where a hardcoded rule gives itself away
Do not use it
it is not independent of the action grader; a status miss and an action miss on the same day are usually one mistake counted twice, so do not sum them
Warehouse Receipt Obligation Tracking Against Licensed Capacity
PresenterOpens the private repo. Visible to admins only.
In one lineLargest counted obligation, exact match
whether the largest obligation category counted under this licence was named exactly as the record writes it
$0.00per 1,000 daily position records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score -- case-insensitive on trimmed text and nothing looser, because a fuzzy match would forgive an invented category
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The record
WH-0447-D4
What the last run left behind
A capacity notice is OPEN against this facility. As of the last run it had been over licensed capacity on 2 consecutive business days.
This facility's own clauses
refer after 3 consecutive business days over capacity; release after 2 back within; storage obligations EXCLUDE company-owned grain
Today's position
694,792 bu counted under this licence against 620,000 bu licensed capacity (the four-column sum is 750,292 bu)
Gold status
UNDER_REFERRAL
Gold action
SUSPENSION_REFERRAL
What the model reported
UNDER_REFERRAL
What the model did
SUSPENSION_REFERRAL
In the model's own words
Excluding company-owned grain, today's storage obligations are 694,792 bu against 620,000 bu licensed capacity, making this the 3rd consecutive business day over capacity and triggering the suspension referral under the licence terms.
Verdict
hit on all three fields
Grader
Verdict
Why
Action, exact match
hit
gold SUSPENSION_REFERRAL, model SUSPENSION_REFERRAL -- exact. This is the reference grader and this is the day the whole kit is about: the third consecutive business day over capacity at a facility whose licence refers after 3.
Position status, exact match
hit
gold UNDER_REFERRAL, model UNDER_REFERRAL -- exact. Reaching it needs the inclusion clause: the four-column sum is 750,292 bu, the counted total is 694,792 bu, and only one of those exceeds the 620,000 bu licence.
Largest counted obligation, exact match
hit
gold 'Delayed-price contracts', model 'Delayed-price contracts' -- exact. ⚠︎ AND THIS IS THE FIELD A FREE ARGMAX AND THE STATELESS CONTROL BOTH BEAT THE SCORED RUN ON across the corpus (98.58 and 99.29 pct against 97.87), so a hit here is one row and not a recommendation.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier, on a schedule that held
scored 92.0%
the fast tier, memory removed (THE CONTROL)
scored 99.3%
free coded rule, usual terms hardcoded
scored 98.7%
free four-column threshold, no memory
scored 98.7%
the wiring stub, no key and no model
scored 98.7%
In operationWhat to monitor
Reference standard: the same data/gold.jsonl.
These rates are UNKNOWN, on purpose
There is no negative class to compute a rate over: naming the largest counted obligation category is a choice among the four printed on the record, so every answer is either the right one or one of three wrong ones and nothing here is a true/false question. The accuracy already published is the only honest figure.
Watch these
top_obligation_accuracy_pct
Alarm on
top_obligation_accuracy_pct falling below the free floor's 98.67 pct, which on this run it already has
How tight can the band be? One row of 150 is 0.67 points; the model-to-floor gap on the shared 141-day scope is 0.71 points, which is under one row and is reported as such rather than as a finding about models.
Cadence: every scored run, free
The decisionWhen to reach for it
Use it
keep it, because it is what lets the run say which fields need a model at all
Do not use it
do not put it in a headline average -- it is the field this kit recommends taking off the model entirely
A living map of modern AI — kept current every morning