Catch a yard trailer's free time before it runs out
Trailers pile up free time while nobody is watching the clock between shift changes. This app checks each trailer against its carrier agreement and flags the moment detention charges start.
PresenterOpens the private repo. Visible to admins only.
For the yard managerLogistics & Transportation · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A yard or dock manager at a trucking or logistics yard, watching every shift.
✕Today's manual process
1Walk the yard report at shift change, one trailer visit at a time.
2Read the gate log to find when the clock really started, and every stop along the way.
3Decide who caused each stop the carrier or the yard, by memory and judgment call.
4Miss one and a detention charge either goes unbilled or gets disputed and lost.
Every trailer walked and judged manually
✓With the app
1The clock is read the moment it really started, straight from the gate log.
2Carrier stops are separated from the yard's own delays, automatically.
3Free time is checked against the carrier agreement on file, every few hours.
4The event opens once the moment it's earned, never twice, for a manager to confirm.
One event opens, exactly once, automatically
See it work
One real case, from the recorded run, step by step
A trailer waiting at Dock 12 for almost six hours has just crossed its four-hour free time limit.
Catch a yard trailer's free time before it runs outReference appBuilt to be shaped to your process
5
1The reading 5.75 hours have now elapsed; before this check the visit was reported clear.
2What doesn't count 0.75 hours were the carrier's own delay; 4.00 hours of free time are allowed.
3What it found Detention has opened on this visit, past its free time.
4The proof 5.75 hours minus the 0.75-hour break minus 4.00 hours free time leaves 1.00.
5Who acts next The event opens now, for a yard manager to confirm.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch a yard trailer's free time before it runs out
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"Which trailers are still on the yard" is a SELECT and nobody needs a language model for it. The question a yard actually has to answer is the next one: which of them are now past the free time their carrier agreement grants, how many hours of that are chargeable, and -- the part that is a monitor question rather than a query -- whether the detention event has already been opened. A detention claim survives a carrier dispute on CONTEMPORANEOUS evidence: the gate log, the appointment record and the yard log as they stood when the free time ran out. Reconstructed three shifts later, it is the thing that gets thrown out. Today a clerk walks the yard report at shift change and works it out from the gate log by eye. Somebody walking the yard report at shift change: opening each open visit, finding the clock start (the later of the appointment and the gate-in, not the arrival), reading the yard log to decide which stops were the carrier's and which were ours, adding what is left to whatever was already on the clock at the last look, and remembering whether a detention event has already been raised on this visit.
Audience
A yard or dock manager deciding what to log this shift, and the transportation analyst who will have to defend the resulting charges to a carrier. Its own answer on the headline band is a draw with free code, which is the point. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual yard snapshots
The corpus is 120 yard snapshots, 0.44 MB (json 1 · jsonl 1 · txt 120). A YARD'S DETENTION FILE IS THE MOST COMMERCIALLY SENSITIVE ARTEFACT AN OPERATION KEEPS. It is the evidence a carrier bills against and the evidence a shipper disputes with, it names drivers, and there is no public one -- for the same reason there is no public corpus of anyone's freight invoices. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/detention.step's output over the planted intervals, so an interval rule this fiddly cannot carry its author's misreading into the score. ⚠︎ AND THE RULES ARE INVENTED. The eight rules, the free-time hours and the per-hour rates reproduce no carrier agreement, no tariff, no broker contract and no regulator's guidance, and name none. The catalogue row this kit was built from (logi:YDA-003) records the free-time value as an OPERATOR-SUPPLIED clock and explicitly BLOCKS a silent default; that is implemented as Rule D-6 and measured as context_incomplete_recall_pct rather than asserted. The rule text is reproduced in full on all 120 pages precisely so it can be read, disbelieved and replaced.
The corpus
The 120 yard snapshotsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your yard snapshots. That is the whole change — there is no database to migrate.
One yard snapshot, as the model receives itVIS-0001-R1.txt · 1 of 120
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated yard snapshot for an AI use-case
kit; it reproduces no carrier agreement, no tariff and no real facility. The detention
terms are ILLUSTRATIVE.
Visit
----------------------------------------------------------------
Facility : DC-ONT-03 (Inland Empire DC)
Visit reference : VIS-0001
Equipment : 53' reefer, unit 510944
Carrier (SCAC) : KNLX
Move type : LIVE LOAD
Gate in (arrival) : 2026-03-04 05:30
Appointment : 2026-03-04 08:45
Clock start : 2026-03-04 08:45 (later of appointment and gate-in)
Gate out (departure) : -- still on site
Door : 136
Detention Terms
----------------------------------------------------------------
Carrier agreement : MSA-KNLX-2025
Free time on file : 3.00 hours from clock start
Detention rate on file : 85.00 USD per hour, 15-minute increments
Rule D-1 Free time runs from the CLOCK START, which is the later of the appointment time and the
gate-in time. A carrier that arrives early does not start its own clock early.
Rule D-2 Detention accrues once free time is exhausted while the equipment is still on site, in
15-minute increments, at the rate on file.
Rule D-3 The clock STOPS for any interval the yard event log attributes to the CARRIER: the driver
leaving the site, no power unit present, an hours-of-service break, or the carrier
declining to work the door it was given. Those hours do not count.
Abridged — the file continues.
The outcomeWhat a good result looks like
Every visit past its free time is on the watchlist with the hours already accrued, and exactly one detention event exists per visit -- opened on the run that first saw the breach, not re-opened on every run after it.
And when it cannot
Two ways, and they cost different things. A MISSED raise is a charge nobody can bill, because the evidence was never captured while it was contemporaneous. A DUPLICATE raise is two charges for one dwell, which is how a shipper loses a carrier's trust and an audit in the same week. The scored run made neither: 0 of 28 and 0 of 92.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
You already know which trailers are over their free time and only need the list — a SQL query, or the b000 floor in this repo Elapsed against a threshold is arithmetic. On this corpus b000 gets 80.83 pct of statuses right for nothing at all.
Your yard log is structured -- every hold carries a coded reason from a closed list — the b002 floor: keyword-free, code-only attribution on your own codes The model's entire margin on this corpus is reading free prose. Take the prose away and there is nothing left to buy: the floor already ties on status and the gap on hours is attribution.
Your yard log is prose an inbound clerk types, and it mentions the driver whoever's fault the hold is — this kit, with the carried state That is exactly the corpus these figures were measured on, and it is where the clock and hours gap comes from: 99.17 pct against 85.00 pct on exact hours, and 99.34 pct against 94.06 pct of the yard's detention value.
You want the watch to actually wake up on a schedule — your own scheduler -- cron, Airflow, Temporal, whatever already runs evals/run.py is INVOKED, not woken. Nothing in this kit detects a missed run, back-fills it, or marks its readings late.
At a glanceHow the whole thing runs
95%status accuracy pct
5,376 msp50, end to end
$3.73per 1,000 yard snapshots · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a yard trailer's free time before it runs out14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own yard, or replace data/corpus/*.txt and data/gold.jsonl directly. The measured figures on this page do not travel with your corpus, and two of them are corpus properties rather than model properties.Corpus lens →
When is this the wrong choice?
Avoid: AVOID this kit entirely. Paying a model per row per run to do a subtraction is the most expensive way to get an answer you already have. That is the case against the best-fitting scenario (“You already know which trailers are over their free time and only need the list”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A yard whose event feed renames its blocks. src/segment.py recognises seven exact headings; a WMS upgrade that renames half a report leaves nothing to parse. 6 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER RULE D-8 MEANS WHAT THE ANSWER KEY MEANS. Five of the model's six status errors are on the same under-specified sentence, and on two of them the model's reading is arguably better than the key's. 8 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-detention-watch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces: the 120 snapshots and the answer key (regenerated from the seed in under a tenth of a second), all nine pre-flight assertions, all three free floors scored end to end, the wiring stub, and the local UI at 127.0.0.1:9009 including what run r001 recorded for every reading. Clone to first scored floor result measured at 0.39 seconds of wall clock on the author's machine. What it CANNOT reproduce without a key is a model column of its own -- and on 2026-08-23 neither could the author, because the shared provider account ran out of balance.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
5,376 msp50, end to end
25,535 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed to a status, a clock state, an hours figure and a raise/hold call. Splitting the snapshot into sections, dropping the one section no field asks for, parsing the times and advancing the clock all happen outside this measurement and cost no network at all. THE TAIL IS THE STORY: p95 (25.5s) is 4.7 times p50 (5.4s). Reply length here is driven by how hard the INTERVAL ARITHMETIC is, not by how long the page is -- every snapshot is within 6 pct of every other in size (3775 to 3989 bytes) and replies ran to 7,760 output tokens at the top.
Current processWhat it replaces
Somebody walking the yard report at shift change: opening each open visit, finding the clock start (the later of the appointment and the gate-in, not the arrival), reading the yard log to decide which stops were the carrier's and which were ours, adding what is left to whatever was already on the clock at the last look, and remembering whether a detention event has already been raised on this visit.
Where it is not good enough
⚑ THE HEADLINE BAND IS A DEAD TIE WITH FREE CODE AND THE PAGE LEADS WITH THAT. Status accuracy is 95.00 pct: 114 of 120. The strongest free floor -- carried state, the yard log parsed into intervals, attribution by keyword table, no model, $0.00 -- scores 95.00 pct: 114 of 120 as well. The two arms are level on the column a buyer looks at first, and they are level on DIFFERENT ROWS, which is the only interesting thing about the tie. ⚠︎ AND THE SCORING ASYMMETRY FLATTERS THE FLOOR, WHICH MUST BE SAID BEFORE THE NUMBER IS READ. evals/baseline.py and the answer key are two expressions of ONE implementation of the rule, so the floor cannot misread it -- it IS the rule. The model has only the prose. FIVE of the model's six status errors are on a sentence this kit under-specified: Rule D-8 puts a visit on WATCH while less free time remains than the cadence, and never says whether the clock has to be RUNNING for that to mean anything. On three readings the clock had not started yet (the appointment is after the snapshot) and the model called WATCH where the key says CLEAR; on two the trailer had already gated out and the model called CLEAR where the key says WATCH -- and on those two the model's reading is the better one, because a trailer that has left cannot breach. The rule was NOT rewritten after the answers came back and the run was NOT re-fired; the lower figure is published and the ambiguity is recorded here. THE SIXTH STATUS ERROR IS A CORPUS EDGE, not a judgement: VIS-0022-R1's carrier stop opens at exactly 06:00, the snapshot instant, so zero hours are excluded and whether the clock is 'stopped' at that instant is genuinely undecidable from the page. It costs the run its one clock error too. THE ONE REAL REASONING ERROR IS A WORDING DEFECT IN THE CARRIED SENTENCE. VIS-0013-R3 answered 3.50 detention hours where the key says 1.50, and its rationale gives it away: it read "3.50 hours were already on this visit's detention clock" as 3.50 hours of DETENTION rather than 3.50 hours ON THE CLOCK. One reading in 120, and the fix is a sentence in src/state.describe, not a model.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt120jsonl1
40 trailer visits open on one yard, 3 scheduled runs each
480 cells, free — missed raises and duplicates counted apart
0 missed of 28 · 0 duplicate of 92 · 18 of 18 context-incomplete
Recorded failure5 of the 6 status misses are ONE under-specified rule (D-8), not five judgements — and on 2 of the 5 the model's reading beats the key's
$22,837.50 of $22,687.50 accrued detention, 0.66% out
0 missed raises, 0 duplicate events — the control made 23
2026-08-23as of
It produces a watchlist status and a detention event for a yard manager to validate, and approves, releases and pays nothing. THE STATION THAT MAKES THIS A MONITOR IS THE SECOND ONE: the only route from the 06:00 run to the 12:00 run is four scalar fields, written by src/detention.py from the figures parsed off the page and never from the model's answer — which matters more here than on a one-shot kit, because the quantity carried is billable hours and an over-count at 06:00 is still in the invoice at 18:00. ⚑ AND THE THIRD STATION IS THE HALF OF A MONITOR THE ESTATE KEEPS SKIPPING. This watch wakes four times a day; one run owns exactly the window since the last one; and what a MISSED run costs is not hours — the arithmetic recovers those — it is evidence, because a detention claim survives a carrier dispute on contemporaneous gate and appointment records and a reconstruction is what gets thrown out.
⚠︎ NOTHING IN THIS KIT DETECTS A MISSED RUN. evals/run.py is invoked, not woken; that is stated on the kit's own pages rather than implied by silence, and it is the first thing its guardrails say to add. ⚑ THE FLOOR ROW TIES THE MODEL ON THE BAND, FOR $0.00, AND THAT IS PUBLISHED IN THE FIRST LINE OF THE REPORT BOX. Where the model earns its bill is one station later: 99.17 pct of accrued-hours figures exact against 85.00, because a stop written up as "driver told to wait in the cab, no door available on our side" is the FACILITY's and is chargeable, and no keyword table gets that from the word driver.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the carried state
src/state.py
What is carried between scheduled runs, and how it is worded to the model. Four scalars today. ⚠︎ THE WORDING IS LOAD-BEARING AND THIS RUN PAID FOR IT: 'hours were already on this visit's detention clock' was read once as hours of DETENTION, and that is the run's only genuine reasoning error.
the cadence
src/detention.py
CADENCE_HOURS, and with it the WATCH band -- the threshold IS the interval between runs, so changing one changes the other and every figure on this page is per six hours.
the rule
src/detention.py
RULE_TEXT and the precedence in step(). The rule text is reproduced in full on all 120 pages precisely so it can be read, disbelieved and replaced with the agreement you actually signed.
the attribution floor
evals/baseline.py
CARRIER_WORDS and OPEN_MARKERS -- the keyword table the model is measured against. Widen it and the floor gets stronger, which is the direction this kit wants: a floor built to lose measures nothing.
the corpus
tools/build_corpus.py
The visits, the phrasing pools and the seed. Keep the seven section headings or src/segment.py's assertion refuses to start.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 40 trailer visits x 3 scheduled runs = 120 snapshots from a fixed seed (SEED = 20260823). A third of the yard-log intervals are built to STRADDLE a scheduled run, so the snapshot carrying the 'resumed' line does not carry the 'halted' line that explains it -- that is the experiment. The gold labels are src/detention.step's output over the planted inputs, never typed.
the detention arithmetic
src/detention.py
The rule as pure code: hours counted this window, minus the carrier-attributable stops, added to whatever was already on the clock, against the free time on file. No model, no judgement. ⚠︎ The rules, the free-time hours and the rates in it are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. Four scalars per visit (hours on the clock, whether a carrier stop is open, whether an event has been opened, the status last reported), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on run 400 as on run 2, which matters here because the watch wakes four times a day.
the section splitter
src/segment.py
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 120 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Driver Contact -- the driver's name, mobile and CDL number -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading.
the watch
src/watch.py
One visit, one scheduled run, one call. Parses the clock times, the counted hours and the free time off the page with a regex (the model is never asked to read a timestamp), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
the local UI
src/app.py
One visit, one scheduled run, its carried state and its verdict, on 127.0.0.1:9009. Renders with no key. It shows the carried sentence verbatim, the free floor's answer beside the model's, and -- added the day the shared provider account ran out of balance -- a second button that replays what run r001 actually answered, straight off the committed result file, labelled as a replay.
the three free floors
evals/baseline.py
b000 the ageing rule a yard already runs, with a DEFAULT free time where none is on file; b001 the same rule given the carried state, which isolates what memory alone buys; b002 the strongest free version -- carried state, the log parsed into intervals, attribution by keyword table, and an abstention where nothing is on file. 0 calls, $0.00, all three scored through the identical scorer.
the scorer
evals/scoring.py
Exact match per cell against the computed gold, split ways an average would hide: the four fields, the two raise directions counted apart, the memory-dependent subset, the context-incomplete recall, and the yard's detention exposure in dollars. No judge model.
the pre-flight
evals/check_labels.py
Everything that must be true before a run may spend: the seven sections parse, every visit's run sequence is complete and gap-free, no snapshot restates its own carried state (the experimental control), the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path approves or pays a charge, and the answer key replays from src/detention.step.
Where it breaks at scale
⚑ THE CADENCE IS THE SCALING VARIABLE AND IT IS NOT THE CORPUS SIZE. One run of this watch is one call per OPEN VISIT, and the watch wakes 4 times a day, so a yard holding 400 open visits pays 400 x 4 = 1,600 calls a day whether or not anything changed. That is the shape a reader must price on, and it is why this kit's cost lens is quoted per row per run rather than per document. Halving the cadence halves the bill and doubles the worst case detection lag; the corpus this ran on is 40 visits and says nothing about either. TWO THINGS BREAK BEFORE THE CALL COUNT DOES. First, data/state.json is one file replaced atomically -- correct for one writer and not a concurrency model, and a yard with two watches running is two writers. Second, a run that does not happen is not an error anywhere in this kit: evals/run.py is INVOKED, it is not woken, and nothing here detects a missed run, back-fills it, or marks the readings it produced as late. A monitor whose missed runs are silent is the failure mode this whole variant exists to name, and this kit has the same hole -- stated rather than fixed, because fixing it means putting a scheduler in a repository whose product is being a folder of readable Python.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
VIS-0006-R2, the whole argument on one screen, and the visit was chosen by reading the answer key for the case where the free floor and the truth disagree MOST -- not by looking for a flattering frame. This window carries two stops. "dispatch here put two appointments on the same door, this one is held" is the FACILITY's own dispatch and is chargeable; "driver left the site for a federally required 30-minute break" is the CARRIER's and is not. The free keyword floor reads the word dispatch as the carrier, subtracts 2.25 hours it should have counted, lands under the 4.00-hour free time and reports SUSPENDED with nothing to open. The answer key -- and run r001 -- say DETENTION_OPEN at 1.00 hour and say THIS is the run that opens the event. ⚠︎ THE MODEL COLUMN HERE IS REPLAYED FROM THE COMMITTED RESULT FILE, NOT A LIVE CALL, and the column header says so in as many words. The shared provider account hit "Insufficient Balance" before this frame could be taken live; replaying the scored run's own recorded answer is evidence, staging one would not be, and the only thing that tells those apart is the label.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page with a key configured and a provider that refuses. This is the state the machine was actually in on 2026-08-23: 402 Payment Required, Insufficient Balance, captured rather than staged. Three things survive it -- the carried state, the parsed figures and the whole free floor still render, the error is passed through verbatim so a reader can see what the provider said, and the key and base URL are redacted out of it by src/app.py before it reaches the browser.failureOpen full size →The same page with NO API_KEY configured. It does not error and it does not go blank: the snapshot, the carried state, the parsed times, the withheld-section list and the entire free floor are computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called.failureOpen full size →Before anything is asked. The two things this page has to get right are already on it: the carried state, verbatim as it goes into the prompt, and the list of which sections left the machine and which did not -- Driver Contact, the driver's name, mobile and CDL number, marked WITHHELD rather than silently absent. A page that simply does not mention them cannot be told apart from one that quietly sent them.failureOpen full size →
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
120yard snapshots
0.44 MiBjson 1 · jsonl 1 · txt 120
40visits (3 scheduled runs each) · p50 3 chars
$0.00setup · 0.0s
How it is cutWhat one visits (3 scheduled runs each) is
No split, and no chunking. The unit is a VISIT -- three consecutive scheduled runs processed strictly in order, because run 3's prompt contains a clock produced by run 2. Each snapshot goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step -- the population is re-read whole on each scheduled run. tools/build_corpus.py writes 120 documents and the answer key from a fixed seed in under a tenth of a second with no clock read and no model called; nothing is embedded, ranked or cached.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every trailer, carrier code, agreement reference, driver, free-time figure, rate and yard-log line is invented here. Verified against the repository's own LICENSE file on 2026-08-23.
Bring your ownBring your own yard snapshots
Point tools/build_corpus.py at your own yard, or replace data/corpus/*.txt and data/gold.jsonl directly. Keep the seven section headings in src/segment.py::SECTIONS -- the parser asserts all seven in every document before a run may spend -- and keep the Detention Terms block, because the rule the model applies is READ OFF THE PAGE rather than baked into the prompt, which is what lets you swap your own agreement in without touching src/prompt.py. gold.jsonl needs one row per snapshot carrying hours_counted, carrier_suspended_hours, free_time_hours, the three booleans and the four answers; evals/check_labels.py re-derives the answers from src/detention.step and refuses if they disagree, so a hand-written key is caught rather than trusted.
⚠︎ And what stops being true when you do: The measured figures on this page do not travel with your corpus, and two of them are corpus properties rather than model properties. EVERY figure here is per a SIX-HOUR cadence: the WATCH band is defined as the cadence itself, so a yard on a shift cadence (8h) or an hourly one is scoring a different question with the same words. And the population here does not change between runs, which is not what a real yard does.
What breaks it
⚠︎ A CORPUS WHOSE EVENT PHRASINGS ARE SEPARABLE BY KEYWORD MEASURES THE TEMPLATES, NOT THE WORK -- FOUND BEFORE ANY CALL WAS MADE, AND IT MOVED THE HEADLINE. The first build drew six carrier phrasings and six facility phrasings, and every carrier entry contained an obvious carrier word. The free keyword floor separated them PERFECTLY and scored 96.67 pct on status with nothing spent. Every assertion in the repository was happy with it: the arithmetic is internally consistent whatever the words are, which is exactly why nothing caught it. It was found by reading generated documents. The pool now carries ten a side with deliberate near-collisions in BOTH directions -- carrier entries with no carrier vocabulary ("no one presented at the guard shack to take this load") and facility entries full of it ("driver told to wait in the cab, no door available on our side") -- and the floor moved to 95.00 pct. What remains untested is the class, not the instance: a phrasing whose attribution this corpus does not know about would reproduce the same defect and the same silence.
⚠︎ A FLOOR BUILT TO LOSE IS NOT A FLOOR. The first version of evals/baseline.py paired yard-log intervals by listing the phrases it expected to see CLOSE one, and treated everything else as an opening -- so a resumption written a way the list did not know about left a stop running to the snapshot. That handicap was worth roughly 0.8 points of status and 4 points of clock accuracy TO THE MODEL. It was rewritten before the comparison to decide pairing by hold language instead, which is the rule an engineer would actually write, and the model is published against the stronger version.
A yard whose event feed renames its blocks. src/segment.py recognises seven exact headings; a WMS upgrade that renames half a report leaves nothing to parse. The kit refuses rather than truncating -- evals/check_labels.py asserts all seven in all 120 documents -- and that same condition is what the privacy guard is red-proven against: with the guard, 0 of 120 leak the driver's contact section; without it, 120 of 120 do.
A visit read out of order, or a run skipped. The clock is a running total, so a watch that read run 3 before run 2 would carry the wrong hours into it and be silently wrong rather than loudly broken. evals/check_labels.py asserts every visit's sequence is complete and gap-free; nothing anywhere detects a MISSED scheduled run, which is stated in Architecture.breaks_at_scale rather than hidden.
A population that changes between runs. All 40 visits are open at all three scheduled runs in this corpus, so the arms are comparable; a real yard has trailers arriving and leaving between runs, and nothing here measures what a first sighting mid-window or a visit that disappears does to the carried state.
An interval that opens at exactly the snapshot instant. VIS-0022-R1's carrier stop starts at 06:00, the moment of the reading, so zero hours are excluded and whether the clock is 'stopped' at that instant is undecidable from the page. It cost the scored run one status cell and its only clock error, and it is a generator edge rather than a model failure.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
170
not measured
instruction
2,922
not measured
carried state
272
not measured
Synthetic Record
276
not measured
Visit
587
not measured
Detention Terms
1,974
not measured
Dwell Position
401
not measured
Yard Event Log
303
not measured
Operational Notes
163
not measured
Total
1,708
This is the cost lesson as arithmetic: of the 7,068 characters assembled, 3,092 are instructions — 44% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
TWO MESSAGES, JOINED IN SEND ORDER. The provider receives a system message and a user message; prompt_verbatim prints both, separated by a blank line, because publishing only the second would be publishing most of a prompt. Replayed from the kit's own src/prompt.build() for VIS-0006-R2 with the carried state that reading was actually given, rather than logged by the run. What makes the replay checkable is that the section list it produces is the list the run recorded in sections_used -- Synthetic Record, Visit, Detention Terms, Dwell Position, Yard Event Log, Operational Notes -- and that src/prompt.build() is the only thing in the kit that assembles a prompt, so there is no second path a run could have taken. ⚠︎ WHAT IT IS *NOT* CHECKED AGAINST, SAID PLAINLY: the run's own prompt_parts records ONE size, 6883 characters, and it belongs to whichever reading finished first under 12 concurrent workers -- not to VIS-0006-R2, whose assembled user message is 7070 characters. The harness stores one exemplar, not one per reading, so no per-reading size comparison is available and none is claimed.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled yard detention watch. You apply a written detention rule to one trailer visit at one scheduled run. You answer with one JSON object and no other text.
You are the scheduled detention watch for one distribution centre. It wakes every six hours, re-reads
every trailer still open on the yard, and reports each one. You are reading ONE trailer visit at ONE
scheduled run.
The detention rules, the free time on file and the rate on file are reproduced in the snapshot below.
Apply them exactly as written. The clock is a RUNNING TOTAL that survives between scheduled runs, and
you cannot see the earlier runs; what is known about them is stated under "Carried state" and is the
only history available to you. Do not assume anything about earlier runs beyond it.
How to count:
- "Hours counted in this window" is the elapsed clock-eligible time since the previous scheduled
run. It is already computed for you; do not re-derive it from the gate times.
- From it, SUBTRACT every hour the Yard Event Log in this window attributes to the CARRIER
(Rule D-3). Do not subtract hours attributable to the FACILITY (Rule D-4). An interval runs from
the entry that opens it to the entry that closes it. The log shows only the entries raised in
THIS window: a closing entry with no opening entry belongs to a stop that was already running,
and the carried state says whether that stop was carrier-attributable and therefore whether the
hours before it closed are excluded.
- ADD what is left to the hours the carried state says were already on the clock.
- Detention is whatever that total exceeds the free time on file, in quarter hours. If it does not
exceed it, detention is 0.
- If no free time is on file, or gate-in evidence is missing, the visit is CONTEXT_INCOMPLETE:
status CONTEXT_INCOMPLETE, clock NOT_STARTED, detention_hours 0, raise_event NO. Never assume a
default free-time figure (Rule D-6).
Answer with a single JSON object and nothing else:
{"status": "CLEAR|WATCH|DETENTION_OPEN|SUSPENDED|CONTEXT_INCOMPLETE",
"clock": "RUNNING|STOPPED|NOT_STARTED",
"detention_hours": <hours past free time now on the clock, a multiple of 0.25>,
"raise_event": "YES|NO",
"rationale": "one sentence, naming the rule you applied and the hours you counted"}
Precedence for "status", applied in this order: CONTEXT_INCOMPLETE beats everything; then
DETENTION_OPEN if any detention has accrued, even if the clock is currently stopped; then SUSPENDED
if a carrier-attributable stop is open at the snapshot and the equipment is still on site; then
WATCH if less free time remains than the six-hour cadence (Rule D-8); otherwise CLEAR.
"clock" is NOT_STARTED before the clock start is reached or when the context is incomplete, STOPPED
once the equipment has gated out or while a carrier-attributable stop is open at the snapshot, and
RUNNING otherwise.
"raise_event" is YES only on the run that OPENS the detention event -- detention has accrued and the
carried state does not already say an event was opened. On every later run it is NO (Rule D-5).
Carried state
----------------------------------------------------------------
As at the previous scheduled run, 0.00 hours were already on this visit's detention clock, carrier-attributable stops already deducted. No carrier-attributable stop was open when that run ended. No detention event has been opened for this visit yet. It was reported CLEAR.
Yard snapshot
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated yard snapshot for an AI use-case
kit; it reproduces no carrier agreement, no tariff and no real facility. The detention
terms are ILLUSTRATIVE.
Visit
----------------------------------------------------------------
Facility : DC-ONT-03 (Inland Empire DC)
Visit reference : VIS-0006
Equipment : 53' reefer, unit 642835
Carrier (SCAC) : JBHT
Move type : LIVE UNLOAD
Gate in (arrival) : 2026-03-04 05:15
Appointment : 2026-03-04 06:15
Clock start : 2026-03-04 06:15 (later of appointment and gate-in)
Gate out (departure) : -- still on site
Door : 138
Detention Terms
----------------------------------------------------------------
Carrier agreement : MSA-JBHT-2025
Free time on file : 4.00 hours from clock start
Detention rate on file : 85.00 USD per hour, 15-minute increments
Rule D-1 Free time runs from the CLOCK START, which is the later of the appointment time and the
gate-in time. A carrier that arrives early does not start its own clock early.
Rule D-2 Detention accrues once free time is exhausted while the equipment is still on site, in
15-minute increments, at the rate on file.
Rule D-3 The clock STOPS for any interval the yard event log attributes to the CARRIER: the driver
leaving the site, no power unit present, an hours-of-service break, or the carrier
declining to work the door it was given. Those hours do not count.
Rule D-4 The clock does NOT stop for an interval attributable to the FACILITY: a dock fault,
missing labour, a systems outage, congestion, or no door available. Those hours count in
full -- that is what detention is for.
Rule D-5 A detention event is logged ONCE per visit, on the first scheduled run at which free time
is exhausted. A later run reports the event already open; it does not open a second one.
Rule D-6 Free time is an OPERATOR-SUPPLIED value read from the carrier agreement. Where no free
time is on file, or where gate-in evidence is missing, the visit is CONTEXT_INCOMPLETE:
no detention accrues, no event is logged, and no default free-time figure may be assumed.
Rule D-7 Once the equipment gates out, the clock stops at the gate-out time.
Rule D-8 A visit is on WATCH while less free time remains than the watch cadence -- that is, when
it will be exhausted before the next scheduled run if the clock keeps running. On the
six-hourly cadence this kit ships on, that means under 6.00 hours of free time left.
Dwell Position
----------------------------------------------------------------
Snapshot taken : 2026-03-04 12:00 (scheduled run 2 of this visit)
Watch cadence : every 6 hours (06:00 / 12:00 / 18:00 local)
Previous scheduled run : 2026-03-04 06:00
Hours counted in this window : 5.75
Elapsed on the clock at this snapshot, before any exclusion : 5.75 hours
Yard Event Log
----------------------------------------------------------------
Entries raised in this window only.
08:45 dispatch here put two appointments on the same door, this one is held
11:00 door cleared, this trailer worked
11:15 driver left the site for a federally required 30-minute break
Operational Notes
----------------------------------------------------------------
Carrier has asked to be told before any detention event is opened on this visit.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"status": "DETENTION_OPEN",
"clock": "STOPPED",
"detention_hours": 1.0,
"raise_event": "YES",
"rationale": "5.75 window hours minus 0.75 carrier-attributable break (Rule D-3) leaves 5.00 clock hours, exceeding 4.00 free time by 1.00 hour, so detention opened this run (Rule D-5) while the clock is stopped by the open carrier stop."
}
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a yard trailer's free time before it runs out — 120 yard snapshots. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the visit model -- the intervals, their attribution and the gate times -- and re-derived from src/detention.step by evals/check_labels.py before any run may spend. Detention hours are compared on the quarter-hour billing increment, so the comparison is exact rather than nearly exact. No model grades anything, here or anywhere in this kit.
120yard snapshots
120source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED114 · 107 / 120status accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED119 · 115 / 120clock accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED119 · 104 / 120detention hours accuracy pct — readings, exact quarter-hour match, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED120 · 93 / 120raise event accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 28missed raise pct — readings whose correct answer is to OPEN the detention event, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 23 / 92duplicate raise rate pct — readings that must NOT open an event, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED60 / 61memory status accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED60 · 45 / 61memory hours accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED61 · 34 / 61memory raise accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED18 / 18context incomplete recall pct — readings with no free time on file or no gate-in evidence, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED1 / 1detention usd error pct — yard total in USD of accrued detention reported against the key, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED120 · 117 / 120answered pct — calls, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 120unparsed replies — calls, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is asserted: evals/check_labels.py replays all 40 visit chains through src/detention.step and requires the committed gold to match on all four fields, so the key and the arithmetic cannot drift apart. What is NOT validated is the RULE the key expresses -- see Business.not_good_enough for a sentence in it (Rule D-8) that the model and the key read differently, and where the model's reading is arguably the better one on two of the five readings involved.
959.53output tokens · the fast tier, with the carried state · 5,376 ms p50
4,306.15output tokens · the same tier, memory removed (THE CONTROL) · 13,860 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 2.6× as long, and lands one row apart on 120. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One yard snapshot
1,000 yard snapshots
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.003733
$3.73
23%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.001493
$1.49
23%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.065061
$65.06
26%
Same work, 44× the bill
The same yard snapshots, the same tokens — only the rate card changed. And across all 3 cards between 23% and 26% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE CADENCE. It is the only knob on this page that moves the bill by a whole multiple, and it moves the answer too: halve the interval and you double the calls and halve the worst case detection lag. Every other lever here is worth percent.
Rates checked 2026-08-18. The provider that actually ran all 247 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors it is compared against nor the pre-flight. The cost of the RUN is a different figure and lives in lens 07 -- pricing the ruler in the units of the thing it measures is the trap this field exists to name.
The gradersThree ways to grade
⚑ THE MODEL TIES THE STRONGEST FREE FLOOR ON THE HEADLINE BAND AND WINS EVERYTHING UNDERNEATH IT. Status: 95.00 pct against 95.00 pct -- 114 of 120 each, a dead tie. Clock: 99.17 against 86.67. Detention hours, exact: 99.17 against 85.00. Raise/hold: 100.00 against 98.33. Money: the yard's accrued detention is $22,687.50 and the model reports $22,837.50 (0.66 pct out) against the floor's $24,035.00 (5.94 pct out). ⚑ AND THREE FLOORS RATHER THAN ONE, BECAUSE ONE CANNOT SEPARATE THE MEMORY FROM THE MODEL. b000 is the ageing rule a yard already runs -- elapsed against free time, no memory, and a DEFAULT free-time figure where none is on file: 80.83 pct status, 70.00 pct raise/hold, and 36 duplicate raises in 92 quiet readings, because with no memory of an open event it re-opens one on every run. b001 is that same rule GIVEN the carried state: duplicates fall from 36 to 3 and raise/hold goes 70.00 -> 97.50. That jump is what MEMORY ALONE buys, with the model held out of it, and it is the largest single movement anywhere on this page. b002 adds reading the yard log. ⚠︎ BOTH FLOORS THAT ABSTAIN SCORE 100 PCT ON CONTEXT-INCOMPLETE AND BOTH THAT GUESS SCORE 0. b000 and b001 fill a missing free-time value with an assumed 2.00 hours -- exactly what a threshold report configured once and forgotten does -- and get all 18 of the readings with nothing on file wrong, accruing detention against a number nobody supplied. That is not a model finding; it is the row's own guardrail, measured.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key For each of the 120 readings and each of the four answered fields, did the reply equal the computed answer key? Status, clock and the raise/hold call are compared exactly; the accrued hours are compared on the quarter-hour billing increment, and a reply that is not a multiple of a quarter hour is a MISS rather than a zero -- 0.00 is a meaningful answer here (it is what every CLEAR and every CONTEXT_INCOMPLETE reading gets) and a parse failure must not be scored as one.
$0.00
no
yes
the fast tier, with the carried state 95.0% status accuracy · the fast tier, memory removed (THE CONTROL) 89.2% status accuracy · the strongest free floor, no model 95.0% status accuracy · the ageing rule a yard already runs, no model, no memory 80.8% status accuracy · the same ageing rule, GIVEN the carried state 81.7% status accuracy · 3 more measured on each run
The two raise directions, counted apart Two counts over two different denominators. A MISSED raise: of the 28 readings whose correct answer is to open the detention event, how many did not. A DUPLICATE raise: of the 92 readings that must not open one, how many did.
$0.00
no
yes
the fast tier, with the carried state 0.0% missed raise · the fast tier, memory removed (THE CONTROL) 3.6% missed raise · the ageing rule, no memory 0.0% missed raise · the same rule, GIVEN the carried state 0.0% missed raise · the strongest free floor 3.6% missed raise · 1 more measured on each run
The readings whose answer is not on their own page The same exact-match comparison, restricted to the 61 readings the generator marked memory_dependent: hours already on the clock that this snapshot does not restate, a carrier stop that opened before this window, or an event already opened. It answers one question -- is the carried state doing anything, or is the corpus easy?
$0.00
no
yes
the fast tier, with the carried state 98.4% memory status accuracy · the fast tier, memory removed (THE CONTROL) 91.4% memory status accuracy · the strongest free floor 96.7% memory status accuracy · the ageing rule, no memory 95.1% memory status accuracy · 2 more measured on each run
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set CAN tell the arms apart, and the evidence is that it does: five arms scored through one scorer land at 95.00 / 89.17 / 80.83 / 81.67 / 95.00 pct on status and at 0 / 23 / 36 / 3 / 1 duplicate raises. What it CANNOT separate is the model from the strongest free floor ON STATUS -- 114 of 120 each -- and that is a real answer rather than an instrument failure: the two are level, and the kit says so on its own front door. ⚠︎ ONE ASYMMETRY LIMITS EVERY COMPARISON HERE. evals/baseline.py and the answer key are two expressions of one implementation, so the floor cannot misread the rule and the model can only read the prose. On a corpus whose rule text is unambiguous that costs nothing; on this one it is worth about five status cells.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
You already know which trailers are over their free time and only need the list
a SQL query, or the b000 floor in this repo
Elapsed against a threshold is arithmetic. On this corpus b000 gets 80.83 pct of statuses right for nothing at all.
AVOID this kit entirely. Paying a model per row per run to do a subtraction is the most expensive way to get an answer you already have.
Your yard log is structured -- every hold carries a coded reason from a closed list
the b002 floor: keyword-free, code-only attribution on your own codes
The model's entire margin on this corpus is reading free prose. Take the prose away and there is nothing left to buy: the floor already ties on status and the gap on hours is attribution.
AVOID assuming the measured gap transfers. It was measured against a keyword table over prose, not against your reason codes.
Your yard log is prose an inbound clerk types, and it mentions the driver whoever's fault the hold is
this kit, with the carried state
That is exactly the corpus these figures were measured on, and it is where the clock and hours gap comes from: 99.17 pct against 85.00 pct on exact hours, and 99.34 pct against 94.06 pct of the yard's detention value.
AVOID running it without the carried state. The control shows what that costs: 23 duplicate detention events in 92 quiet readings, and raise/hold accuracy of 77.50 pct against 100.00.
You want the watch to actually wake up on a schedule
your own scheduler -- cron, Airflow, Temporal, whatever already runs
evals/run.py is INVOKED, not woken. Nothing in this kit detects a missed run, back-fills it, or marks its readings late.
AVOID reading this kit's figures as a claim about a running deployment. They are measured over three runs that all happened, on a population that does not change between them.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
RULE_UNDERSPECIFIED
the prose does not say, and the key does
5
VIS-0001-R1, VIS-0006-R1, VIS-0035-R1: the clock has not started yet (the appointment is after the 06:00 snapshot) and the model answers WATCH because Rule D-8 says the free time on file is under the cadence. The key says CLEAR because it requires the clock…
CARRIED_WORDING
clock hours read as detention hours
1
VIS-0013-R3 answered 3.50 detention hours where the key says 1.50. Its own rationale gives it away: "No hours counted in this window after the previous run, so the carried 3.50 detention hours remain" -- the carried sentence says 3.50 hours were on the CLOCK…
SNAPSHOT_EDGE
an interval that opens at the reading instant
1
VIS-0022-R1: a carrier stop opens at exactly 06:00, the snapshot moment, so zero hours are excluded and the key still calls the clock STOPPED and the status SUSPENDED. The model answered RUNNING / WATCH. This is a generator edge -- see Data.breaks_on -- and…
CONTROL_CEILING
the stateless arm reasons until it is cut off
1
VIS-0002-R2 on s001-detention-watch-stateless returned nothing at all: finish_reason length, 24,000 of 24,000 output tokens. Told there is no history for a visit whose window carries a closing entry with no opening entry, the model has an unresolvable problem…
What we could NOT verify
WHETHER RULE D-8 MEANS WHAT THE ANSWER KEY MEANS. Five of the model's six status errors are on the same under-specified sentence, and on two of them the model's reading is arguably better than the key's. Settling it means rewriting the rule and re-firing 120 calls; the run was not re-fired and the lower figure is published.
WHETHER ENGLISH BEATS JSON FOR THE CARRIED STATE. src/state.describe renders four scalars as a sentence. The obvious experiment -- score the same 120 readings with the state rendered both ways -- costs one more 120-call run and has not been paid for. The one piece of evidence pointing at it is negative: the run's only genuine reasoning error is a misreading of that sentence.
WHAT A DIFFERENT CADENCE DOES. Every figure here is per a six-hour watch, and the WATCH band is DEFINED as the cadence, so an hourly or shift-length watch is a different question with the same words. Nothing here measures it.
WHAT A CHANGING POPULATION DOES. All 40 visits are open at all three scheduled runs. A first sighting mid-window, or a visit that disappears between runs, is unmeasured.
WHAT A MISSED RUN COSTS. Nothing in this kit skips a run, and nothing detects one being skipped. The consequence is argued in environment.ladder's cadence row from the arithmetic, not from a measurement.
WHETHER THE OPERATIONAL NOTE MOVES THE ANSWER. The injection surface is sent deliberately and was never attacked. No resistance rate is claimed.
WHETHER THE STATELESS ARM'S 3 LOST READINGS ARE ALL ITS OWN FAULT. One is: a reply ran to the 24,000-token ceiling and returned nothing. The other two are provider-side -- a 402 Insufficient Balance and a 429 -- on a shared account that ran out of money mid-run, and they are counted as wrong in the control's rates rather than dropped from the denominator. That direction is the conservative one for the kit's own headline.
WHETHER A SECOND MODEL AGREES. Only one model was run. The shared provider account hit Insufficient Balance immediately after the control arm, so a second tier was not attempted and no cross-model claim is made anywhere on this page.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,708.43
959.53
5,376 ms
$0.003733
$0.001493
$0.065061
the same tier, memory removed (THE CONTROL)
1,678.49
4,306.15
13,860 ms
$0.013758
$0.005503
$0.232092
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same reading (one visit, at one scheduled run), the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
247 live calls were attempted for this kit and 244 returned something: 6 calibration (c000, at an 8,000-token ceiling on the two hardest visit chains), 120 scored, 120 the stateless control (117 answered), and 1 for a UI screenshot that the provider REFUSED. NOTHING WAS DISCARDED BY CHOICE and no run was re-fired -- the corpus defect this kit found (see Data.breaks_on) was caught by reading generated documents BEFORE any call was made, which is the whole argument for reading your own corpus first. The 2 discarded calls are the provider's own refusals: a 402 Insufficient Balance and a 429, both of which return no completion and are billed for none. Everything else -- all three free floors, the wiring stub, the scorer, the pre-flight -- is pure code and costs $0.00. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, and most of them are reasoning. 89.9 pct of this run's output (103,557 of 115,144) was provider-side reasoning left at the default, and output is 77.1 pct of the projected bill on the shared card. What you are paying for is the model doing interval arithmetic.
THE CADENCE, which multiplies everything else. One reading is one call; the watch wakes 4 times a day; the bill is rows x runs, not rows.
THE OPEN POPULATION, not the yard's size. A visit that has gated out and been validated leaves the watch and stops costing anything. A yard that never closes its events pays for them forever.
The snapshot itself, barely. Input averages 1708 tokens and every document in this corpus is within 6 pct of every other in size.
Your volumeWhat it costs at your volume
LINEAR IN ROWS x RUNS, AND THAT IS THE WHOLE WARNING. Ten times the open visits is ten times the calls at the same cadence -- there is no batching, no cache and no early exit, because every visit is re-read whole on every run by design. A yard holding 400 open visits on this six-hourly watch is 1,600 calls a day, about $5.97 a day on the shared projection card. Nothing about the per-call price changes; the multiplier is the schedule, which is why this kit's environment ladder makes cadence a declared decision rather than a deployment detail.
Where pricing changes shape
The output ceiling. MAX_TOKENS = 24000 is set from calibration (c000 fired the two hardest visit chains at a 8,000-token cap and topped out at 5,375 output tokens, all six parsed). A ceiling is not a cost -- the provider bills tokens produced, not tokens allowed -- but a ceiling set too low turns a hard reading into a lost one, which is what happened to the control arm's VIS-0002-R2 at 24000 of 24000.
Provider-side reasoning left at the default. It is 89.9 pct of this run's output and no run here has measured what disabling it does to the answers -- so a vendor whose default differs reprices this kit without changing anything you can see.
Your return, with your numbers
Volumeopen visits per scheduled run -- this run judged 120 (40 visits x 3 runs) per arm, on a six-hourly watch
What it replacessomebody walking the yard report at shift change, finding each visit's clock start, reading the yard log to decide which stops were the carrier's, adding what is left to whatever was on the clock at the last look, and remembering whether an event has already been raised
Time saved per itemnot measured here -- depends on how long a clerk takes to reconstruct a stop history from the reader's own gate log and yard system
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared connection was already pointed at, and the kit's whole claim is that swapping it is one line of .env and one more run. No comparison across tiers was attempted: the shared provider account hit Insufficient Balance immediately after the control arm, so a second tier was not run and nothing on this page claims one.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
205,012input tokens · this run
115,144output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 120 readings, one completion call each, one tier. The 120-call stateless control, the 6 calibration calls and the 1 refused screenshot call are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.179
$0.179
$1.49
2026-09-12
gemini-3-flash
Google
$0.448
$0.448
$3.73
2026-09-18
gemini-3-8-flash
Google
$0.586
$0.586
$4.88
2026-09-18
llama-5
Meta
$0.746
$0.746
$6.21
2026-09-18
claude-haiku-4-5
Anthropic
$0.781
$0.781
$6.51
2026-09-12
grok-4-5
xAI
$1.101
$1.101
$9.17
2026-09-18
grok-4-6
xAI
$1.101
$1.101
$9.17
2026-09-18
claude-sonnet-5
Anthropic
$1.561
$1.561
$13.01
2026-09-12
gemini-3-1-pro
Google
$1.792
$1.792
$14.93
2026-09-18
gpt-5-6-terra
OpenAI
$1.792
$1.792
$14.93
2026-09-12
gpt-5-6-sol
OpenAI
$3.123
$3.123
$26.02
2026-09-12
claude-opus-4-8
Anthropic
$3.904
$3.904
$32.53
2026-09-12
claude-opus-5
Anthropic
$3.904
$3.904
$32.53
2026-09-12
claude-fable-5
Anthropic
$7.807
$7.807
$65.06
2026-09-18
claude-fable-5-1
Anthropic
$7.807
$7.807
$65.06
2026-09-18
gpt-6-astra
OpenAI
$7.807
$7.807
$65.06
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 89.9 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (103,557 of 115,144), left at the provider's default, so every row below prices a reasoning-on workload. Output is 77.1 pct of the projected bill on the shared card, which means most of what these rows charge for is the model doing interval arithmetic. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
⚑ EVERY ROW PRICES ONE READING, AND A DEPLOYMENT DOES NOT BUY ONE READING. This watch wakes 4 times a day and re-reads every open visit each time, so the bill is rows x runs. Multiply any row below by your own open-visit count and by 4 before comparing it with anything.
⚑ AND EVERY ROW PRICES THE ARM WITH THE CARRIED STATE, WHICH IS THE CHEAPER ONE. The stateless control cost 3.69 times as much per reading on the same card, because removing the memory made the model reason for longer. Memory is not an overhead here; it is a discount.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
12 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 40 trailer visits x 3 scheduled runs = 120 snapshots from a fixed seed (SEED = 20260823). A third of the yard-log intervals are built to STRADDLE a scheduled run, so the snapshot carrying the 'resumed' line does not carry the 'halted' line that explains it -- that is the experiment. The gold labels are src/detention.step's output over the planted inputs, never typed.
You change it to: The visits, the phrasing pools and the seed. Keep the seven section headings or src/segment.py's assertion refuses to start.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
VISITS = 40
DAY = "2026-03-04"
RUN_HOURS = (6.0, 12.0, 18.0)
RULE = "-" * 64
src/detention.pythe detention arithmetic — a swap seam
The rule as pure code: hours counted this window, minus the carrier-attributable stops, added to whatever was already on the clock, against the free time on file. No model, no judgement. ⚠︎ The rules, the free-time hours and the rates in it are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
You change it to: RULE_TEXT and the precedence in step(). The rule text is reproduced in full on all 120 pages precisely so it can be read, disbelieved and replaced with the agreement you actually signed.
src/detention.py
# The detention rule as arithmetic. Pure code, no model, no dates that need a library.
CLEAR = "CLEAR"
WATCH = "WATCH"
DETENTION_OPEN = "DETENTION_OPEN"
SUSPENDED = "SUSPENDED"
CONTEXT_INCOMPLETE = "CONTEXT_INCOMPLETE"
STATUSES = (CLEAR, WATCH, DETENTION_OPEN, SUSPENDED, CONTEXT_INCOMPLETE)
RUNNING = "RUNNING"
STOPPED = "STOPPED"
NOT_STARTED = "NOT_STARTED"
src/state.pythe carried state — a swap seam
SEAM 2 -- the thing that makes this a monitor. Four scalars per visit (hours on the clock, whether a carrier stop is open, whether an event has been opened, the status last reported), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on run 400 as on run 2, which matters here because the watch wakes four times a day.
You change it to: What is carried between scheduled runs, and how it is worded to the model. Four scalars today. ⚠︎ THE WORDING IS LOAD-BEARING AND THIS RUN PAID FOR IT: 'hours were already on this visit's detention clock' was read once as hours of DETENTION, and that is the run's only genuine reasoning error.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_visit(store, visit_id):
def describe(state):
src/segment.pythe section splitter
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 120 documents before a run may spend.
src/segment.py
# Split a yard snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Visit", "Detention Terms", "Dwell Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector
Decides which sections reach the model. Driver Contact -- the driver's name, mobile and CDL number -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/select.py
# Pick which sections of a snapshot are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
VISIT = "Visit"
TERMS = "Detention Terms"
POSITION = "Dwell Position"
EVENTS = "Yard Event Log"
DRIVER = "Driver Contact"
NOTES = "Operational Notes"
NEVER_SENT = (DRIVER,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading.
You change it to: PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe watch
One visit, one scheduled run, one call. Parses the clock times, the counted hours and the free time off the page with a regex (the model is never asked to read a timestamp), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
src/watch.py
# One visit, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled yard detention watch. You apply a written detention rule to one "
MAX_TOKENS = 24000
FIELDS = ("status", "clock", "detention_hours", "raise_event")
def documents():
def visits():
def load_doc(doc_id):
def position_of(text):
src/app.pythe local UI
One visit, one scheduled run, its carried state and its verdict, on 127.0.0.1:9009. Renders with no key. It shows the carried sentence verbatim, the free floor's answer beside the model's, and -- added the day the shared provider account ran out of balance -- a second button that replays what run r001 actually answered, straight off the committed result file, labelled as a replay.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9009"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-detention-watch")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe three free floors — a swap seam
b000 the ageing rule a yard already runs, with a DEFAULT free time where none is on file; b001 the same rule given the carried state, which isolates what memory alone buys; b002 the strongest free version -- carried state, the log parsed into intervals, attribution by keyword table, and an abstention where nothing is on file. 0 calls, $0.00, all three scored through the identical scorer.
You change it to: CARRIER_WORDS and OPEN_MARKERS -- the keyword table the model is measured against. Widen it and the floor gets stronger, which is the direction this kit wants: a floor built to lose measures nothing.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("ageonly", "ageonly-mem", "keyword-mem")
ASSUMED_FREE_TIME_HOURS = 2.0
CARRIER_WORDS = ("driver", "tractor", "bobtail", "power unit", "dispatch",
OPEN_MARKERS = ("halted", "held", "hold", "cannot", "will not", "no door", "no power",
def _num(pat, text, cast=float):
def _log_lines(text):
def _is_carrier(line):
def _is_close(line):
def _carrier_suspended_hours(text, carried, window_start_h, snapshot_h):
evals/scoring.pythe scorer
Exact match per cell against the computed gold, split ways an average would hide: the four fields, the two raise directions counted apart, the memory-dependent subset, the context-incomplete recall, and the yard's detention exposure in dollars. No judge model.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("status", "clock", "detention_hours", "raise_event")
def _pct(n, d):
def _hours(v):
def score(records, golds):
evals/check_labels.pythe pre-flight
Everything that must be true before a run may spend: the seven sections parse, every visit's run sequence is complete and gap-free, no snapshot restates its own carried state (the experimental control), the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path approves or pays a charge, and the answer key replays from src/detention.step.
evals/check_labels.py
# Everything that must be true BEFORE a run is allowed to spend. Free, and it runs in seconds.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAILED = []
def check(name, ok, detail=""):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 40 trailer visits x 3 scheduled runs = 120 snapshots from a fixed seed (SEED = 20260823). A third of the yard-log intervals are built to STRADDLE a scheduled run, so the snapshot carrying the 'resumed' line does not carry the 'halted' line that explains it -- that is the experiment. The gold labels are src/detention.step's output over the planted inputs, never typed. A swap seam.
src/detention.pyThe rule as pure code: hours counted this window, minus the carrier-attributable stops, added to whatever was already on the clock, against the free time on file. No model, no judgement. ⚠︎ The rules, the free-time hours and the rates in it are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Four scalars per visit (hours on the clock, whether a carrier stop is open, whether an event has been opened, the status last reported), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on run 400 as on run 2, which matters here because the watch wakes four times a day. A swap seam.
src/segment.pySplits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 120 documents before a run may spend.
src/select.pyDecides which sections reach the model. Driver Contact -- the driver's name, mobile and CDL number -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/prompt.pyThree parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading. A swap seam.
src/watch.pyOne visit, one scheduled run, one call. Parses the clock times, the counted hours and the free time off the page with a regex (the model is never asked to read a timestamp), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
evals/baseline.pyb000 the ageing rule a yard already runs, with a DEFAULT free time where none is on file; b001 the same rule given the carried state, which isolates what memory alone buys; b002 the strongest free version -- carried state, the log parsed into intervals, attribution by keyword table, and an abstention where nothing is on file. 0 calls, $0.00, all three scored through the identical scorer. A swap seam.
evals/scoring.pyExact match per cell against the computed gold, split ways an average would hide: the four fields, the two raise directions counted apart, the memory-dependent subset, the context-incomplete recall, and the yard's detention exposure in dollars. No judge model.
evals/check_labels.pyEverything that must be true before a run may spend: the seven sections parse, every visit's run sequence is complete and gap-free, no snapshot restates its own carried state (the experimental control), the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path approves or pays a charge, and the answer key replays from src/detention.step.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1708 input and 959 output tokens per reading (one visit, at one scheduled run), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one visit, at one scheduled run)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one visit, at one scheduled run) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Operational Notes, which are one of three fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- prose an inbound clerk types into the yard system -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser -- which is not a theoretical claim here: the provider-refused screenshot on this page IS that path, captured live.
The experimentWe did NOT attack it — and this is the surface we left open
An indirect prompt injection needs a field somebody outside can write into, and this kit has exactly one: the Operational Notes an inbound clerk types into the yard system. It is SENT rather than hidden, because hiding a surface does not close it. No attack was fired at this kit and none is claimed. The three boundaries below were checked in code on 2026-08-23, and only the first of them is red-proven in both directions rather than argued from an absent code path.
Boundary checked
What could go wrong
What the code guarantees
Does a driver's name, mobile and CDL number ever leave the machine?
Every snapshot carries a Driver Contact section -- the commercial driver's name, mobile number and CDL number. A CDL number is a state-issued licence identifier, and not one field this kit answers asks for any of it. A selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the snapshot MINUS that section rather than to the snapshot. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py over all 120 documents. ⚠︎ AND THE FIRST VERSION OF THAT PROOF MEASURED ZERO AND WAS THE TEST NOT FIRING. Swapping the guard for the naive or list(secs) changes nothing on today's corpus, because every hint names a section all 120 documents carry -- so the fallback is not on any live code path. The hole is CONDITIONAL, so the proof reproduces the condition: a source-system rename (a WMS upgrade renaming half a report) that leaves nothing this kit maps still parsing. With the guard, 0 of 120 leak; with or list(secs), 120 of 120.
Can a wrong answer poison the next scheduled run?
A monitor that fed its own verdict forward would compound one bad reading into every reading after it -- and here the carried quantity is billable hours, so an over-count at 06:00 is still in the invoice at 18:00.
src/detention.step() is the only thing that touches the clock and it reads only the numbers parsed off the page. The model's reply is scored and thrown away. On r001 the 8 wrong cells across 3 fields did not reach a single later reading, because there is no code path by which they could.
Can this kit approve, release or pay a detention charge?
A watch that could open an event could plausibly be extended to approve the charge that follows it -- which is the row's own cap (carrier-payment-release) and the thing a carrier would most object to.
There is no such endpoint, no such function and no configuration flag. evals/check_labels.py greps every .py and .js file in the kit for the names of such paths and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
The first row is the only one measured in both directions. The other two are guarantees about what is ABSENT — no code path approves, releases or pays a charge, and no code path feeds a model reply back into the detention clock — and absence is checked by asserting names it knows, which is weaker than a run and is written here as such.
The result0 attack trials, three boundaries checked -- and the privacy boundary red-proven by removing the guard and watching all 120 snapshots leak a commercial driver's name, mobile and CDL number.
1field an outside party could influence (sent, not hidden)
0 of 0attack trials run
1 of 3boundaries red-proven, not just asserted
0 of 120snapshots leak a driver's name, mobile or CDL
The Operational Notes ARE the field an outside party would influence in a real deployment, and this kit sends them rather than hiding them -- one of the three shipped sentences is even instruction-shaped ("Carrier has asked to be told before any detention event is opened on this visit"), chosen for exactly that reason. But on this corpus they are one of three sentences from a seeded generator, so there is nothing adversarial in them to catch. A version pointed at real clerk prose reopens the question and should be attacked before it ships. What IS measured is the privacy boundary, in both directions, before every run.
Read this twice
The yard log and the operational notes reach the model verbatim, and one of the three shipped notes is instruction-shaped on purpose — “Carrier has asked to be told before any detention event is opened on this visit.” Nothing here filters it and nothing here has measured whether it moves the answer. A model declining to follow an instruction would not be a defence anyway — it is one vendor’s behaviour on one day.
HonestyWhat this does not prove
Whether a real deployment's Operational Notes -- prose an inbound clerk writes freely -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at a real yard. The shipped note asking to be warned before an event is opened is the shape to probe.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two watches on one yard have never been run and would race. On a RUNNING CLOCK a lost run is a permanently wrong number of billable hours rather than a missing row.
Whether a truncated reply can ever be partially trusted. The stateless control lost one call to the ceiling and src/watch._parse rejected it outright, which is the safe behaviour and is not the same as having measured how often a fragment would have been right.
Whether the grep-shaped no-payment assertion would catch an approval path written under a name it does not know. It asserts the absence of seven names; a path called something else would pass it.
What the provider retains. The prompt carries no personal data by construction -- the Driver Contact section never leaves -- but what a vendor keeps of a request is outside this repository and nobody here has verified it.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No carrier-payment release, non-configurable. This kit produces a watchlist status, a clock state, an accrued-hours figure and a raise/hold call for a yard manager to read. It never approves, releases, adjusts, suppresses or pays a detention charge, never negotiates with a carrier and never writes to a TMS, and there is no setting that makes it. Separately, and just as non-configurable: the free-time value is OPERATOR-SUPPLIED. A visit whose agreement carries no detention clause is reported CONTEXT_INCOMPLETE and accrues nothing -- there is no default figure in this kit.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). src/detention.step() is the only thing that touches the clock, and it is called by evals/run.py after every reading including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit approves, releases or pays a charge
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for such names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
No detention accrues against a guessed free-time figure
18 of 18 readings with nothing on file were reported CONTEXT_INCOMPLETE by the model (100.00 pct). The two floors that GUESS a 2.00-hour default got 18 of 18 wrong -- 0.00 pct -- which is the guardrail measured rather than asserted.
Exactly one detention event per visit
0 duplicate raises in 92 readings that must not open one, and 0 missed raises in 28 that must. ⚠︎ Both zeros are results on THIS corpus, not properties of the code: nothing in the kit refuses a second raise, the carried state merely tells the model one is already open. The control arm, with that sentence removed, produced 23 duplicates.
A wrong answer does not propagate into the next scheduled run
The property is guaranteed by the code path -- src/detention.step() never reads the reply -- and r001 is CONSISTENT with it rather than a demonstration of it: every visit carrying a wrong cell was answered correctly at its next run.
The driver's name, mobile and CDL never reach the provider
0 of 120 snapshots leak the Driver Contact section, and 120 of 120 leak it when the guard is removed AND the condition that reaches the fallback is reproduced. Both directions asserted by evals/check_labels.py before any run may spend.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The clock is correct whatever the model says, which means a wrong reading is a wrong watchlist row and not a corrupted history. Those are different problems and only the first one is on this page.
IT IS NOT A SCHEDULER. evals/run.py is invoked; it is not woken. Nothing here detects a missed run, back-fills it or marks its readings late -- which is the half of a monitor this variant exists to name, and this kit has the same hole.
IT IS NOT AN AUTHORITY ON DETENTION. The eight rules, the free-time hours and the rates are invented for this kit and reproduce no carrier agreement, tariff or regulator's guidance.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 34 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
16 measured by the latest run18 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key
alarm
status_accuracy_pct; clock_accuracy_pct; detention_hours_accuracy_pct; raise_event_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. The scored run had 0 of 120 -- but the stateless control at the identical ceiling had 1, so 0 is this arm's result and not a property of the design.
raise-directions
The two raise directions, counted apart
alarm
missed_raise_pct; duplicate_raise_rate_pct — alarm on missed_raise_pct above 0. A duplicate is visible to whoever reads the watchlist; a missed raise is invisible until a carrier invoices for a dwell nobody logged, and by then the contemporaneous evidence is gone.
memory-dependent-subset
The readings whose answer is not on their own page
alarm
memory_status_accuracy_pct; memory_hours_accuracy_pct; memory_raise_accuracy_pct — alarm on memory_raise_accuracy_pct. It is the cell the carried state exists for, and it is the one that moved most when the state was removed -- 58.62 pct against 100.00.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
120
different corpus — nothing is comparable
corpus.bytes
459,108
yard snapshots edited — the count held, the bytes did not
split.count
40
the visits count moved — a different set was scored
split.size_p50
3
the median size of one visit moved
split.size_p95
3
the 95th-percentile size of one visit moved
dataset.rows
120
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cadence_hours 6, context_incomplete_cells 18, documents 120, memory_cells 61, quiet_cells 92, raise_cells 28, readings_scored 120, stateless False, visits 40) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
status accuracy
not yet known
120 readings
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. r001 and s001 are one run each of two DIFFERENT prompts, which is a comparison and not a band.
the raise/hold call, both directions
0 -- saturated at 100.00 pct with 0 missed and 0 duplicate, and that is a statement about the corpus
28 that must open an event, 92 that must not
A grader an arm aces has stopped discriminating. The honest reading is that this corpus's raise decisions are within reach of the configuration WHEN IT HAS THE CARRIED STATE -- remove it and the same grader separates the arms by 22.50 points.
the clock state
99.17 pct -- 119 of 120, and the single miss is a generator edge
120 readings
evals/scoring.py::score, r001-detention-watch. The free floor scores 86.67 pct on the same cells, so this grader DOES discriminate -- 12.50 points.
accrued hours, exact
99.17 pct -- 119 of 120, and the one miss is a wording defect in the carried sentence rather than an arithmetic error
120 readings; 61 of them memory-dependent
evals/scoring.py::score on the quarter-hour billing increment. This is where the model earns its bill: 99.17 pct against the free floor's 85.00.
the memory-dependent subset
98.36 / 100.00 pct over 61 cells
61 readings whose answer is not on their own page
The control, on the same cells, scores 58.62 on the raise column. That gap is the carried state, isolated.
the operator-supplied free-time guardrail
100.00 pct -- 18 of 18
18 readings with no free time on file or no gate-in evidence
Both arms that abstain score 100; both floors that fill a 2.00-hour default score 0. This is the row's own guardrail, measured.
the money
0.66 pct out on a yard total of $22,687.50
1 -- the yard total, which is why the per-reading rates are published beside it
Errors cancel in a total. The free floor is 5.94 pct out and the ageing floor 25.53 pct; a total alone would hide which direction each is wrong in, so the per-reading exact rate is the one to watch.
One run. The control at the identical ceiling answered 97.50 pct and cost 3.69 times as much per reading, so neither the reliability nor the price is a property of the design -- both are properties of the arm.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-detention-watch-ageonly 2026-08-23
b001-detention-watch-ageonlymem 2026-08-23
b002-detention-watch-keywordmem 2026-08-23
answered, %
100.0
100.0
100.0
clock accuracy, %
85.83
85.83
86.67
context incomplete recall, %
0.0
0.0
100.0
detention hours accuracy, %
76.67
80.83
85.00
detention usd error, %
25.53
15.85
5.94
duplicate raise rate, %
39.13
3.26
1.09
input tokens, whole run
0
0
0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
memory hours accuracy, %
59.02
67.21
73.77
memory raise accuracy, %
42.62
96.72
98.36
memory status accuracy, %
95.08
96.72
96.72
missed raise, %
0.00
0.00
3.57
output tokens, whole run
0
0
0
raise event accuracy, %
70.00
97.50
98.33
status accuracy, %
80.83
81.67
95.00
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 3 runs. Columns here are only ever compared with each other.
Metric
c000-detention-watch-calibration 2026-08-23
r001-detention-watch 2026-08-23
s001-detention-watch-stateless 2026-08-23
answered, %
100.0
100.0
97.5
clock accuracy, %
83.33
99.17
95.83
context incomplete recall, %
—
100.0
100.0
detention hours accuracy, %
100.00
99.17
86.67
detention usd error, %
0.00
0.66
1.41
duplicate raise rate, %
0.0
0.0
25.0
input tokens, whole run
10382
205012
196383
model latency p50 ms
32284.00
5376.00
13860.00
model latency p95 ms
42275.00
25535.00
123056.00
memory hours accuracy, %
100.00
98.36
77.59
memory raise accuracy, %
100.00
100.00
58.62
memory status accuracy, %
100.00
98.36
91.38
missed raise, %
0.00
0.00
3.57
output tokens, whole run
19523
115144
503820
raise event accuracy, %
100.0
100.0
77.5
status accuracy, %
83.33
95.00
89.17
not a time series No two of these 3 runs measured the same system — they differ on context_incomplete_cells, documents, max_tokens, memory_cells, quiet_cells, raise_cells, readings_scored, stateless, visits — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-detention-watch-stub 2026-08-23
answered, %
100.0
clock accuracy, %
85.83
context incomplete recall, %
0.0
detention hours accuracy, %
76.67
detention usd error, %
25.53
duplicate raise rate, %
39.13
input tokens, whole run
209014
model latency p50 ms
0.00
model latency p95 ms
0.00
memory hours accuracy, %
59.02
memory raise accuracy, %
42.62
memory status accuracy, %
95.08
missed raise, %
0.0
output tokens, whole run
4678
raise event accuracy, %
70.0
status accuracy, %
80.83
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 16 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
r001-detention-watch against s001-detention-watch-stateless. Same corpus, same model, same grader, same 120 readings; the prompts differ in exactly one block and evals/check_labels.py asserts they are byte-identical everywhere else. ⚠︎ Three of the control's calls produced nothing and are counted as wrong rather than dropped, so every figure on its side of this row carries up to 2.5 points of that.
whether a model is called at all
status 95.00 pct (free code) -> 95.00 pct (NO MOVEMENT -- 114 of 120 each), clock 86.67 -> 99.17, accrued hours 85.00 -> 99.17, yard total error 5.94 pct -> 0.66 pct
measured
r001-detention-watch against b002-detention-watch-keywordmem over the same 120 readings and the same scorer. THE HEADLINE BAND DOES NOT MOVE, and the kit leads with that. ⚠︎ The comparison is also asymmetric in the floor's favour: evals/baseline.py and the answer key are two expressions of one implementation, so the floor cannot misread the rule and the model can only read the prose.
whether memory is available to FREE CODE, with no model involved
b000-detention-watch-ageonly against b001-detention-watch-ageonlymem. Same rule, same corpus, no model in either. This is the single largest movement anywhere on this page and there is no model in it: what a monitor buys is MEMORY, and the model is what you add on top.
b001-detention-watch-ageonlymem (fills 2.00 hours) against b002-detention-watch-keywordmem (abstains) over the same 120 readings. The 18 readings with nothing on file go from all wrong to all right, and they are the reason the row's guardrail exists.
b001 against b002 -- both free, both with the carried state, and the second one parses the log into intervals and attributes them. Reading the log is worth about 4 points of hours to free code and about 14 to the model, which is the shape of the whole comparison.
the cadence
unmeasured -- and it moves the WATCH band's definition as well as the bill, because Rule D-8's threshold IS the interval between runs
reasoning
src/detention.WATCH_MARGIN_HOURS is derived from CADENCE_HOURS. No run at any other cadence exists, so this row is the one lever on this page with no measurement behind it, and it is marked as such rather than estimated.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
status accuracy
nothing yet. ⚠︎ And note what this figure is level with: the free floor scores 95.00 pct on the same 120 readings.
the raise/hold call, both directions
any missed raise at all. The cost of that direction is the whole detention charge, because the evidence stops being contemporaneous
the clock state
any clock error on a reading whose yard log is quiet. This run's one miss is on a stop opening at exactly the snapshot instant; a miss with no stop anywhere near the boundary would be a different failure
accrued hours, exact
any error larger than one quarter hour, or any error on a reading whose window has no yard-log entries at all
the memory-dependent subset
the memory columns falling below the non-memory ones. That would mean the carried sentence is confusing the model rather than informing it, which is exactly the shape of this run's one hours miss
the operator-supplied free-time guardrail
any detention at all on a visit with nothing on file. One such reading is a charge against a number nobody supplied
the money
the total drifting while the per-reading exact rate holds. That is errors stopping cancelling, and it means the population has changed
reliability and the bill
answered_pct below 100, or output tokens rising without the corpus changing. The second one is how removing the carried state announced itself
NextThe three you would add first
A scheduler, and something that notices when a run did not happenThis is the half of a monitor the kit does not ship. evals/run.py is invoked, not woken; nothing detects a missed run, back-fills it or marks its readings late. Every figure on these pages is measured over three runs that all happened, and a detention claim is only as good as the timestamp on its evidence.
A human step in front of anything that becomes a carrier chargeraise_event YES is the highest-consequence output this kit has and it is the first link in a chain that ends with money moving to a carrier. The arms without memory open 23 and 36 duplicate events respectively, and the stateful arm's zero is measured on one run of an invented rule.
The real free-time values, from the real agreementsEvery free-time figure and every rate in this corpus is invented. The row this kit was built from records free time as an operator-supplied clock precisely because it varies per contract, per move type and per lane -- and the kit's own CONTEXT_INCOMPLETE rate is the measurement of what happens when it is missing, not a licence to guess.
A refusal state for a reply that did not parseThe scored run had none, but the stateless control lost one call to the ceiling and it scored as four wrong cells. In a deployment that reading needs to be visibly UNANSWERED rather than silently wrong, and the kit has no vocabulary for that.
A state store with a concurrency modeldata/state.json is one file replaced atomically. That is right for one process and is not a design for a scheduler, and the carried quantity is a running total of billable hours, so a lost write is permanent rather than transient.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free, under a second) on any change to tools/build_corpus.py, src/detention.py, src/segment.py or src/select.py -- it refuses to let a run spend if the corpus, the answer key, the experimental control, the privacy guard or the no-payment property has moved. Re-run the scored eval AND its stateless control together (paid, 240 calls) on any change to src/prompt.py or src/state.py: the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. All three free floors cost nothing and should be re-run on every change, in both directions -- a floor that gets stronger is a finding, not a regression.
What this cannot tell you
One run per arm. Whether the status TIE against the free floor is stable across repeats is not measured, and 114 of 120 each means one cell either way flips which arm reads higher.
Whether the no-propagation property would hold under a wrong answer mid-chain. It is true by construction -- src/detention.step() never reads the reply -- and this run is only consistent with it.
Whether the no-payment guarantee holds against a path named something the assertion does not know. It greps for seven names; the guarantee is about ABSENCE and that is the only mechanical form it can take.
Whether rewriting Rule D-8 to say whether the clock must be running would remove five of this run's six status errors. It has not been tried; it costs one more 120-call run, and the shared provider account had no balance left to try it with.
Whether the zero duplicate raises would survive a corpus where a visit's event is opened and then CLOSED, or disputed, or waived. This corpus has no closed events.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a CLOCK -- four scalar fields written by arithmetic. The cost is flat in history length because nothing here remembers what was said, only what was counted, and a memory layer would give back the growth this design exists to avoid. It would also give back the failure this design exists to avoid: a running total fed from its own model output compounds forever, and here the total is money
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the schedule
evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
THIS is the seam where a framework genuinely earns its place, and the kit says so rather than pretending otherwise: 40 independent chains of 3 strictly-ordered readings, with the cadence, retries, missed-run detection and late back-fill all OUTSIDE the kit. A ThreadPoolExecutor is the right size for an eval and the wrong size for a watch that has to wake four times a day and notice when it did not
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
four exact-match comparisons, three slices of the same cells and one dollar sum is a dict comprehension, not a platform
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each visit is a chain of three readings -- 06:00, 12:00, 18:00 -- with no branching and exactly one edge between consecutive runs, carrying four scalars. Different visits never touch. A framework would add an orchestrator to a for-loop that already runs 40 chains wide.
The other sideWhat a framework costs you
No scheduler. This is a watch that carries state and is still INVOKED, not woken -- and the cadence is the single most important thing about it. That is the honest cost of staying a folder of readable Python, and it is stated on the kit's own page rather than implied by silence.
No missed-run detection, no back-fill, no late marking. All three are what a scheduling engine gives you and none of them exists here.
No concurrency model for the state store. data/state.json is one file replaced atomically -- correct for one writer, and a yard with two watches running is two writers.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
The scheduling seam in particular is argued and not tested. Nothing here has ever run under cron, Airflow or Temporal, so the claim that a framework earns its place there is a design judgement rather than a measurement.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-detention-watch on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
answered_pct below 100, or output tokens rising without the corpus changing. The second one is how removing the carried state announced itself
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-detention-watch-calibration32,284 ms
r001-detention-watch5,376 ms
s001-detention-watch-stateless13,860 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-detention-watch-ageonly, b001-detention-watch-ageonlymem, b002-detention-watch-keywordmem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
yard snapshots
data/corpus/VIS-<n>-R<k>.txt -- 120 files, 459,108 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Driver Contact -- the driver's name, mobile and CDL number -- never does, by src/select.NEVER_SENT
the carried state
data/state.json in a deployment, and in-memory per run in the eval -- four scalars per visit, written by src/detention.step and never by a model
as one English sentence in every prompt, and the UI prints the same sentence verbatim so a reader can audit what the model was told
the answer key
data/gold.jsonl -- 120 rows, the output of src/detention.step over the planted intervals, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the run records
results/eval-*.json in the kit, and one small record per run in the app repo's run register
never -- they are read by the site build, not by a provider
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser -- which is not a theoretical claim here: the provider-refused screenshot on this page IS that path, captured live.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a SIX-HOUR watch: 06:00, 12:00, 18:00 and 00:00 local, 4 runs a day. One run owns exactly the window since the last one -- it is the run that must open any detention event whose free time expired inside it, and the snapshot's Hours counted in this window line is that window measured. The cadence is not a deployment detail here: the WATCH band is DEFINED as the cadence (Rule D-8, "less free time remains than the interval between runs"), so changing the schedule changes the answer as well as the bill.
40 visits x 3 scheduled runs = 120 calls per arm, 0 gaps and 0 out-of-order readings; evals/check_labels.py asserts all 40 sequences are complete before a run may spend. Wall clock 111.3s at 12 workers. A yard of 400 open visits on this cadence is 1,600 calls a day, $5.97 on the shared projection card. (r001-detention-watch, evals/check_labels.py, src/detention.CADENCE_HOURS)
⚠︎ WHAT A MISSED RUN COSTS IS ARGUED HERE FROM THE ARITHMETIC AND IS NOT MEASURED, and the kit says so rather than implying otherwise. Skip a run and the next one's window is twelve hours instead of six, so the counted hours are right and the DETECTION is late: a detention event whose free time expired at 07:00 is opened at 18:00 rather than 12:00, and on this corpus 28 of the 28 readings that must open an event are the FIRST run that saw the breach. What a late open costs is not hours -- the arithmetic recovers those -- it is the evidence: a detention claim survives a carrier dispute on CONTEMPORANEOUS gate and appointment records, and a reconstruction is what gets thrown out.
every figure on this page is per a six-hour watch. An hourly watch pays 6x and re-defines the WATCH band; a shift-length watch pays less and widens the worst-case detection lag to eight hours. Nothing here measures either, and the band moving with the schedule means the two are not comparable runs of one question.
state
four scalar fields per visit -- hours already on the clock (carrier stops already deducted), whether a carrier-attributable stop is open, whether a detention event has already been opened, and the status last reported -- written by src/detention.step from the figures PARSED off the page and never from the model's reply, and rendered by src/state.describe into one English sentence. That sentence is the entire route from one scheduled run to the next, and it is what both scored arms differ by.
272 characters on the worked example. Input tokens 205,012 with it against 196,383 without -- and OUTPUT 115,144 against 503,820, so the carried state CUTS the bill: $0.00373282 a reading with it against $0.01375771 without. Raise/hold 100.00 pct against 77.50; duplicate detention events 0 against 23. (r001-detention-watch against s001-detention-watch-stateless, lenses.LLM.prompt_parts)
four scalars, so run 400 costs what run 2 costs -- the OPPOSITE curve to an intake kit, whose input grows with the square of the turns, and that matters more here because this watch wakes four times a day. What is NOT bounded is accuracy over a longer history, which is unmeasured, and the store itself: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model.
lose the carried state and the raise/hold call collapses -- 77.50 pct against 100.00, 23 duplicate detention events against 0 -- with every reply still well-formed and nothing raising an error. See environment.signatures' traceless row.
model
one completion call per reading, on one provider and one key, through src/adapters/__init__.py. MAX_TOKENS = 24000, set from calibration; thinking is never sent, so every published run left provider-side reasoning at the default and the result files record thinking: null.
120 calls on the scored arm, 115,144 output tokens, 103,557 of them provider-side reasoning (89.9 pct). Largest reply 7,760 tokens against the 24,000 ceiling. p50 5.4s, p95 25.5s. 0 unparsed. (r001-detention-watch, c000-detention-watch-calibration)
the ceiling is the thing that breaks first and it broke on the CONTROL, not on the shipped arm: one stateless call ran to 24000 of 24000 output tokens and returned nothing. On the shipped arm the largest reply was 7,760, so there is roughly 3x headroom -- measured, not assumed.
a provider whose reasoning default differs reprices this kit without changing anything a reader can see, and nothing here has measured what disabling reasoning does to the answers. Only ONE model was run: the shared account hit Insufficient Balance immediately after the control, so no cross-tier claim is made anywhere on this page.
labels
120 readings -- 40 trailer visits re-read at 3 scheduled runs -- with the four answers computed by tools/build_corpus.py from the visit model and re-derived from src/detention.step by evals/check_labels.py before any run may spend. 61 of them are memory-dependent; 18 have no free time on file and must accrue nothing.
it stops scoring at the third scheduled run of each visit. Long enough to hide a carrier stop behind a run boundary and not long enough to test a clock that has been accruing for days -- and the carried quantity is a running total, so a longer chain is where an off-by-a-quarter-hour would accumulate. Nothing here measures that.
the labels encode ONE reading of Rule D-8, and the model disagrees with that reading on five cells (twice, arguably, correctly). A corpus whose rule text is unambiguous would move the status headline; this one does not have that property and the page publishes the lower figure.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a visit reported SUSPENDED whose yard log shows a hold the facility caused
something attributed a facility stop to the carrier and stopped the clock for it. Detention exists precisely to charge for facility delay, so this is the direction that loses money silently
read the yard-log line, not the number. On VIS-0006-R2 the free keyword floor reads "dispatch here put two appointments on the same door" as the carrier, subtracts 2.25 chargeable hours, lands under the 4.00-hour free time and reports SUSPENDED with nothing to open. The key says DETENTION_OPEN at 1.00 hour (results/eval-b002-detention-watch-keywordmem.json against results/eval-r001-detention-watch.json, VIS-0006-R2)
a second detention event opened on a visit that already has one
the reading was made without the carried state -- nothing on the snapshot says an event exists, so a reader who cannot see the previous run opens another one
check what the carried state said before disputing the event. The stateless control opened 23 duplicate events in 92 quiet readings; the stateful arm opened 0 (results/eval-s001-detention-watch-stateless.json against results/eval-r001-detention-watch.json)
detention accruing on a visit whose carrier agreement has no detention clause
something filled the missing free time with a default. That is the one thing this row's guardrail forbids: the figure is operator-supplied and a guessed one produces a charge nobody can defend
look at the Free time on file line. Both floors that guess a 2.00-hour default get all 18 of these readings wrong; both arms that abstain get all 18 right (results/eval-b000-detention-watch-ageonly.json against results/eval-r001-detention-watch.json)
nothing at all -- the watchlist simply does not change between two shifts
TRACELESS. A scheduled run that never fired leaves no artefact anywhere in this kit: no error, no gap marker, no late flag. The readings it would have produced simply do not exist, and the next run's window silently covers twice the time
there is nothing on the watchlist to look at, so look at the SCHEDULER instead: compare the number of readings this kit produced against the number the cadence says it should have (open visits x runs_per_day). On this corpus that is 40 x 3 = 120 and evals/check_labels.py asserts it; in a deployment nothing does, and that gap is the finding rather than a workaround. (argued from evals/run.py's structure, not from a run. Nothing here has ever skipped a scheduled run.)
['Concurrency. data/state.json is replaced atomically, which is correct for one writer; two watches on one yard is two writers and nothing here has tested it.', 'What a missed scheduled run costs. Argued from the arithmetic in the cadence row; never reproduced.', 'A cadence other than six hours, and therefore any WATCH-band figure at another interval.', 'A population that changes between runs -- a first sighting mid-window, or a visit that leaves the watch.', 'Provider-side retention. The prompt carries no personal data by construction, but what the provider keeps of a request is outside this repository and nobody here has verified it.', 'GPU sizing, local inference and anything about running this off a hosted API. Not attempted; not costed.', 'A second model. One tier was run and the shared account ran out of balance immediately after the control arm.', 'Whether the Operational Notes field can move raise_event. The surface is sent deliberately and was never attacked.']
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every trailer, carrier code, agreement reference, driver, free-time figure, rate and yard-log line is invented here. Verified against the repository's own LICENSE file on 2026-08-23. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineThe status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key
For each of the 120 readings and each of the four answered fields, did the reply equal the computed answer key? Status, clock and the raise/hold call are compared exactly; the accrued hours are compared on the quarter-hour billing increment, and a reply that is not a multiple of a quarter hour is a MISS rather than a zero -- 0.00 is a meaningful answer here (it is what every CLEAR and every CONTEXT_INCOMPLETE reading gets) and a parse failure must not be scored as one.
$0.00per 1,000 yard snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function all three free floors are scored through.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
VIS-0006-R2 -- scheduled run 2 of 3, 2026-03-04 12:00, live unload, 4.00 hours free time on file at 85.00 USD/h
What the previous scheduled run left behind
As at the previous scheduled run, 0.00 hours were already on this visit's detention clock, carrier-attributable stops already deducted. No carrier-attributable stop was open when that run ended. No detention event has been opened for this visit yet. It was reported CLEAR.
The yard log, this window only
08:45 dispatch here put two appointments on the same door, this one is held / 11:00 door cleared, this trailer worked / 11:15 driver left the site for a federally required 30-minute break
SUSPENDED, clock STOPPED, 0.00 detention hours, raise_event NO
Why this one
It is the only shape in this corpus where a wrong attribution does not merely change a number -- it loses the event entirely, which is the costly direction.
Grader
Verdict
Why
The status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key
status hit, clock hit, hours hit, raise/hold hit -- four of four
VIS-0006-R2, run 2 of 3 on a live unload with 4.00 hours of free time. The window carries two stops: "dispatch here put two appointments on the same door, this one is held" (08:45-11:00) is the FACILITY's own dispatch and is chargeable, and "driver left the site for a federally required 30-minute break" (11:15, still open at the 12:00 snapshot) is the CARRIER's and is not. 5.75 counted hours minus 0.75 carrier-attributable leaves 5.00 on the clock against 4.00 free: DETENTION_OPEN, clock STOPPED, 1.00 hours, and raise YES because the carried state says no event exists yet. The model answered all four correctly and its rationale names the two rules it applied.
The two raise directions, counted apart
in scope on the RAISE side, and a hit -- one of the 28
This reading's correct answer is to OPEN the event, so it sits inside the 28 that this grader's missed-raise rate is measured over and outside the 92 its duplicate rate is measured over. The model raised it. The free floor did NOT: having subtracted the facility's 2.25 hours as if they were the carrier's, it lands at 2.75 on the clock, under the 4.00 free time, and reports SUSPENDED with nothing to open. That is the costly direction -- a detention event nobody logged is a charge nobody can bill.
The readings whose answer is not on their own page
not in scope
This grader scores only the 61 readings whose answer cannot be reached from their own page. VIS-0006-R2's carried state is all zeros -- 0.00 hours on the clock, no open stop, no event -- so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The visit's NEXT reading, VIS-0006-R3, IS in scope: by then 5.00 hours are carried and a carrier stop is open.
The formulaWhat it computes
accuracy = hits / 120 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
95.0% status accuracy · 3 more measured on this row
the fast tier, memory removed (THE CONTROL)
89.2% status accuracy · 3 more measured on this row
the strongest free floor, no model
95.0% status accuracy · 3 more measured on this row
the ageing rule a yard already runs, no model, no memory
80.8% status accuracy · 3 more measured on this row
the same ageing rule, GIVEN the carried state
81.7% status accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by tools/build_corpus.py over the planted intervals at generation time and re-derived from src/detention.step by evals/check_labels.py before any run may spend. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known to be arguable -- Rule D-8 does not say whether the clock must be RUNNING for a WATCH band to mean anything, and the key assumes it must. Five readings turn on that sentence and on two of them the model's reading is better than the key's. The other three fields have no known ambiguity.
Watch these
status_accuracy_pct
clock_accuracy_pct
detention_hours_accuracy_pct
raise_event_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. The scored run had 0 of 120 -- but the stateless control at the identical ceiling had 1, so 0 is this arm's result and not a property of the design.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because two of them are small: 28 readings that must open an event and 18 with nothing on file, so one row moves the first by 3.6 points and the second by 5.6.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/detention.py, src/segment.py or src/select.py. Re-run the scored eval AND its stateless control together (paid, 240 calls) on any change to src/prompt.py or src/state.py -- the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. All three floors are free and should be re-run on any change at all.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real yard, where whether a stop was the carrier's is decided by a person reading a clerk's note. That is why this corpus is generated rather than captured.
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineThe two raise directions, counted apart
Two counts over two different denominators. A MISSED raise: of the 28 readings whose correct answer is to open the detention event, how many did not. A DUPLICATE raise: of the 92 readings that must not open one, how many did.
$0.00per 1,000 yard snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in the same pass as the exact-match grader.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
VIS-0006-R2 -- scheduled run 2 of 3, 2026-03-04 12:00, live unload, 4.00 hours free time on file at 85.00 USD/h
What the previous scheduled run left behind
As at the previous scheduled run, 0.00 hours were already on this visit's detention clock, carrier-attributable stops already deducted. No carrier-attributable stop was open when that run ended. No detention event has been opened for this visit yet. It was reported CLEAR.
The yard log, this window only
08:45 dispatch here put two appointments on the same door, this one is held / 11:00 door cleared, this trailer worked / 11:15 driver left the site for a federally required 30-minute break
SUSPENDED, clock STOPPED, 0.00 detention hours, raise_event NO
Why this one
It is the only shape in this corpus where a wrong attribution does not merely change a number -- it loses the event entirely, which is the costly direction.
Grader
Verdict
Why
The status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key
status hit, clock hit, hours hit, raise/hold hit -- four of four
VIS-0006-R2, run 2 of 3 on a live unload with 4.00 hours of free time. The window carries two stops: "dispatch here put two appointments on the same door, this one is held" (08:45-11:00) is the FACILITY's own dispatch and is chargeable, and "driver left the site for a federally required 30-minute break" (11:15, still open at the 12:00 snapshot) is the CARRIER's and is not. 5.75 counted hours minus 0.75 carrier-attributable leaves 5.00 on the clock against 4.00 free: DETENTION_OPEN, clock STOPPED, 1.00 hours, and raise YES because the carried state says no event exists yet. The model answered all four correctly and its rationale names the two rules it applied.
The two raise directions, counted apart
in scope on the RAISE side, and a hit -- one of the 28
This reading's correct answer is to OPEN the event, so it sits inside the 28 that this grader's missed-raise rate is measured over and outside the 92 its duplicate rate is measured over. The model raised it. The free floor did NOT: having subtracted the facility's 2.25 hours as if they were the carrier's, it lands at 2.75 on the clock, under the 4.00 free time, and reports SUSPENDED with nothing to open. That is the costly direction -- a detention event nobody logged is a charge nobody can bill.
The readings whose answer is not on their own page
not in scope
This grader scores only the 61 readings whose answer cannot be reached from their own page. VIS-0006-R2's carried state is all zeros -- 0.00 hours on the clock, no open stop, no event -- so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The visit's NEXT reading, VIS-0006-R3, IS in scope: by then 5.00 hours are carried and a carrier stop is open.
The formulaWhat it computes
missed_raise_pct = missed / 28; duplicate_raise_rate_pct = duplicates / 92. Never averaged and never combined into an F-score.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
0.0% missed raise · 1 more measured on this row
the fast tier, memory removed (THE CONTROL)
3.6% missed raise · 1 more measured on this row
the ageing rule, no memory
0.0% missed raise · 1 more measured on this row
the same rule, GIVEN the carried state
0.0% missed raise · 1 more measured on this row
the strongest free floor
3.6% missed raise · 1 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl's raise_event column, computed by src/detention.step. This grader is scored against the reference and is not itself the reference.
These rates are UNKNOWN, on purpose
Its own TPR and TNR are not measured because it does not answer a true/false question about another grader's verdict -- it counts two directions of one field against the key, so there is nothing here for an agreement rate to be about.
Watch these
missed_raise_pct
duplicate_raise_rate_pct
Alarm on
missed_raise_pct above 0. A duplicate is visible to whoever reads the watchlist; a missed raise is invisible until a carrier invoices for a dwell nobody logged, and by then the contemporaneous evidence is gone.
How tight can the band be? 28 readings must open an event and 92 must not, so the finest band either rate can support is 3.6 and 1.1 points respectively. A 'zero missed raises' claim on this corpus is a claim about 28 readings and nothing wider.
Cadence: Free. Re-run with every scored run, and re-run the floors whenever src/state.py changes -- the duplicate count is the most memory-sensitive figure on this page.
The decisionWhen to reach for it
Use it
The two failure directions cost different things and are fixed by different people, which is exactly this kit's situation.
Do not use it
A task where one error direction is free. Here neither is: a missed raise is an unbillable charge, a duplicate is two charges for one dwell.
The readings whose answer is not on their own page
Catch a yard trailer's free time before it runs out
PresenterOpens the private repo. Visible to admins only.
In one lineThe readings whose answer is not on their own page
The same exact-match comparison, restricted to the 61 readings the generator marked memory_dependent: hours already on the clock that this snapshot does not restate, a carrier stop that opened before this window, or an event already opened. It answers one question -- is the carried state doing anything, or is the corpus easy?
$0.00per 1,000 yard snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, same pass, grouped by the flag tools/build_corpus.py set at generation time.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
VIS-0006-R2 -- scheduled run 2 of 3, 2026-03-04 12:00, live unload, 4.00 hours free time on file at 85.00 USD/h
What the previous scheduled run left behind
As at the previous scheduled run, 0.00 hours were already on this visit's detention clock, carrier-attributable stops already deducted. No carrier-attributable stop was open when that run ended. No detention event has been opened for this visit yet. It was reported CLEAR.
The yard log, this window only
08:45 dispatch here put two appointments on the same door, this one is held / 11:00 door cleared, this trailer worked / 11:15 driver left the site for a federally required 30-minute break
SUSPENDED, clock STOPPED, 0.00 detention hours, raise_event NO
Why this one
It is the only shape in this corpus where a wrong attribution does not merely change a number -- it loses the event entirely, which is the costly direction.
Grader
Verdict
Why
The status, the clock, the accrued hours and the raise/hold call, per reading, exact match against the computed answer key
status hit, clock hit, hours hit, raise/hold hit -- four of four
VIS-0006-R2, run 2 of 3 on a live unload with 4.00 hours of free time. The window carries two stops: "dispatch here put two appointments on the same door, this one is held" (08:45-11:00) is the FACILITY's own dispatch and is chargeable, and "driver left the site for a federally required 30-minute break" (11:15, still open at the 12:00 snapshot) is the CARRIER's and is not. 5.75 counted hours minus 0.75 carrier-attributable leaves 5.00 on the clock against 4.00 free: DETENTION_OPEN, clock STOPPED, 1.00 hours, and raise YES because the carried state says no event exists yet. The model answered all four correctly and its rationale names the two rules it applied.
The two raise directions, counted apart
in scope on the RAISE side, and a hit -- one of the 28
This reading's correct answer is to OPEN the event, so it sits inside the 28 that this grader's missed-raise rate is measured over and outside the 92 its duplicate rate is measured over. The model raised it. The free floor did NOT: having subtracted the facility's 2.25 hours as if they were the carrier's, it lands at 2.75 on the clock, under the 4.00 free time, and reports SUSPENDED with nothing to open. That is the costly direction -- a detention event nobody logged is a charge nobody can bill.
The readings whose answer is not on their own page
not in scope
This grader scores only the 61 readings whose answer cannot be reached from their own page. VIS-0006-R2's carried state is all zeros -- 0.00 hours on the clock, no open stop, no event -- so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The visit's NEXT reading, VIS-0006-R3, IS in scope: by then 5.00 hours are carried and a carrier stop is open.
The formulaWhat it computes
accuracy = hits / 61 per field, over the memory_dependent subset only.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
98.4% memory status accuracy · 2 more measured on this row
the fast tier, memory removed (THE CONTROL)
91.4% memory status accuracy · 2 more measured on this row
the strongest free floor
96.7% memory status accuracy · 2 more measured on this row
the ageing rule, no memory
95.1% memory status accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, sliced by the memory_dependent flag. Scored against the reference; not itself the reference.
These rates are UNKNOWN, on purpose
Its own rates are blank for a different reason from the grader above: it is a SLICE of the reference's verdicts rather than a second opinion on them, so an agreement rate between the two would be 1.0 by construction and would mean nothing.
Watch these
memory_status_accuracy_pct
memory_hours_accuracy_pct
memory_raise_accuracy_pct
Alarm on
memory_raise_accuracy_pct. It is the cell the carried state exists for, and it is the one that moved most when the state was removed -- 58.62 pct against 100.00.
How tight can the band be? 61 cells, so the finest band this subset supports is 1.6 points. Two of the three arms on it are within one cell of each other on status, which is why the raise column is the one to read.
Cadence: Free, every run. Re-derive the flag whenever tools/build_corpus.py changes -- it is a property of the corpus, and a corpus with no memory-dependent readings would make this whole kit unfalsifiable (evals/check_labels.py asserts there are at least 40).
The decisionWhen to reach for it
Use it
A kit carries state between runs and has to show the state is load-bearing rather than decorative.
Do not use it
A stateless task. There is nothing for the subset to mean.
A living map of modern AI — kept current every morning