Catch a government contract's deliverable before it's late
Every Monday you re-read the same contract deliverable list, hunting for what actually moved since last week. This app re-reads it for you and tells you which date governs, how many days are at risk, and what's holding each one up.
PresenterOpens the private repo. Visible to admins only.
For the programme analystAerospace & Defense · Government & Public Sector
Why it matters
Today's manual process, and the same job with the app
A programme analyst or contracts administrator at a defense or government contractor, running the Monday deliverable check.
✕Today's manual process
1Read every line of the deliverable register, thirty open items with dates and status notes.
2Check what moved since last Monday, checked manually against an old copy of the register.
3Guess whether a revised date counts when the notes don't say a contract modification backs it up.
4Miss one and a deliverable that's actually overdue quietly drops off this week's list.
Checked by memory, one line at a time
✓With the app
1The register is read whole every Monday, all thirty deliverables, automatically.
2What changed is named which date now governs, and how many days it's been at risk.
3A revised date is checked against whether the notes actually back it with a real modification.
4Nothing slips quietly every deliverable past its gate lands on the watchlist, with its reason.
Every deliverable checked, every Monday, automatically.
See it work
One real case, read by the app, step by step
CDRL-0018-W3, a data item on a defense contract, seven days past its contract due date this Monday.
Catch a government contract's deliverable before it's lateReference appBuilt to be shaped to your process
4
1Which date it's judged against Run W3, 2026-07-20; last week already reported RAISE.
2Days already at risk Seven days past the gate, and still counting.
3What's still missing Government data hasn't arrived, so it stays open.
4The call: escalate Raised last week, unchanged since, so it goes to the analyst.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch a government contract's deliverable before it's late
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A programme office owes a government customer a list of data items, each with a date. The list itself is not hard to read: a spreadsheet with conditional formatting will tell you what is due in the next fortnight. What it will not tell you is what has CHANGED since last Monday — whether the date under a line moved, whether this is the first time a line has been raised or the fourth with nothing having shifted, or whether the 'revised due date' somebody typed in is carried by a contract modification that actually exists. Those are the questions the pass is really for, and none of them is answerable from the row in front of you. The Monday morning pass down a contract deliverable register — thirty open CDRL lines, each with a due date, a submission status and a paragraph of somebody's notes — deciding which ones to put in front of the programme analyst this week and which to leave standing.
Audience
A programme analyst or contracts administrator who runs this pass weekly, and whoever has to justify the raise list to a control account manager. The answer this page gives them is a qualified NO: on this corpus a free rules engine does the whole job, and the only measured reason to pay for a model is robustness to how the notes are worded. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual deliverable records
The corpus is 120 deliverable records, 0.52 MB (json 2 · jsonl 1 · txt 120). A deliverable register is the cleanest case of a population that is re-read whole: every open line is present on every run, nothing arrives as an event, and the finding is never in a row — it is in the difference between this week's row and last week's. It also has a genuine off-page fact. A 'revised due date' prints identically whether or not a contract modification carrying it has been executed, because the field is filled in by whoever asked for the relief and the modification is a different document in a different system. That single boolean decides which date the deliverable is judged against, and it lives only in prose. Without it this corpus would be pure arithmetic and there would be nothing for a model to be measured on at all.
The corpus
The 120 deliverable recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your deliverable records. That is the whole change — there is no database to migrate.
One deliverable record, as the model receives itCDRL-0001-W1.txt · 1 of 120
Synthetic Record
----------------------------------------------------------------
This file is invented for a public evaluation kit. It is not a real contract deliverable record.
Every contract number, programme name, CDRL item, data item title, person, office and date in it
is fictional, and it reproduces no DD Form 1423, no Data Item Description, no contract and no
programme's deliverable register. The watch rule reproduced below was invented here and is not an
authority.
Contract
----------------------------------------------------------------
Contract number : FA8620-24-C-0117
Programme : SENTINEL RIDGE
Scope : Airframe structures upgrade
Contractor : Merrow Aerospace Systems
Administered by : DCMA Springfield
Deliverable
----------------------------------------------------------------
CDRL item : A001
Data item title : Contract Work Breakdown Structure
Recurrence : One-time
Submission medium : Government contractor data portal
Register run : W1 (2026-07-06)
Schedule
----------------------------------------------------------------
Run date : 2026-07-06
Contract due date : 2026-10-05
Revised due date : (none)
Review period : 30 days
Submission Status
----------------------------------------------------------------
Status : IN PROGRESS
Last submitted : (none)
Customer disposition: (none)
Resubmission due : (none)
Watch Rule
----------------------------------------------------------------
DELIVERABLE WATCH RULE (illustrative -- reproduces no DFARS clause, no Data Item Description
and no programme CDRL management plan)
The watch runs every Monday at 06:00 local. On each run every deliverable that has not been
Abridged — the file continues.
The outcomeWhat a good result looks like
A watchlist: for each open deliverable, a band (NONE, WATCH, RAISE, ESCALATE), the date it is being judged against, the days it has spent past its raise gate, and the blocker. A good run raises everything that had to be raised, disturbs nothing that did not, and says which of those answers depended on remembering last week.
And when it cannot
The expensive direction is a MISSED raise: a deliverable that is past its contract date drops off the list because the register printed a revised date nobody executed a modification for. The free age-and-threshold floor does exactly that on 12 of 60 rows that had to be raised (20.00 pct), and it is not a straw man — it is the spreadsheet a programme office already has. The cheap direction is a false alarm, and on this corpus none of the three floors produces one at all (0 of 26 quiet rows).
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
You have a deliverable register and a written lead-time rule, and the notes are written to a house template. — The free rules engine (evals/baseline.py, rules). No provider, no key, no bill. It scored 100.00 pct on every field of this corpus. A model cannot beat 100 and costs money to tie it.
The notes are free prose from many people, and house style changes. — The model arm, with the free floor still computed beside it on every row. The free pattern's whole score on the rows that matter is a fact about wording: reword them and it falls from 100.00 pct to 50.00 pct on the band and to 0.00 pct on the blocker.
You cannot guarantee the watch runs every week. — Monitor the watch itself before you monitor anything with it. One skipped Monday put 16 raises a week late, lost 35 days of accrual and took standing-escalation accuracy from 100.00 pct to 63.16 pct — with the reader unchanged.
And where nothing here is good enough:
You want the raise list to be earlier rather than more accurate. — Neither. Change the CADENCE. A 7-day interval against a 5-day raise gate leaves a 2-day band in which a due date arrives with no run ever having seen it inside the gate: 5 of 30 deliverables here are never warned in time on any schedule, by any reader.
Your register is keyed on something that gets renumbered. — Nothing here yet. Build the identity check first. Every carried scalar is keyed on the deliverable id. A re-keyed line arrives with no history and every previous-run rule silently stops applying to it.
At a glanceHow the whole thing runs
70%band accuracy pct
8,318 msp50, end to end
$0.00per 1,000 deliverable records · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a government contract's deliverable before it's late14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
One record is ONE CDRL LINE AS IT STOOD ON ONE SCHEDULED RUN — not one deliverable and not one document. This kit reads a register and writes a watchlist.Corpus lens →
When is this the wrong choice?
Avoid: Do not add a model here to look modern. You will pay per row per week for an answer a regular expression already had, and you will have added a dependency that can be unavailable on a Monday morning. That is the case against the best-fitting scenario (“You have a deliverable register and a written lead-time rule, and the notes are written to a house template.”). 5 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A register whose line identifiers are renumbered between runs. Every carried scalar is keyed on the deliverable id; a re-keyed line arrives as a deliverable with no history and every previous-run rule silently stops applying to it. 7 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether the free pattern's collapse under rewording (100.00 pct to 50.00 pct on the band) is representative. The paraphrases and the pattern were written by the same author in the same repository, so the probe is an UPPER BOUND on brittleness rather than an estimate. 6 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
6 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — b003-cdrl-watch-rules. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone, run python3 -m tools.build_corpus, and every free arm reproduces from a fixed seed with no key and no network. The recorded results ship in results/, so the kit is browsable offline; python3 -m src.app renders every panel except the model column.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
8,318 msp50, end to end
18,966 msp95
2 minclone to first result
What the clock covers. End to end for ONE deliverable on ONE scheduled run: build the prompt from the selected sections plus the carried state, one completion call, parse the reply. The carried state is computed by pure code before the call and adds no measurable time. Measured over the 120 scored rows of r001-cdrl-watch.
Current processWhat it replaces
The Monday morning pass down a contract deliverable register — thirty open CDRL lines, each with a due date, a submission status and a paragraph of somebody's notes — deciding which ones to put in front of the programme analyst this week and which to leave standing.
Where it is not good enough
⚑ THE HEADLINE IS THAT THE FREE FLOOR WINS, AND THIS KIT SAYS SO RATHER THAN BURYING IT. A rules engine with no model in it — the same date arithmetic, the same six carried scalars, and a regular expression over the analyst's prose — scores 100.00 pct on the band, 100.00 pct on the governing clock, 100.00 pct on the day count and 100.00 pct on the blocker, over all 120 deliverable-runs, for $0.00 (b003-cdrl-watch-rules). There is nothing left for a model to win on this corpus.
⚠︎ AND THE MODEL IS NOT MERELY NO BETTER, IT IS LESS RELIABLE. Its one missed raise on r001-cdrl-watch is on CDRL-0006-W4 — and the repeat probe asked that same row 3 times and got ESCALATE, WATCH, ESCALATE. The free rules engine is deterministic and is right on that row every time; the model is a coin toss on it. Variance is a worse property than a known bias, because a watch that raises a deliverable one Monday and drops it the next teaches its reader to distrust the list.
⚑ AND THE CONTROL SETTLES WHAT THE MODEL IS ADDING: NOTHING. Take the carried state away and the model does not merely score like the memoryless rules engine — it gives the SAME ANSWER on 480 of 480 graded cells, every field, every row (s001-cdrl-watch-stateless against b001-cdrl-watch-notes). Two readers that agree on every cell are not two readers. Whatever this task needs, a regular expression and a date subtraction already supply it.
The one place the two readers come apart is WORDING. Reword the notes without changing a single fact and the free pattern falls to 50.00 pct on the band and 0.00 pct on the blocker across the 24 rows where a revised due date is printed (b004-cdrl-watch-paraphrase-free). That is the whole measured case for paying for a model here, and it is an UPPER BOUND on the free reader's brittleness rather than an estimate, because the paraphrases and the pattern were written by the same author in the same repository.
The model's own side of that comparison is 100.00 pct on the same reworded rows (a002-cdrl-watch-paraphrase).
And two things are not the reader's fault at all. The raise gate is 5 days and the watch runs every 7, so there is a 2-day blind window in every cycle in which a due date can arrive without ever having been inside the gate on a run day: 5 of the 30 deliverables are never warned in time on ANY schedule. Shorten the interval to five days and that closes; no amount of reading does.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt120jsonl1json2
30 open contract deliverables, re-read whole on 4 weekly runs
and 50.00 pct free once the prose is reworded, model 100.00
2026-08-23as of
It produces a watchlist band, the date the deliverable is actually being judged against, the days it has accumulated past its raise gate and the blocker — and it submits nothing. There is no endpoint, no function and no flag anywhere in the kit that submits a deliverable, files an extension request or sends anything to a government point of contact. ⚑ TWO STATIONS MAKE THIS A MONITOR RATHER THAN A CLASSIFIER AND THEY ARE THE SECOND AND THIRD. The register is re-read WHOLE every run — nothing streams and nothing is sliced out — so the finding is never in a row, it is in the difference between this row and last week's; and the only route between the two is six scalars written by src/register.step from the PARSED record, never from the reply. ⚑ THE HEADLINE IS THAT THE FREE COLUMN WON. The rules engine on the fifth station's own terms — same arithmetic, same carried state, plus a regular expression over the prose — is right about the band, the clock, the day count and the blocker on all 120 rows for $0.00; the model at the seventh station is at 99.17 pct and 98.33, and is the only reader here that misses a raise. ⚑ AND THE CONTROL SETTLES WHAT THE SECOND STATION IS WORTH: remove the carried state and the model gives the IDENTICAL ANSWER to the memoryless free floor on 480 of 480 graded cells. Not the same score — the same answer. The two memoryless readers get 0 of 28 standing escalations and 0 of 30 accruing day counts between them, which is the size of the memory and the whole of what either reader is adding.
⚠︎ AND THE ONE EDGE WHERE THE MODEL WINS IS WORDING, MEASURED AS AN UPPER BOUND: reword the prose on the 24 rows that print a revised due date and the pattern falls to 50.00 pct band and 0.00 blocker while the model holds at 100.00 and 100.00 — but the paraphrases and the pattern share an author.
⚠︎ THE THIRD STATION'S FAILURE BELONGS TO THE SCHEDULE AND TO NOBODY ELSE: skipping one Monday put 16 raises a week late, lost 35 days of accrual and took standing-escalation accuracy from 100.00 pct to 63.16 with the reader unchanged, and the 2-day blind window between the 7-day interval and the 5-day raise gate loses 5 of 30 deliverables their warning outright. No amount of reading recovers either; a shorter interval does, and doubles the bill.
⚠︎ THE LADDER, THE REVIEW PERIOD AND ALL THREE ESCALATION RULES ARE INVENTED FOR THIS KIT — they reproduce no DFARS clause, no Data Item Description, no DD Form 1423 and no programme's CDRL management plan, and every band above is agreement with an invented rule.
⚠︎ AND ONE SECTION NEVER LEAVES THE MACHINE: Government Points of Contact — named contracting officers, their office, e-mail and telephone — reaches the provider in 0 of 120 records, subtracted unconditionally rather than by a fallback that happens never to fire.
The swap seams
Seam
File
What changes
The model
src/adapters/__init__.py
PROVIDER, BASE_URL, MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because the cost lens prices them.
The carried state
src/state.py
What is remembered between runs, and how it is put to the model. Six scalars and one English sentence per field. Grow this and the kit stops being cheap on run 400.
The cadence
src/cadence.py
CADENCE_DAYS. It is not a preference: it is set by the shortest gate on the ladder, and blind_window() computes the exposure the current interval leaves. Change one and re-read the other.
The watch rule
src/register.py
WATCH_AT_DAYS, RAISE_AT_DAYS, the governing-date order and the three previous-run rules. The shipped numbers are illustrative and reproduce no clause; this is the file a programme office rewrites first.
What leaves the machine
src/select.py
SECTION_HINTS and NEVER_SENT. A section no field maps to is subtracted unconditionally, so adding a field is also a disclosure decision.
The record shape
src/monitor.py
reading_of — the regular expressions that lift the dates and statuses off a record. Point the kit at a real register export and this is the function that changes.
Components
Component
File
Role
Corpus builder and the register export it produces
tools/build_corpus.py
The population, re-read WHOLE on every scheduled run: 30 open deliverables x 4 weekly runs = 120 records under data/corpus/, 4434 to 4758 bytes each, generated from one seed. Nothing is sliced out and nothing streams — a register is a list, and every open line is on every run. The answer key is the OUTPUT of src/register.replay, never typed, so a hand-written misreading of the rule cannot ride into the score.
Watch rule and arithmetic
src/register.py
The ladder, the governing-date decision and the three previous-run rules, in pure code. The gold is this module's output and so is every arm's carried state.
Cadence
src/cadence.py
SEAM 3 — the clock. What one run owns that the last did not, what a missed run costs, and the blind window between the raise gate and the run interval. Every figure it produces is free: the cost of a schedule is a property of the schedule.
Carried state
src/state.py
SEAM 2 — the thing that makes this a monitor. Six scalars per deliverable (band reported, ladder band observed, governing date, submission status, days at risk, run date), written from the arithmetic and never from the model's reply, rendered into English for the prompt. Constant cost on run 400 as on run 4.
Section splitter
src/segment.py
Splits a record into its eight named sections. Pure code.
Section selector
src/select.py
Decides what leaves the machine. Government Points of Contact is mapped by no field and is subtracted unconditionally — not by a fallback that happens never to fire.
Prompt builder
src/prompt.py
Three parts: the selected sections, the carried state, the question. The stateless control and the note-blind ablation are this builder with one part replaced or removed, and check_labels asserts they differ nowhere else.
Model adapter
src/adapters/__init__.py
SEAM 1 — the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost row.
Free floors
evals/baseline.py
Three of them: dates only, dates plus a pattern over the prose, and the full rules engine with memory. The third is the real competitor and it wins.
Scorer
evals/scoring.py
Exact match per cell against the computed gold, split ten ways an average would hide — the four fields, the watchlist, the quiet rows, the relief rows, the standing escalations, the recovery holds and the accruing rows. No judge model.
Pre-flight
evals/check_labels.py
Ten checks that must pass BEFORE a run may spend, including both directions of the never-sent guard, a re-derivation of every gold label, a proof that every reason code is reachable, and a proof that the one off-page fact is off the page.
Harness
evals/run.py
30 deliverable chains, four strictly-ordered runs each, 12 concurrent workers. The state advances even for a run whose CALL failed, so one transport error cannot turn into three scored failures.
Runtime
src/monitor.py
One deliverable, one scheduled run, one call. Parses every date and status off the record in pure code — the model is never asked to copy a date off a line — and leaves exactly one fact for the reader: whether a printed revised due date is carried by an executed modification. Holds MAX_TOKENS, the published ceiling.
Rewording probe
evals/paraphrase.py
Re-scores the 24 relief rows with the prose reworded and the facts unchanged, running the free pattern and the model over the identical rows. The only experiment on this kit that separates the two readers at all.
Repeat probe
evals/repeat.py
Asks the same rows again and reports how many gave a different answer — scored against ITSELF, not against the key. It is what showed this kit's single missed raise to be variance rather than a misreading.
Local UI
src/app.py
One deliverable, one run, on 127.0.0.1:9006. Shows the carried state verbatim, what this run owns that the last did not, the free floor's answer beside the model's, and which sections left the machine. Renders fully with no key.
Where it breaks at scale
⚑ THE CEILING IS THE CLOCK, NOT THE ROW COUNT. One call per open deliverable per run means a register of 2,000 open lines is 2,000 calls every Monday, which is linear and boring. What is not linear is the interval: the raise gate is 5 days and the run interval is 7, so every cycle carries a 2-day band in which a due date can arrive with no run ever having seen it inside the gate. On this corpus that costs 5 of 30 deliverables their warning outright, and no reader can recover it. Halving the interval halves the exposure and doubles the bill; that trade is the design decision this kit exists to make visible.
⚑ AND A MISSED RUN IS NOT A MISSED ROW, IT IS A MISSED WEEK. Skipping one Monday leaves the state intact — the accrual widens its interval correctly, which is why register.step takes the elapsed days from the carried run date rather than assuming a cadence — but 16 raises arrive a week late, 35 days of accrued risk go uncounted, 7 of 90 reported bands are wrong, and standing-escalation accuracy falls from 100.00 pct to 63.16 pct (b002-cdrl-watch-missedrun). A watch whose reliability nobody monitors is a watch that has already stopped.
⚠︎ THE STATE STORE IS A JSON FILE AND THAT IS THE SEAM THAT BREAKS FIRST. src/state.py writes one file atomically; two watches running against the same register would race it, and a register split across programmes needs one file per programme or a real store. seam:null is not an option here — the seam exists, it is state.load/state.save, and the point it is outgrown is the second concurrent watcher.
⚠︎ AND THE CARRIED STATE IS ONLY AS GOOD AS THE RUN THAT WROTE IT. Nothing in this kit detects a register whose line identifiers have been renumbered between runs; a re-keyed CDRL item arrives as a deliverable with no history, is reported on its ladder band alone, and every previous-run rule silently stops applying to it. That is not measured here and it is the first thing to build if you point this at a real register.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
CDRL-0018-W3, one live call, and it is this kit's whole finding on one screen: the AGREE column reads “same — free code got here too” on all four rows. The model answered ESCALATE against the contract due date, 7 days at risk, blocker AWAITING_GOVERNMENT_DATA — and so did the free rules engine beside it, for $0.00. The model's rationale is correct and worth reading: it applies step 3c, notes the previous run also raised the item and the status has not moved, and says the revised date is only proposed so the contract date governs. That is exactly right, and free code got there without being asked.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same row before anything is run, and the three things this page has to get right are already on it. The carried state, verbatim as it goes into the prompt — reported RAISE last Monday, judged against 2026-07-15, status IN PROGRESS. The cadence panel, naming which deliverables crossed which gate inside this run's own seven days. And the free rules engine's answer, computed locally for nothing. Note the Schedule block: it prints a revised due date of 2026-08-17 marked (proposed), exactly as the four deliverables whose relief is real do. Nothing on the structured page separates them.failureOpen full size →The same page with NO API_KEY configured, after pressing the button. It does not error and it does not go blank: the record, the carried state, the cadence, the parsed dates, the withheld-section list and the entire free rules engine are computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called. This is the honest failure state, not a staged one — and on this kit it is very nearly the whole product, because the free column is the one that wins.failureOpen full size →
How it is cutWhat one deliverables (4 scheduled runs each) is
No split and no chunking. The unit is a DELIVERABLE — one CDRL line read on four consecutive weekly runs, processed strictly in order because run 3's prompt contains state produced by run 2. Each record goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step. One deliverable record goes whole into one call, minus the one section no field asks for.
LicenceLicence
MIT — generated in this repository by tools/build_corpus.py from a fixed seed, so it reproduces byte for byte and belongs to nobody.
Bring your ownBring your own deliverable records
One record is ONE CDRL LINE AS IT STOOD ON ONE SCHEDULED RUN — not one deliverable and not one document. The same line appears again next week as a separate record, and the difference between the two is the whole finding, so a register exported once is a corpus of one run and this kit has nothing to compare it against.
Point src/monitor.reading_of at your own register export and rewrite src/register.RULE_TEXT, WATCH_AT_DAYS and RAISE_AT_DAYS to your programme's own lead windows — those three are the first things to change and the shipped values are illustrative. Then run the free rules floor against your own labelled month before you configure a provider at all: on this corpus it does the entire job, and the only measured reason to add a model is that your analysts' prose does not read like a template.
⚠︎ And what stops being true when you do: This kit reads a register and writes a watchlist. It does not read the deliverables themselves, and it must not: the atlas row it was built from marks the tracking metadata as not controlled technical data while noting that the CONTENT of a data item may well be. Keeping the content out of scope is what makes the export safe to send anywhere. If you extend this to read the documents, that boundary is gone and export control becomes your first design constraint, not your last.
What breaks it
A register whose line identifiers are renumbered between runs. Every carried scalar is keyed on the deliverable id; a re-keyed line arrives as a deliverable with no history and every previous-run rule silently stops applying to it. Not measured here.
A deliverable with no Program Notes at all. The section is where the modification's existence and the blocker both live; with it absent, both fall back to the structured page, which gets the blocker wrong on 73.33 pct of rows (the ladder floor's own figure) and reads a proposed revised date as if it governed.
Two revised due dates in flight at once — a modification executed for one date while a further extension is being requested. governing() takes one revised date and one boolean; it has nowhere to put the second.
A recurring deliverable (monthly, quarterly) where several instances are open at once. This corpus carries recurrence as a printed field and treats every line as one instance; a real register would need the instance, not the CDRL item, as the state key.
Deliverables whose due date is expressed relative to an event rather than as a date — '30 days after PDR'. reading_of parses ISO dates off fixed lines and would return nothing.
A register export where the analyst's prose carries an instruction rather than a description. The Program Notes are sent to the provider and are the one field an outside party can influence; they are this kit's injection surface, and it is deliberately not hidden.
A programme that treats the customer's silence as acceptance. This kit's review clock expires into a RAISE aimed at the government, not into a deemed acceptance, and nothing here can be configured to do the latter.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
160
not measured
instruction
1,632
not measured
carried state
416
not measured
Synthetic Record
478
not measured
Contract
276
not measured
Deliverable
276
not measured
Schedule
215
not measured
Submission Status
203
not measured
Watch Rule
2,520
not measured
Program Notes
412
not measured
Total
1,535
This is the cost lesson as arithmetic: of the 6,588 characters assembled, 4,380 are records — 66% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Both messages, in the order they are sent: the system message, then the user message replayed from the kit's own src/prompt.py::build on CDRL-0018-W3 with the exact carried state the arithmetic produces for that row — the identical code path every arm uses, not a paraphrase. THE SPLIT IS IN CHARACTERS AND THIS LINE SAYS SO: per-part TOKEN counts were not measured, because measuring them means sending nested prefixes of the prompt to a provider and this kit spent no calls on that. The billed input token count in tokens is the run's own average and is not the sum of the parts below.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You run one scheduled pass of a contract deliverable watch. You apply a written watch rule to one deliverable and answer with one JSON object and no other text.
You are running one scheduled pass of a contract deliverable watch. The watch re-reads every open
deliverable on a fixed weekly schedule; you are being shown ONE deliverable as it stands on this
run.
The watch rule, its ladder and the rules that depend on the previous run are reproduced in the
record below. Apply them exactly as written. The rules that depend on earlier runs cannot be
answered from this record: what is known about the previous run is stated under "Carried state" and
is the only history available to you. Do not assume anything about earlier runs beyond it.
Read the Program Notes. They are the only place the record says whether a printed revised due date
is carried by an EXECUTED contract modification or is merely proposed, and that decides which date
this deliverable is judged against.
Answer with a single JSON object and nothing else:
{"raise_band": "NONE|WATCH|RAISE|ESCALATE",
"governing_clock": "CONTRACT_DUE|REVISED_DUE|RESUBMISSION_DUE|CUSTOMER_REVIEW_SLA|NOT_APPLICABLE",
"days_at_risk": <whole number>,
"blocker_cause": "AWAITING_GOVERNMENT_DATA|AWAITING_INTERNAL_APPROVAL|AWAITING_SUBCONTRACTOR_INPUT|IN_CUSTOMER_REVIEW|REWORK_AFTER_REJECTION|RESOURCE_SHORTFALL|NO_BLOCKER",
"rationale": "one sentence, naming the rule or the gate you applied"}
"raise_band" is the band to REPORT after the rules that depend on the previous run have been
applied, which is not always the band the ladder alone would give.
"days_at_risk" is the running total AFTER this run, carried forward from the previous run's total
as the watch rule describes. It is a whole number of days and it is never negative.
Carried state
----------------------------------------------------------------
The previous run of this watch was on 2026-07-13. This run is 7 days after it. On that run this deliverable was reported RAISE. The dates alone already put it past its raise gate then. The date it was judged against was 2026-07-15. Its submission status was IN PROGRESS. 0 days at risk had been accumulated for it by the end of that run.
Deliverable record, as of this run
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This file is invented for a public evaluation kit. It is not a real contract deliverable record.
Every contract number, programme name, CDRL item, data item title, person, office and date in it
is fictional, and it reproduces no DD Form 1423, no Data Item Description, no contract and no
programme's deliverable register. The watch rule reproduced below was invented here and is not an
authority.
Contract
----------------------------------------------------------------
Contract number : W58RGZ-24-C-0088
Programme : HIGH LANTERN
Scope : Rotorcraft avionics refresh
Contractor : Ninebark Aviation
Administered by : DCMA Fort Barrow
Deliverable
----------------------------------------------------------------
CDRL item : A003
Data item title : Software Development Plan
Recurrence : One-time
Submission medium : Government contractor data portal
Register run : W3 (2026-07-20)
Schedule
----------------------------------------------------------------
Run date : 2026-07-20
Contract due date : 2026-07-15
Revised due date : 2026-08-17 (proposed)
Review period : 30 days
Submission Status
----------------------------------------------------------------
Status : IN PROGRESS
Last submitted : (none)
Customer disposition: (none)
Resubmission due : (none)
Watch Rule
----------------------------------------------------------------
DELIVERABLE WATCH RULE (illustrative -- reproduces no DFARS clause, no Data Item Description
and no programme CDRL management plan)
The watch runs every Monday at 06:00 local. On each run every deliverable that has not been
accepted is re-read whole, and each one is placed in one band.
Step 1 -- which date governs.
If the item is ACCEPTED, no date governs.
If the customer has REJECTED it, the resubmission due date governs.
If it has been submitted and the customer has it IN REVIEW, the review period governs and the
date is the submission date plus the review period printed on this record.
If a revised due date is printed AND a contract modification carrying it has been EXECUTED, the
revised date governs.
Otherwise the contract due date governs. A revised date that is only proposed -- a verbal
agreement, a request in flight, a modification not yet issued -- does NOT govern, and the
item is judged against the contract date until the modification exists.
Step 2 -- the ladder, against the governing date.
more than 15 days away NONE
15 days away or fewer WATCH
5 days away or fewer, or past RAISE
Step 3 -- what the previous run makes true, applied in this order.
a. If the previous run raised this item, the ladder no longer says RAISE, and the item has not
been accepted, report WATCH for one run (RECOVERY_HOLD). An item going quiet is the single
thing most worth keeping in front of somebody: it went quiet for a reason, and the reason --
a date that moved, a submission that has not been acknowledged -- is what has to be checked.
b. Otherwise, if the governing date is not the date the previous run recorded, report at least
WATCH and record the reason as DATE_MOVED. Somebody moved it and the modification has to
be seen.
c. Otherwise, if the ladder says RAISE, the previous run also raised it, and the submission
status has not changed since that run, report ESCALATE (STANDING_ESCALATION). One raise is a
reminder; the same raise with nothing having moved is a different message to a different
person.
Days at risk. The count of days this deliverable has spent past its raise gate, accumulated
across runs. It gains the interval between two runs when the ladder said RAISE on both of them,
holds otherwise, and resets to zero when the item is accepted. It is not on this page: it is
carried from the previous run.
Program Notes
----------------------------------------------------------------
Blocked: the government-furnished interface data we need for section 3 has not been delivered.
Requested from the programme office and still not received. We have asked for relief to
2026-08-17. The request went in on 2026-07-08 and is still with the contracting officer; nothing
has been executed, so the contract date still stands.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"raise_band":"NONE","governing_clock":"CONTRACT_DUE","days_at_risk":0,"blocker_cause":"NO_BLOCKER","rationale":"Contract due date 2026-10-05 is 91 days after the run date, so the ladder gives NONE and no previous-run rules apply on this first watch run."}
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a government contract's deliverable before it's late — 120 deliverable records. One model answered, and every answer was then graded Four different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Exact match, per cell, against a key the arithmetic computed. There is no judge model in this kit and that is not a shortcut: three of the four answers are words from a closed list (four bands, five clocks, seven blocker codes) and the fourth is a whole number of days. A judge is for answers whose correctness is a matter of reading; these are matters of equality, and asking a model to grade RAISE == RAISE would add cost, variance and a second thing to be wrong.
120deliverable records
120source documents
1model tier
4grading methods
MeasurementsWhat was measured
COUNTED84 · 88 · 120 · 83 · 119 · 88 / 120band accuracy pct — deliverable-runsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED58 · 62 · 94 · 66 · 93 · 62 / 94band on watchlist pct — deliverable-runs that should not be quietDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED12 · 0 · 0 · 0 · 1 · 0 / 60missed watchlist pct — deliverable-runs that had to be raisedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 0 · 0 · 0 · 0 / 26false alarm rate pct — quiet deliverable-runsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED8 · 12 · 24 · 16 · 24 · 12 / 24authority band pct — deliverable-runs printing a revised due dateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 28 · 12 · 27 · 0 / 28standing band pct — deliverable-runs whose answer is a standing escalationDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED108 · 120 · 120 · 90 · 120 · 120 / 120clock accuracy pct — deliverable-runsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED32 · 120 · 120 · 90 · 118 · 120 / 120blocker cause accuracy pct — deliverable-runsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED90 · 90 · 120 · 85 · 120 · 90 / 120days at risk accuracy pct — deliverable-runsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 30 · 14 · 30 · 0 / 30days at risk when accruing pct — deliverable-runs with a non-zero accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is exact comparison, so what needs validating is the KEY, not the comparator. evals/check_labels.py re-derives every one of the 120 labels from the records themselves through src/register.replay and refuses to let any run start if a single one disagrees. It also proves the three things a key of this shape can get silently wrong: that every reason code is reachable (the first draft's RECOVERY_HOLD was not, under a branch order that made it unreachable — a rule nobody can trigger is a rule nobody can score), that every gold value is inside the vocabulary the prompt names, and that the one off-page fact is genuinely off the page — some printed schedule-block SHAPE must carry both an executed and an unexecuted modification, and the length of the relief must coincide between the two classes so no rule can key on the gap alone. The second of those failed on the first build and was fixed in the corpus, not in the checker.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate — a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One deliverable record
1,000 deliverable records
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.000000
$0.00
0%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.000000
$0.00
0%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.000000
$0.00
0%
Same work, 0× the bill
The same deliverable records, the same tokens — only the rate card changed. And across all 3 cards about 0% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE CADENCE, not the prompt and not the model. Every other knob here moves the bill by percentages; the run interval moves it by a factor. Seven days against fourteen is half the bill and twice the blind window (2 days against 9); five days is a fifth more bill and closes the window entirely. Nothing about the reader changes in any of those.
Rates checked 2026-08-23. The provider that actually ran the calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading is evals/scoring.py in-process: no key, no model, no network, and it runs over all 120 rows in well under a second. Pricing this ruler in the units of the thing it measures — 'one call per deliverable-run' — would put a dollar figure on something that genuinely costs nothing.
The gradersFour ways to grade
⚑ THIS FLOOR IS THE HEADLINE, NOT THE COMPARATOR. It scores 100.00 pct on the band, 100.00 pct on the clock, 100.00 pct on the day count and 100.00 pct on the blocker over all 120 rows, for $0.00. A model can tie it and cannot beat it.
⚠︎ AND ITS PERFECTION IS PARTLY A FACT ABOUT THE WORDING. The pattern in evals/baseline.py was written in this repository against the sixteen note templates in tools/build_corpus.py, also written in this repository. b004-cdrl-watch-paraphrase-free puts a number on that: the same 24 rows with the prose reworded and not one fact changed, and the pattern falls to 50.00 pct on the band and 0.00 pct on the blocker. Its errors under rewording are all in the SAFE direction — with no marker matched it treats relief as unreal, which over-raises rather than under-raises — but that is a property of how it was written, not a guarantee.
⚠︎ THE PARAPHRASES SHARE THAT AUTHOR TOO. An author writing paraphrases knows exactly which words to avoid, so the probe is an UPPER BOUND on brittleness and not an estimate.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The four answered fields, per deliverable-run, exact match against the computed answer key For each of the 120 deliverable-runs and each of the four answered fields, did the reply equal the computed key? Band, clock and blocker are compared exactly after trimming and upper-casing; the day count is compared as a whole number, and a reply that is not a whole number is a MISS rather than a zero — 0 is a meaningful answer here (it is what a deliverable inside its gate gets) and a parse failure must never be scored as one.
$0.00
no
yes
the free rules engine, with memory — THE FLOOR THAT WINS 100.0% band accuracy · the same, with the memory removed 73.3% band accuracy · dates and structured fields only — the spreadsheet 70.0% band accuracy · the winning floor, on a schedule that missed one Monday 92.2% band accuracy · the harness self-test — state-blind and prose-blind, by design 54.2% band accuracy · the fast tier, with the carried state 99.2% band accuracy · the fast tier, memory removed (THE CONTROL) 73.3% band accuracy · the fast tier, CEILING CALIBRATION ONLY 100.0% band accuracy · 3 more measured on each run
Missed raises and false alarms, counted apart Of the 60 deliverable-runs that had to be raised, how many were reported NONE or WATCH (a MISSED RAISE); and of the 26 quiet rows, how many were disturbed (a FALSE ALARM). Two rates over two different denominators, never averaged.
$0.00
no
yes
the fast tier, with the carried state 1.7% value · the free rules engine, with memory 0.0% value · the free floor, no memory 0.0% value · dates and structured fields only 20.0% value · the harness self-test 0.0% value · the fast tier, memory removed (THE CONTROL) 0.0% value
Band accuracy on the rows a single reading cannot answer Band accuracy restricted to the 28 deliverable-runs whose correct answer is a standing escalation — the item was already being raised last run, is still past its gate, and its submission status has not moved. Nothing in the record in front of you says any of that.
$0.00
no
yes
the fast tier, with the carried state 96.4% value · the free rules engine, with memory 100.0% value · the free floor, no memory 0.0% value · dates and structured fields only 0.0% value · the harness self-test 0.0% value · the fast tier, memory removed (THE CONTROL) 0.0% value
Band accuracy where the record prints a revised due date Band accuracy restricted to the 24 deliverable-runs whose Schedule block prints a revised due date marked (proposed). Half are carried by an executed contract modification and half are not; the structured page is the same shape either way and only the analyst's prose separates them. This is the only place on this kit where reading earns anything.
$0.00
no
yes
the fast tier, with the carried state 100.0% value · the free rules engine, with memory 100.0% value · the free floor, no memory 50.0% value · dates and structured fields only 33.3% value · the harness self-test 16.7% value · the fast tier, memory removed (THE CONTROL) 50.0% value
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
⚠︎ THIS LABELLED SET CANNOT TELL A GOOD READER FROM A PERFECT ONE, AND THE REASON IS THAT THE FREE FLOOR ALREADY SCORES 100.00 pct. There is no headroom above it, so any arm that ties it is indistinguishable from any other arm that ties it, and the only discriminating measurement on this corpus is the paraphrase probe — 24 rows, which is a small denominator and the page says so beside every rate.
What the set CAN separate, and does by more than seventy points, is readers that differ in what they consult: the spreadsheet floor at 26.67 pct on the blocker against the pattern floor at 100.00 pct, and every memoryless arm at 0.00 pct on the 28 standing escalations against 100.00 pct with memory. Those are the two axes this corpus was built to separate, and it separates them cleanly.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
You have a deliverable register and a written lead-time rule, and the notes are written to a house template.
The free rules engine (evals/baseline.py, rules). No provider, no key, no bill.
It scored 100.00 pct on every field of this corpus. A model cannot beat 100 and costs money to tie it.
Do not add a model here to look modern. You will pay per row per week for an answer a regular expression already had, and you will have added a dependency that can be unavailable on a Monday morning.
The notes are free prose from many people, and house style changes.
The model arm, with the free floor still computed beside it on every row.
The free pattern's whole score on the rows that matter is a fact about wording: reword them and it falls from 100.00 pct to 50.00 pct on the band and to 0.00 pct on the blocker.
Do not drop the free column when you add the model. It is what tells you, row by row, whether you are buying anything — and on most rows you are not.
You want the raise list to be earlier rather than more accurate.
Neither. Change the CADENCE.
A 7-day interval against a 5-day raise gate leaves a 2-day band in which a due date arrives with no run ever having seen it inside the gate: 5 of 30 deliverables here are never warned in time on any schedule, by any reader.
Do not buy a better reader to fix a schedule problem. It cannot, and the measurement that proves it costs nothing.
You cannot guarantee the watch runs every week.
Monitor the watch itself before you monitor anything with it.
One skipped Monday put 16 raises a week late, lost 35 days of accrual and took standing-escalation accuracy from 100.00 pct to 63.16 pct — with the reader unchanged.
Do not treat a missed run as a gap in coverage you can catch up on. The state carries correctly across the gap; the WARNINGS do not.
Your register is keyed on something that gets renumbered.
Nothing here yet. Build the identity check first.
Every carried scalar is keyed on the deliverable id. A re-keyed line arrives with no history and every previous-run rule silently stops applying to it.
Do not deploy this against a register whose keys move. Not measured here, and the failure is silent.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
trusted-a-proposed-date
Read a printed revised due date as if a modification carried it
12
CDRL-0018-W3: the Schedule block prints Revised due date : 2026-08-17 (proposed) and the prose says the modification has not been issued. The spreadsheet floor reports NONE; the deliverable is five days past its contract date.
no-memory-no-escalation
Reported a first raise where the answer was a standing escalation
28
Every one of the 28 standing-escalation rows, for both memoryless floors. The record says the item is past its date; only the previous run says it was already raised and that nothing has moved since.
no-memory-no-accrual
Reported zero days at risk where days had accumulated
30
All 30 rows with a non-zero accrual, for both memoryless floors. The count is carried, not printed: b001-cdrl-watch-notes scores 0.00 pct on them and b003-cdrl-watch-rules scores 100.00 pct.
blind-window
Never warned in time, on any schedule
5
Two deliverables due six days after a run: outside the 5-day raise gate on the Monday the watch looked, already past due on the Monday after. Three more were already overdue before the watch was stood up. No reader can recover any of the five; a five-day…
raise-a-week-late
A raise the schedule delivered a week after it was due
16
Measured by replaying the same free rules engine on a schedule with the 2026-07-20 run never made (b002-cdrl-watch-missedrun). 16 raises move by a week and 35 days of accrual go uncounted; not one band is wrong for a reason the reader could fix.
model-under-raised
The model left a deliverable standing that the rules escalated
1
CDRL-0006-W4: answered WATCH where the computed key says ESCALATE. ⚠︎ THIS IS THE EXPENSIVE DIRECTION — an under-raise, not a false alarm — and it is the ONLY missed raise anywhere on this page. All three free floors are at 0. ⚑ AND THE REPEAT PROBE LANDED ON…
model-mis-named-the-blocker
The model named a blocker the record does not support
2
CDRL-0013-W4 answered IN_CUSTOMER_REVIEW where the key says NO_BLOCKER; CDRL-0029-W2 answered NO_BLOCKER where the key says RESOURCE_SHORTFALL. The blocker decides WHO gets chased, so this is not a cosmetic field: a deliverable read as sitting with the…
What we could NOT verify
Whether the free pattern's collapse under rewording (100.00 pct to 50.00 pct on the band) is representative. The paraphrases and the pattern were written by the same author in the same repository, so the probe is an UPPER BOUND on brittleness rather than an estimate. A real register's prose sits somewhere between the two corpora and nothing here says where.
Whether rendering the carried state as English rather than as JSON changes anything. src/state.describe argues for English and the argument is a design choice, not a measurement; the experiment is one more 120-row run and has not been paid for.
Whether the branch order in src/register.step matches how a programme office would rank the rules. Recovery is tested before a moved date, so a raised deliverable whose date moves out is reported under RECOVERY_HOLD rather than DATE_MOVED. Both report WATCH, so no band changes; the reason code does.
Run-to-run variance beyond the 10 rows the repeat probe re-asked. p002-cdrl-watch-repeat asked 10 rows 3 times each — 30 asks over 8.3 pct of the corpus — and 1 of them came back with a different answer. That is a measurement of THOSE rows, not a stability rate for the kit, and this page never states it as one. The free arms are deterministic and reproduce exactly.
Whether a deliverable whose identifier is renumbered between runs is detected. It is not, and the failure is silent: the line arrives with no history and every previous-run rule stops applying to it.
Whether the two blind-window deliverables are the right SIZE of problem. The corpus carries two because two make the exposure observable; what fraction of a real register lands in that band depends on how due dates cluster, and nothing here measures it.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the free age-and-threshold floor
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
the free floor, plus a pattern over the prose
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
the free floor, with the carried state and all three rules
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
the fast tier, with the carried state
1,535.27
1,066.67
8,318 ms
$0.003968
$0.001587
$0.068686
the fast tier, memory removed (THE CONTROL)
1,486.47
818.71
6,729 ms
$0.003199
$0.001280
$0.055800
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-23. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same one deliverable on one scheduled run, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
⚑ GRADING THIS KIT IS FREE, AND IT IS A MEASURED ZERO RATHER THAN AN UNPRICED ONE. evals/scoring.py is exact comparison in-process: no key, no model, no network. So are all three floors, the missed-run schedule, the stub, the ten-check pre-flight and the free half of the paraphrase probe — every headline figure on this page was produced without calling a provider at all.
332 provider calls stand behind the figures on these pages. Every one of them is behind a figure above: 8 ceiling calibration, 120 scored, 120 the stateless control, 30 the note-blind ablation, 30 the repeat probe (10 rows asked 3 times) and 24 the paid half of the paraphrase probe. NOTHING WAS DISCARDED and no run was re-fired. ⚠︎ The ledger also carries 121 calls this kit was REFUSED earlier the same day, when the shared account hit zero balance mid-run: 402 Payment Required returns no completion and is billed for none, and they appear on the ledger only because src/budget.py records a call before making it — the direction a spend guard should round.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, and specifically provider-side reasoning. The calibration billed 9879 output tokens against 12399 input across 8 calls, and 9040 of the output (91.51 pct) never reached the reply text at all — it was the model working through the ladder, the governing-date decision and the previous-run rules.
THE NUMBER OF OPEN DELIVERABLES, multiplied by the number of runs. This is the unit that surprises people: the bill is per row PER PERIOD, not per document. A register of 200 open lines watched weekly is 10,400 calls a year, not 200.
THE CADENCE. Halving the interval to close the blind window doubles the bill exactly. That trade is the design decision this kit exists to make visible, and it is a cost decision before it is an accuracy one.
INPUT TOKENS, and they are nearly constant. Every record is between 4434 and 4758 bytes by construction and the carried state is six scalars in one paragraph, so input barely moves — 1535 tokens a row. On this kit input is the small half of the bill.
Your volumeWhat it costs at your volume
Linear in rows and linear in runs, with no cliff in between — there is no index to rebuild, no retrieval to widen and no context that grows. Ten times the register is ten times the calls at the same price per call; ten times the frequency is the same. The only non-linearity is the one below.
Where pricing changes shape
A register large enough that one Monday's run does not finish before the next one starts. At the calibration's p50 of 8318 ms and 12 concurrent chains, that is roughly 872,517 open deliverables — well beyond any single programme, but it is where the design stops being a cron job and starts being a queue.
⚠︎ AND A CLIFF THIS KIT HIT RATHER THAN PREDICTED: the shared provider account reaching zero balance mid-run. 402 Payment Required is not a rate limit and is not retried by src/adapters — it is terminal, it refuses every call in the run, and nothing is billed. A watch on a clock fails ENTIRELY on the Monday its account empties, and the register looks exactly the same as a register with nothing to raise.
Your return, with your numbers
Volume30 open deliverables re-read every 7 days. That is 1560 rows a year on this corpus's cadence; scale it by your own register.
What it replacesThe weekly manual pass down the register — reading each open line, comparing it against last week's list, and deciding what to put in front of the analyst.
Time saved per itemNOT MEASURED. Nobody was timed doing this by hand, so no minutes figure is published. What IS measured is the alternative: the free rules engine does the whole job in under a second for all 120 rows at $0.00, which is the number to put against a person's hour before any model enters the conversation.
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the whole series runs on, so the rows compare across use cases. Nothing about this task argues for a bigger one: the free rules engine already scores 100.00 pct, so a more capable model cannot buy accuracy here — only robustness to wording, which is the one thing this kit has not been able to measure on the model side.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
184,232input tokens · this run
128,001output tokens
—not priced — no committed card for the provider that ran it
120 call(s) on r001-cdrl-watch, one completion each, one tier.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.190
$0.190
$1.59
2026-09-12
gemini-3-flash
Google
$0.476
$0.476
$3.97
2026-09-18
gemini-3-8-flash
Google
$0.618
$0.618
$5.15
2026-09-18
llama-5
Meta
$0.774
$0.774
$6.45
2026-09-18
claude-haiku-4-5
Anthropic
$0.824
$0.824
$6.87
2026-09-12
grok-4-5
xAI
$1.136
$1.136
$9.47
2026-09-18
grok-4-6
xAI
$1.136
$1.136
$9.47
2026-09-18
claude-sonnet-5
Anthropic
$1.648
$1.648
$13.74
2026-09-12
gemini-3-1-pro
Google
$1.904
$1.904
$15.87
2026-09-18
gpt-5-6-terra
OpenAI
$1.904
$1.904
$15.87
2026-09-12
gpt-5-6-sol
OpenAI
$3.297
$3.297
$27.47
2026-09-12
claude-opus-4-8
Anthropic
$4.121
$4.121
$34.34
2026-09-12
claude-opus-5
Anthropic
$4.121
$4.121
$34.34
2026-09-12
claude-fable-5
Anthropic
$8.242
$8.242
$68.69
2026-09-18
claude-fable-5-1
Anthropic
$8.242
$8.242
$68.69
2026-09-18
gpt-6-astra
OpenAI
$8.242
$8.242
$68.69
2026-09-17
Read this against the numbers above
A projection, not a bill. Nobody paid any of these prices and the provider that actually ran the calls is deliberately absent from the table.
Rates are the vendors' published list prices on the date recorded beside each row, and they move.
Reasoning tokens are billed as output on every card here, and on this workload output is the larger half of the bill — so a model that reasons less would be cheaper than its rate card suggests, and one that reasons more would not.
The unit is one deliverable on one scheduled run. Multiply by your register size AND by your run frequency; the second multiplier is the one people forget.
Every figure here is a token count times a rate. Nothing about accuracy is projected, and on this kit accuracy is where the argument is — the free floor is at 100.00 pct for $0.00.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
16 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pyCorpus builder and the register export it produces
The population, re-read WHOLE on every scheduled run: 30 open deliverables x 4 weekly runs = 120 records under data/corpus/, 4434 to 4758 bytes each, generated from one seed. Nothing is sliced out and nothing streams — a register is a list, and every open line is on every run. The answer key is the OUTPUT of src/register.replay, never typed, so a hand-written misreading of the rule cannot ride into the score.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic, seeded, no network, no model.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
RULE = "-" * 64
RUN_DATES = [datetime.date(2026, 7, 6), datetime.date(2026, 7, 13),
RUN_LABELS = ["W1", "W2", "W3", "W4"]
CADENCE_DAYS = 7
BANNER = """\
src/register.pyWatch rule and arithmetic — a swap seam
The ladder, the governing-date decision and the three previous-run rules, in pure code. The gold is this module's output and so is every arm's carried state.
You change it to: WATCH_AT_DAYS, RAISE_AT_DAYS, the governing-date order and the three previous-run rules. The shipped numbers are illustrative and reproduce no clause; this is the file a programme office rewrites first.
src/register.py
# The watch rule, and the arithmetic it implies. Pure code, no model.
NONE, WATCH, RAISE, ESCALATE = "NONE", "WATCH", "RAISE", "ESCALATE"
BANDS = (NONE, WATCH, RAISE, ESCALATE)
CONTRACT_DUE = "CONTRACT_DUE"
REVISED_DUE = "REVISED_DUE"
RESUBMISSION_DUE = "RESUBMISSION_DUE"
CUSTOMER_REVIEW_SLA = "CUSTOMER_REVIEW_SLA"
NOT_APPLICABLE = "NOT_APPLICABLE"
CLOCKS = (CONTRACT_DUE, REVISED_DUE, RESUBMISSION_DUE, CUSTOMER_REVIEW_SLA, NOT_APPLICABLE)
AWAITING_GOVERNMENT_DATA = "AWAITING_GOVERNMENT_DATA"
src/cadence.pyCadence — a swap seam
SEAM 3 — the clock. What one run owns that the last did not, what a missed run costs, and the blind window between the raise gate and the run interval. Every figure it produces is free: the cost of a schedule is a property of the schedule.
You change it to: CADENCE_DAYS. It is not a preference: it is set by the shortest gate on the ladder, and blind_window() computes the exposure the current interval leaves. Change one and re-read the other.
src/cadence.py
# The clock this watch runs on, and what a missed run costs. Pure code, no model.
CADENCE_DAYS = 7
CADENCE_TEXT = "every Monday at 06:00 local, seven days apart"
CADENCE_TRIGGER = ("a scheduler on the machine that holds the register export — cron, a scheduled "
def blind_window(cadence_days=CADENCE_DAYS, raise_at=R.RAISE_AT_DAYS):
def owned_by(run_dates, index, rows_by_unit):
def replay_schedule(run_dates, rows, keep):
def never_warned(run_dates, rows, keep):
def missed_run_cost(run_dates, units, skip):
def next_run(after, cadence_days=CADENCE_DAYS):
src/state.pyCarried state — a swap seam
SEAM 2 — the thing that makes this a monitor. Six scalars per deliverable (band reported, ladder band observed, governing date, submission status, days at risk, run date), written from the arithmetic and never from the model's reply, rendered into English for the prompt. Constant cost on run 400 as on run 4.
You change it to: What is remembered between runs, and how it is put to the model. Six scalars and one English sentence per field. Grow this and the kit stops being cheap on run 400.
src/state.py
# The carried state — the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_unit(store, unit_id):
def advance(store, unit_id, run_date, row):
def _serialise(st):
def _as_date(v):
def describe(state, run_date=None):
src/segment.pySection splitter
Splits a record into its eight named sections. Pure code.
src/segment.py
# Split one deliverable record into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Contract", "Deliverable", "Schedule", "Submission Status",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pySection selector — a swap seam
Decides what leaves the machine. Government Points of Contact is mapped by no field and is subtracted unconditionally — not by a fallback that happens never to fire.
You change it to: SECTION_HINTS and NEVER_SENT. A section no field maps to is subtracted unconditionally, so adding a field is also a disclosure decision.
src/select.py
# Pick which sections of a deliverable record are sent. Pure code — the last deterministic step
BANNER = "Synthetic Record"
CONTRACT = "Contract"
DELIVERABLE = "Deliverable"
SCHEDULE = "Schedule"
STATUS = "Submission Status"
RULEBLOCK = "Watch Rule"
NOTES = "Program Notes"
CONTACTS = "Government Points of Contact"
NEVER_SENT = (CONTACTS,)
src/prompt.pyPrompt builder
Three parts: the selected sections, the carried state, the question. The stateless control and the note-blind ablation are this builder with one part replaced or removed, and check_labels asserts they differ nowhere else.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework — string concatenation
INSTRUCTION = """\
NO_HISTORY = ("No previous run is available for this deliverable. Judge it on this record alone, "
def build(text, carried, run_date=None, stateless=False, note_blind=False):
src/adapters/__init__.pyModel adapter — a swap seam
SEAM 1 — the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost row.
You change it to: PROVIDER, BASE_URL, MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because the cost lens prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
evals/baseline.pyFree floors
Three of them: dates only, dates plus a pattern over the prose, and the full rules engine with memory. The third is the real competitor and it wins.
evals/baseline.py
# The free floors. No model, no key, no network — arithmetic and regular expressions.
def blocker_from_structure(row):
def _flat(notes):
def blocker_from_notes(notes, row):
def authority_from_notes(notes):
def resolve(row, use_notes):
def answer(row, run_date, state, use_notes, use_state):
FLOORS = {
evals/scoring.pyScorer
Exact match per cell against the computed gold, split ten ways an average would hide — the four fields, the watchlist, the quiet rows, the relief rows, the standing escalations, the recovery holds and the accruing rows. No judge model.
evals/scoring.py
# Score a run against the computed gold. Pure code, no model, no judge.
FIELDS = ("raise_band", "governing_clock", "days_at_risk", "blocker_cause")
RAISED = ("RAISE", "ESCALATE")
def _pct(n, d):
def _norm(v):
def _int(v):
def score(records, golds):
def _gold_key(field):
def compare(a, b):
evals/check_labels.pyPre-flight
Ten checks that must pass BEFORE a run may spend, including both directions of the never-sent guard, a re-derivation of every gold label, a proof that every reason code is reachable, and a proof that the one off-page fact is off the page.
evals/check_labels.py
# Everything that must be true BEFORE a run may spend. Free, no key, no network.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def check():
def main():
evals/run.pyHarness
30 deliverable chains, four strictly-ordered runs each, 12 concurrent workers. The state advances even for a run whose CALL failed, so one transport error cannot turn into three scored failures.
evals/run.py
# Run the watch over the 30 deliverables and write one result file. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def load_gold():
GOLDS = load_gold()
def true_row(doc_id):
def stub_complete(cfg, system, user, max_tokens=1024):
def main():
src/monitor.pyRuntime — a swap seam
One deliverable, one scheduled run, one call. Parses every date and status off the record in pure code — the model is never asked to copy a date off a line — and leaves exactly one fact for the reader: whether a printed revised due date is carried by an executed modification. Holds MAX_TOKENS, the published ceiling.
You change it to:reading_of — the regular expressions that lift the dates and statuses off a record. Point the kit at a real register export and this is the function that changes.
src/monitor.py
# One deliverable, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You run one scheduled pass of a contract deliverable watch. You apply a written watch "
MAX_TOKENS = 8000
FIELDS = ("raise_band", "governing_clock", "days_at_risk", "blocker_cause")
def documents():
def units():
def load_doc(doc_id):
def _one(text, label, default=""):
evals/paraphrase.pyRewording probe
Re-scores the 24 relief rows with the prose reworded and the facts unchanged, running the free pattern and the model over the identical rows. The only experiment on this kit that separates the two readers at all.
evals/paraphrase.py
# The paraphrase probe: does the free floor's perfect score survive a rewording?
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PARA = os.path.join(HERE, "data", "paraphrase.json")
def restated(doc_id, note):
def main():
evals/repeat.pyRepeat probe
Asks the same rows again and reports how many gave a different answer — scored against ITSELF, not against the key. It is what showed this kit's single missed raise to be variance rather than a misreading.
evals/repeat.py
# The repeat probe: ask the same rows again and see whether the answer moves. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
GOLDS = {r["doc_id"]: r for r in (json.loads(l) for l in open(GOLD, encoding="utf-8") if l.strip())}
def pick(n=10):
def carried_for(doc_id):
def main():
src/app.pyLocal UI
One deliverable, one run, on 127.0.0.1:9006. Shows the carried state verbatim, what this run owns that the last did not, the free floor's answer beside the model's, and which sections left the machine. Renders fully with no key.
src/app.py
# The minimal local UI. Standard library only — python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9006"))
def _gold():
GOLD_ROWS = _gold()
def _row(doc_id):
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyThe population, re-read WHOLE on every scheduled run: 30 open deliverables x 4 weekly runs = 120 records under data/corpus/, 4434 to 4758 bytes each, generated from one seed. Nothing is sliced out and nothing streams — a register is a list, and every open line is on every run. The answer key is the OUTPUT of src/register.replay, never typed, so a hand-written misreading of the rule cannot ride into the score.
src/register.pyThe ladder, the governing-date decision and the three previous-run rules, in pure code. The gold is this module's output and so is every arm's carried state. A swap seam.
src/cadence.pySEAM 3 — the clock. What one run owns that the last did not, what a missed run costs, and the blind window between the raise gate and the run interval. Every figure it produces is free: the cost of a schedule is a property of the schedule. A swap seam.
src/state.pySEAM 2 — the thing that makes this a monitor. Six scalars per deliverable (band reported, ladder band observed, governing date, submission status, days at risk, run date), written from the arithmetic and never from the model's reply, rendered into English for the prompt. Constant cost on run 400 as on run 4. A swap seam.
src/segment.pySplits a record into its eight named sections. Pure code.
src/select.pyDecides what leaves the machine. Government Points of Contact is mapped by no field and is subtracted unconditionally — not by a fallback that happens never to fire. A swap seam.
src/prompt.pyThree parts: the selected sections, the carried state, the question. The stateless control and the note-blind ablation are this builder with one part replaced or removed, and check_labels asserts they differ nowhere else.
src/adapters/__init__.pySEAM 1 — the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost row. A swap seam.
evals/baseline.pyThree of them: dates only, dates plus a pattern over the prose, and the full rules engine with memory. The third is the real competitor and it wins.
evals/scoring.pyExact match per cell against the computed gold, split ten ways an average would hide — the four fields, the watchlist, the quiet rows, the relief rows, the standing escalations, the recovery holds and the accruing rows. No judge model.
evals/check_labels.pyTen checks that must pass BEFORE a run may spend, including both directions of the never-sent guard, a re-derivation of every gold label, a proof that every reason code is reachable, and a proof that the one off-page fact is off the page.
evals/run.py30 deliverable chains, four strictly-ordered runs each, 12 concurrent workers. The state advances even for a run whose CALL failed, so one transport error cannot turn into three scored failures.
src/monitor.pyOne deliverable, one scheduled run, one call. Parses every date and status off the record in pure code — the model is never asked to copy a date off a line — and leaves exactly one fact for the reader: whether a printed revised due date is carried by an executed modification. Holds MAX_TOKENS, the published ceiling. A swap seam.
evals/paraphrase.pyRe-scores the 24 relief rows with the prose reworded and the facts unchanged, running the free pattern and the model over the identical rows. The only experiment on this kit that separates the two readers at all.
evals/repeat.pyAsks the same rows again and reports how many gave a different answer — scored against ITSELF, not against the key. It is what showed this kit's single missed raise to be variance rather than a misreading.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1535 input and 1066 output tokens per one deliverable on one scheduled run, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
One deliverable on one scheduled runs/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per one deliverable on one scheduled run directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
One outbound HTTPS call per deliverable per scheduled run, to whichever provider is configured, carrying seven of the record's eight sections. Everything else is local: the register is a directory of text files, the carried state is a JSON file beside them, the answer key and every score are computed in-process. There is no database, no queue, no framework and nothing listening on a socket except the optional local UI, which binds 127.0.0.1.
The key is read from <repo>/.env, then the kit's own .env, then the real environment, and it is never written to a result file, a log line or a screenshot. src/app.py strips the key and the base URL out of any provider error before returning it to the browser. The repository has never held a credential: both .env files are gitignored from the first commit and .env.example carries empty values.
The experimentWhat the record looks like with the withheld section, and what leaves
Every record carries eight sections and seven reach the provider. The local UI prints the list on every row — sent, or WITHHELD — because a page that simply does not mention the contacts cannot be told apart from one that quietly sent them. Both screenshots on this page show that table with the withheld row in it. python3 -m evals.check_labels # free, no key, ten checks including this one. Last run on 2026-08-23, on the committed checkout.
The boundary
What a naive build does here
What this kit does, and how far it was proven
Do named government contacting officers ever leave the machine?
Every record carries a Government Points of Contact section — an administrative contracting officer, a procuring contracting officer, an office, an e-mail address and a telephone number. It is the only section whose subject is a PERSON, and it is the customer's people rather than the contractor's. A selector that fell back to the whole record would send all of it to a third-party API on every run of every deliverable.
src/select.NEVER_SENT holds it and _fallback() subtracts it UNCONDITIONALLY, so a field hint that matches nothing falls back to the record MINUS that section rather than to the record. Red-proven in both directions by evals/check_labels.py over all 120 records: 0 send it, and 0 reach it through the unmapped-field fallback. ⚠︎ THE GUARD IS NOT ON A LIVE PATH ON TODAY'S CORPUS — every hint names a section all 120 records carry, so the naive fallback would leak nothing here either. The hole is CONDITIONAL: one deliverable arriving with no Program Notes, which is entirely ordinary, is all it takes.
Can this kit submit a deliverable, request relief, or contact the customer?
A watch that produces a raise list is one short step from a watch that acts on it, and the atlas row this kit was built from marks that step as out of scope absolutely. The obvious build adds a portal client, an extension-request template and a mail step, all behind a flag.
There is no such endpoint, function, flag or argument. The only HTTP client in the kit is src/adapters/__init__.py, which posts to the configured completion endpoint and nowhere else; there is no SMTP, no webhook, and no file write outside data/ and results/. ⚠︎ THIS IS AN ABSENCE ARGUED FROM READING THE KIT, NOT A CHECK THAT RUNS — no gate here greps for it, and guardrails.add_first names adding that grep as work still owed.
Can a model's wrong answer poison the next scheduled run?
The natural way to carry state on a monitor is to feed the model's own last verdict into its next prompt. One bad week then compounds into every week after it, and the eval can no longer tell a single miss from a run that went wrong once and never recovered.
src/register.step() takes the PARSED record and the previous state and never reads a reply; evals/run.py calls it with the pre-run state in every branch, including the exception branch. Measured consequence on r001-cdrl-watch: 3 wrong cells of 480, and no later run in any of those chains inherits them. ⚠︎ The property is guaranteed by the code path; this run is CONSISTENT with it rather than a demonstration of it.
Does a run whose call failed corrupt the chain behind it?
Freezing the carried state when a call errors looks defensive and is not: the next run is then judged against a state that skipped a week, so one transport error becomes three scored failures and hides the real one.
evals/run.py's exception branch calls register.step() before continuing, so the schedule advances even when the provider does not answer. ⚑ EXERCISED FOR REAL, NOT ASSERTED: on 2026-08-23 a shared provider account hit zero balance mid-run and refused all 120 calls with 402; every one of the 30 chains still advanced through all 4 scheduled runs and the result file was structurally complete.
Four gates, all of them code rather than policy. The first three are about one section of one document shape and are testable; the fourth is an absence and is testable by grep. None of them is a model behaviour, which is the point — a guarantee that depends on a model honouring an instruction is not a guarantee.
The resultSeven of eight sections leave the machine. The eighth — named government contacts — is subtracted unconditionally, not by a fallback that happens never to fire.
0 of 120records that send a withheld section
0 of 120records whose unmapped-field fallback reaches one
0 of 0attack trials run against the injection surface
3 of 4boundaries red-proven rather than argued
evals/check_labels.py check 3, run over every record on the committed checkout. It asks two questions of each: does select.sent include anything on NEVER_SENT, and does select.for_field with a field name that maps to nothing include anything on it. Both answer no on all 120.
Read this twice
The Program Notes ARE sent, deliberately, and that is the opposite decision from the one above it. They are the only place the record states whether a contract modification exists, so a kit that withheld them would be measuring nothing at all — and they are prose an outside party could influence, which makes them this kit's own injection surface. Hiding them would hide the surface from the run that is supposed to measure it. The two decisions look inconsistent and are not: one section is data nobody asked for, the other is the data the whole question turns on.
HonestyWhat this does not prove
Provider-side retention. What the configured provider keeps, for how long, and whether it trains on it is a property of the account and the contract, not of this code. Nothing here measures it and nothing here can.
Whether the Program Notes are injectable in practice. The surface is named and no attack was run against it.
Whether a real register export carries personal data in a section this kit sends. The corpus is generated and its contacts are confined to one section by construction; a real export may put a name anywhere, and SECTION_HINTS would then be a privacy decision as well as a prompt-size one.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No submission, no relief request, no customer contact — non-configurable. This kit produces a raise band, a governing clock, an accrued day count and a blocker code for a programme analyst to read. It never submits a deliverable, never files an extension request and never sends anything to a government point of contact, and there is no setting that makes it. Separately: the carried state is written by code from the parsed record and the previous state; the model's four answers are scored and are never written back into the history the next run is judged against.
Everywhere and nowhere — it is a property of what is ABSENT. The only writers in the kit are evals/run.py and evals/paraphrase.py (results/*.json), tools/build_corpus.py and tools/paraphrase.py (data/), and src/state.save (data/state.json). src/state.advance() -> src/register.step() is the only thing that touches the carried state, and evals/run.py calls it after every run including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit submits a deliverable, requests relief, or contacts anybody
0 code paths. There is no HTTP client in the kit except src/adapters/__init__.py, which posts to the configured completion endpoint and nowhere else, and no SMTP, no webhook and no file write outside data/ and results/. ⚠︎ This is an absence argued from reading the kit, not a check that runs: no gate in this kit greps for it.
A wrong answer does not propagate into the next run's state
Guaranteed by the code path: src/register.step() takes the parsed record and the previous state and never reads a model reply. evals/run.py calls it with carried, the state BEFORE this run, in every branch including the exception branch. No model arm has been scored, so there is no run consistent-with-it to cite yet.
A run whose CALL failed still advances the state
evals/run.py's exception branch calls register.step() before continuing, so one transport error cannot turn into three scored failures further down the chain. Exercised for real on 2026-08-23: 120 of 120 calls failed with 402 and every chain still advanced through all four runs.
The one section no field maps to never leaves the machine
0 of 120 records send it, in both directions of the test — see security.trials.
The limitWhat a guardrail is not
It is not a submission tool. It has no portal client, no form, and no upload.
It is not an approval. A band of ESCALATE is a request for a person's attention, not a decision that anything is contractually late — that is a determination a contracting officer makes.
It is not a source of truth about dates. Every clock it uses is a field somebody else maintains; the atlas row this kit was built from names that directly, and the kit never invents a date.
It is not a compliance record. Nothing here is retained, signed or timestamped in a way anybody could rely on later; results/ is evidence about the kit, not about the programme.
WatchedWhat is watched, and why that one
12runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 85 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
14 measured by the latest run71 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The four answered fields, per deliverable-run, exact match against the computed answer key
alarm
band_accuracy_pct; missed_watchlist_pct; false_alarm_rate_pct; unparsed_replies — alarm on unparsed_replies above 0. A row that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Every free arm had 0.
raise-direction
Missed raises and false alarms, counted apart
alarm
missed_watchlist_pct; false_alarm_rate_pct — alarm on missed_watchlist_pct above 0. A deliverable that is past its contract date and not on the list is the failure this kit exists to prevent.
memory-dependent-subset
Band accuracy on the rows a single reading cannot answer
alarm
standing_band_pct — alarm on standing_band_pct falling while plain band accuracy holds. That pattern is what a lost or stale state file looks like, and it is invisible in the headline.
off-page-authority
Band accuracy where the record prints a revised due date
alarm
authority_band_pct — alarm on authority_band_pct falling below the plain band accuracy. It means the reader has started trusting the printed date, which is the expensive direction.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
120
different corpus — nothing is comparable
corpus.bytes
543,190
deliverable records edited — the count held, the bytes did not
split.count
30
the deliverables count moved — a different set was scored
split.size_p50
4
the median size of one deliverable moved
split.size_p95
4
the 95th-percentile size of one deliverable moved
dataset.rows
120
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
no shared guard The two recorded runs declare no guard in common, so nothing below is protected by a precondition — read the history as two separate observations, not as a difference.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
missed raises
0 — the free rules engine is at 0 of 60, and so is the pattern floor
60 deliverable-runs that had to be raised
Free code reaches 0 without a model, so anything above 0 is a regression against something that costs nothing. The spreadsheet floor is at 20.00 pct and that is what the alternative looks like.
false alarms
0 — every free arm reports 0 of 26 quiet rows
26 quiet deliverable-runs
Alert fatigue is the slow failure and it is invisible in an accuracy figure. The stub — which ignores both the memory and the prose — sits at 53.85 pct, which is what an inattentive reader looks like.
band accuracy
0 — saturated at 100.00 pct on the free rules engine, and that is a statement about the corpus
120 deliverable-runs, 94 of them on the watchlist
A grader free code aces has stopped discriminating. The honest reading is that this corpus is within reach of arithmetic, not that nothing could be better. Plain band accuracy is also flattered by the 26 quiet rows, which is why the watchlist figure is banded beside it.
the rows that need memory
0 with memory, and a hard floor of 0.00 pct and 0.00 pct without it — measured, not assumed
28 standing escalations and 30 rows with a non-zero accrual
These two falling while plain band accuracy holds is what a lost or stale state file looks like, and it is invisible in the headline. A missed run alone takes them to 63.16 pct and 73.68 pct with the reader unchanged.
the rows where the page and the truth disagree
0 with the prose read, 33.33 pct without it — the spreadsheet floor's own figure
24 deliverable-runs printing a revised due date
Half of these have a modification behind them and half do not, and the structured page is the same shape either way. Falling here means the reader has started trusting the printed date, which is the expensive direction.
the two smallest authorities
0 — the scored run is at 100.00 pct on both, and so is the rules floor on the recovery half
7 date-moved cells and 6 recovery cells, out of 120 deliverable-runs
These are the two smallest populations the kit bands, which is exactly why they need watching separately: one date-moved cell is 14.29 points and one recovery cell is 16.67, so a single miss is a double-digit fall and averaging them into band accuracy hides it entirely. The ladder and notes floors reach 42.86 pct on the date-moved half and the stateless arm falls to the same 42.86, so this is where carried state is doing work; the recovery half is at 100.00 pct on every free arm and is therefore a statement about the corpus rather than about ability. The stub reports 0.00 on both.
the governing clock and the blocker
0 — both saturated at 100 on the two floors that read the prose; 90.00 pct and 26.67 pct on the one that does not
120 deliverable-runs
Which date governs decides everything downstream of it; the blocker decides WHO gets chased, and a government-side dependency read as an internal one sends the analyst to the wrong person.
the day count
0 — arithmetic over a carried scalar, 100.00 pct on the free rules engine
120 deliverable-runs, 30 of which carry a non-zero count
90 of 120 rows carry a zero, so this headline flatters any reader that guesses zero — the memoryless floors score 75.00 pct here and 0.00 pct on the rows that are actually accruing. Never read it without its companion.
the same answer twice
1 of 10 rows moved across 3 asks each
10 rows x 3 asks = 30 asks, 8.3 pct of the corpus
⚑ THIS IS THE ONE BAND WHERE THE MODEL IS MEASURABLY WORSE THAN FREE CODE, and it is not an accuracy figure. The free rules engine is deterministic: asked a row a hundred times it answers identically a hundred times. The model moved on 1 of 10 rows — and the row it moved on is the same row the scored run got wrong, so this kit's only missed raise is a coin toss rather than a misreading. A watch that raises a deliverable one Monday and drops it the next teaches its reader to distrust the list, which is a slower and worse failure than being consistently wrong.
latency
not yet known as a SPREAD — the repeat probe re-asked 10 rows 3 times but records no per-ask latency, so there is still nothing to difference
120 scored rows on r001-cdrl-watch
p50 8318 ms and p95 18966 ms over 120 rows. A band here would be the spread between two runs of the SAME configuration, and the repeat probe was built to answer whether the ANSWER moves rather than how long it takes — so this stays declared-unknown rather than being filled from one run's own quantiles, which would be a distribution across rows dressed up as a distribution across runs.
tokens
not yet known as a SPREAD — one scored run, and the repeat probe re-asks a 10-row subset rather than the corpus
120 scored rows on r001-cdrl-watch
Input is near-constant by construction — every record is between 4434 and 4758 bytes and the carried state is six scalars — and measured at 1535 tokens a row. Output is the bill and ran 203 to 4330 tokens across the scored run, a 21.3-fold spread, so a narrow band would be dishonest even once there is a second run to build one from.
HistoryRun history
12 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 4 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
ablation · with the model in the path — 2 runs. Columns here are only ever compared with each other.
Metric
a001-cdrl-watch-noteblind 2026-08-23
a002-cdrl-watch-paraphrase 2026-08-23
regex status accuracy, %
—
50.0
regex status correct
—
12
false alarm rate, %
44.44
—
input tokens, whole run
—
37657
output tokens, whole run
—
42864
status accuracy, %
—
100.0
status correct
—
24
not a time series No two of these 2 runs measured the same system — they differ on documents, failures, free, max_tokens, rules_stripped, snapshots, thinking — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 8 runs. Columns here are only ever compared with each other.
Metric
b000-cdrl-watch-ladder 2026-08-23
b001-cdrl-watch-notes 2026-08-23
b002-cdrl-watch-missedrun 2026-08-23
b003-cdrl-watch-rules 2026-08-23
c000-cdrl-watch-calibration 2026-08-23
p002-cdrl-watch-repeat 2026-08-23
r001-cdrl-watch 2026-08-23
s001-cdrl-watch-stateless 2026-08-23
authority band, %
33.33
50.00
100.00
100.00
100.00
—
100.00
50.00
band accuracy, %
70.00
73.33
92.22
100.00
100.00
—
99.17
73.33
band on watchlist, %
61.70
65.96
90.41
100.00
100.00
—
98.94
65.96
blocker cause accuracy, %
26.67
100.00
100.00
100.00
100.00
—
98.33
100.00
clock accuracy, %
90.0
100.0
100.0
100.0
100.0
—
100.0
100.0
date moved band, %
42.86
42.86
100.00
100.00
100.00
—
100.00
42.86
days at risk accuracy, %
75.00
75.00
94.44
100.00
100.00
—
100.00
75.00
days at risk when accruing, %
0.00
0.00
73.68
100.00
100.00
—
100.00
0.00
documents that moved
—
—
—
—
—
1
—
—
false alarm rate, %
0.0
0.0
0.0
0.0
—
—
0.0
0.0
input tokens, whole run
—
—
—
—
12399
46857
184232
178376
model latency p50 ms
0.00
0.00
0.00
0.00
11011.00
—
8318.00
6729.00
model latency p95 ms
0.00
0.00
0.00
0.00
20919.00
—
18966.00
12633.00
missed watchlist, %
20.00
0.00
0.00
0.00
0.00
—
1.67
0.00
output tokens, whole run
—
—
—
—
9879
31688
128001
98245
recovery band, %
100.0
100.0
100.0
100.0
100.0
—
100.0
100.0
standing band, %
0.00
0.00
63.16
100.00
100.00
—
96.43
0.00
not a time series No two of these 8 runs measured the same system — they differ on accruing_cells, authority_cells, date_moved_cells, documents, max_tokens, provider, recovery_cells, rows_scored, standing_cells, stateless, thinking, times, watchlist_cells, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · with the model in the path — 1 run. Columns here are only ever compared with each other.
Metric
b004-cdrl-watch-paraphrase-free 2026-08-23
regex status accuracy, %
50.0
regex status correct
12
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 2 chips that all say so.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-cdrl-watch-stub 2026-08-23
authority band, %
16.67
band accuracy, %
54.17
band on watchlist, %
56.38
blocker cause accuracy, %
11.67
clock accuracy, %
70.0
date moved band, %
0.0
days at risk accuracy, %
75.0
days at risk when accruing, %
0.0
false alarm rate, %
53.85
model latency p50 ms
0.00
model latency p95 ms
0.00
missed watchlist, %
0.0
recovery band, %
0.0
standing band, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 14 chips that all say so.
DeviationsWhat deviated
0 breaches across 12 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the carried state
band 73.33 pct -> 99.17 pct, standing escalations 0.00 pct -> 96.43 pct, accruing day counts 0.00 pct -> 100.00 pct, band on the relief rows 50.00 pct -> 100.00 pct
measured
r001-cdrl-watch against s001-cdrl-watch-stateless. Same corpus, same model, same grader, same 120 rows; the prompts differ in exactly one block and evals/check_labels.py asserts they are byte-identical everywhere else. ⚑ AND THE STATELESS ARM LANDS ON THE MEMORYLESS RULES ENGINE EXACTLY: 480 of 480 graded cells identical to b001-cdrl-watch-notes, answer for answer. Removing the memory does not degrade the model towards free code — it turns it INTO free code.
whether the prompt carries the Program Notes
on the W3 run only: governing clock 86.67 pct -> 100.00 pct, blocker 40.00 pct -> 100.00 pct, band on the relief rows 50.00 pct -> 100.00 pct, false alarms 44.44 pct -> 0.00 pct
measured
a001-cdrl-watch-noteblind against the same W3 rows of r001-cdrl-watch. The prose is the ONLY place the record says whether a printed revised due date is carried by an executed modification, and withholding it breaks exactly the fields that depend on reading: the clock, the blocker, and the relief rows' bands. It also introduces false alarms where the full prompt has none. The notes are load-bearing, and this is the run that proves it rather than asserting it.
whether the analyst's prose is worded the way the pattern expects
on the 24 relief rows: free pattern band 100.00 pct -> 50.00 pct and blocker 100.00 pct -> 0.00 pct, while the model holds at 100.00 pct and 100.00 pct
measured
a002-cdrl-watch-paraphrase. The same 24 rows, the same facts, the same four real modifications and four unreal ones — only the sentences change. ⚑ THIS IS THE ONLY EDGE ON THIS PAGE WHERE THE MODEL BEATS FREE CODE, and it beats it completely. ⚠︎ AND IT IS AN UPPER BOUND: the paraphrases and the pattern share an author, and an author writing paraphrases knows which words to avoid.
whether a scheduled run actually happens
standing escalations 100.00 pct -> 63.16 pct, accruing day counts 100.00 pct -> 73.68 pct, band 100.00 pct -> 92.22 pct; 16 raises arrive a week late and 35 days of accrual go uncounted
measured
b002-cdrl-watch-missedrun against b003-cdrl-watch-rules — the identical free rules engine on the identical corpus with the 2026-07-20 run never made. THE READER IS UNCHANGED between these two arms; only the clock differs, which is why this edge is free to measure and why no model arm was paid for on the reduced schedule.
whether the reader has any memory at all (free arms)
b003-cdrl-watch-rules against b001-cdrl-watch-notes. Same code, same regular expression, same corpus; the only difference is whether register.step is given the previous run's six scalars. Both are free, so this is the size of the memory measured without a provider in the picture at all.
whether the reader reads the prose at all (free arms)
b001-cdrl-watch-notes against b000-cdrl-watch-ladder. The spreadsheet floor trusts the printed revised due date and misses 12 of 60 raises for it; adding a regular expression over the prose takes that to zero. This is the cheapest improvement available on this kit and it costs nothing.
the run interval, CADENCE_DAYS
the blind window is CADENCE_DAYS minus the 5-day raise gate — 2 days at the shipped 7, 9 days at 14, and zero at 5. 5 of 30 deliverables are never warned in time at the shipped interval. The bill scales exactly with it.
measured
src/cadence.blind_window() computes it and src/cadence.never_warned() counts the deliverables it costs, both replayed over the committed corpus in b002-cdrl-watch-missedrun. Arithmetic, not a model figure — no reader of any kind recovers a due date that arrived between two runs.
MAX_TOKENS in src/monitor.py
unparsed replies, which are scored as misses in all four fields. Nothing about the bill — a ceiling is not a cost, because the provider bills tokens produced and not tokens allowed.
reasoning
Argued from the calibration rather than measured: c000-cdrl-watch-calibration set the ceiling at roughly three times the largest reply it saw, and no run since has come within 3670 tokens of it (largest reply on the scored run: 4330 of 8000). No arm was fired at a LOWER ceiling to find where replies start being lost, so the safety margin is reasoned, not measured.
the shape of the carried state in src/state.py
the cost curve. Six scalars cost the same on run 400 as on run 4; a transcript makes input climb with the number of runs, and feeding the MODEL's own answer forward would let one wrong week compound into every week after.
reasoning
Argued from the code path: src/register.step reads the parsed record and the previous state and never a reply, and evals/run.py calls it with the pre-run state in every branch including the exception branch. No arm was fired with a compounding state, because building one to measure it would mean shipping the defect.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
missed raises
any deliverable past its governing date reported NONE or WATCH. This is the failure the kit exists to prevent.
false alarms
a quiet deliverable disturbed. Two in a week and the raise list stops being read.
band accuracy
either figure below 100. Watch the watchlist one first — it moves earlier.
the rows that need memory
either figure falling by more than a row's worth while band accuracy is flat. Check data/state.json before you check anything about the model.
the rows where the page and the truth disagree
below plain band accuracy. On a 24-row denominator one miss is 4.17 pct, so read the count and not the rate.
the two smallest authorities
either below 100. Read the date-moved figure first — it is the one the free floors and the stateless arm both fail, so it moves before anything else does.
the governing clock and the blocker
either below 100. The blocker moves first, because it is the field with seven values rather than five.
the day count
below 100. It is arithmetic; anything less means the memory or the interval is wrong.
the same answer twice
any row moving at all. A zero here would mean 'nothing moved on THESE rows, this many times' — never 'the kit is stable', which is a claim about a population this probe never sampled.
latency
nothing yet. A weekly batch has hours, not seconds — but a p50 that doubles is the first sign of a provider under load, and the tail is what makes a Monday run not finish.
tokens
input moving at all. It means the record shape changed or a withheld section is being sent, which is a disclosure event before it is a cost one.
NextThe three you would add first
An identity check on the register exportEvery carried scalar is keyed on the deliverable id. A renumbered line arrives with no history, is reported on its ladder band alone, and every previous-run rule silently stops applying to it — no error, no empty cell, just a deliverable that quietly stops escalating. It is the first thing to build and it is not built here; Data.breaks_on names it and nothing measures it.
A liveness check on the watch itself, running outside the watchA run that did not happen cannot report that it did not happen, and a register with nothing to raise looks identical to a watch that never fired. Measured: one skipped Monday cost 16 late raises, 35 days of accrual and 36.84 pct of standing-escalation accuracy with the reader unchanged (b002-cdrl-watch-missedrun). This kit's own 120-row run was refused by a provider account earlier the same day and produced a complete, empty, perfectly-shaped result file.
A second pair of eyes on any row whose band came from the prose rather than the datesThat is the 24 rows printing a revised due date, and it is the only place a reader can be wrong about something the structured page cannot check. It is also the only place a model earns anything here: reworded, free code falls to 50.00 pct on those rows and the model holds at 100.00 pct.
A grep in the pre-flight for any write path this kit promises not to haveThe 'never submits, never requests relief, never contacts anybody' guarantee is argued from reading the kit and is NOT enforced by a check — a sibling kit greps its own source for the function names it forbids and this one does not. security.could_not_verify says so rather than implying the absence is proven.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
⚑ THE GUARDRAIL CADENCE AND THE WATCH CADENCE ARE THE SAME CLOCK, and that is unusual enough to say plainly. Every band above is read at the end of each weekly run, because there is a run every week and scoring costs nothing.
Three of them are read as a TREND rather than a level, because their level is fine right up until the moment the design has failed: standing_band_pct falling while plain band accuracy holds is a lost state file; input_tokens_total jumping is a record shape or a disclosure change; days_at_risk_when_accruing_pct falling is an interval that has quietly changed length.
And one thing is watched that is not a metric at all: WHETHER THE RUN HAPPENED. A missed Monday costs 16 late raises, 35 days of accrual and 36.84 pct of standing-escalation accuracy on this corpus, with the reader unchanged — and a watch that did not run produces no numbers, so every band above is silent about it.
What this cannot tell you
The absence guarantee is argued from reading the kit, not enforced by a check. A sibling kit greps its own source for release and write-off function names; this one does not, and adding that grep is a real improvement rather than a formality.
Every band above is set against the FREE FLOOR's numbers, because no scored model arm exists. The green thresholds are therefore what free code already achieves, which is a defensible place to put them and not an empirical distribution of anything.
Whether the ripple table is complete. It was written by reading the kit, and a change nobody has made yet is exactly the kind that is missing from it.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all — argparse, datetime, json, os, random, re, sys, textwrap, urllib, http.client, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior runs)
those carry a TRANSCRIPT; this carries six SCALARS written by arithmetic. The whole reason the cost is flat in history length is that nothing here remembers what was said, only what was counted — and a memory layer would give back both the growth and the compounding this design exists to avoid.
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install.
the clock
src/cadence.py + evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
⚑ THIS IS THE SEAM WHERE A FRAMEWORK GENUINELY EARNS ITS PLACE, and this kit says so rather than pretending otherwise. A ThreadPoolExecutor over 30 chains is the right size for an EVAL and the wrong size for a watch that runs on a programme calendar: retries across days, a run that did not happen, a register export that arrived late, and the alerting that tells you the Monday run failed are all outside this kit. src/cadence.py measures what those cost; it does not provide them.
there is nothing to grade with a model here. Four closed-set values compared for equality against a key the arithmetic computed is a dictionary lookup, and a framework would add a dependency, a config file and a second thing to be wrong.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
register export -> segment -> select (one section subtracted) -> prompt (record + carried state + question) -> one completion call -> parse -> score against the computed key; and, in parallel and for nothing, the same record through the free rules engine. The carried state is written by src/register.step from the PARSED record, never from the reply, and hands the next scheduled run six scalars.
The other sideWhat a framework costs you
No dependency to install, so the fork test is git clone and python3 -m tools.build_corpus.
No scheduler, so running this weekly is somebody else's cron entry — and the kit measures what it costs when that entry does not fire.
No eval framework, so the scorer is 200 lines you can read and the graders are four restrictions of one comparison.
No memory layer, so the state is six scalars a programme office can be shown and can check against its own register.
What we could NOT verify
Whether a scheduling framework would change any published figure. It would not change the answers; it would change how many runs actually happen, which is the thing src/cadence.py prices and nothing here measures in the wild.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on b003-cdrl-watch-rules on the free floor, with the carried state and all three rules, 2026-08-23. This kit records telemetry measured per run — 2 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
0 ms
not yet known as a SPREAD — the repeat probe re-asked 10 rows 3 times but records no per-ask latency, so there is still nothing to difference
nothing yet. A weekly batch has hours, not seconds — but a p50 that doubles is the first sign of a provider under load, and the tail is what makes a Monday run not finish.
Model, p95
0 ms
not yet known as a SPREAD — the repeat probe re-asked 10 rows 3 times but records no per-ask latency, so there is still nothing to difference
nothing yet. A weekly batch has hours, not seconds — but a p50 that doubles is the first sign of a provider under load, and the tail is what makes a Monday run not finish.
Input tokens
—
not measured on this run
Output tokens
—
not measured on this run
No movement column. Not one of the 7 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-cdrl-watch-calibration11,011 ms
r001-cdrl-watch8,318 ms
s001-cdrl-watch-stateless6,729 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
5 runs not plotted. b000-cdrl-watch-ladder, b001-cdrl-watch-notes, b002-cdrl-watch-missedrun, b003-cdrl-watch-rules, p002-cdrl-watch-repeat recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 12 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
deliverable records
data/corpus/-W.txt -- 120 files, 543190 bytes, generated once from a fixed seed by tools/build_corpus.py
seven of the eight sections go to the provider in the prompt; Government Points of Contact -- named contracting officers, their office, e-mail and telephone -- never does, by src/select.NEVER_SENT
the answer key
data/gold.jsonl -- 120 rows, the output of src/register.replay over each deliverable's trajectory, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the carried state
data/state.json in a deployment (src/state.py, written atomically). IN AN EVAL IT IS SCOPED TO THE RUN AND NEVER TOUCHES DISK -- evals/run.py builds a fresh in-memory store per deliverable, because a run that started from the previous run's memory could not be re-run or compared with its own control
one short paragraph per call, produced by state.describe -- six scalar fields, never a prior record and never a prior reply
the schedule
src/cadence.py names the interval and what it costs to miss one; the trigger itself is OUTSIDE the kit -- cron, a scheduled task, or the job that produces the register export. There is no daemon here and nothing listens on a socket
never -- it is four dates in a list
the recorded runs
results/eval-*.json and results/paraphrase-*.json -- the three free floors, the missed-run schedule, the wiring stub, the ceiling calibration and the free half of the paraphrase probe, all committed
never -- they are read by the page and re-scored for $0.00
the key
.env or the shared repo-root .env -- never committed, 0600
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py). src/app.py redacts it out of any error it renders
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The key is read from <repo>/.env, then the kit's own .env, then the real environment, and it is never written to a result file, a log line or a screenshot. src/app.py strips the key and the base URL out of any provider error before returning it to the browser. The repository has never held a credential: both .env files are gitignored from the first commit and .env.example carries empty values.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
Every Monday at 06:00 local, seven days apart, fired by a scheduler OUTSIDE this kit. One run owns every gate crossing in the interval that ended when it started -- not a slice of a stream, a slice of TIME, because the register is re-read whole every time. src/cadence.owned_by returns exactly which deliverables crossed which gate inside a given run's interval, which is the only defensible answer to 'why am I seeing this today and not last week'.
Skipping one run costs 16 raises arriving a week late, 35 days of accrued risk uncounted, 7 of 90 reported bands wrong, and standing-escalation accuracy falling from 100.00 pct to 63.16 pct. Separately the interval itself leaves a 2-day blind window against a 5-day raise gate, and 5 of 30 deliverables are never warned in time on ANY schedule. (b002-cdrl-watch-missedrun against b003-cdrl-watch-rules, both free)
The interval must be shorter than the shortest gate on the ladder or there is a band in every cycle where a due date arrives unseen. At 7 days against a 5-day raise gate the band is 2 days wide; at 14 it is 9. This is the ceiling and it is arithmetic, not a quality figure.
Lengthen the interval and every accrual figure on this page changes meaning, because the accrual counts INTERVALS rather than runs -- src/register.step takes the elapsed days from the carried run date. Shorten it and the bill scales exactly with it. Either way the missed-run arm has to be re-fired, because its numbers are differences against a schedule that no longer exists.
state
Six scalar fields per deliverable -- the band reported last run, the ladder band observed last run, the date it was judged against, its submission status, the accrued days at risk, and the date the last run happened -- written by src/register.step from the PARSED record and never from the model's reply, and rendered by src/state.describe into one English paragraph. That paragraph is the entire route from one scheduled run to the next, and it is the only thing the stateless control differs by.
Removing it costs every previous-run rule: the memoryless floor reports 0.00 pct of the 28 standing escalations and 0.00 pct of the 30 accruing day counts, against 100.00 pct and 100.00 pct with it. Band accuracy falls from 100.00 pct to 73.33 pct on the same 120 rows. ⚑ AND THE STATELESS MODEL ARM LANDS EXACTLY ON THE MEMORYLESS RULES ENGINE: 480 of 480 graded cells identical, not merely equal in aggregate. The carried state is the only thing on this kit that separates any two readers at all. (b003-cdrl-watch-rules against b001-cdrl-watch-notes, both free; and s001-cdrl-watch-stateless against b001-cdrl-watch-notes at cell level)
Six scalars, so run 400 costs what run 4 costs. The ceiling is not size, it is IDENTITY: every field is keyed on the deliverable id, so a register that renumbers a line hands the next run a deliverable with no history and silently stops applying every previous-run rule to it. Not measured here and not detected.
Carry a transcript instead of scalars and the kit stops being flat in history length -- input cost starts climbing with the number of runs, which is the property that separates this from a conversation. Feed the MODEL's own answer forward instead of the arithmetic's and one wrong week compounds into every week after it, and the eval can no longer tell a single miss from a run that went wrong once and never recovered.
model
One completion call per deliverable per scheduled run, over raw HTTP to whichever OpenAI-compatible or Anthropic endpoint .env names. One system message, one user message, a 8000-token ceiling, provider defaults for everything else. The prompt is the selected sections, the carried-state paragraph, and a fixed question with a fixed JSON shape.
Ceiling set from 8 calibration calls: output ran 584 to 2607 tokens, 91.51 pct of it provider-side reasoning, on a corpus whose records are all between 4434 and 4758 bytes. MAX_TOKENS is roughly three times the largest reply seen. (c000-cdrl-watch-calibration)
⚠︎ THE MODEL IS THE ONE DECISION ON THIS LADDER THIS KIT HAS NOT MEASURED PROPERLY. The scored 120-row run was refused by the shared provider account with 402 Payment Required -- Insufficient Balance -- and nothing was billed. 8 calls produced a completion, all of them calibration, on 2 of the 30 deliverables. Every model figure on this page rests on that denominator and the page says so wherever it appears.
A model that cannot return a whole number for the day count invalidates one of the four fields outright; a model whose reasoning exceeds the ceiling returns nothing at all, which is scored as a miss in all four and arrives looking like a quality failure. And a provider outage on a Monday invalidates the RUN, not the row -- which is exactly what happened here.
labels
120 deliverable-runs -- 30 fictional contract deliverables re-read on four weekly runs -- with four graded fields each. The key is COMPUTED by src/register.replay from each deliverable's trajectory and the one off-page fact, never hand-authored, and evals/check_labels.py re-derives every label from the records themselves before any run may spend.
26 quiet rows and 94 on the watchlist; 60 that must be raised; 24 rows printing a revised due date of which 12 have a modification behind them; 28 standing escalations, 6 recovery holds, 7 moved dates; 30 rows with a non-zero accrual. (data/corpus-stats.json and b003-cdrl-watch-rules)
It stops scoring where the corpus stops being anybody's register. Every record here is generated to one shape, so nothing measures a real export's variance, and the analyst's prose was written by the same author as the pattern that reads it. Four runs is long enough to build a standing escalation and not long enough to test an accrual that has been running for a year.
Change a gate (WATCH_AT_DAYS, RAISE_AT_DAYS) or the branch order in src/register.step and every gold label changes -- the key is that module's output. Re-run tools/build_corpus.py and evals/check_labels.py before comparing anything to a figure on this page. Change the note templates and both prose-reading floors stop being measurements of anything, because the pattern was written against them.
corpus refresh
Nothing is cached and nothing is pre-computed, so a refreshed register needs no rebuild: drop new record files in data/corpus/ and the next run reads them. What DOES need care is the carried state, which is keyed on the deliverable id and lives in data/state.json -- it survives a refresh only if the ids do.
Corpus regeneration from the seed is under a second for all 120 records and produces them byte for byte; there is no index build to time, which is why Data.index.built_in_seconds is 0.0. (tools/build_corpus.py, re-run on the committed checkout)
A refresh that renumbers deliverables. The state has no way to recognise a line it has seen before under a new id, so the whole history is silently discarded for that line and every previous-run rule stops applying.
A refresh that changes the record LAYOUT invalidates src/monitor.reading_of, which parses fixed lines. It would fail loudly on the run date and the contract date and silently on the revised date, because an absent revised date is a legitimate value.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
The raise list is shorter this week than last, with no acceptances to explain it
A reader has started trusting printed revised due dates. On this corpus that is worth 12 of 60 raises -- every one of them a deliverable past its contract date with no modification behind the date that replaced it.
Score the 24 rows that print a revised due date on their own. authority_band_pct falling below plain band accuracy is the signature. (eval-b000-cdrl-watch-ladder.json)
Standing escalations collapse to zero while plain band accuracy holds
The state file is missing, stale, or keyed on something that changed. Every deliverable looks like a first sighting, so no previous-run rule can fire.
Read data/state.json and count the entries against the open register. A memoryless reader scores 0.00 pct on the 28 standing escalations and looks perfectly healthy on everything else. (eval-b001-cdrl-watch-notes.json)
Day counts that are all multiples of the interval but too small
A run was missed. The accrual counts intervals correctly across the gap, so the numbers stay plausible -- they are simply short by one interval for every deliverable that was past its gate at both ends of the skipped week.
Compare the run dates in the state file against the schedule. One skipped Monday cost 35 days of accrual across 30 deliverables here. (eval-b002-cdrl-watch-missedrun.json)
No machine symptom — this failure leaves no trace in any output.
Liveness monitoring on the WATCH, outside the watch: a run that did not happen cannot report that it did not happen. Nothing in this kit provides it and guardrails.add_first names it as the second thing to build.
⚠︎ THE MODEL ARM. The 120-row scored run, its stateless control, the note-blind ablation, the repeat probe and the paid half of the paraphrase probe are all wired, all costed and none of them has been fired: the shared provider account returned 402 Payment Required -- Insufficient Balance -- and nothing was billed. 8 calibration calls over 2 of the 30 deliverables are the only model evidence this kit holds.
CONCURRENCY. EVAL_WORKERS=12 was used for the calibration and nothing measured where the provider starts rate-limiting or where the chains stop scaling. The chains are independent by construction; the provider is not.
GPU AND HOSTING. Nothing here runs a model locally, so no sizing figure exists and none is guessed. src/adapters speaks to a local server as readily as to a hosted one and this kit has never pointed it at one.
PROVIDER-SIDE RETENTION. What the configured provider keeps, for how long, and whether it trains on it is a property of the account and the contract. Nothing here measures it.
THE ENGLISH-VERSUS-JSON CARRIED STATE. src/state.describe renders six scalars as prose and argues for it; the experiment that would settle it is one more 120-row run and has not been paid for.
REAL REGISTER VARIANCE. Every record is generated to one shape between 4434 and 4758 bytes, and the analyst's prose was written by the same author as the pattern that reads it. Nothing here measures a real export.
RUN-TO-RUN VARIANCE ON THE MODEL PATH. The free arms are deterministic and reproduce exactly; no model arm has been run twice, so every model figure here is one observation.
The corpus licence, from the Data lens: MIT — generated in this repository by tools/build_corpus.py from a fixed seed, so it reproduces byte for byte and belongs to nobody. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The four answered fields, per deliverable-run, exact match against the computed answer key
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineThe four answered fields, per deliverable-run, exact match against the computed answer key
For each of the 120 deliverable-runs and each of the four answered fields, did the reply equal the computed key? Band, clock and blocker are compared exactly after trimming and upper-casing; the day count is compared as a whole number, and a reply that is not a whole number is a MISS rather than a zero — 0 is a meaningful answer here (it is what a deliverable inside its gate gets) and a parse failure must never be scored as one.
$0.00per 1,000 deliverable records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function every floor, the stub and the missed-run schedule are scored through.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The deliverable and the run
CDRL-0018-W3
What the structured record says
Contract due 2026-07-15; Revised due 2026-08-17 (proposed); Status IN PROGRESS; no disposition.
What the analyst wrote
Blocked: the government-furnished interface data we need for section 3 has not been delivered.
Requested from the programme office and still not received. We have asked for relief to
2026-08-17. The request went in on 2026-07-08 and is still with the contracting officer; nothing
has been executed, so the contract date still stands.
What the previous run left behind
The previous run of this watch was on 2026-07-13. This run is 7 days after it. On that run this deliverable was reported RAISE. The dates alone already put it past its raise gate then. The date it was judged against was 2026-07-15. Its submission status was IN PROGRESS. 0 days at risk had been accumulated for it by the end of that run.
{'raise_band': 'NONE', 'governing_clock': 'REVISED_DUE', 'days_at_risk': 0, 'blocker_cause': 'NO_BLOCKER', 'rationale': 'ladder against the revised due; no history available'}
The free rules engine's answer
{'raise_band': 'ESCALATE', 'governing_clock': 'CONTRACT_DUE', 'days_at_risk': 7, 'blocker_cause': 'AWAITING_GOVERNMENT_DATA', 'rationale': 'ladder against the contract due; carried state applied'}
Why this row
It is the whole argument on one line. The structured page says a revised due date of 2026-08-17 — 28 days of clear air — and prints it exactly as the four deliverables whose relief is real. Only the prose says no modification has been issued. Read the page and this deliverable is quiet; read the prose and it is five days past its contract date, was already raised last Monday, and nothing has moved since.
Grader
Verdict
Why
The four answered fields, per deliverable-run, exact match against the computed answer key
hit
All four fields compared against the computed key: blocker_cause=AWAITING_GOVERNMENT_DATA, days_at_risk=7, governing_clock=CONTRACT_DUE, raise_band=ESCALATE.
Missed raises and false alarms, counted apart
hit
The key raises this row (ESCALATE). A reader that reported NONE or WATCH here would be under-raising a deliverable five days past its contract date.
Band accuracy on the rows a single reading cannot answer
hit
This row's answer is a STANDING_ESCALATION: it is only correct because the previous run already raised the item and the submission status has not moved. Both memoryless floors get it wrong.
Band accuracy where the record prints a revised due date
hit
The record prints a revised due date of 2026-08-17 marked (proposed) and only the prose says no modification carries it. The spreadsheet floor reports NONE here.
The formulaWhat it computes
accuracy = hits / 120 per field. A row whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the free rules engine, with memory — THE FLOOR THAT WINS
100.0% band accuracy · 3 more measured on this row
the same, with the memory removed
73.3% band accuracy · 3 more measured on this row
dates and structured fields only — the spreadsheet
70.0% band accuracy · 3 more measured on this row
the winning floor, on a schedule that missed one Monday
92.2% band accuracy · 3 more measured on this row
the harness self-test — state-blind and prose-blind, by design
54.2% band accuracy · 3 more measured on this row
the fast tier, with the carried state
99.2% band accuracy · 3 more measured on this row
the fast tier, memory removed (THE CONTROL)
73.3% band accuracy · 3 more measured on this row
the fast tier, CEILING CALIBRATION ONLY
100.0% band accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/register.replay over each deliverable's trajectory at generation time. This grader IS the reference, so its own TPR and TNR are not measured — scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one thing about it is arguable — the branch ORDER in src/register.step. Recovery is tested before a moved date, so a raised deliverable whose date moves out is reported WATCH under RECOVERY_HOLD rather than under DATE_MOVED. Both report WATCH, so no band moves either way, but the reason code does and a programme office might order the two rules the other way.
Watch these
band_accuracy_pct
missed_watchlist_pct
false_alarm_rate_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A row that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Every free arm had 0.
How tight can the band be? No tuned threshold anywhere on the model path — the reply decides. The ladder's own gates (15 days, 5 days) are declared on every record and are inputs, not tuning. Denominators are printed beside every rate because three of them are small: 24 relief rows, 28 standing escalations, 6 recovery holds.
Cadence: Re-scored on every run of the watch, because scoring is free. In a deployment the key does not exist, so what is watched weekly is the RAISE RATE and the reason-code mix against the previous run: a monitor whose standing-escalation count collapses has usually lost its state file, not improved.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known — the normal state of a real register, where whether a modification exists is decided by somebody opening the contract file. That is why this corpus is generated rather than captured.
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineMissed raises and false alarms, counted apart
Of the 60 deliverable-runs that had to be raised, how many were reported NONE or WATCH (a MISSED RAISE); and of the 26 quiet rows, how many were disturbed (a FALSE ALARM). Two rates over two different denominators, never averaged.
$0.00per 1,000 deliverable records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The deliverable and the run
CDRL-0018-W3
What the structured record says
Contract due 2026-07-15; Revised due 2026-08-17 (proposed); Status IN PROGRESS; no disposition.
What the analyst wrote
Blocked: the government-furnished interface data we need for section 3 has not been delivered.
Requested from the programme office and still not received. We have asked for relief to
2026-08-17. The request went in on 2026-07-08 and is still with the contracting officer; nothing
has been executed, so the contract date still stands.
What the previous run left behind
The previous run of this watch was on 2026-07-13. This run is 7 days after it. On that run this deliverable was reported RAISE. The dates alone already put it past its raise gate then. The date it was judged against was 2026-07-15. Its submission status was IN PROGRESS. 0 days at risk had been accumulated for it by the end of that run.
{'raise_band': 'NONE', 'governing_clock': 'REVISED_DUE', 'days_at_risk': 0, 'blocker_cause': 'NO_BLOCKER', 'rationale': 'ladder against the revised due; no history available'}
The free rules engine's answer
{'raise_band': 'ESCALATE', 'governing_clock': 'CONTRACT_DUE', 'days_at_risk': 7, 'blocker_cause': 'AWAITING_GOVERNMENT_DATA', 'rationale': 'ladder against the contract due; carried state applied'}
Why this row
It is the whole argument on one line. The structured page says a revised due date of 2026-08-17 — 28 days of clear air — and prints it exactly as the four deliverables whose relief is real. Only the prose says no modification has been issued. Read the page and this deliverable is quiet; read the prose and it is five days past its contract date, was already raised last Monday, and nothing has moved since.
Grader
Verdict
Why
The four answered fields, per deliverable-run, exact match against the computed answer key
hit
All four fields compared against the computed key: blocker_cause=AWAITING_GOVERNMENT_DATA, days_at_risk=7, governing_clock=CONTRACT_DUE, raise_band=ESCALATE.
Missed raises and false alarms, counted apart
hit
The key raises this row (ESCALATE). A reader that reported NONE or WATCH here would be under-raising a deliverable five days past its contract date.
Band accuracy on the rows a single reading cannot answer
hit
This row's answer is a STANDING_ESCALATION: it is only correct because the previous run already raised the item and the submission status has not moved. Both memoryless floors get it wrong.
Band accuracy where the record prints a revised due date
hit
The record prints a revised due date of 2026-08-17 marked (proposed) and only the prose says no modification carries it. The spreadsheet floor reports NONE here.
The formulaWhat it computes
missed = raised-in-key but not-raised-in-reply / raised-in-key. false_alarm = quiet-in-key but not-quiet-in-reply / quiet-in-key.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
1.7% value
the free rules engine, with memory
0.0% value
the free floor, no memory
0.0% value
dates and structured fields only
20.0% value
the harness self-test
0.0% value
the fast tier, memory removed (THE CONTROL)
0.0% value
In operationWhat to monitor
Reference standard: The same computed key. This grader is a re-split of the reference grader's band cells, not a second opinion.
These rates are UNKNOWN, on purpose
Its own rates are not measured against anything because it has nothing to be measured against: it re-partitions the reference grader's cells and inherits that grader's correctness exactly.
Watch these
missed_watchlist_pct
false_alarm_rate_pct
Alarm on
missed_watchlist_pct above 0. A deliverable that is past its contract date and not on the list is the failure this kit exists to prevent.
How tight can the band be? No threshold. The partition is on the KEY's band, not on any score.
Cadence: Weekly, with the run. The missed-raise rate is the number a programme office would put on a wall.
The decisionWhen to reach for it
Use it
Always, on a monitor. The two directions cost differently and an average over them is a number nobody can act on.
Do not use it
Never on this kit. It is here precisely because 26 of 120 rows are quiet, so a reader that reported NONE everywhere would score 21.67 pct on plain band accuracy while raising nothing at all.
Band accuracy on the rows a single reading cannot answer
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineBand accuracy on the rows a single reading cannot answer
Band accuracy restricted to the 28 deliverable-runs whose correct answer is a standing escalation — the item was already being raised last run, is still past its gate, and its submission status has not moved. Nothing in the record in front of you says any of that.
$0.00per 1,000 deliverable records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The deliverable and the run
CDRL-0018-W3
What the structured record says
Contract due 2026-07-15; Revised due 2026-08-17 (proposed); Status IN PROGRESS; no disposition.
What the analyst wrote
Blocked: the government-furnished interface data we need for section 3 has not been delivered.
Requested from the programme office and still not received. We have asked for relief to
2026-08-17. The request went in on 2026-07-08 and is still with the contracting officer; nothing
has been executed, so the contract date still stands.
What the previous run left behind
The previous run of this watch was on 2026-07-13. This run is 7 days after it. On that run this deliverable was reported RAISE. The dates alone already put it past its raise gate then. The date it was judged against was 2026-07-15. Its submission status was IN PROGRESS. 0 days at risk had been accumulated for it by the end of that run.
{'raise_band': 'NONE', 'governing_clock': 'REVISED_DUE', 'days_at_risk': 0, 'blocker_cause': 'NO_BLOCKER', 'rationale': 'ladder against the revised due; no history available'}
The free rules engine's answer
{'raise_band': 'ESCALATE', 'governing_clock': 'CONTRACT_DUE', 'days_at_risk': 7, 'blocker_cause': 'AWAITING_GOVERNMENT_DATA', 'rationale': 'ladder against the contract due; carried state applied'}
Why this row
It is the whole argument on one line. The structured page says a revised due date of 2026-08-17 — 28 days of clear air — and prints it exactly as the four deliverables whose relief is real. Only the prose says no modification has been issued. Read the page and this deliverable is quiet; read the prose and it is five days past its contract date, was already raised last Monday, and nothing has moved since.
Grader
Verdict
Why
The four answered fields, per deliverable-run, exact match against the computed answer key
hit
All four fields compared against the computed key: blocker_cause=AWAITING_GOVERNMENT_DATA, days_at_risk=7, governing_clock=CONTRACT_DUE, raise_band=ESCALATE.
Missed raises and false alarms, counted apart
hit
The key raises this row (ESCALATE). A reader that reported NONE or WATCH here would be under-raising a deliverable five days past its contract date.
Band accuracy on the rows a single reading cannot answer
hit
This row's answer is a STANDING_ESCALATION: it is only correct because the previous run already raised the item and the submission status has not moved. Both memoryless floors get it wrong.
Band accuracy where the record prints a revised due date
hit
The record prints a revised due date of 2026-08-17 marked (proposed) and only the prose says no modification carries it. The spreadsheet floor reports NONE here.
The formulaWhat it computes
hits / 28, over the rows whose key reason is STANDING_ESCALATION.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
96.4% value
the free rules engine, with memory
100.0% value
the free floor, no memory
0.0% value
dates and structured fields only
0.0% value
the harness self-test
0.0% value
the fast tier, memory removed (THE CONTROL)
0.0% value
In operationWhat to monitor
Reference standard: The same computed key, restricted by the key's own reason code.
These rates are UNKNOWN, on purpose
Not scored against anything itself — it is a restriction of the reference grader's cells. What IS unknown is whether 28 rows is enough to separate two readers: it is a small denominator and the page prints it beside the rate for that reason.
Watch these
standing_band_pct
Alarm on
standing_band_pct falling while plain band accuracy holds. That pattern is what a lost or stale state file looks like, and it is invisible in the headline.
How tight can the band be? No threshold.
Cadence: Weekly. This is the row a missed run damages first — the missed-run schedule takes it from 100.00 pct to 63.16 pct while plain band accuracy only falls from 100.00 pct to 92.22 pct.
The decisionWhen to reach for it
Use it
On any kit that carries state. It is the only row that prices the memory.
Do not use it
On a kit whose rows are independent, where it would be empty.
Band accuracy where the record prints a revised due date
Catch a government contract's deliverable before it's late
PresenterOpens the private repo. Visible to admins only.
In one lineBand accuracy where the record prints a revised due date
Band accuracy restricted to the 24 deliverable-runs whose Schedule block prints a revised due date marked (proposed). Half are carried by an executed contract modification and half are not; the structured page is the same shape either way and only the analyst's prose separates them. This is the only place on this kit where reading earns anything.
$0.00per 1,000 deliverable records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass. The paraphrase probe (evals/paraphrase.py) re-scores the identical rows with the prose reworded.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The deliverable and the run
CDRL-0018-W3
What the structured record says
Contract due 2026-07-15; Revised due 2026-08-17 (proposed); Status IN PROGRESS; no disposition.
What the analyst wrote
Blocked: the government-furnished interface data we need for section 3 has not been delivered.
Requested from the programme office and still not received. We have asked for relief to
2026-08-17. The request went in on 2026-07-08 and is still with the contracting officer; nothing
has been executed, so the contract date still stands.
What the previous run left behind
The previous run of this watch was on 2026-07-13. This run is 7 days after it. On that run this deliverable was reported RAISE. The dates alone already put it past its raise gate then. The date it was judged against was 2026-07-15. Its submission status was IN PROGRESS. 0 days at risk had been accumulated for it by the end of that run.
{'raise_band': 'NONE', 'governing_clock': 'REVISED_DUE', 'days_at_risk': 0, 'blocker_cause': 'NO_BLOCKER', 'rationale': 'ladder against the revised due; no history available'}
The free rules engine's answer
{'raise_band': 'ESCALATE', 'governing_clock': 'CONTRACT_DUE', 'days_at_risk': 7, 'blocker_cause': 'AWAITING_GOVERNMENT_DATA', 'rationale': 'ladder against the contract due; carried state applied'}
Why this row
It is the whole argument on one line. The structured page says a revised due date of 2026-08-17 — 28 days of clear air — and prints it exactly as the four deliverables whose relief is real. Only the prose says no modification has been issued. Read the page and this deliverable is quiet; read the prose and it is five days past its contract date, was already raised last Monday, and nothing has moved since.
Grader
Verdict
Why
The four answered fields, per deliverable-run, exact match against the computed answer key
hit
All four fields compared against the computed key: blocker_cause=AWAITING_GOVERNMENT_DATA, days_at_risk=7, governing_clock=CONTRACT_DUE, raise_band=ESCALATE.
Missed raises and false alarms, counted apart
hit
The key raises this row (ESCALATE). A reader that reported NONE or WATCH here would be under-raising a deliverable five days past its contract date.
Band accuracy on the rows a single reading cannot answer
hit
This row's answer is a STANDING_ESCALATION: it is only correct because the previous run already raised the item and the submission status has not moved. Both memoryless floors get it wrong.
Band accuracy where the record prints a revised due date
hit
The record prints a revised due date of 2026-08-17 marked (proposed) and only the prose says no modification carries it. The spreadsheet floor reports NONE here.
The formulaWhat it computes
hits / 24, over the rows where a revised due date is printed.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% value
the free rules engine, with memory
100.0% value
the free floor, no memory
50.0% value
dates and structured fields only
33.3% value
the harness self-test
16.7% value
the fast tier, memory removed (THE CONTROL)
50.0% value
In operationWhat to monitor
Reference standard: The same computed key, restricted to rows whose record prints a revised due date.
These rates are UNKNOWN, on purpose
How much of the free pattern's score on these rows is a fact about the WORDING rather than about reading is measured only as an upper bound: b004-cdrl-watch-paraphrase-free reworded the same rows and the pattern fell from 100.00 pct to 50.00 pct, but the paraphrases and the pattern share an author.
Watch these
authority_band_pct
Alarm on
authority_band_pct falling below the plain band accuracy. It means the reader has started trusting the printed date, which is the expensive direction.
How tight can the band be? No threshold. The subset is defined by a printed field, not by a score.
Cadence: Weekly, and re-check the WORDING whenever the team that writes the notes changes. A pattern-based reader's accuracy here is a property of house style, not of the task.
The decisionWhen to reach for it
Use it
Whenever a kit claims a model earns something a rule cannot. Name the rows where that is true and score only those.
Do not use it
Where the claim is about the whole corpus. This is a subset and the page never quotes it as a headline.
A living map of modern AI — kept current every morning