Open orders sit in a queue that shows how late every task already is, not which ones are about to miss their date. This app reads the day's notes on each order and flags the ones heading for trouble days before the deadline arrives.
PresenterOpens the private repo. Visible to admins only.
For the fulfilment coordinatorTelecommunications · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A fulfilment coordinator at a telecom provider, tracking open orders toward their promised dates.
✕Today's manual process
1Open every stalled order in the queue and check how late its tasks already are.
2Read the day's notes on each one, looking for a supplier delay or a rebooked date.
3Decide from memory which orders were flagged yesterday and which ones are new today.
4Miss the early warning and the customer hears about a missed date with the days of notice already gone.
Every order re-read from memory daily
✓With the app
1Every open order is checked against its promised date, every single day.
2The day's notes are read for a supplier delay or a rebooked date, not just task age.
3A status is set clear, watch, or jeopardy, with the reason and the task behind it.
4Warning arrives early days before the date, while there is still time to act.
Every order checked the same way daily
See it work
One real case, read by the app, step by step
A leased-line order's civils delay looks risky, but the app reads the note and holds it at watch, not jeopardy.
Catch a telecom order before it misses its dateReference appBuilt to be shaped to your process
5
1One check-in day day 3 of this order's five-day watch window.
2Today's status the app reports WATCH — not clear, not urgent yet.
3The evidence a note says the civils gang stood down for weather.
4Nothing to escalate the delay sits inside the allowed grace window, so no rule fires.
5Who acts next a coordinator can open the full snapshot behind this call, any time.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A stuck-order queue tells you how late every task is. It does not tell you which orders are going to MISS. Those are two different questions, and the gap between them is where the working days go: an order whose supplier quoted a fifteen-day lead time this morning is already lost, and nothing on its task table is late yet. By the time the aging thresholds catch up, the committed date is days away and the customer has been told the wrong thing twice. The evidence that would have said so on day two is sitting in the activity notes, in a sentence somebody typed at five o'clock, next to four other sentences that mention a number of working days and commit to nothing. A fulfilment coordinator opening every stalled order in the queue each morning, reading the task ages against the jeopardy thresholds, then reading the activity notes to find the supplier quote or rebooked date that means the committed date is already unreachable -- and remembering which of them were raised yesterday.
Audience
Fulfilment and provisioning operations in telecommunications -- order desks, jeopardy managers and the people who build their daily packs -- and anyone deciding whether a language model earns its place beside an aging report they already have. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual order snapshots
The corpus is 240 order snapshots, 0.72 MB (jsonl 1 · txt 240). A real stuck-order queue names an operator's customers, their site addresses, their mobile numbers, a supplier's commercial lead times and the engineer who did not turn up. There is no public corpus of (order snapshot, activity log, adjudicated jeopardy verdict) for the same reason there is no public corpus of production alert histories. Generating it also bought the one thing this kit could not do without: the OUTCOME of every order is known at generation time and appears on no page. You cannot ask 'how many working days before the miss did this pack surface it' of a corpus that does not know which orders missed.
The corpus
The 240 order snapshotsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your order snapshots. That is the whole change — there is no database to migrate.
One order snapshot, as the model receives itORJ-0001-D1.txt · 1 of 240
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes
no real operator, customer, supplier, order or person. Order ORJ-0001, surveillance day D1.
Order
----------------------------------------------------------------
Order reference : ORJ-0001
Product : Site-to-site connectivity, dual circuit
Segment : Public sector framework
Install type : New build, new duct required
Order placed : Mon 16 Feb 2026
Committed date : Mon 09 Mar 2026
Jeopardy Policy
----------------------------------------------------------------
Overdue threshold : 2 working days
Blocker threshold : 4 working days
Rules stated in the fulfilment jeopardy policy:
1. An order with an open task MORE than the overdue threshold past its due date
is reported JEOPARDY, under MILESTONE_OVERDUE.
2. An order with an open task HELD MORE than the blocker threshold is reported
JEOPARDY, under BLOCKER_AGED.
3. An order whose Activity Log states, on this day, a lead time, a rebooked
appointment, a consent window or a booked date that reaches BEYOND the
committed date -- stated either as a number of working days ahead of today,
or as a date -- is reported JEOPARDY, under LEADTIME_EXCEEDS_REMAINING,
whether or not any task is yet overdue. A commitment landing ON the
committed date does not engage this rule. Nor does a note that states a
number of working days without committing to when anything will be ready --
a retention rule, a cooling-off period, an escalation service level, a
window that has already closed.
4. An order reported JEOPARDY on the previous surveillance day and STILL
Abridged — the file continues.
The outcomeWhat a good result looks like
Per order per surveillance day: the status to REPORT (CLEAR, WATCH or JEOPARDY), which of the five policy rules was applied, and the open task most responsible -- plus a four-field carried state the next day is judged against, written by code from the parsed task ages and never from the model's reply. Across the window: the day each order was first raised, and the working days that left before the committed date.
And when it cannot
It reports WATCH where the policy says CLEAR, because it read a narrative note as a change of task state. Two snapshots of 240, both in the direction that pages nobody, and one of the two did not reproduce on a second ask. The louder failure in this run is not a wrong verdict at all: one reply of 240 hit the output ceiling and returned nothing, which the scorer counts as a miss in all three fields.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
A queue whose jeopardy rules are all about task ages against thresholds — the free aging floor, at $0.00 It applies rules 1, 2 and 5 perfectly and names the top blocker on 240 of 240. It is what your fallout report already does, and this kit measured that it is not improved on by anything.
A rule about a commitment written into an activity note -- a supplier lead time, a rebooked slot, a consent window — the fast tier, with the carried state This is the whole gap. The aging floor misses 9 of the 30 slipped orders outright and loses 45 of the 120 available lead days, all of it on the two trajectories where nothing ever goes overdue. The model wins every one of those days back.
Notes written to a fixed template by a system rather than by a person — a pattern, at $0.00 That is exactly what this corpus is, and the pattern arm scores 100 pct on every axis of it. If your activity log is machine-written -- a status feed, a supplier API, a fixed set of dispositions -- write the pattern and keep the money.
Telling a NEW jeopardy from one that has been running for days — free code, with memory Rule 4 is memory and memory is arithmetic. The floor with memory names the rule correctly on 86.25 pct against 69.58 pct without it, and moves the lead time by exactly zero.
Deciding whether to rebook an engineer or tell a customer their date has moved — a person, reading the status and the rule the pack named The two content errors in this run were both the model reading a narrative note as a change of state, and one of the two did not reproduce on a second ask. Nothing here notifies anybody and nothing should be added that does without a human step in front of it.
And where nothing here is good enough:
An order book of thousands running every working day — nothing here yet -- measure it first The cost is per order per surveillance day. Five thousand open orders on a daily cadence is 25,000 calls a week, about $112 a week on the shared rate card, and the free floor answers most of them identically. The question is which orders are worth pointing a model at, and this kit has not answered it.
At a glanceHow the whole thing runs
100%caught before the miss pct
8,972 msp50, end to end
$4.48per 1,000 order snapshots · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a telecom order before it misses its date14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/jeopardy.py FIRST. Every figure on this page is a property of a 48-order corpus with eight planted trajectories, five surveillance days each and an activity log written from sixteen templates, measured once.Corpus lens →
When is this the wrong choice?
Avoid: Paying a model to do arithmetic a regex already does exactly. That is the case against the best-fitting scenario (“A queue whose jeopardy rules are all about task ages against thresholds”). 6 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Five surveillance days per order. Every rule in this policy resolves inside five, so nothing here measures a jeopardy that has been running for three weeks, a policy with a rolling breach count, or an order whose committed date is amended mid-window. 8 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER THE FREE PATTERN'S 100 PCT MEANS ANYTHING OUTSIDE THIS CORPUS, AND ONLY 24 ROWS OF IT ARE PRICED. The Activity Log is sixteen sentence templates and the pattern was written against them by the same author. 9 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-order-jeopardy. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes every pre-flight assertion (python3 -m evals.check_labels), scores all three free floors (python3 -m evals.run --run-id b002-order-jeopardy-regex --baseline regex, and the other two), re-scores the paraphrase probe's free half (python3 -m evals.paraphrase --free) and serves the whole board in a browser (python3 -m src.app, port 9005). Six commands, plain python3, no install step -- requirements.txt names nothing because the kit imports nothing outside the standard library. What it cannot do without a key is re-run the scored eval or the model's half of the paraphrase probe.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
99.58%rows answered
8,972 msp50, end to end
24,244 msp95
2 minclone to first result
What the clock covers. END-TO-END per surveillance day: one HTTP request carrying the assembled prompt, and the reply parsed to a status, a rule and a top blocker. Splitting the snapshot into sections, dropping the one no field asks for, parsing the task table and the days remaining, and advancing the carried state all happen outside this clock and cost no network at all. p95 (24244 ms) is 2.7 times p50 (8972 ms), and the slowest single call took 62480 ms. Output length tracks how hard the READING is, not how long the page is: every snapshot is between 2,941 and 3,482 bytes and replies ran from 222 to 7096 output tokens.
Current processWhat it replaces
A fulfilment coordinator opening every stalled order in the queue each morning, reading the task ages against the jeopardy thresholds, then reading the activity notes to find the supplier quote or rebooked date that means the committed date is already unreachable -- and remembering which of them were raised yesterday.
Where it is not good enough
⚑ THE HONEST HEADLINE IS THAT FREE CODE ALREADY DOES THIS, ON THIS CORPUS, PERFECTLY. A pattern-matching floor with no model and no calls scores 100 pct on every axis: 30 of 30 slipped orders surfaced before the miss, all 120 available lead days, 100 pct on status, rule and top blocker, 0 false alarms. The model matches it on the lead time -- 30 of 30, 120 of 120 lead days -- and is WORSE on the cells: 98.75 pct status, 99.58 pct rule, 99.17 pct top blocker, and it costs $0.004481 a surveillance day. Anyone reading this page to decide whether to point a model at their fallout queue should read that sentence first. ⚠︎ AND IT IS A FINDING ABOUT THIS CORPUS, NOT ABOUT THE TASK, WHICH THE KIT MEASURED RATHER THAN CAVEATED. The Activity Log is generated from sixteen sentence templates and the pattern that reads them was written by the same author against those templates -- the generator run backwards. evals/paraphrase.py rewrites 24 forward commitments into phrasings the generator never produced, stating the identical fact: the free pattern falls to 54.17 pct and reads only 1 of the 12 commitments that actually engage rule 3, while the model holds at 100.0 pct and reads 12 of 12. The pattern's single hit is not a reading either -- it is the aging threshold catching up on day 5 of an order whose task had also gone overdue. ⚑ THE RUN'S OWN ERRORS ARE THREE CELLS OF 720, AND THEY ARE THE MIRROR OF THE FLOOR'S. On ORJ-0025-D3 and ORJ-0033-D3 the model read a narrative note -- "the joint could not be completed and standing water was reported in the duct", "Civils gang stood down; weather" -- as a HOLD on a task the Open Tasks table shows as merely due, and reported WATCH where the policy says CLEAR. The floor cannot read prose at all; the model reads state changes into it. Both land in the harmless direction: WATCH is not a page, and there were ZERO false JEOPARDY on 143 quiet snapshots. ⚠︎ ONE OF THE TWO DID NOT REPRODUCE -- asked again through the UI, ORJ-0033-D3 came back CLEAR with the right reason, so 98.75 pct is a point estimate whose residual is partly noise and no repeat run exists to say how much. ⚠︎ AND ONE REPLY OF 240 WAS LOST TO THE OUTPUT CEILING, which is the other 3 miss cells. ORJ-0008-D3 produced the full 12000 tokens and no answer at all. It is one of the QUIETEST snapshots in the corpus -- one task due in two working days and one irrelevant note -- so the runaway does not track difficulty, and a ceiling set at 6.5 times the hardest reply the calibration saw was still gone over once.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt240jsonl1
48 open orders, 5 consecutive surveillance days each
30 of 30 surfaced before the miss · median lead 4 days
120 of 120 available lead days won
free code did the same for $0.00 — and falls to 54.17%
100%when the same facts are reworded; the model holds at
2026-08-23as of
It produces a status and a rule for a fulfilment coordinator to read at eight in the morning, and reassigns, rebooks and notifies nothing. THE STATION THAT MAKES THIS A MONITOR IS THE SECOND ONE: the only route from day 2 to day 3 is four scalar fields, written by src/jeopardy.py from the parsed task ages and never from the model's answer — which on a kit whose headline is LEAD TIME is also what stops an order wrongly raised on day 1 scoring a perfect lead time it never earned. ⚑ AND THE HEADLINE IS NOT AN ACCURACY, WHICH IS THE POINT OF THIS PICTURE. Of the 30 orders that missed their committed date, the model surfaced every one before the miss and won all 120 of the working days the policy makes available; the task-age floor a fallout queue already runs surfaced 21 and won 75. The whole 45-day gap sits on the two trajectories where NO TASK EVER GOES PAST ITS DUE DATE and the evidence is one sentence in an activity note — a supplier lead time, a rebooked slot, a booked date beyond the committed date.
⚠︎ A REGEX OVER THOSE NOTES CLOSES THE GAP COMPLETELY, FOR $0.00, AND SCORES 100 PCT ON EVERY FIELD. That row is published rather than hidden, and then tested: rewrite the same 24 commitments in words the corpus generator never produced and the regex reads 1 of the 12 that matter while the model reads 12 of 12. Adding MEMORY to the free floor moves the lead time by exactly zero — rule 4 renames a jeopardy already raised and cannot raise one — so what a model buys here is prose, not history.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the jeopardy policy
src/jeopardy.py
step() is your operator's policy. Change it there and re-run tools/build_corpus.py, which recomputes the whole answer key from the same function -- the gold cannot drift from the rules because it is their output. Change the printed text in the generator to match, because the model reads the policy off the page and not out of the prompt.
the carried state
src/state.py
load()/save() are a JSON file today. Point them at a table, a key-value store or the order management system itself and nothing else changes -- for_unit() and advance() are the whole interface, and describe() is the only thing the prompt sees.
the corpus
tools/build_corpus.py
Point it at your own task families and products, or delete it and drop real snapshots into data/corpus/ named <ORDER>-D<n>.txt with the same seven section headings. Everything downstream reads by order and day and does not care where the documents came from.
what is sent
src/select.py
SECTION_HINTS maps a field to the sections it needs; NEVER_SENT names what is withheld whatever happens. Add a section your snapshots carry and decide, once, which of the two lists it belongs in.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 48 orders x 5 surveillance days = 240 snapshots from a fixed seed (SEED = 20260823) across twelve task families, six products and eight planted trajectories -- five that slip their committed date and three that deliver. It writes the outcome of every order into the answer key and onto no page, which is what makes this a replay rather than a classification set. The gold labels are src/jeopardy.replay's output, never typed.
the jeopardy policy
src/jeopardy.py
The written policy as pure code: two thresholds, five rules, three statuses. Rules 1, 2 and 5 are arithmetic over task ages; rule 3 is a forward commitment stated in prose and this module never reads it -- it is passed in. Rule 4 is memory. The branch order is load-bearing: rule 4 is tested FIRST, so a running jeopardy is reported as sustained rather than as if it had just arrived.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. Four fields per order (the status reported yesterday, the rule applied, consecutive days in jeopardy, the day it was first raised), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on day 300 as on day 3.
the section splitter
src/segment.py
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 240 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Customer Contact is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
the fact parser
src/monitor.py
Reads the two thresholds, the days remaining and every open task's age off the page with a regex -- and DELIBERATELY DOES NOT READ THE ACTIVITY LOG. That asymmetry is the kit: rule 3 lives in prose, and a runtime that could parse it would make the model redundant without measuring anything. evals/check_labels.py asserts the facts this module exposes carry no rule-3 field.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The policy is NOT restated in the instruction -- it is printed on every snapshot, so a forker changes it in one place and the answer key follows.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost document.
the three free floors
evals/baseline.py
aging (task ages, no memory, never opens the Activity Log), agingmem (the same, told yesterday's status) and regex (the same, plus a genuine pattern over the Activity Log in both the count and the date form). Each differs from the one before it by exactly one thing, so the gaps price memory and prose-reading separately. 0 calls, $0.00, scored through the identical scorer.
the lead-time scorer
evals/scoring.py
Replays each arm's OWN answers order by order to find the day it first said JEOPARDY, and turns that into working days before the committed date. The per-cell accuracy is computed in the same pass and reported beside it, never instead of it -- an arm can be wrong on four snapshots and still have surfaced the order eight days early.
the paraphrase probe
evals/paraphrase.py
Rewrites 24 forward commitments into phrasings that appear nowhere in the generator and that no pattern in evals/baseline.py matches, stating the identical fact, and re-scores both arms. It exists because the free pattern scores 100 pct on the corpus as generated, and publishing that number without testing it would be publishing an artefact of the corpus.
the pre-flight
evals/check_labels.py
Everything that must be true before a run may spend: 240 documents, 48 gap-free orders, gold reproducible from jeopardy.replay, the parser agreeing with the key on every arithmetic fact, the privacy guard red-proven in both directions, and eight corpus-honesty assertions each of which measured a non-zero the day it was written.
the local UI
src/app.py
Port 9005. Replays any order through the policy and the three free floors with NO KEY AND NO CALLS -- three of the four rows on the board are pure code -- and adds the model's own row on one click when a key is configured. No route in the file reassigns, rebooks, contacts or writes anything.
Where it breaks at scale
NOT ON WINDOW LENGTH, AND THAT IS THE DESIGN. The carried state is four scalar fields, so day 300 of an order costs exactly what day 3 costs -- unlike a conversation kit, whose input grows with every turn. What it breaks on is four other things. FIRST, THE CHAIN IS SERIAL. Days within one order cannot be parallelised, because day 3's prompt contains state produced by day 2. This run went 48 chains wide with 12 workers and took 280.1 seconds of wall clock for 240 calls; an order book of one order and three hundred days would take three hundred serial calls and no width would help. SECOND, THE STATE STORE IS A FILE. src/state.py writes data/state.json atomically, which is correct for one process and is not a concurrency model. THIRD, THE COST IS PER ORDER PER DAY AND A REAL BOOK IS LARGE. A desk carrying five thousand open orders and running the pack every working day is 25,000 calls a week, which on the shared rate card is about $112 a week -- and the free floor answers 240 of the 240 correctly for nothing, so the question is not whether the model can afford the book but which orders it is worth pointing at. FOURTH, NOTHING HERE SCHEDULES ANYTHING. This is a monitor in the sense that it carries state between days; it is still invoked, not woken. The cadence, the retry of a day whose call failed, and what a surveillance pack should do when a day simply does not arrive are all outside this kit and are not measured by it.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
⚠︎ THIS SHOT WAS AIMED AT A FAILURE AND CAUGHT A CORRECT ANSWER, AND IT SHIPS AS IT CAME BACK. ORJ-0033-D3 is one of the two snapshots the recorded run got wrong: it read "Civils gang stood down; weather" as a hold on a task the table shows as merely due, and reported WATCH where the policy says CLEAR. Asked the same question again through the UI it answered CLEAR and gave the right reason -- "the booked civils date of Fri 08 May 2026 is before the committed date, so none of the jeopardy rules apply". One of the run's two content errors is not stable. That is a measurement about the residual rate and it is worth more than a staged failure would have been.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
THE FREE AGING FLOOR MISSING AN ORDER OUTRIGHT, and this shot costs nothing to take: three of the four rows are pure code. ORJ-0016 missed its committed date by 9 working days. No task on it ever went past its due date, so both aging rows say clear on all five surveillance days and their Lead column reads MISSED IT. The policy and the pattern arm raise it on D4, three working days before the date, off one line in the Activity Log. This is 6 of the 30 slipped orders -- the trajectory where an aging report has nothing to age.failureOpen full size →THE SAME PICTURE SHOWING LOST TIME RATHER THAN A TOTAL MISS. ORJ-0007 missed by 8 working days. The aging floor does catch it -- on D5, with 1 working day left, which is a page nobody can act on. The policy and the pattern arm raise it on D2 with 4 working days left, off a note saying the splice gang is booked for Tue 10 Mar against a committed date of Mon 09 Mar. Same order, same evidence, four days of difference: that gap summed over the whole corpus is the 45 lead days the aging floor loses.failureOpen full size →
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
240order snapshots
0.72 MiBjsonl 1 · txt 240
48orders (5 surveillance days each) · p50 5 chars
$0.00setup · 0.0s
How it is cutWhat one orders (5 surveillance days each) is
No split, and no chunking. The unit is an ORDER -- a run of five consecutive surveillance days processed strictly in order, because day 3's prompt contains state produced by day 2. Each snapshot goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step. tools/build_corpus.py writes 240 documents and the answer key from a fixed seed with no clock read and no model called; nothing is embedded, ranked or cached.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every order, product, task, supplier commitment, activity note and name is invented here. Every contact number is in the 07700 900000-900999 range reserved by Ofcom for fiction and never allocated to a subscriber.
Bring your ownBring your own order snapshots
Replace src/jeopardy.py FIRST. It is your operator's policy, not a utility, and data/gold.jsonl is literally its output -- change the rules and re-run tools/build_corpus.py and the answer key follows. Change the policy TEXT in the generator in the same edit, because the model reads the rules off the page rather than out of the prompt. Then point the generator at your own task families, or delete it and drop real snapshots into data/corpus/ named <ORDER>-D<n>.txt with the same seven section headings.
⚠︎ And what stops being true when you do: Every figure on this page is a property of a 48-order corpus with eight planted trajectories, five surveillance days each and an activity log written from sixteen templates, measured once. Two numbers will not survive contact with a real book. The lead times are a property of the TRAJECTORY MIX -- six orders per trajectory was a design decision, so a book with more slow-aging orders and fewer supplier-quote failures moves the headline without anything having changed about the arms. And the free pattern's 100 pct is a property of the WORDING: it reads 1 of the 12 rule-3 commitments once they are rephrased.
What breaks it
⚠︎ THE ACTIVITY LOG IS SIXTEEN SENTENCE TEMPLATES, AND THAT IS THIS CORPUS'S LOAD-BEARING LIMITATION. It is why the free pattern arm scores 100 pct on every axis: the pattern was written by the same author against the same templates. evals/paraphrase.py measures the size of that artefact on 24 snapshots -- the pattern falls to 54.17 pct -- but 24 rows in two phrasing families is not real desk prose, and the paraphrases share an author with the pattern too. What a corpus of genuinely captured notes would do to either arm is unmeasured.
⚠︎ EVERY FORWARD COMMITMENT ENGAGED RULE 3 ON THE FIRST BUILD -- 36 of 36 -- WHICH MEANS AN ARM COULD HAVE SCORED PERFECTLY ON RULE 3 WITHOUT EVER PERFORMING ITS COMPARISON. Nothing in the repository could see it: the answer key was correct on every page and the pre-flight passed clean. It was found by printing how many forward notes engaged the rule beside how many existed. There are now 66 that do not engage it, 24 of them landing exactly ON the committed date, and evals/check_labels.py carries a permanent assertion written while it measured 36 of 36.
Five surveillance days per order. Every rule in this policy resolves inside five, so nothing here measures a jeopardy that has been running for three weeks, a policy with a rolling breach count, or an order whose committed date is amended mid-window.
Thresholds that are unambiguous by construction. Task ages are drawn clear of their thresholds, so a wrong status is a wrong RULE and never a rounding argument. A real queue has tasks sitting exactly on a threshold and this corpus has few of them.
One clear top blocker. The generator guarantees the leading task leads by at least two working days, and evals/check_labels.py asserts it at 0 ties. A snapshot whose top two tasks are level is a question with no answer, and this corpus contains none of them.
One policy, reproduced identically on every page. A real operator's rules differ by product, by segment, by build type and by whether a third party is in the path; none of that is here.
No public-holiday calendar. Dates skip weekends and nothing else, and every date on these pages would move under a real one. The policy is written in WORKING DAYS REMAINING, printed on the page, so no arm has to do calendar arithmetic -- but the corpus cannot say what happens when the calendar is wrong.
No amended committed dates, no cancellations, no multi-site orders whose sites slip independently, and no order that appears in the window and then leaves it. Every one of those is ordinary in a real book.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
140
not measured
instruction
1,555
not measured
carried state
126
not measured
Synthetic Record
255
not measured
Order
341
not measured
Jeopardy Policy
1,542
not measured
Surveillance Day
201
not measured
Open Tasks
187
not measured
Activity Log
327
not measured
Total
1,138
This is the cost lesson as arithmetic: of the 4,674 characters assembled, 1,695 are instructions — 36% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.py::build on ORJ-0007-D2 with the exact carried state r001-order-jeopardy recorded for that call ({"prev_status": "CLEAR", "prev_rule": "NONE", "days_in_jeopardy": 0, "first_raised_day": null}) -- the identical code path the run used, not a paraphrase. THE SPLIT IS IN CHARACTERS AND THE PAGE SAYS SO: per-part TOKEN counts were not measured, because measuring them means sending nested prefixes of the prompt to the provider and this kit spent no calls on it. The billed totals (273043 in / 312941 out over 240 calls) are measured; nothing here apportions them across the parts.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You apply a written fulfilment jeopardy policy to one open order on one surveillance day. You answer with one JSON object and no other text.
You are reviewing one open telecommunications order on one surveillance day, against the written
fulfilment jeopardy policy reproduced in the snapshot below.
Apply the policy exactly as written. Rule 4 depends on what was reported on EARLIER surveillance
days, which you cannot see; what is known about this order's history is stated under "Carried
state" and is the only history available to you. Do not assume anything about earlier days beyond
it.
Read the Activity Log as carefully as the Open Tasks table. Rule 3 is about a forward commitment
stated in the notes — a supplier lead time, a rebooked appointment, a statutory consent window —
measured against the working days that remain. A note that merely mentions a number of working
days, such as a retention period, a cooling-off window or an escalation service level, is not a
forward commitment and does not engage rule 3.
Answer with a single JSON object and nothing else:
{"status": "CLEAR|WATCH|JEOPARDY",
"rule": "NONE|MILESTONE_OVERDUE|BLOCKER_AGED|LEADTIME_EXCEEDS_REMAINING|ESCALATED_SUSTAINED",
"top_blocker": "the open task most responsible, named exactly as it is written in Open Tasks",
"rationale": "one sentence, naming the rule or the threshold you applied"}
"status" is what to report for this order today. "rule" names which policy rule you applied, or
NONE. "top_blocker" is the open task with the most working days past due or held; if no task is
past due or held, it is the open task nearest to its due date; if the order has no open tasks at
all, it is the word None.
Carried state
----------------------------------------------------------------
On the previous surveillance day this order was reported CLEAR. It has not been in jeopardy on any earlier day of this window.
Order snapshot
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes
no real operator, customer, supplier, order or person. Order ORJ-0007, surveillance day D2.
Order
----------------------------------------------------------------
Order reference : ORJ-0007
Product : Ethernet leased line, 100 Mbit/s
Segment : Small business, single site
Install type : Migration from a third-party circuit
Order placed : Mon 02 Feb 2026
Committed date : Mon 09 Mar 2026
Jeopardy Policy
----------------------------------------------------------------
Overdue threshold : 2 working days
Blocker threshold : 4 working days
Rules stated in the fulfilment jeopardy policy:
1. An order with an open task MORE than the overdue threshold past its due date
is reported JEOPARDY, under MILESTONE_OVERDUE.
2. An order with an open task HELD MORE than the blocker threshold is reported
JEOPARDY, under BLOCKER_AGED.
3. An order whose Activity Log states, on this day, a lead time, a rebooked
appointment, a consent window or a booked date that reaches BEYOND the
committed date -- stated either as a number of working days ahead of today,
or as a date -- is reported JEOPARDY, under LEADTIME_EXCEEDS_REMAINING,
whether or not any task is yet overdue. A commitment landing ON the
committed date does not engage this rule. Nor does a note that states a
number of working days without committing to when anything will be ready --
a retention rule, a cooling-off period, an escalation service level, a
window that has already closed.
4. An order reported JEOPARDY on the previous surveillance day and STILL
breaching any of rules 1 to 3 today is reported JEOPARDY again, under
ESCALATED_SUSTAINED, in place of the rule it is still breaching.
5. An order with an open task past its due date, or held, but within both
thresholds is reported WATCH. An order with neither is reported CLEAR. A task
that has closed stops counting and is not listed.
Surveillance Day
----------------------------------------------------------------
Day : D2
Date : Tue 03 Mar 2026
Working days to committed date : 4
Open Tasks
----------------------------------------------------------------
Open tasks on this order:
Fibre splice Field engineering due in 3 working days
Activity Log
----------------------------------------------------------------
Notes recorded against this order since the previous surveillance day:
- The gang that covers fibre splice is booked for Tue 10 Mar 2026; there is no earlier slot.
- Porting desk validated the account details against the losing carrier's record.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"status": "JEOPARDY", "rule": "LEADTIME_EXCEEDS_REMAINING", "top_blocker": "Fibre splice", "rationale": "Activity Log shows the fibre splice gang is booked for Tue 10 Mar 2026, beyond the committed date of Mon 09 Mar 2026, so rule 3 applies."}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a telecom order before it misses its date — 240 order snapshots. One model answered, and every answer was then graded Four different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
There is no LLM-as-judge in this kit and that is not a shortcut. Every field it answers is a closed set -- three statuses, five rule codes, and a task name copied from a table on the page -- so correctness is a matter of equality, not of reading. And the HEADLINE is not a comparison of answers at all: it is a replay of each arm's own verdicts to find the day it first raised each order, which is arithmetic over a column.
240order snapshots
240source documents
1model tier
4grading methods
MeasurementsWhat was measured
COUNTED30 · 21 · 30 / 30caught before the miss pct — orders that ultimately slipped -- THE HEADLINEDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED1 · 1 · 1 / 30median lead days — slipped orders, working days between the first JEOPARDY and the committed dateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED144 · 22 · 36 / 120total lead days — lead days the policy makes available across the 30 slipped ordersDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED237 · 26 · 30 / 240status accuracy pct — snapshotsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED239 / 240rule accuracy pct — snapshotsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED238 / 240top blocker accuracy pct — snapshotsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 143false jeopardy rate pct — snapshots whose correct status is not JEOPARDYDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED240 / 240unparsed replies — calls -- one reply lost at the 12000-token ceilingDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The comparison cannot be wrong about itself; the risk is in the labels. Those come from src/jeopardy.replay, which the runtime never calls -- src/ asks the model for the status and never computes it -- so the answer key is arithmetic done independently of the code under test. evals/check_labels.py re-derives all 240 gold rows from the corpus values AND asserts that src/monitor.facts_of re-reads every arithmetic fact off the page identically, so a parser drift shows up as a refusal to start. What the key CANNOT settle is whether the policy's own lead time is the earliest a person could have known; see could_not_verify.
1,303.92output tokens · the fast tier, with the carried state · 8,972 ms p50
0.0output tokens · the three free floors -- aging, aging with memory, aging plus a pattern · 0 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 0.0× as long, and lands one row apart on 240. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One order snapshot
1,000 order snapshots
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.004481
$4.48
13%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.001792
$1.79
13%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.076573
$76.57
15%
Same work, 43× the bill
The same order snapshots, the same tokens — only the rate card changed. And across all 3 cards between 13% and 15% of what you pay is the prompt this pipeline sends, not the answer it writes.
provider-side reasoning -- 94.87 pct of this run's output tokens, and output is 87.30 pct of the projected bill. Nothing else on this page is close, and this kit has measured neither what disabling it costs in accuracy nor what a lower ceiling would have saved.
Rates checked 2026-08-18. The provider that actually ran all 276 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. A measured $0.00 -- and three of the four arms on every board here are pure code, so most of what this page compares cost nothing to produce either.
The gradersFour ways to grade
⚑ THERE ARE THREE FREE FLOORS AND THE ONE ABOVE IS THE WEAKEST OF THEM. Read all three before reading the model's row. The aging floor catches 21 of the 30 slipped orders and wins 75 of the 120 available lead days. Adding MEMORY changes neither number by one -- rule 4 renames a jeopardy that is already raised and cannot raise one -- so memory is worth exactly 16.67 points of RULE accuracy and nothing else. Adding a PATTERN over the Activity Log takes it to 30 of 30 and all 120 lead days, at 100 pct on every field, for $0.00: BETTER than the model, which scores 98.75 pct on status. ⚠︎ That last row is why evals/paraphrase.py exists. Rewrite the same commitments in words the generator never produced and the pattern falls to 54.17 pct while the model holds at 100.0 pct. The free arm's perfection is a property of sixteen sentence templates written by the author of the pattern that reads them.
the fast tier, with the carried state 100.0% caught before the miss · the free aging floor, no model 70.0% caught before the miss · the free aging floor, with memory 70.0% caught before the miss · the free aging floor + a pattern 100.0% caught before the miss · 3 more measured on each run
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key For each of the 240 snapshots and each of the three answered fields, did the reply equal the computed answer key? Status and rule are compared exactly; the task name is compared on trimmed lower-cased text and nothing looser, because a fuzzy match would forgive an arm that invented a plausible task.
$0.00
no
yes
the fast tier, with the carried state 98.8% status accuracy · the free aging floor, no model 86.2% status accuracy · the free aging floor, with memory 86.2% status accuracy · the free aging floor + a pattern 100.0% status accuracy · 2 more measured on each run
the fast tier, with the carried state 0.0% false jeopardy rate · the free aging floor, no model 0.0% false jeopardy rate · the free aging floor + a pattern 0.0% false jeopardy rate
The same forward commitment, rewritten in words the generator never produced When a supplier commitment is stated in a phrasing that appears nowhere in the corpus generator and that no pattern in evals/baseline.py matches, does the arm still reach the right status? This is the only grader here that separates reading the SENTENCE from matching the PATTERN.
$0.00
yes — every row
no
the fast tier, with the carried state 100.0% status accuracy · the free aging floor + a pattern 54.2% status accuracy
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Decisively in one direction and not at all in three others. THE AGING FLOOR AND THE MODEL SEPARATE BY 9 ORDERS AND 45 LEAD DAYS on identical snapshots with an identical scorer, and the whole gap sits on the two trajectories where no task ever goes past its due date. What this labelled set CANNOT separate is the model from a PATTERN: on the corpus as generated they tie on the lead time and the pattern is ahead on the cells, and only the paraphrase probe -- 24 rows, two phrasing families, one author -- pulls them apart. Nor can it separate a monitor with memory from one without on anything that matters: the two free arms differ by 16.67 points of RULE accuracy and by nothing else at all. And it cannot separate caution from luck on false jeopardy, because every arm scored 0 of 143.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
A queue whose jeopardy rules are all about task ages against thresholds
the free aging floor, at $0.00
It applies rules 1, 2 and 5 perfectly and names the top blocker on 240 of 240. It is what your fallout report already does, and this kit measured that it is not improved on by anything.
Paying a model to do arithmetic a regex already does exactly.
A rule about a commitment written into an activity note -- a supplier lead time, a rebooked slot, a consent window
the fast tier, with the carried state
This is the whole gap. The aging floor misses 9 of the 30 slipped orders outright and loses 45 of the 120 available lead days, all of it on the two trajectories where nothing ever goes overdue. The model wins every one of those days back.
Reading that as settled. On the corpus AS GENERATED a free pattern does it too, and for nothing; the case for the model rests on the paraphrase probe, which is 24 rows.
Notes written to a fixed template by a system rather than by a person
a pattern, at $0.00
That is exactly what this corpus is, and the pattern arm scores 100 pct on every axis of it. If your activity log is machine-written -- a status feed, a supplier API, a fixed set of dispositions -- write the pattern and keep the money.
Assuming your log is templated because the first fifty notes look alike. Check the tail, which is where the free-text ones are.
Telling a NEW jeopardy from one that has been running for days
free code, with memory
Rule 4 is memory and memory is arithmetic. The floor with memory names the rule correctly on 86.25 pct against 69.58 pct without it, and moves the lead time by exactly zero.
Paying for it. Nothing in this kit's headline moves when memory is added or removed.
Deciding whether to rebook an engineer or tell a customer their date has moved
a person, reading the status and the rule the pack named
The two content errors in this run were both the model reading a narrative note as a change of state, and one of the two did not reproduce on a second ask. Nothing here notifies anybody and nothing should be added that does without a human step in front of it.
Wiring any of this to a customer communication. There is no apply path in the kit and there should not be one.
An order book of thousands running every working day
nothing here yet -- measure it first
The cost is per order per surveillance day. Five thousand open orders on a daily cadence is 25,000 calls a week, about $112 a week on the shared rate card, and the free floor answers most of them identically. The question is which orders are worth pointing a model at, and this kit has not answered it.
Assuming the per-day price is the whole bill. It is the price of one order for one morning.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
narrative_read_as_state
A note about what happened read as a change of task state
2
ORJ-0033-D3: "Civils build is held due to the gang standing down, but the hold is within the 4-working-day blocker threshold, so the order is WATCH." The Open Tasks table says Civils build is DUE IN 3 WORKING DAYS -- it is not held at all, and the note is a…
no_verdict
No reply at all
1
ORJ-0008-D3 produced the full 12000 output tokens with finish_reason 'length' and an empty body -- the whole budget spent on reasoning before any JSON was emitted. It is scored as a miss in all three fields. ⚠︎ THE SNAPSHOT IS ONE OF THE QUIETEST IN THE…
lead_time_lost
The order was surfaced, but later than the policy allows
0
None on this run. The model raised every one of the 30 slipped orders on the exact surveillance day the policy makes available -- 120 of 120 lead days, nothing lost. The free aging floor loses 45 of them.
false_jeopardy
An order raised that the policy does not raise
0
None. 0 of 143 quiet snapshots on every arm, including the pattern arm faced with 106 distractor notes stating more working days than remain.
correct
Every cell correct
237
237 of 240 snapshots answered all three fields correctly, including all 97 snapshots whose correct status is JEOPARDY, all 47 whose correct status is WATCH, and all 18 post-recovery snapshots on the trajectory where a raised order comes back down.
What we could NOT verify
WHETHER THE FREE PATTERN'S 100 PCT MEANS ANYTHING OUTSIDE THIS CORPUS, AND ONLY 24 ROWS OF IT ARE PRICED. The Activity Log is sixteen sentence templates and the pattern was written against them by the same author. evals/paraphrase.py rewrites 24 commitments into unseen phrasings and the pattern falls from 100 pct to 54.17 pct while the model holds at 100.0 pct -- but the paraphrases share that author too, so the probe is an UPPER BOUND on a pattern-matcher's brittleness rather than an estimate of it. What a corpus of genuinely captured desk notes would do to either arm is the single biggest thing this kit does not know.
WHETHER THE POLICY'S LEAD TIME IS THE LEAD TIME THAT MATTERS. Every 'lead days available' figure here is the earliest day the WRITTEN RULES could fire on the evidence present. On several orders a coordinator reading the same note would have called it earlier, and on the aging trajectories the policy itself is arguably late. So 'the model won all 120 available lead days' means it matched the policy exactly, and the policy is a floor on what was knowable rather than a ceiling.
ERROR PROPAGATION, WHICH IS NOT MEASURED AT ALL. The carried state is advanced from the answer key, because rule 3's fact lives only in prose and nothing in src/ may parse it. That is what makes the lead time a property of each day's reading -- an arm wrong on day 2 is not punished on days 3 to 5 for a history it never had -- and it means a deployment's real behaviour, where a wrong day-2 reading corrupts the rest of the order, is outside every figure on this page.
One scored run, fired once. Nothing here is a distribution: whether 98.75 pct status repeats is unmeasured, and one of the two content errors DID NOT REPRODUCE when the same snapshot was re-asked through the UI, so part of the residual is known to be noise and none of it is quantified.
One tier. No second model was run, so nothing on this page compares two models -- the arms differ by what they are allowed to READ and by whether a model was called at all, never by which model.
Whether a stateless arm would score differently. src/prompt.py builds one and evals/check_labels.py asserts it differs from the shipped prompt in exactly one block, so the experiment is available and provably clean -- and it was never fired, because the free floors already price memory at $0.00 and the answer there was 'nothing'. That is an argument, not a measurement, and the page says which.
Whether the 12000-token ceiling is enough. This run lost one reply to it on one of the quietest snapshots in the corpus, and the calibration it was set from (largest reply 1,845 tokens over 10 calls on the two HARDEST orders) gave no warning that a quiet snapshot could produce 12000. Nothing bounds a distribution whose observed spread is 222 to 12000.
Whether disabling provider-side reasoning changes the answers. 94.87 pct of this run's output tokens were reasoning, left at the provider's default. Turning it off would change the price by roughly an order of magnitude and this kit has not measured what it does to the lead time.
Whether rendering the carried state as ENGLISH rather than as JSON matters. src/state.describe argues for English and says plainly that the argument is unmeasured; scoring the same 240 snapshots both ways would cost one more full run and has not been paid for.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,137.68
1,303.92
8,972 ms
$0.004481
$0.001792
$0.076573
the three free floors -- aging, aging with memory, aging plus a pattern
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
⚑ 276 LIVE CALLS WERE MADE FOR THIS KIT AND EVERY ONE OF THEM IS BEHIND A FIGURE ON THIS PAGE. 240 were the scored run, 10 were the calibration that set the output ceiling, 24 were the model's half of the paraphrase probe, and 2 were the UI screenshot -- one aimed at a failure that did not reproduce, and one retaken after the shot was renamed to say so. NOTHING WAS DISCARDED and no run was fired twice. Everything else on this page -- three free floors, the free half of the paraphrase probe, the scorer, the pre-flight and the whole board in the UI -- is pure code and costs $0.00. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING, 94.87 pct of output tokens (296887 of 312941) and therefore most of the bill. It was left at the provider's default and this kit has not measured what turning it off does to the answers.
HOW HARD THE READING IS, NOT HOW LONG THE PAGE IS. Every snapshot is between 2,941 and 3,482 bytes and output tokens ran from 222 to 12000 -- a 54-fold spread on a corpus of one shape. ⚠︎ AND THE SPREAD IS NOT EXPLAINED BY DIFFICULTY: the single call that reached the ceiling was one of the quietest snapshots in the set.
THE FIXED HEAD OF THE PROMPT. The system message, the instruction and the five policy rules are 3237 of 4847 characters -- 67 pct, byte-identical on all 240 calls except for two threshold numbers. The order's own situation is the smaller half of what is sent.
⚑ AND THE LARGEST COST LEVER ON THIS PAGE IS NOT A TOKEN SETTING. It is WHICH ORDERS GET SENT. The free aging floor answers 86.25 pct of snapshots identically to the model for nothing; the whole measured advantage sits on the two trajectories where nothing is overdue. An operator who ran free code first and sent a model only the orders it calls CLEAR would pay for a fraction of the book -- unmeasured here, and the obvious next experiment.
Your volumeWhat it costs at your volume
Linear in ORDER-DAYS, and flat in window length, which is the unusual half. Each surveillance day is one call whose input is one snapshot plus four carried fields, so day 300 of an order costs exactly what day 3 costs. 240 snapshots project to $1.0753 on the shared rate card, so ten times the set is about $10.75. ⚠︎ THE FIGURE THAT MATTERS IS NOT 10x THE CORPUS, IT IS THE BOOK: a desk with five thousand open orders on a daily cadence is 25,000 calls a week, about $112 a week. What does NOT scale is wall clock -- days within one order are strictly serial, so a long window is a long queue no width can shorten.
Where pricing changes shape
THE OUTPUT CEILING IS A CLIFF AND THIS RUN WENT OVER IT ONCE. ORJ-0008-D3 produced the full 12000 tokens and returned nothing at all; it is billed in full and scored as three wrong cells. The ceiling was set at 6.5 times the largest reply a 10-call calibration on the two HARDEST orders had seen (1,845 tokens), and the call that blew it was one of the QUIETEST snapshots in the corpus. A ceiling is not itself billed -- the provider charges tokens produced, not tokens allowed -- so the 211 calls that finish under 2000 tokens pay nothing for headroom; but a call that hits it pays everything and returns nothing.
THE STEP FROM FREE TO PAID IS THE WHOLE COST OF THIS KIT, AND IT IS A STEP RATHER THAN A SLOPE. Three of the four arms cost $0.00 and two of them are within 45 lead days of the model's result. The decision is not which model to buy; it is whether to call one at all.
Your return, with your numbers
Volumeorder-days per surveillance cycle -- this run judged 240 (48 orders x 5 days) in one pass
What it replacesa fulfilment coordinator reading every stalled order's task ages against the jeopardy thresholds and then reading its activity notes for a commitment that already makes the committed date unreachable
Time saved per itemnot measured here -- and the honest comparator is not a person, it is the aging report the desk already runs, which is free and catches 21 of the 30
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start a question whose honest answer turned out to be 'free code already does this on this corpus', which here it does.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
273,043input tokens · this run
312,941output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every model figure on these pages: 240 surveillance days, one completion call each, one tier. The 10 calibration calls and the 24 paraphrase-probe calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here. The three free floors made no calls at all.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.430
$0.430
$1.79
2026-09-12
gemini-3-flash
Google
$1.075
$1.075
$4.48
2026-09-18
gemini-3-8-flash
Google
$1.378
$1.378
$5.74
2026-09-18
llama-5
Meta
$1.671
$1.671
$6.96
2026-09-18
claude-haiku-4-5
Anthropic
$1.838
$1.838
$7.66
2026-09-12
grok-4-5
xAI
$2.424
$2.424
$10.10
2026-09-18
grok-4-6
xAI
$2.424
$2.424
$10.10
2026-09-18
claude-sonnet-5
Anthropic
$3.675
$3.675
$15.31
2026-09-12
gemini-3-1-pro
Google
$4.301
$4.301
$17.92
2026-09-18
gpt-5-6-terra
OpenAI
$4.301
$4.301
$17.92
2026-09-12
gpt-5-6-sol
OpenAI
$7.351
$7.351
$30.63
2026-09-12
claude-opus-4-8
Anthropic
$9.189
$9.189
$38.29
2026-09-12
claude-opus-5
Anthropic
$9.189
$9.189
$38.29
2026-09-12
claude-fable-5
Anthropic
$18.377
$18.377
$76.57
2026-09-18
claude-fable-5-1
Anthropic
$18.377
$18.377
$76.57
2026-09-18
gpt-6-astra
OpenAI
$18.377
$18.377
$76.57
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 94.87 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (296887 of 312941), left at the provider's default, so every row below prices a reasoning-on workload. Output is 87.30 pct of the projected bill on the shared card. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the lead time.
⚠︎ EVERY ROW PRICES AN ARM THAT A FREE ONE MATCHED. The pattern floor scored 30 of 30 orders and all 120 lead days for $0.00 on this corpus, and beat the model on the cells. The cheapest row here is not the cheapest option; the cheapest option is not calling a model. What the paraphrase probe says about that is in Eval.baseline_note and should be read in the same breath.
Output length here is driven by how hard the READING is, not by document size -- 222 tokens to the 12000 ceiling on documents all within 541 bytes of each other -- and the one call that reached the ceiling was on one of the quietest snapshots in the corpus. A corpus with a different mix of trajectories would move every row below.
THESE ROWS PRICE ORDER-DAYS, NOT AN ORDER BOOK. 240 snapshots is 48 orders watched for one working week. A desk carrying five thousand open orders on a daily cadence is 25,000 calls a week -- about $112 on the shared card -- and nothing here has measured which of them are worth sending.
Nothing here includes retries, and this run had none that reached the harness. It DOES include one reply that never arrived: ORJ-0008-D3 produced 12000 output tokens at the ceiling and returned nothing, and those tokens are billed like any others.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
13 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 48 orders x 5 surveillance days = 240 snapshots from a fixed seed (SEED = 20260823) across twelve task families, six products and eight planted trajectories -- five that slip their committed date and three that deliver. It writes the outcome of every order into the answer key and onto no page, which is what makes this a replay rather than a classification set. The gold labels are src/jeopardy.replay's output, never typed.
You change it to: Point it at your own task families and products, or delete it and drop real snapshots into data/corpus/ named <ORDER>-D<n>.txt with the same seven section headings. Everything downstream reads by order and day and does not care where the documents came from.
tools/build_corpus.py
# Generate the synthetic order book: 48 orders x 5 surveillance days = 240 snapshots, plus the
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
SEED = 20260823
ORDERS = 48
DAYS = 5
RULE = "-" * 64
TASKS = [
PRODUCTS = [
src/jeopardy.pythe jeopardy policy — a swap seam
The written policy as pure code: two thresholds, five rules, three statuses. Rules 1, 2 and 5 are arithmetic over task ages; rule 3 is a forward commitment stated in prose and this module never reads it -- it is passed in. Rule 4 is memory. The branch order is load-bearing: rule 4 is tested FIRST, so a running jeopardy is reported as sustained rather than as if it had just arrived.
You change it to: step() is your operator's policy. Change it there and re-run tools/build_corpus.py, which recomputes the whole answer key from the same function -- the gold cannot drift from the rules because it is their output. Change the printed text in the generator to match, because the model reads the policy off the page and not out of the prompt.
SEAM 2 -- the thing that makes this a monitor. Four fields per order (the status reported yesterday, the rule applied, consecutive days in jeopardy, the day it was first raised), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on day 300 as on day 3.
You change it to: load()/save() are a JSON file today. Point them at a table, a key-value store or the order management system itself and nothing else changes -- for_unit() and advance() are the whole interface, and describe() is the only thing the prompt sees.
src/state.py
# The carried state — what makes this a monitor rather than another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_unit(store, order_id):
def advance(store, order_id, max_past_due, longest_held, overdue_at, blocker_at, lead_excess, day):
def describe(state):
src/segment.pythe section splitter
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 240 documents before a run may spend.
src/segment.py
# Split an order snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Order", "Jeopardy Policy", "Surveillance Day",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector — a swap seam
Decides which sections reach the model. Customer Contact is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
You change it to: SECTION_HINTS maps a field to the sections it needs; NEVER_SENT names what is withheld whatever happens. Add a section your snapshots carry and decide, once, which of the two lists it belongs in.
src/select.py
# Pick which sections of an order snapshot are sent. Pure code — the last deterministic step
BANNER = "Synthetic Record"
ORDER = "Order"
POLICY = "Jeopardy Policy"
DAY = "Surveillance Day"
TASKS = "Open Tasks"
CONTACT = "Customer Contact"
LOG = "Activity Log"
NEVER_SENT = (CONTACT,)
SECTION_HINTS = {
src/monitor.pythe fact parser
Reads the two thresholds, the days remaining and every open task's age off the page with a regex -- and DELIBERATELY DOES NOT READ THE ACTIVITY LOG. That asymmetry is the kit: rule 3 lives in prose, and a runtime that could parse it would make the model redundant without measuring anything. evals/check_labels.py asserts the facts this module exposes carry no rule-3 field.
src/monitor.py
# One surveillance day, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You apply a written fulfilment jeopardy policy to one open order on one surveillance "
MAX_TOKENS = 12000
FIELDS = ("status", "rule", "top_blocker")
CALIBRATION_ORDERS = ("ORJ-0007", "ORJ-0016", "ORJ-0024", "ORJ-0032")
def documents():
def units():
def load_doc(doc_id):
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The policy is NOT restated in the instruction -- it is printed on every snapshot, so a forker changes it in one place and the answer key follows.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework — string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost document.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
evals/baseline.pythe three free floors
aging (task ages, no memory, never opens the Activity Log), agingmem (the same, told yesterday's status) and regex (the same, plus a genuine pattern over the Activity Log in both the count and the date form). Each differs from the one before it by exactly one thing, so the gaps price memory and prose-reading separately. 0 calls, $0.00, scored through the identical scorer.
evals/baseline.py
# Three free floors, none of them a straw man. No key, no calls, $0.00.
WORDS = {"one": 1, "two": 2, "three": 3, "four": 4, "five": 5, "six": 6, "seven": 7, "eight": 8,
NUM = r"(\d+|" + "|".join(WORDS) + r")"
FORWARD_PATTERNS = [
FORWARD_RE = [re.compile(p, re.I) for p in FORWARD_PATTERNS]
DATE = r"((?:Mon|Tue|Wed|Thu|Fri|Sat|Sun) \d{2} [A-Z][a-z]{2} \d{4})"
FORWARD_DATE_PATTERNS = [
FORWARD_DATE_RE = [re.compile(p, re.I) for p in FORWARD_DATE_PATTERNS]
COMMITTED_RE = re.compile(r"^Committed date : (.+)$", re.M)
def _value(tok):
evals/scoring.pythe lead-time scorer
Replays each arm's OWN answers order by order to find the day it first said JEOPARDY, and turns that into working days before the committed date. The per-cell accuracy is computed in the same pass and reported beside it, never instead of it -- an arm can be wrong on four snapshots and still have surfaced the order eight days early.
evals/scoring.py
# Score a run against the computed answer key. Pure code, no model, no judge.
STATUSES = ("CLEAR", "WATCH", "JEOPARDY")
RULES = ("NONE", "MILESTONE_OVERDUE", "BLOCKER_AGED", "LEADTIME_EXCEEDS_REMAINING",
def _pct(n, d):
def _norm(v):
def _median(xs):
def score(records, golds):
def _by_pattern(leads):
def compare(arms):
evals/paraphrase.pythe paraphrase probe
Rewrites 24 forward commitments into phrasings that appear nowhere in the generator and that no pattern in evals/baseline.py matches, stating the identical fact, and re-scores both arms. It exists because the free pattern scores 100 pct on the corpus as generated, and publishing that number without testing it would be publishing an artefact of the corpus.
evals/paraphrase.py
# The paraphrase probe: does an arm read the SENTENCE, or the PATTERN?
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
COUNT_FORMS = [
DATE_FORMS = [
def _pick(s):
def _forward_note(g):
def pick(golds, n_each=12):
def rewrite(text, note, form, remaining, committed, day_offset_date):
evals/check_labels.pythe pre-flight
Everything that must be true before a run may spend: 240 documents, 48 gap-free orders, gold reproducible from jeopardy.replay, the parser agreeing with the key on every arithmetic fact, the privacy guard red-proven in both directions, and eight corpus-honesty assertions each of which measured a non-zero the day it was written.
evals/check_labels.py
# Everything that must be true of the corpus BEFORE a run may spend. Free, no key, no model.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAIL = []
DOCS = 240
ORDERS = 48
DAYS = 5
def check(name, ok, detail=""):
def main():
WORDS = {"one": 1, "two": 2, "three": 3, "four": 4, "five": 5, "six": 6, "seven": 7, "eight": 8,
src/app.pythe local UI
Port 9005. Replays any order through the policy and the three free floors with NO KEY AND NO CALLS -- three of the four rows on the board are pure code -- and adds the model's own row on one click when a key is configured. No route in the file reassigns, rebooks, contacts or writes anything.
src/app.py
# The minimal local UI. Standard library only -- python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9005"))
def gold():
def replay_free(order_id):
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 48 orders x 5 surveillance days = 240 snapshots from a fixed seed (SEED = 20260823) across twelve task families, six products and eight planted trajectories -- five that slip their committed date and three that deliver. It writes the outcome of every order into the answer key and onto no page, which is what makes this a replay rather than a classification set. The gold labels are src/jeopardy.replay's output, never typed. A swap seam.
src/jeopardy.pyThe written policy as pure code: two thresholds, five rules, three statuses. Rules 1, 2 and 5 are arithmetic over task ages; rule 3 is a forward commitment stated in prose and this module never reads it -- it is passed in. Rule 4 is memory. The branch order is load-bearing: rule 4 is tested FIRST, so a running jeopardy is reported as sustained rather than as if it had just arrived. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Four fields per order (the status reported yesterday, the rule applied, consecutive days in jeopardy, the day it was first raised), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on day 300 as on day 3. A swap seam.
src/segment.pySplits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 240 documents before a run may spend.
src/select.pyDecides which sections reach the model. Customer Contact is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing. A swap seam.
src/monitor.pyReads the two thresholds, the days remaining and every open task's age off the page with a regex -- and DELIBERATELY DOES NOT READ THE ACTIVITY LOG. That asymmetry is the kit: rule 3 lives in prose, and a runtime that could parse it would make the model redundant without measuring anything. evals/check_labels.py asserts the facts this module exposes carry no rule-3 field.
src/prompt.pyThree parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The policy is NOT restated in the instruction -- it is printed on every snapshot, so a forker changes it in one place and the answer key follows.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost document. A swap seam.
evals/baseline.pyaging (task ages, no memory, never opens the Activity Log), agingmem (the same, told yesterday's status) and regex (the same, plus a genuine pattern over the Activity Log in both the count and the date form). Each differs from the one before it by exactly one thing, so the gaps price memory and prose-reading separately. 0 calls, $0.00, scored through the identical scorer.
evals/scoring.pyReplays each arm's OWN answers order by order to find the day it first said JEOPARDY, and turns that into working days before the committed date. The per-cell accuracy is computed in the same pass and reported beside it, never instead of it -- an arm can be wrong on four snapshots and still have surfaced the order eight days early.
evals/paraphrase.pyRewrites 24 forward commitments into phrasings that appear nowhere in the generator and that no pattern in evals/baseline.py matches, stating the identical fact, and re-scores both arms. It exists because the free pattern scores 100 pct on the corpus as generated, and publishing that number without testing it would be publishing an artefact of the corpus.
evals/check_labels.pyEverything that must be true before a run may spend: 240 documents, 48 gap-free orders, gold reproducible from jeopardy.replay, the parser agreeing with the key on every arithmetic fact, the privacy guard red-proven in both directions, and eight corpus-honesty assertions each of which measured a non-zero the day it was written.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1137 input and 1303 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Activity Log, which is drawn from sixteen sentence templates. src/select.py names the Activity Log as the field an outside party WOULD control in a real deployment -- engineers, suppliers' agents and desk staff write it -- and sends it deliberately, because rule 3 lives there. The surface is visible rather than hidden; on this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after. src/app.py redacts the key and the base URL out of any provider error before returning it to the browser.
The experimentWe did not attack it -- and the boundary that matters most was proven by breaking it on purpose
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus has none. So the five rows below are boundaries, not payloads. Two of them were red-proven rather than asserted: the privacy guard was replaced with the naive fallback every sibling kit once shipped and all 240 snapshots leaked the customer's name and mobile number, then it was restored and all 240 held. A guard that has never been made to fail is a guard nobody has tested. Confirmed by assertion and by reading the recorded run, not by an attack trial, on 2026-08-23 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Can a wrong verdict poison the next surveillance day?
A monitor that fed the model's verdict into its own next prompt would compound one bad day into every day after it -- and on THIS kit it would also destroy the measurement, because an order wrongly raised on day 1 would stay raised for free and score a perfect lead time it never earned.
It cannot. evals/run.py advances the state with src/jeopardy.step() over the task ages parsed from the page and the answer key's rule-3 fact; the model's reply is scored and never written. Measured on the scored run: 3 snapshots were answered wrongly, in 3 different orders, and no order was answered wrongly twice. ⚠︎ Read that honestly -- the property holds by construction, and this run cannot demonstrate it either way, because it also means NOTHING HERE MEASURES ERROR PROPAGATION AT ALL. A deployment must persist the forward-commitment fact itself.
Does the customer's name or number ever leave the machine?
The Customer Contact section names the account contact, the site contact, a mobile number and the account manager -- the only place in the corpus whose subject is a PERSON -- and a selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the document MINUS that section rather than to the document. Red-proven in both directions by evals/check_labels.py over all 240 snapshots: guard on, 0 leak; guard replaced with the naive fallback, 240 of 240 leak.
Can the runtime cheat rule 3?
If anything under src/ could parse a stated lead time off the Activity Log, the model would be measuring nothing and the whole comparison would be between two regexes wearing different labels.
evals/check_labels.py asserts on the FACTS src/monitor.facts_of actually produces -- not on a reading of the code -- that they carry no rule-3 field and no forward commitment anywhere. The only place in the repository that reads a commitment out of a note is evals/baseline.py's pattern arm, which exists to be scored.
Can a snapshot tell the model what happened on an earlier day?
If any note named an earlier day's status or verdict, the carried state would stop being the only route between days and the memory figures would measure nothing.
evals/check_labels.py scans every Activity Log NOTE LINE for a day token other than its own and for a verdict vocabulary, and refuses to let a run spend if it finds one. 0 leaks. ⚠︎ THAT ASSERTION WAS WRONG ON ITS FIRST RUN AND THE RULE WAS FIXED, NOT THE CORPUS: it searched the whole section for 'was reported' and convicted 22 snapshots carrying a field engineer's note about standing water being reported in a duct. A check that fires on correct pages is a check somebody switches off.
Can a figure measured under a non-published output ceiling be mistaken for a scored one?
A calibration fired at a higher cap produces better-looking numbers under a configuration the page does not name.
evals/run.py refuses --max-tokens unless the run id begins with 'c'. The calibration is c000 at a cap of 8,000; the scored run carries the published MAX_TOKENS = 12000 from src/monitor.py.
Each boundary above was checked by running an assertion or by reading the recorded run, not by an attack trial -- there is no untrusted field on this corpus to construct a payload against. Two of the five are red-proven, meaning the guard was removed and the failure observed rather than merely asserted while passing.
The result0 attack trials, and five boundaries checked -- two of them red-proven by removing the guard and watching all 240 snapshots fail.
0untrusted input fields on this corpus
0 of 0attack trials run
2 of 5boundaries red-proven, not just asserted
The Activity Log IS the field an outside party would control in a real deployment -- field engineers, suppliers' agents and desk staff type into it -- and this kit sends it rather than hiding it, because rule 3 lives there. On this corpus it is drawn from sixteen templates by a seeded random number generator, so there is nothing adversarial in it to catch. A version pointed at real notes reopens the question and should be attacked before it ships: an instruction embedded in an engineer's note is the obvious payload, and this kit has not fired one.
Read this twice
The carried state is written from the arithmetic and the answer key, never from the model's reply. That is the difference between a monitor and a system that compounds its own mistakes — and on a kit whose headline is LEAD TIME it is also the difference between a measurement and an artefact, because an order wrongly raised on day 1 would otherwise stay raised for free and score a perfect lead time it never earned. ⚠︎ THE PRICE OF THAT CHOICE IS STATED RATHER THAN HIDDEN: rule 3's fact cannot be parsed by anything in src/, so the harness takes it from the key, which means nothing on this page measures what happens when a deployment persists a WRONG one.
HonestyWhat this does not prove
Whether a real deployment's Activity Log -- prose an engineer or a supplier's agent writes freely -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at real notes. It is also the section rule 3 depends on, so it cannot simply be withheld.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two schedulers advancing the same order have never been run and would race.
Whether a truncated reply can ever be partially trusted. ORJ-0008-D3 returned an EMPTY body at the ceiling, so there was nothing to partially trust; src/monitor._parse rejects it outright and the run records it as a failure. That is the safe behaviour and is not the same as having measured how often a fragment would have been right.
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
The carried state is written by code from the parsed task ages and the answer key's rule-3 fact. The model's status, rule and top blocker are scored and are never written back into the history the next surveillance day is judged against. And nothing in the kit acts on any of it: there is no path that reassigns a task, moves a committed date, chases a supplier or contacts a customer.
evals/run.py's per-day loop calls src/jeopardy.step() after every snapshot including one whose call FAILED. There is no code path anywhere in the kit that writes a model reply into the state store, and no writer of any kind outside results/*.json and data/state.json.
EvidenceDoes it hold?
What
Measured
A wrong verdict does not propagate to the next surveillance day
3 wrongly answered snapshots on r001-order-jeopardy, in 3 DIFFERENT orders -- no order wrong twice. ⚠︎ The property is guaranteed by the code path (jeopardy.step() never reads the reply) and this run is CONSISTENT with it rather than a demonstration of it. It also means the run measures no propagation at all, which is the honest reading and is repeated in Eval.could_not_verify.
A surveillance day whose CALL failed still advances the history
ORJ-0008-D3 lost its reply to the output ceiling and D4 and D5 of that order were still judged against a correct carried state -- both answered correctly. One transport or ceiling failure cannot turn into three scored failures, and unlike most kits in this series THIS BRANCH WAS ACTUALLY EXERCISED on the published run rather than merely written.
Nothing in this kit notifies anybody or changes an order
JEOPARDY is a string in a JSON reply and a column in a result file. There is no notifier, no webhook, no scheduler and no writer outside results/*.json and data/state.json -- confirmed by reading every call site, including src/app.py, whose only POST route returns a verdict to the browser.
The customer's details never reach the provider
0 of 240 snapshots leak the Customer Contact section, and 240 of 240 leak it when the guard is removed. Both directions asserted by evals/check_labels.py before any run may spend.
The runtime cannot answer rule 3 without the model
evals/check_labels.py asserts on the facts src/monitor.facts_of produces that they carry no rule-3 field. Without it the free floor and the model would be reading the same parse and the whole comparison would be circular.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE VERDICT IS RIGHT. The state is correct whatever the model says, which means a wrong status is reported wrongly for that day and nothing catches it -- see Eval.taxonomy's two snapshots where a narrative note was read as a change of task state.
IT IS NOT A RATE LIMIT OR AN APPROVAL GATE. There is no apply path to gate: nothing here escalates, rebooks or contacts anybody, so there is no 'undo' because there is nothing that could 'do'.
⚠︎ IT IS NOT A MEASUREMENT OF ERROR PROPAGATION, AND ON THIS KIT THAT MATTERS MORE THAN USUAL. The state is advanced from the ANSWER KEY, because rule 3's fact lives only in prose. A real deployment must persist that fact itself, from its own reading, and a wrong one corrupts every later day of the order. Every lead-time figure on this page is therefore a property of each day's reading in isolation.
It does not make the state file safe to share. One writer, one file, replaced atomically -- two schedulers on one order would race and the loser's day would vanish.
It does not validate the answer key. The key is arithmetic, and rule 3's boundary -- 'reaches BEYOND the committed date' -- is a reading. Every arm agreed with it on all 24 commitments landing exactly ON the date, which is agreement rather than validation.
WatchedWhat is watched, and why that one
6runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 30 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
10 measured by the latest run20 need the model half
Metric
Owner
Role
Why this one
lead-time-before-the-miss
Was each slipped order surfaced before its committed date, and by how many working days
alarm
caught_before_the_miss_pct; median_lead_days; total_lead_days — alarm on any slipped order dropping out of the caught set, and any lead day lost. Both are single-order events on a denominator of 30, so the row to read is the by-trajectory table rather than the rate.
cell-exact-match
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key
alarm
status_accuracy_pct; rule_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A snapshot that returns nothing scores as a miss in all three fields, so a reliability failure arrives disguised as a quality failure. This run had 1 of 240.
false-jeopardy-on-quiet-snapshots
Snapshots whose correct status is not JEOPARDY that got one anyway
alarm
false_jeopardy_rate_pct; false_jeopardy — alarm on the first false jeopardy on any arm. At zero, a single one is a new failure mode rather than a worse rate -- and the row to read is which distractor note it fired on.
paraphrase-probe
The same forward commitment, rewritten in words the generator never produced
alarm
status_accuracy_pct — alarm on the model dropping below 100 pct here, which would mean its advantage over the pattern is smaller than this probe says.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
240
different corpus — nothing is comparable
corpus.bytes
758,725
order snapshots edited — the count held, the bytes did not
split.count
48
the orders count moved — a different set was scored
split.size_p50
5
the median size of one order moved
split.size_p95
5
the 95th-percentile size of one order moved
dataset.rows
240
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
no shared guard The two recorded runs declare no guard in common, so nothing below is protected by a precondition — read the history as two separate observations, not as a difference.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
orders surfaced before the miss
not yet known
30 orders that ultimately slipped
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. r001 against the three free floors is a comparison, not a band.
lead days won
0 -- at the policy ceiling, and it cannot go higher
120 lead days across 30 slipped orders
120 of 120. The model raised every slipped order on the exact surveillance day the written policy makes available, so this figure is saturated by construction: there is no headroom above it on this corpus. A repeat can only move it down.
the lead-time distribution
median 4.0 working days, and the tail runs down to 1
30 slipped orders, working days between the first JEOPARDY and the committed date
⚑ THE MEDIAN HIDES THE ONLY THING AN OPERATIONS DESK CARES ABOUT, WHICH IS WHY BOTH ENDS ARE BANDED HERE RATHER THAN JUST THE TOTAL. Lead times across the 30 slipped orders were: 1 at 1 working day, 6 at 2 working days, 6 at 3 working days, 6 at 4 working days, 5 at 5 working days, 3 at 6 working days, 2 at 7 working days, 1 at 8 working days. The median is 4.0 -- but 7 of the 30 orders were surfaced with TWO OR FEWER working days left, and one, ORJ-0036 (an AGING order first raised on D5), with a single day. A page that arrives one working day before the committed date is arithmetically a catch and operationally too late to rebook an engineer or re-plan a civils gang. ⚠︎ AND THAT TAIL IS NOT THE MODEL'S FAILURE, IT IS THE POLICY'S CEILING: the model matched the written policy exactly on all 30 orders (120 of 120 lead days), so every short lead here is a day the RULES did not make available rather than a day the model lost. The free aging floor's distribution is worse in the same place and for a different reason: 3 at 1 working day, 4 at 2 working days, 5 at 3 working days, 3 at 4 working days, 2 at 5 working days, 2 at 6 working days, 1 at 7 working days, 1 at 8 working days. 3 of its catches sit at a single working day against the model's 1.
status accuracy
not yet known, and partly noise
240 snapshots
98.75 pct -- 237 of 240. Three of the 6 miss cells are one lost reply and the other three are two snapshots. ⚠︎ ONE OF THOSE TWO DID NOT REPRODUCE when the same snapshot was re-asked through the UI, so some of this residual is variance and no repeat run exists to say how much.
which rule was named
99.58 pct -- 239 of 240, and NOT ONE of the misses is a wrong rule
240 snapshots
evals/scoring.py::score, r001. ⚑ THIS IS A STRONGER RESULT THAN THE STATUS FIGURE BESIDE IT, AND THE DIFFERENCE IS WORTH READING. The single miss is ORJ-0008-D3, the reply lost at the output ceiling, which the scorer counts as a miss in every field. On all 239 snapshots the model actually answered it named the correct policy rule -- including every LEADTIME_EXCEEDS_REMAINING and every ESCALATED_SUSTAINED. Even the two snapshots whose STATUS it got wrong carry gold rule NONE, and it answered NONE on both: the error there was a status band, never a misapplied rule. The free arms sit at 69.58 pct (task ages, no memory), 86.25 pct (with memory -- the only thing memory buys on this kit) and 100.00 pct (with a pattern over the notes).
the top blocker named
99.17 pct -- and every free arm scores 100 pct, so it separates nothing
240 snapshots
evals/scoring.py::score, r001. ⚠︎ THE MODEL IS THE WORST ARM ON THIS GRADER AND THE PAGE SAYS SO RATHER THAN AVERAGING IT IN. Ranking open tasks by working days past due, falling back to the nearest due date, is sorting a table by two integers -- so the aging floor, the floor with memory and the pattern arm all score 240 of 240, and the model scores 238. Its two misses are ORJ-0008-D3 (the lost reply) and ORJ-0025-D3, where the invented hold that produced that snapshot's WATCH also moved the ranking onto the wrong task -- ONE error, TWO cells. A grader a do-nothing baseline aces is measuring that the table is legible, not that the monitor works. The comparison is exact text, case-insensitive on trimmed text and nothing looser, because a fuzzy match would forgive an invented task, which is the only thing this field exists to catch.
false jeopardy
0 -- and 0 on every other arm too, so it separates nothing
143 quiet snapshots
0 of 143 quiet snapshots on the model, the aging floor, the floor with memory and the pattern arm. A grader every arm aces at zero is not discriminating; it is recording that this corpus does not tempt any of them. A rate of 0 on 143 bounds the true rate below roughly 2 pct and says nothing finer.
the free pattern arm under paraphrase
45.83 points, and it is a measurement rather than a spread
24 rewritten snapshots
100 pct on the corpus as generated, 54.17 pct on the same facts restated in unseen words -- 1 of the 12 rule-3 commitments read, against the model's 12. Two runs of one arm over two wordings of one corpus.
the denominator itself
0 -- a constant of the corpus, not a result
240 snapshots
240 snapshots, 30 slipped orders, 18 on-time, 143 quiet snapshots and 120 lead days available, identical on every arm. Which orders slip is decided by tools/build_corpus.py before any model sees anything.
input volume
0 -- prompt assembly is pure code and model-independent
240 calls
273043 input tokens over 240 calls, 1137.68 per snapshot. The prompt is built by string concatenation from the corpus and four carried fields, so this figure cannot move without the corpus or src/prompt.py moving.
output volume
not yet known
240 calls
312941 tokens over 240 calls, 1303.92 average, spread 222 to the 12000 ceiling. One recorded run.
provider-side reasoning share
not yet known
240 calls
94.87 pct of output on the scored run (296887 of 312941), left at the provider's default. A share within one run is not a band between runs.
latency
not yet known
240 calls
p50 8972 ms and p95 24244 ms; the slowest single call took 62480 ms. One recorded run, and the tail is 2.7 times the median.
replies that did not parse
not yet known
240 calls
1 of 240 at the published 12000-token ceiling, and 0 of 10 on the calibration at 8,000. Two points at two configurations, which is a reason to watch the ceiling rather than a band -- especially since the call that blew it was one of the quietest snapshots in the corpus.
coming back down after a recovery
100 pct -- 18 of 18, and it separates nothing
18 post-recovery snapshots on the RECOVER trajectory
evals/scoring.py::score, r001: the three surveillance days after a correctly-raised order's blocking task closes. ⚠︎ ALL FOUR ARMS SCORE 18 OF 18, including the aging floor, which comes down because its threshold stops being crossed and not because it understood anything. A subset every arm aces cannot tell a monitor that knows when to stop shouting from one that never started.
HistoryRun history
6 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
stub · with the model in the path — 1 run. Columns here are only ever compared with each other.
Metric
a000-order-jeopardy-paraphrase-free 2026-08-23
regex status accuracy, %
54.17
regex status correct
13
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 2 chips that all say so.
ablation · with the model in the path — 1 run. Columns here are only ever compared with each other.
Metric
a001-order-jeopardy-paraphrase 2026-08-23
regex status accuracy, %
54.17
regex status correct
13
input tokens, whole run
27512
output tokens, whole run
42217
status accuracy, %
100.0
status correct
24
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 6 chips that all say so.
triage · with the model in the path — 4 runs. Columns here are only ever compared with each other.
Metric
b000-order-jeopardy-aging 2026-08-23
b001-order-jeopardy-agingmem 2026-08-23
b002-order-jeopardy-regex 2026-08-23
r001-order-jeopardy 2026-08-23
caught before the miss, %
70.0
70.0
100.0
100.0
input tokens, whole run
0
0
0
273043
model latency p50 ms
0.00
0.00
0.00
8972.00
model latency p95 ms
0.00
0.00
0.00
24244.00
median lead days
3.0
3.0
4.0
4.0
min lead days
1
1
1
1
output tokens, whole run
0
0
0
312941
rule accuracy, %
69.58
86.25
100.00
99.58
status accuracy, %
86.25
86.25
100.00
98.75
top blocker accuracy, %
100.00
100.00
100.00
99.17
not a time series No two of these 4 runs measured the same system — they differ on documents, max_tokens, provider, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 6 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the arm may read the Activity Log at all
orders surfaced before the miss 21 of 30 -> 30 of 30, lead days won 75 -> 120, status accuracy 86.25 pct -> 98.75 pct
measured
r001-order-jeopardy against b000-order-jeopardy-aging. Same corpus, same answer key, same scorer, same carried state; the only difference is that one arm opens the Activity Log.
whether the arm has memory
nothing on the headline. Orders surfaced 21 -> 21, lead days 75 -> 75, status accuracy 86.25 -> 86.25 pct. The RULE named moves 69.58 -> 86.25 pct and that is all
measured
b001-order-jeopardy-agingmem against b000-order-jeopardy-aging, compared cell by cell on these two run files. Rule 4 renames a jeopardy that is already raised; it cannot raise one.
whether a model is called at all
nothing on this corpus, and the model is slightly WORSE. The free pattern arm ties on every lead-time figure (30 of 30 orders, 120 lead days) and beats the model on the cells: 100 pct against 98.75 pct status. Cost $0.00 against $0.004481 a snapshot
measured
b002-order-jeopardy-regex against r001-order-jeopardy on the identical 240 snapshots.
restating the same commitment in words the corpus generator never produced
the free pattern arm 100 pct -> 54.17 pct status, and 12 of 12 rule-3 commitments read -> 1 of 12. The model does not move: 100.0 pct, 12 of 12
measured
results/ablation-a001-order-jeopardy-paraphrase.json. 24 snapshots, one note rewritten on each, the answer key unchanged. ⚠︎ The paraphrases and the pattern share an author, so this is an upper bound on brittleness.
which trajectory the order is on
for the free aging floor, from 6 of 6 caught on AGING, BLOCKER and SUSTAINED to 3 of 6 on LEADTIME and 0 of 6 on LEADTIME_LATE. For the model, 6 of 6 on all five
measured
results/eval-r001-order-jeopardy.json and eval-b000-order-jeopardy-aging.json, by_pattern. The headline is a property of the trajectory mix, and six orders per trajectory was a design decision rather than a sample of anything.
the published output ceiling
0 of 10 replies lost at 8,000 on the calibration, 1 of 240 at the published 12000
measured
eval-c000-order-jeopardy-calibration (cap 8,000, largest reply 1,845 tokens over the two hardest orders) against eval-r001-order-jeopardy (cap 12000, one reply lost on one of the quietest snapshots in the corpus).
running free code first and sending a model only what it calls CLEAR
unknown -- it would cut most of the bill and it might cut lead time; neither has been measured
reasoning
The free aging floor agrees with the model on 86.25 pct of snapshots at $0.00, and every order it misses is one it called CLEAR. The experiment is free to design and costs one more run to score, and it has not been fired.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
orders surfaced before the miss
nothing yet
lead days won
any drop at all -- and read WHICH order lost days, not the total
the lead-time distribution
the median dropping below 4.0, or ANY order joining the one at a single working day. The second is the one to watch -- it moves by one order, and it is the difference between a warning somebody can act on and a notification
status accuracy
nothing yet
which rule was named
any miss that is NOT a lost reply. A wrong rule on an answered snapshot would be a new failure mode rather than a worse rate, because this run has none
the top blocker named
any drop below the free arms' 100 pct that is not a lost reply. Nothing but one lost reply and one knock-on error has ever moved it
false jeopardy
the first false jeopardy on any arm
the free pattern arm under paraphrase
any change to evals/baseline.py's patterns -- extending them to cover the paraphrases would move this number and settle nothing
the denominator itself
any change at all, on any run
input volume
any change without a corresponding change to the prompt or the corpus
output volume
nothing yet
provider-side reasoning share
nothing yet
latency
nothing yet
replies that did not parse
any second lost reply
coming back down after a recovery
on any JEOPARDY returned on a post-recovery snapshot -- one is enough
NextThe three you would add first
A human step in front of anything that touches an orderThe kit's two content errors were both a narrative note read as a change of task state, and one of the two did not reproduce. Rebooking an engineer or telling a customer their date has moved are the acts a jeopardy verdict leads to, and neither should happen without somebody reading the rule the pack named.
A refusal state for a reply that did not parseThis run lost one snapshot to the output ceiling and it scored as three wrong cells. In a deployment that day needs to be visibly UNANSWERED rather than silently CLEAR -- and on a jeopardy pack a silent CLEAR is the worst possible failure, because it is indistinguishable from good news.
Persistence for the rule-3 fact, with the reading that produced itThe eval takes it from the answer key. A deployment cannot, so it must record what it read and when, or the sustained-jeopardy rule has nothing to stand on and a single bad day's reading silently rewrites the rest of the order's history.
A cheap-first routerThe free aging floor answers 86.25 pct of snapshots identically to the model for $0.00. Sending a model only the orders free code calls CLEAR would cut the bill by most of it; this kit has not measured what it would cost in lead time.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/jeopardy.py, src/monitor.py, src/segment.py or src/select.py -- it refuses to let a run spend if the corpus, the answer key, the parser or the privacy guard has moved. Re-run all three free floors together (free) whenever the policy changes, because the headline is a DIFFERENCE and one arm re-run alone is not comparable with another's old figure. Re-run the scored eval (paid, one call per snapshot) only on a change to src/prompt.py or src/state.py.
What this cannot tell you
One run per arm. Whether the model repeats 98.75 pct status is not measured, and one of its two content errors is KNOWN not to repeat -- the same snapshot answered correctly when re-asked through the UI.
Whether the no-propagation property would hold in a deployment. It is true by construction here because the state is advanced from the answer key, which a deployment cannot do. Nothing on this page measures what a wrong rule-3 reading does to the days that follow it.
Whether a cheap-first router would keep the lead time. It is the obvious next experiment, it is free to design and costs one more run, and it has not been fired.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, zlib, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it. The one optional dependency is not Python -- tools/shoot_ui.mjs needs Node and puppeteer to re-take the committed screenshots.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a VERDICT -- four scalar fields written by arithmetic. The whole reason the cost is flat in window length is that nothing here remembers what was said, only what was decided. And on this kit the free floors measured memory as worth 16.67 points of rule accuracy and nothing else, so a memory layer would be paying a framework's price for a figure that does not move
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the order chain
evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect)
this is the seam where a framework would genuinely earn its place, and the kit says so: 48 independent chains of 5 strictly-ordered steps, with cadence, retries and late-arriving days all outside the kit. A ThreadPoolExecutor is the right size for an eval and the wrong size for a surveillance pack that actually runs every morning
the routing decision
evals/baseline.py
model routers and cascade libraries (RouteLLM, semantic routers)
the honest use of a router here is not cheap-model-then-expensive-model, it is NO-MODEL-then-model: free code agrees with the model on 86.25 pct of snapshots. A router library would give somewhere to put that policy, and this kit has not measured what it costs in lead time
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
three exact-match comparisons and a replay over one column is a dict comprehension, not a platform -- and the lead-time replay is the part no general harness has a primitive for, because it scores a SEQUENCE of verdicts rather than each verdict
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each order is a chain of five steps -- one surveillance day each -- with no branching and exactly one edge between consecutive days, carrying four fields. Different orders never touch. A framework would add an orchestrator to a for-loop that already runs 48 chains wide.
The other sideWhat a framework costs you
No scheduler. This is a monitor that carries state and is still INVOKED, not woken -- the cadence, the retry policy and the question of what to do when a surveillance day does not arrive are all outside the kit.
No state store worth the name. data/state.json is one file replaced atomically; a framework would bring a checkpointer with a concurrency model, which this does not have.
No routing. Every snapshot goes to the model, and this kit's own free floors say most of them need not.
No observability beyond what evals/run.py prints and writes to results/. There is no tracing and no dashboard integration.
No retry/backoff beyond src/adapters' own bounded retry, which covers a busy provider and a dropped connection and nothing else -- and notably NOT a reply that runs away to the ceiling, which is the failure this run actually had.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-order-jeopardy on the fast tier, with the carried state, 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
8,972 ms
not yet known
nothing yet
Model, p95
24,244 ms
not yet known
nothing yet
Input tokens
273,043
0 -- prompt assembly is pure code and model-independent
any change without a corresponding change to the prompt or the corpus
Output tokens
312,941
not yet known
nothing yet
No movement column. Not one of the 3 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
No drift chart. Only one run has recorded a call latency; the other 3 runs on record made no model call at all — a rules-only floor stores 0, which is an absence, not a zero-millisecond answer. One timed run is a reading, not a trend, and it is in the table above. A chart appears the day a second run times itself.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — one period report goes whole into one call, and the prior period's carried state goes with it.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 6 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
order snapshots
data/corpus/<ORDER>-D<n>.txt -- 240 files, 758725 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Customer Contact never does, by src/select.NEVER_SENT
the answer key
data/gold.jsonl -- 240 rows, the output of src/jeopardy.replay over the corpus values, never hand-authored. It also carries the OUTCOME of every order -- which slipped and by how much -- which appears on no snapshot
never -- evals/scoring.py is pure code, no model and no key
the carried state
data/state.json in a deployment (src/state.py, written atomically). ⚠︎ IN AN EVAL IT IS SCOPED TO THE RUN AND NEVER TOUCHES DISK -- evals/run.py builds a fresh in-memory store per order, because a run that started from the previous run's memory could not be re-run or compared with the free floors
one 126-character English sentence per call, produced by state.describe -- four scalar fields, never a prior snapshot and never a prior reply
the recorded runs
results/eval-*.json and results/ablation-*.json -- the scored run, three free floors, the calibration, the wiring stub and both halves of the paraphrase probe, all committed
never -- they are read by the page and by nothing that calls a provider
the key
.env or the shared repo-root .env -- never committed, 0600
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py). src/app.py redacts it out of any provider error before returning one to the browser
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after. src/app.py redacts the key and the base URL out of any provider error before returning it to the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
windowing
the scoreable unit, made before anything judges it. 240 snapshots become 48 chains of five consecutive surveillance days, processed strictly in order by evals/run.py because day 3's prompt contains state produced by day 2. Different orders are independent and run concurrently, 48 chains wide. ⚠︎ NOTHING IS CUT FROM A STREAM HERE -- the corpus arrives as discrete daily snapshots and the 'window' is simply which five belong to one order. The decision keeps the variant's name because the contract requires it; what it actually decides is the chain.
48 chains x 5 days = 240 calls, 0 gaps and 0 out-of-order days; evals/check_labels.py asserts all 48 sequences are complete before a run may spend. Wall clock 280.1 s at 12 workers against 240 serial calls (r001-order-jeopardy, evals/check_labels.py)
five surveillance days. Every rule in this policy resolves inside five, so nothing here measures a jeopardy that has been running for three weeks -- and a longer window is exactly where the sustained rule and the state store would start to matter
a different TRAJECTORY MIX invalidates the headline rather than extending it. Six orders per trajectory was a design decision, and the lead-time figures are a property of it: a book with more slow-aging orders and fewer supplier-quote failures closes the gap between the model and the free floor without anything about either arm having changed
state
four scalar fields per order -- the status reported yesterday, the rule applied, consecutive days in jeopardy, and the day it was first raised -- written by src/jeopardy.step from the parsed task ages and the answer key's rule-3 fact, never from the model's reply, and rendered by src/state.describe into one English sentence. That single sentence is the entire route from one day to the next.
126 characters on the worked example, about 32 input tokens of the 1137.68 a snapshot -- 2.8 pct of input. ⚑ AND ITS MEASURED VALUE IS ZERO ON THE HEADLINE: the free floor WITH memory and the free floor WITHOUT it caught the same 21 orders with the same 75 lead days, differing only by 16.67 points of rule accuracy (b001-order-jeopardy-agingmem against b000-order-jeopardy-aging)
the state is four scalars, so day 300 costs what day 3 costs -- the OPPOSITE curve to an intake kit. What is NOT bounded is the store itself: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model
⚠︎ AND THE SOURCE OF THE STATE IS THE LIMITATION THIS PAGE REPEATS MOST. It is advanced from the ANSWER KEY, because rule 3's fact lives only in prose and nothing in src/ may parse it. That makes every lead time a property of each day's reading in isolation, and it means no figure here measures what a deployment's own wrong reading would do to the days after it
model
one call per ORDER-DAY carrying the instruction, the carried-state sentence and six of the snapshot's seven sections, behind src/adapters/__init__.py, at the published MAX_TOKENS = 12000 with provider-side reasoning left at the default
239 of 240 replies parsed; the one that did not produced the full 12000 output tokens and an empty body. p50 8972 ms, p95 24244 ms, slowest call 62480 ms (lenses.Eval.scores, r001-order-jeopardy)
the 12000-token ceiling was set at 6.5 times the largest reply a 10-call calibration on the two HARDEST orders had seen (1,845 tokens), and one call went over it anyway -- on one of the QUIETEST snapshots in the corpus. Output length here does not track difficulty, so calibrating on the hardest cases was necessary and was not sufficient
one tier was run, so nothing here compares two models; the four arms on this page differ by what they may READ and by whether a model was called at all. Point .env at your own model and the free scorer re-runs on your numbers
labels
data/gold.jsonl, 240 rows, computed by src/jeopardy.replay over the corpus values at generation time -- a function the runtime never calls -- plus the outcome record (which orders slipped, and by how many working days) that makes the lead-time backtest possible and appears on no page
97 snapshots gold JEOPARDY, 47 WATCH, 96 CLEAR; 30 orders slipped and 18 delivered on time; 120 lead days available across the slipped orders. Rule codes: 143 NONE, 18 MILESTONE_OVERDUE, 6 BLOCKER_AGED, 12 LEADTIME_EXCEEDS_REMAINING, 61 ESCALATED_SUSTAINED (lenses.Eval.dataset, order-jeopardy-2026-08-23-48orders-240snapshots)
the key is only as good as its reading of rule 3's boundary -- 'reaches BEYOND the committed date'. The corpus places 24 commitments exactly ON the date, where the key says the rule does not fire, and every arm agreed with it on all 24. That is agreement, not validation
your own order book: hand-label the gold, which is the real work -- and unlike most kits you must also record the OUTCOME of every order, because without it there is no lead time to measure. This kit's key is a luxury of controlling both the generator and the policy
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
an order reported CLEAR on every surveillance day that then misses its committed date
nothing on its task table ever went past a threshold, and the evidence was in the Activity Log. This is the free aging floor's whole failure mode and it is silent: every verdict it gave was arithmetically correct
read the notes on the days it called CLEAR, looking for a commitment -- a lead time, a rebooked date, a consent window -- rather than for lateness. 9 of the 30 slipped orders in this corpus are that shape (results/eval-b000-order-jeopardy-aging.json, leads: the LEADTIME_LATE orders show arm_first_day null against a gold_first_day of 4)
WATCH on a snapshot where the Open Tasks table shows nothing past due and nothing held
the model read a narrative note -- a lost day, a failed visit, weather -- as a change of TASK STATE. It is the mirror of the floor's failure: the floor cannot read prose, and the model reads state into it
compare the rationale against the task table before disputing the verdict. Two snapshots of 240 in this run, both in the harmless direction, and ONE OF THE TWO DID NOT REPRODUCE when the same snapshot was re-asked (results/eval-r001-order-jeopardy.json misses -- ORJ-0025-D3 and ORJ-0033-D3, and docs/shots/verdict-second-ask.png for the re-ask)
a surveillance day that returned nothing at all
the reply ran to the output ceiling with the whole budget spent on reasoning. It scores as three wrong cells, so a reliability failure arrives disguised as a quality failure
check output_tokens against src/monitor.MAX_TOKENS and finish_reason before blaming the prompt. ⚠︎ DO NOT ASSUME IT WAS A HARD SNAPSHOT: the one this run lost was among the quietest in the corpus, one task and two irrelevant notes (results/eval-r001-order-jeopardy.json, failures -- ORJ-0008-D3, 12000 output tokens, finish_reason 'length', empty body)
No machine symptom — this failure leaves no trace in any output.
There is none, and it is the failure this whole kit is built to name: an order reported CLEAR on every day of its window looks exactly like an order that is fine. A deployment that quietly lost its ability to read the Activity Log -- a changed section heading, a feed that stopped carrying notes, a selector edit -- would keep answering, keep parsing, keep costing money and keep scoring 86.25 pct on status, with every reply well-formed and nothing anywhere raising an error. The only symptom is orders missing their dates having never been raised, which is a statistic nobody looks at until a customer escalates. No check in this kit compares an arm's raises against the outcome at run time; the lead-time grader does it afterwards, on a corpus whose outcomes are known.
⚠︎ THIS KIT'S FLOW VARIANT IS 'triage' AND IT IS BORROWED, AS IT IS FOR EVERY MONITOR IN THIS SERIES. 'monitor' is a pattern in build/domain/readiness.py with no flow variant of its own, so this kit borrows 'triage' and inherits two station labels that do not describe it -- 'Cut into windows' over a thing that is not a window, and 'Page threshold' over arithmetic that pages nobody. They are used unchanged rather than reworded, because inventing station words silently breaks lens reachability while a wrong-sounding label with an honest caption under it is merely visible. ⚑ AND index_answer IS CARRIED VERBATIM FROM THE CLOSED LIST AND IS NOT QUITE TRUE HERE: this kit's unit is a daily order SNAPSHOT rather than a 'period report', and the wording was taken from the list rather than invented because the standard requires it verbatim. The mismatch is reported rather than papered over; a fourth wording naming a snapshot would fit every monitor in this series. Concurrency and GPU sizing -- 48 order chains, 12 workers, nothing measured past 240 calls. Whether days within an order can be run out of order -- they cannot, by construction, and no throughput figure accounts for that at longer windows. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- every figure prices the cache-miss rate for exactly this reason, and the fixed head of this prompt (3237 of 4847 characters, byte-identical on all 240 calls) is precisely what a cache would have discounted. Whether any arm repeats -- every configuration ran once.
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every order, product, task, supplier commitment, activity note and name is invented here. Every contact number is in the 07700 900000-900999 range reserved by Ofcom for fiction and never allocated to a subscriber. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Was each slipped order surfaced before its committed date, and by how many working days
Catch a telecom order before it misses its date
PresenterOpens the private repo. Visible to admins only.
In one lineWas each slipped order surfaced before its committed date, and by how many working days
For each of the 30 orders that ultimately missed its committed date, on which surveillance day did this arm first say JEOPARDY, and how many working days was that before the date? An arm that never said it missed the order entirely.
$0.00per 1,000 order snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function all three free floors are scored through.
Every grader on these pages scored the same 240 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The snapshot
ORJ-0007-D2 -- Ethernet leased line, 100 Mbit/s, small business, single site
Committed date and days remaining
Mon 09 Mar 2026, working day 6 of the window. 4 working days remain on this snapshot
Open Tasks, in full
Fibre splice, Field engineering, due in 3 working days. Nothing is past due and nothing is held
The line in the Activity Log that decides it
"The gang that covers fibre splice is booked for Tue 10 Mar 2026; there is no earlier slot."
Carried state, in the prompt
On the previous surveillance day this order was reported CLEAR. It has not been in jeopardy on any earlier day of this window.
The free aging floor
CLEAR / NONE -- correct task arithmetic, and no rule, because it never opens the Activity Log
The model
JEOPARDY / LEADTIME_EXCEEDS_REMAINING -- "Activity Log shows the fibre splice gang is booked for Tue 10 Mar 2026, beyond the committed date of Mon 09 Mar 2026, so rule 3 applies"
Ground truth
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, and the order went on to miss by 8 working days
Scored as
3 of 3 cells correct, and 4 working days of lead won against the aging floor's 1
Grader
Verdict
Why
Was each slipped order surfaced before its committed date, and by how many working days
raised on D2 -- 4 working days before the committed date, the earliest the policy allows
ORJ-0007 missed its committed date by 8 working days. On D2 nothing on it is overdue: the fibre splice is due in three working days and four remain. The free aging floor answers CLEAR here and does not raise this order until D5, with one working day left. The model raised it on D2 off one line in the Activity Log.
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key
status hit, rule hit, top blocker hit
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, all three matching the computed key. The rationale does the comparison rule 3 is made of and names both dates.
Snapshots whose correct status is not JEOPARDY that got one anyway
not in scope
This grader scores only the 143 snapshots whose correct status is CLEAR or WATCH. This one is correctly JEOPARDY, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The same forward commitment, rewritten in words the generator never produced
in scope, and both arms were re-asked
ORJ-0007-D2 is one of the 12 rule-3 commitments the paraphrase probe rewrites. With the note restated as "No slot until Tue 10 Mar 2026. The regional plan is full until then." the model still answered JEOPARDY; the free pattern answered CLEAR, because none of its patterns matches that sentence.
The formulaWhat it computes
For each slipped order, the first day where the arm's own answer is JEOPARDY; lead = committed working day - that day. Caught = the count with a lead of at least 1. The denominator is ORDERS, not snapshots.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% caught before the miss · 3 more measured on this row
the free aging floor, no model
70.0% caught before the miss · 3 more measured on this row
the free aging floor, with memory
70.0% caught before the miss · 3 more measured on this row
the free aging floor + a pattern
100.0% caught before the miss · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl's outcome record -- which orders slipped, computed at generation time and printed on no snapshot. The lead time itself is read off each arm's OWN answers, so this grader compares an arm to the world rather than to another arm.
These rates are UNKNOWN, on purpose
What the RIGHT lead time is. The gold lead is the policy's own -- the earliest day the written rules could fire on the evidence present -- not the earliest day a human being could have known. On several LEADTIME orders a coordinator reading the note would have called it before the policy does, and nothing here measures that ceiling. So 'the model won all 120 available lead days' means it matched the POLICY, and the policy is a floor on what was knowable.
Watch these
caught_before_the_miss_pct
median_lead_days
total_lead_days
Alarm on
any slipped order dropping out of the caught set, and any lead day lost. Both are single-order events on a denominator of 30, so the row to read is the by-trajectory table rather than the rate.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. 30 orders means the finest honest band is about 3.3 points.
Cadence: Re-run the free floors (free) on any change to tools/build_corpus.py, src/jeopardy.py, src/segment.py or src/select.py. Re-run the scored eval (paid, one call per snapshot) on any change to src/prompt.py or src/state.py.
The decisionWhen to reach for it
Use it
Always, and first. It is the only figure that answers the question a monitor exists to answer.
Do not use it
As a measure of caution. It ignores the on-time orders entirely, so an arm that reported JEOPARDY on everything would score 100 pct here -- which is why the false jeopardy grader is printed beside it and never under it.
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key
Catch a telecom order before it misses its date
PresenterOpens the private repo. Visible to admins only.
In one lineThe status, the rule and the top blocker, per snapshot, exact match against the computed answer key
For each of the 240 snapshots and each of the three answered fields, did the reply equal the computed answer key? Status and rule are compared exactly; the task name is compared on trimmed lower-cased text and nothing looser, because a fuzzy match would forgive an arm that invented a plausible task.
$0.00per 1,000 order snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the lead-time grader.
Every grader on these pages scored the same 240 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The snapshot
ORJ-0007-D2 -- Ethernet leased line, 100 Mbit/s, small business, single site
Committed date and days remaining
Mon 09 Mar 2026, working day 6 of the window. 4 working days remain on this snapshot
Open Tasks, in full
Fibre splice, Field engineering, due in 3 working days. Nothing is past due and nothing is held
The line in the Activity Log that decides it
"The gang that covers fibre splice is booked for Tue 10 Mar 2026; there is no earlier slot."
Carried state, in the prompt
On the previous surveillance day this order was reported CLEAR. It has not been in jeopardy on any earlier day of this window.
The free aging floor
CLEAR / NONE -- correct task arithmetic, and no rule, because it never opens the Activity Log
The model
JEOPARDY / LEADTIME_EXCEEDS_REMAINING -- "Activity Log shows the fibre splice gang is booked for Tue 10 Mar 2026, beyond the committed date of Mon 09 Mar 2026, so rule 3 applies"
Ground truth
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, and the order went on to miss by 8 working days
Scored as
3 of 3 cells correct, and 4 working days of lead won against the aging floor's 1
Grader
Verdict
Why
Was each slipped order surfaced before its committed date, and by how many working days
raised on D2 -- 4 working days before the committed date, the earliest the policy allows
ORJ-0007 missed its committed date by 8 working days. On D2 nothing on it is overdue: the fibre splice is due in three working days and four remain. The free aging floor answers CLEAR here and does not raise this order until D5, with one working day left. The model raised it on D2 off one line in the Activity Log.
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key
status hit, rule hit, top blocker hit
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, all three matching the computed key. The rationale does the comparison rule 3 is made of and names both dates.
Snapshots whose correct status is not JEOPARDY that got one anyway
not in scope
This grader scores only the 143 snapshots whose correct status is CLEAR or WATCH. This one is correctly JEOPARDY, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The same forward commitment, rewritten in words the generator never produced
in scope, and both arms were re-asked
ORJ-0007-D2 is one of the 12 rule-3 commitments the paraphrase probe rewrites. With the note restated as "No slot until Tue 10 Mar 2026. The regional plan is full until then." the model still answered JEOPARDY; the free pattern answered CLEAR, because none of its patterns matches that sentence.
The formulaWhat it computes
accuracy = hits / 240 per field. A snapshot whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
98.8% status accuracy · 2 more measured on this row
the free aging floor, no model
86.2% status accuracy · 2 more measured on this row
the free aging floor, with memory
86.2% status accuracy · 2 more measured on this row
the free aging floor + a pattern
100.0% status accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/jeopardy.replay over the corpus values at generation time. This grader IS the reference, so its own error rate is not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate cannot be measured: it IS the reference. What CAN be wrong is the key's reading of rule 3's boundary -- 'reaches BEYOND the committed date' -- and the corpus places 24 commitments exactly ON the date, where the key says the rule does not fire. Every arm agreed with the key on all 24, so nothing in this run disputes the reading; that is agreement, not validation.
Watch these
status_accuracy_pct
rule_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A snapshot that returns nothing scores as a miss in all three fields, so a reliability failure arrives disguised as a quality failure. This run had 1 of 240.
How tight can the band be? 720 cells, so one cell is 0.139 points. Nothing here is tuned; there is no threshold to sweep.
Cadence: Whenever the lead-time grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Beside the lead time, to say whether an arm that surfaced the right orders also described them correctly.
Do not use it
As the headline. An arm can be wrong on four snapshots of an order and still have surfaced it eight working days early, and an arm can be right on four and have surfaced it one day before the date. This kit is about the second number.
Snapshots whose correct status is not JEOPARDY that got one anyway
Catch a telecom order before it misses its date
PresenterOpens the private repo. Visible to admins only.
In one lineSnapshots whose correct status is not JEOPARDY that got one anyway
Of the 143 snapshots whose correct status is CLEAR or WATCH, how many were raised to JEOPARDY? This is the price of being early, and it is the direction that costs a coordinator their attention.
$0.00per 1,000 order snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 240 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The snapshot
ORJ-0007-D2 -- Ethernet leased line, 100 Mbit/s, small business, single site
Committed date and days remaining
Mon 09 Mar 2026, working day 6 of the window. 4 working days remain on this snapshot
Open Tasks, in full
Fibre splice, Field engineering, due in 3 working days. Nothing is past due and nothing is held
The line in the Activity Log that decides it
"The gang that covers fibre splice is booked for Tue 10 Mar 2026; there is no earlier slot."
Carried state, in the prompt
On the previous surveillance day this order was reported CLEAR. It has not been in jeopardy on any earlier day of this window.
The free aging floor
CLEAR / NONE -- correct task arithmetic, and no rule, because it never opens the Activity Log
The model
JEOPARDY / LEADTIME_EXCEEDS_REMAINING -- "Activity Log shows the fibre splice gang is booked for Tue 10 Mar 2026, beyond the committed date of Mon 09 Mar 2026, so rule 3 applies"
Ground truth
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, and the order went on to miss by 8 working days
Scored as
3 of 3 cells correct, and 4 working days of lead won against the aging floor's 1
Grader
Verdict
Why
Was each slipped order surfaced before its committed date, and by how many working days
raised on D2 -- 4 working days before the committed date, the earliest the policy allows
ORJ-0007 missed its committed date by 8 working days. On D2 nothing on it is overdue: the fibre splice is due in three working days and four remain. The free aging floor answers CLEAR here and does not raise this order until D5, with one working day left. The model raised it on D2 off one line in the Activity Log.
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key
status hit, rule hit, top blocker hit
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, all three matching the computed key. The rationale does the comparison rule 3 is made of and names both dates.
Snapshots whose correct status is not JEOPARDY that got one anyway
not in scope
This grader scores only the 143 snapshots whose correct status is CLEAR or WATCH. This one is correctly JEOPARDY, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The same forward commitment, rewritten in words the generator never produced
in scope, and both arms were re-asked
ORJ-0007-D2 is one of the 12 rule-3 commitments the paraphrase probe rewrites. With the note restated as "No slot until Tue 10 Mar 2026. The regional plan is full until then." the model still answered JEOPARDY; the free pattern answered CLEAR, because none of its patterns matches that sentence.
The formulaWhat it computes
cells where gold status != JEOPARDY and the reply said JEOPARDY, divided by 143.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
0.0% false jeopardy rate
the free aging floor, no model
0.0% false jeopardy rate
the free aging floor + a pattern
0.0% false jeopardy rate
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to the 143 cells whose status is not JEOPARDY.
These rates are UNKNOWN, on purpose
Its own rate cannot be separated from the reference grader's, and with ZERO positives on every arm there is nothing here to estimate a rate FROM. Zero of 143 bounds the true rate below roughly 2 pct and says nothing finer.
Watch these
false_jeopardy_rate_pct
false_jeopardy
Alarm on
the first false jeopardy on any arm. At zero, a single one is a new failure mode rather than a worse rate -- and the row to read is which distractor note it fired on.
How tight can the band be? 143 rows, so one row is 0.70 points.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Always, and beside the lead time rather than under it. The two move in opposite directions and a single accuracy figure hides the trade entirely.
Do not use it
⚠︎ AS A DISCRIMINATOR ON THIS RUN, BECAUSE IT IS SATURATED AT ZERO ON ALL FOUR ARMS. Nothing here raised a single false jeopardy. That is a real result and it is also a grader that currently tells the arms apart not at all, which the page says rather than banking.
The same forward commitment, rewritten in words the generator never produced
Catch a telecom order before it misses its date
PresenterOpens the private repo. Visible to admins only.
In one lineThe same forward commitment, rewritten in words the generator never produced
When a supplier commitment is stated in a phrasing that appears nowhere in the corpus generator and that no pattern in evals/baseline.py matches, does the arm still reach the right status? This is the only grader here that separates reading the SENTENCE from matching the PATTERN.
$0.00per 1,000 order snapshots
yesdata leaves your network
nosame answer every time
MethodHow the test was run
evals/paraphrase.py. The free half re-scores the pattern arm for $0.00; the model half cost 24 live calls.
Every grader on these pages scored the same 240 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The snapshot
ORJ-0007-D2 -- Ethernet leased line, 100 Mbit/s, small business, single site
Committed date and days remaining
Mon 09 Mar 2026, working day 6 of the window. 4 working days remain on this snapshot
Open Tasks, in full
Fibre splice, Field engineering, due in 3 working days. Nothing is past due and nothing is held
The line in the Activity Log that decides it
"The gang that covers fibre splice is booked for Tue 10 Mar 2026; there is no earlier slot."
Carried state, in the prompt
On the previous surveillance day this order was reported CLEAR. It has not been in jeopardy on any earlier day of this window.
The free aging floor
CLEAR / NONE -- correct task arithmetic, and no rule, because it never opens the Activity Log
The model
JEOPARDY / LEADTIME_EXCEEDS_REMAINING -- "Activity Log shows the fibre splice gang is booked for Tue 10 Mar 2026, beyond the committed date of Mon 09 Mar 2026, so rule 3 applies"
Ground truth
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, and the order went on to miss by 8 working days
Scored as
3 of 3 cells correct, and 4 working days of lead won against the aging floor's 1
Grader
Verdict
Why
Was each slipped order surfaced before its committed date, and by how many working days
raised on D2 -- 4 working days before the committed date, the earliest the policy allows
ORJ-0007 missed its committed date by 8 working days. On D2 nothing on it is overdue: the fibre splice is due in three working days and four remain. The free aging floor answers CLEAR here and does not raise this order until D5, with one working day left. The model raised it on D2 off one line in the Activity Log.
The status, the rule and the top blocker, per snapshot, exact match against the computed answer key
status hit, rule hit, top blocker hit
JEOPARDY / LEADTIME_EXCEEDS_REMAINING / Fibre splice, all three matching the computed key. The rationale does the comparison rule 3 is made of and names both dates.
Snapshots whose correct status is not JEOPARDY that got one anyway
not in scope
This grader scores only the 143 snapshots whose correct status is CLEAR or WATCH. This one is correctly JEOPARDY, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The same forward commitment, rewritten in words the generator never produced
in scope, and both arms were re-asked
ORJ-0007-D2 is one of the 12 rule-3 commitments the paraphrase probe rewrites. With the note restated as "No slot until Tue 10 Mar 2026. The regional plan is full until then." the model still answered JEOPARDY; the free pattern answered CLEAR, because none of its patterns matches that sentence.
The formulaWhat it computes
status accuracy over 24 rewritten snapshots -- 12 whose commitment reaches past the committed date and 12 whose does not -- against the unchanged answer key.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% status accuracy
the free aging floor + a pattern
54.2% status accuracy
In operationWhat to monitor
Reference standard: the unchanged data/gold.jsonl. The rewrite states the identical fact, so no gold label moves and the two arms are scored on the same cells they were scored on before.
These rates are UNKNOWN, on purpose
Everything about real desk prose. Two phrasing families and 24 rows, written by the author of the thing they are testing. What it establishes is that the free arm's 100 pct is not robust to wording; what it cannot establish is how much worse either arm would be on notes nobody in this repository wrote.
Watch these
status_accuracy_pct
Alarm on
the model dropping below 100 pct here, which would mean its advantage over the pattern is smaller than this probe says.
How tight can the band be? 24 rows means one row is 4.17 points. Nothing is tuned.
Cadence: Re-run the free half on any change to evals/baseline.py's patterns. Re-run the paid half only when the prompt changes.
The decisionWhen to reach for it
Use it
Whenever a free pattern arm scores well on a generated corpus. It is the check that says whether the score is about the task or about the wording.
Do not use it
⚠︎ AS A MEASUREMENT OF REAL-WORLD BRITTLENESS. The paraphrases and the pattern share an author, so this is an UPPER BOUND on how brittle a pattern-matcher is, not an estimate of it. The pattern was deliberately NOT extended to cover these phrasings: extending it moves the boundary and settles nothing.
A living map of modern AI — kept current every morning