Every evening your desk has to decide which failing trades are close to forcing a costly buy-in, market by market, holiday by holiday. This app counts each fail's age at its own market and flags exactly which ones are due.
PresenterOpens the private repo. Visible to admins only.
For the settlements deskWealth & Asset Management · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A settlements manager at a brokerage or asset manager, working the evening fails list.
✕Today's manual process
1Open each fail and find its market, then pull that market's own holiday calendar.
2Count the days manually, deciding case by case which ones the market itself caused.
3Check the deadline against the buy-in terms in the agreement, line by line.
4One wrong count forces a buy-in too soon, or misses the deadline and leaves a breach on the file.
Fails counted and dated manually, every evening
✓With the app
1Each fail is opened and its own market's calendar read for you.
2The days are counted, market shutdowns and counterparty delays told apart.
3The deadline is set straight from the buy-in terms on file.
4Hold or instruct is decided, one clear call, every evening.
Every fail counted and dated automatically, every evening
See it work
One real case: what the app found, step by step
A failing trade at Northgate Exchange, six settlement days into its clock, one day short of a forced buy-in.
Catch failing trades before the buy-in deadlineReference appBuilt to be shaped to your process
6
1The fail being watched checked again this evening, at its own market.
2The clock it read seven settlement days is the deadline, at this market.
3What it found still suspended, not a live delivery failure today.
4What doesn't count a closed line at the market doesn't add to the count.
5The count six of seven settlement days used, one day short of the deadline.
6The call hold. No buy-in ordered on this evening's run.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"Which trades are still failing" is a SELECT and nobody needs a language model for it. The question a settlements desk actually has to answer every evening is the next one, and it has a hard, one-way clock on it: how many SETTLEMENT DAYS has this fail aged AT ITS OWN PLACE OF SETTLEMENT -- not calendar days, and not the desk's own business days -- and is today the day the mandatory buy-in falls due. Three things make that unanswerable from the row in front of you. The market where it settles may have been shut on a day the desk was open. The line may have been un-settleable for reasons that are nobody's counterparty's fault, which do not age the fail, mixed in with days the delivering party simply did not deliver, which do. And the exceptions feed keeps one day, so an event that opened before this window is not on this page at all. Somebody working the evening fails list by hand: opening each open fail, finding its place of settlement, pulling that market's calendar, counting settlement days from the intended settlement date, deciding for each exceptions line whether the market was shut or the counterparty was short, subtracting the first kind and not the second, projecting the remaining days forward across the next holiday, and remembering which fails already have a buy-in instructed. On this book that is 23 readings an evening, five evenings a week.
Audience
A settlements or middle-office manager deciding what to instruct before the cut-off, and the operations analyst who will have to defend the date to a counterparty who disputes it. Both of them are reading the same evening list and both of them need the arithmetic shown, not asserted -- a buy-in is a forced purchase in the market and the first question anybody asks afterwards is which days you counted. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual fail snapshots
The corpus is 112 fail snapshots, 0.68 MB (json 1 · jsonl 2 · txt 201). A SETTLEMENT FAILS FILE IS BOTH THE MOST COMMERCIALLY SENSITIVE AND ONE OF THE MOST REGULATED BOOKS ON A DESK. It names the counterparty, the instrument, the size and the price of every trade a firm could not settle -- a list of who cannot deliver what is information that moves a price -- and there is no public one at any granularity that would let a monitor be tested. The alternative, a scrubbed real book, is worse: scrubbing is exactly the step that changes what the model is asked to read, and on this book the counterparty is half of what the line is about. So the corpus is invented and the generator ships beside it. See data/SOURCES.md.
The corpus
The 112 fail snapshotsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your fail snapshots. That is the whole change — there is no database to migrate.
One fail snapshot, as the model receives itFAIL-0001-R1.txt · 1 of 112
Synthetic Record
------------------------------------------------------------------------
Every field below is invented. This is a generated settlement-fail snapshot for an AI
use-case kit. It reproduces no real trade, counterparty, fund, depository or market, and
the buy-in terms are ILLUSTRATIVE.
Fail
------------------------------------------------------------------------
Fail reference : FAIL-0001
Trade reference : TRD-898073
Counterparty reference : CPTY-3259
Instrument reference : SYN-1954 (ordinary shares)
Instrument class : Liquid equity
Place of settlement : NGX -- Northgate Exchange (Northgate Central Depository)
Direction : RECEIVE -- we are the receiving party; the
delivering party has not delivered
Intended settlement date (ISD) : 2026-02-27
Original quantity : 20,000
Residual failing quantity : 20,000
Price on file : 167.74 USD
Residual exposure at the price on file : 3,354,800.00 USD
Buy-in Terms
------------------------------------------------------------------------
Settlement agreement : SSA-5801-2025
Extension period on file : 4 settlement days from the intended settlement date
Buy-in agent : standing appointment; instructed by this desk
Rule S-1 A fail ages in SETTLEMENT DAYS at its PLACE OF SETTLEMENT, counted from the day AFTER the
intended settlement date up to and including the snapshot day. A day on which that market
does not settle -- a weekend, or a holiday in that market -- does not age the fail, even
Abridged — the file continues.
The outcomeWhat a good result looks like
Every open fail carries a counted age in settlement days, an exact buy-in deadline date, and one clear instruct-or-hold call, and exactly one buy-in is instructed per fail on the run at which the extension period is first found exhausted. A fail whose agreement carries no extension period is reported CONTEXT_INCOMPLETE and ages nothing, rather than being aged against a default nobody agreed to.
And when it cannot
Two ways, and they cost different things. A MISSED buy-in leaves a mandatory purchase un-instructed past its deadline: the exposure stays open, the fail keeps ageing, and the breach is on the file. A PREMATURE buy-in is worse in a way that is easy to miss -- it is a forced purchase in the market, executed by an agent at that morning's price, charged to a counterparty who still had days left. The aged-fails report a desk runs today produces 35 of them on this book of 112 readings; the model produces 0.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
You already know which fails are past their extension period and only need the list — a SQL query against your settlements book, or the b000 floor in this repo Days elapsed against a threshold is arithmetic. On this corpus b000 gets 59.82 pct of statuses right for nothing at all -- it is only the DEADLINE DATE it cannot produce.
You need a safety score for the whole kit rather than one abstention rule — a red-team run this kit does not have context_incomplete_recall_pct measures ONE stated abstention rule on 10 readings. It is a guardrail score, not a safety score.
You want a number for what YOUR missed runs cost — evals/cadence.py, re-run on your own book -- it costs nothing and needs no key The crossings-lost figure is a fact about which fails in THIS book settle in full the morning after they cross. Yours will differ and the code is free.
At a glanceHow the whole thing runs
100%status accuracy pct
9,815 msp50, end to end
$5.86per 1,000 fail snapshots · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch failing trades before the buy-in deadline14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own book, or replace data/corpus/*.txt and data/gold.jsonl directly. The measured figures on this page do not travel with your corpus, and two of them are corpus properties rather than model properties.Corpus lens →
When is this the wrong choice?
Avoid: Anything where the truth is contested. This grader is correct BY CONSTRUCTION on this corpus because the same script wrote the pages and the key. It says nothing about whether the nine rules are the right rules. That is the case against the best-fitting scenario (“You already know which fails are past their extension period and only need the list”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
The corpus is 24 invented fails over five business days in four invented markets, generated from one seed by a committed script. Every figure on this page is a fact about that book before it is a fact about the task. 8 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether rendering the carried state as ENGLISH rather than as JSON matters. src/state.describe writes a paragraph because a model handed {"age_settlement_days": 6} has to be told what the field counts. 6 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-settle-fail. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces: both corpora and both answer keys (regenerated from the seed by python3 tools/build_corpus.py, byte-identical), every pre-flight check in evals/check_labels.py, all four free floors, the free missed-run floor arm, the whole of evals/cadence.py including the crossings lost permanently, and the UI on both schedules with the free-floor column populated and the scored runs replayable off the committed result files. What it cannot reproduce without a key is a live model column.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
9,815 msp50, end to end
36,665 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed to a status, an age, a deadline date and an instruct-or-hold call. It does NOT include the corpus build, the label checks or the scorer, none of which call anything. Measured over the 112 readings of r001-settle-fail with 16 concurrent fail chains, so it is the latency a single reading saw under that concurrency, not a serial figure.
Current processWhat it replaces
Somebody working the evening fails list by hand: opening each open fail, finding its place of settlement, pulling that market's calendar, counting settlement days from the intended settlement date, deciding for each exceptions line whether the market was shut or the counterparty was short, subtracting the first kind and not the second, projecting the remaining days forward across the next holiday, and remembering which fails already have a buy-in instructed. On this book that is 23 readings an evening, five evenings a week.
Where it is not good enough
⚑ THE MODEL IS AT 100 PCT ON ALL FOUR FIELDS AND THAT IS THE LEAST INTERESTING NUMBER ON THIS PAGE. It answered every one of the 112 readings, and it was right about the status, the counted age, the exact buy-in deadline date and the instruct-or-hold call on every one. But free code that costs $0.00 is at 95.54 pct on the deadline, so the model's whole margin over the strongest floor is 4.46 points, on five readings out of 112. A team that cannot spend should run evals/baseline.py's marketcal-mem and be nearly all the way there.
⚠︎ WHAT IT IS NOT GOOD ENOUGH FOR. (1) A 100 pct on a 112-reading synthetic book is not evidence about a real book. The corpus was generated by a committed script from one seed, and its hardest judgement -- was the line shut, or was the counterparty short -- is drawn from thirteen phrasings. A real exceptions feed has thousands. (2) It is a WATCHLIST. The kit never instructs, executes or settles a buy-in, and a 100 pct reading is not a reason to let it. (3) The first reading of every fail is easier here than in life: no suspensive event in this corpus opens before the watch is stood up, so the first run can be answered from the calendar alone. A real stand-up inherits a backlog of suspensions nobody has a record of. (4) The one place the model is measurably fragile is not the model at all -- skip one scheduled run and it drops to 95.12 pct on the deadline and 92.68 pct on the age over the readings it still takes, and 8 of its 11 answers that changed changed from right to wrong.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt201jsonl2json1
24 open receive fails, re-read whole across two schedules — 112 intact readings, 89 with one business day never read
every business day, 18:00 local — five runs, 2026-03-05 to 2026-03-11
one run owns the settlement days accrued since the last one
Recorded failureskip one business day and 2 of 18 mandatory buy-ins are gone for good — $8,315,720 of residual exposure, settled in full before the next run
caldays alone: deadline 8.93% · +carried state: still 8.93%
swap ONE calendar, bizdays -> marketcal: 75.00% -> 95.54%
Recorded failurecounting the desk's OWN business days instead of the market's costs 20.54 points on the deadline — 28 dates wrong against 5, 26 of them EARLY by 49 days total
missed and premature buy-ins counted apart, never averaged
Recorded failure0 misses, 0 premature buy-ins — but the whole margin over the strongest free floor is 4.46 points, on 5 readings where prose, not structure, decides
free market-calendar floor did 95.54% on the deadline for $0.00
one skipped run: 2 of 18 buy-ins gone for good, $8,315,720
2026-08-23as of
It produces a watchlist row — a counted age, an exact buy-in deadline date and an instruct-or-hold call — for a settlements desk to validate. It never instructs, executes or settles a buy-in, and there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is four scalars written by src/state.py from the parsed arithmetic and never from the model's reply, so a wrong reading is a wrong row on the desk's list and never a corrupted history. The clock is the second: one run owns every settlement day accrued since the last one, and Rule S-6 freezes a deadline the moment the count reaches the extension period, so a run that does not happen is not a late fix — it is a day the record never gets to have an opinion about.
⚠︎ THE FREE FLOOR IS THE STORY AND THE PAGE LEADS WITH IT. The model is 100.00% on all four fields over all 112 readings, and free code that costs nothing is at 95.54% on the field that matters most — the exact deadline date — PROVIDED it counts days on the market's own calendar rather than the watching desk's. Counting the desk's Monday-to-Friday days instead is a one-line swap that costs 20.54 points, 28 wrong dates against 5, 26 of them early by 49 days in total.
⚠︎ AND THE MODEL'S REAL MARGIN IS NARROW AND LOCALISED: 4.46 points over the best free floor, on five readings where a market-wide shutdown is written entirely in counterparty language and the floor's keyword table has nothing in it to catch that. ⚑ THE SHARPEST NUMBER ON THIS PAGE IS NOT ON THE ACCURACY ROW AT ALL: skip one scheduled business day and 2 of the 18 mandatory buy-ins in this book cross their deadline and settle in full before the next run — $8,315,720 of residual exposure that no reading, on any schedule, at any price, will ever record. Accuracy on the readings the degraded schedule still takes barely moves (95.12% status); a monitor grades itself only on the readings it took.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because the Cost lens prices them.
the market calendars
src/marketcal.py
MARKETS is four dictionaries of holidays. Replace them with your own markets, or with a real calendar feed -- everything downstream, including both day-counting floors, reads through settles() and business_days().
the buy-in rule
src/settlement.py
RULE_TEXT is what the model reads and step() is what the answer key uses; they are deliberately the same nine rules in two forms. Change the extension periods, the precedence or the Rule S-5 projection convention here and the harness, the UI and all four floors move together.
what the model is asked
src/prompt.py
INSTRUCTION and the JSON shape. Add or drop an answered field here and in evals/scoring.FIELDS; the four are scored independently.
what leaves the machine
src/select.py
SECTION_HINTS and NEVER_SENT. Moving a section out of NEVER_SENT is the only way to widen what is sent, and check_labels.py will convict it.
the corpus
tools/build_corpus.py
PLAN, the two event pools and SEED. Keep the eight section headings and the two dated calendar blocks and every reader in the kit still parses.
the free floors
evals/baseline.py
MODES plus one branch in review(). A fifth floor is roughly thirty lines and costs nothing to run.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 24 open fails x 5 scheduled runs = 112 snapshots from a fixed seed (SEED = 20260823), AND a second rendering of the same book with the third business day never read (89 snapshots). Every gold value is computed from the fail model through src/settlement.step, never read back off the rendered page. Each fail is given a TARGET crossing run and the generator works backwards over that market's calendar to find the intended settlement date that produces it -- because a corpus whose crossings fell where the seed put them could contain none at all on the day the ablation skips, and a missed-run measurement over a day nothing crossed measures nothing.
the market calendars
src/marketcal.py
Four invented markets, each with its own holidays, and the two day-counting rules the kit compares: settlement days at the place of settlement, and the watching desk's own Monday-to-Friday days. This file IS the trap the kit measures. Weekends are closed in every market modelled, which is a simplification and is stated.
the buy-in rule
src/settlement.py
The nine illustrative rules as arithmetic: the running count, the suspensive deduction, the forward projection of the deadline over the market calendar, the freeze once the extension period is reached, and the once-per-fail instruction. step() takes the LIST of settlement days in the window rather than a count, because when the window is wide -- the first run, or any run after a missed one -- the count can cross the extension period part-way through it and the deadline is a day INSIDE the window.
the carried state
src/state.py
Four scalars between runs -- the counted age, whether a suspensive event was open, the deadline once frozen, and the previous status -- rendered as one English paragraph. Written from the arithmetic, never from the model's reply, so one bad reading cannot compound into every reading after it.
the section splitter
src/segment.py
Splits a snapshot into its eight named sections. All eight are asserted present in every document on both schedules before a run may spend.
the privacy gate
src/select.py
Decides which sections reach the prompt. Counterparty Desk Contact is mapped by nothing and subtracted unconditionally in the fallback, so it never leaves the machine -- personal data, and a market-sensitive fact about who cannot deliver.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state paragraph, and the sections select.py let through. The middle part is the experiment; --stateless replaces it and changes nothing else, which check_labels.py asserts line by line.
the runtime
src/watch.py
One fail, one scheduled run, one call. Parses the arithmetic inputs off the page with regexes so the scorer can tell a reading failure from a reasoning failure, and holds the published output ceiling.
the model adapter
src/adapters/__init__.py
Raw HTTP to any OpenAI-compatible endpoint, plus per-vendor adapters. No SDK, no dependency; transient statuses retried, terminal ones not.
the spend guard
src/budget.py
Counts live calls against a cap read from the shared .env, ledgered before the call is made. Shared by every kit under the same repo root, because they share a key.
the four free floors
evals/baseline.py
caldays, caldays-mem, bizdays-mem, marketcal-mem -- all $0.00, all scored through the identical scorer. bizdays-mem and marketcal-mem differ in exactly one thing, which calendar the days come from, so the price of the trap is a measured number.
the scorer
evals/scoring.py
Exact match per cell, no model. Missed buy-ins and premature buy-ins counted apart and never averaged; the deadline scored by direction as well as by hit rate.
the cadence comparison
evals/cadence.py
Compares the intact schedule against one where a business day is never read, twice -- once with the model taken out entirely, once with it. Free. Prints the crossings lost and the cells that moved rather than an accuracy delta, because the accuracy delta is the small half of the damage.
the pre-flight checks
evals/check_labels.py
Everything that must be true before a run may spend: the answer key replays through the same advance() the harness uses on both schedules, no snapshot restates its own carried state, the privacy guard holds and is red-proven in both directions, the two calendars actually disagree on this corpus, and a skipped run widens the calendar window and NOT the activity log.
the local UI
src/app.py
http.server, no framework. Shows the carried state verbatim beside the answer and the free floor beside both, and carries a SCHEDULE SELECTOR that re-reads the same fail off the ablation corpus -- so a reader can watch a deadline move because nobody looked on Monday.
Where it breaks at scale
⚑ THE CADENCE IS THE SCALING VARIABLE, NOT THE BOOK SIZE, AND IT PULLS IN BOTH DIRECTIONS. One run of this watch is one call per OPEN fail, so the bill is (open fails) x (runs per day) -- and unlike a document kit, the population does not shrink as you work it. This run cost $0.00585772 a reading projected onto the shared rate card and 112 readings covered 24 fails over 5 runs. Ten times the book at the same daily cadence is ten times the bill, linearly, with no cliff. Running twice a day doubles it again.
⚠︎ AND SLOWING DOWN IS NOT FREE, WHICH IS THE WHOLE POINT OF THIS KIT. Skipping ONE business day on this book lost 2 mandatory buy-ins permanently -- $8,315,720 of residual exposure in fails that crossed their deadline on the unread day and settled in full before the next run, so no reading on any schedule records that a buy-in was required. Three more were instructed a day late. The free floor's accuracy over the readings it still took fell from 97.75 to 92.13 pct on the deadline, and 16 of the 21 cells that moved moved from right to wrong with not one moving the other way.
⚠︎ THE STATE DOES NOT GROW, AND THAT IS DESIGNED. Four scalars a fail, the same on run 500 as on run 2. A monitor that carried a transcript would price its input at the square of the number of runs; this one is flat, and evals/run.py re-derives the count from the calendar rather than inheriting it from a reply, which is why the blind window does not poison every reading after it.
⚠︎ WHERE IT ACTUALLY BREAKS. The forward calendar printed on each page is 16 calendar days long. A fail whose remaining extension runs past the end of it cannot be projected from the page -- the free floor abstains on exactly two readings for that reason, and a book of long-dated growth-market fails would hit it constantly. Widening the block widens every prompt, so it is an input-token decision, not a free one.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
FAIL-0016-R3, the whole argument on one screen, and the reading was chosen by scanning both answer keys for the case where free code and the truth disagree MOST -- not by looking for a flattering frame. Its window carries one exceptions line: "neither we nor the counterparty can put this through: the line has been shut to all instructions since the notice went out this morning". That is the MARKET being shut, written entirely in counterparty language. The free keyword floor sees no market vocabulary, ages the fail a day it did not age, reaches the 7-day extension period and says INSTRUCT THE BUY-IN NOW with the deadline pulled to 2026-03-09. The answer key says SUSPENDED at 6 days with the deadline still on 2026-03-11, and the scored run agrees with the key, citing Rule S-3 and Rule S-5 by name. The agreement column says "differs" on all four rows, which is the page admitting the model earned its bill here and, on 107 other readings, did not.successOpen full size →The page before anything is asked, with no key configured. The carried state, the parsed arithmetic inputs, the free floor and the withheld-section list are all computed locally and render offline -- so a forker who has just cloned this sees the whole argument, and the only blank column is the one a call would fill.emptyOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
THE SAME FAIL, ONE DAY LATER, ON THE SCHEDULE WHERE NOBODY LOOKED ON MONDAY -- and the model is now wrong in the expensive direction. The schedule selector is a real control: it re-reads FAIL-0016 at run 4 off the ablation corpus, where the previous scheduled run is 2026-03-06, the calendar window is two settlement days wide, and the activity log shows ZERO entries, because the shutdown opened on the day that was never read and this feed keeps one day. Given a carried state that says nothing was open and two days that look ordinary, the run answers BUY-IN DUE, 8 settlement days, deadline 2026-03-09, instruct it now. The truth is SUSPENDED, 6 days, deadline 2026-03-12, hold. All four fields wrong, and the free floor agrees with it -- both columns read "same", which on this frame means both were blind rather than both were right. The amber banner is the kit saying so: it detects the gap for nothing and cannot recover the entry.failureOpen full size →The read button pressed with no API_KEY. It returns a plain sentence saying nothing was called, not an error and not a stack trace -- and the free floor beside it still has an answer. This is the state most forkers are in for their first ten minutes, and a kit whose page is broken until somebody pays teaches nothing.failureOpen full size →
How it is cutWhat one fails (up to 5 scheduled runs each) is
No split, and no chunking. The unit is a FAIL -- up to five consecutive scheduled runs processed strictly in order, because run 4's prompt contains a count produced by run 3. Different fails are independent and run concurrently. There is no train or dev set: nothing is fitted, and the four free floors were written from settlements vocabulary rather than tuned against this corpus.
SetupWhat the setup figure measured
There is no index and no retrieval step -- every open fail is re-read whole on each scheduled run. tools/build_corpus.py writes both corpora, both answer keys and the stats file in about a second, entirely offline.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every field was written by the generator committed beside it.
Bring your ownBring your own fail snapshots
Point tools/build_corpus.py at your own book, or replace data/corpus/*.txt and data/gold.jsonl directly. Keep the eight section headings (src/segment.SECTIONS -- the parser asserts all eight before a run may spend), keep the two dated blocks in the Settlement Calendar section in the printed form, and put your own markets in src/marketcal.MARKETS. The gold row shape is one line of data/gold.jsonl.
⚠︎ And what stops being true when you do: The measured figures on this page do not travel with your corpus, and two of them are corpus properties rather than model properties. The 95.54 pct the strongest free floor scores is a fact about thirteen exceptions phrasings and four holiday calendars. The 2 crossings lost to one missed run is a fact about which fails happen to settle in full the next morning. Re-run all seven arms on your book before quoting anything here.
What breaks it
The corpus is 24 invented fails over five business days in four invented markets, generated from one seed by a committed script. Every figure on this page is a fact about that book before it is a fact about the task.
⚠︎ A CORPUS WHOSE EXCEPTIONS LINES ARE SEPARABLE BY KEYWORD MEASURES THE TEMPLATES, NOT THE WORK -- FOUND BEFORE ANY CALL. The first pool had clean vocabulary on both sides and a free keyword table separated it perfectly. The shipped pool carries suspensive events written entirely in counterparty language and counterparty events written entirely in market language; the keyword floor is now wrong on four of thirteen phrasings and wrong in BOTH directions.
⚠︎ NO SUSPENSIVE EVENT OPENS BEFORE THE FIRST SCHEDULED RUN, SO THE FIRST READING IS EASIER HERE THAN IN LIFE. The watch is stood up over an already-failing book and must count back to the intended settlement date from the calendar alone. A real stand-up inherits a backlog of suspensions nobody has a record of, and this corpus does not model it. evals/check_labels.py asserts the simplification rather than hiding it.
⚠︎ EVERY FAIL IS A RECEIVE FAIL. The desk is owed securities and is the party entitled to buy in. A DELIVERY fail -- where the desk is the one being bought in against -- is a different reading with a different action, and it is not in this corpus and not measured.
⚠︎ EVERY MARKET SETTLES MONDAY TO FRIDAY. src/marketcal.WEEKEND is a constant, so a market whose week is not Monday to Friday would be counted wrong by the kit and by every floor built on it.
⚠︎ THE ARITHMETIC INPUTS ARE PRE-COMPUTED ON THE PAGE. "Settlement days at NGX in this window" is printed for the reader, because it is a calendar fact a database answers exactly and for nothing. A feed that did not carry it would make this a harder task than the one measured here.
⚠︎ THE PROJECTION CONVENTION IN RULE S-5 IS INVENTED. An open suspension is projected to cost exactly one further settlement day. That is a stated convention, not a prediction, and a book whose suspensions typically run a week would need a different one -- every deadline on a suspended fail would be published too early.
⚠︎ ONLY FOUR MARKETS, AND THE HOLIDAYS WERE PLACED TO FALL IN THE WINDOW. The gap between the two day-counting floors is therefore a property of this calendar as much as of the rule. On a book settling in one market with no holiday in the period, bizdays-mem and marketcal-mem would be identical and the trap would cost nothing.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
164
not measured
instruction
4,043
not measured
carried state
358
not measured
Synthetic Record
299
not measured
Fail
874
not measured
Buy-in Terms
3,061
not measured
Ageing Position
481
not measured
Settlement Calendar
882
not measured
Activity Log
311
not measured
Operational Notes
179
not measured
Total
2,550
This is the cost lesson as arithmetic: of the 10,652 characters assembled, 4,207 are instructions — 39% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
TWO MESSAGES, JOINED IN SEND ORDER. The provider receives a system message and a user message; prompt_verbatim prints both, separated by a blank line, because publishing only the second would be publishing most of a prompt. Replayed from the kit's own src/prompt.build() for FAIL-0016-R3 with the carried state that reading was actually given, rather than logged by the run. What makes the replay checkable is that the section list it produces is the list the run recorded in sections_used -- Synthetic Record, Fail, Buy-in Terms, Ageing Position, Settlement Calendar, Activity Log, Operational Notes -- and that src/prompt.build() is the only thing in the kit that assembles a prompt, so there is no second path a run could have taken. Every one of the 10 parts below is a literal substring of the verbatim prompt, in ascending order, and the builder asserts it. ⚠︎ WHAT IT IS NOT CHECKED AGAINST, SAID PLAINLY: the run's own prompt_parts records ONE size, 10663 characters, belonging to whichever reading finished first under 16 concurrent workers -- not to FAIL-0016-R3, whose assembled user message is 10678 characters. The harness stores one exemplar, not one per reading, so no per-reading size comparison is available and none is claimed.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled settlement-fail watch. You apply a written buy-in rule to one open fail at one scheduled run. You answer with one JSON object and no other text.
You are the scheduled settlement-fail watch for one asset manager's settlements desk. It wakes once
every business day, after the settlement cycle has closed, re-reads EVERY fail still open on the
book, and reports each one. You are reading ONE fail at ONE scheduled run.
The buy-in rules, the extension period on file and the place-of-settlement calendar are reproduced
in the snapshot below. Apply them exactly as written. The fail's AGE is a RUNNING COUNT of
settlement days that survives between scheduled runs, and you cannot see the earlier runs; what is
known about them is stated under "Carried state" and is the only history available to you. Do not
assume anything about earlier runs beyond it.
How to count:
- "Settlement days at <market> in this window" is how many days in this window that market
settled. It is already computed for you from the calendar; do not re-derive it.
- From those days, SUBTRACT every one covered by a SUSPENSIVE event (Rule S-3). Do NOT subtract a
day lost to the delivering counterparty -- no stock, a blocked account, a mis-instruction,
funding that did not arrive. Those days age the fail in full, because they are what the buy-in
exists for. Deciding which kind an entry is means reading what it says, not which words it
uses: an entry can be written entirely in counterparty language and still describe a market
that was closed to everybody, and the other way round.
- A suspensive event runs from its opening entry INCLUSIVE to its closing entry EXCLUSIVE. The
Activity Log shows only the entries raised on THIS snapshot day. A closing entry with no opening
entry belongs to an event that was already running: the carried state says whether one was open
and therefore whether the earlier settlement days in this window were suspended.
- ADD what is left to the age the carried state says was already counted.
- The buy-in deadline is the settlement day on which that count reaches the extension period.
Project it forward over the "Forward" calendar in the snapshot, counting only days marked
SETTLES. If a suspensive event is open at this snapshot, skip the FIRST such day (Rule S-5).
- Once the count has reached the extension period, the deadline is FIXED at the day it reached it
and never moves again (Rule S-6). If the carried state already gives you a fixed deadline,
repeat it. If the count reaches the extension period on THIS run, identify the day it reached
it from the "Since the previous scheduled run" calendar in this snapshot -- when more than one
settlement day fell in this window, that day may be earlier than the snapshot day.
- A partial settlement does NOT re-base the intended settlement date (Rule S-8). The residual
keeps the original ISD, the original age and the original deadline.
- If no extension period is on file, or no intended settlement date is captured, the fail is
CONTEXT_INCOMPLETE: status CONTEXT_INCOMPLETE, age 0, buyin_deadline NOT_DETERMINABLE,
raise_buyin NO. Never assume a default extension period (Rule S-2).
Answer with a single JSON object and nothing else:
{"status": "AGEING|DUE_NEXT_RUN|BUYIN_DUE|SUSPENDED|CONTEXT_INCOMPLETE",
"age_settlement_days": <whole number of settlement days now counted against this fail>,
"buyin_deadline": "YYYY-MM-DD", or "NOT_DETERMINABLE",
"raise_buyin": "YES|NO",
"rationale": "one sentence, naming the rule you applied and the days you counted"}
Precedence for "status", applied in this order: CONTEXT_INCOMPLETE beats everything; then BUYIN_DUE
if the counted age has reached the extension period, even if a suspensive event is open right now;
then SUSPENDED if a suspensive event is open at this snapshot; then DUE_NEXT_RUN if the deadline
falls on or before the next scheduled run named in the snapshot; otherwise AGEING.
"raise_buyin" is YES only on the run that FIRST finds the extension period exhausted -- the count
has reached it and the carried state does not already give a fixed deadline. On every later run it
is NO (Rule S-7).
Carried state
------------------------------------------------------------------------
As at the previous scheduled run, 6 settlement day(s) had already been counted against this fail at its place of settlement, suspensive days already excluded. No suspensive event was open when that run ended. The extension period has not yet been found exhausted, so no deadline has been fixed and no buy-in has been instructed. It was reported DUE_NEXT_RUN.
Fail snapshot
------------------------------------------------------------------------
Synthetic Record
------------------------------------------------------------------------
Every field below is invented. This is a generated settlement-fail snapshot for an AI
use-case kit. It reproduces no real trade, counterparty, fund, depository or market, and
the buy-in terms are ILLUSTRATIVE.
Fail
------------------------------------------------------------------------
Fail reference : FAIL-0016
Trade reference : TRD-220266
Counterparty reference : CPTY-2967
Instrument reference : SYN-2285 (senior unsecured notes)
Instrument class : Corporate debt
Place of settlement : NGX -- Northgate Exchange (Northgate Central Depository)
Direction : RECEIVE -- we are the receiving party; the
delivering party has not delivered
Intended settlement date (ISD) : 2026-02-26
Original quantity : 5,000
Residual failing quantity : 5,000
Price on file : 42.34 USD
Residual exposure at the price on file : 211,700.00 USD
Buy-in Terms
------------------------------------------------------------------------
Settlement agreement : SSA-2623-2025
Extension period on file : 7 settlement days from the intended settlement date
Buy-in agent : standing appointment; instructed by this desk
Rule S-1 A fail ages in SETTLEMENT DAYS at its PLACE OF SETTLEMENT, counted from the day AFTER the
intended settlement date up to and including the snapshot day. A day on which that market
does not settle -- a weekend, or a holiday in that market -- does not age the fail, even
if the watching desk is open that day.
Rule S-2 The extension period is an AGREEMENT-SUPPLIED value read from the settlement agreement.
Where no extension period is on file, or where no intended settlement date is captured,
the fail is CONTEXT_INCOMPLETE: it does not age, no deadline is determinable, no buy-in
is instructed, and no default extension period may be assumed.
Rule S-3 A settlement day on which the line could not settle AT ALL -- a depository blackout, a
market suspension, a settlement-engine outage, a corporate-action freeze -- is SUSPENSIVE
and does not age the fail. A day lost to the delivering party -- no stock in the account,
a blocked account, a mis-instruction, funding that did not arrive -- is NOT suspensive and
ages the fail in full. That is what the buy-in exists for.
Rule S-4 A suspensive event runs from the day its opening entry is raised, INCLUSIVE, to the day
its closing entry is raised, EXCLUSIVE. The activity log shows only the entries raised on
the snapshot day, so an event that opened earlier is known only from the carried state.
Rule S-5 The buy-in deadline is the settlement day on which the counted age reaches the extension
period, projected forward over the place-of-settlement calendar. A suspensive event open
at the snapshot is projected to cost ONE further settlement day; if it runs longer, a
later run moves the deadline out.
Rule S-6 Once the counted age has reached the extension period, the deadline is FIXED at the day
it reached it and never moves again. A suspension after that day does not extend it.
Rule S-7 A buy-in is instructed ONCE per fail, on the first scheduled run at which the extension
period is found exhausted. A later run reports the requirement already live; it does not
instruct a second buy-in.
Rule S-8 A PARTIAL SETTLEMENT DOES NOT RE-BASE THE INTENDED SETTLEMENT DATE. The residual quantity
keeps the original ISD, the original counted age and the original deadline.
Rule S-9 A fail is DUE_NEXT_RUN while its deadline falls on or before the next scheduled run --
that is, when the requirement will become live before the desk looks again.
Precedence, applied in this order: CONTEXT_INCOMPLETE, then BUYIN_DUE, then SUSPENDED, then
DUE_NEXT_RUN, then AGEING.
Ageing Position
------------------------------------------------------------------------
Snapshot taken : 2026-03-09 18:00 (scheduled run 3 of this watch)
Watch cadence : every business day, 18:00, at the watching desk
Previous scheduled run : 2026-03-06 18:00
Next scheduled run : 2026-03-10 18:00
Settlement days at NGX in this window : 1
Settlement days since ISD, before any suspension : 7
Settlement Calendar
------------------------------------------------------------------------
Place of settlement : NGX -- Northgate Exchange
NGX settles Monday to Friday except the dates marked CLOSED below.
Since the previous scheduled run
2026-03-07 Sat CLOSED (weekend)
2026-03-08 Sun CLOSED (weekend)
2026-03-09 Mon SETTLES
Forward, from the next calendar day
2026-03-10 Tue SETTLES
2026-03-11 Wed SETTLES
2026-03-12 Thu SETTLES
2026-03-13 Fri SETTLES
2026-03-14 Sat CLOSED (weekend)
2026-03-15 Sun CLOSED (weekend)
2026-03-16 Mon SETTLES
2026-03-17 Tue SETTLES
2026-03-18 Wed SETTLES
2026-03-19 Thu SETTLES
2026-03-20 Fri CLOSED (market holiday: Spring Bank Holiday)
2026-03-21 Sat CLOSED (weekend)
2026-03-22 Sun CLOSED (weekend)
2026-03-23 Mon SETTLES
2026-03-24 Tue SETTLES
2026-03-25 Wed SETTLES
Activity Log
------------------------------------------------------------------------
Entries raised on this snapshot day only. This feed does not retain earlier days.
07:41 neither we nor the counterparty can put this through: the line has been shut to all instructions since the notice went out this morning
Operational Notes
------------------------------------------------------------------------
Nothing outstanding on the pre-matching check; shape and place of settlement both agree.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"status": "SUSPENDED",
"age_settlement_days": 6,
"buyin_deadline": "2026-03-11",
"raise_buyin": "NO",
"rationale": "Applied Rule S-3 to exclude 2026-03-09 as a suspensive shutdown and Rule S-5 to skip 2026-03-10, leaving age at 6 of 7 extension days with the projected deadline at 2026-03-11."
}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch failing trades before the buy-in deadline — 112 fail snapshots. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the fail model and the market calendars, never read back off the rendered page. Four fields per reading: the status, the counted age in settlement days, the buy-in deadline as an exact ISO date, and the instruct-or-hold call. A date is compared as a string, so being one day early is exactly as wrong as being a month early on the hit rate -- which is why the scorer ALSO counts the signed error in days and separates early from late.
112fail snapshots
112source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED112 · 102 · 39 / 112status accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED112 · 96 · 38 / 112age accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED112 · 96 · 39 / 112buyin deadline accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED112 · 77 · 40 / 112raise buyin accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 18missed buyin pct — readings whose correct answer is to INSTRUCT the buy-in, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 94premature buyin rate pct — readings whose correct answer is NOT a live buy-in requirement, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED80 / 80memory status accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED80 / 80memory age accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED80 / 80memory deadline accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED10 / 10context incomplete recall pct — readings with no extension period on file or no intended settlement date captured, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 1exposure usd error pct — book total in USD of residual exposure reported as under a live buy-in requirement, against the key, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is asserted. evals/check_labels.py replays every fail chain on BOTH schedules through evals/run.advance -- the same function the harness uses to build the carried state the model is given -- and requires it to reproduce the committed answer key exactly; it does. It also asserts that no snapshot restates its own carried state, that the on-page "settlement days since ISD" figure is NOT the counted age on 19 readings (so a reader that copies it is wrong), that the stateful and stateless prompts differ on exactly one line, and that the privacy guard holds in both directions. All of it runs in seconds and none of it spends.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One fail snapshot
1,000 fail snapshots
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.005858
$5.86
22%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.002343
$2.34
22%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.101879
$101.88
25%
Same work, 43× the bill
The same fail snapshots, the same tokens — only the rate card changed. And across all 3 cards between 22% and 25% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE CARRIED STATE, AND IT IS THE UNUSUAL KIND OF LEVER THAT MOVES BOTH WAYS AT ONCE. Removing it costs 2.92x the money AND drops the deadline from 100.0 to 85.71 pct AND produces 29 duplicate buy-in instructions AND loses 6 readings to the output ceiling. There is no trade-off to reason about here: the paragraph is 358 characters, it is the cheapest thing in the prompt, and it is the difference between a monitor and a classifier that keeps forgetting.
Rates checked 2026-08-18. The provider that actually ran all 275 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the four free floors, evals/cadence.py or evals/check_labels.py. No model judges anything anywhere in this kit. The only money is the readings themselves.
The gradersThree ways to grade
⚑ THE MODEL BEATS THE STRONGEST FREE FLOOR, AND THE MARGIN IS 4.46 POINTS ON THE DEADLINE. 100.0 pct against 95.54 pct on the exact buy-in deadline date, 100.0 against 97.32 on the status. That is the honest headline and it is a narrow one: a team that cannot spend should run marketcal-mem and be nearly all the way there.
⚑ THERE ARE FOUR FLOORS BECAUSE THERE ARE THREE SEPARATE THINGS TO PRICE.
b000 caldays -- the aged-fails report a desk already runs. 8.93 pct on the deadline. It is not a strawman; it is what a threshold report configured once and forgotten actually does, including guessing a default extension period where the agreement carries none, which is the guardrail Rule S-2 forbids: 0.0 pct context-incomplete recall.
b001 caldays-mem -- the same rule with memory. Identical on status, age and deadline, to the decimal. THE ONLY THING MEMORY BUYS A CALENDAR-DAY RULE is not re-instructing the same buy-in every evening: 39.29 -> 68.75 pct on instruct-or-hold. A rule that needs no memory is a rule nobody has to operate, and that is exactly why desks run one -- and it is wrong on 91.07 pct of the deadlines.
b002 bizdays-mem -- THE TRAP. Identical to b003 in every respect but ONE: it counts the watching desk's Monday-to-Friday days instead of the market's settlement days. 75.0 pct against b003's 95.54. 28 deadlines wrong against b003's 5, 26 of them a real date landing EARLY by 49 days in total, and 5 premature buy-ins against b003's 2. That is the price of counting a settlement fail in the wrong kind of day, and it is a measured number rather than a warning. It is also the defect that ships, because it is right on every day both calendars agree.
b003 marketcal-mem -- the strongest floor. 95.54 pct on the deadline.
⚠︎ AND THE FLOOR'S 13 WRONG CELLS SIT ON FIVE READINGS, IN TWO SHAPES. Three are readings where the keyword table read a market-wide shutdown written in counterparty language as the delivering party's fault, aged the fail a day it did not age, and called a premature buy-in with the deadline pulled early. Two are readings where the remaining extension ran past the end of the page's forward calendar and the floor abstained with NOT_DETERMINABLE; the model extended the weekly pattern past the page and got the date right. Those five readings are the entire case for spending money on this task.
⚠︎ ONE ASYMMETRY FLATTERS THE FLOORS, AND IT IS THE SAME ONE EVERY KIT IN THIS SERIES HAS: b002 and b003 share the day-counting implementation with the answer key, so they cannot misread the RULE -- they are the rule -- while the model can only read the prose. What they can and do misread is the English.
the fast tier, with the carried state 100.0% status accuracy · the same tier, memory removed (THE CONTROL) 91.1% status accuracy · the same tier, one business day never read (THE ABLATION) 95.1% status accuracy · free floor: b000 -- caldays 59.8% status accuracy · free floor: b001 -- caldays-mem 59.8% status accuracy · free floor: b002 -- bizdays-mem 90.2% status accuracy · free floor: b003 -- marketcal-mem 97.3% status accuracy · 3 more measured on each run
Context-incomplete recall -- the guardrail, scored Of the 10 readings whose agreement carries no extension period or whose feed captured no intended settlement date, how many were reported CONTEXT_INCOMPLETE rather than aged against a guessed number.
$0.00
no
yes
the fast tier, with the carried state 100.0% context incomplete recall · free floor: b000 -- caldays 0.0% context incomplete recall · free floor: b003 -- marketcal-mem 100.0% context incomplete recall
What a missed run costs, in crossings rather than in points How many mandatory buy-ins fell due on the day nobody looked, how many of those were merely late, and how many are gone permanently because the fail settled in full before the next run.
$0.00
no
yes
no headline metric on any of its 3 runs — they record crossings lost permanently · crossings seen late
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
The four arms are separated by wide margins on the discriminator, and the discriminator is the exact buy-in deadline date: 100.0 pct for the model, 95.54 for the strongest free floor, 75.0 for the same floor counting the desk's own business days, and 8.93 for the aged-fails report a desk runs today. Between the two day-counting floors the ONLY difference is which calendar the days come from, so that 20.54-point gap is attributable to one line of code. The stateless control separates on a different axis entirely -- 85.71 pct on the deadline and 68.75 pct on instruct-or-hold -- so the two experiments do not confound each other.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
You already know which fails are past their extension period and only need the list
a SQL query against your settlements book, or the b000 floor in this repo
Days elapsed against a threshold is arithmetic. On this corpus b000 gets 59.82 pct of statuses right for nothing at all -- it is only the DEADLINE DATE it cannot produce.
Anything where the truth is contested. This grader is correct BY CONSTRUCTION on this corpus because the same script wrote the pages and the key. It says nothing about whether the nine rules are the right rules.
You need a safety score for the whole kit rather than one abstention rule
a red-team run this kit does not have
context_incomplete_recall_pct measures ONE stated abstention rule on 10 readings. It is a guardrail score, not a safety score.
Reading it as a safety score. It measures one rule on 10 readings. It does not mean the kit abstains correctly on anything else.
You want a number for what YOUR missed runs cost
evals/cadence.py, re-run on your own book -- it costs nothing and needs no key
The crossings-lost figure is a fact about which fails in THIS book settle in full the morning after they cross. Yours will differ and the code is free.
Transferring the number. 2 lost crossings is a fact about which fails in THIS book happen to settle the next morning. Re-derive it on yours; the code is free and needs no key.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
BLIND_WINDOW
a suspensive event that opened on a run nobody made
8
FAIL-0016-R4 on the missed-run schedule. The activity feed keeps one day, the shutdown opened on 2026-03-09, and run 3 never happened. The reply: "Carried age 6 plus 2 non-suspensive NGX settlement days (Mar 9, Mar 10) reaches the 7-day extension on Mar 9, so…
NO_MEMORY_DUPLICATE
the same buy-in instructed again, memory removed
29
The stateless control. Told no history is available, the watch cannot know a buy-in has already been instructed on this fail, so it instructs it again on every later run: 29 duplicate instructions across 94 readings whose correct answer is to hold. Downstream…
CEILING_NO_ANSWER
reply ran off the output ceiling and returned nothing
6
FAIL-0002-R4, FAIL-0003-R3, FAIL-0003-R4 and 3 more in the stateless control, finish_reason length at the full 24000-token ceiling. Given a closing entry with no opening entry and told there is no history at all, the model has an unresolvable problem and…
KEYWORD_ATTRIBUTION
the free floor's only real failure shape
13
FAIL-0016-R3. The line reads "neither we nor the counterparty can put this through: the line has been shut to all instructions since the notice went out this morning". The floor's table holds depository, blackout, book closing, settlement engine, record date…
What we could NOT verify
Whether rendering the carried state as ENGLISH rather than as JSON matters. src/state.describe writes a paragraph because a model handed {"age_settlement_days": 6} has to be told what the field counts. The obvious experiment -- the same 112 readings with the state rendered both ways -- costs one more full run and was not paid for. It is a design choice, not a finding.
Run-to-run variance. No reading was fired twice. A 100 pct on all four fields could hide a reading that is 60/40 and happened to land right; nothing here would show it.
Whether a second model agrees. Every paid arm ran on one model at one tier. The kit's whole architecture claim is that swapping is .env plus another run, and that claim is untested here -- no cross-model number is published because none was measured.
Whether the model would still be at 100 pct on a real exceptions feed. Its hardest judgement is drawn from thirteen phrasings written by one author. A real feed has thousands, typed by people under time pressure, and this measures none of them.
What the ablation would cost on the schedule's OTHER days. One business day was skipped, the third. Skipping run 2 or run 4 would hit different fails and a different number of crossings; the 2 lost crossings is one draw, not a distribution.
Whether a desk could detect the blind window in time to act. The kit reports the gap, and the reading it reports it on is already wrong. Nothing here measures whether a human seeing that banner recovers the answer.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
2,550.5
1,527.49
9,815 ms
$0.005858
$0.002343
$0.101879
the same tier, memory removed (THE CONTROL)
2,486.69
5,291.04
22,304 ms
$0.017116
$0.006847
$0.289419
the same tier, one business day never read (THE ABLATION)
2,551.39
2,568.34
13,481 ms
$0.008981
$0.003592
$0.153931
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same reading (one open fail, at one scheduled run), the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
275 live calls were attempted for this kit and 269 returned something usable: 10 calibration (c000, under an 8,000-token ceiling, both hardest chains), 112 scored (r001), 41 for the missed-run ablation (a001 -- 48 readings were REPLAYED without calling, because upstream of the skip the prompts are byte-identical to r001's), and 112 for the stateless control (s001), of which 6 ran to the ceiling and came back unusable. Nothing was discarded and no run was overwritten; every result file is committed, including the free floors, which cost nothing at all. The four FREE arms -- b000, b001, b002, b003 -- plus the free missed-run floor and the whole cadence comparison add $0.00.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, AND MOST OF THEM ARE REASONING. 92.8 pct of this run's output (158,842 of 171,079 tokens) was provider-side reasoning. The prompt is 2550.5 input tokens and costs $0.00127525 of the $0.00585772 per reading; the thinking costs the other $0.00458247.
THE CADENCE, WHICH IS THE ONLY MULTIPLIER ON THE PAGE. One run is one call per OPEN fail. Daily on this book is 120 calls a week; twice a day is 240. Halving it halves the bill and costs what the cadence lens measures -- on one skipped business day, 2 mandatory buy-ins gone permanently.
THE RULE SHEET, REPEATED IN EVERY PROMPT. The nine buy-in rules are 3061 characters of the 10844-character prompt -- roughly 28 pct of the input, on every single reading. Moving them to a system message the provider can cache is the obvious lever and it was not measured here.
THE FORWARD CALENDAR. Sixteen dated lines on every page so the deadline can be projected without a tool call. Shorten it and the free floor abstains more often and the model has to extrapolate; lengthen it and every reading pays.
Your volumeWhat it costs at your volume
LINEAR IN FAILS x RUNS, AND THAT IS THE WHOLE WARNING. Ten times the open fails is ten times the calls at the same cadence -- $0.00585772 a reading, so a 240-fail book read every business day is about $7.03 a week on the shared card. There is no retrieval step to amortise and no index to reuse: the population is re-read whole every run, by design, because a monitor that sliced a stream would miss the fails that did not move. The two things that do NOT grow are the carried state, which is four scalars per fail forever, and the grader, which is free.
Where pricing changes shape
The output ceiling. MAX_TOKENS = 24000 is set from calibration (c000 fired the two hardest fail chains under an 8,000 ceiling and the largest reply was 5043 tokens). The stateful run's largest reply was 6594. The STATELESS control hit the full ceiling 6 times and returned nothing for those readings -- billed in full. A ceiling is not a cost until an arm starts reaching it, and then it is a cost with no answer attached.
A missed run reprices the readings that follow it. Over the same 41 readings the ablation spent 1.46x the output tokens of the intact run and its p95 latency was 95.1s against 36.7s. Skipping a run to save money buys back less than it looks like it does.
Your return, with your numbers
Volumeopen fails per scheduled run -- this run judged 112 readings across 24 fails and 5 business days, on a once-a-business-day cadence
What it replacessomebody working the evening fails list by hand: pulling each market's calendar, counting settlement days from the intended settlement date, deciding for each exceptions line whether the market was shut or the counterparty was short, projecting the remainder forward across the next holiday, and remembering which fails already have a buy-in instructed
Time saved per itemnot measured here -- depends on how long an analyst takes to reconstruct a suspension history from a feed that keeps one day, which is exactly the part this kit shows is not always possible at all
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared connection was already pointed at, and the kit's whole claim is that swapping is .env plus another run. No cross-model comparison was paid for and none is published; could_not_verify says so.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
285,656input tokens · this run
171,079output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 112 readings, one completion call each, one tier. The 112-call stateless control, the 41-call missed-run ablation and the 10 calibration calls are NOT in this figure and are priced separately under Cost.cost_of_evaluation_usd. The four free floors, the free missed-run floor arm and the whole cadence comparison cost nothing at all.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.262
$0.262
$2.34
2026-09-12
gemini-3-flash
Google
$0.656
$0.656
$5.86
2026-09-18
gemini-3-8-flash
Google
$0.856
$0.856
$7.64
2026-09-18
llama-5
Meta
$1.084
$1.084
$9.68
2026-09-18
claude-haiku-4-5
Anthropic
$1.141
$1.141
$10.19
2026-09-12
grok-4-5
xAI
$1.598
$1.598
$14.27
2026-09-18
grok-4-6
xAI
$1.598
$1.598
$14.27
2026-09-18
claude-sonnet-5
Anthropic
$2.282
$2.282
$20.38
2026-09-12
gemini-3-1-pro
Google
$2.624
$2.624
$23.43
2026-09-18
gpt-5-6-terra
OpenAI
$2.624
$2.624
$23.43
2026-09-12
gpt-5-6-sol
OpenAI
$4.564
$4.564
$40.75
2026-09-12
claude-opus-4-8
Anthropic
$5.705
$5.705
$50.94
2026-09-12
claude-opus-5
Anthropic
$5.705
$5.705
$50.94
2026-09-12
claude-fable-5
Anthropic
$11.411
$11.411
$101.88
2026-09-18
claude-fable-5-1
Anthropic
$11.411
$11.411
$101.88
2026-09-18
gpt-6-astra
OpenAI
$11.411
$11.411
$101.88
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 92.8 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (158,842 of 171,079), left at the provider's default, so every row prices a reasoning-on workload. Output is 78 pct of the projected bill on the shared card, which means most of what these rows charge for is the model walking a calendar. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
⚑ EVERY ROW PRICES ONE READING, AND A DEPLOYMENT DOES NOT BUY ONE READING. This watch wakes every business day and re-reads every open fail each time, so the bill is open fails x runs. Multiply any row below by your own open-fail count and by five before comparing it with anything.
⚑ AND EVERY ROW PRICES THE ARM WITH THE CARRIED STATE, WHICH IS THE CHEAPER ONE. The stateless control cost 2.92 times as much per reading on the same card, because removing the memory made the model reason for longer and pushed 6 readings past the output ceiling. Memory is not an overhead here; it is a discount.
⚑ A MISSED RUN REPRICES WHAT FOLLOWS IT. Over the same 41 readings, the ablation spent 1.46 times the output tokens of the intact run. Skipping a run to save money buys back less than it looks like it does, and costs 2 mandatory buy-ins.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
15 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Seven of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 24 open fails x 5 scheduled runs = 112 snapshots from a fixed seed (SEED = 20260823), AND a second rendering of the same book with the third business day never read (89 snapshots). Every gold value is computed from the fail model through src/settlement.step, never read back off the rendered page. Each fail is given a TARGET crossing run and the generator works backwards over that market's calendar to find the intended settlement date that produces it -- because a corpus whose crossings fell where the seed put them could contain none at all on the day the ablation skips, and a missed-run measurement over a day nothing crossed measures nothing.
You change it to: PLAN, the two event pools and SEED. Keep the eight section headings and the two dated calendar blocks and every reader in the kit still parses.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
CORPUS_MISSED = os.path.join(HERE, "data", "corpus-missed-r3")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
GOLD_MISSED = os.path.join(HERE, "data", "gold-missed-r3.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
RULE = "-" * 72
RUN_DATES = ["2026-03-05", "2026-03-06", "2026-03-09", "2026-03-10", "2026-03-11"]
src/marketcal.pythe market calendars — a swap seam
Four invented markets, each with its own holidays, and the two day-counting rules the kit compares: settlement days at the place of settlement, and the watching desk's own Monday-to-Friday days. This file IS the trap the kit measures. Weekends are closed in every market modelled, which is a simplification and is stated.
You change it to: MARKETS is four dictionaries of holidays. Replace them with your own markets, or with a real calendar feed -- everything downstream, including both day-counting floors, reads through settles() and business_days().
src/marketcal.py
# The place-of-settlement calendars. Pure code, standard library, no market data feed.
WEEKEND = (5, 6)
MARKETS = {
MICS = tuple(sorted(MARKETS))
def d(s):
def s(x):
def holiday(mic, day):
def settles(mic, day):
def why_closed(mic, day):
def days(start, end):
src/settlement.pythe buy-in rule — a swap seam
The nine illustrative rules as arithmetic: the running count, the suspensive deduction, the forward projection of the deadline over the market calendar, the freeze once the extension period is reached, and the once-per-fail instruction. step() takes the LIST of settlement days in the window rather than a count, because when the window is wide -- the first run, or any run after a missed one -- the count can cross the extension period part-way through it and the deadline is a day INSIDE the window.
You change it to: RULE_TEXT is what the model reads and step() is what the answer key uses; they are deliberately the same nine rules in two forms. Change the extension periods, the precedence or the Rule S-5 projection convention here and the harness, the UI and all four floors move together.
src/settlement.py
# The buy-in rule as arithmetic. Pure code, integers and dates, no model and no library.
AGEING = "AGEING"
DUE_NEXT_RUN = "DUE_NEXT_RUN"
BUYIN_DUE = "BUYIN_DUE"
SUSPENDED = "SUSPENDED"
CONTEXT_INCOMPLETE = "CONTEXT_INCOMPLETE"
STATUSES = (AGEING, DUE_NEXT_RUN, BUYIN_DUE, SUSPENDED, CONTEXT_INCOMPLETE)
YES = "YES"
NO = "NO"
RAISES = (YES, NO)
src/state.pythe carried state
Four scalars between runs -- the counted age, whether a suspensive event was open, the deadline once frozen, and the previous status -- rendered as one English paragraph. Written from the arithmetic, never from the model's reply, so one bad reading cannot compound into every reading after it.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_fail(store, fail_id):
def describe(state):
src/segment.pythe section splitter
Splits a snapshot into its eight named sections. All eight are asserted present in every document on both schedules before a run may spend.
src/segment.py
# Split a fail snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Fail", "Buy-in Terms", "Ageing Position",
RULE = "-" * 72
def split(text):
def render(secs):
src/select.pythe privacy gate — a swap seam
Decides which sections reach the prompt. Counterparty Desk Contact is mapped by nothing and subtracted unconditionally in the fallback, so it never leaves the machine -- personal data, and a market-sensitive fact about who cannot deliver.
You change it to: SECTION_HINTS and NEVER_SENT. Moving a section out of NEVER_SENT is the only way to widen what is sent, and check_labels.py will convict it.
src/select.py
# Pick which sections of a snapshot are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
FAIL = "Fail"
TERMS = "Buy-in Terms"
POSITION = "Ageing Position"
CALENDAR = "Settlement Calendar"
ACTIVITY = "Activity Log"
CONTACT = "Counterparty Desk Contact"
NOTES = "Operational Notes"
NEVER_SENT = (CONTACT,)
src/prompt.pythe prompt — a swap seam
Three parts: the fixed instruction, the carried-state paragraph, and the sections select.py let through. The middle part is the experiment; --stateless replaces it and changes nothing else, which check_labels.py asserts line by line.
You change it to: INSTRUCTION and the JSON shape. Add or drop an answered field here and in evals/scoring.FIELDS; the four are scored independently.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/watch.pythe runtime
One fail, one scheduled run, one call. Parses the arithmetic inputs off the page with regexes so the scorer can tell a reading failure from a reasoning failure, and holds the published output ceiling.
src/watch.py
# One fail, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
CORPUS_MISSED = os.path.join(HERE, "data", "corpus-missed-r3")
SYSTEM = ("You are a scheduled settlement-fail watch. You apply a written buy-in rule to one open "
MAX_TOKENS = 24000
FIELDS = ("status", "age_settlement_days", "buyin_deadline", "raise_buyin")
def corpus_dir(schedule="intact"):
def documents(schedule="intact"):
def fails(schedule="intact"):
src/adapters/__init__.pythe model adapter — a swap seam
Raw HTTP to any OpenAI-compatible endpoint, plus per-vendor adapters. No SDK, no dependency; transient statuses retried, terminal ones not.
You change it to: PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because the Cost lens prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/budget.pythe spend guard
Counts live calls against a cap read from the shared .env, ledgered before the call is made. Shared by every kit under the same repo root, because they share a key.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
evals/baseline.pythe four free floors — a swap seam
caldays, caldays-mem, bizdays-mem, marketcal-mem -- all $0.00, all scored through the identical scorer. bizdays-mem and marketcal-mem differ in exactly one thing, which calendar the days come from, so the price of the trap is a measured number.
You change it to: MODES plus one branch in review(). A fifth floor is roughly thirty lines and costs nothing to run.
evals/baseline.py
# THE FREE FLOORS. Four of them, none a strawman, all of them $0.00.
MODES = ("caldays", "caldays-mem", "bizdays-mem", "marketcal-mem")
ASSUMED_EXTENSION_DAYS = 7
SUSPENSIVE_WORDS = ("depository", "market operator", "settlement engine", "blackout",
OPEN_MARKERS = ("cannot", "will not", "has not", "have not", "no stock", "holds no",
def _num(pat, text, cast=float):
def _calendar_block(text, heading):
def _log_entries(text):
def _is_suspensive(line):
def _is_open(line):
evals/scoring.pythe scorer
Exact match per cell, no model. Missed buy-ins and premature buy-ins counted apart and never averaged; the deadline scored by direction as well as by hit rate.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("status", "age_settlement_days", "buyin_deadline", "raise_buyin")
DATE_FIELDS = ("buyin_deadline",)
def _pct(n, d):
def _age(v):
def _norm(field, v):
def _daydiff(got, want):
def score(records, golds):
def answers_that_moved(a_cells, b_cells):
evals/cadence.pythe cadence comparison
Compares the intact schedule against one where a business day is never read, twice -- once with the model taken out entirely, once with it. Free. Prints the crossings lost and the cells that moved rather than an accuracy delta, because the accuracy delta is the small half of the damage.
evals/cadence.py
# What a missed run costs -- twice, once without the model and once with it. Free -- no calls.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
PAIRS = [
def _load(run_id):
def _acc(cells, field, keys):
def main():
evals/check_labels.pythe pre-flight checks
Everything that must be true before a run may spend: the answer key replays through the same advance() the harness uses on both schedules, no snapshot restates its own carried state, the privacy guard holds and is red-proven in both directions, the two calendars actually disagree on this corpus, and a skipped run widens the calendar window and NOT the activity log.
evals/check_labels.py
# Everything that must be true BEFORE a run is allowed to spend. Free, and it runs in seconds.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
FAILED = []
def check(name, ok, detail=""):
def load(path):
def main():
src/app.pythe local UI
http.server, no framework. Shows the carried state verbatim beside the answer and the free floor beside both, and carries a SCHEDULE SELECTOR that re-reads the same fail off the ablation corpus -- so a reader can watch a deadline move because nobody looked on Monday.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "9014"))
RECORDED = {"intact": os.environ.get("RECORDED_RUN", "r001-settle-fail"),
GOLD = {s: RUN.load_gold(s) for s in RUN.SCHEDULES}
SCHEDULE_LABEL = {"intact": "all 5 runs made",
def carried_for(doc_id, schedule):
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 24 open fails x 5 scheduled runs = 112 snapshots from a fixed seed (SEED = 20260823), AND a second rendering of the same book with the third business day never read (89 snapshots). Every gold value is computed from the fail model through src/settlement.step, never read back off the rendered page. Each fail is given a TARGET crossing run and the generator works backwards over that market's calendar to find the intended settlement date that produces it -- because a corpus whose crossings fell where the seed put them could contain none at all on the day the ablation skips, and a missed-run measurement over a day nothing crossed measures nothing. A swap seam.
src/marketcal.pyFour invented markets, each with its own holidays, and the two day-counting rules the kit compares: settlement days at the place of settlement, and the watching desk's own Monday-to-Friday days. This file IS the trap the kit measures. Weekends are closed in every market modelled, which is a simplification and is stated. A swap seam.
src/settlement.pyThe nine illustrative rules as arithmetic: the running count, the suspensive deduction, the forward projection of the deadline over the market calendar, the freeze once the extension period is reached, and the once-per-fail instruction. step() takes the LIST of settlement days in the window rather than a count, because when the window is wide -- the first run, or any run after a missed one -- the count can cross the extension period part-way through it and the deadline is a day INSIDE the window. A swap seam.
src/state.pyFour scalars between runs -- the counted age, whether a suspensive event was open, the deadline once frozen, and the previous status -- rendered as one English paragraph. Written from the arithmetic, never from the model's reply, so one bad reading cannot compound into every reading after it.
src/segment.pySplits a snapshot into its eight named sections. All eight are asserted present in every document on both schedules before a run may spend.
src/select.pyDecides which sections reach the prompt. Counterparty Desk Contact is mapped by nothing and subtracted unconditionally in the fallback, so it never leaves the machine -- personal data, and a market-sensitive fact about who cannot deliver. A swap seam.
src/prompt.pyThree parts: the fixed instruction, the carried-state paragraph, and the sections select.py let through. The middle part is the experiment; --stateless replaces it and changes nothing else, which check_labels.py asserts line by line. A swap seam.
src/watch.pyOne fail, one scheduled run, one call. Parses the arithmetic inputs off the page with regexes so the scorer can tell a reading failure from a reasoning failure, and holds the published output ceiling.
src/adapters/__init__.pyRaw HTTP to any OpenAI-compatible endpoint, plus per-vendor adapters. No SDK, no dependency; transient statuses retried, terminal ones not. A swap seam.
src/budget.pyCounts live calls against a cap read from the shared .env, ledgered before the call is made. Shared by every kit under the same repo root, because they share a key.
evals/baseline.pycaldays, caldays-mem, bizdays-mem, marketcal-mem -- all $0.00, all scored through the identical scorer. bizdays-mem and marketcal-mem differ in exactly one thing, which calendar the days come from, so the price of the trap is a measured number. A swap seam.
evals/scoring.pyExact match per cell, no model. Missed buy-ins and premature buy-ins counted apart and never averaged; the deadline scored by direction as well as by hit rate.
evals/cadence.pyCompares the intact schedule against one where a business day is never read, twice -- once with the model taken out entirely, once with it. Free. Prints the crossings lost and the cells that moved rather than an accuracy delta, because the accuracy delta is the small half of the damage.
evals/check_labels.pyEverything that must be true before a run may spend: the answer key replays through the same advance() the harness uses on both schedules, no snapshot restates its own carried state, the privacy guard holds and is red-proven in both directions, the two calendars actually disagree on this corpus, and a skipped run widens the calendar window and NOT the activity log.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2550 input and 1527 output tokens per reading (one open fail, at one scheduled run), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one open fail, at one scheduled run)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one open fail, at one scheduled run) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Operational Notes, which are one of three fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- prose a settlements clerk types into the exceptions screen -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser, and tools/shoot_ui.mjs blanks the key entirely and refuses to start if something is already listening on the port, so no screenshot can spend.
The experimentWe did NOT attack it — and this is the surface we left open
An indirect prompt injection needs a field somebody outside can write into, and this kit has exactly one: the Operational Notes, prose a settlements clerk types into the exceptions screen. It is SENT, on purpose, and one of the three shipped sentences is instruction-shaped for exactly that reason -- "The counterparty's settlements desk has asked to be told before any buy-in is instructed on this line". On this corpus that sentence comes from a seeded generator, so there is nothing adversarial in it to catch and no trial is claimed. A version pointed at real clerk prose reopens the question and should be attacked before it ships. No attack was fired at this kit and none is claimed. The three boundaries below were checked in code on 2026-08-23, and only the first of them is red-proven in both directions rather than argued from an absent code path.
Boundary checked
What could go wrong
What this kit does
Does a counterparty's settlements contact -- their name, direct line, mobile and email -- ever leave the machine?
Every snapshot carries a Counterparty Desk Contact section. The naive or list(secs) fallback every sibling kit once shipped sends the WHOLE document when no mapped section parses -- which is what a depository upgrade that renames the report's blocks produces. Reproduced in evals/check_labels.py: 112 of 112 documents leak with that fallback.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a feed whose blocks are all renamed still cannot send it. Measured before every run: 0 of 112 leak with the guard, 112 of 112 without it.
Can anything here instruct, execute or settle a buy-in?
A monitor that can act on its own reading of a deadline can force a purchase in the market off a mis-parsed calendar line -- and this kit is demonstrably capable of reading a blind window and answering INSTRUCT IT NOW with total confidence. The missed-run screenshot on this page is that exact reply.
There is no such endpoint, no such button and no configuration flag. evals/check_labels.py greps every .py and .js file for the names of such paths (instruct_buyin, execute_buyin, place_order, cancel_trade, release_cash, and three def-name patterns) and passes at zero. Nothing outside results/*.json and data/state.json is written at all.
Can a run be started that spends more than the operator expects?
One key funds every kit in this repository. A loop with a bad exit condition, a --limit left off, or two terminals at once, and the first sign is the provider's dashboard tomorrow.
src/budget.py counts live calls against MAX_CALLS_PER_DAY from the shared .env, ledgered BEFORE the call is made so a crash over-counts rather than under-counts, and shared across every kit under the same root. Every paid run prints what it is about to spend and says out loud when no cap is configured. --max-tokens refuses to run on any run id that does not begin with c.
The first row is the only one measured in both directions. The other two are guarantees about what is ABSENT from the code, and an absent path is checked by grepping for the names we know -- a path called something else would pass.
The result0 attack trials, three boundaries checked -- and the privacy boundary red-proven by removing the guard and watching all 112 documents leak.
1field an outside party could influence (sent, not hidden)
0 of 0attack trials run
1 of 3boundaries red-proven, not just asserted
0 of 112snapshots leak a counterparty contact's name, line or email
The Operational Notes ARE the field an outside party would influence in a real deployment, and this kit sends them rather than hiding them -- one of the three shipped sentences is even instruction-shaped, chosen for exactly that reason. But on this corpus they are one of three sentences from a seeded generator, so there is nothing adversarial in them to catch. A version pointed at real clerk prose reopens the question and should be attacked before it ships. What IS measured is the privacy boundary, in both directions, before every run.
Read this twice
The activity log and the operational notes reach the model verbatim, and one of the three shipped notes is instruction-shaped. That is deliberate: hiding the one field an outsider could write into would hide the surface from the run that is supposed to measure it. It also means a real deployment of this kit is reading attacker-writable prose next to a rule that decides whether to force a purchase in the market.
HonestyWhat this does not prove
Whether a real deployment's Operational Notes -- prose a settlements clerk writes freely -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at a real book. The shipped note asking to be told before any buy-in is instructed is the shape to probe.
Whether the privacy guard covers anything beyond the section it names. It subtracts Counterparty Desk Contact unconditionally, and it is red-proven in both directions -- but a counterparty's identity also appears as a reference in the Fail section, which IS sent. A reference is not a name; whether it is re-identifiable against a real book is not something this repository can test.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two watches on one book have never been run and would race. On a RUNNING COUNT a lost run is a permanently wrong deadline rather than a missing row.
Whether a truncated reply can ever be partially trusted. The stateless control lost 6 calls to the output ceiling and src/watch._parse rejected all of them outright; nothing here has looked at whether a cut-off reply's first field was usable.
Whether the absence checks catch a path named something else. evals/check_labels.py greps for eight names it knows. A function called send_to_agent would pass it.
What the provider retains. The prompt carries no personal data by construction, but what happens to a request after it leaves this machine is outside this repository.
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No buy-in instruction, non-configurable. This kit produces a watchlist status, a counted age, a deadline date and an instruct-or-hold call for a settlements desk to validate. It never instructs a buy-in agent, places an order, cancels a trade or releases cash, and there is no setting that makes it.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). src/watch.py states the rule where the calls are made rather than only on a page.
EvidenceDoes it hold?
What
Measured
Nothing in this kit instructs, executes or settles a buy-in
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for instruct_buyin, execute_buyin, place_order, cancel_trade, release_cash and three def-name patterns, and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
A counterparty settlements contact's name, direct line, mobile and email never leave the machine
0 of 112 snapshots leak with the guard; 112 of 112 leak with the naive or list(secs) fallback, reproduced in evals/check_labels.py by renaming the report's blocks the way a depository upgrade would.
No default extension period is ever assumed (Rule S-2)
100.0 pct context-incomplete recall on the 10 readings whose agreement carries no extension period or whose feed captured no intended settlement date. The b000 and b001 floors guess seven days and score 0.0 pct, which is what the rule exists to forbid.
The model's reply never becomes the next run's input
src/settlement.step writes the carried state from the arithmetic. evals/check_labels.py replays every chain on both schedules through the same advance() the harness uses and requires the committed key exactly: 0 mismatches over 201 readings.
A run says what it is about to spend, and refuses past the cap
src/budget.py ledgers each call BEFORE it is made and is shared across every kit under the same repo root. Every paid run prints the plan line, and says out loud when no cap is configured rather than being silent about it.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The arithmetic is correct whatever the model says, which means a wrong reading is a wrong watchlist row -- and on this task a wrong watchlist row can read INSTRUCT THE BUY-IN NOW with a rationale that cites two rules by name.
IT IS NOT A DEFENCE AGAINST A MISSED RUN. The guardrails are about what the kit will not do. The most expensive failure measured on this page is something nobody did at all: 2 mandatory buy-ins worth $8,315,720.00 vanished from the record because one scheduled run was not made.
IT IS NOT AN INJECTION DEFENCE. The Operational Notes are attacker-writable in any real deployment and are sent verbatim, deliberately, so the surface is visible. Nothing here has been attacked.
WatchedWhat is watched, and why that one
10runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 35 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
16 measured by the latest run19 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The status, the counted age, the buy-in deadline date and the instruct-or-hold call, per reading, exact match against the computed key
alarm
status_accuracy_pct; age_accuracy_pct; buyin_deadline_accuracy_pct; raise_buyin_accuracy_pct; missed_buyins; premature_buyins; unparsed_replies — alarm on unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. The stateless control is exactly that case: 6 of its 112 readings came back unusable.
guardrail-context-incomplete
Context-incomplete recall -- the guardrail, scored
alarm
context_incomplete_recall_pct — alarm on anything below 100. A default extension period is the one thing Rule S-2 forbids.
cadence-crossings
What a missed run costs, in crossings rather than in points
alarm
crossings_lost_permanently; crossings_seen_late; moved_right_to_wrong; runs_missed_detected — alarm on crossings_lost_permanently above 0 -- because that number can never be recovered and never shows up in any accuracy figure.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
112
different corpus — nothing is comparable
corpus.bytes
713,762
fail snapshots edited — the count held, the bytes did not
split.count
24
the fails count moved — a different set was scored
split.size_p50
5
the median size of one fail moved
split.size_p95
5
the 95th-percentile size of one fail moved
dataset.rows
112
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cadence every business day, 18:00, at the watching desk, fails 24, stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
buy-in deadline, exact date
not yet known
112 readings
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. r001, s001 and a001 are one run each of three DIFFERENT conditions, so none of them bounds the others.
status and age
not yet known
112 readings, 80 of them memory-dependent
No repeat run. Same reason as the deadline band.
missed buy-ins
not yet known
18 readings whose correct answer is to instruct
0 on one run. A count of zero over eighteen cells bounds nothing.
premature buy-ins
not yet known
94 readings whose correct answer is not a live requirement
0 on one run.
the guardrail
not yet known
10 readings with no extension period or no ISD on file
100 pct on one run over ten readings. Ten cells bound very little.
readings that returned an answer at all
100.00 with the carried state, 94.64 without it
112 readings
MEASURED across two arms: r001 lost none, s001 lost 6 to the 24000-token output ceiling.
latency
p50 9815-22304 ms and p95 36665-162950 ms across the three paid arms
readings that returned an answer, at 16 concurrent fail chains
MEASURED, but across three DIFFERENT conditions rather than three runs of one -- so it is a range, not a variance band. The spread is caused by the prompt, not by the provider: the stateless arm reasons far longer.
tokens, and therefore the bill
output 171,079-592,596 tokens over a full 112-reading arm
112 readings per full arm
MEASURED on r001 and s001, which are the same corpus and the same model with one prompt block different. Input barely moves (285,656 against 278,509); output is 3.46x.
crossings lost to a missed run
2 on one skipped business day
18 buy-in requirements on this book
MEASURED, and free: read off the two answer keys by tools/build_corpus.py and printed by evals/cadence.py. One day was skipped, so this is one draw rather than a distribution.
HistoryRun history
10 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 5 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
a000-settle-fail-missedrun-floor 2026-08-23
b000-settle-fail-caldays 2026-08-23
b001-settle-fail-caldaysmem 2026-08-23
b002-settle-fail-bizdaysmem 2026-08-23
b003-settle-fail-marketcalmem 2026-08-23
age accuracy, %
93.26
16.96
16.96
83.93
97.32
answered, %
100.0
100.0
100.0
100.0
100.0
buyin deadline accuracy, %
92.13
8.93
8.93
75.00
95.54
context incomplete recall, %
100.0
0.0
0.0
100.0
100.0
exposure usd error, %
10.44
66.66
66.66
8.63
4.00
input tokens, whole run
0
0
0
0
0
model latency p50 ms
0.00
0.00
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
0.00
0.00
memory age accuracy, %
89.83
7.50
7.50
80.00
96.25
memory deadline accuracy, %
89.83
5.00
5.00
76.25
95.00
memory status accuracy, %
94.92
68.75
68.75
90.00
96.25
missed buyin, %
0.0
0.0
0.0
0.0
0.0
output tokens, whole run
0
0
0
0
0
premature buyin rate, %
2.74
37.23
37.23
5.32
2.13
raise buyin accuracy, %
97.75
39.29
68.75
95.54
98.21
status accuracy, %
96.63
59.82
59.82
90.18
97.32
not a time series No two of these 5 runs measured the same system — they differ on context_incomplete_cells, crossing_cells, documents, floor, memory_cells, quiet_cells, readings_scored, run_skipped, runs_missed_detected, schedule, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 4 runs. Columns here are only ever compared with each other.
Metric
a001-settle-fail-missedrun 2026-08-23
c000-settle-fail-calibration 2026-08-23
r001-settle-fail 2026-08-23
s001-settle-fail-stateless 2026-08-23
age accuracy, %
92.68
100.00
100.00
85.71
answered, %
100.00
100.00
100.00
94.64
buyin deadline accuracy, %
95.12
100.00
100.00
85.71
context incomplete recall, %
100.0
—
100.0
100.0
exposure usd error, %
0.46
0.00
0.00
5.56
input tokens, whole run
104607
25622
285656
278509
model latency p50 ms
13481.00
24798.00
9815.00
22304.00
model latency p95 ms
95132.00
42978.00
36665.00
162950.00
memory age accuracy, %
91.89
100.00
100.00
86.49
memory deadline accuracy, %
94.59
100.00
100.00
86.49
memory status accuracy, %
94.59
100.00
100.00
94.59
missed buyin, %
0.0
0.0
0.0
0.0
output tokens, whole run
105302
25685
171079
592596
premature buyin rate, %
3.12
0.00
0.00
2.13
raise buyin accuracy, %
97.56
100.00
100.00
68.75
status accuracy, %
95.12
100.00
100.00
91.07
not a time series No two of these 4 runs measured the same system — they differ on context_incomplete_cells, crossing_cells, documents, fails, max_tokens, memory_cells, quiet_cells, readings_scored, run_skipped, runs_missed_detected, schedule, stateless — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-settle-fail-stub 2026-08-23
age accuracy, %
16.96
answered, %
100.0
buyin deadline accuracy, %
8.93
context incomplete recall, %
0.0
exposure usd error, %
66.66
input tokens, whole run
297464
model latency p50 ms
0.00
model latency p95 ms
0.00
memory age accuracy, %
7.5
memory deadline accuracy, %
5.0
memory status accuracy, %
68.75
missed buyin, %
0.0
output tokens, whole run
4756
premature buyin rate, %
37.23
raise buyin accuracy, %
39.29
status accuracy, %
59.82
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 16 chips that all say so.
DeviationsWhat deviated
0 breaches across 10 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the carried state
deadline 85.71 pct -> 100.0 pct, age 85.71 -> 100.0, status 91.07 -> 100.0, instruct-or-hold 68.75 -> 100.0, duplicate buy-in instructions 0 -> 29, readings that returned nothing 0 -> 6, output tokens 592,596 -> 171,079
measured
r001-settle-fail against s001-settle-fail-stateless. Same corpus, same model, same grader, same 112 readings; the prompts differ in exactly one block and evals/check_labels.py asserts they differ on exactly one line.
whether a scheduled run is made
2 buy-in requirements lost permanently and 3 instructed late; free-floor deadline 97.75 pct -> 92.13 pct and age 100.0 -> 93.26 over the readings both schedules took; model deadline 100.0 pct -> 95.12 pct and age 100.0 -> 92.68 over the 41 readings both arms took; 16 and 8 cells right-to-wrong respectively, with 0 wrong-to-right on either
measured
evals/cadence.py over b003 against a000 (free) and r001 against a001 (41 paid calls). The crossings figures come from the two answer keys and cost nothing.
b002-settle-fail-bizdaysmem against b003-settle-fail-marketcalmem. Free, no model. The two floors are the same code but for which calendar supplies the days, so the whole gap is attributable to that one choice.
whether the rule may assume a default extension period
context-incomplete recall 0.0 pct -> 100.0 pct over 10 readings, and 35 premature buy-ins -> 2 across the book
measured
b000-settle-fail-caldays against b003-settle-fail-marketcalmem. Both free. b000 guesses seven days where the agreement is silent, which is what a threshold report configured once and forgotten does.
the output ceiling
readings that returned nothing 0 -> 6, but only on the stateless arm. The stateful arm's largest reply was 6594 tokens against the same 24000 ceiling.
measured
c000-settle-fail-calibration set the ceiling from the two hardest chains under an 8,000-token probe (largest reply 5043); r001 and s001 both ran at 24000.
how far forward the page's calendar block reaches
the free floor abstains with NOT_DETERMINABLE on 2 readings whose remaining extension runs past the end of it; the model extrapolates the weekly pattern and gets both dates right
measured
results/eval-b003-settle-fail-marketcalmem.json, FAIL-0022-R1 and FAIL-0022-R2, against r001. Lengthening the block widens every prompt, so it is an input-token decision rather than a free one -- and that trade was not measured.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
buy-in deadline, exact date
nothing yet. ⚠︎ And note what this figure is level with: the strongest free floor scores 95.54 pct on the same 112 readings, so a one-cell wobble is most of the margin.
status and age
nothing yet. The strongest free floor is at 97.32 pct on status and 97.32 on age over the same readings.
missed buy-ins
anything above 0 should. A missed buy-in is a mandatory purchase left un-instructed past a deadline that does not move.
premature buy-ins
anything above 0. The aged-fails report a desk runs today produces 35 of them on this book, and reports $132,560,636 of exposure as buy-in-due against a true $79,540,506.
the guardrail
anything below 100. A default extension period is the one thing Rule S-2 forbids, and the b000 and b001 floors score 0.0 pct here.
readings that returned an answer at all
anything below 100. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure.
latency
nothing automatically. A p95 climbing towards the stateless arm's 163s is a sign the model is being asked something it cannot resolve.
tokens, and therefore the bill
output climbing towards the stateless figure. On this kit that is not a cost signal, it is a correctness signal: the model reasons longest when the answer is not reachable from what it was given.
crossings lost to a missed run
any value above 0 is unrecoverable by definition -- and this one does not appear in any accuracy figure on the page, which is the whole reason it is a band of its own.
NextThe three you would add first
A scheduler, and something that refuses to publish a reading taken through a blind windowThis is the half of a monitor the kit does not ship. It DETECTS a missed run for free -- 21 of 41 readings on the ablation schedule -- and then answers anyway. A deployment should mark those readings provisional and go and get the feed.
An exceptions feed that retains more than one dayThe single change that would remove the whole arithmetic half of the missed-run cost. The crossings lost permanently would still be lost -- those need the run itself.
A maintained market-calendar sourceA stale holiday produces exactly the defect the b002 floor measures: a deadline that is right on every day two calendars agree and early on every foreign holiday.
Two-person review before any buy-in is instructed downstreamA buy-in is a forced purchase at market. The kit is decision-free by construction; whatever consumes its watchlist must not be.
A red-team pass over the Operational NotesThe surface is real, named and untested. 0 trials were run and none is claimed.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free, under a second) on any change to tools/build_corpus.py, src/settlement.py, src/marketcal.py, src/segment.py or src/select.py -- it refuses to let a run spend if the answer key stops replaying, if a snapshot starts restating its own carried state, if the privacy guard stops holding, or if the two calendars stop disagreeing on this corpus. Re-run evals/cadence.py (also free) on any change to the schedule.
What this cannot tell you
One run per arm. Whether a 100 pct on all four fields is stable across repeats is not measured, and on a 112-reading book one cell either way is most of the margin over free code.
Whether any of the five guardrail holds survives a real deployment. Four of them are properties of absent code, checked by grepping for names we know.
What skipping a different business day would cost. One was ablated.
Whether a desk seeing the blind-window banner recovers the answer. The kit reports the gap; the reading it reports it on is already wrong.
Whether the Operational Notes can move raise_buyin. 0 attack trials.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and the emptiness is load-bearing: a forker runs this on whichever key they already hold, without installing a client for a vendor they will never call.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a COUNTER -- four scalar fields written by arithmetic. The cost is flat in history length because nothing here remembers what was said, only what was computed, and that is why removing it makes the kit MORE expensive rather than less.
the schedule
evals/run.py
a scheduler or workflow engine (Airflow, Temporal, cron plus a state store)
the kit is INVOKED, not woken, and that is the honest gap. What it does ship that a bare cron does not is detection: every reading checks the page's own 'previous scheduled run' line against the cadence and reports a gap, free. On the ablation schedule that fired on 21 of 41 readings.
the market calendars
src/marketcal.py
a business-calendar library (pandas market calendars, workalendar, a vendor feed)
four dictionaries and two functions. A real deployment should use a maintained calendar -- getting a holiday wrong is precisely the defect this kit measures -- and everything downstream reads through settles() and business_days(), so it is one file to replace.
the model call
src/adapters/__init__.py
a provider SDK or a router (LiteLLM, OpenRouter, a vendor client)
raw HTTP with a per-vendor adapter and transient-versus-terminal retry. A router would buy failover and lose the property that the whole kit installs nothing.
the grader
evals/scoring.py
an eval framework (promptfoo, DeepEval, Braintrust, an LLM-judge harness)
nothing here is judged by a model. The answer is a status from a closed list, an integer, an ISO date and a yes/no, all computed by the corpus generator -- so the grader is an equality test and a framework would add a dependency and a service to run it.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each fail is a chain of up to five readings -- one per business day -- with no branching and no retries beyond the transport layer. Different fails are independent, which is where the concurrency is. The only edge that matters is the one carrying four scalars from one reading to the next, and it carries them through pure arithmetic rather than through the model.
The other sideWhat a framework costs you
No scheduler. This is a watch that carries state and is still INVOKED, not woken -- and the cadence is the single most important thing about it. That is the honest cost of shipping no framework, and the kit narrows it by DETECTING a missed run rather than by preventing one.
No calendar service. Four invented markets in a Python dict. A production book settles in dozens of markets whose holidays change, and a stale calendar produces exactly the defect b002 measures.
No persistence beyond a JSON file. data/state.json is one writer, replaced atomically, and nothing here has been tested with two.
No retry policy above the transport. A reading that errors advances the count anyway -- deliberately, so one transport error does not turn into four scored failures -- but nothing re-attempts it.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
Whether a checkpointer-backed memory would score differently from four scalars. Not attempted; the comparison that WAS run is four scalars against nothing at all.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-settle-fail on the same tier, one business day never read (THE ABLATION), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
9,815 ms
p50 9815-22304 ms and p95 36665-162950 ms across the three paid arms
nothing automatically. A p95 climbing towards the stateless arm's 163s is a sign the model is being asked something it cannot resolve.
Model, p95
36,665 ms
p50 9815-22304 ms and p95 36665-162950 ms across the three paid arms
nothing automatically. A p95 climbing towards the stateless arm's 163s is a sign the model is being asked something it cannot resolve.
Input tokens
285,656
output 171,079-592,596 tokens over a full 112-reading arm
output climbing towards the stateless figure. On this kit that is not a cost signal, it is a correctness signal: the model reasons longest when the answer is not reachable from what it was given.
Output tokens
171,079
output 171,079-592,596 tokens over a full 112-reading arm
output climbing towards the stateless figure. On this kit that is not a cost signal, it is a correctness signal: the model reasons longest when the answer is not reachable from what it was given.
No movement column. Not one of the 8 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
a001-settle-fail-missedrun13,481 ms
c000-settle-fail-calibration24,798 ms
r001-settle-fail9,815 ms
s001-settle-fail-stateless22,304 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
5 runs not plotted. a000-settle-fail-missedrun-floor, b000-settle-fail-caldays, b001-settle-fail-caldaysmem, b002-settle-fail-bizdaysmem, b003-settle-fail-marketcalmem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 10 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
fail snapshots
data/corpus/FAIL-<n>-R<k>.txt -- 112 files, and a second rendering of the same book under data/corpus-missed-r3/ with one business day never read (89 files). 713,762 bytes in total, generated from a fixed seed by tools/build_corpus.py.
seven of the eight sections go to the provider in the prompt; Counterparty Desk Contact -- the settlements contact's name, direct line, mobile and email, and with it the identity of who cannot deliver -- never does, by src/select.NEVER_SENT
the market calendars
src/marketcal.py -- four invented markets and their holidays, in Python. There is no calendar service and no network call; the relevant days are printed into each snapshot as two dated blocks so the reading is self-contained
as those two dated blocks, inside the Settlement Calendar section of every prompt -- roughly twenty lines a reading, and the reason a deadline can be projected at all
the carried state
data/state.json in a deployment, and in-memory per run in the eval -- four scalars per fail, written by src/settlement.step and never by a model
as one English paragraph in every prompt, and the UI prints the same paragraph verbatim so a reader can audit what the model was told
the answer keys
data/gold.jsonl (112 rows) and data/gold-missed-r3.jsonl (89 rows) -- the output of src/settlement.step over the planted intervals, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the run records
results/eval-*.json in the kit -- eleven of them, including four free floors, a free missed-run floor arm and the cadence comparison, all of which cost nothing
never -- they are read by the site build, not by a provider
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser, and tools/shoot_ui.mjs blanks the key entirely and refuses to start if something is already listening on the port, so no screenshot can spend.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a DAILY watch: once every business day at 18:00, after the settlement cycle has closed. One run owns exactly the window since the last one, and the snapshot's Settlement days at <market> in this window line is that window measured against the place-of-settlement calendar -- which is why a Monday run over a weekend still owns one settlement day and not three. The cadence is not a deployment detail here: the DUE_NEXT_RUN band is DEFINED as the cadence (Rule S-9, "the deadline falls on or before the next scheduled run"), so changing the schedule changes the answer as well as the bill.
24 fails over 5 business days = 112 readings per full arm, in strict per-fail order; evals/check_labels.py asserts every sequence is ordered and starts at run 1 before a run may spend. Wall clock 207.7s at 16 workers. A book of 500 open fails on this cadence is 2,500 readings a week, $14.64 on the shared projection card. (r001-settle-fail, a001-settle-fail-missedrun, evals/cadence.py, evals/check_labels.py)
⚑ WHAT A MISSED RUN COSTS IS MEASURED HERE, NOT ARGUED. The corpus ships twice and evals/cadence.py compares them, free. Skip ONE business day and: 5 of the 18 buy-in requirements on this book fell on that day; 3 were instructed a day late; 2 are GONE PERMANENTLY, worth $8,315,720.00 of residual exposure, because those fails settled in full before the next run and no reading on any schedule records that a buy-in was required. The accuracy barely moves -- which is the finding, not a reassurance: a monitor grades itself on the readings it took. On the arithmetic side 10 activity entries become unreachable, 5 settlement days are counted as ageing that never aged the fail, and 16 of the free floor's 21 moved cells moved from right to wrong with not one moving the other way.
every figure on this page is per a once-a-business-day watch. A twice-daily watch pays double and narrows the DUE_NEXT_RUN band to a few hours; a weekly watch pays a fifth and would have lost far more than two crossings on this book. Neither was run, and the band moving with the schedule means the two are not comparable runs of one question.
state
four scalars per fail -- the counted age in settlement days, whether a suspensive event was open at the end of the last run, the buy-in deadline once Rule S-6 has frozen it, and the previous status -- written by src/settlement.step from the arithmetic and rendered to the prompt as one English paragraph by src/state.describe. Never the previous snapshot, never the previous reply.
80 of 112 readings are memory-dependent: their correct answer is not reachable from their own snapshot. With the state: 100.0 pct on the deadline over those readings. Without it (s001, the same prompt with that one block replaced): 86.49 pct, 29 duplicate buy-in instructions, and 6 readings that ran off the output ceiling and returned nothing. (r001-settle-fail against s001-settle-fail-stateless, src/state.py)
The state is flat in history length by construction -- four fields on run 500 as on run 2 -- so there is no ceiling to hit on size. The ceiling is on TRUST: because the state is written from the arithmetic and not from the reply, a wrong reading cannot poison the next one, which is also why the missed-run damage stays confined to the readings taken through the blind window. A monitor that fed its own verdict forward would compound it, and nothing here measures how badly.
everything the kit claims about monitoring. Remove the paragraph and this is an ordinary one-shot classifier over a single snapshot, which is exactly what the control run is and exactly why it is published beside the scored one.
model
one model, one key, one call per reading, over an OpenAI-compatible endpoint by raw HTTP. max_tokens 24000, no temperature sent, no thinking parameter sent. 16 concurrent fail chains; a fail's own runs are strictly serial.
112 readings, 100.0 pct on all four fields, p50 9815ms and p95 36665ms end to end, $0.00585772 a reading on the shared projection card. Largest reply 6594 output tokens against the 24000 ceiling. (r001-settle-fail, c000-settle-fail-calibration, src/watch.MAX_TOKENS)
The ceiling is generous on purpose and it is not decorative: the stateless arm hit it 6 times and returned nothing for those readings, billed in full. A ceiling costs nothing until an arm starts reaching it, and then it costs the whole reading.
every accuracy and every dollar. One tier was run; no cross-model number is published because none was measured.
labels
a computed answer key, not a hand-authored one. tools/build_corpus.py produces data/gold.jsonl from the fail model through src/settlement.step, and a second key for the missed-run schedule in which only the run that must instruct each buy-in moves. Four fields per reading, exact match, no model grades anything.
112 rows and 89 rows. evals/check_labels.py replays every chain on both schedules through the same advance() the harness uses and requires it to reproduce the committed key exactly: 0 mismatches. It also asserts that the on-page "settlement days since ISD" figure is NOT the counted age on 19 readings, so a reader that copies it is wrong. (tools/build_corpus.py, evals/check_labels.py, data/gold.jsonl)
The key is correct BY CONSTRUCTION on this corpus, which is its strength and its limit: it says nothing about whether the nine illustrative rules are the right rules, and Rule S-5's one-day projection convention is invented. Every deadline on a suspended fail inherits it.
nothing on this page transfers to a book whose labels are contested -- which is the normal state of a real settlements book, where whether the market was shut or the counterparty was short is exactly what the dispute is about.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a fail reported BUY-IN DUE whose activity log shows the market was shut, not the counterparty
something attributed a market-wide shutdown to the delivering party and aged the fail a day it did not age. That is the direction that forces a purchase nobody was entitled to make
read the exceptions line, not the number. On FAIL-0016-R3 the free keyword floor reads "neither we nor the counterparty can put this through: the line has been shut to all instructions" as the counterparty's, ages the fail to 7 of 7, and answers BUY-IN DUE with the deadline pulled to 2026-03-09. The key says SUSPENDED at 6 with 2026-03-11 (results/eval-b003-settle-fail-marketcalmem.json against results/eval-r001-settle-fail.json, FAIL-0016-R3)
a buy-in deadline that is EARLIER than the desk expected, on a fail settling abroad
the days were counted on the wrong calendar. A rule using the watching desk's Monday-to-Friday days instead of the place of settlement's settlement days is right on every day the two agree and early on every foreign market holiday
compare the two floors on the same reading. b002 and b003 differ in exactly one thing -- which calendar -- and b002 gets 28 deadlines wrong against b003's 5, 26 of them landing early by 49 days in total (results/eval-b002-settle-fail-bizdaysmem.json against results/eval-b003-settle-fail-marketcalmem.json)
the same buy-in instructed on a fail that already has one open
the carried state is not reaching the prompt. Rule S-7 is a once-per-fail rule and once-per-anything cannot be answered from one snapshot
check that src/state.describe's paragraph is in the prompt. Removing it produced 29 duplicate instructions across 94 readings whose correct answer was to hold (results/eval-s001-settle-fail-stateless.json)
a reading whose page says the previous scheduled run was more than one business day ago, or whose activity log is empty when the window is two settlement days wide
a scheduled run was missed and this reading is being taken through a blind window. The exceptions feed keeps one day, so whatever was raised on the unread day is gone
do not trust the age or the deadline on that reading. The kit flags it for free -- 21 of 41 readings on the ablation schedule -- and the UI shows an amber banner. Recovery needs the feed, not the model (results/eval-cadence-settle-fail.json, results/eval-a001-settle-fail-missedrun.json)
['Concurrency. data/state.json is replaced atomically, which is correct for one writer; two watches on one book is two writers and nothing here has tested it.', 'A cadence other than once a business day, and therefore any DUE_NEXT_RUN figure at another interval.', 'Skipping a business day OTHER than the third. One day was ablated; runs 2 and 4 would hit different fails and a different number of crossings.', 'Delivery fails -- where the desk is the party being bought in AGAINST. Every fail in this corpus is a receive fail.', 'A market whose settlement week is not Monday to Friday. src/marketcal.WEEKEND is a constant.', 'Whether rendering the carried state as JSON rather than English changes anything. It is a design choice and the page says so.', 'Run-to-run variance. No reading was fired twice, so a 100 pct could hide an unstable reading.', 'A second model or a second tier. One was run.', 'Provider-side retention. The prompt carries no personal data by construction, but what the provider keeps of a request is outside this repository and nobody here has verified it.', 'GPU sizing, local inference and anything about running this off a hosted API. Not attempted; not costed.', 'Whether the Operational Notes field can move raise_buyin. The surface is sent deliberately and was never attacked.']
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every field was written by the generator committed beside it. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The status, the counted age, the buy-in deadline date and the instruct-or-hold call, per reading, exact match against the computed key
Catch failing trades before the buy-in deadline
PresenterOpens the private repo. Visible to admins only.
In one lineThe status, the counted age, the buy-in deadline date and the instruct-or-hold call, per reading, exact match against the computed key
For each of the 112 readings and each of the four answered fields, did the reply equal the value tools/build_corpus.py computed from the fail model.
$0.00per 1,000 fail snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function all four free floors, the stateless control and the missed-run ablation are scored through, so the seven arms are comparable by construction.
Every grader on these pages scored the same 112 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
FAIL-0016-R3 -- scheduled run 3 of 5, 2026-03-09, Corporate debt at NGX, extension period 7 settlement days, residual 5,000 at 42.34 USD (211,700.00 USD of exposure)
What the previous scheduled run left behind
As at the previous scheduled run, 6 settlement day(s) had already been counted against this fail at its place of settlement, suspensive days already excluded. No suspensive event was open when that run ended. The extension period has not yet been found exhausted, so no deadline has been fixed and no buy-in has been instructed. It was reported DUE_NEXT_RUN.
The exceptions feed, this snapshot day only
07:41 neither we nor the counterparty can put this through: the line has been shut to all instructions since the notice went out this morning
The answer key
SUSPENDED, aged 6 settlement days, buy-in deadline 2026-03-11, instruct NO
What run r001 answered
SUSPENDED, aged 6 settlement days, buy-in deadline 2026-03-11, instruct NO
It is the reading where a wrong attribution does not merely move a number -- it flips the action. The one exceptions line in this window says the line was shut to ALL instructions, and says it in counterparty language: "neither we nor the counterparty can put this through". Read as the counterparty's fault, the day ages the fail, the count reaches the 7-day extension period, Rule S-6 freezes the deadline three days early at 2026-03-09 and Rule S-7 says instruct the buy-in TODAY. Read correctly it is a suspensive day, the count stays at 6 of 7, the fail is SUSPENDED and the deadline is still ahead at 2026-03-11. One sentence, $211,700.00 of residual exposure, and the difference between holding and forcing a purchase in the market.
Grader
Verdict
Why
The status, the counted age, the buy-in deadline date and the instruct-or-hold call, per reading, exact match against the computed key
status hit, age hit, deadline hit, instruct/hold hit -- four of four
FAIL-0016-R3, run 3 of 5 on a corporate-debt line settling at NGX with a 7 settlement day extension period. The carried state gives 6 days counted and no event open. The reply excluded 2026-03-09 under Rule S-3 and applied Rule S-5's one-day projection skip, landing on 6 of 7 with the deadline at 2026-03-11 -- the key exactly. The strongest free floor, on the same carried state and the same page, read the line as the delivering party's and answered BUY-IN DUE at 7 days with the deadline pulled to 2026-03-09: status miss, age miss, deadline miss by three days, instruct/hold miss. And on the schedule where run 3 is never made, the model itself makes the identical mistake one day later on FAIL-0016-R4 -- BUY-IN DUE, 8 days, deadline 2026-03-09, instruct now, against a truth of SUSPENDED, 6 days, 2026-03-12, hold -- because the shutdown opened on the day nobody read and the activity feed keeps one day.
Context-incomplete recall -- the guardrail, scored
not applicable to this reading -- FAIL-0016-R3 carries both an extension period and an ISD
This grader is a rate over the 10 readings whose agreement carries no extension period or whose feed captured no intended settlement date, and the example row is not one of them. It is named here rather than left blank because a grader with no verdict on the worked example reads as a grader nobody ran: it WAS run, on every reading, and scored 100.0 pct. The rows it judges are the two fails the generator gives a missing clause and a missing ISD, and the b000 floor gets all of them wrong by guessing seven days.
What a missed run costs, in crossings rather than in points
this reading exists on both schedules and its own crossing was not lost -- but the fail is wrong one day later on the ablation
FAIL-0016 crosses nothing on run 3; what happens to it is the ARITHMETIC half of the damage, and this grader counts the other half. On the missed-run schedule FAIL-0016-R4 sees a two-day calendar window, a carried state that correctly says nothing was open at the end of run 2, and an activity log with ZERO entries. Both the model and the free floor answer BUY-IN DUE at 8 days with the deadline at 2026-03-09; the truth is SUSPENDED at 6 with 2026-03-12. The crossings this grader actually counts are elsewhere: 5 fell on the unread day, 3 were instructed late and 2 are gone permanently, worth $8,315,720.00.
The formulaWhat it computes
accuracy = hits / 112 per field. A reading whose reply did not parse counts as a MISS in every field rather than being dropped -- a run that lost readings must not score better for having lost them.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% status accuracy · 3 more measured on this row
the same tier, memory removed (THE CONTROL)
91.1% status accuracy · 3 more measured on this row
the same tier, one business day never read (THE ABLATION)
95.1% status accuracy · 3 more measured on this row
free floor: b000 -- caldays
59.8% status accuracy · 3 more measured on this row
free floor: b001 -- caldays-mem
59.8% status accuracy · 3 more measured on this row
free floor: b002 -- bizdays-mem
90.2% status accuracy · 3 more measured on this row
free floor: b003 -- marketcal-mem
97.3% status accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl and data/gold-missed-r3.jsonl, computed by tools/build_corpus.py from the fail model and the market calendars at generation time and re-derived through evals/run.advance by evals/check_labels.py before any run may spend.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one thing about it is known to be arguable -- Rule S-5's convention that an open suspension is projected to cost exactly ONE further settlement day. That is invented. Every deadline on a suspended fail inherits it.
Watch these
status_accuracy_pct
age_accuracy_pct
buyin_deadline_accuracy_pct
raise_buyin_accuracy_pct
missed_buyins
premature_buyins
unparsed_replies
Alarm on
unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. The stateless control is exactly that case: 6 of its 112 readings came back unusable.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides, and the scorer compares strings and integers. The one threshold in the kit is Rule S-9's DUE_NEXT_RUN band, and it is not tuned: it IS the cadence.
Cadence: Every run, automatically -- evals/run.py scores in the same process that made the calls, so no arm can exist without its grade. evals/check_labels.py runs BEFORE any run may spend and re-derives the whole key on both schedules; re-run it free on any change to the generator, the rule, the calendars, the splitter or the privacy gate.
The decisionWhen to reach for it
Use it
The truth is known and two of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real book, where whether a line was shut to everybody or the counterparty was simply short is exactly what the desk is arguing about with the counterparty.
Context-incomplete recall -- the guardrail, scored
Catch failing trades before the buy-in deadline
PresenterOpens the private repo. Visible to admins only.
In one lineContext-incomplete recall -- the guardrail, scored
Of the 10 readings whose agreement carries no extension period or whose feed captured no intended settlement date, how many were reported CONTEXT_INCOMPLETE rather than aged against a guessed number.
$0.00per 1,000 fail snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the cell match.
Every grader on these pages scored the same 112 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
FAIL-0016-R3 -- scheduled run 3 of 5, 2026-03-09, Corporate debt at NGX, extension period 7 settlement days, residual 5,000 at 42.34 USD (211,700.00 USD of exposure)
What the previous scheduled run left behind
As at the previous scheduled run, 6 settlement day(s) had already been counted against this fail at its place of settlement, suspensive days already excluded. No suspensive event was open when that run ended. The extension period has not yet been found exhausted, so no deadline has been fixed and no buy-in has been instructed. It was reported DUE_NEXT_RUN.
The exceptions feed, this snapshot day only
07:41 neither we nor the counterparty can put this through: the line has been shut to all instructions since the notice went out this morning
The answer key
SUSPENDED, aged 6 settlement days, buy-in deadline 2026-03-11, instruct NO
What run r001 answered
SUSPENDED, aged 6 settlement days, buy-in deadline 2026-03-11, instruct NO
It is the reading where a wrong attribution does not merely move a number -- it flips the action. The one exceptions line in this window says the line was shut to ALL instructions, and says it in counterparty language: "neither we nor the counterparty can put this through". Read as the counterparty's fault, the day ages the fail, the count reaches the 7-day extension period, Rule S-6 freezes the deadline three days early at 2026-03-09 and Rule S-7 says instruct the buy-in TODAY. Read correctly it is a suspensive day, the count stays at 6 of 7, the fail is SUSPENDED and the deadline is still ahead at 2026-03-11. One sentence, $211,700.00 of residual exposure, and the difference between holding and forcing a purchase in the market.
Grader
Verdict
Why
The status, the counted age, the buy-in deadline date and the instruct-or-hold call, per reading, exact match against the computed key
status hit, age hit, deadline hit, instruct/hold hit -- four of four
FAIL-0016-R3, run 3 of 5 on a corporate-debt line settling at NGX with a 7 settlement day extension period. The carried state gives 6 days counted and no event open. The reply excluded 2026-03-09 under Rule S-3 and applied Rule S-5's one-day projection skip, landing on 6 of 7 with the deadline at 2026-03-11 -- the key exactly. The strongest free floor, on the same carried state and the same page, read the line as the delivering party's and answered BUY-IN DUE at 7 days with the deadline pulled to 2026-03-09: status miss, age miss, deadline miss by three days, instruct/hold miss. And on the schedule where run 3 is never made, the model itself makes the identical mistake one day later on FAIL-0016-R4 -- BUY-IN DUE, 8 days, deadline 2026-03-09, instruct now, against a truth of SUSPENDED, 6 days, 2026-03-12, hold -- because the shutdown opened on the day nobody read and the activity feed keeps one day.
Context-incomplete recall -- the guardrail, scored
not applicable to this reading -- FAIL-0016-R3 carries both an extension period and an ISD
This grader is a rate over the 10 readings whose agreement carries no extension period or whose feed captured no intended settlement date, and the example row is not one of them. It is named here rather than left blank because a grader with no verdict on the worked example reads as a grader nobody ran: it WAS run, on every reading, and scored 100.0 pct. The rows it judges are the two fails the generator gives a missing clause and a missing ISD, and the b000 floor gets all of them wrong by guessing seven days.
What a missed run costs, in crossings rather than in points
this reading exists on both schedules and its own crossing was not lost -- but the fail is wrong one day later on the ablation
FAIL-0016 crosses nothing on run 3; what happens to it is the ARITHMETIC half of the damage, and this grader counts the other half. On the missed-run schedule FAIL-0016-R4 sees a two-day calendar window, a carried state that correctly says nothing was open at the end of run 2, and an activity log with ZERO entries. Both the model and the free floor answer BUY-IN DUE at 8 days with the deadline at 2026-03-09; the truth is SUSPENDED at 6 with 2026-03-12. The crossings this grader actually counts are elsewhere: 5 fell on the unread day, 3 were instructed late and 2 are gone permanently, worth $8,315,720.00.
The formulaWhat it computes
recall = correctly abstaining readings / 10
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% context incomplete recall
free floor: b000 -- caldays
0.0% context incomplete recall
free floor: b003 -- marketcal-mem
100.0% context incomplete recall
In operationWhat to monitor
Reference standard: the 10 readings whose gold status is CONTEXT_INCOMPLETE, from data/gold.jsonl.
These rates are UNKNOWN, on purpose
Whether the kit abstains correctly on anything OTHER than these two shapes -- a missing extension period and a missing intended settlement date. It has not been given a third.
Watch these
context_incomplete_recall_pct
Alarm on
anything below 100. A default extension period is the one thing Rule S-2 forbids.
How tight can the band be? No threshold. The rule is categorical: an absent value is an abstention.
Cadence: Every run, in the same pass as the cell match. It is a rate over the 10 readings whose gold status is CONTEXT_INCOMPLETE, so it is only meaningful while the corpus keeps carrying them -- and evals/check_labels.py does not assert that it does, which is stated here rather than assumed.
The decisionWhen to reach for it
Use it
There is a stated rule forbidding a default and a population that triggers it.
Do not use it
The population contains no such rows -- then this is a rate over zero and the guardrail is untested, not passed.
What a missed run costs, in crossings rather than in points
Catch failing trades before the buy-in deadline
PresenterOpens the private repo. Visible to admins only.
In one lineWhat a missed run costs, in crossings rather than in points
How many mandatory buy-ins fell due on the day nobody looked, how many of those were merely late, and how many are gone permanently because the fail settled in full before the next run.
$0.00per 1,000 fail snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
tools/build_corpus.py writes the counts; evals/cadence.py prints them.
Every grader on these pages scored the same 112 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
FAIL-0016-R3 -- scheduled run 3 of 5, 2026-03-09, Corporate debt at NGX, extension period 7 settlement days, residual 5,000 at 42.34 USD (211,700.00 USD of exposure)
What the previous scheduled run left behind
As at the previous scheduled run, 6 settlement day(s) had already been counted against this fail at its place of settlement, suspensive days already excluded. No suspensive event was open when that run ended. The extension period has not yet been found exhausted, so no deadline has been fixed and no buy-in has been instructed. It was reported DUE_NEXT_RUN.
The exceptions feed, this snapshot day only
07:41 neither we nor the counterparty can put this through: the line has been shut to all instructions since the notice went out this morning
The answer key
SUSPENDED, aged 6 settlement days, buy-in deadline 2026-03-11, instruct NO
What run r001 answered
SUSPENDED, aged 6 settlement days, buy-in deadline 2026-03-11, instruct NO
It is the reading where a wrong attribution does not merely move a number -- it flips the action. The one exceptions line in this window says the line was shut to ALL instructions, and says it in counterparty language: "neither we nor the counterparty can put this through". Read as the counterparty's fault, the day ages the fail, the count reaches the 7-day extension period, Rule S-6 freezes the deadline three days early at 2026-03-09 and Rule S-7 says instruct the buy-in TODAY. Read correctly it is a suspensive day, the count stays at 6 of 7, the fail is SUSPENDED and the deadline is still ahead at 2026-03-11. One sentence, $211,700.00 of residual exposure, and the difference between holding and forcing a purchase in the market.
Grader
Verdict
Why
The status, the counted age, the buy-in deadline date and the instruct-or-hold call, per reading, exact match against the computed key
status hit, age hit, deadline hit, instruct/hold hit -- four of four
FAIL-0016-R3, run 3 of 5 on a corporate-debt line settling at NGX with a 7 settlement day extension period. The carried state gives 6 days counted and no event open. The reply excluded 2026-03-09 under Rule S-3 and applied Rule S-5's one-day projection skip, landing on 6 of 7 with the deadline at 2026-03-11 -- the key exactly. The strongest free floor, on the same carried state and the same page, read the line as the delivering party's and answered BUY-IN DUE at 7 days with the deadline pulled to 2026-03-09: status miss, age miss, deadline miss by three days, instruct/hold miss. And on the schedule where run 3 is never made, the model itself makes the identical mistake one day later on FAIL-0016-R4 -- BUY-IN DUE, 8 days, deadline 2026-03-09, instruct now, against a truth of SUSPENDED, 6 days, 2026-03-12, hold -- because the shutdown opened on the day nobody read and the activity feed keeps one day.
Context-incomplete recall -- the guardrail, scored
not applicable to this reading -- FAIL-0016-R3 carries both an extension period and an ISD
This grader is a rate over the 10 readings whose agreement carries no extension period or whose feed captured no intended settlement date, and the example row is not one of them. It is named here rather than left blank because a grader with no verdict on the worked example reads as a grader nobody ran: it WAS run, on every reading, and scored 100.0 pct. The rows it judges are the two fails the generator gives a missing clause and a missing ISD, and the b000 floor gets all of them wrong by guessing seven days.
What a missed run costs, in crossings rather than in points
this reading exists on both schedules and its own crossing was not lost -- but the fail is wrong one day later on the ablation
FAIL-0016 crosses nothing on run 3; what happens to it is the ARITHMETIC half of the damage, and this grader counts the other half. On the missed-run schedule FAIL-0016-R4 sees a two-day calendar window, a carried state that correctly says nothing was open at the end of run 2, and an activity log with ZERO entries. Both the model and the free floor answer BUY-IN DUE at 8 days with the deadline at 2026-03-09; the truth is SUSPENDED at 6 with 2026-03-12. The crossings this grader actually counts are elsewhere: 5 fell on the unread day, 3 were instructed late and 2 are gone permanently, worth $8,315,720.00.
The formulaWhat it computes
Read off the two answer keys: crossings on the skipped run, minus those whose fail still appears on a later reading of the ablation schedule. No model, no call.
The analysisWhat it actually did
Model
Result
the two answer keys, no model
no headline metric on this row — it records crossings lost permanently 2 · crossings seen late 3
free floor, intact against missed-run
no headline metric on this row — it records crossings lost permanently 2 · crossings seen late 3
the model, intact against missed-run
no headline metric on this row — it records crossings lost permanently 2 · crossings seen late 3
In operationWhat to monitor
Reference standard: data/gold.jsonl against data/gold-missed-r3.jsonl -- the same truth, with the run that must instruct each buy-in moved.
These rates are UNKNOWN, on purpose
What skipping a DIFFERENT business day would cost. Run 3 was skipped; runs 2 and 4 would hit different fails and a different number of crossings. This is one draw.
Watch these
crossings_lost_permanently
crossings_seen_late
moved_right_to_wrong
runs_missed_detected
Alarm on
crossings_lost_permanently above 0 -- because that number can never be recovered and never shows up in any accuracy figure.
How tight can the band be? No threshold. A crossing is lost when the fail has no later reading on the ablation schedule, which is a set membership test.
Cadence: On demand and free: python3 -m evals.cadence. It reads the committed result files, so it is re-runnable forever without a key. Re-run it whenever the schedule, the corpus or the ablation changes -- and note that the crossings half needs no result file at all: it comes from the two answer keys, which tools/build_corpus.py regenerates from the seed.
The decisionWhen to reach for it
Use it
A monitor has a cadence and a deadline that does not move.
Do not use it
The population is stable and nothing leaves it -- then a late reading is late and nothing is lost, and an accuracy delta would be the whole story.
A living map of modern AI — kept current every morning