Catch payment risk to your subcontractor's suppliers
Paying your subcontractor in full doesn't clear what they owe a supplier below them. Only a waiver does, and this app checks each pay cycle, remembers what happened last cycle, and flags the payment risk before you sign off.
PresenterOpens the private repo. Visible to admins only.
For the subcontract administratorConstruction & Engineering · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A subcontract administrator or project accountant on a construction job, reviewing pay applications before they're signed.
✕Today's manual process
1Read this cycle's statement for one supplier balance under a subcontractor.
2Dig through old pay applications to see if this balance was already open last cycle.
3Decide by memory whether a joint check or hold should already be in place.
4Miss the pattern once and a subcontractor gets paid in full while its supplier stays unpaid.
Every cycle checked against memory
✓With the app
1Every cycle is read and the balance and waiver status show up already filled in.
2Last cycle is remembered so nothing has to be dug up from old pay applications.
3The status is reported watch, exposed or clear, worked out from the history, not memory.
4The pattern is caught so a joint check goes on before a payment goes out wrong.
Every cycle checked against its own history
See it work
One real case, read by the app, step by step
Obligation STO-0022 clears in cycle three, but the app holds it open one more cycle because it was already flagged as exposed the cycle before.
Catch payment risk to your subcontractor's suppliersReference appBuilt to be shaped to your process
Catch payment risk to your subcontractor's suppliers
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A general contractor's exposure to a second-tier party — a supplier or sub-subcontractor working under a first-tier subcontractor — is not discharged by paying the first-tier subcontractor. What discharges it is a waiver, and a waiver that has not arrived is an administrative annoyance the first time and a payment problem the second time. The difference between those two is not in the pay application you are holding. It is in whether the same obligation was open last cycle, which nobody can see without going back through prior packs — so on a project with a few hundred open obligations, nobody does, and the control that exists on paper is a control nobody runs. A subcontract administrator reading this cycle's exposure statement for one second-tier party and then going back through the prior pay applications to work out whether the escalation rules have been tripped — for every open obligation, every pay cycle, before the pay meeting.
Audience
Subcontract administrators, project accountants and the risk functions above them: anyone who owns a joint-check policy with consecutive-cycle rules in it and currently applies it by memory. It is also for the person deciding whether to BUILD this, and for them the most useful number on the page is not the headline — it is that the free ageing rule they probably already run scores WORSE than doing nothing. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual pay-cycle statements
The corpus is 150 pay-cycle statements, 0.45 MB (jsonl 1 · txt 150). A real second-tier exposure pack names a real general contractor's real subcontractors, their unpaid balances, their waiver history and the project accountant who owns each one. None of that can be committed to a public repository, and a redacted version of it would still be somebody's payment position. So the corpus is generated — and generating it bought three things a real pack could not have given:
· THE GOLD IS COMPUTED. src/exposure.replay produces every label. A hand-marked set on a three-branch rule with a resetting counter would carry its author's misreading into the score.
· THE PATTERN MIX IS CHOSEN. 11 of the 50 obligations are the counter-reset trap and 2 are its mirror (two open cycles where the waiver DID arrive, so no joint check is owed). A real sample would have been mostly quiet cycles and would have said nothing about either.
· HISTORY LEAKS NOWHERE ELSE. evals/check_labels.py check 9 asserts that no section describing the obligation mentions another cycle, so the carried state is provably the only route from one cycle to the next. That is the experimental control on the entire kit and it cannot be established on found data.
The corpus
The 150 pay-cycle statementsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your pay-cycle statements. That is the whole change — there is no database to migrate.
One pay-cycle statement, as the model receives itSTO-0001-C1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated by tools/build_corpus.py from a fixed
seed for the AI Foundry use-case kit "subtier-exposure". No real project,
subcontractor, supplier, payment or waiver is described here, and the exposure
policy reproduced below is invented rather than taken from anybody's procedure.
Obligation
----------------------------------------------------------------
Obligation ID : STO-0001
Project : Northgate Distribution Centre Phase II
Prime contract : GC-2024-0655 (general contractor, lump sum)
First-tier sub : Ardent Fire Protection (subcontract SC-14, fire sprinkler and standpipe)
Second-tier : Thornbury Insulation Co (sub-subcontractor, tier 2)
Scope : mechanical insulation
Subcontract val : $ 842,000.00
Pay cycle : 1 of 3 (pay period 2025-05)
Exposure Policy
----------------------------------------------------------------
Exposure is the second-tier unpaid balance expressed as a percentage
of the amount certified to the first-tier subcontractor this cycle.
Measure : unpaid_exposure
Watch at : 19.5
Exposed at : 35.5
Direction : higher_is_worse
Escalation rules, applied in this order:
1. RELEASE HOLD -- an obligation reported EXPOSED last cycle is reported WATCH
this cycle even when the reading alone is CLEAR. One clean cycle does not
close the file; the unconditional waiver for the prior payment must land
first.
2. BOND CLAIM NOTICE -- a second consecutive cycle reported EXPOSED is escalated
to a bond-claim or stop-notice preservation letter.
3. JOINT CHECK -- a second consecutive cycle at WATCH or worse, where the
Abridged — the file continues.
The outcomeWhat a good result looks like
Per pay cycle: the posture to REPORT after the escalation rules have been applied, the name of the rule that was applied, and the component of the balance driving the exposure — with the carried state that produced the verdict printed beside it, so the answer can be explained to the subcontractor it is about.
And when it cannot
It drops an escalation that was already running. All 4 wrong cells in 150 are third cycles and 3 of the 4 are the same shape: a joint check live since cycle 2, and cycle 3 answered as though the obligation were new. It never invented one — 0 false alarms across the 100 quiet cycles.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
You have a joint-check or withholding policy with CONSECUTIVE-CYCLE rules in it — this kit That is exactly the dependence four scalar fields capture, and the 30.66-point gap between the scored run and its stateless control is a measurement of it.
You want to know whether a model is needed at all — this kit The page answers it in the unflattering direction on three of the four graders. Posture, the counter-reset subset and the top driver are all matched exactly by a free regex; only the action grader separates.
You want to know what happens if the monitor does not run one month — this kit a000-subtier-exposure-missedcycle measures it for free: 33 of 100 cycles move, 22 owed escalations never fire, all 11 counter-reset obligations go wrong.
And where nothing here is good enough:
You want to run it over a real second-tier exposure pack — not this kit, not yet Every regex in src/monitor.py is written to this corpus's exact line shapes and no real accounting export will match one.
You want a stability guarantee before acting on a raise — not this kit The repeat probe was written (evals/repeat.py) and NOT run — the shared provider balance was exhausted before it could fire. Nothing here says whether the same cycle answers the same way twice.
At a glanceHow the whole thing runs
97%escalation accuracy pct
6,835 msp50, end to end
$5.99per 1,000 pay-cycle statements · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch payment risk to your subcontractor's suppliers14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/exposure.py FIRST, then src/cadence.py. Every figure on this page is a property of a three-cycle corpus with unambiguous drivers, clean thresholds and exactly one second-tier party per subcontract.Corpus lens →
When is this the wrong choice?
Avoid: …unless your policy's rules depend on more than the previous cycle. Everything here is first-order: the state remembers last cycle and two counters, nothing else. That is the case against the best-fitting scenario (“You have a joint-check or withholding policy with CONSECUTIVE-CYCLE rules in it”). 5 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A statement whose Exposure Breakdown has two components within 5 points of each other. The corpus forbids it and the checker asserts it, because a top-driver grader that is a coin toss is not a grader — but real breakdowns tie all the time, and this kit has no rule for what to do when they do. 7 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER THE ANSWER IS STABLE. evals/repeat.py is written, committed and runnable — 6 obligations x 3 cycles x 2 fresh asks, chosen for the rules they exercise — and it was NOT RUN. 8 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-subtier-exposure. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes all nine pre-flight checks, runs BOTH free floors and the cadence ablation, and serves the whole UI board with four of its five arms live — including the model's own row, which is read off the committed result file rather than re-asked. Measured at 0.2 seconds from corpus build to a scored free floor. The only thing a key adds is a live row in the UI and the ability to re-run the two paid arms.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
6,835 msp50, end to end
64,333 msp95
1 minclone to first result
What the clock covers. Measured on the fast tier — END-TO-END per pay cycle: one HTTP request carrying the assembled prompt, and the reply parsed to a posture, an action code, a driver and its share. It does NOT include the threshold arithmetic or the state advance, both of which are pure code and complete in under a millisecond. The p95 is 9.4x the p50 and that gap is the story: 95 pct of this run's output tokens were provider-side reasoning, and the reasoning is long exactly on the cycles where a rule fires.
Current processWhat it replaces
A subcontract administrator reading this cycle's exposure statement for one second-tier party and then going back through the prior pay applications to work out whether the escalation rules have been tripped — for every open obligation, every pay cycle, before the pay meeting.
Where it is not good enough
⚑ THE FAILURE IS ONE PATTERN AND IT IS NAMED. 4 of 150 cycles are wrong, and ALL FOUR ARE THIRD CYCLES. Three of them are the WWW shape — three consecutive open cycles with no waiver, where a joint check was already running from cycle 2 and the correct answer for cycle 3 is to keep it running. On STO-0037, STO-0039 and STO-0040 the reader dropped it: it answered WATCH / NONE (and on STO-0040, WATCH / RELEASE_HOLD, which is a rule that does not apply at all) where the policy says EXPOSED / JOINT_CHECK. The fourth, STO-0029-C3, is a RELEASE_HOLD it did not raise. Every one of the four is a MISSED escalation; the run raised zero it should not have.
⚠︎ THAT DIRECTION IS THE COMFORTABLE ONE AND IT IS STILL NOT GOOD ENOUGH. On this use case a missed joint check is a general contractor paying a first-tier sub in full while a supplier below them stays unpaid and keeps its lien rights — which is the entire thing the control exists to prevent. 92.0 pct on the cycles where a rule actually fires means 4 of 50 real escalations were not raised. Nobody should run this without a person reading the raises AND the non-raises on any obligation that has been open more than one cycle.
⚠︎ AND THE 100 PCT FIGURES ARE NOT THE MODEL'S. Posture on the 100 quiet cycles, the counter-reset subset and the top driver are all perfect — and THE FREE ARITHMETIC FLOOR SCORES 100 PCT ON ALL THREE TOO, for $0.00. Measured cell by cell, not assumed: 100 of 100, 11 of 11, 150 of 150 on both arms. Do not read any of the three as evidence about a reader.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1
50 open second-tier obligations, 3 consecutive pay cycles each
66.67pct from one statement alone — the free floor's score, and the control's
0 false alarms in 100 quiet cycles · 0 replies lost
2026-08-23as of
It produces a posture, an action code and the component driving the balance for a subcontract administrator to read before the pay meeting, and it acts on none of them — there is no endpoint, no function and no flag anywhere in the kit that issues a joint check, withholds a payment, files a bond claim, serves a notice or contacts a subcontractor, and every one of those is a contractual act against a third party's cash flow. THE TWO STATIONS THAT MAKE THIS A MONITOR ARE THE SECOND AND THE THIRD, AND MOST STATEFUL KITS SHIP ONLY THE SECOND. The carried state is four scalar fields, written by src/exposure.py from the reading a regex parsed off the page and never from the model's answer — so a wrong cycle costs you that cycle and nothing after it. The clock is the other half: a run must HAPPEN, and a missed one is not recoverable by running twice, which is why claim() refuses a period already in the register. ⚑ WHAT A MISSED RUN COSTS WAS MEASURED RATHER THAN ASSERTED, for $0.00: the policy replayed with one monthly run skipped moves 33 of the 100 surviving verdicts, silently drops 5 joint checks, 12 release holds and 5 bond-claim notices, and gets every one of the eleven counter-reset obligations wrong against 11 of 11 when the run happens. A missed run leaves NO DOCUMENT to read and a state file that is perfectly self-consistent, so nothing but the register can see it — and no evaluation of the reader can find it at all.
⚠︎ THE FLOOR STATION IS UNFLATTERING TO THE INCUMBENT ON PURPOSE: the ageing report most contractors already run scores below answering NONE to everything, and the free arithmetic floor matches the model EXACTLY on the top driver (150 of 150), on the counter-reset subset (11 of 11) and on the posture of all 100 quiet cycles. It is only on the 50 cycles where a rule fires that the two separate at all — 26.0 pct against 92.0 — which is the whole of what this kit measures.
⚠︎ THE POLICY IS INVENTED — two thresholds, three escalation rules, four action codes and every day count, written for this kit. No lien statute is reproduced and the 60-day ageing horizon is not a legal deadline. Replace src/exposure.py and the answer key recomputes, which invalidates every figure here by design.
The swap seams
Seam
File
What changes
the policy
src/exposure.py
REPLACE THIS FIRST. It is your procedure's thresholds and escalation rules, not a utility. The gold is recomputed from whatever you put here, so changing it invalidates every published number on this page and says so.
the cadence
src/cadence.py
how often a run is owed, and what a period is. Monthly is stated, not measured.
the carried state
src/state.py
which facts survive between cycles, and how they are rendered in English. Keep it scalar — state that grows with history turns the cost curve quadratic.
the section selector
src/select.py
what leaves your machine. NEVER_SENT is the list to read before pointing this at real statements.
the provider
.env plus src/adapters/__init__.py
PROVIDER, BASE_URL, MODEL. A new vendor is one function and one dict entry.
Components
Component
File
Role
corpus builder
tools/build_corpus.py
writes the 150 exposure statements and the computed gold from one seed; free, offline, reproducible
exposure policy
src/exposure.py
the thresholds and the three escalation rules, as arithmetic. Produces the gold and the state; never calls a model
carried state
src/state.py
four scalar fields between cycles, written from the arithmetic and rendered into one English sentence for the prompt
cadence
src/cadence.py
which pay period is owed, which have been run, and the refusal that stops one period being run twice
segmenter
src/segment.py
splits a statement into its seven named sections
section selector
src/select.py
decides which sections reach the prompt; Internal Contacts is mapped by nothing and therefore never sent
prompt builder
src/prompt.py
instruction + carried state + selected sections, by string concatenation
model adapter
src/adapters/__init__.py
one completion over raw HTTP; provider chosen by .env
runtime
src/monitor.py
one cycle, one call, the reply parsed to four fields
budget guard
src/budget.py
counts calls against a shared per-day cap before spending
eval harness
evals/run.py
50 obligation chains, strictly serial within a chain, concurrent across them
scorer
evals/scoring.py
four graders, exact match, pure code
free floors
evals/baseline.py
threshold arithmetic, and the 60-day ageing rule
cadence ablation
evals/missed_cycle.py
replays every obligation with one run skipped and scores what moved
pre-flight checks
evals/check_labels.py
nine assertions that must hold before a run may spend
local UI
src/app.py
one obligation, three cycles, five arms — four of them free
Where it breaks at scale
NOT ON HISTORY LENGTH, AND THAT IS THE DESIGN. The carried state is four scalar fields, so cycle 30 costs exactly what cycle 3 costs — 1,071.77 input tokens on this corpus, flat. What it breaks on is BREADTH and the CLOCK.
⚑ BREADTH: the unit of work is one obligation per pay period, so the bill is rows x periods, forever. 400 open second-tier obligations is 400 calls every month, not 400 calls once. Any reader who prices this per document has the wrong number.
⚑ THE CLOCK: a run must happen, and a missed one is not recoverable by running twice. src/cadence.claim() refuses a second run of the same period precisely because advancing a counter is not idempotent — run 2025-05 twice and every obligation at WATCH for one cycle is reported EXPOSED and put on a joint check because the monitor ran twice. Measured in a000-subtier-exposure-missedcycle: skip ONE monthly run and 33 of the remaining 100 cycles change verdict, 22 escalations that were owed never fire at all, and all 11 counter-reset obligations get the wrong answer.
⚠︎ AND CONCURRENCY IS NOT MEASURED. This run used 14 workers across 50 independent chains; nothing here has been tried at 400 chains, and the provider's own rate limits are not known to this kit.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
WHAT THE CARRIED STATE BOUGHT, IN ONE ROW. STO-0022 runs WATCH, WATCH, clear. The policy says NONE, then JOINT_CHECK (second consecutive open cycle, waiver still missing), then RELEASE_HOLD (came back clear from EXPOSED, so the file stays open one more cycle). The recorded run matches all three. BOTH free floors miss both escalations — neither is reachable without memory — and the ageing floor also false-alarms a joint check on cycle 1. The missed-run arm loses both as well, from the identical corpus and the identical policy: that one is the schedule's fault, not the reader's, and no eval of the model could find it.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
TWO ARMS IMPOSING A JOINT CHECK NOBODY EARNED, for two different reasons, on the same obligation. STO-0011 is the counter-reset trap: WATCH, then a clear cycle, then WATCH. The clear cycle resets the consecutive count, so cycle 3 is WATCH / NONE — which the policy, the arithmetic floor and the recorded model run all get right. The 60-day ageing floor raises a JOINT_CHECK on cycle 1 on nothing but a day count. The missed-run replay raises one on cycle 3, because it never saw the clear cycle that reset the counter. Both are contractual acts against a subcontractor that the procedure does not authorise.failureOpen full size →ONE OF THE FOUR WRONG CELLS IN 150, AS IT CAME BACK. STO-0037 is three consecutive open cycles with no waiver. The recorded run got cycle 2 right — EXPOSED / JOINT_CHECK — and then dropped it on cycle 3, answering WATCH / NONE on an obligation whose joint check was already running. ⚠︎ NOTE THE BOTTOM ROW: on this obligation the missed-run replay gets cycle 3 RIGHT, by coincidence, because skipping cycle 2 left a count that happened to land on the same branch. A cadence arm that is sometimes right is exactly why its damage is invisible without the ablation.failureOpen full size →
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150pay-cycle statements
0.45 MiBjsonl 1 · txt 150
50obligations (3 pay cycles each) · p50 3 chars
$0.00setup · 0.0s
How it is cutWhat one obligations (3 pay cycles each) is
No split, and no chunking. The unit is an OBLIGATION — a run of three consecutive pay-cycle statements processed strictly in order, carrying state. A statement is 3,142 to 3,219 bytes and goes whole into one call minus the one section the selector withholds.
SetupWhat the setup figure measured
There is no index and no retrieval step. tools/build_corpus.py writes 150 statements and the gold file from a seed in well under a second, and that is the whole of the preparation. The 0.0 is a measured zero, not an unmeasured one: nothing is embedded, nothing is stored and nothing is looked up.
LicenceLicence
MIT — this repository's own licence. Nothing is fetched and no third-party grant is involved; every statement is generated in-process on a fresh clone.
Bring your ownBring your own pay-cycle statements
Replace src/exposure.py FIRST, then src/cadence.py. The first is your subcontract administration procedure — its thresholds, its escalation rules and their order — and the gold in data/gold.jsonl is recomputed from it, so every number on this page becomes a number about YOUR policy the moment you change it. The second is your billing calendar. Only then point tools/build_corpus.py at your own statements, or replace the regexes in src/monitor.py with a reader for your accounting export.
⚠︎ And what stops being true when you do: Every figure on this page is a property of a three-cycle corpus with unambiguous drivers, clean thresholds and exactly one second-tier party per subcontract. A real project has none of those. Nothing here has been run on a real exposure pack, and 97.33 pct on this corpus is not a prediction about yours — it is a statement that the escalation rules are learnable from four scalars, which is a different and smaller claim.
What breaks it
A statement whose Exposure Breakdown has two components within 5 points of each other. The corpus forbids it and the checker asserts it, because a top-driver grader that is a coin toss is not a grader — but real breakdowns tie all the time, and this kit has no rule for what to do when they do.
A reading exactly on a threshold. Every value here is at least 1.5 points clear of the band it is not in, so no published number depends on whether the rule reads >= or >. Real readings land on thresholds.
An obligation whose statement omits a section. src/select.py subtracts NEVER_SENT unconditionally rather than falling back to the whole document, so a missing Exposure Breakdown degrades the answer instead of leaking the contacts block — but it still degrades it, and nothing here measures by how much.
A pay cycle with no statement at all. The harness iterates the corpus; a deployment iterates open obligations, and an obligation that was not billed this period has no statement to read. src/cadence.py has no rule for it and neither does the policy.
More than one second-tier party under one first-tier subcontract. Every obligation here is one pair. Real subcontracts carry a dozen suppliers below them and their exposures interact — a joint check to one does not discharge another.
A first-tier subcontractor who is themselves the second tier on another package. Nothing here models a chain longer than two.
Statements not produced by this builder. Every regex in src/monitor.py is written to this corpus's exact line shapes; a real accounting export will not match one of them.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
134
not measured
instruction
1,220
not measured
carried state
211
not measured
Synthetic Record
311
not measured
Obligation
421
not measured
Exposure Policy
1,055
not measured
Cycle Position
206
not measured
Exposure Breakdown
339
not measured
Field Notes
137
not measured
Total
1,071.77
This is the cost lesson as arithmetic: of the 4,034 characters assembled, 2,469 are documents — 61% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.py::build on STO-0022-C3 with the exact carried state r001 fed it — {prev_reported: EXPOSED, prev_raw: WATCH, consec_open: 2, consec_exposed: 0} — which renders as the 209-character sentence above. Total assembled prompt: 4,567 characters.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You apply a written second-tier exposure policy to one pay cycle of one obligation. You answer with one JSON object and no other text.
You are reviewing one second-tier payment obligation for one pay cycle against a general
contractor's exposure policy.
The policy, its thresholds and its escalation rules are reproduced in the statement below. Apply
them exactly as written. The escalation rules depend on what was reported in EARLIER pay cycles,
which you cannot see; what is known about this obligation's history is stated under "Carried state"
and is the only history available to you. Do not assume anything about earlier cycles beyond it.
Answer with a single JSON object and nothing else:
{"posture": "CLEAR|WATCH|EXPOSED",
"action": "NONE|JOINT_CHECK|BOND_CLAIM_NOTICE|RELEASE_HOLD",
"top_driver": "the component contributing most to this cycle's balance, named exactly as written",
"top_driver_share_pct": <number>,
"rationale": "one sentence, naming the rule or the threshold you applied"}
"posture" is the posture to REPORT after the escalation rules have been applied, which is not always
the posture the reading alone would imply. "action" names which policy rule you applied, or NONE.
Watch the direction of the measure. On some obligations a HIGHER reading is worse and on others a
LOWER reading is worse; the statement says which.
Carried state
----------------------------------------------------------------
Last cycle this obligation was reported EXPOSED. Its reading alone last cycle was WATCH, before the escalation rules were applied. It has now been at WATCH or worse for 2 consecutive cycles, counting last cycle.
Pay-cycle exposure statement
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated by tools/build_corpus.py from a fixed
seed for the AI Foundry use-case kit "subtier-exposure". No real project,
subcontractor, supplier, payment or waiver is described here, and the exposure
policy reproduced below is invented rather than taken from anybody's procedure.
Obligation
----------------------------------------------------------------
Obligation ID : STO-0022
Project : Fairmount Regional Water Reclamation
Prime contract : GC-2023-1188 (general contractor, lump sum)
First-tier sub : Ardent Fire Protection (subcontract SC-28, fire sprinkler and standpipe)
Second-tier : Steadfast Crane Services (sub-subcontractor, tier 2)
Scope : crane and rigging
Subcontract val : $ 1,240,000.00
Pay cycle : 3 of 3 (pay period 2025-10)
Exposure Policy
----------------------------------------------------------------
Coverage is the share of cumulative payments through this tier that
is covered by executed second-tier waivers. A LOWER number is worse.
Measure : waiver_coverage
Watch at : 85.6
Exposed at : 65.1
Direction : lower_is_worse
Escalation rules, applied in this order:
1. RELEASE HOLD -- an obligation reported EXPOSED last cycle is reported WATCH
this cycle even when the reading alone is CLEAR. One clean cycle does not
close the file; the unconditional waiver for the prior payment must land
first.
2. BOND CLAIM NOTICE -- a second consecutive cycle reported EXPOSED is escalated
to a bond-claim or stop-notice preservation letter.
3. JOINT CHECK -- a second consecutive cycle at WATCH or worse, where the
conditional waiver for the prior payment has still not been received, is
reported EXPOSED and the next payment is issued on a joint check made out to
both tiers.
The consecutive count is a count of CONSECUTIVE cycles at WATCH or worse. A cycle
whose reading alone is CLEAR resets it to zero.
Cycle Position
----------------------------------------------------------------
Cumulative payments to first tier : $ 447,863.77
Covered by executed waivers : $ 428,919.13
Reading : 95.77
Waiver received : yes
Days since last furnishing : 31
Preliminary notice served : yes
Exposure Breakdown
----------------------------------------------------------------
Components of the uncovered amount this cycle:
Payments with no waiver on file 60.5 % of the balance
Waivers rejected for wrong through-date 18.2 % of the balance
Progress payments still in transit 13.8 % of the balance
Conditional waivers never made unconditional 7.5 % of the balance
Field Notes
----------------------------------------------------------------
Project accountant note: the owner's payment for this period landed four working days late, which moved every downstream release with it.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"posture":"CLEAR","action":"NONE","top_driver":"Retainage held by first tier","top_driver_share_pct":46.0,"rationale":"The reading of 4.21 is below the 19.5 watch threshold, and no prior-cycle escalation rule applies."}
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch payment risk to your subcontractor's suppliers — 150 pay-cycle statements. One model answered, and every answer was then graded Four different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
There is no LLM-as-judge in this kit and that is not a shortcut. Every field it answers is a closed set — three postures, four action codes, a component name copied from a table on the page. A judge is for answers whose correctness is a matter of reading; these are matters of equality, and asking a model to grade WATCH == WATCH would add cost, variance and a second thing to be wrong.
150pay-cycle statements
150source documents
1model tier
4grading methods
MeasurementsWhat was measured
COUNTED146 · 100 · 100 · 86 · 67 / 150escalation accuracy pct — pay cycles, with the carried state -- WHICH POLICY RULE APPLIED, the reference grader and the only one of the six that cannot be answered without memoryDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED46 · 0 / 50escalation when a rule fires pct — cycles where a rule actually fires -- a joint check, a bond-claim notice or a release hold -- with the carried state. 46 of 50; all 4 misses are third cyclesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED146 · 113 · 113 / 150band accuracy pct — pay cycles, with the carried state -- the POSTURE reported after the escalation rulesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED11 · 0 / 11counter reset accuracy pct — third cycles of the WATCH-CLEAR-WATCH obligations, with the carried state. Read it as a trap not fallen into: the FREE arithmetic floor also scores 11 of 11 hereDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 25 · 11 / 100false alarm rate pct — quiet cycles, with the carried state -- 0 joint checks imposed that the policy does not call for. The free 60-day ageing floor's rate on the same 100 cycles is 25.0Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 · 150 / 150top driver accuracy pct — pay cycles, with the carried state -- AND A SORT OF THE TABLE SCORES THE SAME 150. This field needs no model at allDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 150unparsed replies — calls -- every reply parsed; largest 16828 output tokens under a 20,000 cap, which the 15-call calibration had put at 7,681Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The comparison cannot be wrong about itself; the risk is in the labels. Those come from src/exposure.replay, which is 30 lines of arithmetic with no model in it, and evals/check_labels.py re-derives every gold row from the corpus's own parsed numbers and refuses to let a run start if any disagree. It also re-runs the corpus builder into a temporary directory and diffs, so a silent corpus drift is a refusal rather than a different score.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate — a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One pay-cycle statement
1,000 pay-cycle statements
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.005989
$5.99
9%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.002395
$2.40
9%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.101599
$101.60
11%
Same work, 42× the bill
The same pay-cycle statements, the same tokens — only the rate card changed. And across all 3 cards between 9% and 11% of what you pay is the prompt this pipeline sends, not the answer it writes.
Provider-side reasoning — 95.3 pct of this run's output tokens, and output is the majority of the projected bill. The adapter has a documented thinking field and this kit's harness never sends it, so every published figure is at the provider's default. Turning it off is one argument and an unmeasured accuracy risk.
Rates checked 2026-08-18. The provider that actually ran all 315 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. A measured $0.00, not an unpriced one — and the two free floors and the cadence ablation are free too, so four of the six arms on this page cost nothing to reproduce.
The gradersFour ways to grade
⚑ THIS KIT PUBLISHES TWO FREE FLOORS AND THE SECOND ONE IS UNFLATTERING TO THE INCUMBENT.
b001 — arithmetic only posture 75.33 · action 66.67 · fires 0.0 · false alarms 0.0
b000 — the 60-day ageing rule posture 72.67 · action 57.33 · fires 22.0 · false alarms 25.0
r001 — the scored run posture 97.33 · action 97.33 · fires 92.0 · false alarms 0.0
s001 — no carried state posture 75.33 · action 66.67 · fires 0.0 · false alarms 0.0
⚠︎ THE AGEING RULE — the control a subcontract administration group most likely already runs, 'nothing furnished for 60 days and no waiver on file, put it on a joint check' — SCORES NINE POINTS BELOW DOING NOTHING on action, and three points below on posture. It reaches only 22 pct of the cycles where a rule genuinely fires, and it pays for that with a 25 pct false-alarm rate on the quiet ones: 25 joint checks imposed on subcontractors who did not earn one. It is not a straw man and it is not tuned against the kit — evals/baseline.py --sweep prints the whole curve from 20 to 120 days, and 60 is NOT its best point (120 days scores 66.67, which is the ageing rule switching itself off). The published number is deliberately not the peak.
⚑ AND THE STATELESS ARM TIES THE FREE FLOOR EXACTLY — NOT ON THE SCORES, ON THE ANSWERS. s001 and b001 agree to the decimal on all six graders (75.33, 66.67, 0.0, 100.0, 0.0, 100.0), and the stronger statement is the measured one: they give the IDENTICAL ANSWER ON ALL 450 SCORED CELLS — 0 differences, checked cell by cell against the two committed result files rather than inferred from the totals. With no memory the model is worth precisely nothing over a regex and two comparisons: it does not lose, and it does not win. That is the honest reading of what this kit measures. The 30.66-point gap between r001 and s001 is not 'what a model is worth here' — it is what FOUR SCALAR FIELDS are worth here, and they are written by 30 lines of arithmetic that anyone could have written.
the fast tier, with the carried state 97.3% band accuracy · the fast tier, memory removed (THE CONTROL) 75.3% band accuracy · the free arithmetic floor, no model 75.3% band accuracy · the free 60-day ageing rule, no model 72.7% band accuracy · the same policy, one scheduled run skipped 72.0% band accuracy · 2 more measured on each run
Action accuracy restricted to the cycles where a rule actually fires Of the cycles whose correct action is not NONE, how many were named correctly? This is the number the kit is actually about; the 150-cycle average is two-thirds NONE and a monitor that says nothing scores 66.67 pct on it.
$0.00
no
yes
the fast tier, with the carried state 92.0% accuracy · the fast tier, memory removed (THE CONTROL) 0.0% accuracy · the free arithmetic floor, no model 0.0% accuracy · the free 60-day ageing rule, no model 22.0% accuracy · the same policy, one scheduled run skipped 31.2% accuracy
False alarms on the cycles where no rule fires Of the cycles whose correct action is NONE, how many got an escalation anyway? On this use case that is not a nuisance page -- it is a joint check imposed on a subcontractor's cash flow, and it is the direction the free ageing rule fails in.
$0.00
no
yes
the fast tier, with the carried state 0.0% false alarm rate · the fast tier, memory removed (THE CONTROL) 0.0% false alarm rate · the free arithmetic floor, no model 0.0% false alarm rate · the free 60-day ageing rule, no model 25.0% false alarm rate · the same policy, one scheduled run skipped 16.2% false alarm rate
Action accuracy on the eleven counter-reset obligations On the third cycle of the eleven WATCH-CLEAR-WATCH obligations, was the action named correctly? WATCH-CLEAR-WATCH looks like two open cycles to anything counting open cycles rather than counting CONSECUTIVE ones, so it is the shape where a plausible-looking monitor puts a recovered obligation on a joint check.
$0.00
no
yes
the fast tier, with the carried state 100.0% accuracy · the fast tier, memory removed (THE CONTROL) 100.0% accuracy · the free arithmetic floor, no model 100.0% accuracy · the free 60-day ageing rule, no model 72.7% accuracy · the same policy, one scheduled run skipped 0.0% accuracy
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Decisively on one grader, not at all on three. The stateful and stateless arms separate by 30.66 points on action and 22.0 on posture, and by 92.0 points on the cycles where a rule fires — from 0.0 to 92.0, which is the whole of that grader's range. On top_driver they are identical at 100.0, on the counter-reset subset identical at 100.0, and on false alarms identical at 0.0. A page that averaged the six would report a comfortable improvement and hide that three of them measure nothing about memory at all.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
You have a joint-check or withholding policy with CONSECUTIVE-CYCLE rules in it
this kit
That is exactly the dependence four scalar fields capture, and the 30.66-point gap between the scored run and its stateless control is a measurement of it.
…unless your policy's rules depend on more than the previous cycle. Everything here is first-order: the state remembers last cycle and two counters, nothing else.
You want to know whether a model is needed at all
this kit
The page answers it in the unflattering direction on three of the four graders. Posture, the counter-reset subset and the top driver are all matched exactly by a free regex; only the action grader separates.
…do not read the headline 97.33 as the model's contribution. Its contribution over the free floor is on ONE grader.
You want to know what happens if the monitor does not run one month
this kit
a000-subtier-exposure-missedcycle measures it for free: 33 of 100 cycles move, 22 owed escalations never fire, all 11 counter-reset obligations go wrong.
…the ablation skips exactly one cycle in a three-cycle history. Nothing here measures two consecutive missed runs, or a run missed at a different point in a longer chain.
You want to run it over a real second-tier exposure pack
not this kit, not yet
Every regex in src/monitor.py is written to this corpus's exact line shapes and no real accounting export will match one.
…and 97.33 pct is not a prediction about your pack. It is a statement about a corpus with unambiguous drivers, clean thresholds and one supplier per subcontract.
You want a stability guarantee before acting on a raise
not this kit
The repeat probe was written (evals/repeat.py) and NOT run — the shared provider balance was exhausted before it could fire. Nothing here says whether the same cycle answers the same way twice.
…and on this use case that matters more than usual: a joint check imposed one month and not the next, on identical facts, is a dispute rather than a control.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
standing-joint-check-dropped
A joint check already running, answered as though the obligation were new
2
STO-0037-C3 and STO-0039-C3, both WWW obligations on their third consecutive open cycle with the waiver still missing. STO-0039-C3 says it outright: "Reading 28.30 is WATCH (between 20.2 and 38.3), and no escalation rule applies because the second-consecutive…
waiver-flag-used-to-close-a-hold
A release hold discharged by a waiver the rule never mentions
1
STO-0029-C3, an EEC obligation reported EXPOSED for two cycles and now reading 3.30 against a Watch threshold of 16.0. The run answered CLEAR / NONE: "Reading alone is 3.30, below the 16.0 Watch threshold, and the prior unconditional waiver has been received…
wrong-rule-named
The right instinct, the wrong rule, from the same misreading as the first class
1
STO-0040-C3, the one waiver_coverage obligation among the four and the only cell in 150 where a rule was named that does not apply: "...having been reported EXPOSED last cycle, RELEASE HOLD reports WATCH; the JOINT CHECK rule for a second consecutive…
What we could NOT verify
WHETHER THE ANSWER IS STABLE. evals/repeat.py is written, committed and runnable — 6 obligations x 3 cycles x 2 fresh asks, chosen for the rules they exercise — and it was NOT RUN. It failed at the first call with 402 Insufficient Balance on the shared provider key, which nine kits were drawing on the same day. No figure on this page rests on it and none is claimed.
WHETHER c001's SIX CALIBRATION CALLS ARE IN THE RUN REGISTER. They are not. results/eval-c001-subtier-exposure-calibration.json is committed and readable, but the estate's extractor keys this result shape on scores.counter_reset_accuracy_pct and that two-obligation slice contained no counter-reset obligation, so the value is null and the record refuses. The published ceiling cites c000, which does have a record.
WHETHER THE ENGLISH STATE SENTENCE BEATS A JSON ONE. src/state.describe renders the carried state as prose because the rule text is prose. Scoring the same 50 obligations with it rendered as JSON would cost one more 150-cycle run and has not been paid for.
WHETHER TWO CONSECUTIVE MISSED RUNS ARE WORSE THAN ONE, OR DIFFERENTLY WRONG. The cadence ablation skips exactly one cycle of three. A longer chain with a gap in the middle is the real operational case and this corpus is too short to hold one.
WHETHER THE 4 WRONG CELLS ARE THE MODEL OR THE PROMPT. Three of the four are the same pattern (WWW third cycle) and the obvious next experiment — an ablation stripping the sentence that warns the rules depend on unseen history — was not run, for the same balance reason as the repeat probe.
CONCURRENCY AND RATE LIMITS. 14 workers over 50 chains completed in 228.4 seconds. Nothing here has been tried at the 400-chain breadth the Architecture lens names as the real scaling axis.
ANYTHING ABOUT A REAL EXPOSURE PACK. No real statement has been through this kit.
PROVIDER-SIDE RETENTION. What the provider keeps of the 150 prompts is not knowable from here and is not claimed.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,071.77
1,817.62
6,835 ms
$0.005989
$0.002395
$0.101599
the same tier, memory removed (THE CONTROL)
1,055.41
585.23
5,313 ms
$0.002283
$0.000913
$0.039816
the free arithmetic floor (no model at all)
0
0
0 ms
$0.000000
$0.000000
$0.000000
the free 60-day ageing rule (no model at all)
0
0
0 ms
$0.000000
$0.000000
$0.000000
the same policy, one scheduled run skipped
0
0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
⚑ 339 LINES ARE ON THE SHARED CALL LEDGER FOR THIS KIT AND ONLY 315 OF THEM RETURNED ANYTHING. The ledger is written BEFORE each call so a crash over-counts rather than under-counts, and 24 of the lines are the repeat probe's attempts, every one of which came back 402 Insufficient Balance and produced no completion. Those are billed for nothing and priced at nothing — discarded_usd is null rather than 0.0 because no token count exists for a call that never happened. The 315 that did run are: 150 scored, 150 stateless control, 15 calibration. The four free arms — both floors, the stub and the cadence ablation — cost nothing and are not in this total.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING. 259736 of 272643 output tokens (95.3 pct) on the scored run were reasoning, and output is 91.1 pct of the projected bill on the shared card. It is the single biggest lever on this page and this kit never touched it — every run left thinking at the provider default.
THE RULEBOOK, SENT EVERY CALL. The Exposure Policy section is 1,055 of the 4,567 prompt characters and it is identical on all 150 cycles. Nothing here caches it; a provider with prompt caching would take that off the input bill and this kit has not measured whether it would.
THE CADENCE, WHICH IS THE ONE PEOPLE GET WRONG. One call per obligation per pay period, forever. 400 open obligations is 400 calls a month, not 400 calls.
NOT THE HISTORY. The carried state is 209 characters and does not grow. Cycle 30 costs what cycle 3 costs — which is the opposite of a conversation kit, where input cost climbs with the square of the turns.
Your volumeWhat it costs at your volume
Linear in OBLIGATION-CYCLES, and flat in history length, which is the unusual half. Each cycle is one call whose input is one statement plus 209 characters of state, so ten times the obligations is ten times the bill and ten times the pay cycles is ten times the bill — but ten times the HISTORY is free. On the shared card, 400 obligations x 12 monthly cycles is 4,800 calls a year at $0.0060 each: $28.75. The eval is what costs money to reproduce, not the running of it.
Where pricing changes shape
A policy with more rules makes the reasoning longer, not the prompt. The 15-fold spread in output tokens on this corpus tracks which BRANCH fired, not how long the statement was.
A second second-tier party under the same subcontract is a second call, not a longer one. Nothing here batches, and batching would break the per-obligation state.
Your return, with your numbers
Volumeobligation-cycles per pay period — this run judged 150 (50 obligations x 3 cycles) in one pass per arm
What it replacesa subcontract administrator reading each open obligation's current statement and then going back through prior pay applications to work out whether a consecutive-cycle rule has been tripped
Time saved per itemnot measured here — depends on how long applying a joint-check policy by hand takes at the reader's own contractor, which this kit has not observed
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one — the right place to start a question whose answer might be 'a regex is enough', which on three of this kit's four graders it is.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
160,765input tokens · this run
272,643output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 150 pay cycles, one completion call each, one tier. The stateless control is a second 150 and is priced separately below.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.359
$0.359
$2.40
2026-09-12
gemini-3-flash
Google
$0.898
$0.898
$5.99
2026-09-18
gemini-3-8-flash
Google
$1.143
$1.143
$7.62
2026-09-18
llama-5
Meta
$1.360
$1.360
$9.06
2026-09-18
claude-haiku-4-5
Anthropic
$1.524
$1.524
$10.16
2026-09-12
grok-4-5
xAI
$1.957
$1.957
$13.05
2026-09-18
grok-4-6
xAI
$1.957
$1.957
$13.05
2026-09-18
claude-sonnet-5
Anthropic
$3.048
$3.048
$20.32
2026-09-12
gemini-3-1-pro
Google
$3.593
$3.593
$23.95
2026-09-18
gpt-5-6-terra
OpenAI
$3.593
$3.593
$23.95
2026-09-12
gpt-5-6-sol
OpenAI
$6.096
$6.096
$40.64
2026-09-12
claude-opus-4-8
Anthropic
$7.620
$7.620
$50.80
2026-09-12
claude-opus-5
Anthropic
$7.620
$7.620
$50.80
2026-09-12
claude-fable-5
Anthropic
$15.240
$15.240
$101.60
2026-09-18
claude-fable-5-1
Anthropic
$15.240
$15.240
$101.60
2026-09-18
gpt-6-astra
OpenAI
$15.240
$15.240
$101.60
2026-09-17
Read this against the numbers above
PROJECTED, NOT PAID. Every dollar here is a measured token count times a published rate, on a card the run did not use. The provider that actually ran the 315 calls is not named on this page and is not in these tables.
OUTPUT DOMINATES, AND MOST OF THE OUTPUT IS REASONING. 95.3 pct of this run's output tokens were provider-side reasoning. A card with a cheap input rate and an expensive output rate prices this kit much worse than its input length suggests, and the spread across these twelve rows (0.0024 to 0.1016 per cycle, 42x) is mostly a spread in output pricing.
A CEILING IS NOT A COST. MAX_TOKENS is 20,000 and the median reply is nowhere near it. The provider bills tokens produced, not tokens allowed.
THE UNIT IS AN OBLIGATION-CYCLE, NOT A DOCUMENT. Multiply by pay periods per year before comparing this to any one-shot kit's per-document figure.
REASONING WAS NEVER TURNED OFF. Every published figure is at the provider's default. The adapter has the field; this kit never sent it, and every result file records thinking: null so a future run cannot be quietly compared against these.
NOTHING WAS MEASURED ON ANY OF THESE TWELVE MODELS. The token counts are from one tier. A different model would produce different output lengths — probably very different, given how much of this output is reasoning — so these rows are a price comparison, not a performance one.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
16 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus builder
writes the 150 exposure statements and the computed gold from one seed; free, offline, reproducible
the thresholds and the three escalation rules, as arithmetic. Produces the gold and the state; never calls a model
You change it to: REPLACE THIS FIRST. It is your procedure's thresholds and escalation rules, not a utility. The gold is recomputed from whatever you put here, so changing it invalidates every published number on this page and says so.
src/exposure.py
# The exposure policy, and the escalation arithmetic it implies. Pure code, no model.
CLEAR, WATCH, EXPOSED = "CLEAR", "WATCH", "EXPOSED"
NONE = "NONE"
JOINT_CHECK = "JOINT_CHECK"
BOND_CLAIM_NOTICE = "BOND_CLAIM_NOTICE"
RELEASE_HOLD = "RELEASE_HOLD"
POSTURES = (CLEAR, WATCH, EXPOSED)
ACTIONS = (NONE, JOINT_CHECK, BOND_CLAIM_NOTICE, RELEASE_HOLD)
HIGHER = "higher_is_worse"
LOWER = "lower_is_worse"
src/state.pycarried state — a swap seam
four scalar fields between cycles, written from the arithmetic and rendered into one English sentence for the prompt
You change it to: which facts survive between cycles, and how they are rendered in English. Keep it scalar — state that grows with history turns the cost curve quadratic.
src/state.py
# The carried state — the thing that makes this a monitor and not another one-shot kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_unit(store, unit_id):
def advance(store, unit_id, value, watch_at, exposed_at, direction, waiver_received):
def describe(state):
src/cadence.pycadence — a swap seam
which pay period is owed, which have been run, and the refusal that stops one period being run twice
You change it to: how often a run is owed, and what a period is. Monthly is stated, not measured.
src/cadence.py
# The clock. When this monitor wakes, what one run owns that the last did not, and what a skipped
PERIOD_FORMAT = "YYYY-MM"
CADENCE = {
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_REGISTER = os.path.join(HERE, "data", "runs.json")
def _key(period):
def _ord(period):
def _unord(n):
def periods_between(first, last):
def load_register(path=DEFAULT_REGISTER):
src/segment.pysegmenter
splits a statement into its seven named sections
src/segment.py
# Split a pay-cycle exposure statement into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Obligation", "Exposure Policy", "Cycle Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pysection selector — a swap seam
decides which sections reach the prompt; Internal Contacts is mapped by nothing and therefore never sent
You change it to: what leaves your machine. NEVER_SENT is the list to read before pointing this at real statements.
src/select.py
# Pick which sections of an exposure statement are sent. Pure code — the last deterministic step
BANNER = "Synthetic Record"
OBLIGATION = "Obligation"
POLICY = "Exposure Policy"
POSITION = "Cycle Position"
BREAKDOWN = "Exposure Breakdown"
CONTACTS = "Internal Contacts"
NOTES = "Field Notes"
NEVER_SENT = (CONTACTS,)
SECTION_HINTS = {
src/prompt.pyprompt builder
instruction + carried state + selected sections, by string concatenation
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework — string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pymodel adapter
one completion over raw HTTP; provider chosen by .env
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/monitor.pyruntime
one cycle, one call, the reply parsed to four fields
src/monitor.py
# One pay cycle, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You apply a written second-tier exposure policy to one pay cycle of one obligation. "
MAX_TOKENS = 20000
FIELDS = ("posture", "action", "top_driver", "top_driver_share_pct")
def documents():
def units():
def load_doc(doc_id):
def reading_of(text):
src/budget.pybudget guard
counts calls against a shared per-day cap before spending
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
evals/run.pyeval harness
50 obligation chains, strictly serial within a chain, concurrent across them
evals/run.py
# Run the monitor over the 50 obligations and write one result file. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def load_gold():
def stub_complete(cfg, system, user, max_tokens=1024):
def main():
evals/scoring.pyscorer
four graders, exact match, pure code
evals/scoring.py
# Score a run against the computed gold. Pure code, no model, no judge.
POSTURES = ("CLEAR", "WATCH", "EXPOSED")
ACTIONS = ("NONE", "JOINT_CHECK", "BOND_CLAIM_NOTICE", "RELEASE_HOLD")
RESET_PATTERN = "WCW"
RESET_CYCLE = "C3"
def _pct(n, d):
def _norm(v):
def score(records, golds):
def compare(stateful, stateless):
evals/baseline.pyfree floors
threshold arithmetic, and the 60-day ageing rule
evals/baseline.py
# The free, no-model floor: threshold arithmetic plus the ageing rule a subcontract administration
AGEING_DAYS = 60
KINDS = ("ageing", "arith")
def review(text, ageing_days=AGEING_DAYS, kind="ageing"):
def _sweep():
evals/missed_cycle.pycadence ablation
replays every obligation with one run skipped and scores what moved
evals/missed_cycle.py
# What a SKIPPED RUN costs, measured. Free, deterministic, no model call.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def load_gold():
def replay(skip_cycle):
def main():
evals/check_labels.pypre-flight checks
nine assertions that must hold before a run may spend
evals/check_labels.py
# Everything that must be true BEFORE a run may spend. Free, offline, no model call.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def main():
src/app.pylocal UI
one obligation, three cycles, five arms — four of them free
src/app.py
# The minimal local UI. Standard library only — `python3 -m src.app`, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
SCORED = os.path.join(HERE, "results", "eval-r001-subtier-exposure.json")
PORT = int(os.environ.get("PORT", "9008"))
def gold():
def recorded():
def replay_free(unit_id):
class H(BaseHTTPRequestHandler):
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pywrites the 150 exposure statements and the computed gold from one seed; free, offline, reproducible
src/exposure.pythe thresholds and the three escalation rules, as arithmetic. Produces the gold and the state; never calls a model A swap seam.
src/state.pyfour scalar fields between cycles, written from the arithmetic and rendered into one English sentence for the prompt A swap seam.
src/cadence.pywhich pay period is owed, which have been run, and the refusal that stops one period being run twice A swap seam.
src/segment.pysplits a statement into its seven named sections
src/select.pydecides which sections reach the prompt; Internal Contacts is mapped by nothing and therefore never sent A swap seam.
src/prompt.pyinstruction + carried state + selected sections, by string concatenation
src/adapters/__init__.pyone completion over raw HTTP; provider chosen by .env
src/monitor.pyone cycle, one call, the reply parsed to four fields
src/budget.pycounts calls against a shared per-day cap before spending
evals/run.py50 obligation chains, strictly serial within a chain, concurrent across them
evals/scoring.pyfour graders, exact match, pure code
evals/baseline.pythreshold arithmetic, and the 60-day ageing rule
evals/missed_cycle.pyreplays every obligation with one run skipped and scores what moved
evals/check_labels.pynine assertions that must hold before a run may spend
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1071 input and 1817 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py, which generates all 150 statements from a fixed seed with no network access and no user input. The one field that WOULD be attacker-controlled in a real deployment — Field Notes, prose a project accountant types, and in practice the box where a subcontractor's own correspondence gets pasted — is sent to the model deliberately rather than hidden, so that a version pointed at real statements inherits a surface that is visible instead of one nobody looked for.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never printed. config.sources() reports WHICH file contributed and never a value. src/app.py redacts api_key and base_url out of any exception before it reaches the browser. config.save() opens the file with mode 0600 before the write rather than chmod'ing after, so it never exists at the default umask even for an instant. This repository has never held a credential.
The experimentWe did not attack it — and the boundary that matters most was proven by breaking it on purpose
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus generates every one of its own fields — so the trial would have measured nothing. What WAS worth proving is the privacy boundary, and that was proven by breaking it: evals/check_labels.py removes the top_driver hint from src/select.SECTION_HINTS, which is exactly what a statement with no Exposure Breakdown would do to it, and asserts that Internal Contacts still does not reach the prompt. With the unconditional subtraction removed, the naive fallback sends the accountant's name, their employee reference and their desk extension — and no published figure moves, which is why it needs a check rather than a code review. The cadence guard was broken the same way: claim one period twice and the check fails. Confirmed by assertion and by reading the recorded run, not by an attack trial, on 2026-08-23 — no attack run was fired; see redteam.why.
Boundary checked
Boundary
What was actually done
Can a wrong answer poison the next pay cycle?
A monitor that fed the model's verdict back into its own next prompt would compound one bad cycle into every cycle after it. On this use case that means a bond claim filed against a subcontractor because of an arithmetic error two months earlier.
It cannot. src/state.advance() calls src/exposure.step() on the RAW READING parsed from the page; the model's reply is scored and never written. Measured on the scored run: 4 cycles were answered wrongly, in 4 DIFFERENT obligations, and every one of them is a THIRD cycle — the last in its chain. ⚠︎ Read that honestly: the property holds by construction, but this corpus cannot demonstrate it, because nothing follows a C3. What it shows is the absence of the other half — no wrong answer at C1 or C2 went on to corrupt a later cycle, because there were none.
Can the kit issue a joint check, withhold money or file a claim?
An exposure monitor that could act on its own verdict would be imposing contractual consequences on a third party from a model's reading of a spreadsheet.
There is no write path of any kind. Grep the kit: nothing under src/ or evals/ opens a socket except src/adapters/__init__.py, which POSTs one completion request; the only files written are result JSON under results/ and the state and run-register files under data/. src/app.py has exactly one POST route and it asks the model a question. Red-proven by reading every route in the file.
Does the project accountant's name leave the machine?
Every statement names the accountant, their employee reference, a desk extension and a named risk manager. A prompt builder that sends 'the document' sends all of it.
RED-PROVEN IN BOTH DIRECTIONS. src/select.NEVER_SENT holds Internal Contacts and src/select._fallback subtracts it UNCONDITIONALLY rather than falling back to the whole document. evals/check_labels.py check 5 removes the top_driver hint — simulating a statement with no Exposure Breakdown — and asserts the contacts section still does not appear. Without that guard the naive or list(secs) fallback would leak it and no published figure would move.
Can the monitor be made to run a period twice?
A stateless kit can be re-run all day for nothing worse than a second bill. This one advances counters, and advancing a counter is not idempotent.
src/cadence.claim() refuses a period already in the register and raises AlreadyRun rather than returning falsy. evals/check_labels.py check 7 red-proves it: it claims 2025-04 twice and fails the build if the second one succeeds. ⚠︎ The EVAL harness deliberately does not use the register — an eval must be re-runnable — and that exemption is stated in evals/run.py's docstring rather than left implicit.
Each boundary above was checked by running an assertion or by reading the recorded run, not by an attack trial. Two of the four are red-proven — the privacy guard and the cadence guard both fail the pre-flight check if the guard is removed. The other two are arguments from the code path plus a run that is consistent with them.
The result0 attack trials, and four boundaries checked — two of them red-proven by removing the guard and watching the pre-flight check fail.
0untrusted input fields on this corpus
0 of 0attack trials run
2 of 4boundaries red-proven, not just asserted
The Field Notes section IS the field an outside party would control in a real deployment — on a live project it is where a subcontractor's own emailed excuse gets pasted — and this kit sends it rather than hiding it. But on this corpus it is one of eight sentences chosen by a seeded random number generator, so there is nothing adversarial in it to catch. A version pointed at real statements reopens the question and should be attacked before it ships.
Read this twice
The state is written from the arithmetic, never from the model's reply. That is the difference between a monitor you can put in front of a payment decision and one you cannot. src/exposure.step() takes the reading parsed off the page by a regex and the counters the last run left; the model's posture, action and driver are scored and thrown away. A wrong answer costs you that cycle and nothing after it. And the second half of the same sentence: a run that does not happen costs you every cycle after it — measured at 33 of 100 verdicts moved and 22 owed escalations never fired, from one skipped month.
HonestyWhat this does not prove
Whether the Field Notes section can be used to change a verdict. It is the injection surface, it is sent on purpose, and nothing has attacked it.
What the provider retains of the 150 prompts. Not knowable from here and not claimed.
Whether a real accounting export carries fields this kit's selector has no rule for. Every statement here has exactly seven sections because one file wrote them all.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
The carried state is written by code from the parsed reading and the previous state. The model's posture, action and driver are scored against the gold and are never fed forward. Separately: a pay period may be claimed by exactly one run, and a skipped period is visible in the register rather than silently absorbed.
src/state.advance() -> src/exposure.step(), called by evals/run.py after every cycle including one whose call failed. src/cadence.claim() for the period guard.
EvidenceDoes it hold?
What
Measured
A wrong answer does not propagate to the next pay cycle
4 wrong cycles on r001-subtier-exposure, in 4 DIFFERENT obligations — no obligation wrong twice. ⚠︎ Every one of them is a THIRD cycle, so nothing followed any of them: the property is guaranteed by the code path (exposure.step() never reads the reply) and this run is CONSISTENT with it rather than a demonstration of it.
A failed call does not freeze the history
0 failures on this run, so the branch was not exercised live. evals/run.py advances the state on the exception path as well as the success path, because a cycle whose call failed is still a cycle that happened — freezing there would turn one transport error into three scored failures and hide the real one.
A pay period cannot be run twice
Red-proven free in evals/check_labels.py check 7: claiming 2025-04 twice raises cadence.AlreadyRun, and the check fails the build if it does not. This is the guard that stops a monitor imposing a joint check because it ran twice rather than because anything happened on the project.
The contacts section never reaches the provider
Red-proven free in evals/check_labels.py check 5, in both directions: present on today's corpus, and still absent when a field's hint is removed to simulate a statement with no Exposure Breakdown.
The limitWhat a guardrail is not
It is not a check on whether the ANSWER is right. The guardrail stops a wrong answer spreading; it does nothing about the 4 cycles that were wrong in the first place.
It is not a scheduler. src/cadence.py records which periods were owed and run and refuses a double-run; it does not install a cron entry, run a daemon or wake anything up.
It is not a rate limiter or a spend cap on its own. src/budget.py counts calls against a shared per-day ceiling, and on the day this kit was built that ceiling was NOT SET — the harness said so out loud before every run.
It is not a substitute for a person. Every raise on this page is a recommendation to look at an obligation before a pay meeting. Issuing the joint check remains a named human act.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 40 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
10 measured by the latest run30 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
alarm
escalation_accuracy_pct; escalation_when_a_rule_fires_pct; false_alarms; unparsed_replies — alarm on unparsed_replies above 0. A cycle that returns nothing is scored as a miss in all three fields, so a reliability failure arrives disguised as a quality failure. This run had 0 of 150 -- under a 20,000 ceiling its largest reply was 16,828, so the margin was 16 pct and not comfortable.
action-when-a-rule-fires
Action accuracy restricted to the cycles where a rule actually fires
alarm
escalation_when_a_rule_fires_pct; escalations_that_fire — alarm on escalations_that_fire moving away from 50 on a full arm. That denominator is a property of the corpus and is identical on every full arm, so if it moves, no accuracy on this page is comparable with the last one. ⚠︎ It is 32 rather than 50 on the missed-run arm BY CONSTRUCTION and that is not a regression.
false-alarms-on-quiet-cycles
False alarms on the cycles where no rule fires
alarm
false_alarm_rate_pct; false_alarms; unparsed_replies — alarm on any false alarm at all on the scored arm. It is at 0 of 100, so a single one is unambiguous -- and it would be a contractual act against a third party.
counter-reset-subset
Action accuracy on the eleven counter-reset obligations
alarm
counter_reset_accuracy_pct; counter_reset_units — alarm on any drop below 100 pct on a full arm -- an obligation put on a joint check after it recovered, which is the loudest wrong answer this kit can give. ⚑ CHECK src/cadence.missed() FIRST: a missed run produces this exact failure on all eleven, and a reader produces it on none.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
477,059
pay-cycle statements edited — the count held, the bytes did not
split.count
50
the obligations count moved — a different set was scored
split.size_p50
3
the median size of one obligation moved
split.size_p95
3
the 95th-percentile size of one obligation moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (counter_reset_units 11, stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
action accuracy
not yet known
150 pay cycles
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. evals/repeat.py exists to produce one and did not run — the shared provider balance was exhausted at its first call (402 Insufficient Balance). r001 and s001 are one run each of two DIFFERENT prompts, which is a comparison and not a band.
action accuracy where a rule fires
not yet known
50 cycles where a rule fires
No repeat — and this figure is the wrong thing to band anyway. It is 46 correct answers and 4 wrong ones concentrated on one pattern, so its variance across runs would be a property of whether that pattern flips, not of the model's steadiness.
posture accuracy
not yet known
150 pay cycles
No repeat. Note also that the free arithmetic floor scores 75.33 here with zero variance by construction, so a band on this metric is only interesting above that.
false alarms
not yet known
100 quiet cycles
0 of 100 on one run. A zero from a single run is a measurement of that run, not a property of the system, and this page does not claim otherwise.
counter-reset obligations
not yet known
11 third cycles of the WATCH-CLEAR-WATCH obligations
No repeat — and this one is SATURATED rather than steady, which is a different reason to distrust a band on it. The correct answer on all eleven rows is NONE, so the scored run, the stateless control and the free arithmetic floor all sit at 100.0 and a band across them would measure nothing. The only arm that moves it is the cadence ablation, at 0.0, and that is a property of the schedule.
top driver named
not yet known
150 pay cycles
No repeat — and SATURATED on every arm, including both free floors, at 150 of 150. A sort of the Exposure Breakdown table answers this field exactly, so a band here would be a band on a regex. It is charted because the run measured it, not because it discriminates anything.
latency
not yet known
150 completed calls
No repeat, and the p95 is the figure to watch rather than the p50: 64,333 ms against 6,835 — a 9.4x tail, because the reasoning is long exactly on the cycles where a rule fires. A band on the p50 would look reassuring and describe the wrong half of the distribution. ⚠︎ Both figures are END-TO-END per cycle at 14 concurrent chains; they are not a property of the model alone and would move with the worker count.
token volume
not yet known
150 calls, 160,765 in and 272,643 out
No repeat. Input is close to deterministic here — one statement plus a 209-character state sentence — so a band on it would be narrow and uninteresting. OUTPUT is not: 95.3 pct of it is provider-side reasoning left at the provider's default, the per-call spread on this run ran from a few hundred tokens to 16,828, and output is 91.1 pct of the projected bill. That is where a band would actually say something, and no repeat run exists to draw one.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 2 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · with the model in the path — 6 runs. Columns here are only ever compared with each other.
Metric
a000-subtier-exposure-missedcycle 2026-08-23
b000-subtier-exposure-ageing 2026-08-23
b001-subtier-exposure-arith 2026-08-23
c000-subtier-exposure-calibration 2026-08-23
r001-subtier-exposure 2026-08-23
s001-subtier-exposure-stateless 2026-08-23
band accuracy, %
72.00
72.67
75.33
100.00
97.33
75.33
counter reset accuracy, %
0.00
72.73
100.00
100.00
100.00
100.00
escalation accuracy, %
67.00
57.33
66.67
100.00
97.33
66.67
escalation when a rule fires, %
31.25
22.00
0.00
100.00
92.00
0.00
false alarm rate, %
16.18
25.00
0.00
0.00
0.00
0.00
input tokens, whole run
0
0
0
9627
160765
158312
model latency p50 ms
—
0.00
0.00
6240.00
6835.00
5313.00
model latency p95 ms
—
0.00
0.00
60579.00
64333.00
8058.00
output tokens, whole run
0
0
0
15100
272643
87784
top driver accuracy, %
100.0
100.0
100.0
100.0
100.0
100.0
not a time series No two of these 6 runs measured the same system — they differ on counter_reset_units, documents, escalations_that_fire, max_tokens, provider, stateless, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-subtier-exposure-stub 2026-08-23
band accuracy, %
75.33
counter reset accuracy, %
100.0
escalation accuracy, %
66.67
escalation when a rule fires, %
0.0
false alarm rate, %
0.0
input tokens, whole run
169229
model latency p50 ms
0.00
model latency p95 ms
0.00
output tokens, whole run
5197
top driver accuracy, %
100.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 10 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the carried state
action accuracy 66.67 pct -> 97.33 pct, action where a rule fires 0.0 pct -> 92.0 pct, posture accuracy 75.33 pct -> 97.33 pct, false alarms 0.0 pct -> 0.0 pct, cost per cycle $0.0027 -> $0.0060 on the shared card
measured
r001-subtier-exposure against s001-subtier-exposure-stateless. Same corpus, same model, same grader, same 150 cycles; the prompts differ in exactly one line and evals/check_labels.py check 6 asserts they are identical everywhere else.
whether the scheduled run actually happens
action accuracy 97.33 pct -> 67.0 pct on the surviving cycles, action where a rule fires -> 31.25 pct, counter-reset obligations 100.0 pct -> 0.0 pct, and 22 owed escalations never fire at all
measured
a000-subtier-exposure-missedcycle. The POLICY replayed with one monthly run skipped — no model on either side, so this is a property of the schedule and not of any reader. 33 of 100 surviving cycles changed verdict.
the free ageing rule instead of a reader
action accuracy 66.67 pct (do nothing) -> 57.33 pct, false alarms 0.0 pct -> 25.0 pct, in exchange for reaching 22.0 pct of the cycles where a rule fires
measured
b000-subtier-exposure-ageing against b001-subtier-exposure-arith. Both free, both over the same 150 cycles. The incumbent control is worse than doing nothing on this corpus.
the ageing horizon, 20 days to 120 days
action accuracy 46.0 pct -> 66.67 pct, monotonically, reaching its maximum exactly where the rule stops firing at all
measured
evals/baseline.py --sweep, free, no calls. Published at 60 days, which is NOT the peak — the peak is the rule switched off.
the max_tokens ceiling
at 8,000 the hardest calibration cycle came back at 7,681 — 319 tokens from empty. The scored run's largest reply was 16,828, which a 16,000 ceiling would have lost
measured
c000-subtier-exposure-calibration (9 calls at 8,000) and r001-subtier-exposure (150 calls at 20,000, output_tokens_max 16828).
provider-side reasoning
not measured. 95.3 pct of this run's output tokens were reasoning and output is the majority of the projected bill, so this is the largest untouched lever on the page
reasoning
src/adapters has a documented thinking field; this kit's harness never sent it and every result file records thinking: null so a future run that DOES send it cannot be quietly compared against one that did not.
the measure direction
action accuracy 96.0 pct on the 75 higher-is-worse cycles, 98.67 pct on the 75 LOWER-is-worse cycles
measured
r001-subtier-exposure by_measure. Reported apart rather than averaged: a reader that silently assumed one direction would score near-perfectly on one half and near-zero on the other, and the average of those looks like mediocrity on both.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
action accuracy
nothing yet
action accuracy where a rule fires
nothing yet
posture accuracy
nothing yet
false alarms
nothing yet
counter-reset obligations
nothing yet
top driver named
nothing yet
latency
nothing yet
token volume
nothing yet
NextThe three you would add first
A human step in front of any joint check, withholding or bond-claim noticeThe run missed 4 of 50 escalations that were owed, and 3 of those 4 were a joint check already running that it dropped. Both an over-raise and an under-raise land on a third party's cash flow; a workflow without a person in front of it would be imposing contractual consequences from an unreviewed reading.
An alert when a scheduled run does not happenTHIS IS THE ONE MOST TEAMS WILL NOT BUILD. A missed run leaves no document to read and a state file that is perfectly self-consistent. Measured: one skipped monthly run moves 33 of 100 verdicts, silently drops 5 joint checks, 12 release holds and 5 bond-claim notices, and gets all 11 counter-reset obligations wrong. src/cadence.missed() is the three lines that see it; nothing in this kit pages anybody.
A second reader on any obligation open more than one cycleAll 4 wrong cells are third cycles. The error rate on first cycles is 0 of 50.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/exposure.py, src/segment.py, src/select.py, src/state.py, src/cadence.py or src/prompt.py — it re-derives the gold, re-generates the corpus into a temporary directory and diffs, and red-proves the privacy and double-run guards. Re-run both free floors and the cadence ablation (also free) whenever the policy changes, because all three are properties of the policy rather than of any model. The two paid arms only need re-running when the model, the ceiling or the prompt moves — and if either moves, BOTH must be re-run, or the 30.66-point headline is comparing two different experiments.
What this cannot tell you
That the no-propagation guarantee holds in practice on a longer chain. Every wrong answer on this run is a third cycle, so nothing followed one.
That the double-run guard is enough in a real deployment. It stops a second run of the same PERIOD; it does nothing about two processes running the same period concurrently, and src/cadence.py has no lock.
Whether the verdicts are stable. No repeat probe was run — the shared balance ran out.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all — json, os, re, random, argparse, urllib, http.server and concurrent.futures from the standard library, and nothing else. requirements.txt is a comment explaining why it is empty. The whole kit is 16 Python files you can read in an hour, and the two that carry the argument (src/exposure.py and src/cadence.py) are under 200 lines between them.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a VERDICT — four scalar fields written by arithmetic. The whole reason the cost is flat in history length is that nothing here remembers what was said, only what was decided, and a memory layer would give back the growth this design exists to avoid.
the cadence
src/cadence.py
schedulers and orchestrators (Airflow, Dagster, Temporal, a cron entry)
this is deliberately NOT one. It is the bookkeeping a scheduler would drive — which period is owed, which have been run, whether one was skipped, and the refusal that stops a double-run. The scheduler is the adopter's platform and this kit does not install one; what it does is make the thing a scheduler must not get wrong explicit and testable.
the policy
src/exposure.py
rules engines (Drools, JSON Logic, a decision table in a BI tool)
three branches and two counters is not a rules engine's problem. What a rules engine would add here is a place for the rules to be edited by somebody who cannot read Python — which is a real need on a subcontract administration team and is a reason to reach for one, not a reason this kit ships without one.
the model call
src/adapters/__init__.py
vendor SDKs and multi-provider clients (LiteLLM, LangChain chat models)
one completion over raw HTTP, 200 lines, because a forker runs this on whichever key they already hold. Adding a provider is one function and one dict entry; adding an SDK would bake a vendor preference into the exact file whose purpose is not having one.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each obligation is a chain of three cycles processed strictly in order, carrying four scalars. There is no branching, no tool selection, no retry-with-reflection and no sub-agent. Drawing it as a graph would suggest a control flow that does not exist and would hide the only structural fact that matters: cycles inside one obligation are serial and different obligations are independent, which is why the run is 50 chains wide rather than 150 documents deep.
The other sideWhat a framework costs you
No framework means no version to pin and nothing to break on someone else's release. It also means no retry policy, no tracing, no cost dashboard and no queue — every one of which a real deployment of this would want.
A rules engine would let a subcontract administrator edit the escalation rules without editing Python. That is a genuine cost of shipping the policy as code, and on this use case the person who owns the policy is usually not the person who owns the repository.
A scheduler is the piece this kit most obviously lacks, and the cadence ablation is the measurement of what its absence costs. Anyone deploying this needs one; what they should NOT do is assume the scheduler is the easy part, because the failure it causes is invisible.
No tracing means the 4 wrong cells can only be examined through the recorded result file. That was enough here because the run is 150 rows; at 400 obligations a month it would not be.
What we could NOT verify
Whether a memory layer would score better than four scalars. Not run — it would need a second 150-cycle arm and the balance was gone.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-subtier-exposure on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
6,835 ms
not yet known
nothing yet
Model, p95
64,333 ms
not yet known
nothing yet
Input tokens
160,765
not yet known
nothing yet
Output tokens
272,643
not yet known
nothing yet
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-subtier-exposure-calibration6,240 ms
r001-subtier-exposure6,835 ms
s001-subtier-exposure-stateless5,313 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. a000-subtier-exposure-missedcycle, b000-subtier-exposure-ageing, b001-subtier-exposure-arith recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
pay-cycle statements
data/corpus/<OBLIGATION>-C<n>.txt — 150 files, 477,059 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Internal Contacts never does, by src/select.NEVER_SENT
the answer key
data/gold.jsonl — 150 rows, the output of src/exposure.replay over the corpus values, never hand-authored
never — evals/scoring.py is pure code, no model and no key
the carried state
data/state.json in a deployment, written atomically by src/state.save; in an eval it is scoped to the run and never touches disk, so a run can be repeated
one rendered English sentence of it goes into each prompt — 209 characters. The JSON itself never leaves the machine
the run register
data/runs.json — which pay period was run by which run id. The only thing in the kit that can see a MISSED run, because a missed run leaves no document
never
the provider credential
.env (gitignored) or the shared repo-root .env, merged key by key by src/config.load with the real environment winning
as an Authorization header on the one POST src/adapters makes, and nowhere else. src/app.py redacts it out of any exception before it reaches a browser
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never printed. config.sources() reports WHICH file contributed and never a value. src/app.py redacts api_key and base_url out of any exception before it reaches the browser. config.save() opens the file with mode 0600 before the write rather than chmod'ing after, so it never exists at the default umask even for an instant. This repository has never held a credential.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
ONE RUN PER PAY PERIOD, and one run owns exactly one period for every obligation open in it: read this period's statement, apply the policy against the state the last run left, write the state the next run will read. src/cadence.claim() REFUSES a period already in the register, because advancing a counter is not idempotent — run 2025-05 twice and every obligation at WATCH for one cycle is reported EXPOSED and put on a joint check because the monitor ran twice, not because anything happened on the project. src/cadence.missed() is the only thing in the kit that can see a run that did not happen, since a missed run leaves no document to read and a state file that is perfectly self-consistent.
One skipped monthly run moves 33 of the 100 surviving cycles, drops 5 joint checks, 12 release holds and 5 bond-claim notices that were owed, and gets 0 of 11 counter-reset obligations right. Action accuracy on the surviving cycles falls from 97.33 pct to 67.0 pct and posture from 97.33 pct to 72.0 pct. The double-run refusal is red-proven free in evals/check_labels.py check 7. (a000-subtier-exposure-missedcycle, evals/check_labels.py)
monthly is STATED, NOT MEASURED — nothing here observed a real billing calendar, and a contractor billing twice a month or on a rolling 30-day cycle changes what one run owns. The ablation also skips exactly ONE cycle of three; two consecutive missed runs, or a gap in a longer chain, is the real operational case and this corpus is too short to hold one.
a different period length invalidates the cadence figures rather than scaling them: the escalation rules are written in CONSECUTIVE CYCLES, so halving the period halves the time a joint check takes to fire and changes what the policy means, not just what it costs.
state
FOUR SCALAR FIELDS between cycles — prev_reported, prev_raw, consec_open, consec_exposed — written by src/exposure.step() from the RAW READING parsed off the page, and rendered into one English sentence for the prompt by src/state.describe(). ⚠︎ THE MODEL'S REPLY IS NEVER WRITTEN BACK. It is scored and discarded, so one bad cycle cannot compound into every cycle after it. That is the difference between a monitor you can put in front of a payment decision and one you cannot.
The state sentence is 209 characters, 4.6 pct of a 4,567-character prompt, and it is worth 30.66 points of action accuracy and 92.0 points on the cycles where a rule fires (0.0 -> 92.0). Cost is FLAT in history length: cycle 30 costs what cycle 3 costs, 1,071.77 input tokens. 4 wrong cycles on the scored run, in 4 different obligations, none wrong twice. (r001-subtier-exposure against s001-subtier-exposure-stateless)
first-order only. The state remembers last cycle plus two counters, so a policy whose rules reach back three cycles cannot be expressed at all. It is also per-obligation: nothing here models two suppliers under one subcontract whose exposures interact.
growing the state past scalars invalidates the flat cost curve, which is the whole reason this is not an intake kit. A state that carried prior statements would put input cost on a quadratic in the number of cycles.
model
ONE completion call per pay cycle, over raw HTTP in src/adapters/__init__.py, to whatever PROVIDER / BASE_URL / MODEL the merged .env names. No SDK, no framework, no second call and no retry-with-reflection. MAX_TOKENS = 20000, and the harness refuses to override it for any run id not beginning with c, so a figure measured under a ceiling the page does not name cannot be mistaken for a scored one.
150 of 150 replies parsed, 0 failures. Output tokens ran to 16,828 against the 20,000 ceiling; the 15-call calibration had topped out at 7,681, so the scored run exceeded its own calibration by 2.2x and a 16,000 ceiling — which the calibration would have justified — would have lost that reply entirely. p50 6,835 ms, p95 64,333 ms, 95.3 pct of output tokens provider-side reasoning. (r001-subtier-exposure, c000-subtier-exposure-calibration)
one tier, one provider, one run. Nothing here says how any other model behaves, and given that 95.3 pct of the output is reasoning, output length on a different model is not predictable from these numbers. Concurrency was 14 workers over 50 chains; 400 chains is untested and the provider's rate limits are not known to this kit.
changing the model or the ceiling invalidates BOTH paid arms at once — the 30.66-point headline is a difference between two runs, so re-running one without the other compares two different experiments.
labels
150 labelled pay cycles, computed by src/exposure.replay rather than authored, over 50 obligations in 10 deliberately chosen patterns. 11 of the 50 are the counter-reset trap and 2 are its mirror. Half the corpus is higher-is-worse and half is LOWER-is-worse, scored apart. The scoring stops at exact match on three closed sets — there is no judge anywhere in this kit.
150 rows, 50 cycles where a rule fires and 100 quiet, 42 CLEAR / 61 WATCH / 47 EXPOSED postures, and an action mix of 100 NONE / 23 JOINT_CHECK / 17 RELEASE_HOLD / 10 BOND_CLAIM_NOTICE. Every gold row is re-derived from the corpus's own parsed numbers by evals/check_labels.py before a run may spend. (data/gold.jsonl, evals/check_labels.py, dataset version subtier-exposure-2026-08-23-50obligations-150cycles)
these are labels about a corpus this kit wrote. They are correct by construction, which means they measure whether the rules are applicable and not whether they are applied to anything real. No real exposure pack has been through this kit and no figure here is a prediction about one.
replacing src/exposure.py replaces the gold, and therefore every number on this page. That is stated rather than hidden: the kit recomputes the labels from whatever policy you put there, so a forker's first edit invalidates the headline by design.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a third consecutive open cycle reported WATCH with action NONE
the reader treated the third cycle as though the obligation were new, dropping a joint check that had been running since the second cycle. This is the run's dominant failure and it is one shape: 3 of the 4 wrong cells.
check the pattern before disputing the answer. Every wrong cell on this run is a THIRD cycle; the error rate on first and second cycles is 0 of 100. If your own obligation is on cycle 3 or later of a continuous run, read the non-raises as carefully as the raises. (results/eval-r001-subtier-exposure.json misses — STO-0037-C3, STO-0039-C3, and STO-0040-C3 which answers RELEASE_HOLD instead)
an obligation reported CLEAR the cycle after it was EXPOSED
the release hold did not fire. The policy holds the file open for one cycle after an EXPOSED reading even when the reading alone is clear, because the unconditional waiver for the prior payment has not landed. Answering CLEAR closes it a cycle early.
look at the carried state printed beside the verdict. If prev_reported is EXPOSED and the action is NONE, the rule was not applied — 1 of 17 release holds on this run. (results/eval-r001-subtier-exposure.json misses — STO-0029-C3)
a joint check on an obligation whose middle cycle read clear
something is counting open cycles rather than CONSECUTIVE open cycles — or a scheduled run did not happen and the clear cycle was never observed. Both produce the identical wrong answer and the state file looks correct either way.
check src/cadence.missed() FIRST, before suspecting the reader. On this corpus the model gets all 11 counter-reset obligations right and a single missed run gets all 11 wrong, so a raise of this shape is far more likely to be a schedule failure than a reading failure. (results/eval-a000-subtier-exposure-missedcycle.json — counter_reset accuracy 0.0 pct, against 100.0 pct on results/eval-r001-subtier-exposure.json)
a joint check raised on an obligation with a waiver on file
an ageing rule is firing on a day count alone. The 60-day free floor does exactly this and it is the commonest incumbent control on this problem.
compare against the arithmetic floor before assuming the ageing rule helps. On this corpus it imposes 25 joint checks on quiet cycles and reaches only 22 pct of the ones that were owed — nine points WORSE than doing nothing. (results/eval-b000-subtier-exposure-ageing.json against results/eval-b001-subtier-exposure-arith.json)
⚠︎ THE REPEAT PROBE WAS NOT RUN AND THAT IS THE BIGGEST HOLE ON THIS PAGE. evals/repeat.py is written, committed and runnable — 6 obligations chosen for the rules they exercise, 3 cycles each, 2 fresh asks on top of the recorded one — and it failed at its FIRST call with 402 Insufficient Balance on the shared provider key, which nine kits were drawing on the same day. So nothing here says whether the same pay cycle answers the same way twice, and on this use case that matters more than usual: a joint check imposed one month and not the next, on identical facts, is a dispute rather than a control. 24 ledger lines exist for the attempt and no completion came back for any of them.
⚠︎ ALSO NOT MEASURED: whether provider-side reasoning can be turned off without losing accuracy (95.3 pct of output tokens, and the largest untouched cost lever here); whether the English state sentence beats a JSON one; whether two consecutive missed runs are worse than one or differently wrong; concurrency beyond 14 workers and 50 chains, against the 400-chain breadth the Architecture lens names as the real scaling axis; the provider's own rate limits; what the provider retains of the 150 prompts; GPU sizing, because nothing here is self-hosted; and anything at all about a real second-tier exposure pack, because none has been through this kit.
⚠︎ AND ONE CALIBRATION RUN IS NOT IN THE RUN REGISTER. c001-subtier-exposure-calibration's 6 calls are committed as a result file and readable, but the estate's extractor keys this shape on a counter-reset figure and that two-obligation slice contained no counter-reset obligation, so the value is null and the record refuses. The published ceiling cites c000, which does have one.
The corpus licence, from the Data lens: MIT — this repository's own licence. Nothing is fetched and no third-party grant is involved; every statement is generated in-process on a fresh clone. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineThe reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
For each pay cycle and each of the three answered fields, did the reply equal the computed answer key? Posture and action are compared exactly; the driver name is compared on trimmed lower-cased text and nothing looser, because a fuzzy match would forgive a reader that invented a plausible component.
$0.00per 1,000 pay-cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/missed_cycle.py call.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The pay cycle
STO-0022-C3 — Lakeshore Precast under Halloran Sitework, waiver_coverage
Reading and thresholds
95.77 pct coverage, against Watch at 85.6 and Exposed at 65.1 (LOWER is worse), waiver received this cycle
Carried state, in the prompt
Last cycle this obligation was reported EXPOSED. Its reading alone last cycle was WATCH, before the escalation rules were applied. It has now been at WATCH or worse for 2 consecutive cycles, counting last cycle.
The free floors
CLEAR / NONE — correct threshold arithmetic, and no rule, because it has no last cycle. The ageing floor says the same: 31 days, waiver received.
The model
WATCH / RELEASE_HOLD — "Reading alone is CLEAR at 95.77% coverage, but because last cycle was reported EXPOSED, Rule 1 requires reporting WATCH with a RELEASE_HOLD until the prior unconditional waiver lands."
The same policy, one run missed
CLEAR / NONE — the same policy, no model involved, with the cycle-2 run skipped. It never observed the EXPOSED cycle, so there is nothing to hold and the file closes.
Top driver
Payments with no waiver on file, 60.5 pct
Ground truth
WATCH / RELEASE_HOLD / Payments with no waiver on file
Scored as
3 of 3 cells correct — and the posture both free floors got right on 113 of 150 cycles is wrong here for exactly the reason this kit exists.
Grader
Verdict
Why
The reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
posture hit, action hit, top driver hit
STO-0022-C3, a precast supplier under a sitework subcontract, measured on waiver_coverage where a LOWER number is worse. The reading is 95.77 pct against Watch at 85.6 and Exposed at 65.1, so the reading alone is CLEAR and BOTH free floors answer CLEAR / NONE. The carried state said the obligation was reported EXPOSED last cycle and has been at WATCH or worse for two consecutive cycles; the model answered WATCH / RELEASE_HOLD, citing rule 1, and named Payments with no waiver on file at 60.5 pct. All three cells match the computed key.
Action accuracy restricted to the cycles where a rule actually fires
hit -- one of the 50
This cycle's gold action is RELEASE_HOLD, so it is inside the 50 this grader scores. It is one of the 46 the stateful arm got right. Both free floors and the stateless control answered NONE here -- three of their fifty structural zeros -- and so did the missed-run arm, which is one of its 12 lost release holds.
False alarms on the cycles where no rule fires
not in scope
This grader scores only the 100 cycles whose correct action is NONE. A rule fires here, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
Action accuracy on the eleven counter-reset obligations
not in scope
This grader scores only the third cycle of the eleven WATCH-CLEAR-WATCH obligations. STO-0022 is WATCH-WATCH-CLEAR, so the row is outside its denominator. Its eleven rows all have gold NONE, which is why three of the five arms score 11 of 11 on it and the page reads it as saturated -- and why the missed-run arm's 0 of 11 is the only thing it genuinely discriminates.
The formulaWhat it computes
accuracy = hits / cycles scored, per field. A cycle whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
97.3% band accuracy · 2 more measured on this row
the fast tier, memory removed (THE CONTROL)
75.3% band accuracy · 2 more measured on this row
the free arithmetic floor, no model
75.3% band accuracy · 2 more measured on this row
the free 60-day ageing rule, no model
72.7% band accuracy · 2 more measured on this row
the same policy, one scheduled run skipped
72.0% band accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/exposure.replay over the corpus values at generation time. This grader IS the reference, so its own error rate is not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate cannot be measured: it IS the reference. What can be wrong is the KEY, and the key is src/exposure.py -- an invented procedure. A forker who replaces it replaces every number this grader has ever produced, which the kit says out loud rather than hiding.
Watch these
escalation_accuracy_pct
escalation_when_a_rule_fires_pct
false_alarms
unparsed_replies
Alarm on
unparsed_replies above 0. A cycle that returns nothing is scored as a miss in all three fields, so a reliability failure arrives disguised as a quality failure. This run had 0 of 150 -- under a 20,000 ceiling its largest reply was 16,828, so the margin was 16 pct and not comfortable.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because two of them are small: 50 cycles where a rule fires and 11 counter-reset cycles, so one row moves the first by 2.0 points and the second by 9.1.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/exposure.py, src/segment.py, src/select.py or src/cadence.py. Re-run BOTH paid arms (150 calls each) on any change to src/prompt.py or src/state.py -- both change what the model is told, and re-running one without the other compares two different experiments.
The decisionWhen to reach for it
Use it
The truth is known and every answer is a word from a closed list.
Do not use it
The truth is not known -- the normal state of a real subcontract administration procedure, whose escalations are decided by a person at a pay meeting and not by a replayable function. That is why this corpus is generated rather than captured.
Action accuracy restricted to the cycles where a rule actually fires
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineAction accuracy restricted to the cycles where a rule actually fires
Of the cycles whose correct action is not NONE, how many were named correctly? This is the number the kit is actually about; the 150-cycle average is two-thirds NONE and a monitor that says nothing scores 66.67 pct on it.
$0.00per 1,000 pay-cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader -- it is a slice of the same cells, not a second comparison.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The pay cycle
STO-0022-C3 — Lakeshore Precast under Halloran Sitework, waiver_coverage
Reading and thresholds
95.77 pct coverage, against Watch at 85.6 and Exposed at 65.1 (LOWER is worse), waiver received this cycle
Carried state, in the prompt
Last cycle this obligation was reported EXPOSED. Its reading alone last cycle was WATCH, before the escalation rules were applied. It has now been at WATCH or worse for 2 consecutive cycles, counting last cycle.
The free floors
CLEAR / NONE — correct threshold arithmetic, and no rule, because it has no last cycle. The ageing floor says the same: 31 days, waiver received.
The model
WATCH / RELEASE_HOLD — "Reading alone is CLEAR at 95.77% coverage, but because last cycle was reported EXPOSED, Rule 1 requires reporting WATCH with a RELEASE_HOLD until the prior unconditional waiver lands."
The same policy, one run missed
CLEAR / NONE — the same policy, no model involved, with the cycle-2 run skipped. It never observed the EXPOSED cycle, so there is nothing to hold and the file closes.
Top driver
Payments with no waiver on file, 60.5 pct
Ground truth
WATCH / RELEASE_HOLD / Payments with no waiver on file
Scored as
3 of 3 cells correct — and the posture both free floors got right on 113 of 150 cycles is wrong here for exactly the reason this kit exists.
Grader
Verdict
Why
The reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
posture hit, action hit, top driver hit
STO-0022-C3, a precast supplier under a sitework subcontract, measured on waiver_coverage where a LOWER number is worse. The reading is 95.77 pct against Watch at 85.6 and Exposed at 65.1, so the reading alone is CLEAR and BOTH free floors answer CLEAR / NONE. The carried state said the obligation was reported EXPOSED last cycle and has been at WATCH or worse for two consecutive cycles; the model answered WATCH / RELEASE_HOLD, citing rule 1, and named Payments with no waiver on file at 60.5 pct. All three cells match the computed key.
Action accuracy restricted to the cycles where a rule actually fires
hit -- one of the 50
This cycle's gold action is RELEASE_HOLD, so it is inside the 50 this grader scores. It is one of the 46 the stateful arm got right. Both free floors and the stateless control answered NONE here -- three of their fifty structural zeros -- and so did the missed-run arm, which is one of its 12 lost release holds.
False alarms on the cycles where no rule fires
not in scope
This grader scores only the 100 cycles whose correct action is NONE. A rule fires here, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
Action accuracy on the eleven counter-reset obligations
not in scope
This grader scores only the third cycle of the eleven WATCH-CLEAR-WATCH obligations. STO-0022 is WATCH-WATCH-CLEAR, so the row is outside its denominator. Its eleven rows all have gold NONE, which is why three of the five arms score 11 of 11 on it and the page reads it as saturated -- and why the missed-run arm's 0 of 11 is the only thing it genuinely discriminates.
The formulaWhat it computes
hits among cells where gold action != NONE, divided by the count of those cells.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
92.0% accuracy
the fast tier, memory removed (THE CONTROL)
0.0% accuracy
the free arithmetic floor, no model
0.0% accuracy
the free 60-day ageing rule, no model
22.0% accuracy
the same policy, one scheduled run skipped
31.2% accuracy
In operationWhat to monitor
Reference standard: the same data/gold.jsonl the reference grader uses, restricted to the cells whose action is not NONE. It has no independent standard of its own.
These rates are UNKNOWN, on purpose
It has nothing to be scored against, because it is not answering a different question from the reference grader -- it is the same comparison over a subset.
Watch these
escalation_when_a_rule_fires_pct
escalations_that_fire
Alarm on
escalations_that_fire moving away from 50 on a full arm. That denominator is a property of the corpus and is identical on every full arm, so if it moves, no accuracy on this page is comparable with the last one. ⚠︎ It is 32 rather than 50 on the missed-run arm BY CONSTRUCTION and that is not a regression.
How tight can the band be? 50 rows means the finest honest band is roughly 2 points. Nothing is tuned; there is no threshold on the model path to sweep.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
The corpus is mostly quiet, which every real one is.
Do not use it
As a headline on its own. It ignores the 100 quiet cycles entirely, so a monitor that escalated everything would score 100 pct here and impose a joint check on every subcontractor on the job.
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineFalse alarms on the cycles where no rule fires
Of the cycles whose correct action is NONE, how many got an escalation anyway? On this use case that is not a nuisance page -- it is a joint check imposed on a subcontractor's cash flow, and it is the direction the free ageing rule fails in.
$0.00per 1,000 pay-cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The pay cycle
STO-0022-C3 — Lakeshore Precast under Halloran Sitework, waiver_coverage
Reading and thresholds
95.77 pct coverage, against Watch at 85.6 and Exposed at 65.1 (LOWER is worse), waiver received this cycle
Carried state, in the prompt
Last cycle this obligation was reported EXPOSED. Its reading alone last cycle was WATCH, before the escalation rules were applied. It has now been at WATCH or worse for 2 consecutive cycles, counting last cycle.
The free floors
CLEAR / NONE — correct threshold arithmetic, and no rule, because it has no last cycle. The ageing floor says the same: 31 days, waiver received.
The model
WATCH / RELEASE_HOLD — "Reading alone is CLEAR at 95.77% coverage, but because last cycle was reported EXPOSED, Rule 1 requires reporting WATCH with a RELEASE_HOLD until the prior unconditional waiver lands."
The same policy, one run missed
CLEAR / NONE — the same policy, no model involved, with the cycle-2 run skipped. It never observed the EXPOSED cycle, so there is nothing to hold and the file closes.
Top driver
Payments with no waiver on file, 60.5 pct
Ground truth
WATCH / RELEASE_HOLD / Payments with no waiver on file
Scored as
3 of 3 cells correct — and the posture both free floors got right on 113 of 150 cycles is wrong here for exactly the reason this kit exists.
Grader
Verdict
Why
The reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
posture hit, action hit, top driver hit
STO-0022-C3, a precast supplier under a sitework subcontract, measured on waiver_coverage where a LOWER number is worse. The reading is 95.77 pct against Watch at 85.6 and Exposed at 65.1, so the reading alone is CLEAR and BOTH free floors answer CLEAR / NONE. The carried state said the obligation was reported EXPOSED last cycle and has been at WATCH or worse for two consecutive cycles; the model answered WATCH / RELEASE_HOLD, citing rule 1, and named Payments with no waiver on file at 60.5 pct. All three cells match the computed key.
Action accuracy restricted to the cycles where a rule actually fires
hit -- one of the 50
This cycle's gold action is RELEASE_HOLD, so it is inside the 50 this grader scores. It is one of the 46 the stateful arm got right. Both free floors and the stateless control answered NONE here -- three of their fifty structural zeros -- and so did the missed-run arm, which is one of its 12 lost release holds.
False alarms on the cycles where no rule fires
not in scope
This grader scores only the 100 cycles whose correct action is NONE. A rule fires here, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
Action accuracy on the eleven counter-reset obligations
not in scope
This grader scores only the third cycle of the eleven WATCH-CLEAR-WATCH obligations. STO-0022 is WATCH-WATCH-CLEAR, so the row is outside its denominator. Its eleven rows all have gold NONE, which is why three of the five arms score 11 of 11 on it and the page reads it as saturated -- and why the missed-run arm's 0 of 11 is the only thing it genuinely discriminates.
The formulaWhat it computes
cells where gold action == NONE and the reply did not match, divided by the count of those cells.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
0.0% false alarm rate
the fast tier, memory removed (THE CONTROL)
0.0% false alarm rate
the free arithmetic floor, no model
0.0% false alarm rate
the free 60-day ageing rule, no model
25.0% false alarm rate
the same policy, one scheduled run skipped
16.2% false alarm rate
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to the cells whose action is NONE.
These rates are UNKNOWN, on purpose
The scored run raised ZERO in 100. A zero from a single run bounds the true rate loosely and no repeat run exists to narrow it -- evals/repeat.py was written for exactly that and could not be paid for.
Watch these
false_alarm_rate_pct
false_alarms
unparsed_replies
Alarm on
any false alarm at all on the scored arm. It is at 0 of 100, so a single one is unambiguous -- and it would be a contractual act against a third party.
How tight can the band be? 100 rows, so one row is 1.0 point. No threshold is tuned.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Always, and beside the fires rate rather than under it. The two move in opposite directions and a single accuracy figure hides the trade entirely -- which is exactly what it does to the ageing floor, whose 22 pct reach costs 25 false alarms.
Do not use it
As a measure of caution on an arm that answers NONE by construction -- there it is 0.00 pct and means nothing at all. Two of the five arms below are in that state.
Action accuracy on the eleven counter-reset obligations
Catch payment risk to your subcontractor's suppliers
PresenterOpens the private repo. Visible to admins only.
In one lineAction accuracy on the eleven counter-reset obligations
On the third cycle of the eleven WATCH-CLEAR-WATCH obligations, was the action named correctly? WATCH-CLEAR-WATCH looks like two open cycles to anything counting open cycles rather than counting CONSECUTIVE ones, so it is the shape where a plausible-looking monitor puts a recovered obligation on a joint check.
$0.00per 1,000 pay-cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader. It is deliberately NOT added to any headline -- it double-counts eleven cycles already inside the action figure.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The pay cycle
STO-0022-C3 — Lakeshore Precast under Halloran Sitework, waiver_coverage
Reading and thresholds
95.77 pct coverage, against Watch at 85.6 and Exposed at 65.1 (LOWER is worse), waiver received this cycle
Carried state, in the prompt
Last cycle this obligation was reported EXPOSED. Its reading alone last cycle was WATCH, before the escalation rules were applied. It has now been at WATCH or worse for 2 consecutive cycles, counting last cycle.
The free floors
CLEAR / NONE — correct threshold arithmetic, and no rule, because it has no last cycle. The ageing floor says the same: 31 days, waiver received.
The model
WATCH / RELEASE_HOLD — "Reading alone is CLEAR at 95.77% coverage, but because last cycle was reported EXPOSED, Rule 1 requires reporting WATCH with a RELEASE_HOLD until the prior unconditional waiver lands."
The same policy, one run missed
CLEAR / NONE — the same policy, no model involved, with the cycle-2 run skipped. It never observed the EXPOSED cycle, so there is nothing to hold and the file closes.
Top driver
Payments with no waiver on file, 60.5 pct
Ground truth
WATCH / RELEASE_HOLD / Payments with no waiver on file
Scored as
3 of 3 cells correct — and the posture both free floors got right on 113 of 150 cycles is wrong here for exactly the reason this kit exists.
Grader
Verdict
Why
The reported posture, the action code and the top driver, per pay cycle, exact match against the computed answer key
posture hit, action hit, top driver hit
STO-0022-C3, a precast supplier under a sitework subcontract, measured on waiver_coverage where a LOWER number is worse. The reading is 95.77 pct against Watch at 85.6 and Exposed at 65.1, so the reading alone is CLEAR and BOTH free floors answer CLEAR / NONE. The carried state said the obligation was reported EXPOSED last cycle and has been at WATCH or worse for two consecutive cycles; the model answered WATCH / RELEASE_HOLD, citing rule 1, and named Payments with no waiver on file at 60.5 pct. All three cells match the computed key.
Action accuracy restricted to the cycles where a rule actually fires
hit -- one of the 50
This cycle's gold action is RELEASE_HOLD, so it is inside the 50 this grader scores. It is one of the 46 the stateful arm got right. Both free floors and the stateless control answered NONE here -- three of their fifty structural zeros -- and so did the missed-run arm, which is one of its 12 lost release holds.
False alarms on the cycles where no rule fires
not in scope
This grader scores only the 100 cycles whose correct action is NONE. A rule fires here, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
Action accuracy on the eleven counter-reset obligations
not in scope
This grader scores only the third cycle of the eleven WATCH-CLEAR-WATCH obligations. STO-0022 is WATCH-WATCH-CLEAR, so the row is outside its denominator. Its eleven rows all have gold NONE, which is why three of the five arms score 11 of 11 on it and the page reads it as saturated -- and why the missed-run arm's 0 of 11 is the only thing it genuinely discriminates.
The formulaWhat it computes
hits among action cells where pattern == WCW and cycle == C3, divided by 11.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% accuracy
the fast tier, memory removed (THE CONTROL)
100.0% accuracy
the free arithmetic floor, no model
100.0% accuracy
the free 60-day ageing rule, no model
72.7% accuracy
the same policy, one scheduled run skipped
0.0% accuracy
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to the 11 WCW third cycles.
These rates are UNKNOWN, on purpose
What a corpus that could discriminate a READER here would look like. This one cannot: because the reset rule's correct answer is NONE, a monitor that has understood it and one that has never heard of it produce the same eleven answers. Measuring it properly needs a pattern where getting it right means SAYING something, and this corpus has none.
Watch these
counter_reset_accuracy_pct
counter_reset_units
Alarm on
any drop below 100 pct on a full arm -- an obligation put on a joint check after it recovered, which is the loudest wrong answer this kit can give. ⚑ CHECK src/cadence.missed() FIRST: a missed run produces this exact failure on all eleven, and a reader produces it on none.
How tight can the band be? 11 rows. One row is 9.1 points.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Reading a monitor's failure modes rather than its score. The eleven rows are invisible in a 150-cycle average.
Do not use it
⚠︎ AS EVIDENCE THAT A MONITOR UNDERSTANDS THE RESET RULE, WHICH IS WHAT IT LOOKS LIKE AND IS NOT WHAT IT MEASURES. The correct answer on all eleven is NONE, so three of the five arms score 100 pct -- including the free arithmetic floor, which answers NONE to everything and has never heard of a counter. It is saturated, and the page says so rather than banking it. ⚑ WHAT IT DOES DISCRIMINATE IS THE SCHEDULE: the missed-run arm scores 0 of 11 on exactly these rows.
A living map of modern AI — kept current every morning