Catch a dealer's warranty claims turning into a pattern
Every week, warranty audit has to work out which dealers are gaming claims and which are just unlucky, mostly from memory. This app rereads the whole claim book, names the pattern behind each dealer's claims and says whether they're getting worse or quieter.
PresenterOpens the private repo. Visible to admins only.
For the warranty auditorAutomotive · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A warranty auditor at a vehicle manufacturer, checking every dealer's open claims each week before deciding who needs a closer look.
✕Today's manual process
1Pull each dealer's claim book and work out which duplicates, repeat repairs and inflated lines are already known.
2Read the investigator's notes to see which findings were declined, and whether the reason was real.
3Decide if the dealer is escalating or quietly closing out, mostly from memory and a spreadsheet.
4Miss a real escalation and the 30-day window to act on it closes.
Every dealer checked manually each week
✓With the app
1The claim book is re-read in full each week, so nothing already known is raised again.
2Declined notes are read for you and weighed for whether the reason actually holds up.
3A level and a direction are named escalating, easing or unchanged, against last week's position.
4Every dealer that needs a closer look is on the queue before the window to act closes.
Only dealers needing a look reach the queue
See it work
One real case: what the app reads, step by step
Dealer DLR-0040-W2 picks up three new duplicate claims and a declined note this week, and moves from WATCH to REVIEW.
Catch a dealer's warranty claims turning into a patternReference appBuilt to be shaped to your process
6
1Last week's position WATCH last time; the book ended at CLM-100255.
2Newly countable $4,119.09 of claims, all still inside the chargeback window.
3The level it lands on REVIEW, the first time this dealer has reached it.
4Three new claims counted since last week, so the dealer is getting worse.
5Which way it's moving WORSENING, up from WATCH last week.
6The reason in plain words the claims named by number, and why the declined note didn't clear them.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch a dealer's warranty claims turning into a pattern
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
'Which dealers have duplicate warranty claims' is a SELECT and nobody needs a language model for it -- and it hands an investigator the same dealer at the top of the list every Monday, because a claim that was countable last week is still countable this week. The question a warranty audit programme actually has to answer is the next one: of everything in this book, what was NOT countable the last time anybody looked, and is this dealer escalating or closing out? Three things make that unanswerable from the page in front of you. A claim can become countable with nothing about it having changed, because a clock crossed. A dealer's level depends on how many consecutive runs it has already been at REVIEW. And a finding an investigator declined last week may or may not have taken its claims out of the population, depending on a sentence somebody typed. Somebody working a warranty audit queue by hand every Monday: pulling each dealer's open claim book, working out which of the duplicates, repeat repairs, inflated labour lines and out-of-population recall claims were already raised weeks ago, reading the investigator's notes to see which findings were declined and whether the reason was a real one, and only then deciding whether this dealer is getting worse or quietly closing out.
Audience
Warranty audit and dealer-network compliance at a vehicle manufacturer or importer, the investigators who work the queue this produces, and the audit manager who owns any money movement -- plus anyone deciding whether a language model has any business near a queue that names dealers. This kit's own answer, on this corpus, is a heavily qualified yes: the model is better, by +0.63 points and one reading of 160, and free code with the same memory does 98.12 pct of the job for nothing. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual dealer readings
The corpus is 160 dealer readings, 1.00 MB (json 2 · jsonl 2 · txt 160). A WARRANTY FRAUD QUEUE IS A LIST OF NAMED DEALERS WITH AN ALLEGATION BESIDE EACH ONE. There is no public corpus of that and there never will be, for the same reason there is no public corpus of anybody's internal audit file: publishing it IS the disclosure, and a dealer principal is an identifiable person against whom nothing has yet been proven. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/programme.replay's output over the RENDERED pages, so a set-and-interval rule this fiddly cannot carry its author's misreading into the score -- and evals/check_labels.py re-derives all 160 rows from the shipped files before any run may spend. It also bought the nine storylines. A captured book contains whatever it contains; this one deliberately separates a dealer whose problem is growing from one whose identical-looking pile is three weeks old and closing, because those two are indistinguishable on a single page and are the entire reason this kit carries state. ⚠︎ AND IT SHIPS BLOCKED PENDING ANCHOR. The 30-day duplicate span, the 180-day repeat span, the 1.5x labour multiple, the 21-day parts clock, the 500 USD parts floor, the 30-day chargeback window, the four levels and the three-consecutive-runs rule are INVENTED for this kit. They reproduce no manufacturer's warranty policy, no dealer agreement, no regulator's guidance and no industry code, and name none. The catalogue row this kit was built from records the guardrail as BLOCKED-PENDING-ANCHOR.
The corpus
The 160 dealer readingsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your dealer readings. That is the whole change — there is no database to migrate.
One dealer reading, as the model receives itDLR-0001-W1.txt · 1 of 160
Synthetic Record
------------------------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes no real
dealer, person, vehicle, warranty claim, campaign or manufacturer. Dealer DLR-0001, run W1 (2026-07-06).
Dealer
------------------------------------------------------------------------------
Dealer code : DLR-0001
Region : North
Franchise : volume marque
Claim volume : small (under 400 claims a year)
Enrolled : 2018-01-01
Run date : 2026-07-06
Previous run : none
Claims in book : 4 (every open claim paid in the last 90 days, re-read whole this run)
Audit Programme Rules
------------------------------------------------------------------------------
Countable claims. A claim in the book below is COUNTABLE against one of five patterns:
DUPLICATE_CLAIM two or more paid claims on the same VIN for the same failure code
within 30 days. The later claim of each pair is countable.
REPEAT_REPAIR 3 or more paid claims on the same VIN for the same component
group within 180 days. The third and each later claim is countable.
LABOUR_TIME_INFLATION labour hours claimed above 1.5 times the published time.
PARTS_NOT_RETURNED a claim above 500.00 USD recording parts returned N, more than
21 days after payment. It becomes countable on the run that
crosses that clock, not on the run the claim first appeared.
OUT_OF_POPULATION_RECALL a claim citing a campaign whose VIN sequence falls outside the
campaign's affected range, or whose build date falls outside the
Abridged — the file continues.
The outcomeWhat a good result looks like
Per dealer per run: the level (NONE / WATCH / REVIEW / INVESTIGATE), the pattern attributed to it, how many claims became countable since the previous run, and the movement against the level last reported -- plus a four-field carried position the next run is judged against, written by code from the parsed claim book and never from the model's reply.
And when it cannot
The two directions cost different things and are never averaged. A dealer that earned REVIEW or INVESTIGATE and is reported quiet leaves claims unworked while the chargeback window runs out -- and on this corpus the window is 30 days, so the run that misses it may be the last one that could have acted. A quiet dealer put on the queue is a named dealer asked to explain claims that were fine. The model, both free floors and the calibration all sit at 0 missed raises of 57; the ONLY arm that ever misses one is the stateless control, at 1 -- the arm with no memory, which is the point. The one-reading floor buys its own zero by raising 44 false alarms in 103 quiet readings, against the model's 2 and the carry-aware floor's 3.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Running a weekly duplicate-and-pattern watch over a dealer network at all — the free carry-aware rule floor first, at $0.00 -- then add the model only if you have measured the margin on your own book 98.12 pct on the level, 100.0 pct on the pattern, 96.88 pct on the count, 98.12 pct on the movement, 0 missed raises and 3 false alarms in 103 quiet readings -- with no model, no key and no network. It is evals/baseline.py and you can run it in a fresh clone right now. The model reaches 98.75 pct on the level for 0.014193 USD a reading; that margin is +0.63 points.
Deciding whether an investigator's declination note removes claims from the count — the model, on that subset only This is the one clause a rule cannot express, and it is the only place the model earns its price on this kit. It reads 22 of the 24 prose-dependent readings against the free carry-aware floor's 21 -- 91.67 pct against 87.5. The keyword table classifies 12 of 16 wordings and the four it misses are substantive reasons written without the obvious word. EVERY error in both best arms lives here.
Choosing how often to run it — weekly, on the shipped constants -- and re-derive it for your own evals/missed_run.py settles it arithmetically and for free. The parts clock is 21 days and the chargeback window 30, so a countable claim has 9 days of headroom; a weekly run is inside it and a fortnightly one is not. One missed run costs 6 claims worth 9,779.29 USD on this corpus, and changes 32 of 40 run-4 answers.
Standing a monitor up for the first time on a book nobody has watched — a person, working the first run's output as a backlog rather than as a queue A monitor's first run has no previous reading, so every standing claim reads as new. On this corpus run 1 surfaces 20 claims worth 30,380.79 USD that were ALREADY past the chargeback window on the day they were found. That is not a false positive and it is not a model failure -- it is a backlog, and treating it as a weekly queue would put a network's worth of unactionable findings in front of an investigator on day one.
Debiting a dealer, raising a chargeback or opening an audit — the audit manager, on evidence an investigator has worked Nothing in this kit does any of those and nothing should be added that does. The cap on the catalogue row is money-movement, hard and non-configurable, and evals/check_labels.py asserts the absence of ten money-moving call names before any run may spend.
And where nothing here is good enough:
Running the watch with no memory between runs -- a weekly query, joined to nothing — nothing here. Do not do it The free one-reading floor is exactly that, and it is what most programmes actually run. Level 51.88 pct, 44 false alarms in 103 quiet readings, a structural 0.0 pct on the counts that need memory, and UNCHANGED as the answer to 'what moved' on every reading of every week. It surfaces every real raise and buries them in 44 spurious ones.
A dealer network with years of weekly history, or thousands of dealers — nothing here yet -- measure it first The carried position is four scalars, so the COST per run does not grow with history. What is unmeasured is accuracy over a long chain (every rule in this corpus resolves inside four runs), and what does grow is the PROMPT, because the population is re-read whole -- the claim book is already 24 pct of this prompt's characters at 536 claims.
At a glanceHow the whole thing runs
99%level accuracy pct
26,085 msp50, end to end
$14.19per 1,000 dealer readings · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a dealer's warranty claims turning into a pattern14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/programme.py FIRST, and not as an option. Every figure here is a property of a four-run corpus of 40 invented dealers, measured once.Corpus lens →
When is this the wrong choice?
Avoid: Paying a model to do set arithmetic over a printed table that a dict comprehension does exactly and for nothing -- and equally, dismissing the model because the margin is small. +0.63 points is 1 reading of 160 here; on a book where one missed dealer is expensive that may be worth it. It is a decision, not a default, and this page publishes the number rather than the conclusion. That is the case against the best-fitting scenario (“Running a weekly duplicate-and-pattern watch over a dealer network at all”). 7 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Four weekly runs per dealer. Long enough to hide an escalation behind the consecutive-runs rule and to let a parts clock cross mid-series; not long enough to test a position that has been running for two years, a dealer that escalates and clears twice, or a claim that ages out of the 90-day book while a finding on it is still open. 7 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
One run per arm, each fired once. Nothing here is a distribution: whether 98.75 pct level accuracy repeats on a second identical run is unmeasured, and the small denominators (24 prose-dependent readings, 13 INVESTIGATE readings, 2 errors) would make a repeat noisy. 12 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 3 models on the free one-reading rule floor and the free carry-aware rule floor and the fast tier, with the carried position, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-warranty-dupe-watch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes all eleven pre-flight assertions (python3 -m evals.check_labels), writes the missed-run key and prices the cadence gap (python3 -m evals.missed_run), scores BOTH free floors (python3 -m evals.run --run-id b000-... --baseline, and again with --carry-floor), proves the wiring end to end (--stub) and serves the whole UI (python3 -m src.app). Six commands, plain python3, no install step -- requirements.txt names nothing because the kit imports nothing outside the standard library. ⚑ ON THIS KIT THE FREE PATH IS NOT A CONSOLATION PRIZE: it reproduces every published quality figure, because every published quality figure was measured without a provider.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
26,085 msp50, end to end
92,596 msp95
2 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed to a level, a pattern, a count and a movement. Parsing the claim book, the campaigns and the dispositions, and advancing the carried position, all happen outside this clock and cost no network at all. Measured over all 160 readings of r001-warranty-dupe-watch at the published MAX_TOKENS = 24,000. ⚠︎ THE TAIL IS THE STORY AND IT IS NOT DRIVEN BY PAGE LENGTH: p95 (92,596 ms) is 3.5 times p50 (26,085 ms), while every reading is within a narrow band of every other in size. What drives a long reply is how much set arithmetic the run needs -- a dealer with two patterns and a declination to read is several times the work of a dealer with an empty book. ⚠︎ AND THESE FIGURES WERE MEASURED WHILE THREE OTHER KITS WERE RUNNING AGAINST THE SAME PROVIDER KEY, so they carry queueing that a solo run would not; they are an honest upper bound on this configuration rather than a clean measurement of it.
Current processWhat it replaces
Somebody working a warranty audit queue by hand every Monday: pulling each dealer's open claim book, working out which of the duplicates, repeat repairs, inflated labour lines and out-of-population recall claims were already raised weeks ago, reading the investigator's notes to see which findings were declined and whether the reason was a real one, and only then deciding whether this dealer is getting worse or quietly closing out.
Where it is not good enough
⚑ THE MODEL WINS, AND IT WINS BY ONE READING. On the level it scores 98.75 pct against the free carry-aware rule floor's 98.12 -- 158 correct readings against 157, a margin of +0.63 points. On all four fields at once it is 98.75 against 96.88, which is 158 readings against 155. The pattern field is a dead heat at 100.0 pct on both. That is what 0.014193 USD a reading buys over a rule engine that costs nothing, and this page leads with it rather than with the accuracy. ⚑ AND THE THING THAT ACTUALLY MOVED THE NUMBER IS FREE. Handing that same rule engine the four scalars of carried state -- and changing nothing else -- moves the level from 51.88 pct to 98.12, the new-finding count from 41.88 to 96.88 and false raises from 44 to 3, at $0.00 on both sides. The carried state is worth 46.24 points; the model is worth 0.63 more. A reader deciding what to build here should read those two sentences in that order. ⚠︎ BOTH OF THE MODEL'S REMAINING ERRORS ARE THE SAME SENTENCE THE FLOOR ALSO FAILS. DLR-0033-W4 and DLR-0036-W4 each carry a DECLINED finding whose reason is substantive, written without any marker a keyword table looks for; both arms keep counting two claims that rule 6 removed, and report REVIEW / 2 new / UNCHANGED where the key says WATCH / 0 / EASING. The model reads 22 of the 24 prose-dependent readings against the floor's 21 -- it is better at this and it is not solved. Every error in the best two arms of this kit is one clause of one rule. ⚑⚑ AND ONE UNPLANNED REPEAT SAYS THE MARGIN IS NOT ESTABLISHED AS STABLE. A single live call fired for this kit's screenshot, on DLR-0035-W2, answered REVIEW / 3 / WORSENING where the scored run answered the key's WATCH / 1 / UNCHANGED -- on a prompt built by the identical code path -- and classified the same declination note as administrative where the scored run called it substantive. That is the prose subset, which is where the entire margin lives. One repeat is evidence and not a measurement; it does not overturn the scored run, and it is enough to say the +0.63 points are not shown to be reliable. The frame is published as it came back and the next experiment worth paying for is a repeat of the 24 prose readings, not a bigger model. ⚠︎ THE ONE-READING FLOOR IS THE OTHER LESSON AND IT IS WORSE THAN IT LOOKS. Without memory the same rule engine scores 51.88 pct on the level and raises 44 false alarms in 103 quiet readings -- and it MISSES NOTHING, because it raises nearly everything. On a dealer with a large standing pile and nothing new it reports INVESTIGATE every week for ever. Its memory-dependent count accuracy is 0.0 pct, a structural zero. ⚠︎ AN EARLIER ATTEMPT AT THIS RUN WAS ABANDONED AND ITS SPEND IS RECORDED RATHER THAN HIDDEN. The provider account emptied mid-pass and returned 402 Payment Required after 39 of 160 readings. That partial file is preserved unaltered at kits/UC0087-warranty-dupe-watch/docs/incomplete/ with a note explaining why its numbers are an arithmetic consequence of 121 absent replies (39 of 160 is 24.375 pct, and every 'accuracy' in it is that same number wearing a different label). NOTHING on this page is drawn from it, the scored run below was written to its own file, and the run register refuses the abandoned one by itself -- runlog's triage extractor declines any record whose answered count is not a majority of its windows. ⚠︎ A SECOND INCIDENT COST MORE AND IS RECORDED TOO: when the account was refunded, a launch that appeared to have died had not, so two copies of the same run raced each other against one result file and double-billed every call until one was killed. src/budget.py counts calls and cannot see that two runs of one run-id are in flight -- its own docstring names that hazard and does not solve it.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt160jsonl2
40 franchise dealers, 4 consecutive weekly runs each
Recorded failureall 5 of the carry-aware floor's errors are one shape — a DECLINED note whose reason its keyword table could not classify (12 of 16 wordings). It is 0.63 points behind the model and free
Recorded failure⚠︎ one unplanned repeat of DLR-0035-W2 came back WRONG on a byte-identical prompt, landing on the free floor's answer — the margin is not shown to be stable
4 fields exact, then 6 slices, free — no judge model
missed raises and false raises reported apart, never averaged
Recorded failure0 missed raises of 57 on the model and both free arms — but the STATELESS control misses 1, which is where losing the memory finally shows up as a miss rather than as noise
model + 4 scalars: level 98.75% · 158 of 160 fully correct
free code + the SAME 4 scalars: 98.12% · 155 of 160 · $0.00
46.24the carried state is worth points and costs nothing
+0.63the model is worth more — one reading — and one repeat flipped
2026-08-23as of
It produces a level, a pattern, a count of what is new and a movement for a warranty investigator to work, and debits, charges back and opens nothing — the cap on the catalogue row is money-movement, hard and non-configurable, and evals/check_labels.py asserts the absence of ten money-moving call names before any run may spend. THE STATION THAT MATTERS IS THE SECOND ONE. Take those four scalars away and the identical rule engine reports INVESTIGATE with 5 new findings on the same dealer every week for ever, because a claim it counted last Monday looks exactly like one filed on Friday; level accuracy falls from 98.12 pct to 51.88 and false raises rise from 3 to 44, with every answer still well-formed and nothing raising an error. ⚑ AND THE THIRD STATION IS THE ONE SIBLING MONITORS LEAVE OUT. Four cadence kits shipped the day before this one and none of them ran on a clock. This one does, and the schedule is measured rather than asserted: a claim has nine days between becoming countable and leaving the chargeback window, so a weekly run catches it and a fortnightly one does not — 6 claims worth $9,779.29 lost to a single skipped Monday.
⚠︎ A ⚑ AND THE CADENCE ARM ANSWERS THE GAP: 97.50 pct on run 4 with run 3 skipped, scored against its own answer key. The kit copes; the programme still loses the claims.
⚠︎ A MONITOR'S FIRST RUN IS ITS WORST RUN AND NO CADENCE FIXES IT: with no previous reading every standing claim reads as new, and run 1 surfaces 20 claims worth $30,380.79 that were already past the chargeback window on the day they were found. That is a backlog, not a queue.
⚠︎ EVERY THRESHOLD IN THIS PROGRAMME IS INVENTED — the catalogue row reads BLOCKED-PENDING-ANCHOR — and the full rule text is reprinted on all 160 readings so it can be read, disbelieved and replaced.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the carried position
src/state.py
load()/save() are a JSON file today. Point them at a table, a key-value store or the audit programme's own database and nothing else in the kit changes -- for_dealer() is the whole read interface and describe() is the only thing the prompt sees.
the audit programme
src/programme.py
Every threshold, the level table, the consecutive-runs rule and the chargeback window are your programme's, not ours. Change them there and re-run tools/build_corpus.py, which recomputes the whole answer key from the same functions -- the key cannot drift from the rules because it is their output. ⚠︎ THIS IS THE FIRST FILE TO REPLACE, not an optional one: everything shipped is invented.
the corpus
tools/build_corpus.py
Point it at your own dealers and claim extracts, or delete it and drop real readings into data/corpus/ named <DEALER>-W<n>.txt with the same seven section headings. Everything downstream reads by dealer and run and does not care where the readings came from.
the cadence
evals/run.py
SCHEDULES maps a name to the sequence of runs the programme actually performs. 'weekly' and 'missed-w3' ship; add your own and evals/missed_run.py prices it against the weekly one for nothing. ⚠︎ This is the seam most kits in this series do not have, and it is where a monitor's real behaviour lives.
what is sent
src/select.py
SECTION_HINTS maps a field to the sections it needs; NEVER_SENT names what is withheld whatever happens. ⚠︎ On a dealer network this is the load-bearing seam: check NEVER_SENT covers every section your readings carry that names a person, before anything spends.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 40 dealers x 4 weekly runs = 160 readings and 536 claims from a fixed seed (SEED = 20260823), across nine deliberately separated storylines. It writes each reading, reads it BACK through the runtime's own parser, and computes the answer key from that -- so the page and the key cannot disagree about what is on the page.
the audit programme
src/programme.py
The five countable-claim tests, the level table, the consecutive-runs rule and the declination clause, as pure code -- plus assess() and replay(), which ARE the answer key and which the runtime never calls. ⚠︎ Every threshold in it is INVENTED and the kit ships that as a stated blocker; the full rule text is reproduced on all 160 readings so it can be read, disbelieved and replaced.
the carried position
src/state.py
SEAM 2 -- the thing that makes this a monitor. Four scalars per dealer (the level last reported, consecutive runs at REVIEW or above, the pattern attributed, and the claim number the last book ended at), written from the arithmetic and never from the model's reply, and rendered into English for the prompt. It costs the same on run 40 as on run 3.
the section splitter
src/segment.py
Splits a reading into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 160 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Dealer Contact -- the principal's name, trading address, telephone and email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-position sentence, and the selected sections in document order. The stateless build replaces the position block with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading. ⚠︎ It is also what refused this kit's scored run: a 402 Insufficient Balance is terminal, not transient, so it is raised rather than retried.
the watch
src/watch.py
One dealer, one run, one call. Parses the claim book, the campaigns and the dispositions off the page with regexes (the model is never asked to read a number), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
the local UI
src/app.py
One dealer, one run, its carried position and four columns, on 127.0.0.1:9001. Renders with no key. It shows the carried sentence verbatim and BOTH free floors beside the model's answer and the key, so a reader can see on which rows a model bought nothing.
the free one-reading floor
evals/baseline.py
Everything a rule engine reading THIS week's book can compute, with no idea what last week said. 0 calls, $0.00, scored through the identical scorer.
the free carry-aware floor
evals/baseline.py
The same rule engine given the same four scalars the model is given. It implements every printed rule except one -- whether a DECLINED note is substantive or administrative -- for which it uses an ordered keyword table. 0 calls, $0.00. ⚑ THIS IS THE HONEST COMPARATOR AND CURRENTLY THE BEST ARM THIS KIT HAS.
the scorer
evals/scoring.py
Exact match per cell against the computed key, split ways an average would hide: the four fields, the raise directions counted apart, the memory-dependent subset, the prose-dependent subset, the first run against the later ones, and the recoverable exposure in dollars. No judge model.
the cadence measure
evals/missed_run.py
Writes the SECOND answer key -- what the correct answers become when one weekly run never happens -- and prices the gap in claims and dollars. Free, no calls. ⚑ This is the half of a monitor that usually goes unmeasured.
the pre-flight
evals/check_labels.py
Eleven things that must be true before a run may spend: the corpus shape, the seven sections, claim numbers in payment order, every declination note in exactly one pool, the key re-derived from the rendered pages, the memory-dependent flag verified in both directions, the two prompt builds byte-identical outside one block, the privacy guard red-proven in both directions, no money-moving code path, both floors answering every reading, and the missed-run key reproducing.
the run harness
evals/run.py
40 dealer chains, four strictly-ordered runs each, 12 concurrent workers. The position advances even for a run whose CALL failed, so one transport error cannot turn into four scored failures -- which is exactly the branch the 402 exercised.
Where it breaks at scale
NOT ON HISTORY LENGTH, AND THAT IS THE DESIGN. The carried position is four scalars, so run 40 costs exactly what run 3 costs -- unlike a conversation kit, whose input grows with every turn. What it breaks on is five other things. FIRST, THE CHAIN IS SERIAL. Runs within one dealer cannot be parallelised, because run 3's prompt contains a position produced by run 2. This corpus went 40 chains wide; a dealer with two years of weekly history would be 104 serial calls and no width would help. ⚑ EXCEPT FOR RE-FIRING ONE RUN, which IS cheap here precisely because the position is arithmetic rather than model output -- --runs W4 reproduces the chain for free and spends only on the run being measured. SECOND, THE STATE STORE IS A FILE. src/state.py writes data/state.json atomically, which is correct for one process and is not a concurrency model; two schedulers advancing one dealer would race, and the loser's run would vanish -- and because a lost run resets nothing but ADVANCES nothing either, the next run reads stale claim numbers and re-raises claims already counted. THIRD, THE POPULATION IS RE-READ WHOLE EVERY RUN. That is the shape of the problem, not an implementation choice, and it means the prompt grows with the dealer's OPEN BOOK -- 536 claims here at 6543 bytes a reading; a large dealer's 90-day book is several hundred, and the claim book is already 24 pct of this prompt's characters. FOURTH, NOTHING HERE SCHEDULES ANYTHING. The kit measures a cadence and does not implement one; it is invoked, not woken. FIFTH, A REAL DEALER NETWORK IS NOT 40 DEALERS. A volume manufacturer's is thousands, and at one call per dealer per week the arithmetic that matters becomes the bill rather than the wall clock: see Cost.cost_at_10x.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
DLR-0029-W3, and the finding that does NOT depend on a model at all. Five duplicate pairs, every one paid before the programme's first run, and nothing new since. The free ONE-READING floor reports INVESTIGATE with 5 new findings -- exactly what it reported last week and exactly what it will report next week. The answer key says WATCH, 0 new, EASING: a dealer on its way off the queue. The free CARRY-AWARE floor in the middle column gets it right for nothing. ⚑ THE GAP BETWEEN THE TWO FREE COLUMNS IS 46.24 POINTS OF LEVEL ACCURACY AND COSTS NOTHING; the model adds +0.63 more. This frame is taken with API_KEY blanked, so the model column is empty by construction and cannot have spent.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
⚠︎ ONE LIVE CALL ON DLR-0035-W2, AND IT DISAGREES WITH THIS KIT'S OWN SCORED RUN ON THE SAME READING. The investigator declined a finding with the note 'the sequence number falls inside the range as amended in June and the reading predates the amendment' -- substantive, so rule 6 takes two claims out of the count and the key says WATCH / 1 / UNCHANGED. The SCORED run answered exactly that, and its rationale said so: 'after the substantive declination removed CLM-100330 and CLM-100353'. THIS call, on a prompt assembled by the identical code path, answered REVIEW / 3 / WORSENING and reasoned 'the investigator's declination was administrative so the claims stay countable' -- landing on the same answer as the free keyword table it is supposed to beat. ⚑ THE MODEL'S ENTIRE +0.63-POINT MARGIN OVER FREE CODE IS EARNED ON THIS SUBSET, and a single unplanned repeat of one of its readings flipped. That is one repeat, which is evidence and not a measurement -- and it is enough to say the margin is not established as stable. Recorded in Eval.could_not_verify rather than re-shot.failureOpen full size →The same page with the button pressed and NO API_KEY configured. It does not error and it does not go blank: the reading, the carried position, both free floors, the answer key and the withheld-section list are all computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called. This is the honest failure state -- and this kit spent a day in it, when the provider account emptied mid-run and the page rendered a 402 in exactly this slot, with the key redacted out of it by src/app.py.failureOpen full size →
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
160dealer readings
1.00 MiBjson 2 · jsonl 2 · txt 160
40dealers (4 weekly runs each) · p50 4 chars
$0.00setup · 0.0s
How it is cutWhat one dealers (4 weekly runs each) is
No split, and no chunking. The unit is a DEALER -- a run of four consecutive weekly readings processed strictly in order, because run 3's prompt contains a position produced by run 2. Each reading goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step. tools/build_corpus.py writes 160 readings and two answer keys from a fixed seed with no clock read and no model called; nothing is embedded, ranked or cached.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every dealer, principal, address, VIN, claim, operation code, recall campaign, investigator finding and note is invented here. Verified against the repository's own LICENSE file on 2026-08-23.
Bring your ownBring your own dealer readings
Replace src/programme.py FIRST, and not as an option. It is your programme's thresholds, level table and chargeback window, the shipped ones are invented, and data/gold.jsonl is literally its output -- change the rules and re-run tools/build_corpus.py and the answer key follows. Then point the corpus generator at your own dealers, or delete it and drop real readings into data/corpus/ named <DEALER>-W<n>.txt with the same seven section headings. src/segment.py, src/select.py, src/prompt.py and src/state.py read sections by name and do not care where the readings came from. Then set SCHEDULES in evals/run.py to the cadence you actually run, and evals/missed_run.py will price a missed one for nothing.
⚠︎ And what stops being true when you do: Every figure here is a property of a four-run corpus of 40 invented dealers, measured once. Three denominators are small: 24 prose-dependent readings, 13 INVESTIGATE readings, and 2 total errors in the best arm -- two rows is not a rate. ⚠︎ AND THE MARGIN THAT DECIDES WHETHER TO PAY FOR A MODEL IS ONE READING. 98.75 pct against 98.12 is 158 correct readings against 157, and both arms' errors sit in one storyline of nine. A book with a different mix of investigator wording moves that margin without either system having changed, and it could move it either way. The thing most likely not to survive contact with a real programme is the declination taxonomy -- the field every arm is worst on, and the one the catalogue row already records as unconfirmed.
What breaks it
⚠︎ A SILENT REGEX FAILURE ATE THE ONLY DECLINATION THAT MATTERED -- FOUND BY THE CORPUS BUILDER, FIXED, AND NOW GATED. segment.split strips the trailing newline off every section body, so the disposition regex -- which required a newline after the note line -- parsed every finding's note EXCEPT the last one in the section. The last one is the most recently raised, which is precisely the one whose claims are new this run and therefore the only one that can change an answer. The symptom was that the DECLINED_SUBSTANTIVE and DECLINED_ADMIN groups -- the pair the corpus exists to separate -- produced identical gold. Nothing errored. It was caught only because the key is derived from the rendered page and the two groups were then compared; a key built from the in-memory structures would have agreed with itself perfectly.
Four weekly runs per dealer. Long enough to hide an escalation behind the consecutive-runs rule and to let a parts clock cross mid-series; not long enough to test a position that has been running for two years, a dealer that escalates and clears twice, or a claim that ages out of the 90-day book while a finding on it is still open.
The declination notes are drawn from two pools of eight. A real investigator writes freely, and the substantive/administrative distinction in the wild will be blurrier than sixteen sentences written by one author on one afternoon. ⚠︎ AND THE SAME AUTHOR WROTE THE KEYWORD TABLE THAT SCORES THEM, which is a conflict named out loud in evals/baseline.py: the markers are generic audit vocabulary chosen before the notes were scored, written once and not tuned afterwards, and they land on 12 of 16. A table fitted to those sixteen would have got all sixteen and proven nothing.
The parts clock is the only ageing-in mechanism. Every other pattern becomes countable the moment its claim appears, so the corpus exercises 'the population changed under a fixed rule' on one pattern of five. A programme whose duplicate span or repeat window also aged claims in would have far more memory-dependent readings than this one's 93 of 160.
One claim book shape and one rule text, reproduced identically on every reading. A real programme's rules differ by region and franchise, carry exceptions, and get amended mid-series -- and an amendment that lands between two runs is exactly the case this corpus has none of.
Free-text analyst prose. The Programme Notes section is this kit's injection surface and it is deliberately sent, but on this corpus it is one of five fixed sentences chosen by a seeded generator. A programme whose notes are written freely by analysts is the real surface and this corpus does not exercise it.
A dealer that leaves the network, or a claim book that empties. Every dealer here has a book on every run. src/programme.assess carries last_claim_seq forward when a book is empty, and nothing exercises it.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
144
not measured
instruction
2,552
not measured
carried position
265
not measured
Synthetic Record
290
not measured
Dealer
391
not measured
Audit Programme Rules
3,227
not measured
Claim Book
2,435
not measured
Investigator Disposition
472
not measured
Programme Notes
169
not measured
Total
2,560
This is the cost lesson as arithmetic: of the 9,945 characters assembled, 3,757 are datas — 38% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.py::build on DLR-0033-W2 with the exact carried position r001-warranty-dupe-watch recorded for that call ({"prev_level": "WATCH", "runs_at_review_or_above": 0, "prev_pattern": "DUPLICATE_CLAIM", "last_claim_seq": 100270}) -- the identical code path the run used, not a paraphrase. THE SPLIT IS IN CHARACTERS AND THE PAGE SAYS SO: per-part TOKEN counts were not measured, because measuring them means sending nested prefixes of the prompt to a provider and this kit spent no calls on it. tokens.input and tokens.output are the MEASURED averages over the scored run's 160 calls (409,625 in / 688,689 out). The stateless control's build of the same reading differs only inside the carried-position block, and evals/check_labels.py asserts every other byte is identical -- that is the entire experiment, and both arms of it are now fired.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You apply a written warranty audit programme to one dealer's claim book on one scheduled run. You answer with one JSON object and no other text.
You are reading one franchise dealer's warranty claim book on one scheduled run of an audit
programme, and deciding what the programme should report about that dealer THIS RUN.
The programme's rules, its thresholds and its chargeback window are reproduced in the reading below.
Apply them exactly as written. The rules depend on what the PREVIOUS run of this programme recorded
about this dealer, which you cannot see; what is known is stated under "Carried state" and is the
only history available to you. Do not assume anything about earlier runs beyond it.
How to decide:
- Work out which claims in the Claim Book are COUNTABLE under each of the five patterns, using the
thresholds in the rules.
- Read the Investigator Disposition. A DECLINED finding removes the claims it names from the
countable population only when the investigator's recorded reason is substantive -- the claims
are genuinely separate work, correctly authorised, or already accepted by the manufacturer. A
declination for an administrative reason -- the evidence was insufficient, nobody could be
reached, the file was closed unworked -- leaves those claims countable. OPEN and QUEUED change
nothing.
- Of what is left, work out which claims are NEW this run: claim numbers above the number the
carried state gives, plus any claim that crossed the parts clock between the previous run date
and this one.
- Take the pattern with the most new countable claims, apply the level table, then the consecutive
runs rule, then decide the movement against the level last reported.
Answer with a single JSON object and nothing else:
{"level": "NONE|WATCH|REVIEW|INVESTIGATE",
"pattern": "DUPLICATE_CLAIM|REPEAT_REPAIR|LABOUR_TIME_INFLATION|PARTS_NOT_RETURNED|OUT_OF_POPULATION_RECALL|NONE_DETECTED",
"new_findings": <whole number of NEW countable claims attributed to that pattern this run>,
"movement": "NEW|WORSENING|UNCHANGED|EASING|CLEARED",
"rationale": "one sentence, naming the claims you counted and the rule you applied"}
"new_findings" counts only claims that were not already countable at the previous run; a claim that
was countable last run and is countable now has already been counted and is not new. "movement" is
this run's level compared with the level the carried state reports, and is CLEARED only when this
run's level is NONE and the previous one was not.
This programme surfaces a pattern for an investigator to work. Do not recommend a debit, a
chargeback or an audit -- none of those is yours or this programme's to make.
Carried state
------------------------------------------------------------------------------
At the previous run on 2026-07-06 this dealer was reported WATCH for DUPLICATE_CLAIM. It has not been at REVIEW or above on any run up to and including that one. The claim book then ended at CLM-100270; every claim numbered above that is new since the last reading.
Dealer reading
------------------------------------------------------------------------------
Synthetic Record
------------------------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes no real
dealer, person, vehicle, warranty claim, campaign or manufacturer. Dealer DLR-0033, run W2 (2026-07-13).
Dealer
------------------------------------------------------------------------------
Dealer code : DLR-0033
Region : South West
Franchise : volume marque
Claim volume : large (over 1500 claims a year)
Enrolled : 2022-09-06
Run date : 2026-07-13
Previous run : 2026-07-06
Claims in book : 14 (every open claim paid in the last 90 days, re-read whole this run)
Audit Programme Rules
------------------------------------------------------------------------------
Countable claims. A claim in the book below is COUNTABLE against one of five patterns:
DUPLICATE_CLAIM two or more paid claims on the same VIN for the same failure code
within 30 days. The later claim of each pair is countable.
REPEAT_REPAIR 3 or more paid claims on the same VIN for the same component
group within 180 days. The third and each later claim is countable.
LABOUR_TIME_INFLATION labour hours claimed above 1.5 times the published time.
PARTS_NOT_RETURNED a claim above 500.00 USD recording parts returned N, more than
21 days after payment. It becomes countable on the run that
crosses that clock, not on the run the claim first appeared.
OUT_OF_POPULATION_RECALL a claim citing a campaign whose VIN sequence falls outside the
campaign's affected range, or whose build date falls outside the
campaign's build window. Both tests must pass for the claim to be in
population.
New this run. A countable claim is NEW when its claim number is above the number the previous
reading's book ended at -- stated under Carried state -- or when it crossed the parts clock since
the previous run date. Nothing else is new; a claim that was countable last run and is countable
now has already been counted.
Level. From the number of NEW countable claims attributed to the pattern:
4 or more INVESTIGATE
2 or more REVIEW
1 WATCH
0 the level EASES one step from the level last reported
(INVESTIGATE -> REVIEW -> WATCH -> NONE).
A dealer at REVIEW or above for 3 consecutive runs including this one is INVESTIGATE
whatever the count.
Pattern. The pattern with the most new countable claims. On a tie, the earlier of the two in the
list above. When there are no new countable claims and the level is still above NONE, the pattern
last reported stands; at NONE the pattern is NONE_DETECTED.
Disposition (rule 6). A finding the investigator has DECLINED removes the claims it names from the
countable population ONLY when the recorded reason is SUBSTANTIVE -- the claims are genuinely
separate work, correctly authorised, or already accepted by the manufacturer. A declination for an
ADMINISTRATIVE reason -- the evidence was insufficient, nobody could be reached, the file was closed
unworked -- leaves the claims countable, because "we could not prove it" is not "it did not happen".
Findings marked OPEN or QUEUED change nothing.
Chargeback window. A claim cannot be charged back to the dealer more than 30 days after it was
paid. A finding raised after that date is a finding nobody can act on.
This programme surfaces patterns for an investigator. It debits nothing, charges back nothing and
opens no audit by itself.
Recall campaigns in force and their affected populations:
RCL-2451 affected VIN sequence 010000-019999 built 2025-03-01 to 2025-08-31
RCL-2478 affected VIN sequence 040000-044999 built 2025-09-01 to 2026-01-31
Claim Book
------------------------------------------------------------------------------
Every open claim on this dealer, as at the run date. This is the whole population, re-read;
it is not a list of what changed.
CLM-100098 VIN 2FADP82633A500272 op 64-1150 fail F22 comp BRK labour 1.0h (pub 1.3h) parts Y built 2025-05-14 paid 2026-05-25 campaign - 340.29 USD
CLM-100133 VIN SHGCM70298C500276 op 51-2007 fail F58 comp ENG labour 1.3h (pub 1.7h) parts Y built 2025-05-14 paid 2026-06-08 campaign - 187.80 USD
CLM-100201 VIN 2FADP36641D500277 op 64-1150 fail F58 comp BDY labour 3.6h (pub 4.8h) parts Y built 2025-05-14 paid 2026-06-26 campaign - 1386.88 USD
CLM-100223 VIN 1WBAV36641D500279 op 64-1150 fail F80 comp BDY labour 1.0h (pub 1.1h) parts Y built 2025-05-14 paid 2026-06-28 campaign - 1178.60 USD
CLM-100234 VIN 7HGCM41155B500278 op 64-1150 fail F11 comp HVA labour 2.3h (pub 2.5h) parts Y built 2025-05-14 paid 2026-06-29 campaign - 1223.18 USD
CLM-100247 VIN 1WBAV36641D500279 op 42-1180 fail F80 comp BDY labour 3.4h (pub 3.3h) parts Y built 2025-05-14 paid 2026-07-01 campaign - 994.99 USD
CLM-100270 VIN WWBAV36641D500273 op 42-1180 fail F58 comp HVA labour 3.2h (pub 3.9h) parts Y built 2025-05-14 paid 2026-07-04 campaign - 325.14 USD
CLM-100288 VIN 8VWZZ70298C500280 op 13-0455 fail F17 comp HVA labour 3.0h (pub 2.3h) parts Y built 2025-05-14 paid 2026-07-08 campaign - 1579.07 USD
CLM-100311 VIN 9WBAV70298C500281 op 33-4821 fail F44 comp BDY labour 1.4h (pub 1.2h) parts Y built 2025-05-14 paid 2026-07-09 campaign - 1503.27 USD
CLM-100326 VIN 8VWZZ70298C500280 op 33-4821 fail F17 comp HVA labour 1.1h (pub 1.4h) parts Y built 2025-05-14 paid 2026-07-10 campaign - 1480.38 USD
CLM-100327 VIN 5VWZZ36641D500282 op 27-3310 fail F17 comp SUS labour 4.5h (pub 4.9h) parts Y built 2025-05-14 paid 2026-07-10 campaign - 2013.57 USD
CLM-100348 VIN WWBAV70298C500275 op 27-3310 fail F80 comp BDY labour 1.1h (pub 1.0h) parts Y built 2025-05-14 paid 2026-07-11 campaign - 240.35 USD
CLM-100349 VIN 9WBAV70298C500281 op 27-3310 fail F44 comp BDY labour 4.1h (pub 4.5h) parts Y built 2025-05-14 paid 2026-07-11 campaign - 848.34 USD
CLM-100363 VIN 5VWZZ36641D500282 op 27-3310 fail F17 comp SUS labour 4.3h (pub 4.4h) parts Y built 2025-05-14 paid 2026-07-12 campaign - 916.70 USD
Investigator Disposition
------------------------------------------------------------------------------
Findings raised on this dealer since the programme last read its book, and what the
investigator did with them. A finding names the claims it was raised on.
FND-2006 raised 2026-07-09 DUPLICATE_CLAIM status DECLINED
claims CLM-100326, CLM-100349
note "the second visit was a goodwill repair authorised separately; the two are correctly distinct claims"
Programme Notes
------------------------------------------------------------------------------
The programme extract ran clean; no records were rejected for this dealer.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"level": "WATCH", "pattern": "DUPLICATE_CLAIM", "new_findings": 1, "movement": "UNCHANGED", "rationale": "Counted CLM-100363 as the only new countable claim (later duplicate on VIN 5VWZZ36641D500282 for F17 within 30 days; CLM-100326 and CLM-100349 were removed by the substantive declination), so one new finding yields WATCH, unchanged from the prior WATCH."}
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a dealer's warranty claims turning into a pattern — 160 dealer readings. Three tiers of one model family answered, and every answer was then graded Five different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
There is no LLM-as-judge in this kit and that is not a shortcut. Three of the four answered fields are closed sets -- four levels, six patterns, five movements -- and the fourth is a whole number of claims. A judge is for answers whose correctness is a matter of reading; these are matters of equality and set arithmetic, and asking a model to grade 'REVIEW == REVIEW' would add cost, variance and a second thing to be wrong. evals/scoring.py compares each cell exactly against a key src/programme.replay computed, and scores both free floors and the stub through the same function. The rationale field is not scored by anything, and this page says so where it publishes a reply.
160dealer readings
160source documents
3model tiers
480graded answers
5grading methods
MeasurementsWhat was measured
COUNTED158 · 157 · 91 · 83 · 39 / 160level accuracy pct — readings, THE PUBLISHED CONFIGURATION -- one call per dealer per runDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED160 · 160 · 146 · 145 / 160pattern accuracy pct — readings, THE PUBLISHED CONFIGURATION -- one call per dealer per runDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED158 · 155 · 78 · 67 / 160new findings accuracy pct — readings, exact whole number, THE PUBLISHED CONFIGURATION -- one call per dealer per runDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED158 · 157 · 72 · 57 / 160movement accuracy pct — readings, THE PUBLISHED CONFIGURATION -- one call per dealer per runDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED155 · 35 · 158 / 160all four correct pct — readings with every one of the four fields correct, the free carry-aware floorDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED90 · 16 · 91 · 24 / 93memory level accuracy pct — readings whose answer is NOT derivable from their own page, the free carry-aware floorDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED21 · 4 · 22 / 24prose level accuracy pct — readings carrying a DECLINED finding, the free carry-aware floor -- THE ONLY FIELD WITH HEADROOM LEFTDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED3 · 44 · 2 / 103false raise rate pct — quiet readings, the free carry-aware floor -- 3 false raisesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 · 0 / 57missed raises — readings that earned a raise, the free carry-aware floor -- THE EXPENSIVE DIRECTIONDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 160unparsed replies — calls -- every reply parsed; largest 18,053 tokens under the published 24,000-token ceilingDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The comparison cannot be wrong about itself; the risk is in the labels. Those come from src/programme.assess, which the runtime never calls -- src/watch.py asks the model for the level and never computes it -- so the key is arithmetic done independently of the code under test rather than the code under test marking its own homework. It is computed from the RENDERED pages, and evals/check_labels.py re-derives all 160 rows from the shipped files and refuses to let a run spend if any disagrees. It also verifies the memory-dependent flag in BOTH directions -- a reading is in that subset when the same page with no history produces a different answer, measured rather than labelled. ⚠︎ WHAT THE KEY CANNOT SETTLE is whether the substantive / administrative split of a declination is the split a real audit programme would draw; see could_not_verify.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One dealer reading
1,000 dealer readings
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.014193
$14.19
9%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.005677
$5.68
9%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.240817
$240.82
11%
Same work, 42× the bill
The same dealer readings, the same tokens — only the rate card changed. And across all 3 cards between 9% and 11% of what you pay is the prompt this pipeline sends, not the answer it writes.
whether a model is called at all. The free carry-aware rule floor scores 98.12 pct on the level for a measured $0.00; the model scores 98.75 for 0.014193 USD a reading. The whole budget question here is whether +0.63 points -- 1 reading in 160, and 3 readings on all four fields at once -- is worth about 14.19 USD per thousand dealer-runs on the shared card. This kit publishes the margin and does not pretend it is large.
Rates checked 2026-08-18. The provider that ran the 8 calibration calls and the 39 answered calls of the abandoned scored run is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. A measured $0.00, not an unpriced one -- and on this kit the same is true of every published QUALITY figure, because both floors are pure code. evals/missed_run.py prices a missed run for $0.00 as well.
The gradersFive ways to grade
READ THIS ROW BEFORE ANY MODEL FIGURE ON THIS PAGE, BECAUSE IT IS 0.63 POINTS BEHIND THE MODEL AND FREE. b001 is not a straw man and it is not a model in disguise: it is 167 lines of ordinary Python implementing the printed rules, given the same four scalars the model's prompt carries. It scores 98.12 pct on the level against the model's 98.75, ties it at 100.0 pct on the pattern, misses 0 of 57 raises against the model's 0, and raises 3 false alarms in 103 quiet readings against 2. ⚑ THE GAP BETWEEN THE TWO FLOORS IS THE ARGUMENT FOR CARRYING STATE, AND IT IS FOUR TIMES LARGER THAN THE GAP THE MODEL BUYS. b000 is the identical engine with the memory taken away: level 51.88 pct, 44 false raises, and a structural 0.0 pct on the counts that need memory. Handing a rule engine four scalars moves the level 46.24 points; handing the same problem to a language model moves it +0.63 more. ⚑ AND BOTH ARMS FAIL ON THE SAME CLAUSE. All 5 of b001's errors and both of the model's sit in the DECLINED_SUBSTANTIVE storyline, where a declination's reason has to be read as substantive or administrative. The keyword table lands on 12 of 16 wordings; the model reads 22 of the 24 prose-dependent readings against the floor's 21. That is the whole of what a model is buying here.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key For each of the 160 readings and each of the four answered fields, did the reply equal the computed key? Level, pattern and movement are compared exactly; the count is compared as a whole number, and a reply that is not a whole number is a MISS rather than a zero -- 0 is a meaningful answer here (it is what an easing dealer gets) and a parse failure must never be scored as one.
$0.00
no
yes
the fast tier, with the carried position 98.8% level accuracy · the fast tier, memory removed (THE CONTROL) 56.9% level accuracy · the free carry-aware rule floor, no model 98.1% level accuracy · the free one-reading rule floor, no model 51.9% level accuracy · 3 more measured on each run
the fast tier, with the carried position 100.0% detection · the fast tier, memory removed (THE CONTROL) 98.2% detection · the free carry-aware rule floor, no model 100.0% detection · the free one-reading rule floor, no model 100.0% detection · 1 more measured on each run
The readings whose answer cannot be derived from their own page On the 93 readings where the same page with no carried position produces a different answer, was the level right and was the count right? This is the subset the whole monitor argument rests on. ⚑ IT IS A MEASUREMENT, NOT A LABEL: tools/build_corpus.py flags it by computing the answer twice, and evals/check_labels.py verifies the flag in BOTH directions, so the subset can be neither padded nor quietly trimmed.
$0.00
no
yes
the fast tier, with the carried position 97.8% level accuracy · the fast tier, memory removed (THE CONTROL) 25.8% level accuracy · the free carry-aware rule floor, no model 96.8% level accuracy · the free one-reading rule floor, no model 17.2% level accuracy · 2 more measured on each run
The readings where an investigator's sentence decides the answer On the 24 readings carrying a DECLINED finding, was the level right? Rule 6 says a declination removes its claims from the countable population only when the reason is substantive, and the reason is a sentence. ⚑ THIS IS THE ONLY GRADER ON WHICH A LANGUAGE MODEL HAS ANYTHING TO SELL IN THIS KIT.
$0.00
no
yes
the fast tier, with the carried position 91.7% level accuracy · the fast tier, memory removed (THE CONTROL) 16.7% level accuracy · the free carry-aware rule floor, no model 87.5% level accuracy · the free one-reading rule floor, no model 16.7% level accuracy
What one missed run costs, in claims and in dollars For every claim the weekly schedule counts as new somewhere, was the chargeback window still open on the run that first counted it -- and is it still open when one weekly run never happens? ⚑ THIS IS THE HALF OF A MONITOR THAT USUALLY GOES UNMEASURED. Four sibling kits in this estate carry state and none of them runs on a clock.
$0.00
no
yes
the model, ANSWERING over a missed run 97.5% level accuracy · the weekly schedule, as shipped: no headline metric, 2 measurements · one missed run -- run 3 never happens: no headline metric, 3 measurements · the programme's FIRST run, whatever the schedule: no headline metric, 2 measurements · 2 more measured on each run
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
DECISIVELY IN ONE DIRECTION AND NARROWLY IN THE OTHER, AND THE PAGE LEADS WITH THE WIDE ONE. b001 against b000 is the same rule engine with and without four scalars of carried state, on the same 160 readings and the same scorer: level 98.12 pct against 51.88 (+46.24), count 96.88 against 41.88, movement 98.12 against 35.62, all-four 96.88 against 21.88, false raises 3 against 44, and the memory-dependent subset 96.77 pct against 17.2. NO MODEL IS INVOLVED ON EITHER SIDE of that comparison. ⚑ THE MODEL AGAINST THE FREE CARRY-AWARE FLOOR IS THE NARROW ONE: level 98.75 against 98.12, a margin of +0.63 points, which is 1 reading of 160. On all four fields at once it is 98.75 against 96.88 -- 3 readings. The pattern field is a dead heat at 100.0 pct. The margin is real, it is measured, and it is roughly one seventieth of the gap that carrying state opens. ⚑ AND THE CONTROL SEPARATES FURTHER STILL: the same model with the carried-position block removed scores 56.88 pct on the level and 48.75 on the count, against 98.75 and 98.75 with it. That is the experiment this kit was built for and both arms are now fired. ⚠︎ AND ONE REPEAT UNDERMINES THE NARROW ONE. A single live call on DLR-0035-W2, fired for a screenshot, flipped to the free floor's answer on the prose subset the margin depends on. It is one reading and it is published as it came back. ⚠︎ WHAT THIS LABELLED SET CANNOT SEPARATE: the model from the free carry-aware floor on the pattern field (both 100.0 pct -- saturated), or on missed raises (both 0 of 57). A grader every arm aces says the corpus stopped discriminating on that axis rather than that the axis is solved. ⚠︎ NOR CAN IT SEPARATE the two readings of a declination note in the general case; it can only price this corpus's sixteen sentences, which the author of the kit also wrote.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Running a weekly duplicate-and-pattern watch over a dealer network at all
the free carry-aware rule floor first, at $0.00 -- then add the model only if you have measured the margin on your own book
98.12 pct on the level, 100.0 pct on the pattern, 96.88 pct on the count, 98.12 pct on the movement, 0 missed raises and 3 false alarms in 103 quiet readings -- with no model, no key and no network. It is evals/baseline.py and you can run it in a fresh clone right now. The model reaches 98.75 pct on the level for 0.014193 USD a reading; that margin is +0.63 points.
Paying a model to do set arithmetic over a printed table that a dict comprehension does exactly and for nothing -- and equally, dismissing the model because the margin is small. +0.63 points is 1 reading of 160 here; on a book where one missed dealer is expensive that may be worth it. It is a decision, not a default, and this page publishes the number rather than the conclusion.
Deciding whether an investigator's declination note removes claims from the count
the model, on that subset only
This is the one clause a rule cannot express, and it is the only place the model earns its price on this kit. It reads 22 of the 24 prose-dependent readings against the free carry-aware floor's 21 -- 91.67 pct against 87.5. The keyword table classifies 12 of 16 wordings and the four it misses are substantive reasons written without the obvious word. EVERY error in both best arms lives here.
Reading 91.67 pct as solved. The model still gets two of these wrong (DLR-0033-W4, DLR-0036-W4), the subset is 24 readings drawn from sixteen sentences one author wrote, and that author also wrote the keyword table it is being compared against.
Running the watch with no memory between runs -- a weekly query, joined to nothing
nothing here. Do not do it
The free one-reading floor is exactly that, and it is what most programmes actually run. Level 51.88 pct, 44 false alarms in 103 quiet readings, a structural 0.0 pct on the counts that need memory, and UNCHANGED as the answer to 'what moved' on every reading of every week. It surfaces every real raise and buries them in 44 spurious ones.
Reading its 100.0 pct detection as a good result. It misses nothing because it raises nearly everything.
Choosing how often to run it
weekly, on the shipped constants -- and re-derive it for your own
evals/missed_run.py settles it arithmetically and for free. The parts clock is 21 days and the chargeback window 30, so a countable claim has 9 days of headroom; a weekly run is inside it and a fortnightly one is not. One missed run costs 6 claims worth 9,779.29 USD on this corpus, and changes 32 of 40 run-4 answers.
Carrying these two constants into your own programme. They are invented, and the whole conclusion is the difference between them.
Standing a monitor up for the first time on a book nobody has watched
a person, working the first run's output as a backlog rather than as a queue
A monitor's first run has no previous reading, so every standing claim reads as new. On this corpus run 1 surfaces 20 claims worth 30,380.79 USD that were ALREADY past the chargeback window on the day they were found. That is not a false positive and it is not a model failure -- it is a backlog, and treating it as a weekly queue would put a network's worth of unactionable findings in front of an investigator on day one.
Comparing run 1's numbers with later runs'. This kit reports them apart for that reason -- first-run level accuracy is 100.0 pct against 35.83 pct on later runs for the free one-reading floor, and the difference is the backlog, not the reader.
Debiting a dealer, raising a chargeback or opening an audit
the audit manager, on evidence an investigator has worked
Nothing in this kit does any of those and nothing should be added that does. The cap on the catalogue row is money-movement, hard and non-configurable, and evals/check_labels.py asserts the absence of ten money-moving call names before any run may spend.
Wiring any of this to a debit. There is no such code path and that is the guardrail.
A dealer network with years of weekly history, or thousands of dealers
nothing here yet -- measure it first
The carried position is four scalars, so the COST per run does not grow with history. What is unmeasured is accuracy over a long chain (every rule in this corpus resolves inside four runs), and what does grow is the PROMPT, because the population is re-read whole -- the claim book is already 24 pct of this prompt's characters at 536 claims.
Assuming any figure here transfers. The corpus mix decides every headline on this page.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
declination_read_as_administrative
A substantive declination read as an administrative one, so cleared claims keep counting
7
EVERY ERROR IN BOTH BEST ARMS OF THIS KIT IS THIS ONE CLAUSE -- the model's 2 and the free carry-aware floor's 5, and the model's two (DLR-0033-W4, DLR-0036-W4) are a subset of the floor's five. Taking the floor's in full: DLR-0033-W4, DLR-0034-W4…
no_memory_every_claim_reads_as_new
Standing claims counted again every week, because nothing says they were counted before
93
DLR-0029, the STABLE_HIGH storyline, is the cleanest instance. Five duplicate pairs, all paid before run 1, nothing new afterwards. The free one-reading floor reports INVESTIGATE with 5 new findings on run 1, run 2, run 3 and run 4 -- the same answer for…
movement_unanswerable_without_memory
"Which way did this dealer move" answered UNCHANGED because there is nothing to compare to
103
The free one-reading floor answers UNCHANGED on all 160 readings, by design -- a rule engine with no state cannot say whether anything moved, and evals/baseline.py says so rather than guessing. It is right 35.62 pct of the time, which is exactly the share of…
false_raise_on_a_quiet_dealer
A dealer with nothing new put on an investigator's queue
44
44 of 103 quiet readings for the free one-reading floor, against 3 for the carry-aware one. Every one is a dealer whose pile was already counted, or whose declination was already accepted. ⚠︎ THIS IS THE DIRECTION THIS VERTICAL PAYS FOR TWICE: an…
missed_raise
A dealer that earned a raise, reported quiet
1
None on the model or on either free floor, in 57 readings that earned a raise -- all three err in the other direction. ⚑ THE STATELESS CONTROL IS THE EXCEPTION AND IT IS THE WHOLE POINT: it misses 1, and it is the one arm told there is no history. Read the…
no_verdict
No reply at all
0
None on any completed run. The scored run answered all 160 readings, its stateless control 160 of 160, both free floors all 160, and the calibration all 8. The largest reply on the scored run reached 18,053 tokens against the published 24,000-token ceiling…
correct
Every cell correct
158
158 of 160 readings answered on all four fields by the scored model run, against 155 of 160 by the free carry-aware floor, including every one of the 32 CLEAN readings, all 16 STABLE_HIGH readings, all 16 PARTS_AGEING readings and all 16 DECLINED_ADMIN…
What we could NOT verify
⚑⚑ THE MODEL'S MARGIN OVER FREE CODE IS NOT ESTABLISHED AS STABLE, AND THE EVIDENCE AGAINST IT IS THIS KIT'S OWN SCREENSHOT. One live call was fired on DLR-0035-W2 to illustrate a reading the scored run got right. On a prompt assembled by the identical code path it got it WRONG -- REVIEW / 3 / WORSENING against the key's WATCH / 1 / UNCHANGED -- and its rationale classified the same investigator's note as 'administrative' where the scored run had called it 'substantive'. That is the PROSE subset, and the prose subset is where the whole +0.63-point margin is earned: on the arithmetic subsets the model and the free rule engine are within a reading of each other. One repeat is evidence, not a measurement, and it does not overturn the scored run -- but it is enough to withdraw any claim that the margin is reliable, and it says the thing worth measuring next is a repeat of the 24 prose-dependent readings rather than a bigger model. The frame is published as it came back.
⚠︎ AN EARLIER ATTEMPT AT THIS RUN WAS ABANDONED AND ITS SPEND IS RECORDED RATHER THAN HIDDEN. The provider account emptied mid-pass and returned 402 Payment Required after 39 of 160 readings. That partial file is preserved unaltered at kits/UC0087-warranty-dupe-watch/docs/incomplete/ with a note explaining why its numbers are an arithmetic consequence of 121 absent replies (39 of 160 is 24.375 pct, and every 'accuracy' in it is that same number wearing a different label). NOTHING on this page is drawn from it, the scored run below was written to its own file, and the run register refuses the abandoned one by itself -- runlog's triage extractor declines any record whose answered count is not a majority of its windows. ⚠︎ A SECOND INCIDENT COST MORE AND IS RECORDED TOO: when the account was refunded, a launch that appeared to have died had not, so two copies of the same run raced each other against one result file and double-billed every call until one was killed. src/budget.py counts calls and cannot see that two runs of one run-id are in flight -- its own docstring names that hazard and does not solve it.
⚠︎ THE MODEL'S MARGIN OVER FREE CODE IS ONE READING, AND ONE READING IS NOT A RATE. +0.63 points of level accuracy is the difference between 158 and 157 correct readings of 160. A corpus with a different mix of investigator wording moves that margin without either system having changed, and it could move it either way. The direction is measured; its stability is not.
One run per arm, each fired once. Nothing here is a distribution: whether 98.75 pct level accuracy repeats on a second identical run is unmeasured, and the small denominators (24 prose-dependent readings, 13 INVESTIGATE readings, 2 errors) would make a repeat noisy. The free arms are deterministic pure code and repeat exactly; the model arm has never been repeated.
One tier. No second model was run, so nothing on this page compares two of them -- the arms differ by MEMORY and by whether a model was called at all. The margin left above the free floor was 1.88 points before this run and 1.25 after it, so there is very little room for a better model to demonstrate anything on this corpus.
⚠︎ THE AUTHOR OF THIS KIT WROTE BOTH THE SIXTEEN DECLINATION NOTES AND THE KEYWORD TABLE THAT SCORES THEM. That is a conflict and it is named rather than hidden: a table fitted to those sixteen sentences would score 16 of 16 and prove nothing, and a table written to fail them would be a straw man. The markers in evals/baseline.py are generic audit vocabulary, written once before the notes were scored and not touched afterwards, and they land on 12 of 16. Whether that ratio -- or the model's advantage over it -- survives real investigator prose is unmeasured and unmeasurable without a real book.
The latency figures were measured while three other kits were running against the same provider key. p50 26,085 ms and p95 92,596 ms therefore include queueing a solo run would not have; they are an honest upper bound on this configuration and not a clean measurement of it. The token counts are unaffected.
⚠︎ THE PUBLISHED CEILING WAS SET FROM A CALIBRATION THAT UNDER-ESTIMATED THE TAIL BY 1.6x. MAX_TOKENS is 24,000, chosen as 2.1x the largest reply of an 8-call calibration on the two hardest dealers (11,346 at a 32,000 cap). The scored run's largest reply reached 18,053 -- 75 pct of the ceiling -- with 0 of 160 lost. It held, and it held with less room than the calibration implied.
Whether rendering the carried position as English rather than JSON matters. src/state.describe argues for English and says plainly the argument is unmeasured; scoring the same 160 readings with it rendered both ways costs one more full run and has not been paid for.
Whether disabling provider-side reasoning changes the answers. 97.8 pct of the scored run's output tokens were reasoning, left at the provider's default. Turning it off would change the price by close to an order of magnitude and this kit has not measured what it does to the four fields.
Whether the substantive / administrative split of a declination is the split a real warranty programme draws at all. The catalogue row records the guardrail as BLOCKED-PENDING-ANCHOR and every threshold behind every figure here is invented.
How any arm behaves on a chain longer than four runs. The consecutive-runs rule is only just exercised at four, and the failure shape a long chain would amplify -- a count that drifts and never resets -- has no instance in this corpus.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried position
2,560.16
4,304.31
26,085 ms
$0.014193
$0.005677
$0.240817
the same tier, memory removed (THE CONTROL)
2,522.68
7,250.21
50,927 ms
$0.023012
$0.009205
$0.387737
the free carry-aware rule floor
0
0
0 ms
$0.000000
$0.000000
$0.000000
the free one-reading rule floor
0
0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same dealer reading, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
⚑ THIS KIT DISCARDED CALLS AND PUBLISHES THE NUMBER, WHICH MOST DO NOT. The runs behind every figure on this page are 8 calibration + 160 scored + 160 stateless control + 40 missed-run arm. On top of those, 173 calls were BILLED and produced nothing publishable, reconciled against the shared ledger's 664 dispatches for this kit: 39 answered before the account emptied (preserved unaltered in docs/incomplete/, never scored), 22 from a launch that reported success and wrote no file, and 112 from its surviving twin racing the relaunch. A further 121 calls were REFUSED with 402 and billed nothing at all. Everything else -- both free floors, the wiring stub, the scorer, the pre-flight and the cadence measure -- is pure code and costs a measured $0.00. The dollar figures here are projections onto the shared Google Gemini 3 Flash card, not what anybody paid.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING, 97.8 pct of output tokens on the scored run (673,517 of 688,689) and therefore most of the bill. It was left at the provider's default and this kit has not measured what turning it off does to the four fields.
HOW MANY CLAIMS ARE IN THE BOOK, WHICH IS THE INPUT SIDE AND IS THE ONE THAT GROWS. This kit re-reads the whole open population every run, so the prompt scales with the dealer's 90-day book: input averaged 2560 tokens a reading across 160 readings, and the Claim Book section is 24 pct of the worked example's characters at 14 claims. A large dealer's book is several hundred, and unlike a per-document kit there is no chunking step to bound it.
HOW HARD THE SET ARITHMETIC IS, ON THE OUTPUT SIDE. Output averaged 4304 tokens and the largest reply reached 18,053 -- 75 pct of the published ceiling -- on readings that differ little in length. The long ones are the runs carrying a declination to read.
THE FIXED HEAD OF THE PROMPT, WHICH IS UNUSUALLY LARGE HERE BECAUSE THE RULEBOOK IS REPRINTED ON EVERY READING. The system message, the instruction and the programme rules are 5,923 of 10,148 characters -- 58 pct, byte-identical on all 160 readings. That is a deliberate transparency decision (the rules are invented, so they are printed to be disbelieved) and it is also the single biggest thing a prompt cache would discount.
⚑ AND THE LARGEST LEVER ON THIS PAGE IS STILL WHETHER YOU CALL A MODEL AT ALL. The free carry-aware floor scores 98.12 pct on the level for $0.00 against the model's 98.75. Every dollar here buys +0.63 points -- 1 reading of 160.
Your volumeWhat it costs at your volume
Linear in DEALER-RUNS, and the multiplier is the cadence rather than the corpus. One call per dealer per run: 40 dealers weekly is 2080 calls a year, and this run's 160 readings cost 2.2709 USD on the shared card, so ten times the set is about 22.71 USD. ⚠︎ A REAL DEALER NETWORK IS WHERE THIS ARITHMETIC STOPS BEING COMFORTABLE: at 0.014193 USD a reading, two thousand dealers watched weekly is about 28.39 USD a week -- 1,476.07 a year -- to buy +0.63 points of level accuracy over free code. The honest deployment is the model on the subset the floor cannot answer, which on this corpus is 24 readings of 160 and where the model is genuinely better (22 of 24 against 21). ⚠︎ AND THE INPUT SIDE GROWS, because the population is re-read whole: input averaged 2560 tokens on books of around 14 claims, and a large dealer's is several hundred. What does NOT scale is wall clock: runs within one dealer are strictly serial, and this run took 575 seconds of wall time for 160 calls at 12 workers.
Where pricing changes shape
⚠︎ THE OUTPUT CEILING WAS CLOSER THAN THE CALIBRATION IMPLIED. The published MAX_TOKENS is 24,000, set at 2.1x the largest reply the 8-call calibration saw (11,346 at a 32,000 cap). The scored run's largest reply reached 18,053 -- 75 pct of the ceiling -- with 0 of 160 replies lost. A calibration on the two hardest dealers under-estimated the tail of the full corpus by 1.6x, which is an argument for calibrating high rather than for calibrating cleverly. A ceiling is not billed: the provider charges tokens produced, not tokens allowed, so short readings pay nothing for the headroom the long ones need.
⚠︎ THE CLIFF THIS KIT ACTUALLY WENT OVER TWICE WAS NOT A TOKEN CEILING, IT WAS OPERATIONAL. First the account balance emptied mid-run (402 Payment Required, 39 of 160 answered, preserved in docs/incomplete/). Then, on the retry, a launch that appeared dead was not, and two copies of the same run raced one result file and double-billed every call until one was killed. src/budget.py caps calls per day and can see neither a balance nor a second run of the same run-id -- its own docstring names the second hazard and does not solve it. A 402 returns no completion and is billed for none, so the first incident cost 39 calls; the second cost roughly a hundred.
⚑ REMOVING THE MEMORY MADE THE MODEL COST MORE, NOT LESS -- 1,160,033 output tokens against 688,689, 1.68x, so the stateless control is 1.62x the price per reading (0.023012 against 0.014193 USD) to score 56.88 pct on the level instead of 98.75. Input is within 1.5 pct between the two arms, because the carried-position block is a few hundred characters; the entire difference is OUTPUT. Told there is no history for a dealer whose book plainly contains claims it has seen before, the model has an unresolvable problem and reasons about it at length. THE MEMORY IS A DISCOUNT AND A QUALITY GAIN AT THE SAME TIME, which is the opposite of the usual trade.
THE MEMORY STEP IS CHEAP ON BOTH SIDES OF THE LEDGER. The carried position is four scalars rendered as 265 characters -- 2.6 pct of the prompt -- so it costs almost nothing to send, and on the free floor it costs literally nothing and buys 46.24 points of level accuracy. What it does to OUTPUT tokens is measured here against the stateless control; see Cost.cost_by_model and the ripple table.
Your return, with your numbers
Volumedealer-runs per scheduled pass -- this corpus is 40 dealers x 4 weekly runs = 160 readings per arm
What it replacessomebody pulling each dealer's open claim book every Monday, working out which findings were already raised, reading the investigator's notes to see which declinations were real, and deciding whether the dealer is escalating or closing out
Time saved per itemnot measured here -- depends on how long reconstructing a dealer's raised-and-declined history takes in the reader's own audit system. ⚠︎ AND ON THIS KIT THE ROI QUESTION IS UNUSUAL: the free carry-aware floor delivers 98.12 pct of the level accuracy for $0.00, so the return being asked about is the model's MARGINAL return over free code -- +0.63 points -- not its return over the manual process.
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start a question whose honest answer might be 'free code already does this'. On this corpus free code does 98.12 pct of it, and the cheap tier buys +0.63 points more. A more expensive tier was not run: the margin available above the free floor was 1.88 points before this run started, so there is very little room for a better model to demonstrate anything, and buying that room was not worth another 160 calls.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
409,625input tokens · this run
688,689output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every model number on these pages: 160 readings, one completion call each, one tier. The 160-call stateless control, the 40-call missed-run arm, the 8-call calibration and roughly 173 discarded calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.908
$0.045
$5.68
2026-09-12
gemini-3-flash
Google
$2.271
$0.114
$14.19
2026-09-18
gemini-3-8-flash
Google
$2.890
$0.144
$18.06
2026-09-18
llama-5
Meta
$3.439
$0.172
$21.49
2026-09-18
claude-haiku-4-5
Anthropic
$3.853
$0.193
$24.08
2026-09-12
grok-4-5
xAI
$4.951
$0.248
$30.95
2026-09-18
grok-4-6
xAI
$4.951
$0.248
$30.95
2026-09-18
claude-sonnet-5
Anthropic
$7.706
$0.385
$48.16
2026-09-12
gemini-3-1-pro
Google
$9.084
$0.454
$56.77
2026-09-18
gpt-5-6-terra
OpenAI
$9.084
$0.454
$56.77
2026-09-12
gpt-5-6-sol
OpenAI
$15.412
$0.771
$96.33
2026-09-12
claude-opus-4-8
Anthropic
$19.265
$0.963
$120.41
2026-09-12
claude-opus-5
Anthropic
$19.265
$0.963
$120.41
2026-09-12
claude-fable-5
Anthropic
$38.531
$1.927
$240.82
2026-09-18
claude-fable-5-1
Anthropic
$38.531
$1.927
$240.82
2026-09-18
gpt-6-astra
OpenAI
$38.531
$1.927
$240.82
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are the scored run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 97.8 PCT OF THE SCORED RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (673,517 of 688,689), left at the provider's default, so every row below prices a reasoning-on workload. Output is 91 pct of the projected bill on the shared card, which means most of what these rows charge for is the model doing set arithmetic. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
⚑ THE CHEAPEST OPTION ON THIS PAGE IS NOT ON THIS TABLE, AND IT IS 0.63 POINTS BEHIND THE ROW THAT RAN. The free carry-aware rule floor scores 98.12 pct on the level for a measured $0.00 against the model's 98.75. Every row below is being asked to justify +0.63 points -- 1 reading of 160. Read the table with that number in hand.
⚠︎ A MORE EXPENSIVE TIER WAS NOT RUN, AND ON THIS KIT THAT IS A REASONED OMISSION RATHER THAN A GAP. The margin available above the free floor was 1.88 points before the scored run started; the cheap tier took +0.63 of it and left 1.25. There is very little room for a frontier model to demonstrate anything, and the rows below price that room rather than measure it.
⚠︎ AT SCALE THESE ROWS ARGUE AGAINST THEMSELVES. At 0.014193 USD a reading on the shared card, two thousand dealers watched weekly is about 1,476.07 USD a year -- to buy +0.63 points over free code. The honest deployment is the model on the subset the floor cannot answer, which on this corpus is 24 readings of 160.
Nothing here includes retries, and the scored run reported none that reached the harness -- though it was measured while three other kits shared the same provider key, so the LATENCY it carries includes queueing that the token counts do not. It also includes no tokens billed for an answer that never arrived: all 160 replies parsed. The discarded calls from the two spend incidents are not in these rows.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
15 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 40 dealers x 4 weekly runs = 160 readings and 536 claims from a fixed seed (SEED = 20260823), across nine deliberately separated storylines. It writes each reading, reads it BACK through the runtime's own parser, and computes the answer key from that -- so the page and the key cannot disagree about what is on the page.
You change it to: Point it at your own dealers and claim extracts, or delete it and drop real readings into data/corpus/ named <DEALER>-W<n>.txt with the same seven section headings. Everything downstream reads by dealer and run and does not care where the readings came from.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic, no clock read, no model, no network.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
RUN_DATES = ("2026-07-06", "2026-07-13", "2026-07-20", "2026-07-27")
RULE = "-" * 78
CAMPAIGNS = {
REGIONS = ("North", "Midlands", "South West", "South East", "Scotland", "Wales")
src/programme.pythe audit programme — a swap seam
The five countable-claim tests, the level table, the consecutive-runs rule and the declination clause, as pure code -- plus assess() and replay(), which ARE the answer key and which the runtime never calls. ⚠︎ Every threshold in it is INVENTED and the kit ships that as a stated blocker; the full rule text is reproduced on all 160 readings so it can be read, disbelieved and replaced.
You change it to: Every threshold, the level table, the consecutive-runs rule and the chargeback window are your programme's, not ours. Change them there and re-run tools/build_corpus.py, which recomputes the whole answer key from the same functions -- the key cannot drift from the rules because it is their output. ⚠︎ THIS IS THE FIRST FILE TO REPLACE, not an optional one: everything shipped is invented.
src/programme.py
# The warranty audit programme, as pure code. The answer key, and the rule text the reading prints.
DUPLICATE_SPAN_DAYS = 30
REPEAT_SPAN_DAYS = 180
REPEAT_MIN_CLAIMS = 3
LABOUR_MULTIPLE = 1.5
PARTS_MIN_USD = 500.0
PARTS_CLOCK_DAYS = 21
CHARGEBACK_WINDOW_DAYS = 30
BOOK_WINDOW_DAYS = 90
LEVELS = ("NONE", "WATCH", "REVIEW", "INVESTIGATE")
src/state.pythe carried position — a swap seam
SEAM 2 -- the thing that makes this a monitor. Four scalars per dealer (the level last reported, consecutive runs at REVIEW or above, the pattern attributed, and the claim number the last book ended at), written from the arithmetic and never from the model's reply, and rendered into English for the prompt. It costs the same on run 40 as on run 3.
You change it to: load()/save() are a JSON file today. Point them at a table, a key-value store or the audit programme's own database and nothing else in the kit changes -- for_dealer() is the whole read interface and describe() is the only thing the prompt sees.
src/state.py
# The carried state -- the thing that makes this a monitor and not another classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_dealer(store, dealer_id):
def describe(state, prev_run_date=None):
def _ordinal(n):
src/segment.pythe section splitter
Splits a reading into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 160 documents before a run may spend.
src/segment.py
# Split one dealer reading into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Dealer", "Audit Programme Rules", "Claim Book",
RULE = "-" * 78
def split(text):
def body_of(text, name):
def render(secs):
src/select.pythe section selector — a swap seam
Decides which sections reach the model. Dealer Contact -- the principal's name, trading address, telephone and email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
You change it to: SECTION_HINTS maps a field to the sections it needs; NEVER_SENT names what is withheld whatever happens. ⚠︎ On a dealer network this is the load-bearing seam: check NEVER_SENT covers every section your readings carry that names a person, before anything spends.
src/select.py
# Pick which sections of a dealer reading are sent. Pure code -- the last deterministic step
BANNER = "Synthetic Record"
DEALER = "Dealer"
RULES = "Audit Programme Rules"
BOOK = "Claim Book"
DISPOSITION = "Investigator Disposition"
CONTACT = "Dealer Contact"
NOTES = "Programme Notes"
NEVER_SENT = (CONTACT,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-position sentence, and the selected sections in document order. The stateless build replaces the position block with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, prev_run_date=None, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading. ⚠︎ It is also what refused this kit's scored run: a 402 Insufficient Balance is terminal, not transient, so it is raised rather than retried.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe watch
One dealer, one run, one call. Parses the claim book, the campaigns and the dispositions off the page with regexes (the model is never asked to read a number), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
src/watch.py
# One dealer, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You apply a written warranty audit programme to one dealer's claim book on one "
MAX_TOKENS = 24000
FIELDS = ("level", "pattern", "new_findings", "movement")
def documents():
def dealers():
def load_doc(doc_id):
CLAIM_RE = re.compile(
src/app.pythe local UI
One dealer, one run, its carried position and four columns, on 127.0.0.1:9001. Renders with no key. It shows the carried sentence verbatim and BOTH free floors beside the model's answer and the key, so a reader can see on which rows a model bought nothing.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9001"))
GOLD_ROWS = {}
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe free one-reading floor
Everything a rule engine reading THIS week's book can compute, with no idea what last week said. 0 calls, $0.00, scored through the identical scorer.
evals/baseline.py
# The two free floors. No model, no key, no network, and both scored through the same scorer.
ADMIN_MARKERS = ("insufficient", "out of time", "unable", "could not be reached",
SUBSTANTIVE_MARKERS = ("goodwill", "authorised", "approved", "campaign", "accepted",
def classify_note(note):
def _countable(book, run_date, campaigns, cleared):
def review(text, carried=None):
evals/baseline.pythe free carry-aware floor
The same rule engine given the same four scalars the model is given. It implements every printed rule except one -- whether a DECLINED note is substantive or administrative -- for which it uses an ordered keyword table. 0 calls, $0.00. ⚑ THIS IS THE HONEST COMPARATOR AND CURRENTLY THE BEST ARM THIS KIT HAS.
evals/baseline.py
# The two free floors. No model, no key, no network, and both scored through the same scorer.
ADMIN_MARKERS = ("insufficient", "out of time", "unable", "could not be reached",
SUBSTANTIVE_MARKERS = ("goodwill", "authorised", "approved", "campaign", "accepted",
def classify_note(note):
def _countable(book, run_date, campaigns, cleared):
def review(text, carried=None):
evals/scoring.pythe scorer
Exact match per cell against the computed key, split ways an average would hide: the four fields, the raise directions counted apart, the memory-dependent subset, the prose-dependent subset, the first run against the later ones, and the recoverable exposure in dollars. No judge model.
evals/scoring.py
# Score one arm against the computed answer key. Pure code -- no judge, no key, no network.
FIELDS = ("level", "pattern", "new_findings", "movement")
RAISED = ("REVIEW", "INVESTIGATE")
LEVEL_VALUES = set(PR.LEVELS)
PATTERN_VALUES = set(PR.PATTERNS) | {PR.NONE_DETECTED}
MOVEMENT_VALUES = set(PR.MOVEMENTS)
def _pct(n, d):
def _int_or_none(v):
def score(records, golds):
evals/missed_run.pythe cadence measure
Writes the SECOND answer key -- what the correct answers become when one weekly run never happens -- and prices the gap in claims and dollars. Free, no calls. ⚑ This is the half of a monitor that usually goes unmeasured.
evals/missed_run.py
# What a missed run costs. Free -- no model, no key, no network.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
OUT_GOLD = os.path.join(HERE, "data", "gold-missed-w3.jsonl")
OUT_REPORT = os.path.join(HERE, "results", "cadence-missed-w3.json")
SCHEDULE = ("W1", "W2", "W4")
SKIPPED = "W3"
def chain_for(dealer_id, labels):
def main():
evals/check_labels.pythe pre-flight
Eleven things that must be true before a run may spend: the corpus shape, the seven sections, claim numbers in payment order, every declination note in exactly one pool, the key re-derived from the rendered pages, the memory-dependent flag verified in both directions, the two prompt builds byte-identical outside one block, the privacy guard red-proven in both directions, no money-moving code path, both floors answering every reading, and the missed-run key reproducing.
evals/check_labels.py
# Everything that must be true before a run may spend. Free -- no model, no key, no network.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
GOLD_MISSED = os.path.join(HERE, "data", "gold-missed-w3.jsonl")
RUNS = ("W1", "W2", "W3", "W4")
EXPECT_DEALERS = 40
EXPECT_DOCS = 160
FORBIDDEN_CALLS = ("debit(", "charge_back(", "raise_chargeback(", "post_chargeback(",
FORBIDDEN_ON_PAGE = ("last reported", "previously reported", "previous level", "prev_level",
class Refused(SystemExit):
evals/run.pythe run harness — a swap seam
40 dealer chains, four strictly-ordered runs each, 12 concurrent workers. The position advances even for a run whose CALL failed, so one transport error cannot turn into four scored failures -- which is exactly the branch the 402 exercised.
You change it to: SCHEDULES maps a name to the sequence of runs the programme actually performs. 'weekly' and 'missed-w3' ship; add your own and evals/missed_run.py prices it against the weekly one for nothing. ⚠︎ This is the seam most kits in this series do not have, and it is where a monitor's real behaviour lives.
evals/run.py
# Run the watch over the 40 dealers and write one result file. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
GOLD_MISSED = os.path.join(HERE, "data", "gold-missed-w3.jsonl")
SCHEDULES = {
def load_gold(path):
def stub_complete(cfg, system, user, max_tokens=1024):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 40 dealers x 4 weekly runs = 160 readings and 536 claims from a fixed seed (SEED = 20260823), across nine deliberately separated storylines. It writes each reading, reads it BACK through the runtime's own parser, and computes the answer key from that -- so the page and the key cannot disagree about what is on the page. A swap seam.
src/programme.pyThe five countable-claim tests, the level table, the consecutive-runs rule and the declination clause, as pure code -- plus assess() and replay(), which ARE the answer key and which the runtime never calls. ⚠︎ Every threshold in it is INVENTED and the kit ships that as a stated blocker; the full rule text is reproduced on all 160 readings so it can be read, disbelieved and replaced. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Four scalars per dealer (the level last reported, consecutive runs at REVIEW or above, the pattern attributed, and the claim number the last book ended at), written from the arithmetic and never from the model's reply, and rendered into English for the prompt. It costs the same on run 40 as on run 3. A swap seam.
src/segment.pySplits a reading into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 160 documents before a run may spend.
src/select.pyDecides which sections reach the model. Dealer Contact -- the principal's name, trading address, telephone and email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing. A swap seam.
src/prompt.pyThree parts: the fixed instruction, the carried-position sentence, and the selected sections in document order. The stateless build replaces the position block with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading. ⚠︎ It is also what refused this kit's scored run: a 402 Insufficient Balance is terminal, not transient, so it is raised rather than retried. A swap seam.
src/watch.pyOne dealer, one run, one call. Parses the claim book, the campaigns and the dispositions off the page with regexes (the model is never asked to read a number), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
evals/baseline.pyEverything a rule engine reading THIS week's book can compute, with no idea what last week said. 0 calls, $0.00, scored through the identical scorer.
evals/baseline.pyThe same rule engine given the same four scalars the model is given. It implements every printed rule except one -- whether a DECLINED note is substantive or administrative -- for which it uses an ordered keyword table. 0 calls, $0.00. ⚑ THIS IS THE HONEST COMPARATOR AND CURRENTLY THE BEST ARM THIS KIT HAS.
evals/scoring.pyExact match per cell against the computed key, split ways an average would hide: the four fields, the raise directions counted apart, the memory-dependent subset, the prose-dependent subset, the first run against the later ones, and the recoverable exposure in dollars. No judge model.
evals/missed_run.pyWrites the SECOND answer key -- what the correct answers become when one weekly run never happens -- and prices the gap in claims and dollars. Free, no calls. ⚑ This is the half of a monitor that usually goes unmeasured.
evals/check_labels.pyEleven things that must be true before a run may spend: the corpus shape, the seven sections, claim numbers in payment order, every declination note in exactly one pool, the key re-derived from the rendered pages, the memory-dependent flag verified in both directions, the two prompt builds byte-identical outside one block, the privacy guard red-proven in both directions, no money-moving code path, both floors answering every reading, and the missed-run key reproducing.
evals/run.py40 dealer chains, four strictly-ordered runs each, 12 concurrent workers. The position advances even for a run whose CALL failed, so one transport error cannot turn into four scored failures -- which is exactly the branch the 402 exercised. A swap seam.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2560 input and 4304 output tokens per dealer reading, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Dealer readings/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per dealer reading directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Programme Notes, which are one of five fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- prose an analyst types into a review queue -- and sends them deliberately, so the surface is visible rather than hidden. ⚠︎ ON THIS VERTICAL THAT SURFACE IS UNUSUALLY ATTRACTIVE: a dealer under investigation has a direct interest in a note that talks the level down, and rule 6 means one sentence can legitimately remove claims from the countable population. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser -- which mattered here: the error the page actually renders today is a 402 body from the provider.
The experimentWe did not attack it -- and the boundary that matters most on this vertical was proven by breaking it on purpose
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus has none. So the five rows above are boundaries, not payloads. One of them was red-proven: src/select._fallback was replaced with the naive or list(secs) every sibling kit once shipped, the condition that actually reaches the fallback was reproduced -- a reading with no Synthetic Record banner and no Programme Notes, which is what a raw dealer-management-system extract looks like -- and all 160 readings leaked the principal's name, trading address, telephone and email. Restoring the guard took it back to 0 of 160. ⚠︎ THE REPRODUCTION IS THE HALF THAT MATTERS. Swapping the guard alone changes nothing on today's corpus, because every hint names a section every reading carries, so the fallback is unreachable and the red-proof would have measured ZERO and reported a green. A guard that has never been MADE to fail is a guard nobody has tested. Confirmed by assertion and by reading the recorded runs, not by an attack trial, on 2026-08-23 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a dealer principal's name and address ever leave the machine?
Every reading carries a Dealer Contact section -- the principal, their trading address, a telephone number and an email. It is the only place in this corpus whose subject is a PERSON, and the output of this kit is an ALLEGATION against that person that no investigator has yet worked. A selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the reading MINUS that section rather than to the reading. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py over all 160 documents, and the red-proof REPRODUCES THE CONDITION rather than merely swapping the guard: on today's corpus every hint names a section every reading carries, so the fallback is not on any live code path and a naive swap would measure zero and look like a pass. The proof strips the Synthetic Record banner and the Programme Notes -- which is exactly what a raw extract out of a dealer management system looks like before anybody adds a header -- and then: with the guard, 0 of 160 leak; with or list(secs), 160 of 160.
Can a wrong answer poison the next run?
A monitor that fed its own verdict into its next prompt would compound one bad run into every run after it -- and here the compounding has teeth in BOTH directions: a false INVESTIGATE would keep escalating itself through the consecutive-runs rule, and a missed finding would let a dealer's level ease away to NONE while the claims stood and the chargeback window ran out.
It cannot. evals/run.py advances the position by calling src/programme.assess on the parsed claim book; the model's reply is scored and never written. ⚠︎ Read that honestly: the property holds BY CONSTRUCTION and no run demonstrates it, because demonstrating it needs a wrong answer mid-chain and both of the scored run's errors land on run 4, the last of their chains. The stateless control's errors do sit mid-chain and do not propagate -- consistent with the property, not a proof of it.
Can a reading tell the model what the previous run decided?
If any reading named the programme's own previous level, its consecutive count or the claim number it stopped at, the stateless control could read the history off the page and the gap this kit publishes would be measuring nothing.
evals/check_labels.py scans the Dealer section of all 160 readings for six forbidden phrasings and refuses to let a run spend if it finds one. 0 hits. It also asserts that the stateful and stateless prompt builds are byte-identical outside the carried-position block, and that the memory-dependent subset is memory-dependent by MEASUREMENT -- computed by answering the same page twice, once with the position and once without, and verified in both directions.
Can this kit debit a dealer, raise a chargeback or open an audit?
A fraud watch that grew an apply path would be a system that decides, unsupervised, that a named dealer owes money -- and the cap on the catalogue row is money-movement, hard and non-configurable: no debit, chargeback or audit-opening act ever originates from the pack.
There is no endpoint, no function, no flag and no button anywhere in src/, evals/, tools/ or ui/ that debits, charges back, claws back, adjusts a claim or opens a case. The only things this kit writes are results/*.json and data/state.json. evals/check_labels.py asserts it mechanically over every .py, .js and .mjs file by looking for what would have to be present -- ten call names, crude and grep-shaped, and the only mechanical form a guarantee about ABSENCE can take.
Can a figure measured under a non-published output ceiling be mistaken for a scored one?
A calibration fired at a higher cap produces better-looking numbers under a configuration the page does not name. This kit's calibration ran at 32,000 tokens and its largest reply was 11,346; the scored run at the published 24,000 reached 18,053, so the calibration was not merely a different setting, it under-described the workload.
evals/run.py refuses --max-tokens unless the run id begins with 'c'. The calibration is c000, fired at 32,000; every scored arm -- r001, s001 and a001 -- carries the published MAX_TOKENS = 24,000 from src/watch.py, and no figure on this page mixes the two configurations.
Each boundary above was checked by running an assertion or by reading a recorded run, not by an attack trial -- there is no untrusted field on this corpus to construct a payload against. One of the five is red-proven, meaning the guard was removed and the failure observed rather than merely asserted while passing, and its red-proof reproduces the condition that reaches the guard rather than relying on today's corpus to reach it.
The result0 attack trials, five boundaries checked -- and the privacy boundary red-proven by removing the guard, reproducing the condition that reaches it, and watching all 160 readings leak a dealer principal's name, trading address, telephone and email.
0untrusted input fields on this corpus
0 of 0attack trials run
1 of 5boundaries red-proven, not just asserted
0 of 160readings leak a dealer principal's name or address
The Programme Notes ARE the field an outside party would influence in a real deployment, and this kit sends them rather than hiding them -- but on this corpus they are one of five sentences chosen by a seeded generator, so there is nothing adversarial in them to catch. ⚠︎ A version pointed at real analyst prose reopens the question sharply, because rule 6 gives a sentence the power to remove claims from the count. What IS measured is the privacy boundary, in both directions, before every run.
Read this twice
The carried position is written from the arithmetic, never from the model's reply, and on this kit that matters in two directions rather than one. A monitor that fed its own verdict forward would not lose one run to a bad answer: a false INVESTIGATE would keep re-escalating itself through the consecutive-runs rule, and a missed finding would let a dealer ease away to NONE while its claims stood and the chargeback window ran out. src/programme.assess never reads the reply. ⚠︎ AND NO RUN HERE DEMONSTRATES IT, WHICH THE PAGE SAYS RATHER THAN LETTING A NUMBER IMPLY OTHERWISE: the guarantee is in the code path and the scored run was never completed. What the boundary gives you is narrower than correctness -- it guarantees only that a wrong answer stays where it is.
HonestyWhat this does not prove
Whether a real programme's analyst notes -- prose somebody writes freely into a review queue -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at a real book. It is a sharper question here than on most kits because rule 6 lets one sentence remove claims from the countable population.
Whether the no-propagation property holds under a wrong answer MID-CHAIN. Both of the scored run's errors are on run 4, the last of their chains, so nothing downstream of them exists to be corrupted. The stateless control errs mid-chain and does not propagate, which is consistent with the property and is not a demonstration of it.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two schedulers advancing one dealer have never been run and would race. A lost run does not reset the position, it fails to advance it -- so the next run reads a stale claim number and re-raises claims already counted.
Whether the grep-shaped no-money-movement assertion would catch a path written under a name it does not know. It asserts the absence of ten call names; a path called something else would pass it.
Whether a truncated reply can ever be partially trusted. No run has hit the published ceiling -- the largest reply reached 75 pct of it -- so src/watch._parse's outright rejection of an unparseable reply has never been exercised on a real truncation.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No debit, no chargeback, no audit opened, non-configurable. This kit produces a level, a pattern, a count of what is new and a movement, for an investigator to work. It never debits a dealer, raises or posts a chargeback, claws back a claim, adjusts a claim value or opens a case, and there is no setting that makes it. Separately: the carried position is written by code from the parsed claim book and the previous position; the model's four answers are scored and are never written back into the history the next run is judged against.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json), evals/missed_run.py (a second answer key and a cadence report) and src/state.save (data/state.json). evals/run.py -> src/programme.assess is the only thing that touches the position, and it is called after every run including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit debits a dealer, charges back, claws back or opens an audit
0 code paths. evals/check_labels.py greps every .py, .js and .mjs file in the kit for ten money-moving call names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
A wrong answer does not propagate into the next run's position
6 wrong cells on r001-warranty-dupe-watch, and every dealer carrying one was judged at its next run against a position produced by the arithmetic rather than by the reply. Both of the model's errors are on RUN 4, the last of each chain, so this run cannot demonstrate the property on the model's own mistakes -- and the stateless control's 253 wrong cells sit across all four runs and do not propagate either. ⚠︎ The property is guaranteed by the code path (programme.assess never reads the reply); these runs are CONSISTENT with it rather than a demonstration of it.
A run whose CALL failed still advances the position
EXERCISED FOR REAL, and not by design. evals/run.py's exception branch calls programme.assess before continuing, so one transport error cannot turn into four scored failures. The abandoned scored run hit it 121 times when the provider returned 402, and every affected dealer's later runs were still judged against a correct position -- which is why the 39 answered readings are scored correctly in the preserved file even though most of their chain-mates were refused.
A dealer principal's name and address never reach the provider
0 of 160 readings leak the Dealer Contact section, and 160 of 160 leak it when the guard is removed AND the condition that reaches the fallback is reproduced. Both directions asserted by evals/check_labels.py before any run may spend.
The answer key cannot disagree with the page it is computed from
All 160 gold rows re-derived from the RENDERED files by src/programme.replay in evals/check_labels.py, plus all 120 rows of the missed-run key. This caught a real defect: a regex that silently dropped the last disposition note in every section -- the only one that could change an answer -- and made the two DECLINED storylines produce identical gold.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The position advances correctly whatever any reader says, which means a wrong level is reported wrongly for that run and nothing catches it -- see Eval.taxonomy's five declination misses, which this guardrail does nothing to prevent.
IT IS NOT AN APPROVAL GATE, BECAUSE THERE IS NOTHING TO APPROVE. There is no apply path to gate: nothing here moves money, so there is no 'undo' because there is nothing that could 'do'. A deployment that adds one adds the gate at the same time or it has no guardrail at all -- and on this vertical the thing being gated is a debit against a named business.
⚠︎ IT IS NOT A GUARANTEE ABOUT THE PROGRAMME ITSELF. Every threshold, the level table, the consecutive-runs rule and the chargeback window in src/programme.py are invented. A watch that is arithmetically perfect against an invented programme is arithmetically perfect against nothing.
IT IS NOT A GUARANTEE THAT THE QUEUE IS FAIR TO THE DEALER. The output names a business and asserts a pattern. The free one-reading floor raises 44 false alarms in 103 quiet readings and the carry-aware one raises 3; nothing in this kit tells a dealer they were raised, and nothing measures what a wrongly-raised dealer experiences.
It does not make the state file safe to share. One writer, one file, replaced atomically -- two schedulers on one dealer would race, and a lost write means the next run re-raises claims already counted.
It does not validate the declination taxonomy, which is the field every arm is worst on and the one the catalogue row already records as unconfirmed.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 56 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
7 measured by the latest run49 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
alarm
model.triage_detection; model.triage_page_precision; taxonomy.false_page; taxonomy.missed_incident; taxonomy.no_verdict — alarm on any unparsed reply on the published configuration, and any movement at all in taxonomy.missed_incident. Both free floors sit at 0 missed raises, so a first movement there is a new failure mode rather than a worse rate.
raises-missed-and-raises-invented
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
alarm
model.triage_detection; model.triage_page_precision; taxonomy.missed_incident; taxonomy.false_page — alarm on any missed raise at all. One is enough: the cost of that direction is the whole recoverable claim value, and the chargeback window keeps running.
memory-dependent-subset
The readings whose answer cannot be derived from their own page
alarm
memory_level_accuracy_pct; memory_new_findings_accuracy_pct — alarm on any drop on the carry-aware floor. It is at 96.77 pct, so a first movement is a new failure rather than a worse rate.
prose-dependent-subset
The readings where an investigator's sentence decides the answer
alarm
prose_level_accuracy_pct — alarm on a rise in this figure with no change to evals/baseline.py's marker lists -- it would mean the corpus moved, not that the reader improved.
the-cadence
What one missed run costs, in claims and in dollars
alarm
claims_lost_to_the_missed_run; usd_lost_to_the_missed_run — alarm on any missed run at all, on the shipped constants. There is no partial credit: the headroom is 9 days and a skipped week is 7.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
160
different corpus — nothing is comparable
corpus.bytes
1,046,916
dealer readings edited — the count held, the bytes did not
split.count
40
the dealers count moved — a different set was scored
split.size_p50
4
the median size of one dealer moved
split.size_p95
4
the 95th-percentile size of one dealer moved
dataset.rows
160
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (gate_recall 1.0) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
level accuracy
98.75 pct (model) against 98.12 (free carry-aware floor) -- a comparison, not a spread
160 readings
r001-warranty-dupe-watch against b001-warranty-dupe-watch-carryrule. ⚠︎ NEITHER IS A BAND. A band is the spread between runs of the SAME system at the same settings; the free arms are deterministic and repeat exactly, and the model arm has been run once. The +0.63-point difference is one reading of 160.
raises missed -- the expensive direction
0 of 57 on the model and on both free floors; 1 on the stateless control
57 readings that earned a raise
evals/scoring.py over r001, s001, b001 and b000. ⚠︎ THREE ARMS SIT AT ZERO AND ONE DOES NOT, WHICH IS THE ONLY REASON THIS GRADER SAYS ANYTHING. The corpus's raise threshold is reachable by a threshold rule, so 0 is mostly a statement about the corpus -- but the stateless control misses 1, and it is the one arm with no memory. The expensive direction is where losing the carried state finally shows up as a MISS rather than as noise.
false raises -- the cheap direction, which is not free
evals/scoring.py. This grader DOES discriminate, and the memory buys far more on it than the model does: 44 false alarms without the carried position, 3 with it, 2 with a model on top.
evals/scoring.py, over a subset flagged by MEASUREMENT (answer the same page twice, once with the position and once without) and verified in both directions by evals/check_labels.py. ⚑ THE FOUR-WAY SPREAD IS THE KIT'S CENTRAL RESULT: the two arms WITH the carried state are close together and the two without it are far below, whether or not a model is involved.
the prose-dependent subset -- where the model's margin is won
91.67 pct (model) against 87.5 (free carry-aware floor)
24 readings carrying a DECLINED finding
evals/scoring.py. The keyword table lands on 12 of 16 declination wordings; the model reads 22 of 24 readings against the floor's 21. ⚑ EVERY ERROR IN BOTH BEST ARMS IS HERE -- the model's 2 and the floor's 5.
the cadence arm -- answering over a missed run
97.50 pct on run 4 against the missed-run answer key
40 readings (run 4 only, run 3 skipped)
a001-warranty-dupe-watch-missedrun, scored against data/gold-missed-w3.jsonl -- a SECOND answer key, because the correct answers over a fortnightly gap are different answers and not worse ones. 32 of 40 run-4 answers genuinely differ between the two schedules.
claims lost to one missed run
6 claims, 9,779.29 USD -- a count, not a rate
171 claims first counted under the weekly schedule
evals/missed_run.py, free. Entirely determined by which claims cross the 21-day parts clock in the skipped week against a 30-day chargeback window -- 9 days of headroom against a 7-day slip.
readings correctly raised
not yet known -- and the one repeat this kit has FLIPPED
57 readings that earned a raise
57 of 57 on the model, the free carry-aware floor and the free one-reading floor alike; 56 on the stateless control. ⚠︎ NONE OF THAT IS A BAND. A band is the spread between runs of the SAME system at the same settings, and nothing here ran twice: the free arms are deterministic pure code and repeat byte-identically, and the model arm was fired once. The only repeat this kit has is accidental -- one live screenshot call re-asking a reading r001 got right -- and it came back with the OTHER answer. So the honest content of this row is that the spread is unmeasured and the single datapoint on it disagreed.
quiet readings correctly held
not yet known -- no arm has been run twice
103 quiet readings
101 of 103 on the model, 100 on the free carry-aware floor, 59 on the free one-reading floor and 65 on the stateless control. That four-way spread is a COMPARISON between four different systems, not a spread between repeats of one, and this row says so rather than dressing it up as a band. It is the complement of taxonomy.false_page and carries the same finding from the other side: the memory is what separates the arms, not the model.
the denominators themselves
0 -- constants of the corpus, not results
160 readings
57 readings that earned a raise, 103 quiet, 93 memory-dependent, 24 prose-dependent, 13 INVESTIGATE, on every arm. Which readings land where is decided by tools/build_corpus.py before any reader sees anything.
input volume
1.5 pct between the two model arms -- a measurement, not a spread
160 calls
409,625 input tokens with the carried position, 403,628 without -- the difference is the position block. Prompt assembly is pure code, so this figure is model-independent.
output volume and reasoning share
688,689 tokens with the carried position against 1,160,033 without
160 calls
97.8 pct of the scored run's output was provider-side reasoning (673,517 of 688,689). Two different prompts, not two runs of one, so this is a comparison and not a band.
latency
p50 26,085 ms, p95 92,596 ms -- and both include queueing from three other kits on one key
160 calls
One recorded run at the published configuration, measured under contention. The tail is 3.5 times the median.
replies that did not parse
0 of 160 (model) and 0 of 160 (the control), at the same ceiling
160 calls
Two configurations at the published 24,000-token ceiling. The largest reply reached 18,053, 75 pct of it.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 2 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · with the model in the path — 6 runs. Columns here are only ever compared with each other.
Metric
a001-warranty-dupe-watch-missedrun 2026-08-23
b000-warranty-dupe-watch-onereading 2026-08-23
b001-warranty-dupe-watch-carryrule 2026-08-23
c000-warranty-dupe-watch-calibration 2026-08-23
r001-warranty-dupe-watch 2026-08-23
s001-warranty-dupe-watch-stateless 2026-08-23
Verdict
input tokens, whole run
112327
—
—
24890
409625
403628
moved -1.5%
model latency p50 ms
47066.00
—
—
43032.00
26085.00
50927.00
moved +95.2%
model latency p95 ms
156577.00
—
—
87055.00
92596.00
138327.00
moved +49.4%
output tokens, whole run
310129
—
—
45785
688689
1160033
moved +68.4%
triage detection
100.00
100.00
100.00
100.00
100.00
98.25
moved -1.8%
triage page precision
100.00
56.44
95.00
100.00
96.61
59.57
moved -38.3%
false page
0
44
3
0
2
38
moved +1800.0%
held correct
19
59
100
4
101
65
moved -35.6%
missed incident
0
0
0
0
0
1
noise +0.0%
no verdict
0
0
0
0
0
0
exact
paged correct
21
57
57
4
57
56
moved -1.8%
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-warranty-dupe-watch-stub 2026-08-23
triage detection
100.0
triage page precision
56.44
false page
44
held correct
59
missed incident
0
no verdict
0
paged correct
57
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 7 chips that all say so.
What two runs bought
The deterministic metrics reproduced exactly — the same five rows fail, by name (L26 L28 L31 L32 L46) — so their band is exact-match: any movement at all is real. Latency was the only thing that drifted on identical inputs, by at most 0.0%, so its band must sit wider than that. Both bands above are those two facts.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the reader is given the four scalars of carried state
level 51.88 pct -> 98.12 pct, count 41.88 -> 96.88, movement 35.62 -> 98.12, all four 21.88 -> 96.88, false raises 44 -> 3, memory-dependent level 17.2 -> 96.77
measured
b000-warranty-dupe-watch-onereading against b001-warranty-dupe-watch-carryrule. Same corpus, same rule engine, same scorer, same 160 readings; the two differ in exactly one thing -- whether four scalars from the previous run are available. ⚑ NO MODEL WAS INVOLVED ON EITHER SIDE, which is the point: this is the argument for carrying state, and it is not an argument for a language model.
whether the reader can classify an investigator's declination note
prose-dependent level 87.5 pct for a keyword table against 91.67 pct for the model -- and EVERY remaining error in both best arms is here: the model's 2 and the free carry-aware floor's 5
measured
r001 against b001 over the 24 readings carrying a DECLINED finding. The floor's ordered keyword table classifies 12 of the 16 wordings; the 4 misses are substantive reasons written without the obvious word, and each keeps one or two claims countable that rule 6 removed. ⚑ THIS IS THE ONLY LEVER A MODEL PULLS ON THIS KIT, AND IT IS WHERE THE ENTIRE +0.63-POINT MARGIN COMES FROM.
whether a weekly run happens
6 claims worth 9,779.29 USD leave the chargeback window permanently, and 32 of 40 run-4 answers become different answers
measured
results/cadence-missed-w3.json, produced by evals/missed_run.py for $0.00. It compares the weekly schedule against one with run 3 removed, per CLAIM rather than per reading, and asks of both whether the chargeback window was open on the run that first counted the claim.
whether this is the programme's FIRST run
20 claims worth 30,380.79 USD arrive already out of time; first-run level accuracy 100.0 pct against 35.83 pct on later runs for the one-reading floor
measured
data/gold.jsonl grouped by run label, and results/cadence-missed-w3.json. On run 1 the carried position IS the empty position, so both floors agree and both score 100 pct -- which is exactly why the first run is reported apart. The backlog it surfaces is real work and mostly unactionable.
whether a model is called at all
level 98.12 pct -> 98.75, count 96.88 -> 98.75, movement 98.12 -> 98.75, all four 96.88 -> 98.75, false raises 3 -> 2, prose-dependent level 87.5 -> 91.67, cost per reading $0.00 -> 0.014193
measured
r001-warranty-dupe-watch against b001-warranty-dupe-watch-carryrule over the same 160 readings and the same scorer. ⚠︎ THE MARGIN IS ONE READING ON THE LEVEL AND THREE ON ALL-FOUR, and this row prints it beside the free row rather than on its own. It is roughly one seventieth of what the carried state buys, and it is the only lever on this table that costs money.
whether the MODEL is given the carried state (THE CONTROL)
r001-warranty-dupe-watch against s001-warranty-dupe-watch-stateless. Same corpus, same model, same grader, same 160 readings; the prompts differ in exactly one block and evals/check_labels.py asserts they are byte-identical everywhere else. ⚑ THIS IS THE SAME LEVER AS THE FIRST ROW OF THIS TABLE, PULLED ON A MODEL INSTEAD OF A RULE ENGINE -- and it moves the numbers by a comparable amount, which is the point: the carried state is what does the work, whoever reads it.
rendering the carried position as English rather than as JSON
unknown -- the experiment costs one more full run and has not been paid for
reasoning
src/state.describe's own docstring argues for English on the grounds that the rule text says 'every claim numbered above that is new' in English, and states plainly that this is a design choice rather than a measurement.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
level accuracy
any repeat of r001 that moves the level by more than one reading -- there is no repeat yet, so nothing has established what one reading of noise looks like
raises missed -- the expensive direction
any missed raise at all. One is enough: the cost is the whole recoverable claim value and the chargeback window keeps running
false raises -- the cheap direction, which is not free
any rise above 2 on a carry-aware arm. Every one is a named dealer asked to explain claims that were fine
the memory-dependent subset
any drop on a carry-aware arm
the prose-dependent subset -- where the model's margin is won
a rise with no change to evals/baseline.py's marker lists -- that would mean the corpus moved, not that the reader improved
the cadence arm -- answering over a missed run
a gap between this and the weekly arm's run-4 accuracy -- that would mean the kit answers a two-week gap worse, which is different from the gap COSTING more
claims lost to one missed run
any missed run at all, on the shipped constants
readings correctly raised
any drop below 57. It is the complement of taxonomy.missed_incident, so the two move together and a fall here IS a missed raise
quiet readings correctly held
any fall on a carry-aware arm. Every reading that leaves this count arrives in taxonomy.false_page as a named dealer asked to explain claims that were fine
the denominators themselves
any change at all, on any run
input volume
any change without a corresponding change to the prompt or the corpus
output volume and reasoning share
nothing yet
latency
nothing yet -- there is no contention-free measurement to compare against
replies that did not parse
any unparsed reply on the published configuration
NextThe three you would add first
A human step in front of anything that debits a dealerINVESTIGATE is the highest-consequence output this kit has and it is an assertion about a named business. The catalogue row's cap is money-movement, hard and non-configurable, and the row's own eval intent asks to see 'every dealer debited or charged back and the audit manager who decided it -- no debit traces to the pack alone'.
The real programme rules, before anything elseEvery threshold in src/programme.py is invented and the catalogue row reads BLOCKED-PENDING-ANCHOR. Every level on these pages is agreement with numbers written for this kit.
A declination reason recorded as a CODE beside its sentenceThe single clause a rule engine cannot express is whether a declination was substantive or administrative, and it costs 5 of the best arm's 5 errors. A programme that records the reason as a code alongside the free text removes the entire remaining gap -- and removes most of the argument for a model here at the same time. That is the honest recommendation even though it makes this kit smaller.
A first-run mode that presents the backlog as a backlogA monitor's first run has no previous reading, so every standing claim reads as new. On this corpus that surfaces 20 claims worth 30,380.79 USD that were already out of time on the day they were found. A queue that does not distinguish 'new this week' from 'the backlog we inherited' will be abandoned in week one.
A scheduler, and an alarm when a run does not happenThis kit measures a cadence and does not implement one. On the shipped constants a single missed weekly run costs 6 claims worth 9,779.29 USD, permanently. A monitor with no alarm on its own schedule has an undetectable failure mode whose symptom is a quiet week -- which is what everybody hopes to see.
A state store with a concurrency modeldata/state.json is one file replaced atomically. That is right for one process and is not a design for a scheduler.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
⚑ THIS KIT HAS TWO CADENCES AND THEY ARE DIFFERENT THINGS. THE PROGRAMME'S: weekly, and evals/missed_run.py measures why -- a countable claim has 9 days of headroom against the chargeback window, a weekly run is inside it and a fortnightly one is not, and one missed run costs 6 claims worth 9,779.29 USD permanently. Re-derive it for your own constants; it is free and takes a second. THE EVAL'S: re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/programme.py, src/segment.py or src/select.py -- it refuses to let a run spend if the corpus, either answer key, the experimental control, the privacy guard or the no-money-movement property has moved. Re-run both free floors together (free) on any change to evals/baseline.py: the headline is a DIFFERENCE, so one floor re-run alone is not comparable with the other's old figure. Re-run evals/missed_run.py (free) whenever a threshold or a schedule changes. Re-run the scored eval and its stateless control together (paid, 320 calls) on any change to src/prompt.py or src/state.py.
What this cannot tell you
Whether any model figure above repeats. Each arm was fired once; the free arms are deterministic pure code and repeat byte-identically, so the only thing on this page with a possible spread is also the only thing never measured twice. The model's margin over the free carry-aware floor is one reading of 160, which is inside the noise a repeat might show.
Whether the free floors' figures are stable. They are deterministic pure code, so a repeat is byte-identical -- which means they have no variance and also that nothing here estimates the variance a model arm would have.
Whether the no-propagation property would hold under a wrong answer mid-chain. It is true by construction -- programme.assess never reads the reply -- and no completed model run exists to be consistent with it.
Whether the no-money-movement guarantee holds against a path named something the assertion does not know. It greps for ten call names; the guarantee is about ABSENCE and that is the only mechanical form it can take.
Whether recording a declination reason as a code -- the change this kit recommends first -- would remove the whole remaining gap. It would remove the 5 measured errors; whether a real programme's codes are applied consistently enough to trust is a question about that programme and not about this kit.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried position
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a POSITION -- four scalars written by arithmetic. The whole reason the cost is flat in history length is that nothing here remembers what was said, only where the dealer stands. A memory layer would give back the growth this design exists to avoid, and the failure too: a level fed from its own model output would re-escalate itself through the consecutive-runs rule for ever
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install. ⚠︎ It also earned its keep on this kit: the 402 that stopped the scored run is classified as TERMINAL rather than transient, so 121 refusals cost nothing. A router with a generous retry policy would have made 121 more requests for the same answer
the schedule
evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect)
⚑ THIS IS THE SEAM WHERE A FRAMEWORK WOULD GENUINELY EARN ITS PLACE, AND THIS KIT IS THE ONE THAT CAN PRICE IT. 40 independent chains of 4 strictly-ordered steps, with cadence, retries, late-arriving runs and 'what do we do when Monday does not happen' all outside the kit -- and evals/missed_run.py measures that a missed Monday costs 6 claims worth 9,779.29 USD. A ThreadPoolExecutor is the right size for an eval and the wrong size for a monitor that actually runs on an audit calendar
the state store
src/state.py
durable state (Postgres, DynamoDB, a checkpointer)
data/state.json is one file replaced atomically -- correct for one writer, and not a concurrency model. A real programme's position belongs in the audit system's own database, and load()/for_dealer() are the whole read interface
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
four exact-match comparisons and six slices of the same cells is a dict comprehension, not a platform. ⚠︎ The one thing a harness WOULD have given for free is a second answer key per schedule, which this kit had to write itself in evals/missed_run.py
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each dealer is a chain of four steps -- run 1, 2, 3, 4 -- with no branching and exactly one edge between consecutive runs, carrying four scalars. Different dealers never touch. A framework would add an orchestrator to a for-loop that already runs 40 chains wide. ⚑ AND THE EDGE IS CHEAP TO REBUILD, which is the unusual part: because the position is arithmetic rather than model output, --runs W4 reproduces the whole chain for free and spends only on the step being measured.
The other sideWhat a framework costs you
No scheduler. This is a monitor that measures a cadence and does not implement one -- it is invoked, not woken. ⚠︎ On this kit that gap is priced rather than hand-waved: 6 claims worth 9,779.29 USD per missed run.
No state store worth the name. data/state.json is one file replaced atomically; a framework would bring a checkpointer with a concurrency model, which this does not have.
No observability beyond what evals/run.py prints and writes to results/. There is no tracing and no dashboard integration.
No retry/backoff beyond src/adapters' own bounded retry, which covers a busy provider and a dropped connection and nothing else -- and deliberately does NOT cover a 402, which is terminal.
No budget guard beyond src/budget.py's daily call cap, which counts calls, not dollars, and can see neither an account balance nor a second copy of the same run. ⚠︎ THAT IS THE GAP THIS KIT PAID FOR TWICE: the balance ran out mid-run (39 calls billed for nothing), and then a launch that appeared dead was not, so two copies of one run raced each other and double-billed until one was killed. The second hazard is named in budget.py's own docstring and not solved there.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
Whether a scheduling engine would actually prevent a missed run, as opposed to relocating the question. The cost of the missed run is measured; the reliability of any fix for it is not.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-warranty-dupe-watch on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
26,085 ms
p50 26,085 ms, p95 92,596 ms -- and both include queueing from three other kits on one key
nothing yet -- there is no contention-free measurement to compare against
Model, p95
92,596 ms
p50 26,085 ms, p95 92,596 ms -- and both include queueing from three other kits on one key
nothing yet -- there is no contention-free measurement to compare against
Input tokens
409,625
1.5 pct between the two model arms -- a measurement, not a spread
any change without a corresponding change to the prompt or the corpus
Output tokens
688,689
688,689 tokens with the carried position against 1,160,033 without
nothing yet
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
a001-warranty-dupe-watch-missedrun47,066 ms
c000-warranty-dupe-watch-calibration43,032 ms
r001-warranty-dupe-watch26,085 ms
s001-warranty-dupe-watch-stateless50,927 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
dealer readings
data/corpus/<DEALER>-W<n>.txt -- 160 files, 1,046,916 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Dealer Contact -- the principal's name, trading address, telephone and email -- never does, by src/select.NEVER_SENT
the answer keys, both of them
data/gold.jsonl (160 rows, the weekly schedule) and data/gold-missed-w3.jsonl (120 rows, the schedule with one run removed) -- both the output of src/programme.replay over the RENDERED pages, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the carried position
data/state.json in a deployment (src/state.py, written atomically). ⚠︎ IN AN EVAL IT IS SCOPED TO THE RUN AND NEVER TOUCHES DISK -- evals/run.py builds a fresh in-memory position per dealer, because a run that started from the previous run's memory could not be re-run or compared with its own control
two or three short English sentences per call, produced by state.describe -- four scalars, never a prior reading and never a prior reply
the recorded runs
results/eval-*.json -- the scored run, its stateless control, the missed-run arm, both free floors, the wiring stub and the calibration, plus results/cadence-missed-w3.json, all committed. ⚠︎ AND ONE FILE THAT IS DELIBERATELY NOT IN results/: docs/incomplete/eval-r001-...-INCOMPLETE-402.json, an abandoned earlier attempt at the scored run, moved aside rather than deleted or overwritten so the 39 calls it billed stay recoverable from the repository
never -- they are read by the page and by nothing else
the key
.env or the shared repo-root .env -- never committed, 0600
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py). src/app.py redacts it out of any error it renders, which is live rather than theoretical here: for a day the error this page rendered was a provider 402 body
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser -- which mattered here: the error the page actually renders today is a 402 body from the provider.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
WEEKLY, and the kit measures why rather than asserting it. One run owns exactly what the previous one did not: claims filed since the previous reading's highest claim number, plus claims that crossed the 21-day parts clock between the previous run date and this one. Both facts are on the page (the previous run date) or in the carried state (the claim number). ⚠︎ THE KIT MEASURES A CADENCE AND DOES NOT IMPLEMENT ONE: there is no scheduler here, and evals/run.py's SCHEDULES map is the seam where one would attach.
a countable claim has 9 days of headroom (chargeback window 30 minus parts clock 21). One missed weekly run costs 6 claims worth 9,779.29 USD permanently, changes 32 of 40 run-4 answers, and drops recoverable claims from 151 to 145 of 171. The model still ANSWERS the two-week gap correctly -- 97.50 pct on run 4 against the missed-run key, against 98.75 pct on the same run weekly (results/cadence-missed-w3.json and evals/missed_run.py (free, 0 calls), plus a001-warranty-dupe-watch-missedrun (40 calls))
a fortnight. The headroom is nine days and a skipped week is seven, so the weekly schedule is inside it with two days to spare and the fortnightly one is five days outside. There is no partial credit and no tuning: the answer is settled by two constants, both invented, and re-derived for free by re-running the module
⚑ ANSWERING THE GAP AND PAYING FOR IT ARE DIFFERENT QUESTIONS, and only the first is a model result. The kit copes with a two-week gap; the programme still loses 6 claims worth 9,779.29 USD, permanently, because a claim that goes out of time stops being an alert rather than becoming a late one. ⚠︎ AND A MONITOR'S FIRST RUN IS ITS WORST RUN, WHICH NO CADENCE FIXES. Run 1 has no previous reading, so every standing claim reads as new: 20 claims worth 30,380.79 USD arrived already past the chargeback window on the day they were found. That is a backlog, not a queue, and a deployment that does not separate them will be abandoned in week one.
state
four scalars per dealer -- the level last reported, how many consecutive runs it has been at REVIEW or above, the pattern attributed, and the claim number the last reading's book ended at -- written by src/programme.assess from the PARSED claim book and never from any reply, and rendered by src/state.describe into two or three English sentences. That block is the entire route from one run to the next, and it is what the two model arms differ by.
265 characters on the worked example -- 2.6 pct of the prompt. Removing it from the MODEL's prompt moves level accuracy 56.88 pct -> 98.75, the count 48.75 -> 98.75 and the memory-dependent subset 25.81 -> 97.85. Handing the same four scalars to a free RULE ENGINE instead moves the level 51.88 -> 98.12 at $0.00 on both sides. The carried state is what does the work, whoever reads it. (r001-warranty-dupe-watch against s001-warranty-dupe-watch-stateless, and b001-warranty-dupe-watch-carryrule against b000-warranty-dupe-watch-onereading; lenses.LLM.prompt_parts)
the position is four scalars, so run 40 costs what run 3 costs -- the OPPOSITE curve to an intake kit, whose input grows with the square of the turns. What is NOT bounded is the READING: the population is re-read whole every run, so the prompt grows with the dealer's open book, and the claim book is already 24 pct of this prompt's characters at 14 claims. Nor is accuracy over a long chain bounded -- every rule in this corpus resolves inside four runs. The store itself is data/state.json, one file replaced atomically: correct for one writer and not a concurrency model
take the carried position away and the count field collapses on the readings that need it, movement becomes near-unanswerable, and false raises rise -- with every answer still well-formed and nothing raising an error. See environment.signatures' traceless row
model
one call per DEALER PER RUN carrying the instruction, the carried-position block and six of the reading's seven sections, behind src/adapters/__init__.py, at the published MAX_TOKENS = 24,000 with provider-side reasoning left at the default
160 of 160 replies parsed on the scored run and 160 of 160 on its stateless control; the largest reply reached 18,053 output tokens, 75 pct of the published ceiling. 97.8 pct of output was provider-side reasoning. Input averaged 2560 tokens a reading and output 4304. The model beats a free rule engine given the same carried state by +0.63 points of level accuracy (r001-warranty-dupe-watch and s001-warranty-dupe-watch-stateless; lenses.Eval.scores)
⚠︎ THE CEILING HELD WITH LESS ROOM THAN THE CALIBRATION IMPLIED -- 18,053 of 24,000, 75 pct -- because the calibration measured the two hardest DEALERS and not the hardest READINGS. But the binding constraints on this kit were both operational rather than numeric: an empty account balance (402, 39 readings lost) and then two copies of one run racing each other. src/budget.py caps calls per day and can see neither
one tier only, so nothing here ranks two models -- the arms differ by MEMORY and by whether a model was called at all. And the margin a better model could win is 1.25 points, because free code already takes 0.63 of the 1.88 that were available. Point .env at your own model and the free scorer re-runs on your numbers for nothing
labels
data/gold.jsonl, 160 rows, computed by src/programme.assess over the RENDERED pages at generation time -- a function the runtime never calls, so the key is arithmetic done independently of the code under test. Plus data/gold-missed-w3.jsonl, 120 rows: a SECOND key for the schedule with one run removed, because the correct answers under a fortnightly gap are different answers and not worse ones
50 NONE / 53 WATCH / 44 REVIEW / 13 INVESTIGATE over 160 readings; 57 that earned a raise, 103 quiet, 93 memory-dependent (verified in both directions), 24 prose-dependent, 48 DECLINED findings drawn from two pools of eight wordings (lenses.Eval.dataset, warranty-dupe-watch-2026-08-23-40dealers-160readings; evals/check_labels.py re-derives every row of both keys before a run may spend)
⚠︎ THE KEY IS ONLY AS GOOD AS AN INVENTED PROGRAMME. Every threshold, the level table, the consecutive-runs rule and the chargeback window were written for this kit and confirmed against nothing; the catalogue row reads BLOCKED-PENDING-ANCHOR. And the one judgement in the key that is not arithmetic -- whether a declination reason is substantive -- is a lookup against sixteen sentences the kit's own author wrote, which is also the only subset where the model beats free code
your own programme: hand-label the gold, which is the real work. This kit's key is a luxury of controlling both the generator and the rules, and a hand-labelled key has an error rate nothing here has measured
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
the same dealer at INVESTIGATE with the same finding count, week after week after week
something is reading each week's book with no memory of the last one, so every standing claim counts as new every time. It is not an escalating dealer; it is the same pile being counted again
compare the finding count against the previous run's before working the dealer. On DLR-0029 the free one-reading floor reports INVESTIGATE with 5 new findings on all four runs; the key says INVESTIGATE, REVIEW, WATCH, NONE -- a dealer working its way off the queue (results/eval-b000-warranty-dupe-watch-onereading.json against results/eval-b001-warranty-dupe-watch-carryrule.json, DLR-0029)
a dealer raised again on claims an investigator already declined, with the declination sitting right there on the page
the declination's reason was read as administrative when it was substantive. Rule 6 removes the claims only for a substantive reason, and the reason is a sentence -- a keyword table that finds no marker defaults to 'keep counting'
read the note, not the status. Every error in both best arms of this kit is this -- the model's 2 and the free carry-aware floor's 5 -- and all of them are wordings without an obvious marker: 'the labour hours are the published time for the revised procedure', 'the file shows two unrelated faults'. Group the raised dealers by their declination wording before working any of them (results/eval-r001-warranty-dupe-watch.json and results/eval-b001-warranty-dupe-watch-carryrule.json, misses on DLR-0033 to DLR-0036)
a monitor's first run producing a huge queue, most of which cannot be acted on
nothing is wrong. Run 1 has no previous reading, so the whole standing population reads as new -- and a claim's chargeback window has been running since it was paid, not since anybody looked
check the payment dates before staffing the queue. On this corpus run 1 surfaces 20 claims worth 30,380.79 USD that were already out of time on the day they were found. Work it as a backlog with its own triage, and read week 2 as the first real week (results/cadence-missed-w3.json, claims_already_out_of_time_on_the_run_that_found_them)
a quiet week, right after a week the scheduled run did not happen
possibly nothing, and possibly that findings which should have been raised are now outside the chargeback window and will never appear as anything at all. A claim that goes out of time does not become an alert; it stops being one
re-run evals/missed_run.py against your own constants before assuming the week was quiet. On the shipped ones a single missed run silently removes 6 claims worth 9,779.29 USD from what any later run can recover -- and the kit will keep answering correctly the whole time, at 97.50 pct on the run after the gap (results/cadence-missed-w3.json and results/eval-a001-warranty-dupe-watch-missedrun.json)
No machine symptom — this failure leaves no trace in any output.
There is none, and this is the failure with no machine symptom. A deployment that quietly lost its carried position -- a state file truncated, a store returning empty, a schedule that silently re-ran run 1 -- would keep answering, keep parsing and keep costing money, and would report a BIGGER queue rather than a smaller one. Every standing claim would read as new, so the symptom is a busy week, which looks exactly like a dealer network getting worse. No check in this kit compares two runs at run time. Both pairs of arms measure how invisible it is: the model loses 41.87 points of level accuracy when the position is removed and the free rule engine loses 46.24, and every one of those wrong answers is well-formed.
⚠︎ THE LARGEST UNMEASURED THING IS WHETHER ANY OF IT REPEATS. Every arm was fired once. The two free arms are deterministic pure code and repeat byte-identically, so the only figures with a possible spread -- the model's -- are the ones never measured twice, and the model's margin over the free carry-aware floor is one reading of 160. ⚑ THIS KIT IS AMONG THE FIRST TO USE THE monitor FLOW VARIANT, AND THREE THINGS ABOUT THE SURROUNDING VOCABULARY DID NOT FIT AND ARE RECORDED RATHER THAN PAPERED OVER. FIRST, app.tabs[0].kind IS triage BECAUSE APP_UI HAS NO monitor RENDERER -- the panel this kit draws is four columns over one reading, not a page/hold decision on a window, and the nearest available kind is used unchanged rather than a station word being invented. SECOND, THE RUN REGISTER HAS NO monitor RUN KIND: every record this kit writes is kind: "triage", and runlog's triage extractor records the two directions under model.triage_detection and model.triage_page_precision and counts them in a windows guard. A reading is a whole population re-read on a schedule, not a slice cut out of a stream, and nothing pages anybody. THIRD, THAT EXTRACTOR CARRIES ONLY TWO OF THIS KIT'S SIX METRICS -- the level, the pattern, the new-finding count and the movement are in the committed result files and not in the register -- and it has no field for a run's SCHEDULE, which is this kit's real comparability guard the way window_seconds is triage's. r001 and a001 are scored against DIFFERENT answer keys and are not one series, and their register records cannot currently say so. WHAT IS ALSO NOT MEASURED: concurrency and worker sizing -- 40 dealer chains at 12 workers, and the one timing figure this kit has was taken while three other kits shared the provider key. Whether runs within a dealer can ever be reordered -- they cannot, by construction. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- every figure prices the cache-miss rate for exactly that reason, and the fixed head of this prompt (5,923 of 10,148 characters, byte-identical on all 160 readings, because the invented rulebook is reprinted so it can be disbelieved) is precisely what a cache would have discounted. Whether GPU sizing or a local model changes any of it -- untested.
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every dealer, principal, address, VIN, claim, operation code, recall campaign, investigator finding and note is invented here. Verified against the repository's own LICENSE file on 2026-08-23. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineThe level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
For each of the 160 readings and each of the four answered fields, did the reply equal the computed key? Level, pattern and movement are compared exactly; the count is compared as a whole number, and a reply that is not a whole number is a MISS rather than a zero -- 0 is a meaningful answer here (it is what an easing dealer gets) and a parse failure must never be scored as one.
$0.00per 1,000 dealer readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function both free floors and the wiring stub are scored through.
Every grader on these pages scored the same 480 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DLR-0035-W2 -- a DECLINED_SUBSTANTIVE dealer, run 2 of 4
The book and the schedule
14 claims in the book, re-read whole. Run date 2026-07-13; the programme last read this dealer on 2026-07-06 and its book then ended at CLM-100249.
What the disposition carries
a DECLINED finding naming two claims, and a one-sentence reason that carries none of the words an audit keyword table would look for
Carried position, in the prompt
At the previous run on 2026-07-06 this dealer was reported WATCH for DUPLICATE_CLAIM. It has not been at REVIEW or above on any run up to and including that one. The claim book then ended at CLM-100249; every claim numbered above that is new since the last reading.
Free rule floor, no memory
INVESTIGATE, 4 new findings, UNCHANGED -- it has no memory at all
Free rule floor, same memory
REVIEW, 3 new findings, WORSENING -- right about everything except the sentence
The model, with the carried position
{"level": "WATCH", "pattern": "DUPLICATE_CLAIM", "new_findings": 1, "movement": "UNCHANGED"} -- CORRECT, and it is one of the readings that separates the model from the free floor on this page
Ground truth
WATCH, 1 new finding, UNCHANGED
Scored as
Both free floors put this dealer on an investigator's queue when the programme's own rule says it should not be there, and the model does not. It is the shape of row this kit was built to isolate and the only shape a model is paid for here -- and the model still gets two of these wrong elsewhere (DLR-0033-W4, DLR-0036-W4), so reading a declination is where it is better rather than where it is finished.
Grader
Verdict
Why
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
model: all four correct. Both free floors: level, count and movement MISS
DLR-0035-W2. Between run 1 and run 2 the investigator declined a finding naming two claims, with the note 'the sequence number falls inside the range as amended in June and the reading predates the amendment'. Rule 6 says a substantive reason removes those claims, so the key is WATCH / 1 / UNCHANGED and the model answered exactly that. The free carry-aware floor's keyword table carries no marker that fires on that sentence, defaults to ADMINISTRATIVE -- the conservative direction -- keeps both claims countable and answers REVIEW / 3 / WORSENING. The free one-reading floor is worse still at INVESTIGATE / 4.
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
a FALSE RAISE by both free floors; the model held it correctly
The key is WATCH, which is below the REVIEW threshold, so this reading sits in the QUIET denominator of 103 and not in the raised one. Both free floors report REVIEW or above and put a named dealer on an investigator's queue when the programme's own rule says it should not be there; the model does not. Across the whole set the carry-aware floor does this 3 times and the one-reading floor 44, against the model's 2. Neither direction is averaged into the other.
The readings whose answer cannot be derived from their own page
in scope -- one of the 93, and the model is right on it
This reading is memory-dependent by MEASUREMENT rather than by label: the same page answered with no carried position produces a different answer, which is how tools/build_corpus.py flags the subset and how evals/check_labels.py verifies it in both directions. The model scores 97.85 pct here, the free carry-aware floor 96.77, and the two arms with no memory at all collapse to 25.81 and 17.20.
The readings where an investigator's sentence decides the answer
in scope -- and this is the row the model's whole margin is made of
One of the 24 readings carrying a DECLINED finding, and one of the 5 the free carry-aware floor gets wrong. The model reads 22 of them against the floor's 21. ⚠︎ AND THIS EXACT READING IS ALSO THE ONE THAT FLIPPED: a single unplanned live call, on a prompt built by the identical code path, came back REVIEW / 3 / WORSENING -- the free floor's answer -- and called the same note administrative. The scored run's answer stands; the stability of it does not.
What one missed run costs, in claims and in dollars
not in scope -- this grader scores CLAIMS, not readings
The cadence grader compares two schedules per CLAIM, asking whether the chargeback window was still open on the run that first counted it. It has no verdict to give about a single reading, and this one is shown as out of scope rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things. What it measures instead: one missed weekly run costs 6 claims worth 9,779.29 USD permanently, and the model still answered the degraded schedule at 97.50 pct.
The formulaWhat it computes
accuracy = hits / 160 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried position
98.8% level accuracy · 3 more measured on this row
the fast tier, memory removed (THE CONTROL)
56.9% level accuracy · 3 more measured on this row
the free carry-aware rule floor, no model
98.1% level accuracy · 3 more measured on this row
the free one-reading rule floor, no model
51.9% level accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/programme.replay over the RENDERED pages at generation time. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known to be arguable -- whether the substantive / administrative split of a declination is the split a real programme would draw. The other three fields are arithmetic over a printed table and have no known ambiguity.
Watch these
model.triage_detection
model.triage_page_precision
taxonomy.false_page
taxonomy.missed_incident
taxonomy.no_verdict
Alarm on
any unparsed reply on the published configuration, and any movement at all in taxonomy.missed_incident. Both free floors sit at 0 missed raises, so a first movement there is a new failure mode rather than a worse rate.
How tight can the band be? No tuned threshold anywhere. The denominators are printed beside every rate because three of them are small: 24 prose-dependent readings, 13 INVESTIGATE readings, and 5 total errors in the best arm.
Cadence: Re-run evals/check_labels.py (free) on any change to the corpus generator, src/programme.py, src/segment.py or src/select.py. Re-run both free floors (free) on any change to evals/baseline.py. Re-run the scored eval (paid, one call per reading) on any change to src/prompt.py or src/state.py -- both change what the model is told. Re-run evals/missed_run.py (free) whenever a threshold or the schedule moves, because the cadence answer is arithmetic over both.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real audit programme, where whether a declination was substantive is decided by a person reading a note. That is why this corpus is generated rather than captured.
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineThe two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
Of the 57 readings whose correct level is REVIEW or INVESTIGATE, how many were reported at REVIEW or above? And of the 103 quiet readings, how many were put on the queue anyway? 103 of 160 readings are quiet, so an arm reporting NONE to everything would score 64.4 pct on any per-reading accuracy while surfacing nothing.
$0.00per 1,000 dealer readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader -- a slice of the same cells, not a second comparison.
Every grader on these pages scored the same 480 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DLR-0035-W2 -- a DECLINED_SUBSTANTIVE dealer, run 2 of 4
The book and the schedule
14 claims in the book, re-read whole. Run date 2026-07-13; the programme last read this dealer on 2026-07-06 and its book then ended at CLM-100249.
What the disposition carries
a DECLINED finding naming two claims, and a one-sentence reason that carries none of the words an audit keyword table would look for
Carried position, in the prompt
At the previous run on 2026-07-06 this dealer was reported WATCH for DUPLICATE_CLAIM. It has not been at REVIEW or above on any run up to and including that one. The claim book then ended at CLM-100249; every claim numbered above that is new since the last reading.
Free rule floor, no memory
INVESTIGATE, 4 new findings, UNCHANGED -- it has no memory at all
Free rule floor, same memory
REVIEW, 3 new findings, WORSENING -- right about everything except the sentence
The model, with the carried position
{"level": "WATCH", "pattern": "DUPLICATE_CLAIM", "new_findings": 1, "movement": "UNCHANGED"} -- CORRECT, and it is one of the readings that separates the model from the free floor on this page
Ground truth
WATCH, 1 new finding, UNCHANGED
Scored as
Both free floors put this dealer on an investigator's queue when the programme's own rule says it should not be there, and the model does not. It is the shape of row this kit was built to isolate and the only shape a model is paid for here -- and the model still gets two of these wrong elsewhere (DLR-0033-W4, DLR-0036-W4), so reading a declination is where it is better rather than where it is finished.
Grader
Verdict
Why
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
model: all four correct. Both free floors: level, count and movement MISS
DLR-0035-W2. Between run 1 and run 2 the investigator declined a finding naming two claims, with the note 'the sequence number falls inside the range as amended in June and the reading predates the amendment'. Rule 6 says a substantive reason removes those claims, so the key is WATCH / 1 / UNCHANGED and the model answered exactly that. The free carry-aware floor's keyword table carries no marker that fires on that sentence, defaults to ADMINISTRATIVE -- the conservative direction -- keeps both claims countable and answers REVIEW / 3 / WORSENING. The free one-reading floor is worse still at INVESTIGATE / 4.
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
a FALSE RAISE by both free floors; the model held it correctly
The key is WATCH, which is below the REVIEW threshold, so this reading sits in the QUIET denominator of 103 and not in the raised one. Both free floors report REVIEW or above and put a named dealer on an investigator's queue when the programme's own rule says it should not be there; the model does not. Across the whole set the carry-aware floor does this 3 times and the one-reading floor 44, against the model's 2. Neither direction is averaged into the other.
The readings whose answer cannot be derived from their own page
in scope -- one of the 93, and the model is right on it
This reading is memory-dependent by MEASUREMENT rather than by label: the same page answered with no carried position produces a different answer, which is how tools/build_corpus.py flags the subset and how evals/check_labels.py verifies it in both directions. The model scores 97.85 pct here, the free carry-aware floor 96.77, and the two arms with no memory at all collapse to 25.81 and 17.20.
The readings where an investigator's sentence decides the answer
in scope -- and this is the row the model's whole margin is made of
One of the 24 readings carrying a DECLINED finding, and one of the 5 the free carry-aware floor gets wrong. The model reads 22 of them against the floor's 21. ⚠︎ AND THIS EXACT READING IS ALSO THE ONE THAT FLIPPED: a single unplanned live call, on a prompt built by the identical code path, came back REVIEW / 3 / WORSENING -- the free floor's answer -- and called the same note administrative. The scored run's answer stands; the stability of it does not.
What one missed run costs, in claims and in dollars
not in scope -- this grader scores CLAIMS, not readings
The cadence grader compares two schedules per CLAIM, asking whether the chargeback window was still open on the run that first counted it. It has no verdict to give about a single reading, and this one is shown as out of scope rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things. What it measures instead: one missed weekly run costs 6 claims worth 9,779.29 USD permanently, and the model still answered the degraded schedule at 97.50 pct.
The formulaWhat it computes
detection = raised-correct / readings that earned a raise. false-raise rate = quiet readings reported at REVIEW or above / quiet readings. Never averaged.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried position
100.0% detection · 1 more measured on this row
the fast tier, memory removed (THE CONTROL)
98.2% detection · 1 more measured on this row
the free carry-aware rule floor, no model
100.0% detection · 1 more measured on this row
the free one-reading rule floor, no model
100.0% detection · 1 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, sliced by whether the gold level is REVIEW or above.
These rates are UNKNOWN, on purpose
Whether 0 missed raises survives a harder mix. Both floors sit at 0, which says the corpus's raise threshold is within reach of a threshold rule -- a saturated grader tells you the corpus stopped discriminating on that axis, not that the direction is solved. The false-raise side DOES discriminate: 44 against 3.
Watch these
model.triage_detection
model.triage_page_precision
taxonomy.missed_incident
taxonomy.false_page
Alarm on
any missed raise at all. One is enough: the cost of that direction is the whole recoverable claim value, and the chargeback window keeps running.
How tight can the band be? 57 raised readings and 103 quiet ones, so one row is 1.75 and 0.97 points respectively. Nothing is tuned.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Always, and side by side. The two move in opposite directions and a single accuracy hides the trade entirely.
Do not use it
Detection alone, on an arm that raises nearly everything. The free one-reading floor scores 100.0 pct detection and it means almost nothing, because it also raises 44 false alarms in 103 quiet readings.
The readings whose answer cannot be derived from their own page
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineThe readings whose answer cannot be derived from their own page
On the 93 readings where the same page with no carried position produces a different answer, was the level right and was the count right? This is the subset the whole monitor argument rests on. ⚑ IT IS A MEASUREMENT, NOT A LABEL: tools/build_corpus.py flags it by computing the answer twice, and evals/check_labels.py verifies the flag in BOTH directions, so the subset can be neither padded nor quietly trimmed.
$0.00per 1,000 dealer readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 480 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DLR-0035-W2 -- a DECLINED_SUBSTANTIVE dealer, run 2 of 4
The book and the schedule
14 claims in the book, re-read whole. Run date 2026-07-13; the programme last read this dealer on 2026-07-06 and its book then ended at CLM-100249.
What the disposition carries
a DECLINED finding naming two claims, and a one-sentence reason that carries none of the words an audit keyword table would look for
Carried position, in the prompt
At the previous run on 2026-07-06 this dealer was reported WATCH for DUPLICATE_CLAIM. It has not been at REVIEW or above on any run up to and including that one. The claim book then ended at CLM-100249; every claim numbered above that is new since the last reading.
Free rule floor, no memory
INVESTIGATE, 4 new findings, UNCHANGED -- it has no memory at all
Free rule floor, same memory
REVIEW, 3 new findings, WORSENING -- right about everything except the sentence
The model, with the carried position
{"level": "WATCH", "pattern": "DUPLICATE_CLAIM", "new_findings": 1, "movement": "UNCHANGED"} -- CORRECT, and it is one of the readings that separates the model from the free floor on this page
Ground truth
WATCH, 1 new finding, UNCHANGED
Scored as
Both free floors put this dealer on an investigator's queue when the programme's own rule says it should not be there, and the model does not. It is the shape of row this kit was built to isolate and the only shape a model is paid for here -- and the model still gets two of these wrong elsewhere (DLR-0033-W4, DLR-0036-W4), so reading a declination is where it is better rather than where it is finished.
Grader
Verdict
Why
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
model: all four correct. Both free floors: level, count and movement MISS
DLR-0035-W2. Between run 1 and run 2 the investigator declined a finding naming two claims, with the note 'the sequence number falls inside the range as amended in June and the reading predates the amendment'. Rule 6 says a substantive reason removes those claims, so the key is WATCH / 1 / UNCHANGED and the model answered exactly that. The free carry-aware floor's keyword table carries no marker that fires on that sentence, defaults to ADMINISTRATIVE -- the conservative direction -- keeps both claims countable and answers REVIEW / 3 / WORSENING. The free one-reading floor is worse still at INVESTIGATE / 4.
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
a FALSE RAISE by both free floors; the model held it correctly
The key is WATCH, which is below the REVIEW threshold, so this reading sits in the QUIET denominator of 103 and not in the raised one. Both free floors report REVIEW or above and put a named dealer on an investigator's queue when the programme's own rule says it should not be there; the model does not. Across the whole set the carry-aware floor does this 3 times and the one-reading floor 44, against the model's 2. Neither direction is averaged into the other.
The readings whose answer cannot be derived from their own page
in scope -- one of the 93, and the model is right on it
This reading is memory-dependent by MEASUREMENT rather than by label: the same page answered with no carried position produces a different answer, which is how tools/build_corpus.py flags the subset and how evals/check_labels.py verifies it in both directions. The model scores 97.85 pct here, the free carry-aware floor 96.77, and the two arms with no memory at all collapse to 25.81 and 17.20.
The readings where an investigator's sentence decides the answer
in scope -- and this is the row the model's whole margin is made of
One of the 24 readings carrying a DECLINED finding, and one of the 5 the free carry-aware floor gets wrong. The model reads 22 of them against the floor's 21. ⚠︎ AND THIS EXACT READING IS ALSO THE ONE THAT FLIPPED: a single unplanned live call, on a prompt built by the identical code path, came back REVIEW / 3 / WORSENING -- the free floor's answer -- and called the same note administrative. The scored run's answer stands; the stability of it does not.
What one missed run costs, in claims and in dollars
not in scope -- this grader scores CLAIMS, not readings
The cadence grader compares two schedules per CLAIM, asking whether the chargeback window was still open on the run that first counted it. It has no verdict to give about a single reading, and this one is shown as out of scope rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things. What it measures instead: one missed weekly run costs 6 claims worth 9,779.29 USD permanently, and the model still answered the degraded schedule at 97.50 pct.
The formulaWhat it computes
hits among cells where gold.memory_dependent is true, divided by 93, reported separately for the level and the count.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried position
97.8% level accuracy · 2 more measured on this row
the fast tier, memory removed (THE CONTROL)
25.8% level accuracy · 2 more measured on this row
the free carry-aware rule floor, no model
96.8% level accuracy · 2 more measured on this row
the free one-reading rule floor, no model
17.2% level accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to cells the generator computed as memory-dependent and check_labels re-verified in both directions.
These rates are UNKNOWN, on purpose
How either arm behaves on a history longer than four runs. Every rule in this corpus resolves inside four, and the consecutive-runs rule is only just exercised. A dealer with two years of weekly readings is unmeasured.
Watch these
memory_level_accuracy_pct
memory_new_findings_accuracy_pct
Alarm on
any drop on the carry-aware floor. It is at 96.77 pct, so a first movement is a new failure rather than a worse rate.
How tight can the band be? 93 readings. One row is 1.08 points.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Reading whether the carried position is worth its complexity, rather than reading a headline. These 93 readings are 58 pct of the corpus and they are where every arm separates.
Do not use it
As the kit's headline. It double-counts cells already inside the level figure.
The readings where an investigator's sentence decides the answer
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineThe readings where an investigator's sentence decides the answer
On the 24 readings carrying a DECLINED finding, was the level right? Rule 6 says a declination removes its claims from the countable population only when the reason is substantive, and the reason is a sentence. ⚑ THIS IS THE ONLY GRADER ON WHICH A LANGUAGE MODEL HAS ANYTHING TO SELL IN THIS KIT.
$0.00per 1,000 dealer readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 480 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DLR-0035-W2 -- a DECLINED_SUBSTANTIVE dealer, run 2 of 4
The book and the schedule
14 claims in the book, re-read whole. Run date 2026-07-13; the programme last read this dealer on 2026-07-06 and its book then ended at CLM-100249.
What the disposition carries
a DECLINED finding naming two claims, and a one-sentence reason that carries none of the words an audit keyword table would look for
Carried position, in the prompt
At the previous run on 2026-07-06 this dealer was reported WATCH for DUPLICATE_CLAIM. It has not been at REVIEW or above on any run up to and including that one. The claim book then ended at CLM-100249; every claim numbered above that is new since the last reading.
Free rule floor, no memory
INVESTIGATE, 4 new findings, UNCHANGED -- it has no memory at all
Free rule floor, same memory
REVIEW, 3 new findings, WORSENING -- right about everything except the sentence
The model, with the carried position
{"level": "WATCH", "pattern": "DUPLICATE_CLAIM", "new_findings": 1, "movement": "UNCHANGED"} -- CORRECT, and it is one of the readings that separates the model from the free floor on this page
Ground truth
WATCH, 1 new finding, UNCHANGED
Scored as
Both free floors put this dealer on an investigator's queue when the programme's own rule says it should not be there, and the model does not. It is the shape of row this kit was built to isolate and the only shape a model is paid for here -- and the model still gets two of these wrong elsewhere (DLR-0033-W4, DLR-0036-W4), so reading a declination is where it is better rather than where it is finished.
Grader
Verdict
Why
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
model: all four correct. Both free floors: level, count and movement MISS
DLR-0035-W2. Between run 1 and run 2 the investigator declined a finding naming two claims, with the note 'the sequence number falls inside the range as amended in June and the reading predates the amendment'. Rule 6 says a substantive reason removes those claims, so the key is WATCH / 1 / UNCHANGED and the model answered exactly that. The free carry-aware floor's keyword table carries no marker that fires on that sentence, defaults to ADMINISTRATIVE -- the conservative direction -- keeps both claims countable and answers REVIEW / 3 / WORSENING. The free one-reading floor is worse still at INVESTIGATE / 4.
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
a FALSE RAISE by both free floors; the model held it correctly
The key is WATCH, which is below the REVIEW threshold, so this reading sits in the QUIET denominator of 103 and not in the raised one. Both free floors report REVIEW or above and put a named dealer on an investigator's queue when the programme's own rule says it should not be there; the model does not. Across the whole set the carry-aware floor does this 3 times and the one-reading floor 44, against the model's 2. Neither direction is averaged into the other.
The readings whose answer cannot be derived from their own page
in scope -- one of the 93, and the model is right on it
This reading is memory-dependent by MEASUREMENT rather than by label: the same page answered with no carried position produces a different answer, which is how tools/build_corpus.py flags the subset and how evals/check_labels.py verifies it in both directions. The model scores 97.85 pct here, the free carry-aware floor 96.77, and the two arms with no memory at all collapse to 25.81 and 17.20.
The readings where an investigator's sentence decides the answer
in scope -- and this is the row the model's whole margin is made of
One of the 24 readings carrying a DECLINED finding, and one of the 5 the free carry-aware floor gets wrong. The model reads 22 of them against the floor's 21. ⚠︎ AND THIS EXACT READING IS ALSO THE ONE THAT FLIPPED: a single unplanned live call, on a prompt built by the identical code path, came back REVIEW / 3 / WORSENING -- the free floor's answer -- and called the same note administrative. The scored run's answer stands; the stability of it does not.
What one missed run costs, in claims and in dollars
not in scope -- this grader scores CLAIMS, not readings
The cadence grader compares two schedules per CLAIM, asking whether the chargeback window was still open on the run that first counted it. It has no verdict to give about a single reading, and this one is shown as out of scope rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things. What it measures instead: one missed weekly run costs 6 claims worth 9,779.29 USD permanently, and the model still answered the degraded schedule at 97.50 pct.
The formulaWhat it computes
hits among cells where gold.prose_dependent is true, divided by 24.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried position
91.7% level accuracy
the fast tier, memory removed (THE CONTROL)
16.7% level accuracy
the free carry-aware rule floor, no model
87.5% level accuracy
the free one-reading rule floor, no model
16.7% level accuracy
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to the 24 readings whose Investigator Disposition carries a DECLINED finding.
These rates are UNKNOWN, on purpose
Whether the substantive / administrative distinction is the one a real programme draws, and whether real investigator prose separates as cleanly as sixteen written sentences. Both are unmeasured and the second is unmeasurable without a real book.
Watch these
prose_level_accuracy_pct
Alarm on
a rise in this figure with no change to evals/baseline.py's marker lists -- it would mean the corpus moved, not that the reader improved.
How tight can the band be? 24 readings. One row is 4.17 points, so nothing finer than about four points means anything here.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Deciding whether to pay for a model here at all.
Do not use it
As a general claim about reading comprehension. It is 24 readings drawn from sixteen sentences written by one author, and the same author wrote the keyword table they are scored against -- a conflict named in evals/baseline.py and in Data.breaks_on rather than left implicit.
What one missed run costs, in claims and in dollars
Catch a dealer's warranty claims turning into a pattern
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one missed run costs, in claims and in dollars
For every claim the weekly schedule counts as new somewhere, was the chargeback window still open on the run that first counted it -- and is it still open when one weekly run never happens? ⚑ THIS IS THE HALF OF A MONITOR THAT USUALLY GOES UNMEASURED. Four sibling kits in this estate carry state and none of them runs on a clock.
$0.00per 1,000 dealer readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/missed_run.py, in-process, free. It also writes data/gold-missed-w3.jsonl -- a SECOND answer key, because the correct answers under a fortnightly gap are different answers and not worse ones.
Every grader on these pages scored the same 480 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DLR-0035-W2 -- a DECLINED_SUBSTANTIVE dealer, run 2 of 4
The book and the schedule
14 claims in the book, re-read whole. Run date 2026-07-13; the programme last read this dealer on 2026-07-06 and its book then ended at CLM-100249.
What the disposition carries
a DECLINED finding naming two claims, and a one-sentence reason that carries none of the words an audit keyword table would look for
Carried position, in the prompt
At the previous run on 2026-07-06 this dealer was reported WATCH for DUPLICATE_CLAIM. It has not been at REVIEW or above on any run up to and including that one. The claim book then ended at CLM-100249; every claim numbered above that is new since the last reading.
Free rule floor, no memory
INVESTIGATE, 4 new findings, UNCHANGED -- it has no memory at all
Free rule floor, same memory
REVIEW, 3 new findings, WORSENING -- right about everything except the sentence
The model, with the carried position
{"level": "WATCH", "pattern": "DUPLICATE_CLAIM", "new_findings": 1, "movement": "UNCHANGED"} -- CORRECT, and it is one of the readings that separates the model from the free floor on this page
Ground truth
WATCH, 1 new finding, UNCHANGED
Scored as
Both free floors put this dealer on an investigator's queue when the programme's own rule says it should not be there, and the model does not. It is the shape of row this kit was built to isolate and the only shape a model is paid for here -- and the model still gets two of these wrong elsewhere (DLR-0033-W4, DLR-0036-W4), so reading a declination is where it is better rather than where it is finished.
Grader
Verdict
Why
The level, the pattern, the new-finding count and the movement, per reading, exact match against the computed answer key
model: all four correct. Both free floors: level, count and movement MISS
DLR-0035-W2. Between run 1 and run 2 the investigator declined a finding naming two claims, with the note 'the sequence number falls inside the range as amended in June and the reading predates the amendment'. Rule 6 says a substantive reason removes those claims, so the key is WATCH / 1 / UNCHANGED and the model answered exactly that. The free carry-aware floor's keyword table carries no marker that fires on that sentence, defaults to ADMINISTRATIVE -- the conservative direction -- keeps both claims countable and answers REVIEW / 3 / WORSENING. The free one-reading floor is worse still at INVESTIGATE / 4.
The two directions, counted apart: dealers that earned a raise and were reported quiet, and quiet dealers put on the queue
a FALSE RAISE by both free floors; the model held it correctly
The key is WATCH, which is below the REVIEW threshold, so this reading sits in the QUIET denominator of 103 and not in the raised one. Both free floors report REVIEW or above and put a named dealer on an investigator's queue when the programme's own rule says it should not be there; the model does not. Across the whole set the carry-aware floor does this 3 times and the one-reading floor 44, against the model's 2. Neither direction is averaged into the other.
The readings whose answer cannot be derived from their own page
in scope -- one of the 93, and the model is right on it
This reading is memory-dependent by MEASUREMENT rather than by label: the same page answered with no carried position produces a different answer, which is how tools/build_corpus.py flags the subset and how evals/check_labels.py verifies it in both directions. The model scores 97.85 pct here, the free carry-aware floor 96.77, and the two arms with no memory at all collapse to 25.81 and 17.20.
The readings where an investigator's sentence decides the answer
in scope -- and this is the row the model's whole margin is made of
One of the 24 readings carrying a DECLINED finding, and one of the 5 the free carry-aware floor gets wrong. The model reads 22 of them against the floor's 21. ⚠︎ AND THIS EXACT READING IS ALSO THE ONE THAT FLIPPED: a single unplanned live call, on a prompt built by the identical code path, came back REVIEW / 3 / WORSENING -- the free floor's answer -- and called the same note administrative. The scored run's answer stands; the stability of it does not.
What one missed run costs, in claims and in dollars
not in scope -- this grader scores CLAIMS, not readings
The cadence grader compares two schedules per CLAIM, asking whether the chargeback window was still open on the run that first counted it. It has no verdict to give about a single reading, and this one is shown as out of scope rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things. What it measures instead: one missed weekly run costs 6 claims worth 9,779.29 USD permanently, and the model still answered the degraded schedule at 97.50 pct.
The formulaWhat it computes
a claim is recoverable when the run that first counted it falls within 30 days of the claim's payment date. Compared between the weekly schedule and a schedule with run 3 removed. No model on either side.
The analysisWhat it actually did
Model
Result
the weekly schedule, as shipped
no headline metric on this row — it records claims recoverable 151 · usd recoverable 217830.89
one missed run -- run 3 never happens
no headline metric on this row — it records claims recoverable 145 · claims lost 6 · usd lost 9779.29
the model, ANSWERING over a missed run
97.5% level accuracy · 2 more measured on this row
the programme's FIRST run, whatever the schedule
no headline metric on this row — it records claims out of time 20 · usd out of time 30380.79
In operationWhat to monitor
Reference standard: data/gold.jsonl and data/gold-missed-w3.jsonl, both computed by src/programme.replay and both re-derived by evals/check_labels.py.
These rates are UNKNOWN, on purpose
Whether the COST of a missed run generalises. The 6 claims lost are entirely determined by which ones cross the parts clock in the skipped week, so a different week would lose a different number, and both constants behind the arithmetic are invented. What IS settled is the direction and the mechanism: nine days of headroom against a seven-day slip leaves no margin for a fortnight.
Watch these
claims_lost_to_the_missed_run
usd_lost_to_the_missed_run
Alarm on
any missed run at all, on the shipped constants. There is no partial credit: the headroom is 9 days and a skipped week is 7.
How tight can the band be? 6 claims lost is a count, not a rate. It is entirely determined by which claims cross the parts clock in the skipped week, which is a property of the generator.
Cadence: Re-run whenever a threshold in src/programme.py or a schedule in evals/run.py changes. Free.
The decisionWhen to reach for it
Use it
Before deciding a schedule. The arithmetic is settled by two constants and is not a matter of opinion.
Do not use it
As a claim about any real programme. Both constants are invented.
A living map of modern AI — kept current every morning