Catch a customer deduction before its window closes
Each week a customer's deduction sits on the ledger, and the clock to dispute it is already running against a date buried in the trade agreement. This app finds that date, reads the countdown, and tells the desk which claims must be filed now.
PresenterOpens the private repo. Visible to admins only.
For the deductions deskManufacturing & CPG · Payments & Fintech
Why it matters
Today's manual process, and the same job with the app
A deductions analyst at a manufacturer, closing the week's aging report before claims expire.
✕Today's manual process
1Pull the aging report and sort deductions by how long they've sat on the ledger.
2Open each trade agreement to find the dispute clause and work out which date starts the clock.
3Work out the days left manually, checking the clause against today's date for every claim.
4Miss one and a valid claim becomes a write-off with no appeal.
Every claim aged and read manually
✓With the app
1The queue is pulled automatically and every deduction is aged against its own window.
2The agreement is read for you and the anchor date its clause names is applied every time.
3The countdown is computed the exact days left, with no manual subtraction.
4The desk sees it in time so the claim is filed before the window closes.
Every claim aged against its own clause
See it work
One real case: what the app found, step by step
A deduction with no start date named in its clause, so the countdown runs from the delivery date, with two days left to file.
Catch a customer deduction before its window closesReference appBuilt to be shaped to your process
4
1The case, remembered No dispute lodged yet, and $9302 is on the line this week.
2The verdict, and its anchor FILE NOW: the countdown runs from the date of delivery.
3The dates on file The invoice and delivery dates the anchor choice was based on.
4The countdown, and the call Two days left, so the claim is filed now, not written off.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch a customer deduction before its window closes
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"Which deductions are oldest" is a SELECT and nobody needs a language model for it. The question a receivables desk actually has to answer is the next one: which of these claims still has a dispute window open, and is this the week that window closes. A deduction may be disputed only inside the window the customer's trade agreement grants -- and that window does not run from the day the customer took the money. It runs from an ANCHOR the agreement's dispute clause names: the proof of delivery on a shortage, the original invoice on a pricing claim, the remittance on a compliance fine -- and a customer whose clause names one anchor for everything overrides all of it. Past the close, an invalid deduction stops being a claim and becomes a write-off; no escalation reopens it. Today an analyst works the ageing report by eye and reads the agreements one at a time. An analyst working the deductions ageing report at week's end: opening each claim, finding the trade agreement, reading its dispute clause to work out which date the window runs from, subtracting the elapsed days from the window granted, and remembering which of these already has a dispute lodged against it.
Audience
A deductions or trade-claims analyst deciding what to lodge this week, and the receivables manager who signs off the write-offs at month end. Its own answer beats free code by 8.67 points on the countdown and loses to it outright on one clause phrasing, which is the point. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual deduction records
The corpus is 150 deduction records, 0.61 MB (json 1 · jsonl 1 · txt 150). A DEDUCTIONS QUEUE IS A LIST OF A MANUFACTURER'S CUSTOMERS, WHAT EACH OF THEM SHORT-PAID, WHY THEY SAID THEY DID IT, AND WHICH CLAIMS THE MANUFACTURER INTENDS TO FIGHT. It is commercially sensitive in both directions at once and there is no public one. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/deduction.step's output over the planted terms, so an anchor rule this fiddly cannot carry its author's misreading into the score. ⚠︎ AND THE RULES ARE INVENTED. The eight rules, the window lengths (30/45/60/90 days) and the anchor defaults reproduce no trade agreement, no customer deduction policy, no retailer's vendor guide and no regulator's guidance, and name none. The catalogue row this kit was built from (mfg:TPC-006) records the dispute window as an OPERATOR-SUPPLIED clock; that is implemented as Rule Q-8 and measured as context_incomplete_recall_pct rather than asserted. The rule text is reproduced in full on all 150 pages precisely so it can be read, disbelieved and replaced.
The corpus
The 150 deduction recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your deduction records. That is the whole change — there is no database to migrate.
One deduction record, as the model receives itDED-0001-R1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated deductions-queue snapshot for an AI
use-case kit; it reproduces no trade agreement, no customer's deduction policy, no
retailer's vendor guide and no real company. The dispute terms are ILLUSTRATIVE.
Deduction
----------------------------------------------------------------
Customer : CUST-4021 (Harbourgate Markets)
Deduction reference : DED-0001
Deduction type : SHORTAGE
Amount deducted : 1,248.68 USD
Reason code as taken : 812 -- quantity short at receipt
Original invoice : INV-748230
Purchase order : PO-5162616
Ship-to : DC-4 Pinewell
Dispute Terms
----------------------------------------------------------------
Trade agreement : TA-4021-2025
Dispute window granted : 60 days
Dispute clause (verbatim) : The customer will consider a written dispute lodged within the
window above. Partial disputes are permitted; the window is
not extended by a partial dispute.
Rule Q-1 A deduction may be disputed only INSIDE the dispute window the customer's trade agreement
grants. Once that window closes an undisputed deduction is a WRITE-OFF: it is not
recoverable by any later process, and no escalation reopens it.
Rule Q-2 The window runs from an ANCHOR DATE set by the agreement's dispute clause -- not from the
age of the deduction on the ledger. Read the clause before counting anything.
Rule Q-3 Where the clause is silent about a deduction type, the default anchors are: SHORTAGE and
Abridged — the file continues.
The outcomeWhat a good result looks like
Every deduction whose window closes before the next scheduled run is on the worklist with the correct number of days left, and exactly one dispute exists per deduction -- lodged on the run that first saw the deadline inside the watch interval, not re-lodged on every run after it.
And when it cannot
Two ways, and they cost different things. A MISSED filing is a write-off: the window shuts and the money is gone. A DUPLICATE filing is two claims for one charge on the customer's desk, which gets both bounced as conflicting and, on customers who treat a resubmission as a new claim, restarts their review clock. The scored run made 1 of 34 and 0 of 116 -- one missed filing worth $1,428.34, and no duplicates.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Every deduction on your book is disputed from the same date, under one standard clause — a SQL query, or the b000 floor in this repo Days on the ledger against one threshold is arithmetic. On the slice of THIS corpus where that is true, b000 scores 100.0 pct of countdowns exactly right for nothing at all.
Your agreements set the anchor with a coded field, or your clauses come from a small closed set of templates — the b002 floor: a lookup on your codes, or a keyword table over your templates The model's entire margin here is reading clause prose that a table cannot parse. Take the prose away and there is nothing left to buy: the floor already reaches 88.0 pct on the countdown for $0.00, and it is PERFECT on the reverse-collision clause where the model is not.
Your dispute clauses are negotiated per customer and say what they say — this kit, with the carried state That is exactly the corpus these figures were measured on. The countdown goes 96.67 pct against 88.0, the missed-filing bill goes $1,428.34 against $19,683.89, and the at-risk total lands 0.22 pct out against 12.97.
You want the watch to actually wake up on a schedule — your own scheduler -- cron, Airflow, Temporal, whatever already runs -- and an alarm on a run that did not fire evals/run.py is INVOKED, not woken, and evals/cadence.py measures what that gap costs: skipping ONE weekly run loses 30 of 34 filings permanently, worth $453,621.42. That is the largest number this kit publishes and it is about the schedule, not the model.
At a glanceHow the whole thing runs
97%dispute window accuracy pct
4,974 msp50, end to end
$3.00per 1,000 deduction records · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a customer deduction before its window closes14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own queue, or replace data/corpus/*.txt and data/gold.jsonl directly. The measured figures on this page do not travel with your corpus, and two of them are corpus properties rather than model properties.Corpus lens →
When is this the wrong choice?
Avoid: AVOID this kit entirely. Paying a model per row per run to do one subtraction is the most expensive way to get an answer you already have. That is the case against the best-fitting scenario (“Every deduction on your book is disputed from the same date, under one standard clause”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
AN AGREEMENT CLAUSE THAT NAMES A DATE THE LEDGER DOES NOT CARRY. Nine readings carry "the date the carrier's signed receipt REACHED US", which is the receipt's arrival, not the delivery -- a different date, and one this corpus never prints. 6 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER "THE DATE THE CARRIER'S SIGNED RECEIPT REACHED US" MEANS THE DELIVERY. Five of the model's seven anchor errors turn on that one clause and on all five the model's reading is arguably better than the key's. 9 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-deduction-age. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces: the 150 snapshots and the answer key (regenerated from the seed in well under a second), all fifteen pre-flight assertions, all three free floors scored end to end, the whole cadence and missed-run arithmetic, the wiring stub, and the local UI at 127.0.0.1:9016 including what run r001-deduction-age recorded for every reading. What it CANNOT reproduce without a key is a model column of its own -- and every figure this kit publishes about the SCHEDULE, which is the largest number on the page, needs no key at all.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
4,974 msp50, end to end
9,795 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed to a posture, an anchor, a whole-day countdown and a file/hold call. Splitting the snapshot into sections, dropping the one section no field asks for, computing the three candidate elapsed-day figures and advancing the carried state all happen outside this measurement and cost no network at all. p95 (9.8s) is 2.0 times p50 (5.0s), and the spread is the CLAUSE rather than the page: every snapshot in this corpus is within 9 pct of every other in size (4060 to 4423 bytes), while a clause that sets an anchor in words naming none of the three candidate dates is a reading task with three plausible answers. The longest reply ran to 14538 output tokens.
Current processWhat it replaces
An analyst working the deductions ageing report at week's end: opening each claim, finding the trade agreement, reading its dispute clause to work out which date the window runs from, subtracting the elapsed days from the window granted, and remembering which of these already has a dispute lodged against it.
Where it is not good enough
⚑ THE MODEL WINS THE HEADLINE AND THE PAGE STILL LEADS WITH WHAT THAT COST AND WHERE IT IS WRONG. The discriminator is the countdown, exact whole days: 96.67 pct against the strongest free floor's 88.0 pct -- 145 readings of 150, a 8.67-point margin bought for $0.00300 a reading. ⚠︎ ALL FIVE OF THE MODEL'S ANCHOR ERRORS SIT ON ONE SENTENCE, NOT FIVE JUDGEMENTS. The clause "Whichever date the deduction was taken, the window for every category runs from the date the carrier's signed receipt REACHED US" scores 4 of 9 while every other clause phrasing in the corpus scores 100 pct -- and on all five the model's reading is arguably the better one. That clause names the date the receipt reached us, which is a different date from the delivery and is genuinely NOT ON THE PAGE, so the model applies Rule Q-8 and abstains. The answer key assumes the delivery date. THE CLAUSE WAS NOT REWRITTEN AND THE RUN WAS NOT RE-FIRED; the lower figure is published and the ambiguity is recorded here. ⚠︎ AND ONE OF THOSE FIVE IS THE RUN'S ONLY MISSED FILING, WORTH $1,428.34: DED-0030-R2, where the guardrail abstained in the safe direction and a live claim with 7 days left was not lodged. Abstention is not free on a hard clock. ⚠︎ TWO FURTHER ANCHOR CELLS WERE SCORED WRONG FOR A VOCABULARY SLIP, NOT A MISREAD: the reply said ORIGINAL_INVOICE_DATE where the closed list says INVOICE_DATE. The reasoning was right, the countdown was right, and both are counted as misses anyway -- that is a prompt defect this kit owns, and it is 2 of the 7 anchor errors. ⚠︎ AND THE COMPARISON IS ASYMMETRIC IN THE FLOOR'S FAVOUR ON EXACTLY THOSE ROWS. evals/baseline.py keys on the phrase "signed receipt" and gets all nine of them right BECAUSE IT DOES NOT READ THE SENTENCE. A keyword table cannot misread a clause; it can only match or fail to.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1json1
50 open deductions on one manufacturer's receivables desk, each re-read whole at three weekly runs
It produces a countdown and a file/hold recommendation for a deductions analyst to action. It never lodges a dispute, closes a window, writes off a claim or credits an account, and there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is four scalars written by src/state.py from the arithmetic and never from the model's reply — whether a dispute is already lodged, last run's countdown, the posture last reported, how many runs have seen it — so a wrong reading is a wrong worklist row this run, never a corrupted history for the next one. The clock is the second: Rule Q-1 makes a window that closes unfiled a permanent write-off, and a run that does not happen is not a late fix — measured at 30 of 34 last-chance filings lost for good, $453,621.42, if a single weekly run is skipped.
⚠︎ AND THE SHARPER CADENCE FINDING RUNS THE OTHER WAY: widening the interval from weekly to MONTHLY, over the same 17-week horizon, loses only ONE claim, $3,278.53, while cutting the provider calls 4.5x, 900 to 200. The risk on this kit's schedule is a run that fails to fire, not one that fires too rarely.
⚠︎ THE FLOOR STATION CARRIES THE FREE NUMBER'S OWN CONTRADICTION, NOT A CLEAN LOSS. b000, the ageing report a desk already runs, is 100.0 pct right on the 51 readings whose anchor is the remittance date and 0.0 pct right on the other 99 — an aggregate of 34.0 pct that is an average of a perfect score and a zero, not a measure of the rule. Read the clause for free (b002) and the disjoint disappears: 92.68 / 83.33 / 82.35 pct across the same three anchors, still $0.00.
⚠︎ AND THE MODEL'S OWN ERRORS ARE FIVE READINGS OF ONE CLAUSE, NOT A SPREAD OF JUDGEMENT. "Whichever date the deduction was taken, the window for every category runs from the date the carrier's signed receipt REACHED US" names a date genuinely not on the page; Rule Q-8 makes the model abstain (CONTEXT_INCOMPLETE) on all five readings that carry it, and one of the five — DED-0030-R2 — is the run's only missed filing, $1,428.34. Two further anchor cells are wrong for a vocabulary slip rather than a misread: the model answered ORIGINAL_INVOICE_DATE where the closed list reads INVOICE_DATE, with the countdown and the posture both correct underneath. The clause was not rewritten and the run was not re-fired.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the carried state
src/state.py
What is carried between scheduled runs, and how it is worded to the model. Four scalars today. ⚠︎ ONE OF THEM IS A KNOWN TRADE-OFF: dropping days_to_deadline_last_run would force every reading to re-derive the countdown from the clause, which is slower and possibly more accurate on the anchor -- the control arm scores 96.67 pct on the anchor against this arm's 95.33. Two cells; not a claim.
the cadence
src/deduction.py
CADENCE_DAYS, and with it BOTH bands -- FILE_BAND_DAYS is the interval and WATCH_BAND_DAYS is twice it, so changing the schedule changes what the postures mean as well as what the run costs. Every figure on this page is per seven days.
the anchor rule
src/deduction.py
TYPE_DEFAULT_ANCHOR and RULE_TEXT. The rule text is reproduced in full on all 150 pages precisely so it can be read, disbelieved and replaced with the agreements you actually signed.
the clause floor
evals/baseline.py
POD_WORDS, INVOICE_WORDS and DEDUCTION_WORDS, and the order they are tried in -- the keyword table the model is measured against. Widen it and the floor gets stronger, which is the direction this kit wants: a floor built to lose measures nothing.
the corpus
tools/build_corpus.py
The deductions, the clause pools and the seed. Keep the seven section headings or src/segment.py's assertion refuses to start.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 50 open deductions x 3 weekly runs = 150 snapshots from a fixed seed (SEED = 990823). 13 of the 50 customers carry a dispute clause that OVERRIDES the type default anchor, which moves 39 readings off it -- that is the experiment. The gold labels are src/deduction.step's output over the planted terms, never typed.
the dispute arithmetic
src/deduction.py
The rule as pure code: the anchor resolved from the clause override or the type default, the window granted minus the days elapsed from that anchor, and the two bands -- both of which ARE the cadence. No model, no judgement. ⚠︎ The eight rules, the window lengths and the anchor defaults are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. Four scalars per deduction (whether a dispute is already lodged, last run's countdown, the posture last reported, how many runs have seen it), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. ⚠︎ One of the four is double-edged and the kit measures it rather than defending it: carrying last week's countdown lets a reading reach the right number by subtracting 7 without re-reading the clause.
the section splitter
src/segment.py
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Customer Contact -- the customer's A/P analyst, her direct dial and her work email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading.
the watch
src/watch.py
One deduction, one scheduled run, one call. Parses the window, the three candidate anchor dates and their elapsed-day figures off the page with a regex (the model is never asked to subtract two dates), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
the local UI
src/app.py
One deduction, one scheduled run, its carried state and its verdict, on 127.0.0.1:9016. Renders with no key. It shows the carried sentence verbatim, the free floor's answer beside the model's, and a second button that replays what run r001-deduction-age actually answered, straight off the committed result file, labelled as a replay.
the three free floors
evals/baseline.py
b000 the deductions ageing report every ERP prints -- days since the deduction posted against one assumed 60-day window; b001 the same rule given the carried state, which isolates what memory alone buys; b002 the strongest free version -- carried state, the window read off the page, the anchor resolved by an ordered keyword table over the clause with a fall-back to the type default, and an abstention where nothing is on file. 0 calls, $0.00, all three scored through the identical scorer.
the cadence arithmetic
evals/cadence.py
⚑ THE HALF OF A MONITOR THE ACCURACY TABLES CANNOT SEE, and it costs $0.00. It answers two questions from the deduction model with no key and no network: what does skipping ONE scheduled run cost (30 of 34 filings lost permanently, $453,621.42), and what does the interval itself cost (weekly to monthly loses one claim worth $3,278.53 and cuts the bill 4.5x). A reading that never happened produces no cell for a scorer to grade, which is why this is a separate file rather than a column.
the scorer
evals/scoring.py
Exact match per cell against the computed gold, split ways an average would hide: the four fields, the two filing directions counted apart AND in dollars, the memory-dependent subset, the context-incomplete recall, and -- the cut this kit exists for -- every rate again PER ANCHOR and PER SCHEDULED RUN. No judge model.
the pre-flight
evals/check_labels.py
Fifteen things that must be true before a run may spend: the seven sections parse, every sequence is complete and gap-free, no run-varying section says a dispute was lodged and the non-varying ones are byte-identical across runs (the experimental control), no page carries a deadline field or a countdown, THE ANCHOR IS NOT DERIVABLE FROM THE DEDUCTION TYPE, every anchor slice is big enough to report, the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path writes off or credits a deduction, and the answer key replays from src/deduction.step.
Where it breaks at scale
⚑ THE CADENCE IS THE SCALING VARIABLE AND IT IS NOT THE CORPUS SIZE. One run of this watch is one call per OPEN DEDUCTION, and the watch wakes weekly, so a desk holding 4,000 open deductions pays 4,000 calls a week whether or not anything moved -- which is why this kit's cost lens is quoted per row per run rather than per document. ⚑ AND THE THING THAT BREAKS FIRST IS NOT THE VOLUME, IT IS A RUN THAT DID NOT FIRE, MEASURED: skipping ONE weekly run loses 30 of the 34 last-chance filings PERMANENTLY, worth $453,621.42 (evals/cadence.py, $0.00). Widening the cadence from weekly to monthly, by contrast, loses ONE claim worth $3,278.53 and cuts the call count from 900 to 200 over the same horizon. The schedule risk on this task is reliability, not frequency. TWO THINGS BREAK BEFORE THE CALL COUNT DOES. First, data/state.json is one file replaced atomically -- correct for one writer and not a concurrency model, and a desk with two watches running is two writers; here a lost write forgets which disputes are already lodged, and the next run files a second one against every claim in the queue. Second, a run that does not happen is not an error anywhere in this kit: evals/run.py is INVOKED, it is not woken, and nothing here detects a missed run, back-fills it or marks its readings late. The kit measures what that costs and does not fix it, because fixing it means putting a scheduler in a repository whose product is being a folder of readable Python.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
DED-0026-R2, the whole argument on one screen, and the deduction was chosen by reading the answer key for the case where the free floor and the truth disagree MOST -- not by looking for a flattering frame. Its dispute clause reads "Disputes must quote the original invoice number and be submitted inside the window stated above", which says what a dispute must CONTAIN and sets no anchor at all -- so Rule Q-3 governs and a DAMAGE claim runs from the proof of delivery. The free keyword floor sees the word invoice, anchors there, computes -1 day and reports EXPIRED with nothing to file. The answer key -- and run r001-deduction-age -- say 2 days left and say THIS is the run that has to file. That is $9,302.06 the floor writes off on a claim still alive. ⚠︎ THE MODEL COLUMN HERE IS REPLAYED FROM THE COMMITTED RESULT FILE, NOT A LIVE CALL, and the column header says so in as many words: replaying the scored run's own recorded answer is evidence, staging one would not be, and the only thing that tells those apart is the label.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page on DED-0030-R2, which is the scored run's ONLY missed filing and its most expensive single error. The clause anchors every category on "the date the carrier's signed receipt reached us" -- a date that is genuinely not among the three on the page -- so the model applies Rule Q-8 and abstains: CONTEXT INCOMPLETE, no anchor, no countdown. The free floor keys on the phrase "signed receipt", anchors on the proof of delivery, and gets it right: 7 days left, FILE IT NOW. $1,428.34 the model would have let lapse and free code would have saved. This frame is a replay of the committed run, labelled as one, and it is on the page because a kit that only screenshots the rows it wins is an advert.failureOpen full size →The same page with NO API_KEY configured. It does not error and it does not go blank: the snapshot, the carried state, the three candidate anchor dates, the withheld-section list and the entire free floor are computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called.failureOpen full size →Before anything is asked. The two things this page has to get right are already on it: the carried state, verbatim as it goes into the prompt, and the list of which sections left the machine and which did not -- Customer Contact, a named individual at the customer with her direct dial and work email, marked WITHHELD rather than silently absent. A page that simply does not mention them cannot be told apart from one that quietly sent them.failureOpen full size →
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150deduction records
0.61 MiBjson 1 · jsonl 1 · txt 150
50deductions (3 weekly runs each) · p50 3 chars
$0.00setup · 0.0s
How it is cutWhat one deductions (3 weekly runs each) is
No split, and no chunking. The unit is a DEDUCTION -- three consecutive weekly runs processed strictly in order, because week 3 has to know whether week 1 lodged a dispute. Each snapshot goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step -- the population is re-read whole on each scheduled run. tools/build_corpus.py writes 150 documents and the answer key from a fixed seed in well under a second with no clock read and no model called; nothing is embedded, ranked or cached.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every customer, deduction, trade agreement, dispute clause, window length, amount and activity-log line is invented here. Verified against the repository's own LICENSE file on 2026-08-23.
Bring your ownBring your own deduction records
Point tools/build_corpus.py at your own queue, or replace data/corpus/*.txt and data/gold.jsonl directly. Keep the seven section headings in src/segment.py::SECTIONS -- the parser asserts all seven in every document before a run may spend -- and keep the Dispute Terms block, because the clause the model reads is ON THE PAGE rather than baked into the prompt, which is what lets you swap your own agreements in without touching src/prompt.py. gold.jsonl needs one row per snapshot carrying window_days, days_elapsed_from_anchor, the resolved anchor, context_complete and the four answers; evals/check_labels.py re-derives the answers from src/deduction.step and refuses if they disagree, so a hand-written key is caught rather than trusted.
⚠︎ And what stops being true when you do: The measured figures on this page do not travel with your corpus, and two of them are corpus properties rather than model properties. EVERY figure here is per a SEVEN-DAY cadence: both posture bands are DEFINED as the cadence, so a desk on a daily or monthly watch is scoring a different question with the same words. And the population here does not change between runs, which is not what a real deductions queue does. The per-anchor split is the figure most likely to move: it depends entirely on how many of YOUR agreements carry an overriding clause, and here that is 13 of 50.
What breaks it
⚠︎ THE CORPUS DELIBERATELY WITHHOLDS FILING EVIDENCE FROM THE PAGE, AND THAT FLATTERS MEMORY. A real deductions log carries "customer's claims desk acknowledged our packet" about half the time -- and with that on the page, "the carried state helped" would be unfalsifiable, because a stateless reading could infer the filing. So the activity log is restricted to portal reconciliations, reason-code remaps, document chases and receipt arrivals, and evals/check_labels.py asserts that over the three sections that vary between runs. The cost is that the control arm's 72.67 pct posture accuracy against this arm's 96.67 is an UPPER BOUND on what memory is worth; against a log that leaks, the gap would be smaller and nothing here measures by how much.
⚠︎ A CORPUS WHOSE CLAUSES ARE SEPARABLE BY KEYWORD MEASURES THE TEMPLATES, NOT THE WORK. The clause pools carry deliberate near-collisions in BOTH directions -- a delivery-anchored clause that says "The invoice date is immaterial for this purpose", a delivery-anchored clause phrased as "the consignee's stamped paperwork reaches our claims inbox" with no delivery vocabulary at all, and two clauses that set NO anchor while mentioning a proof of delivery and an invoice number respectively. The free keyword floor loses 18 of its 20 window errors on exactly those two decoys, and it BEATS the model 9-4 on the reverse collision. What remains untested is the class, not the instance: a phrasing whose anchor this corpus does not know about would reproduce the same defect and the same silence.
AN AGREEMENT CLAUSE THAT NAMES A DATE THE LEDGER DOES NOT CARRY. Nine readings carry "the date the carrier's signed receipt REACHED US", which is the receipt's arrival, not the delivery -- a different date, and one this corpus never prints. The answer key treats it as the delivery date and the model abstains under Rule Q-8 on five of the nine. The model's reading is arguably better; the key was not changed and the run was not re-fired.
A RECEIVABLES SYSTEM THAT RENAMES ITS BLOCKS. src/segment.py recognises seven exact headings; an upgrade that renames half a report leaves nothing to parse. The kit refuses rather than truncating -- evals/check_labels.py asserts all seven in all 150 documents -- and that same condition is what the privacy guard is red-proven against: with the guard, 0 of 150 leak the customer contact section; without it, 150 of 150 do.
A DEDUCTION READ OUT OF ORDER, OR A RUN SKIPPED. The filing history is carried, so a watch that read week 3 before week 2 would not know a dispute existed and would lodge a second one. evals/check_labels.py asserts every sequence is complete and gap-free; nothing anywhere detects a MISSED scheduled run, and evals/cadence.py measures what one costs (30 of 34 filings, $453,621.42) rather than fixing it.
A POPULATION THAT CHANGES BETWEEN RUNS. All 50 deductions are open at all three runs in this corpus, so the arms are comparable; a real queue takes on new deductions daily and closes others, and nothing here measures what a first sighting mid-window, or a claim that settles between runs, does to the carried state.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
174
not measured
instruction
2,728
not measured
carried state
149
not measured
Synthetic Record
332
not measured
Deduction
474
not measured
Dispute Terms
2,188
not measured
Aging Position
570
not measured
Claim Activity Log
213
not measured
Deductions Desk Notes
199
not measured
Total
1,705
This is the cost lesson as arithmetic: of the 7,027 characters assembled, 2,902 are instructions — 41% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
TWO MESSAGES, JOINED IN SEND ORDER. The provider receives a system message and a user message; prompt_verbatim prints both, separated by a blank line, because publishing only the second would be publishing most of a prompt. Replayed from the kit's own src/prompt.build() for DED-0026-R2 with the carried state that reading was actually given, rather than logged by the run. What makes the replay checkable is that the section list it produces is the list the run recorded in sections_used -- Synthetic Record, Deduction, Dispute Terms, Aging Position, Claim Activity Log, Deductions Desk Notes -- and that src/prompt.build() is the only thing in the kit that assembles a prompt, so there is no second path a run could have taken. ⚠︎ WHAT IT IS *NOT* CHECKED AGAINST, SAID PLAINLY: the run's own prompt_parts records ONE size, 7080 characters, and it belongs to whichever reading finished first under 12 concurrent workers -- not to DED-0026-R2, whose assembled user message is 7037 characters. The harness stores one exemplar, not one per reading, so no per-reading size comparison is available and none is claimed.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled deduction-queue watch. You apply a written dispute rule to one customer deduction at one scheduled run. You answer with one JSON object and no other text.
You are the scheduled deduction-queue watch for one manufacturer's receivables desk. It wakes every
seven days, re-reads every deduction still open in the queue, and reports each one. You are reading
ONE deduction at ONE scheduled run.
The dispute rules, the window granted and the trade agreement's dispute clause are reproduced in the
snapshot below. Apply them exactly as written. A dispute window is a HARD DEADLINE: once it closes,
an undisputed deduction is a write-off and no later run recovers it. You cannot see the earlier runs;
what is known about them is stated under "Carried state" and is the only history available to you.
Do not assume anything about earlier runs beyond it.
How to count:
- First decide the ANCHOR. The window does NOT run from the age of the deduction on the ledger. It
runs from whichever date the agreement's dispute clause names (Rule Q-2). If the clause names one
anchor for all deductions, that overrides the type default for this deduction too (Rule Q-4). If
the clause is silent about which date starts the clock -- for instance it only says what must be
attached to a dispute, or which channel to use -- then the type default applies (Rule Q-3).
- Then read the "days elapsed" figure printed beside THAT anchor date. It is already computed for
you; do not re-derive it from the calendar.
- days_to_deadline = the window granted MINUS that elapsed figure. It may be negative.
- If the agreement carries no dispute clause, or the anchor date the clause needs is NOT ON FILE,
the deduction is CONTEXT_INCOMPLETE: posture CONTEXT_INCOMPLETE, anchor NONE, days_to_deadline
null, file_dispute NO. Never assume a default window and never substitute another date
(Rule Q-8).
Answer with a single JSON object and nothing else:
{"posture": "CLEAR|WATCH|FILE_NOW|FILED|EXPIRED|CONTEXT_INCOMPLETE",
"anchor": "INVOICE_DATE|POD_DATE|DEDUCTION_DATE|NONE",
"days_to_deadline": <whole days left on the window, negative if it has closed, null if unknown>,
"file_dispute": "YES|NO",
"rationale": "one sentence, naming the anchor you chose and the rule that set it"}
Precedence for "posture", applied in this order: CONTEXT_INCOMPLETE beats everything; then FILED if
the carried state says a dispute has already been lodged, even if the window has since closed
(Rule Q-7); then EXPIRED if days_to_deadline is negative; then FILE_NOW if days_to_deadline is 7 or
fewer, because the deadline will pass before the next scheduled run; then WATCH if it is 14 or fewer;
otherwise CLEAR.
"file_dispute" is YES only on the run that LODGES the dispute -- the posture is FILE_NOW and the
carried state does not already say one was lodged. On every later run it is NO (Rule Q-6).
Carried state
----------------------------------------------------------------
No dispute has been lodged against this deduction yet. At the previous scheduled run, 9 day(s) remained on the dispute window. It was reported WATCH.
Deductions queue snapshot
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated deductions-queue snapshot for an AI
use-case kit; it reproduces no trade agreement, no customer's deduction policy, no
retailer's vendor guide and no real company. The dispute terms are ILLUSTRATIVE.
Deduction
----------------------------------------------------------------
Customer : CUST-4315 (Northbank Cash and Carry)
Deduction reference : DED-0026
Deduction type : DAMAGE
Amount deducted : 9,302.06 USD
Reason code as taken : 834 -- product damaged in transit
Original invoice : INV-882678
Purchase order : PO-5456732
Ship-to : XD-3 Larkmere
Dispute Terms
----------------------------------------------------------------
Trade agreement : TA-4315-2025
Dispute window granted : 90 days
Dispute clause (verbatim) : Disputes must quote the original invoice number and be
submitted inside the window stated above. The customer's
claims portal is the only accepted channel.
Rule Q-1 A deduction may be disputed only INSIDE the dispute window the customer's trade agreement
grants. Once that window closes an undisputed deduction is a WRITE-OFF: it is not
recoverable by any later process, and no escalation reopens it.
Rule Q-2 The window runs from an ANCHOR DATE set by the agreement's dispute clause -- not from the
age of the deduction on the ledger. Read the clause before counting anything.
Rule Q-3 Where the clause is silent about a deduction type, the default anchors are: SHORTAGE and
DAMAGE from the PROOF OF DELIVERY date; PRICING and TRADE_PROMO from the ORIGINAL INVOICE
date; COMPLIANCE_FINE and UNSALEABLES from the DEDUCTION (remittance) date.
Rule Q-4 A clause that names ONE anchor for all deductions OVERRIDES Rule Q-3 for every type on
that customer, whichever type this deduction happens to be.
Rule Q-5 Days to deadline = the window granted MINUS the days elapsed from the anchor date. Both
numbers are printed on this page; do not re-derive the elapsed days from the calendar.
Rule Q-6 A dispute is filed ONCE per deduction, on the first scheduled run at which the deadline
falls inside the watch interval. A later run reports it already filed; it does not file a
second dispute.
Rule Q-7 A deduction whose dispute was filed inside the window stays FILED after the window closes.
Only an UNDISPUTED deduction past its deadline is EXPIRED.
Rule Q-8 Where the agreement carries no dispute clause, or the anchor date the clause needs is not
on file, the deduction is CONTEXT_INCOMPLETE: no deadline is computed, nothing is filed,
and NO default window and NO substitute anchor may be assumed.
Aging Position
----------------------------------------------------------------
Watch run : 2026-03-09 (scheduled run 2 of this deduction)
Watch cadence : every 7 days (Mondays, 08:00 local)
Previous scheduled run : 2026-03-02
Candidate anchor dates, and days elapsed as at this run
Original invoice date : 2025-12-08 ( 91 days elapsed )
Proof of delivery date : 2025-12-11 ( 88 days elapsed )
Deduction / remittance date : 2025-12-25 ( 74 days elapsed )
Days on the ledger since taken : 74
Claim Activity Log
----------------------------------------------------------------
Entries raised in this window only.
2026-03-05 customer's claims portal reconciled against our open items, this line still open
Deductions Desk Notes
----------------------------------------------------------------
Customer's claims desk has asked that we raise nothing further on this deduction while their own review is open.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"posture": "FILE_NOW",
"anchor": "POD_DATE",
"days_to_deadline": 2,
"file_dispute": "YES",
"rationale": "The dispute clause names no start date, so Rule Q-3's default anchor for DAMAGE applies (POD date); 90-day window minus 88 days elapsed leaves 2 days, so this run files the dispute."
}
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a customer deduction before its window closes — 150 deduction records. One model answered, and every answer was then graded Four different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the deduction model -- the resolved anchor, the window granted and the evidence dates -- and re-derived from src/deduction.step by evals/check_labels.py before any run may spend. The countdown is compared as a whole number of days, and a null countdown MATCHES a null one, because "no deadline can be computed" is the correct answer on 22 readings rather than a missing value. No model grades anything, here or anywhere in this kit.
150deduction records
150source documents
1model tier
4grading methods
MeasurementsWhat was measured
COUNTED145 · 145 · 132 · 51 / 150dispute window accuracy pct — readings, exact whole-day match on the countdown, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED143 · 145 · 132 / 150anchor accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED145 · 109 · 139 · 88 / 150posture accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED149 · 147 / 150file dispute accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED1 / 34missed filing pct — readings whose correct answer is to LODGE the dispute -- the last run that can, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 3 / 116duplicate filing rate pct — readings that must NOT lodge a dispute, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED36 · 0 / 39memory posture accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED36 / 39memory window accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 / 22context incomplete recall pct — readings with no dispute clause on file or no anchor date on file, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED13 / 13expired recall pct — readings whose window has already closed undisputed, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 1 / 1at risk usd error pct — queue total in USD of deductions reported as last-chance or written off, against the key, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 · 150 / 150answered pct — calls, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 150unparsed replies — calls, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is asserted: evals/check_labels.py replays all 50 deduction chains through src/deduction.step and requires the committed gold to match on all four fields, so the key and the arithmetic cannot drift apart. What is NOT validated is one clause the KEY reads differently from the model -- see Business.not_good_enough and the CLAUSE_NAMES_AN_ABSENT_DATE taxonomy row, where the model's reading is arguably the better one on all five affected readings.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One deduction record
1,000 deduction records
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.003002
$3.00
28%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.001201
$1.20
28%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.052869
$52.87
32%
Same work, 44× the bill
The same deduction records, the same tokens — only the rate card changed. And across all 3 cards between 28% and 32% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE CADENCE -- it is the only knob here that moves the bill by a whole multiple. But ⚑ THIS KIT MEASURED WHAT MOVING IT BUYS AND THE ANSWER IS ALMOST NOTHING: weekly to monthly is 4.5x cheaper and loses one claim worth $3,278.53. What actually costs money is a run that does not fire -- $453,621.42 for a single skipped week. Spend on the scheduler, not on the interval.
Rates checked 2026-08-18. The provider that actually ran all 306 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors it is compared against, the pre-flight, or evals/cadence.py -- which produces the largest number on this page. The cost of the RUN is a different figure and lives in lens 07; pricing the ruler in the units of the thing it measures is the trap this field exists to name.
The gradersFour ways to grade
⚑ THE MODEL BEATS THE STRONGEST FREE FLOOR ON THE DISCRIMINATOR BY 13 READINGS, AND THE FLOOR IS NOT A STRAW MAN. Dispute window, exact whole days: 96.67 pct against 88.0. Anchor: 95.33 against 88.0. Posture: 96.67 against 92.67. Missed filings: 1 of 34 ($1,428.34) against 3 of 34 ($19,683.89). Duplicate filings: 0 against 2. The queue's at-risk value is $651,445.25 and the model reports $650,016.91 (0.22 pct out) against the floor's 12.97 pct. ⚑ AND THREE FLOORS RATHER THAN ONE, BECAUSE ONE CANNOT SEPARATE THE MEMORY FROM THE READING. b000 is the deductions ageing report an ERP already prints -- days on the ledger against one assumed 60-day window, no memory, no clause read: 32.67 pct posture, 82.67 pct file/hold, 4 duplicate filings and 22 missed ones worth $256,640.23. b001 is that same rule GIVEN the carried state: duplicates fall 4 -> 0 and posture goes 32.67 -> 58.67. That jump is what MEMORY ALONE buys with no model anywhere near it. b002 adds reading the clause, and it is worth 54.0 points of window accuracy on top -- the largest single movement on this page, and still free. ⚑ READ THE THREE SLICES BEFORE THE HEADLINE. The incumbent's aggregate window accuracy is 34.0 pct, which reads like a system with a general accuracy problem. It is nothing of the sort: it scores 100.0 pct on the 51 readings anchored on the remittance and 0.0 pct on the 36 anchored on the invoice and the 41 anchored on the delivery. Perfectly disjoint -- one third exactly right, two thirds exactly wrong, and the aggregate sitting in between telling you nothing about either. ⚠︎ BOTH FLOORS THAT GUESS SCORE 0 ON THE GUARDRAIL AND BOTH ARMS THAT ABSTAIN SCORE 100. b000 and b001 fill a missing window with an assumed 60 days and a missing clause with the remittance date -- exactly what a threshold report configured once and forgotten does -- and get all 22 of the readings with nothing on file wrong, ageing a deduction against a number nobody supplied. That is not a model finding; it is the row's own guardrail, measured.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key For each of the 150 readings and each of the four answered fields, did the reply equal the computed answer key? The posture, the anchor and the file/hold call are compared exactly against closed lists; the countdown is compared as a whole number of days, and a null MATCHES a null because "no deadline can be computed" is the correct answer on 22 readings and must not be scored as a parse failure.
$0.00
no
yes
the fast tier, with the carried state 96.7% posture accuracy · the fast tier, memory removed (THE CONTROL) 72.7% posture accuracy · the strongest free floor, no model 92.7% posture accuracy · the deductions ageing report an ERP already prints, no memory 32.7% posture accuracy · the same ageing rule, GIVEN the carried state 58.7% posture accuracy · 3 more measured on each run
The same comparison, cut per anchor -- the grader this kit exists for The identical exact-match comparison, restricted to the readings whose CORRECT anchor is the proof of delivery (41), the original invoice (36) or the remittance (51). It answers the one question the aggregate cannot: is this arm uniformly mediocre, or is it wired to the wrong date for part of the population?
$0.00
no
yes
the fast tier, with the carried state 87.8% pod window · the strongest free floor 92.7% pod window · the deductions ageing report an ERP already prints 0.0% pod window · the fast tier, memory removed (THE CONTROL) 87.8% pod window · 2 more measured on each run
The two filing directions, counted apart and in dollars Two counts over two different denominators, and one of them is also counted in money. A MISSED filing: of the 34 readings whose correct answer is to lodge the dispute, how many did not -- and what were they worth, because on a hard clock a missed filing is a write-off rather than a late alert. A DUPLICATE filing: of the 116 readings that must not lodge one, how many did.
$0.00
no
yes
the fast tier, with the carried state 2.9% missed filing · the fast tier, memory removed (THE CONTROL) 0.0% missed filing · the strongest free floor 8.8% missed filing · the deductions ageing report, no memory 64.7% missed filing · the same rule, GIVEN the carried state 64.7% missed filing · 1 more measured on each run
The readings whose answer is not on their own page The same exact-match comparison, restricted to the 39 readings where a dispute had ALREADY been lodged before the reading began. Nothing on those pages says so, so the correct posture (FILED, not EXPIRED and not FILE_NOW) is reachable only through the carried state. It answers one question -- is the memory doing anything, or is the corpus easy?
$0.00
no
yes
the fast tier, with the carried state 92.3% memory posture accuracy · the fast tier, memory removed (THE CONTROL) 0.0% memory posture accuracy · the strongest free floor 100.0% memory posture accuracy · the deductions ageing report, no memory 0.0% memory posture accuracy · 2 more measured on each run
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set CAN tell the arms apart, and the evidence is that it does: five arms scored through one scorer land at 96.67 / 96.67 / 88.0 / 34.0 / 34.0 pct on the dispute window and at 0 / 3 / 2 / 0 / 4 duplicate filings. It also separates them where an aggregate would not: the incumbent's 34.0 pct splits into 0.0 / 0.0 / 100.0 across the three anchors, which is the difference between "this needs tuning" and "two thirds of this is wired to the wrong date". ⚠︎ WHAT IT CANNOT SEPARATE IS THE MODEL FROM ITS OWN CONTROL ON THE ARITHMETIC -- 96.67 pct against 96.67 pct on the window, a dead tie -- and that is a real answer rather than an instrument failure: the carried state buys the posture and the money, not the subtraction. ⚠︎ AND ONE ASYMMETRY LIMITS EVERY COMPARISON HERE. evals/baseline.py resolves the anchor by matching phrases; it cannot MISREAD a clause, only match or fail to. On the nine readings whose clause names a date the page does not carry, that is worth nine cells to the floor and five errors to the model.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Every deduction on your book is disputed from the same date, under one standard clause
a SQL query, or the b000 floor in this repo
Days on the ledger against one threshold is arithmetic. On the slice of THIS corpus where that is true, b000 scores 100.0 pct of countdowns exactly right for nothing at all.
AVOID this kit entirely. Paying a model per row per run to do one subtraction is the most expensive way to get an answer you already have.
Your agreements set the anchor with a coded field, or your clauses come from a small closed set of templates
the b002 floor: a lookup on your codes, or a keyword table over your templates
The model's entire margin here is reading clause prose that a table cannot parse. Take the prose away and there is nothing left to buy: the floor already reaches 88.0 pct on the countdown for $0.00, and it is PERFECT on the reverse-collision clause where the model is not.
AVOID assuming the measured gap transfers. It was measured against a keyword table over free prose, not against your reason codes.
Your dispute clauses are negotiated per customer and say what they say
this kit, with the carried state
That is exactly the corpus these figures were measured on. The countdown goes 96.67 pct against 88.0, the missed-filing bill goes $1,428.34 against $19,683.89, and the at-risk total lands 0.22 pct out against 12.97.
AVOID running it without the carried state. The control shows what that costs: 33 live disputes reported as write-offs, 3 duplicate filings, and an at-risk figure 71.39 pct out.
You want the watch to actually wake up on a schedule
your own scheduler -- cron, Airflow, Temporal, whatever already runs -- and an alarm on a run that did not fire
evals/run.py is INVOKED, not woken, and evals/cadence.py measures what that gap costs: skipping ONE weekly run loses 30 of 34 filings permanently, worth $453,621.42. That is the largest number this kit publishes and it is about the schedule, not the model.
AVOID buying a faster cadence instead. Weekly to monthly loses ONE claim worth $3,278.53 and cuts the bill 4.5x; skipping one weekly run loses eleven. On this population the schedule risk is reliability, not frequency.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
CLAUSE_NAMES_AN_ABSENT_DATE
the clause anchors on a date the page does not carry
5
DED-0009-R3, DED-0018-R2 and DED-0030-R1/R2/R3 all carry the clause "Whichever date the deduction was taken, the window for every category runs from the date the carrier's signed receipt REACHED US." The model answers CONTEXT_INCOMPLETE with anchor NONE, and…
ABSTENTION_COSTS_MONEY
the guardrail fires in the safe direction, at a price
1
DED-0030-R2 is one of the five above and is also the scored run's ONLY missed filing. Gold says FILE_NOW with 7 days left on a $1,428.34 claim; the model abstains under Rule Q-8 and lodges nothing, and the free floor -- which keys on the phrase "signed…
VOCABULARY_OUTSIDE_THE_CLOSED_LIST
right answer, wrong token
2
DED-0003-R3 and DED-0040-R2 both replied ORIGINAL_INVOICE_DATE where the closed list says INVOICE_DATE. Both rationales are exactly right ("Rule Q-3 defaults PRICING to the original invoice date") and both countdowns are exactly right, so neither reading cost…
CONTROL_REPORTS_LIVE_CLAIMS_AS_WRITE_OFFS
the stateless arm cannot tell FILED from EXPIRED
33
On s001-deduction-age-stateless the posture confusion is FILED -> EXPIRED 33 times. Told there is no history, the arm sees a window that closed and reports a write-off; in fact the dispute was lodged in time and the claim is live. That single confusion is…
What we could NOT verify
WHETHER "THE DATE THE CARRIER'S SIGNED RECEIPT REACHED US" MEANS THE DELIVERY. Five of the model's seven anchor errors turn on that one clause and on all five the model's reading is arguably better than the key's. Settling it means rewriting the clause and re-firing 150 calls; the run was not re-fired and the lower figure is published.
WHETHER MEMORY WOULD STILL BE WORTH 24.0 POSTURE POINTS AGAINST A LOG THAT LEAKS. This corpus deliberately keeps filing evidence off the page so the memory claim is falsifiable, and a real deductions log would often carry it. The measured gap is therefore an UPPER BOUND and nothing here measures the real one.
WHETHER ENGLISH BEATS JSON FOR THE CARRIED STATE. src/state.describe renders four scalars as a sentence. The obvious experiment -- score the same readings with the state rendered both ways -- costs one more 150-call run and has not been paid for.
WHETHER CARRYING LAST WEEK'S COUNTDOWN HURTS THE ANCHOR. The control scores 96.67 pct on the anchor against this arm's 95.33, which is two cells and is reported as two cells. The mechanism is plausible -- a reading that can subtract 7 from a carried number need never re-read the clause -- and two cells cannot establish it.
WHAT A DIFFERENT CADENCE DOES TO THE ANSWERS. evals/cadence.py measures what other intervals do to the SCHEDULE ($0.00, no model), but both posture bands are DEFINED as the cadence, so a daily or monthly watch is a different scoring question and no arm has been run at one.
WHAT A CHANGING POPULATION DOES. All 50 deductions are open at all three runs. A claim first seen mid-window, or one that settles between runs, is unmeasured.
WHETHER THE DESK NOTE MOVES THE ANSWER. The injection surface is sent deliberately and was never attacked. No resistance rate is claimed.
WHETHER A SECOND MODEL AGREES. One tier was run. No cross-model claim is made anywhere on this page.
WHETHER ANY OF THIS IS STABLE ACROSS REPEATS. One run per arm, so every difference under about two readings is inside the noise nobody has measured.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,704.59
716.47
4,974 ms
$0.003002
$0.001201
$0.052869
the same tier, memory removed (THE CONTROL)
1,681.62
562.45
4,853 ms
$0.002528
$0.001011
$0.044939
the three free floors, no model
0
0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same reading (one deduction, at one scheduled run), the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
306 live calls were attempted for this kit and all 306 returned something: 6 calibration (c000, at an 8,000-token ceiling on two deduction chains), 150 scored, and 150 for the stateless control. NOTHING WAS DISCARDED, no run was re-fired and no screenshot spent anything -- all four UI frames were taken with API_KEY blanked in the child process's environment and the two that show a verdict replay the committed result file. Everything else -- all three free floors, the cadence and missed-run arithmetic, the wiring stub, the scorer, the fifteen pre-flight assertions -- is pure code and costs $0.00. ⚑ AND THE LARGEST FIGURE THIS KIT PUBLISHES IS IN THAT FREE HALF: $453,621.42 lost to one skipped run, computed with no key. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, and most of them are reasoning. 88.2 pct of this run's output (94782 of 107470) was provider-side reasoning left at the default, and output is 71.6 pct of the projected bill on the shared card. What you are paying for is the anchor decision.
THE CADENCE, which multiplies everything else. One reading is one call; the watch wakes weekly; the bill is rows x runs, not rows. ⚑ AND THE CHEAPER CADENCE IS ALMOST FREE OF CONSEQUENCE HERE, MEASURED: monthly costs 200 calls over the horizon against weekly's 900 and loses one claim worth $3,278.53 (evals/cadence.py, $0.00).
THE OPEN QUEUE, not the desk's turnover. A deduction that is disputed, settled or written off leaves the watch and stops costing anything. A desk that never closes its claims pays for them forever.
The snapshot itself, barely. Input averages 1705 tokens and every document in this corpus is within 9 pct of every other in size.
Your volumeWhat it costs at your volume
LINEAR IN ROWS x RUNS, AND THAT IS THE WHOLE WARNING. Ten times the open deductions is ten times the calls at the same cadence -- there is no batching, no cache and no early exit, because every open claim is re-read whole on every run by design. A desk holding 4,000 open deductions on this weekly watch is 4,000 calls a week, about $12.01 a week on the shared projection card. Nothing about the per-call price changes; the multiplier is the schedule, which is why this kit's environment ladder makes cadence a declared decision rather than a deployment detail.
Where pricing changes shape
The output ceiling. MAX_TOKENS = 24000, and the calibration run UNDER-ESTIMATED what was needed: c000 fired two deduction chains at an 8,000-token cap and topped out at 1004 output tokens, all six parsed -- which reads as 24x headroom. The scored run's largest reply was 14538. A ceiling is not a cost (the provider bills tokens produced, not tokens allowed) so the generous figure was kept; a ceiling set from that calibration would have been six times too low and would have turned the hardest readings into lost ones.
Provider-side reasoning left at the default. It is 88.2 pct of this run's output and no run here has measured what disabling it does to the answers -- so a vendor whose default differs reprices this kit without changing anything you can see.
Your return, with your numbers
Volumeopen deductions per scheduled run -- this run judged 150 (50 deductions x 3 runs) per arm, on a weekly watch
What it replacesan analyst working the deductions ageing report at week's end: opening each claim, finding the trade agreement, reading its dispute clause to work out which date the window runs from, subtracting the elapsed days, and remembering which claims already have a dispute lodged
Time saved per itemnot measured here -- depends on how long an analyst takes to find and read a customer's dispute clause, which is the whole task
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared connection was already pointed at, and the kit's whole claim is that swapping it is one line of .env and one more run. No comparison across tiers was attempted and nothing on this page claims one.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
255,689input tokens · this run
107,470output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 150 readings, one completion call each, one tier. The 150-call stateless control and the 6 calibration calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here. Every figure this kit publishes about the SCHEDULE cost nothing at all and is not on this table.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.180
$0.180
$1.20
2026-09-12
gemini-3-flash
Google
$0.450
$0.450
$3.00
2026-09-18
gemini-3-8-flash
Google
$0.595
$0.595
$3.97
2026-09-18
llama-5
Meta
$0.776
$0.776
$5.18
2026-09-18
claude-haiku-4-5
Anthropic
$0.793
$0.793
$5.29
2026-09-12
grok-4-5
xAI
$1.156
$1.156
$7.71
2026-09-18
grok-4-6
xAI
$1.156
$1.156
$7.71
2026-09-18
claude-sonnet-5
Anthropic
$1.586
$1.586
$10.57
2026-09-12
gemini-3-1-pro
Google
$1.801
$1.801
$12.01
2026-09-18
gpt-5-6-terra
OpenAI
$1.801
$1.801
$12.01
2026-09-12
gpt-5-6-sol
OpenAI
$3.172
$3.172
$21.15
2026-09-12
claude-opus-4-8
Anthropic
$3.965
$3.965
$26.43
2026-09-12
claude-opus-5
Anthropic
$3.965
$3.965
$26.43
2026-09-12
claude-fable-5
Anthropic
$7.930
$7.930
$52.87
2026-09-18
claude-fable-5-1
Anthropic
$7.930
$7.930
$52.87
2026-09-18
gpt-6-astra
OpenAI
$7.930
$7.930
$52.87
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 88.2 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (94782 of 107470), left at the provider's default, so every row below prices a reasoning-on workload. Output is 71.6 pct of the projected bill on the shared card. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
⚑ EVERY ROW PRICES ONE READING, AND A DEPLOYMENT DOES NOT BUY ONE READING. This watch wakes weekly and re-reads every open deduction each time, so the bill is rows x runs. Multiply any row below by your own open-claim count and by 52 before comparing it with anything.
⚑ AND THE CHEAPER ARM IS THE ONE WITHOUT MEMORY, WHICH IS THE OPPOSITE OF THE SIBLING MONITOR. The stateless control cost 15.8 pct LESS per reading on the same card, because a reading with no carried countdown simply subtracts while one with a carried countdown reconciles. Memory here is a 18.7 pct premium that buys 24.0 posture points.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
13 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 50 open deductions x 3 weekly runs = 150 snapshots from a fixed seed (SEED = 990823). 13 of the 50 customers carry a dispute clause that OVERRIDES the type default anchor, which moves 39 readings off it -- that is the experiment. The gold labels are src/deduction.step's output over the planted terms, never typed.
You change it to: The deductions, the clause pools and the seed. Keep the seven section headings or src/segment.py's assertion refuses to start.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 990823
COUNT = 50
RULE = "-" * 64
RUN_DATES = (dt.date(2026, 3, 2), dt.date(2026, 3, 9), dt.date(2026, 3, 16))
HORIZON_START = dt.date(2026, 1, 5)
src/deduction.pythe dispute arithmetic — a swap seam
The rule as pure code: the anchor resolved from the clause override or the type default, the window granted minus the days elapsed from that anchor, and the two bands -- both of which ARE the cadence. No model, no judgement. ⚠︎ The eight rules, the window lengths and the anchor defaults are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
You change it to: TYPE_DEFAULT_ANCHOR and RULE_TEXT. The rule text is reproduced in full on all 150 pages precisely so it can be read, disbelieved and replaced with the agreements you actually signed.
src/deduction.py
# The dispute-window rule as arithmetic. Pure code, no model, standard library only.
CLEAR = "CLEAR"
WATCH = "WATCH"
FILE_NOW = "FILE_NOW"
FILED = "FILED"
EXPIRED = "EXPIRED"
CONTEXT_INCOMPLETE = "CONTEXT_INCOMPLETE"
POSTURES = (CLEAR, WATCH, FILE_NOW, FILED, EXPIRED, CONTEXT_INCOMPLETE)
INVOICE_DATE = "INVOICE_DATE"
POD_DATE = "POD_DATE"
src/state.pythe carried state — a swap seam
SEAM 2 -- the thing that makes this a monitor. Four scalars per deduction (whether a dispute is already lodged, last run's countdown, the posture last reported, how many runs have seen it), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. ⚠︎ One of the four is double-edged and the kit measures it rather than defending it: carrying last week's countdown lets a reading reach the right number by subtracting 7 without re-reading the clause.
You change it to: What is carried between scheduled runs, and how it is worded to the model. Four scalars today. ⚠︎ ONE OF THEM IS A KNOWN TRADE-OFF: dropping days_to_deadline_last_run would force every reading to re-derive the countdown from the clause, which is slower and possibly more accurate on the anchor -- the control arm scores 96.67 pct on the anchor against this arm's 95.33. Two cells; not a claim.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_deduction(store, deduction_id):
def describe(state):
src/segment.pythe section splitter
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 documents before a run may spend.
src/segment.py
# Split a deduction snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Deduction", "Dispute Terms", "Aging Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector
Decides which sections reach the model. Customer Contact -- the customer's A/P analyst, her direct dial and her work email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/select.py
# Pick which sections of a snapshot are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
DEDUCTION = "Deduction"
TERMS = "Dispute Terms"
POSITION = "Aging Position"
ACTIVITY = "Claim Activity Log"
CONTACT = "Customer Contact"
NOTES = "Deductions Desk Notes"
NEVER_SENT = (CONTACT,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading.
You change it to: PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe watch
One deduction, one scheduled run, one call. Parses the window, the three candidate anchor dates and their elapsed-day figures off the page with a regex (the model is never asked to subtract two dates), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
src/watch.py
# One deduction, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled deduction-queue watch. You apply a written dispute rule to one "
MAX_TOKENS = 24000
FIELDS = ("posture", "anchor", "days_to_deadline", "file_dispute")
def documents():
def chains():
def load_doc(doc_id):
def position_of(text):
src/app.pythe local UI
One deduction, one scheduled run, its carried state and its verdict, on 127.0.0.1:9016. Renders with no key. It shows the carried sentence verbatim, the free floor's answer beside the model's, and a second button that replays what run r001-deduction-age actually answered, straight off the committed result file, labelled as a replay.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9016"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-deduction-age")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe three free floors — a swap seam
b000 the deductions ageing report every ERP prints -- days since the deduction posted against one assumed 60-day window; b001 the same rule given the carried state, which isolates what memory alone buys; b002 the strongest free version -- carried state, the window read off the page, the anchor resolved by an ordered keyword table over the clause with a fall-back to the type default, and an abstention where nothing is on file. 0 calls, $0.00, all three scored through the identical scorer.
You change it to: POD_WORDS, INVOICE_WORDS and DEDUCTION_WORDS, and the order they are tried in -- the keyword table the model is measured against. Widen it and the floor gets stronger, which is the direction this kit wants: a floor built to lose measures nothing.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("agedays", "agedays-mem", "anchor-mem")
ASSUMED_WINDOW_DAYS = 60
POD_WORDS = ("signed receipt", "proof of delivery", "receiving dock", "goods were signed",
INVOICE_WORDS = ("billing document", "invoice date", "we billed", "billed the order",
DEDUCTION_WORDS = ("remittance advice", "short payment posted", "deduction is taken",
def _num(pat, text, cast=int):
def _clause(text):
def _anchor_from_clause(clause, deduction_type):
def _elapsed(text, label):
evals/cadence.pythe cadence arithmetic
⚑ THE HALF OF A MONITOR THE ACCURACY TABLES CANNOT SEE, and it costs $0.00. It answers two questions from the deduction model with no key and no network: what does skipping ONE scheduled run cost (30 of 34 filings lost permanently, $453,621.42), and what does the interval itself cost (weekly to monthly loses one claim worth $3,278.53 and cuts the bill 4.5x). A reading that never happened produces no cell for a scorer to grade, which is why this is a separate file rather than a column.
evals/cadence.py
# WHAT A MISSED RUN COSTS, AND WHAT THE CADENCE COSTS. Free. No key, no model, no network.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
SWEEP_CADENCES = (1, 3, 7, 14, 30)
def population():
def evidence_ready(v):
def schedule(start, end, cadence_days):
def catchable(v, runs):
def main():
evals/scoring.pythe scorer
Exact match per cell against the computed gold, split ways an average would hide: the four fields, the two filing directions counted apart AND in dollars, the memory-dependent subset, the context-incomplete recall, and -- the cut this kit exists for -- every rate again PER ANCHOR and PER SCHEDULED RUN. No judge model.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("posture", "anchor", "days_to_deadline", "file_dispute")
ANCHOR_SLICES = ("POD_DATE", "INVOICE_DATE", "DEDUCTION_DATE")
def _pct(n, d):
def _days(v):
def score(records, golds):
evals/check_labels.pythe pre-flight
Fifteen things that must be true before a run may spend: the seven sections parse, every sequence is complete and gap-free, no run-varying section says a dispute was lodged and the non-varying ones are byte-identical across runs (the experimental control), no page carries a deadline field or a countdown, THE ANCHOR IS NOT DERIVABLE FROM THE DEDUCTION TYPE, every anchor slice is big enough to report, the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path writes off or credits a deduction, and the answer key replays from src/deduction.step.
evals/check_labels.py
# Everything that must be true BEFORE a run is allowed to spend. Free, and it runs in seconds.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAILED = []
def check(name, ok, detail=""):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 50 open deductions x 3 weekly runs = 150 snapshots from a fixed seed (SEED = 990823). 13 of the 50 customers carry a dispute clause that OVERRIDES the type default anchor, which moves 39 readings off it -- that is the experiment. The gold labels are src/deduction.step's output over the planted terms, never typed. A swap seam.
src/deduction.pyThe rule as pure code: the anchor resolved from the clause override or the type default, the window granted minus the days elapsed from that anchor, and the two bands -- both of which ARE the cadence. No model, no judgement. ⚠︎ The eight rules, the window lengths and the anchor defaults are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Four scalars per deduction (whether a dispute is already lodged, last run's countdown, the posture last reported, how many runs have seen it), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. ⚠︎ One of the four is double-edged and the kit measures it rather than defending it: carrying last week's countdown lets a reading reach the right number by subtracting 7 without re-reading the clause. A swap seam.
src/segment.pySplits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 documents before a run may spend.
src/select.pyDecides which sections reach the model. Customer Contact -- the customer's A/P analyst, her direct dial and her work email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/prompt.pyThree parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading. A swap seam.
src/watch.pyOne deduction, one scheduled run, one call. Parses the window, the three candidate anchor dates and their elapsed-day figures off the page with a regex (the model is never asked to subtract two dates), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
evals/baseline.pyb000 the deductions ageing report every ERP prints -- days since the deduction posted against one assumed 60-day window; b001 the same rule given the carried state, which isolates what memory alone buys; b002 the strongest free version -- carried state, the window read off the page, the anchor resolved by an ordered keyword table over the clause with a fall-back to the type default, and an abstention where nothing is on file. 0 calls, $0.00, all three scored through the identical scorer. A swap seam.
evals/cadence.py⚑ THE HALF OF A MONITOR THE ACCURACY TABLES CANNOT SEE, and it costs $0.00. It answers two questions from the deduction model with no key and no network: what does skipping ONE scheduled run cost (30 of 34 filings lost permanently, $453,621.42), and what does the interval itself cost (weekly to monthly loses one claim worth $3,278.53 and cuts the bill 4.5x). A reading that never happened produces no cell for a scorer to grade, which is why this is a separate file rather than a column.
evals/scoring.pyExact match per cell against the computed gold, split ways an average would hide: the four fields, the two filing directions counted apart AND in dollars, the memory-dependent subset, the context-incomplete recall, and -- the cut this kit exists for -- every rate again PER ANCHOR and PER SCHEDULED RUN. No judge model.
evals/check_labels.pyFifteen things that must be true before a run may spend: the seven sections parse, every sequence is complete and gap-free, no run-varying section says a dispute was lodged and the non-varying ones are byte-identical across runs (the experimental control), no page carries a deadline field or a countdown, THE ANCHOR IS NOT DERIVABLE FROM THE DEDUCTION TYPE, every anchor slice is big enough to report, the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path writes off or credits a deduction, and the answer key replays from src/deduction.step.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1704 input and 716 output tokens per reading (one deduction, at one scheduled run), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one deduction, at one scheduled run)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one deduction, at one scheduled run) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Deductions Desk Notes, which are one of three fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- in a working desk much of what lands there is pasted verbatim out of a customer's claim portal or a chaser email -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
The experimentWe did NOT attack it — and this is the surface we left open
An indirect prompt injection needs a field somebody outside can write into, and this kit has exactly one: the Deductions Desk Notes. It is SENT rather than hidden, because hiding a surface does not close it -- and one of the three shipped notes asks, in the customer's voice, for the thing the kit exists to do. No attack was fired at this kit and none is claimed. The three boundaries below were checked in code on 2026-08-23, and only the first of them is red-proven in both directions rather than argued from an absent code path.
Boundary checked
What could go wrong
What the code guarantees
Does a named individual at the customer ever leave the machine?
Every snapshot carries a Customer Contact section -- the customer's accounts-payable analyst by name, her direct dial and her work email. Not one field this kit answers asks for any of it, and a named contact at a trading partner is the sort of field that turns a receivables worklist into a personal-data disclosure the moment somebody pastes it somewhere. A selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the snapshot MINUS that section rather than to the snapshot. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py over all 150 documents. ⚠︎ AND THE FIRST VERSION OF THAT PROOF WOULD HAVE MEASURED ZERO AND BEEN THE TEST NOT FIRING. Swapping the guard for the naive or list(secs) changes nothing on today's corpus, because every hint names a section all 150 documents carry -- so the fallback is not on any live code path. The hole is CONDITIONAL, so the proof reproduces the condition: a receivables system renaming its blocks on upgrade. With the guard, 0 of 150 leak; with or list(secs), 150 of 150.
Can a wrong answer poison the next scheduled run?
A monitor that fed its own verdict forward would compound one bad reading into every reading after it -- and here the compounded quantity is a DEADLINE, so a claim mis-anchored in week 1 is a write-off in week 3 with no run left to recover it.
src/deduction.step() is the only thing that writes the carried state and it reads only the figures parsed off the page. The model's reply is scored and thrown away. On r001-deduction-age the 18 wrong cells did not reach a single later reading, because there is no code path by which they could.
Can this kit write off, credit or adjust a deduction?
A watch that can decide a window has closed could plausibly be extended to post the write-off that follows -- which is this row's own cap and the thing a finance controller would most object to.
There is no such endpoint, no such function and no configuration flag. evals/check_labels.py greps every .py and .js file in the kit for the names of such paths and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
The first row is the only one measured in both directions. The other two are guarantees about what is ABSENT — no code path writes off, credits or adjusts a deduction, and no code path feeds a model reply back into the carried state — and absence is checked by asserting names it knows, which is weaker than a run and is written here as such.
The result0 attack trials, three boundaries checked -- and the privacy boundary red-proven by removing the guard and watching all 150 snapshots leak a named individual's direct dial and work email.
1field an outside party could influence (sent, not hidden)
0 of 0attack trials run
1 of 3boundaries red-proven, not just asserted
0 of 150snapshots leak a customer contact's name, dial or email
The Deductions Desk Notes ARE the field an outside party would influence in a real deployment, and this kit sends them rather than hiding them -- one of the three shipped sentences is even instruction-shaped, and it asks for the omission that costs the money. But on this corpus they are one of three sentences from a seeded generator, so there is nothing adversarial in them to catch. A version pointed at real desk prose reopens the question and should be attacked before it ships. What IS measured is the privacy boundary, in both directions, before every run.
Read this twice
The claim activity log and the desk notes reach the model verbatim, and one of the three shipped notes is instruction-shaped on purpose — “Customer’s claims desk has asked that we raise nothing further on this deduction while their own review is open.” That sentence asks the watch not to file, which is exactly the action whose omission costs the money. Nothing here filters it and nothing here has measured whether it moves the answer. A model declining to follow an instruction would not be a defence anyway — it is one vendor’s behaviour on one day.
HonestyWhat this does not prove
Whether a real desk's notes -- prose an analyst writes freely, often pasted out of a customer's portal -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at a real queue. The shipped note asking that nothing further be raised is the shape to probe, because the action it asks to suppress is the one whose omission is unrecoverable.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two watches on one queue have never been run and would race. Here a lost write forgets which disputes are lodged, and the next run files a second one against every claim in the queue.
Whether a truncated reply can ever be partially trusted. Neither scored arm lost a call, so src/watch._parse's outright rejection was never exercised on this kit.
Whether the grep-shaped no-write-off assertion would catch a path written under a name it does not know. It asserts the absence of nine names; the guarantee is about ABSENCE and that is the only mechanical form it can take.
What the provider retains. The prompt carries no personal data by construction -- the Customer Contact section never leaves -- but what a vendor keeps of a request is outside this repository and nobody here has verified it.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No write-off, no credit, no customer contact -- non-configurable. This kit produces a worklist posture, an anchor, a countdown and a file/hold call for a deductions analyst to read. It never writes off, credits, adjusts, clears or suppresses a deduction, never posts to a customer's account, never lodges anything with a customer's claims portal and never sends a message to a customer, and there is no setting that makes it. Separately, and just as non-configurable: the dispute window and its anchor are OPERATOR-SUPPLIED. A deduction whose agreement carries no dispute clause, or whose anchor date is not on file, is reported CONTEXT_INCOMPLETE and aged against nothing -- there is no default window and no substitute anchor in this kit.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json), evals/cadence.py (one results file) and src/state.save (data/state.json). src/deduction.step() is the only thing that touches the carried state, and it is called by evals/run.py after every reading including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit writes off, credits or adjusts a deduction
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for nine such names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
No deduction is aged against a guessed window or a substitute anchor
22 of 22 readings with nothing on file were reported CONTEXT_INCOMPLETE by the model (100.0 pct), and 22 of 22 by the strongest free floor. The two floors that GUESS -- an assumed 60-day window from the remittance date -- got all 22 wrong, 0.0 pct, which is the guardrail measured rather than asserted.
Exactly one dispute per deduction
0 duplicate filings in 116 readings that must not lodge one, and 1 missed in 34 that must. ⚠︎ The zero is a result on THIS corpus, not a property of the code: nothing in the kit refuses a second filing, the carried state merely tells the model one exists. The control arm, with that sentence removed, produced 3.
A wrong answer does not propagate into the next scheduled run
The property is guaranteed by the code path -- src/deduction.step() never reads the reply -- and r001-deduction-age is CONSISTENT with it rather than a demonstration of it.
The customer contact's name, direct dial and email never reach the provider
0 of 150 snapshots leak the Customer Contact section, and 150 of 150 leak it when the guard is removed AND the condition that reaches the fallback is reproduced. Both directions asserted by evals/check_labels.py before any run may spend.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The carried state is correct whatever the model says, which means a wrong reading is a wrong worklist row and not a corrupted history. Those are different problems and only the first one is on this page.
IT IS NOT A SCHEDULER, AND ON THIS KIT THAT IS THE MOST EXPENSIVE THING IT IS NOT. evals/run.py is invoked; it is not woken. Nothing here detects a missed run, back-fills it or marks its readings late -- and evals/cadence.py measures what one skipped week costs: 30 of 34 filings gone permanently, $453,621.42.
IT IS NOT AN AUTHORITY ON DEDUCTION DISPUTES. The eight rules, the window lengths and the anchor defaults are invented for this kit and reproduce no trade agreement, customer deduction policy, retailer's vendor guide or regulator's guidance.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 42 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
24 measured by the latest run18 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
alarm
dispute_window_accuracy_pct; anchor_accuracy_pct; posture_accuracy_pct; file_dispute_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Both scored arms had 0 of 150.
per-anchor-slice
The same comparison, cut per anchor -- the grader this kit exists for
alarm
pod_window_pct; invoice_window_pct; remittance_window_pct — alarm on any slice more than 10 points below the aggregate. That is the signature of a rule wired to the wrong date for part of the population, and it is invisible on a single number.
filing-directions
The two filing directions, counted apart and in dollars
alarm
missed_filing_pct; duplicate_filing_rate_pct; missed_filing_usd — alarm on missed_filing_usd above zero. A duplicate is visible to whoever reads the worklist and to the customer's claims desk; a missed filing is invisible until a month-end write-off report, by which point the window has been shut for weeks.
memory-dependent-subset
The readings whose answer is not on their own page
alarm
memory_posture_accuracy_pct; memory_window_accuracy_pct; memory_file_accuracy_pct — alarm on memory_posture_accuracy_pct falling below the whole-population posture figure. That would mean the carried sentence is confusing the model rather than informing it.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
636,931
deduction records edited — the count held, the bytes did not
split.count
50
the deductions count moved — a different set was scored
split.size_p50
3
the median size of one deduction moved
split.size_p95
3
the 95th-percentile size of one deduction moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cadence_days 7, context_incomplete_cells 22, deductions 50, documents 150, expired_cells 13, memory_cells 39, must_file_cells 34, quiet_cells 116, readings_scored 150, stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
the dispute window, exact days
not yet known
150 readings
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. r001-deduction-age and s001-deduction-age-stateless are one run each of two DIFFERENT prompts, which is a comparison and not a band.
the anchor, per slice
95.33 pct whole -- and 87.8 / 94.44 / 100.0 across the delivery, invoice and remittance slices
150 readings; 41 / 36 / 51 in the three slices
evals/scoring.py::score, r001-deduction-age. The incumbent free rule scores 0.0 / 0.0 / 100.0 on the same slices, which is why the slice is the band and the aggregate is not.
the file/hold call, both directions
99.33 pct with 1 missed ($1,428.34) and 0 duplicate
34 that must lodge a dispute, 116 that must not
The free floor makes 3 and 2 on the same cells, and the control 0 and 3. This grader discriminates in both directions.
the posture
96.67 pct -- 92.31 over the 39 memory-dependent cells
150 readings; 39 of them memory-dependent
The control, on the same cells, scores 72.67 and 0.0. That gap is the carried state, isolated.
the operator-supplied window guardrail
100.0 pct -- 22 of 22
22 readings with no dispute clause on file or no anchor date on file
Both arms that abstain score 100; both floors that fill a 60-day default score 0.0. This is the row's own guardrail, measured.
the money
0.22 pct out on a queue at-risk total of $651,445.25
1 -- the queue total, which is why the per-reading rates are published beside it
The free floor is 12.97 pct out and the control 71.39 pct -- and the control is out in the DANGEROUS direction, overstating the write-off by reporting live claims as expired.
One run. The control at the identical ceiling also answered 100.0 pct and cost 15.8 pct LESS per reading, so neither the reliability nor the price is a property of the design.
⚑ the schedule -- the band nothing on the model path can see
30 of 34 filings and $453,621.42 lost to ONE skipped weekly run
3 scheduled runs, 34 last-chance filings
evals/cadence.py over the deduction model. $0.00, no key, no network. Widening the cadence from weekly to monthly loses ONE claim ($3,278.53) and cuts the call count 4.5x, so the risk here is reliability rather than frequency.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-deduction-age-agedays 2026-08-23
b001-deduction-age-agedaysmem 2026-08-23
b002-deduction-age-anchormem 2026-08-23
anchor accuracy, %
34.0
34.0
88.0
anchor invoice, %
0.00
0.00
83.33
anchor pod, %
0.00
0.00
92.68
anchor remittance, %
100.00
100.00
82.35
answered, %
100.0
100.0
100.0
at risk usd error, %
30.53
59.58
12.97
context incomplete recall, %
0.0
0.0
100.0
dispute window accuracy, %
34.0
34.0
88.0
duplicate filing rate, %
3.45
0.00
1.72
expired recall, %
46.15
46.15
100.00
file dispute accuracy, %
82.67
85.33
96.67
input tokens, whole run
0
0
0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
memory file accuracy, %
89.74
100.00
100.00
memory posture accuracy, %
0.0
100.0
100.0
memory window accuracy, %
28.21
28.21
94.87
missed filing, %
64.71
64.71
8.82
missed filing usd
256640.23
256640.23
19683.89
output tokens, whole run
0
0
0
posture accuracy, %
32.67
58.67
92.67
window invoice, %
0.00
0.00
83.33
window pod, %
0.00
0.00
92.68
window remittance, %
100.00
100.00
82.35
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 3 runs. Columns here are only ever compared with each other.
Metric
c000-deduction-age-calibration 2026-08-23
r001-deduction-age 2026-08-23
s001-deduction-age-stateless 2026-08-23
anchor accuracy, %
83.33
95.33
96.67
anchor invoice, %
—
94.44
100.00
anchor pod, %
75.0
87.8
87.8
anchor remittance, %
—
100.0
100.0
answered, %
100.0
100.0
100.0
at risk usd error, %
0.00
0.22
71.39
context incomplete recall, %
100.0
100.0
100.0
dispute window accuracy, %
83.33
96.67
96.67
duplicate filing rate, %
0.00
0.00
2.59
expired recall, %
—
100.00
92.31
file dispute accuracy, %
100.00
99.33
98.00
input tokens, whole run
10250
255689
252243
model latency p50 ms
5632.00
4974.00
4853.00
model latency p95 ms
8984.00
9795.00
8566.00
memory file accuracy, %
100.00
100.00
92.31
memory posture accuracy, %
50.00
92.31
0.00
memory window accuracy, %
50.00
92.31
92.31
missed filing, %
0.00
2.94
0.00
missed filing usd
0.00
1428.34
0.00
output tokens, whole run
3978
107470
84367
posture accuracy, %
83.33
96.67
72.67
window invoice, %
—
100.0
100.0
window pod, %
75.0
87.8
87.8
window remittance, %
—
100.0
100.0
not a time series No two of these 3 runs measured the same system — they differ on context_incomplete_cells, deductions, documents, expired_cells, max_tokens, memory_cells, must_file_cells, quiet_cells, readings_scored, stateless — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-deduction-age-stub 2026-08-23
anchor accuracy, %
34.0
anchor invoice, %
0.0
anchor pod, %
0.0
anchor remittance, %
100.0
answered, %
100.0
at risk usd error, %
30.53
context incomplete recall, %
0.0
dispute window accuracy, %
34.0
duplicate filing rate, %
3.45
expired recall, %
46.15
file dispute accuracy, %
82.67
input tokens, whole run
265916
model latency p50 ms
0.00
model latency p95 ms
0.00
memory file accuracy, %
89.74
memory posture accuracy, %
0.0
memory window accuracy, %
28.21
missed filing, %
64.71
missed filing usd
256640.23
output tokens, whole run
6634
posture accuracy, %
32.67
window invoice, %
0.0
window pod, %
0.0
window remittance, %
100.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 24 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
r001-deduction-age against s001-deduction-age-stateless. Same corpus, same model, same grader, same 150 readings; the prompts differ in exactly one block and evals/check_labels.py asserts they are byte-identical everywhere else. ⚑ THE MEMORY BUYS THE POSTURE AND THE MONEY AND NOTHING ELSE, AND IT IS NOT FREE -- which is the opposite of what the sibling monitor found, and is published as measured rather than reconciled.
r001-deduction-age against b002-deduction-age-anchormem over the same 150 readings and the same scorer. The model wins the discriminator by 13 readings. ⚠︎ The comparison is asymmetric in the floor's favour on nine of them: evals/baseline.py matches the phrase "signed receipt" and cannot misread the sentence it never reads, so it takes 9 of 9 where the model takes 4.
whether memory is available to FREE CODE, with no model involved
b000-deduction-age-agedays against b001-deduction-age-agedaysmem. Same rule, same corpus, no model in either. Memory stops you filing twice and does not move the countdown by a single reading -- which is the cleanest statement on this page of what a monitor's memory is actually for.
b001-deduction-age-agedaysmem against b002-deduction-age-anchormem -- both free, both with the carried state, and the second one resolves the anchor from the clause and abstains where nothing is on file. 54.0 points of window accuracy and $236,956.34 of write-off, for $0.00. This is the single largest movement anywhere on this page and there is no model in it.
whether a scheduled run actually fires
30 of 34 last-chance filings, worth $453,621.42, from lodged to permanently lost
measured
evals/cadence.py over the deduction model, $0.00 and no provider. It checks from the answer key that the NEXT run finds the window already closed rather than asserting it. This is the largest single figure this kit publishes and the only lever on this list that no model change can touch.
the cadence itself
weekly -> monthly: 900 calls -> 200 over the same horizon, and 0 -> 1 claims lost ($3,278.53)
measured
evals/cadence.py's sweep over 1, 3, 7, 14 and 30-day schedules against the same population and horizon. ⚠︎ THE HONEST READING IS THAT THE FREQUENCY BARELY MATTERS HERE AND THE RELIABILITY IS EVERYTHING: a 4.5x cheaper cadence costs one claim, and one skipped run costs eleven. It is a property of how these deadlines and evidence dates are distributed, and a queue whose evidence arrives later would behave differently.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
the dispute window, exact days
nothing yet. ⚠︎ And note what this figure is measured against: the free floor scores 88.0 pct on the same 150 readings.
the anchor, per slice
any slice more than 10 points below the aggregate. That is a rule wired to the wrong date for part of the population, and a single number cannot see it
the file/hold call, both directions
any missed filing at all. The cost of that direction is the whole deduction, because the window does not reopen
the posture
FILED being reported as EXPIRED at all. That is a live claim being called a write-off, and the control does it 33 times
the operator-supplied window guardrail
any countdown at all on a deduction with nothing on file. One such reading is a claim aged against a number nobody supplied
the money
the total drifting while the per-reading exact rate holds. That is errors stopping cancelling, and it means the population has changed
reliability and the bill
answered_pct below 100, or output tokens rising without the corpus changing
⚑ the schedule -- the band nothing on the model path can see
a run that did not fire. There is no symptom for this INSIDE the kit -- see environment.signatures' traceless row -- so the alarm has to live in whatever invokes it
NextThe three you would add first
A scheduler, and something that notices when a run did not happenThis is the half of a monitor the kit does not ship, and here it is measured rather than argued: skipping ONE weekly run loses 30 of 34 last-chance filings PERMANENTLY, worth $453,621.42. Every accuracy figure on these pages is measured over three runs that all happened.
A human step in front of anything that becomes a write-offThe consequential output here is the ABSENCE of a filing, which is the hardest kind of decision to review because it leaves no artefact. The arms without a clause read miss 22 of 34 filings worth $256,640.23 and nothing anywhere flags it.
The real dispute clauses, from the real trade agreementsEvery window length and every anchor in this corpus is invented. The row this kit was built from records the window as an operator-supplied clock precisely because it varies per customer -- and the kit's own CONTEXT_INCOMPLETE rate is the measurement of what happens when it is missing, not a licence to guess.
A closed vocabulary the reply is validated against before it is scoredTwo of this run's seven anchor errors were the right answer under a name outside the list (ORIGINAL_INVOICE_DATE for INVOICE_DATE). In a deployment that reading needs to be visibly REJECTED and retried rather than silently counted as a different anchor.
A state store with a concurrency modeldata/state.json is one file replaced atomically. That is right for one process and is not a design for a scheduler, and what is carried is which disputes exist -- so a lost write does not degrade the answer, it re-files every claim in the queue.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free, under a second) on any change to tools/build_corpus.py, src/deduction.py, src/segment.py or src/select.py -- it refuses to let a run spend if the corpus, the answer key, the anchor trap, the experimental control, the privacy guard or the no-write-off property has moved. Re-run evals/cadence.py on any change to the corpus or the cadence; it is free and it carries the largest figure on the page. Re-run the scored eval AND its stateless control together (paid, 300 calls) on any change to src/prompt.py or src/state.py: the memory claim is a difference, so one arm re-run alone is not comparable with the other's old figure. All three free floors cost nothing and should be re-run on every change, in both directions -- a floor that gets stronger is a finding, not a regression.
What this cannot tell you
One run per arm. Whether the 8.67-point win over the free floor is stable across repeats is not measured, and 13 readings either way is the whole margin.
Whether the no-propagation property would hold under a wrong answer mid-chain. It is true by construction -- src/deduction.step() never reads the reply -- and this run is only consistent with it.
Whether the no-write-off guarantee holds against a path named something the assertion does not know. It greps for nine names; the guarantee is about ABSENCE and that is the only mechanical form it can take.
Whether rewriting the "reached us" clause would remove five of this run's seven anchor errors. It has not been tried; it costs one more 150-call run.
Whether the zero duplicate filings would survive a corpus where a dispute is lodged and then REJECTED by the customer, or withdrawn, or partially allowed. This corpus has no closed disputes, and a re-filing after a rejection is a legitimate second dispute that this kit's vocabulary cannot express.
Whether the missed-run arithmetic holds on a queue that takes on new deductions between runs. evals/cadence.py's population is fixed; a claim first seen after its own window has closed is a different failure and is not counted here.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries four scalars written by arithmetic. The cost is flat in history length because nothing here remembers what was said, only what was decided -- and a deduction can sit in a queue for a quarter, so run 52 is an ordinary number here. It would also give back the failure this design exists to avoid: a deadline fed from the model's own output compounds into a write-off three runs later
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the schedule
evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
THIS is the seam where a framework genuinely earns its place, and this kit has the number to prove it rather than the usual assertion: 50 independent chains of 3 strictly-ordered readings, with the cadence, retries, missed-run detection and late back-fill all OUTSIDE the kit -- and one missed run costs $453,621.42. A ThreadPoolExecutor is the right size for an eval and the wrong size for a watch that has to wake every week and notice when it did not
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
four exact-match comparisons, three slices of the same cells, a per-anchor partition and one dollar sum is a dict comprehension, not a platform
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each deduction is a chain of three readings -- three consecutive Mondays -- with no branching and exactly one edge between consecutive runs, carrying four scalars. Different deductions never touch. A framework would add an orchestrator to a for-loop that already runs 50 chains wide.
The other sideWhat a framework costs you
No scheduler. This is a watch that carries state and is still INVOKED, not woken -- and this kit measured what that gap costs rather than describing it: $453,621.42 for one skipped weekly run. That is the honest cost of staying a folder of readable Python, and it is stated on the kit's own page rather than implied by silence.
No missed-run detection, no back-fill, no late marking. All three are what a scheduling engine gives you and none of them exists here.
No concurrency model for the state store. data/state.json is one file replaced atomically -- correct for one writer, and a desk with two watches running is two writers.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
The scheduling seam is the one row with a measured CONSEQUENCE and no measured ALTERNATIVE: evals/cadence.py prices the gap, and nothing here has ever run under cron, Airflow or Temporal to show that closing it is as easy as the row implies.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-deduction-age on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
answered_pct below 100, or output tokens rising without the corpus changing
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-deduction-age-calibration5,632 ms
r001-deduction-age4,974 ms
s001-deduction-age-stateless4,853 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-deduction-age-agedays, b001-deduction-age-agedaysmem, b002-deduction-age-anchormem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
deductions-queue snapshots
data/corpus/DED-<n>-R<k>.txt -- 150 files, 636931 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Customer Contact -- the customer's A/P analyst by name, her direct dial and her work email -- never does, by src/select.NEVER_SENT
the carried state
data/state.json in a deployment, and in-memory per run in the eval -- four scalars per deduction, written by src/deduction.step and never by a model
as one English sentence in every prompt, and the UI prints the same sentence verbatim so a reader can audit what the model was told
the answer key
data/gold.jsonl -- 150 rows, the output of src/deduction.step over the planted terms, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the schedule arithmetic
results/cadence-deduction-age.json -- what one skipped run costs and what five other cadences cost, computed from the deduction model
never -- evals/cadence.py opens no socket and needs no key
the run records
results/eval-*.json in the kit, and one small record per run in the app repo's run register
never -- they are read by the site build, not by a provider
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a SEVEN-DAY watch: every Monday at 08:00 local, 52 runs a year. One run owns exactly the window since the last one -- it is the run that must lodge any dispute whose window closes inside it. The cadence is not a deployment detail here: BOTH posture bands are DEFINED as the cadence (FILE_BAND_DAYS is the interval, WATCH_BAND_DAYS is twice it), so changing the schedule changes what the answers mean as well as what they cost.
50 deductions x 3 scheduled runs = 150 calls per arm, 0 gaps and 0 out-of-order readings; evals/check_labels.py asserts all 50 sequences are complete before a run may spend. Wall clock 215.3s at 12 workers. A desk of 4,000 open deductions on this cadence is 4,000 calls a week, $12.01 on the shared projection card. (r001-deduction-age, evals/cadence.py, evals/check_labels.py, src/deduction.CADENCE_DAYS)
⚑ WHAT A MISSED RUN COSTS IS MEASURED HERE, NOT ARGUED, AND IT IS THE LARGEST FIGURE THIS KIT PUBLISHES. Skip one weekly run and 30 of the 34 last-chance filings are lost PERMANENTLY -- $453,621.42 of $491,143.79. Per run: run 1 -> 11 lost, $165,072.59; run 2 -> 10 lost, $113,385.89; run 3 -> 9 lost, $175,162.94. It is permanent rather than late because the FILE band IS the interval: the run that owns a deadline is the only run that can act on it, and by the next one Rule Q-1 has made the deduction a write-off. ⚠︎ AND THE OBVIOUS RESPONSE IS THE WRONG ONE. Running MORE often barely helps: the same script sweeps 1, 3, 7, 14 and 30-day schedules over the same population and finds 1 claims lost at monthly and 0 at every other interval. Monthly is 4.5x cheaper and costs one claim worth $3,278.53. The schedule risk here is reliability, not frequency.
every accuracy figure on this page is per a seven-day watch. Both posture bands move with the schedule, so a daily or monthly watch is not a comparable run of the same question -- and no arm has been run at one. The SCHEDULE figures, by contrast, are swept and published for five cadences because they cost nothing to compute.
state
four scalar fields per deduction -- whether a dispute is already lodged, the countdown as at the previous run, the posture last reported, and how many runs have seen it -- written by src/deduction.step from figures PARSED off the page and never from the model's reply, and rendered by src/state.describe into one English sentence. That sentence is the entire route from one scheduled run to the next, and it is what both scored arms differ by.
149 characters on the worked example. Input tokens 255689 with it against 252243 without, and OUTPUT 107470 against 84367 -- so the carried state COSTS money here: $0.00300170 a reading with it against $0.00252815 without. What it buys: posture 96.67 pct against 72.67, 33 live claims no longer reported as write-offs, duplicate filings 0 against 3, at-risk error 0.22 pct against 71.39. What it does NOT buy: the countdown, 96.67 against 96.67. (r001-deduction-age against s001-deduction-age-stateless, lenses.LLM.prompt_parts)
four scalars, so run 52 costs what run 2 costs -- the OPPOSITE curve to an intake kit, and it matters here because a disputed deduction can sit in a queue for a quarter. What is NOT bounded is accuracy over a longer history, which is unmeasured, and the store itself: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model. ⚠︎ AND ONE OF THE FOUR SCALARS IS A KNOWN TRADE-OFF: carrying last week's countdown lets a reading reach the right number by subtracting 7 without re-reading the clause, and the control arm -- which cannot -- scores 96.67 pct on the anchor against this arm's 95.33. Two cells. Reported, not claimed.
lose the carried state and the posture collapses -- 72.67 pct against 96.67, with 33 live disputes reported as write-offs and an at-risk total 71.39 pct out -- while every reply stays well-formed and nothing raises an error. See environment.signatures' traceless row.
model
one completion call per reading, on one provider and one key, through src/adapters/__init__.py. MAX_TOKENS = 24000; thinking is never sent, so every published run left provider-side reasoning at the default and the result files record thinking: null.
150 calls on the scored arm, 107470 output tokens, 94782 of them provider-side reasoning (88.2 pct). Largest reply 14538 tokens against the 24000 ceiling. p50 5.0s, p95 9.8s. 0 unparsed. (r001-deduction-age, c000-deduction-age-calibration)
⚠︎ THE CALIBRATION UNDER-ESTIMATED THIS BY SIX TIMES AND THE GENEROUS CEILING IS WHY NOTHING WAS LOST. c000 fired two deduction chains at an 8,000-token cap and topped out at 1004 output tokens, all six parsed -- which reads as 24x headroom against the published 24000. The scored run's largest reply was 14538. A ceiling sized from that calibration would have been below the hardest readings, and a ceiling is not a cost: the provider bills tokens produced, not tokens allowed. Calibrating on two chains is the defect; keeping the high ceiling is what covered it.
a provider whose reasoning default differs reprices this kit without changing anything a reader can see, and nothing here has measured what disabling reasoning does to the answers. Only ONE model was run and no cross-tier claim is made anywhere on this page.
labels
150 readings -- 50 open deductions re-read at 3 weekly runs -- with the four answers computed by tools/build_corpus.py from the deduction model and re-derived from src/deduction.step by evals/check_labels.py before any run may spend. 39 of them are memory-dependent; 22 have no clause or no anchor date on file and must be aged against nothing; 34 are the last run that can lodge a dispute.
150 rows in data/gold.jsonl, 0 replay mismatches, 0 sequence gaps. Posture spread: CLEAR 22, CONTEXT_INCOMPLETE 22, EXPIRED 13, FILED 39, FILE_NOW 34, WATCH 20. Anchor spread: DEDUCTION_DATE 51, INVOICE_DATE 36, NONE 22, POD_DATE 41. 13 of the 50 customers carry a clause that OVERRIDES the type default, and evals/check_labels.py asserts all six deduction types carry more than one anchor -- without that the whole task collapses to a lookup. (data/gold.jsonl, evals/check_labels.py, data/corpus-stats.json)
it stops scoring at the third weekly run of each deduction. Long enough for a dispute lodged in week 1 to have to be remembered in week 3, and not long enough to test a claim that has sat in a queue for a quarter -- which is where a wrong anchor does its real damage, because the error does not grow, the deadline simply arrives.
the labels encode ONE reading of one clause, and the model disagrees with that reading on five cells -- arguably correctly. A corpus whose clauses were all unambiguous would move the anchor headline; this one does not have that property and the page publishes the lower figure.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a deduction reported EXPIRED whose dispute was in fact lodged in time
the reading was made without the carried state. Nothing on a snapshot says a dispute exists, so a reader who cannot see the previous run sees a closed window and calls a live claim a write-off
check what the carried state said before writing anything off. The stateless control makes this exact confusion 33 times in 150 readings; the stateful arm makes it 0 (results/eval-s001-deduction-age-stateless.json against results/eval-r001-deduction-age.json)
a countdown that is weeks out on one KIND of deduction and exact on the others
something is ageing from the wrong date for part of the population. The window runs from an anchor the clause names, and a rule keyed on days-since-posted is right for the remittance-anchored claims and wrong for everything else
read the per-anchor cut, not the aggregate. The incumbent ageing rule scores 34.0 pct whole -- and 0.0 / 0.0 / 100.0 across the delivery, invoice and remittance slices. The aggregate looks like a tuning problem; the slices say two thirds of the queue is wired to the wrong date (results/eval-b000-deduction-age-agedays.json, by_anchor)
a deduction ageing against a window on an agreement with no dispute clause
something filled a missing window with a default, or substituted an anchor date it could find for the one the clause asked for. That is the one thing this row's guardrail forbids: a guessed window produces a filing deadline nobody can defend
look at the Dispute window granted line and the three candidate dates. Both floors that guess get all 22 of these readings wrong; both arms that abstain get all 22 right (results/eval-b000-deduction-age-agedays.json against results/eval-r001-deduction-age.json)
nothing at all -- the worklist simply does not change between two weeks
TRACELESS. A scheduled run that never fired leaves no artefact anywhere in this kit: no error, no gap marker, no late flag. The readings it would have produced simply do not exist, and the deductions whose windows closed inside that week are gone
there is nothing on the worklist to look at, so look at the SCHEDULER instead: compare the number of readings this kit produced against the number the cadence says it should have (open deductions x runs). On this corpus that is 50 x 3 = 150 and evals/check_labels.py asserts it; in a deployment nothing does. ⚑ AND THE PRICE OF THAT GAP IS MEASURED RATHER THAN FEARED: one skipped weekly run is 30 filings and $453,621.42, permanently (results/cadence-deduction-age.json, computed from the deduction model with no provider and no key)
['Concurrency. data/state.json is replaced atomically, which is correct for one writer; two watches on one queue is two writers and nothing here has tested it. Here a lost write does not degrade an answer, it re-files every claim in the queue.', 'A run that fires LATE rather than not at all. evals/cadence.py models a run that did not happen; a run that happens two days after its slot is a different question.', 'A cadence other than seven days AS A SCORING QUESTION. The schedule consequences are swept for five cadences and cost nothing; no ARM has been scored at any other interval, and both posture bands move with it.', "A population that changes between runs -- a deduction first seen mid-window, or one that settles or is rejected between runs. A re-filing after a customer rejects a dispute is a legitimate second filing that this kit's vocabulary cannot express.", 'Provider-side retention. The prompt carries no personal data by construction, but what the provider keeps of a request is outside this repository and nobody here has verified it.', 'GPU sizing, local inference and anything about running this off a hosted API. Not attempted; not costed.', 'A second model. One tier was run.', "Whether the Deductions Desk Notes field can move file_dispute. The surface is sent deliberately and was never attacked -- and one of the three shipped notes asks, in the customer's voice, for exactly the omission that costs the money."]
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every customer, deduction, trade agreement, dispute clause, window length, amount and activity-log line is invented here. Verified against the repository's own LICENSE file on 2026-08-23. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineThe posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
For each of the 150 readings and each of the four answered fields, did the reply equal the computed answer key? The posture, the anchor and the file/hold call are compared exactly against closed lists; the countdown is compared as a whole number of days, and a null MATCHES a null because "no deadline can be computed" is the correct answer on 22 readings and must not be scored as a parse failure.
$0.00per 1,000 deduction records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function all three free floors are scored through.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DED-0026-R2 -- weekly run 2 of 3, 2026-03-09, a DAMAGE deduction of 9,302.06 USD on a 90-day dispute window
What the previous scheduled run left behind
No dispute has been lodged against this deduction yet. At the previous scheduled run, 9 day(s) remained on the dispute window. It was reported WATCH.
The agreement's dispute clause, and what it does not say
"Disputes must quote the original invoice number and be submitted inside the window stated above. The customer's claims portal is the only accepted channel." -- it says what a dispute must CONTAIN and sets no anchor at all, so Rule Q-3's type default governs and a DAMAGE claim runs from the proof of delivery.
The answer key
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What run r001-deduction-age answered
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What the free floor answered
EXPIRED, anchored on INVOICE_DATE, -1 days, file_dispute NO
Why this one
It is the shape where a wrong anchor does not merely change a number -- it turns a live claim with two days left into a write-off, which is the costly direction and is not recoverable on any later run.
Grader
Verdict
Why
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
posture hit, anchor hit, countdown hit, file/hold hit -- four of four
DED-0026-R2, run 2 of 3. The window is 90 days; the page prints 91 days elapsed from the invoice, 88 from the delivery and 74 from the remittance. The clause mentions the invoice but does not anchor on it, so Rule Q-3 puts a DAMAGE claim on the delivery: 90 minus 88 leaves 2, inside the 7-day file band, and the carried state says nothing is lodged yet -- so this run files. The model answered all four correctly and its rationale names the rule it applied.
The two filing directions, counted apart and in dollars
in scope on the FILE side, and a hit -- one of the 34
This reading's correct answer is to lodge, so it sits inside the 34 that the missed-filing rate is measured over and outside the 116 the duplicate rate is measured over. The model lodged it. The free floor did NOT: having anchored on the invoice date it lands at -1 day, reports EXPIRED and files nothing. That is 9,302.06 USD written off on a claim with two days still on it, and it is one of the 3 such claims worth $19,683.89 that the floor loses.
The same comparison, cut per anchor -- the grader this kit exists for
in scope on the PROOF OF DELIVERY slice -- one of 41
This reading's correct anchor is the proof of delivery, so it sits in the slice where the model scores 87.8 pct and the free floor 92.68 pct. It is also the slice where the incumbent ageing rule scores 0.0 pct -- every one of its 41 readings wrong, because that rule ages from the remittance and nothing else.
The readings whose answer is not on their own page
not in scope
This grader scores only the 39 readings where a dispute had already been lodged before the reading began, so nothing on the page could reveal the correct posture. DED-0026-R2's carried state says nothing is lodged, so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The deduction's NEXT reading, DED-0026-R3, IS in scope: by then the dispute exists and only the carried state knows.
The formulaWhat it computes
accuracy = hits / 150 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
96.7% posture accuracy · 3 more measured on this row
the fast tier, memory removed (THE CONTROL)
72.7% posture accuracy · 3 more measured on this row
the strongest free floor, no model
92.7% posture accuracy · 3 more measured on this row
the deductions ageing report an ERP already prints, no memory
32.7% posture accuracy · 3 more measured on this row
the same ageing rule, GIVEN the carried state
58.7% posture accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by tools/build_corpus.py over the planted terms at generation time and re-derived from src/deduction.step by evals/check_labels.py before any run may spend. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known to be arguable -- nine readings carry a clause anchoring on "the date the carrier's signed receipt reached us", the key treats that as the delivery date, and the receipt's arrival date is a different date that this corpus never prints. The other three fields have no known ambiguity.
Watch these
dispute_window_accuracy_pct
anchor_accuracy_pct
posture_accuracy_pct
file_dispute_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Both scored arms had 0 of 150.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because two of them are small: 34 readings that must lodge a dispute and 22 with nothing on file, so one row moves the first by 2.9 points and the second by 4.5.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/deduction.py, src/segment.py or src/select.py. Re-run the scored eval AND its stateless control together (paid, 300 calls) on any change to src/prompt.py or src/state.py -- the memory claim is a difference, so one arm re-run alone is not comparable with the other's old figure. All three floors and evals/cadence.py are free and should be re-run on any change at all.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real deductions desk, where which date a clause anchors on is settled by a person reading a contract and sometimes by the customer disagreeing. That is why this corpus is generated rather than captured.
The same comparison, cut per anchor -- the grader this kit exists for
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineThe same comparison, cut per anchor -- the grader this kit exists for
The identical exact-match comparison, restricted to the readings whose CORRECT anchor is the proof of delivery (41), the original invoice (36) or the remittance (51). It answers the one question the aggregate cannot: is this arm uniformly mediocre, or is it wired to the wrong date for part of the population?
$0.00per 1,000 deduction records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, same pass, grouped by the anchor tools/build_corpus.py resolved at generation time.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DED-0026-R2 -- weekly run 2 of 3, 2026-03-09, a DAMAGE deduction of 9,302.06 USD on a 90-day dispute window
What the previous scheduled run left behind
No dispute has been lodged against this deduction yet. At the previous scheduled run, 9 day(s) remained on the dispute window. It was reported WATCH.
The agreement's dispute clause, and what it does not say
"Disputes must quote the original invoice number and be submitted inside the window stated above. The customer's claims portal is the only accepted channel." -- it says what a dispute must CONTAIN and sets no anchor at all, so Rule Q-3's type default governs and a DAMAGE claim runs from the proof of delivery.
The answer key
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What run r001-deduction-age answered
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What the free floor answered
EXPIRED, anchored on INVOICE_DATE, -1 days, file_dispute NO
Why this one
It is the shape where a wrong anchor does not merely change a number -- it turns a live claim with two days left into a write-off, which is the costly direction and is not recoverable on any later run.
Grader
Verdict
Why
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
posture hit, anchor hit, countdown hit, file/hold hit -- four of four
DED-0026-R2, run 2 of 3. The window is 90 days; the page prints 91 days elapsed from the invoice, 88 from the delivery and 74 from the remittance. The clause mentions the invoice but does not anchor on it, so Rule Q-3 puts a DAMAGE claim on the delivery: 90 minus 88 leaves 2, inside the 7-day file band, and the carried state says nothing is lodged yet -- so this run files. The model answered all four correctly and its rationale names the rule it applied.
The two filing directions, counted apart and in dollars
in scope on the FILE side, and a hit -- one of the 34
This reading's correct answer is to lodge, so it sits inside the 34 that the missed-filing rate is measured over and outside the 116 the duplicate rate is measured over. The model lodged it. The free floor did NOT: having anchored on the invoice date it lands at -1 day, reports EXPIRED and files nothing. That is 9,302.06 USD written off on a claim with two days still on it, and it is one of the 3 such claims worth $19,683.89 that the floor loses.
The same comparison, cut per anchor -- the grader this kit exists for
in scope on the PROOF OF DELIVERY slice -- one of 41
This reading's correct anchor is the proof of delivery, so it sits in the slice where the model scores 87.8 pct and the free floor 92.68 pct. It is also the slice where the incumbent ageing rule scores 0.0 pct -- every one of its 41 readings wrong, because that rule ages from the remittance and nothing else.
The readings whose answer is not on their own page
not in scope
This grader scores only the 39 readings where a dispute had already been lodged before the reading began, so nothing on the page could reveal the correct posture. DED-0026-R2's carried state says nothing is lodged, so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The deduction's NEXT reading, DED-0026-R3, IS in scope: by then the dispute exists and only the carried state knows.
The formulaWhat it computes
accuracy = hits / slice size, per anchor, on all four fields. The slices partition the 128 readings whose anchor is not NONE and never overlap.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
87.8% pod window · 2 more measured on this row
the strongest free floor
92.7% pod window · 2 more measured on this row
the deductions ageing report an ERP already prints
0.0% pod window · 2 more measured on this row
the fast tier, memory removed (THE CONTROL)
87.8% pod window · 2 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, partitioned by the anchor the generator resolved. Scored against the reference; not itself the reference.
These rates are UNKNOWN, on purpose
Its own rates are blank because it is a PARTITION of the reference's verdicts rather than a second opinion on them; an agreement rate between the two would be 1.0 by construction.
Watch these
pod_window_pct
invoice_window_pct
remittance_window_pct
Alarm on
any slice more than 10 points below the aggregate. That is the signature of a rule wired to the wrong date for part of the population, and it is invisible on a single number.
How tight can the band be? The three slices are 41, 36 and 51 readings, so the finest bands they support are 2.4, 2.8 and 2.0 points. A one-cell difference is not a finding on any of them; a hundred-point gap is.
Cadence: Free, every run. Re-derive the slices whenever tools/build_corpus.py changes -- they are a property of the corpus, and a corpus whose anchors were a function of the deduction type would make this whole grader meaningless (evals/check_labels.py asserts they are not).
The decisionWhen to reach for it
Use it
One rule is applied to a population that is not homogeneous in the thing the rule turns on. That is this kit exactly, and it is far more common than the single-number board admits.
Do not use it
A population where the rule really is uniform. There is nothing for the slices to disagree about and three identical numbers say so.
The two filing directions, counted apart and in dollars
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineThe two filing directions, counted apart and in dollars
Two counts over two different denominators, and one of them is also counted in money. A MISSED filing: of the 34 readings whose correct answer is to lodge the dispute, how many did not -- and what were they worth, because on a hard clock a missed filing is a write-off rather than a late alert. A DUPLICATE filing: of the 116 readings that must not lodge one, how many did.
$0.00per 1,000 deduction records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in the same pass as the exact-match grader.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DED-0026-R2 -- weekly run 2 of 3, 2026-03-09, a DAMAGE deduction of 9,302.06 USD on a 90-day dispute window
What the previous scheduled run left behind
No dispute has been lodged against this deduction yet. At the previous scheduled run, 9 day(s) remained on the dispute window. It was reported WATCH.
The agreement's dispute clause, and what it does not say
"Disputes must quote the original invoice number and be submitted inside the window stated above. The customer's claims portal is the only accepted channel." -- it says what a dispute must CONTAIN and sets no anchor at all, so Rule Q-3's type default governs and a DAMAGE claim runs from the proof of delivery.
The answer key
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What run r001-deduction-age answered
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What the free floor answered
EXPIRED, anchored on INVOICE_DATE, -1 days, file_dispute NO
Why this one
It is the shape where a wrong anchor does not merely change a number -- it turns a live claim with two days left into a write-off, which is the costly direction and is not recoverable on any later run.
Grader
Verdict
Why
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
posture hit, anchor hit, countdown hit, file/hold hit -- four of four
DED-0026-R2, run 2 of 3. The window is 90 days; the page prints 91 days elapsed from the invoice, 88 from the delivery and 74 from the remittance. The clause mentions the invoice but does not anchor on it, so Rule Q-3 puts a DAMAGE claim on the delivery: 90 minus 88 leaves 2, inside the 7-day file band, and the carried state says nothing is lodged yet -- so this run files. The model answered all four correctly and its rationale names the rule it applied.
The two filing directions, counted apart and in dollars
in scope on the FILE side, and a hit -- one of the 34
This reading's correct answer is to lodge, so it sits inside the 34 that the missed-filing rate is measured over and outside the 116 the duplicate rate is measured over. The model lodged it. The free floor did NOT: having anchored on the invoice date it lands at -1 day, reports EXPIRED and files nothing. That is 9,302.06 USD written off on a claim with two days still on it, and it is one of the 3 such claims worth $19,683.89 that the floor loses.
The same comparison, cut per anchor -- the grader this kit exists for
in scope on the PROOF OF DELIVERY slice -- one of 41
This reading's correct anchor is the proof of delivery, so it sits in the slice where the model scores 87.8 pct and the free floor 92.68 pct. It is also the slice where the incumbent ageing rule scores 0.0 pct -- every one of its 41 readings wrong, because that rule ages from the remittance and nothing else.
The readings whose answer is not on their own page
not in scope
This grader scores only the 39 readings where a dispute had already been lodged before the reading began, so nothing on the page could reveal the correct posture. DED-0026-R2's carried state says nothing is lodged, so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The deduction's NEXT reading, DED-0026-R3, IS in scope: by then the dispute exists and only the carried state knows.
The formulaWhat it computes
missed_filing_pct = missed / 34; duplicate_filing_rate_pct = duplicates / 116; missed_filing_usd = the summed face value of the missed claims. Never averaged and never combined into an F-score.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
2.9% missed filing · 1 more measured on this row
the fast tier, memory removed (THE CONTROL)
0.0% missed filing · 1 more measured on this row
the strongest free floor
8.8% missed filing · 1 more measured on this row
the deductions ageing report, no memory
64.7% missed filing · 1 more measured on this row
the same rule, GIVEN the carried state
64.7% missed filing · 1 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl's file_dispute column, computed by src/deduction.step. Scored against the reference and not itself the reference.
These rates are UNKNOWN, on purpose
Its own TPR and TNR are not measured because it does not answer a true/false question about another grader's verdict -- it counts two directions of one field against the key.
Watch these
missed_filing_pct
duplicate_filing_rate_pct
missed_filing_usd
Alarm on
missed_filing_usd above zero. A duplicate is visible to whoever reads the worklist and to the customer's claims desk; a missed filing is invisible until a month-end write-off report, by which point the window has been shut for weeks.
How tight can the band be? 34 readings must lodge a dispute and 116 must not, so the finest band either rate can support is 2.9 and 0.9 points respectively. A "one missed filing" claim on this corpus is a claim about 34 readings and nothing wider.
Cadence: Free. Re-run with every scored run, and re-run the floors whenever src/state.py changes -- the duplicate count is the most memory-sensitive figure on this page.
The decisionWhen to reach for it
Use it
The two failure directions cost different things and are fixed by different people, which is exactly this kit's situation.
Do not use it
A task where one error direction is free. Here neither is: a missed filing is money that cannot be recovered, a duplicate is two claims for one charge on the customer's desk.
The readings whose answer is not on their own page
Catch a customer deduction before its window closes
PresenterOpens the private repo. Visible to admins only.
In one lineThe readings whose answer is not on their own page
The same exact-match comparison, restricted to the 39 readings where a dispute had ALREADY been lodged before the reading began. Nothing on those pages says so, so the correct posture (FILED, not EXPIRED and not FILE_NOW) is reachable only through the carried state. It answers one question -- is the memory doing anything, or is the corpus easy?
$0.00per 1,000 deduction records
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, same pass, grouped by the flag tools/build_corpus.py set at generation time -- before the step that produced the answer, so it records what was known going in.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
DED-0026-R2 -- weekly run 2 of 3, 2026-03-09, a DAMAGE deduction of 9,302.06 USD on a 90-day dispute window
What the previous scheduled run left behind
No dispute has been lodged against this deduction yet. At the previous scheduled run, 9 day(s) remained on the dispute window. It was reported WATCH.
The agreement's dispute clause, and what it does not say
"Disputes must quote the original invoice number and be submitted inside the window stated above. The customer's claims portal is the only accepted channel." -- it says what a dispute must CONTAIN and sets no anchor at all, so Rule Q-3's type default governs and a DAMAGE claim runs from the proof of delivery.
The answer key
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What run r001-deduction-age answered
FILE_NOW, anchored on POD_DATE, 2 days left, file_dispute YES
What the free floor answered
EXPIRED, anchored on INVOICE_DATE, -1 days, file_dispute NO
Why this one
It is the shape where a wrong anchor does not merely change a number -- it turns a live claim with two days left into a write-off, which is the costly direction and is not recoverable on any later run.
Grader
Verdict
Why
The posture, the anchor, the countdown and the file/hold call, per reading, exact match against the computed answer key
posture hit, anchor hit, countdown hit, file/hold hit -- four of four
DED-0026-R2, run 2 of 3. The window is 90 days; the page prints 91 days elapsed from the invoice, 88 from the delivery and 74 from the remittance. The clause mentions the invoice but does not anchor on it, so Rule Q-3 puts a DAMAGE claim on the delivery: 90 minus 88 leaves 2, inside the 7-day file band, and the carried state says nothing is lodged yet -- so this run files. The model answered all four correctly and its rationale names the rule it applied.
The two filing directions, counted apart and in dollars
in scope on the FILE side, and a hit -- one of the 34
This reading's correct answer is to lodge, so it sits inside the 34 that the missed-filing rate is measured over and outside the 116 the duplicate rate is measured over. The model lodged it. The free floor did NOT: having anchored on the invoice date it lands at -1 day, reports EXPIRED and files nothing. That is 9,302.06 USD written off on a claim with two days still on it, and it is one of the 3 such claims worth $19,683.89 that the floor loses.
The same comparison, cut per anchor -- the grader this kit exists for
in scope on the PROOF OF DELIVERY slice -- one of 41
This reading's correct anchor is the proof of delivery, so it sits in the slice where the model scores 87.8 pct and the free floor 92.68 pct. It is also the slice where the incumbent ageing rule scores 0.0 pct -- every one of its 41 readings wrong, because that rule ages from the remittance and nothing else.
The readings whose answer is not on their own page
not in scope
This grader scores only the 39 readings where a dispute had already been lodged before the reading began, so nothing on the page could reveal the correct posture. DED-0026-R2's carried state says nothing is lodged, so everything needed is on the snapshot and the row is outside its denominator entirely. It is shown rather than omitted, because a reader watching one reading receive every verdict has to be able to see which graders had nothing to say about it. The deduction's NEXT reading, DED-0026-R3, IS in scope: by then the dispute exists and only the carried state knows.
The formulaWhat it computes
accuracy = hits / 39 per field, over the memory_dependent subset only.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
92.3% memory posture accuracy · 2 more measured on this row
the fast tier, memory removed (THE CONTROL)
0.0% memory posture accuracy · 2 more measured on this row
the strongest free floor
100.0% memory posture accuracy · 2 more measured on this row
the deductions ageing report, no memory
0.0% memory posture accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, sliced by the memory_dependent flag. Scored against the reference; not itself the reference.
These rates are UNKNOWN, on purpose
Its own rates are blank because it is a SLICE of the reference's verdicts rather than a second opinion on them.
Watch these
memory_posture_accuracy_pct
memory_window_accuracy_pct
memory_file_accuracy_pct
Alarm on
memory_posture_accuracy_pct falling below the whole-population posture figure. That would mean the carried sentence is confusing the model rather than informing it.
How tight can the band be? 39 cells, so the finest band this subset supports is 2.6 points.
Cadence: Free, every run. Re-derive the flag whenever tools/build_corpus.py changes -- a corpus with no memory-dependent readings would make this kit unfalsifiable (evals/check_labels.py asserts there are at least 30).
The decisionWhen to reach for it
Use it
A kit carries state between runs and has to show the state is load-bearing rather than decorative.
Do not use it
A stateless task. There is nothing for the subset to mean.
A living map of modern AI — kept current every morning