Catch unbilled power accounts before the charge expires
Some power accounts go months without a bill, and past a point the charge can no longer be collected. This app checks each account against that deadline and flags the ones running out of time.
PresenterOpens the private repo. Visible to admins only.
For the billing supervisorEnergy & Utilities · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A billing supervisor at an energy retailer, working the unbilled-account queue each month.
✕Today's manual process
1Open each account manually and find the oldest meter read that was never billed, not just the last bill sent.
2Check the exception log cycle by cycle, for a stop the customer caused, so the clock isn't held against them by mistake.
3Decide by memory whether this account is still safe, worth chasing, or already too late to bill.
4Miss one stoppage and a recoverable bill gets written off, or a lost one gets chased for nothing.
Every account judged from memory
✓With the app
1The account is read for you and the oldest unbilled meter read is pulled out, not the last bill sent.
2Every earlier stoppage is checked and only the days the customer caused are taken off the clock.
3A risk level comes back safe, worth chasing, or already past saving, with the reason stated.
4Nothing is missed every account gets the same check, so a stoppage is never overlooked and nothing is wrongly written off.
Every account judged the same way
See it work
One real case, read by the app, step by step
Account ACC-0024-C3 carries 319 days from last cycle and now shows fifteen days of headroom.
Catch unbilled power accounts before the charge expiresReference appBuilt to be shaped to your process
4
1Carried from last cycle 319 days already counted, with no suspension left open.
2This cycle's days 31 new days added, for 350 of the 365 days used.
3The blocker an exception hold, with its own lead time to clear.
4The outcome approaching: fifteen days left, under twice the 14 needed to clear it.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch unbilled power accounts before the charge expires
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
'Which accounts are unbilled beyond cycle' is a SELECT and nobody needs a language model for it. The question a billing operation actually has to answer is the next one: how close is each of them to the point where the charge cannot be recovered at all, and how much of this month's unbilled balance is already past it. That is not the elapsed time since the last bill -- on all 150 statements here the last bill is an estimate issued weeks ago that moves nothing. It is the time since the oldest UNBILLED ACTUAL meter read, minus every stretch in which the delay was the customer's own, on an account still inside the protection, judged against the remediation lead time for its blocker. Three of those four terms can sit in a cycle whose statement is not on your screen. Somebody working an unbilled-account queue by hand: opening each account, finding the oldest read that was actually billed rather than the last bill issued, going back through earlier cycles' exception entries to work out how much of the delay was the customer's own, and only then deciding whether this is an account to chase this week or one the charge is already lost on -- and totalling the lost ones for the month-end pack.
Audience
Billing operations and revenue assurance in an energy retailer, and the finance function that receives the unbilled-exposure number at month end -- anyone who owns a back-billing window with a suspension rule in it, and anyone deciding whether a language model has any business near one. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual cycle statements
The corpus is 150 cycle statements, 0.53 MB (json 1 · jsonl 1 · txt 150). AN UNBILLED-ACCOUNT WATCHLIST IS A LIST OF HOUSEHOLDS. Not contracts -- homes: who lives there, where the meter is, how much energy they have used that nobody has billed them for, and an agent's note about why. There is no public corpus of that and there never will be one, for the same reason there is no public corpus of anyone's bank statements; a supplier's own unbilled book is among the most sensitive operational files it holds, because it joins consumption to identity to money owed. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/backbill.replay's output over the planted patterns, so an interval rule this fiddly cannot carry its author's misreading into the score. ⚠︎ AND IT SHIPS BLOCKED PENDING ANCHOR. BACKBILL_WINDOW_DAYS = 365, all six remediation lead times, the segment threshold and the dispute rule are INVENTED for this kit. They reproduce no statute, no regulator's guidance, no licence condition, no industry code and no supplier's billing policy, and name none. The catalogue row this kit was built from records the real window value as an OPEN QUESTION shared with a sibling row -- 'one clock, not two independently answered ones' -- and nobody here has confirmed it. The rule text is reproduced in full on all 150 pages precisely so it can be read, disbelieved and replaced.
The corpus
The 150 cycle statementsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your cycle statements. That is the whole change — there is no database to migrate.
One cycle statement, as the model receives itACC-0001-C1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes no real
customer, household, address, meter, supplier or charge. Account ACC-0001, cycle C1.
Account
----------------------------------------------------------------
Account : ACC-0001
Segment : domestic (inside the back-billing protection)
Product : Prepayment converted to credit, electricity
Meter type : smart, non-communicating
Supply since : 2020-05-14
Billing cycle : monthly, cycle end on day 20
Back-Billing Rule
----------------------------------------------------------------
Back-billing window : 365 days
Remediation lead time published for each blocker:
EXCEPTION_HOLD 14 days
STUCK_ORDER 7 days
METER_PROBLEM 45 days
NO_ACCESS 60 days
READ_DISPUTE 30 days
CAUSE_UNKNOWN 30 days
1. A charge for energy consumed more than 365 days before the date a valid bill is issued cannot
be recovered from the customer and must be written off.
2. The 365 days run from the date of the OLDEST UNBILLED ACTUAL METER READ. They do not run from
the last bill issued, and they do not run from the date the billing exception was raised.
3. The window is SUSPENDED only for a period the Exception Log records as a customer-attributable
suspension -- access to the meter refused or not arranged, or a reading the customer undertook
to provide and did not. A suspension runs from the day it is OPENED up to but not including the
day it is CLOSED, and days inside it do not count towards the window. A standing block whose
recorded reason sounds customer-attributable does NOT suspend the window unless a suspension
was opened for it.
Abridged — the file continues.
The outcomeWhat a good result looks like
Per cycle: the proximity band after the suspension rule has been applied, whether the back-billing clock is RUNNING, SUSPENDED or NOT_PROTECTED, how many days now count towards the window, and which blocker is standing -- plus a four-field carried accrual the next cycle is judged against, written by code from the parsed dates and never from the model's reply.
And when it cannot
It names the wrong blocker, and on this run the wrong blocker is what produces the only wrong band. 13 of 150 cycles carry a blocker code the key disagrees with; 10 of them answer CAUSE_UNKNOWN where a code was mappable and 3 answer NO_ACCESS on a sentence about a reading the customer never sent. Because the lead time is what turns a day count into a band, a mis-mapped blocker with a different lead time silently moves the band -- which is exactly what happened once. Under- and over-stating proximity cost differently and are never averaged: this run reported 0 of 54 watchlist cycles as quiet (none) and raised 1 false alarm on 96 quiet cycles.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Sorting an unbilled queue where nothing was ever suspended and every account's history is on its own page — the free one-page floor, at $0.00 The floor gets the band right on 86.67 pct of all cycles and the blocker code on 98.0 pct -- BETTER than the model's 91.33. On the clean-drift, suspension-inside-one-cycle and not-protected-by-segment patterns it is exactly right, because the elapsed calendar IS the risk clock there.
An account whose suspension began in a cycle you are not holding -- a CLOSED entry with no OPENED, or an empty exception log — the fast tier WITH the carried accrual 100.0 pct of the bands and 92.0 pct of the day counts on the 25 memory-dependent cells, against the floor's 20.0 pct and 0 pct. This is the half the accrual plainly buys, and it is the difference between AT_RISK at 360 days and BREACHED at 396 on the same statement.
Producing the month-end unbilled-exposure figure for finance — the fast tier WITH the carried accrual, with the per-account list published beside the total The total is exact ($29,024.35 of $29,024.35, 0.0 pct error) where the floor is 49.39 pct OVER and the stateless arm 65.98 pct UNDER. ⚠︎ But publish the list too: this grader is where errors cancel, and a total can be right while rows are wrong.
Deciding to write a balance off — a person, reading the band, the day count and the carried accrual the kit printed The highest-consequence output here is BREACHED, and both no-memory arms get the total wrong in opposite directions -- the floor by 49.39 pct over, the stateless arm by 65.98 pct under. Nothing in this kit writes anything off and nothing should be added that does.
And where nothing here is good enough:
Deciding which exception code an account's recorded reason belongs to — neither, as built -- free code beats the model and both are worse than a curated list The floor scores 98.0 pct and the model 91.33 pct (96.0 pct under the most generous reading of its own answers). The model invented ten CAUSE_UNKNOWNs on reasons that map cleanly, while getting 18 of 18 genuinely unrecorded ones right. And the catalogue row already records the real taxonomy as unconfirmed, so this field is measuring agreement with an invented list.
An account with years of history rather than three cycles — nothing here yet -- measure it first The accrual is four scalar fields, so the COST does not grow. The accuracy is unmeasured: every rule in this corpus resolves inside three cycles, and the one error shape this run has (the half-open boundary, one day out) is exactly the shape a longer history gives more chances to accumulate -- on a RUNNING TOTAL, one day out per cycle is forty days out after forty cycles.
At a glanceHow the whole thing runs
99%band accuracy pct
12,343 msp50, end to end
$6.88per 1,000 cycle statements · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch unbilled power accounts before the charge expires14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/backbill.py FIRST, and not as an option. Every figure here is a property of a three-cycle corpus of 50 invented households, measured once.Corpus lens →
When is this the wrong choice?
Avoid: Paying a model to do date subtraction a regex already does perfectly. That is the case against the best-fitting scenario (“Sorting an unbilled queue where nothing was ever suspended and every account's history is on its own page”). 6 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Three cycles per account. Long enough to hide a suspension behind a boundary; not long enough to test an accrual that has been running for years, a window that closes mid-history, or a suspension that opens and closes twice. 8 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
One run per arm, each fired once. Nothing here is a distribution: whether 99.33 pct band accuracy repeats on a second identical run is unmeasured, and the small denominators (54 watchlist cycles, 25 memory-dependent cells, 11 AT_RISK cells, 1 false alarm) make any repeat noisy. 11 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried accrual, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-unbilled-watch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes every pre-flight assertion (python3 -m evals.check_labels), scores the free one-page floor (python3 -m evals.run --run-id b000-unbilled-watch-onepage --baseline), re-scores the recorded run under the other reading of one blocker code (python3 -m evals.ambiguity) and serves the whole UI (python3 -m src.app). Five commands, plain python3, no install step -- requirements.txt names nothing because the kit imports nothing outside the standard library. What it cannot do without a key is re-run the scored eval or its stateless control.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
12,343 msp50, end to end
55,907 msp95
2 minclone to first result
What the clock covers. END-TO-END per cycle: one HTTP request carrying the assembled prompt, and the reply parsed to a band, a clock status, a day count and a blocker code. Splitting the statement into sections, dropping the one section no field asks for, parsing the dates and advancing the accrual all happen outside this clock and cost no network at all. THE TAIL IS THE STORY: p95 (55.9s) is 4.5 times p50 (12.3s). Output length here is driven by how hard the INTERVAL ARITHMETIC is, not by how long the page is -- every statement is within 6 pct of every other in size and replies ran from 282 to 12664 tokens, a 44.9x spread.
Current processWhat it replaces
Somebody working an unbilled-account queue by hand: opening each account, finding the oldest read that was actually billed rather than the last bill issued, going back through earlier cycles' exception entries to work out how much of the delay was the customer's own, and only then deciding whether this is an account to chase this week or one the charge is already lost on -- and totalling the lost ones for the month-end pack.
Where it is not good enough
⚑ THE PROXIMITY JUDGEMENT IS ESSENTIALLY SOLVED ON THIS CORPUS AND THE TAXONOMY FIELD IS NOT -- AND THE MODEL LOSES THAT ONE TO FREE CODE. Band accuracy is 99.33 pct: ONE miss in 150 cycles, and it is not a proximity failure. ACC-0034-C2 computed the day count exactly right (304 days, 61 remaining) and then applied a 60-day NO_ACCESS remediation lead time where the key says 30-day READ_DISPUTE; its own rationale states the whole chain. Every other band cell in the run is correct, so the single miss is DOWNSTREAM of the blocker-code field. BLOCKER CAUSE IS 91.33 PCT AGAINST THE FREE FLOOR'S 98.0 PCT -- the model is 6.67 points WORSE than an ordered keyword table on the one field a keyword table can do. And the 13 misses are TWO recorded phrasings, not 13 judgements: ten are 'the customer undertook to send a reading and has not done so' and three are 'cycle run completed with this account held back for checking'. THE CAUSE_UNKNOWN BUCKET BECAME A REFUGE: the model got 18 of 18 genuinely unrecorded reasons right and INVENTED TEN MORE -- a bucket for 'nobody wrote down why' is the honest way to answer an open question about a taxonomy, and it is also, measurably, where a model puts anything it does not want to commit on. ⚠︎ AND ONE OF THOSE TWO PHRASINGS IS THIS KIT'S OWN DEFECT. READ_DISPUTE is named for a disagreement; 'undertook to send a reading and has not done so' describes an omission, and a reader who declines to force it into that code is doing exactly what the prompt asks. The code was NOT renamed after the answers came back; evals/ambiguity.py prices the disagreement for $0.00 (blocker cause 91.33 pct -> 96.0 pct, seven cycles recovered, band unmoved because both codes carry the same 30-day lead time) and the LOWER figure is published everywhere. Note what survives either reading: at 96.0 pct the model STILL does not beat a keyword table. THE DAY COUNT IS OFF BY ONE, THREE TIMES. 98.0 pct exact, mean absolute error 0.02 days, worst error 1 day -- all three misses are boundary-suspension accounts. ⚠︎ AND ONE UNPLANNED REPEAT SAYS THIS IS NOT A STABLE MISREADING. The UI screenshot fires one live call on ACC-0007-C3, on a prompt verified byte-identical to the scored run's for that cycle. The scored run answered 359 days; that call answered 360, which is right. A single repeat is evidence and not a measurement -- but it is enough to withdraw 'the model is a day out on boundary suspensions' as a finding. What is measured is that 3 of 150 day counts were one day out ON THIS RUN; whether that is a convention error or sampling noise is UNRESOLVED, and resolving it costs one repeat run.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1
50 unbilled accounts, 3 consecutive billing cycles each
exposure $29,024.35 of $29,024.35 · 0 replies lost
2026-08-23as of
It produces a watchlist band, a clock status, a day count and a blocker code for a billing supervisor to read, and releases nothing — there is no endpoint, no function and no flag anywhere in the kit that issues, releases, adjusts, suppresses or writes off a charge, and evals/check_labels.py asserts it. The station that makes this a monitor is the second one: the only route from cycle 2 to cycle 3 is four scalar fields, written by src/backbill.py from the dates PARSED off the page and never from the model's answer. That matters more here than on a verdict kit, because the carried quantity is a RUNNING TOTAL — one over-count in March would still be in the number in December. ⚑ WHAT THE ACCRUAL BUYS, MEASURED AGAINST FREE PYTHON RATHER THAN AGAINST A HANDICAPPED CONTROL: band 99.33 pct against 86.67, the 54 watchlist cycles 100 pct against 81.48, the day count 98.0 pct against 83.33, and the 25 memory-dependent cells 100 pct against 20. The two no-memory arms get the month-end exposure total wrong in OPPOSITE directions — the floor 49.39 pct over, the stateless control 65.98 pct under — so 'no memory' is not a bias anybody could correct for with a constant.
⚠︎ AND THE HEADLINE HAS ONE MISS IN IT, WHICH IS NOT A PROXIMITY FAILURE: ACC-0034-C2 computed 304 days and 61 remaining exactly right and then applied the wrong blocker's lead time. The band field has no independent failure anywhere in the run.
⚠︎ THE 365-DAY WINDOW AND ALL SIX LEAD TIMES ARE INVENTED FOR THIS KIT — they reproduce no statute, no regulator's guidance, no licence condition and no supplier's policy, the catalogue row records the real value as an open question, and every band above is agreement with an invented rule.
⚠︎ AND EVERY ACCOUNT IS A HOUSEHOLD: the Customer Contact section — name, supply address, phone, email — reaches the provider in 0 of 150 documents and in 150 of 150 with the guard removed.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the carried accrual
src/state.py
load()/save() are a JSON file today. Point them at a table, a key-value store or a billing data mart and nothing else in the kit changes -- for_account() and advance() are the whole interface, and describe() is the only thing the prompt sees.
the back-billing rule
src/backbill.py
BACKBILL_WINDOW_DAYS, LEAD_DAYS and step() are your regulator's window and your operation's lead times. Change them there and re-run tools/build_corpus.py, which recomputes the whole answer key from the same function -- the gold cannot drift from the rules because it is their output. ⚠︎ THIS IS THE FIRST FILE TO REPLACE, not an optional one: the shipped window is invented.
the corpus
tools/build_corpus.py
Point it at your own accounts, tariffs and exception vocabulary, or delete it and drop real cycle statements into data/corpus/ named <ACCOUNT>-C<n>.txt with the same seven section headings. Everything downstream reads statements by account and cycle and does not care where they came from.
what is sent
src/select.py
SECTION_HINTS maps a field to the sections it needs; NEVER_SENT names what is withheld whatever happens. ⚠︎ On a domestic book this is the load-bearing seam: check NEVER_SENT covers every section your statements carry that names a person, before anything spends.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 50 accounts x 3 consecutive monthly cycles = 150 statements from a fixed seed (SEED = 20260823), across two fuels and both segments. It plants eight patterns, of which 20 accounts carry a suspension and 9 of those span a cycle boundary -- the trap. Consumption, unit rates and standing charges are drawn PER FUEL; see Data.breaks_on for the defect that rule was written from. The gold labels are src/backbill.replay's output, never typed.
the back-billing arithmetic
src/backbill.py
The window rule as pure code: days accrued, minus dated customer-attributable suspensions, against the remediation lead time for the standing blocker. No model, no judgement. ⚠︎ The window length and every lead time in it are INVENTED and the kit ships that as a stated blocker -- see Data.why_this_corpus.
the carried accrual
src/state.py
SEAM 2 -- the thing that makes this a monitor. Four fields per account (days already counted, whether a suspension is open, the band last reported, whether it is still protected), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on cycle 40 as on cycle 3.
the section splitter
src/segment.py
Splits a cycle statement into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Customer Contact -- the household's name, supply address, phone and email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-accrual sentence, and the selected sections in document order. The stateless build replaces the accrual sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost document.
the watch
src/monitor.py
One cycle, one call. Parses the cycle dates and the unbilled value off the page with a regex (the model is never asked to read a date), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
the local UI
src/app.py
One account, one cycle, its carried state and its proximity, on 127.0.0.1:9004. Renders with no key. It shows the carried sentence verbatim and the free one-page floor's answer beside the model's, so a reader can see on which rows the model bought nothing.
the free one-page floor
evals/baseline.py
Everything a reader of ONE statement can compute: the oldest unbilled actual read, the suspended days this page shows, an ordered keyword map of the standing reason, and the segment line. 0 calls, $0.00, scored through the identical scorer.
the scorer
evals/scoring.py
Exact match per cell against the computed gold, split five ways an average would hide: the four fields, the cycles on the watchlist, the quiet cycles, the memory-dependent subset, and the portfolio exposure in dollars. No judge model.
the pre-flight
evals/check_labels.py
Everything that must be true before a run may spend: 150 documents, 50 gap-free accounts, every document declaring itself synthetic on line 1, gold reproducible from backbill.replay, no statement naming another cycle, the two prompt builds differing in exactly one block, the privacy guard red-proven in both directions, per-fuel consumption inside a household's range, and no code path anywhere that releases a bill or writes a charge off.
the ambiguity re-score
evals/ambiguity.py
Re-scores the recorded run against the OTHER defensible reading of one blocker code. Free, no new calls. It exists because ten of the thirteen blocker misses are one recorded sentence whose code name does not describe it, and the honest response is to publish both numbers rather than reword the taxonomy after seeing the answers.
the run harness
evals/run.py
50 account chains, three strictly-ordered cycles each, 12 concurrent workers. The accrual advances even for a cycle whose CALL failed, so one transport error cannot turn into three scored failures.
Where it breaks at scale
NOT ON HISTORY LENGTH, AND THAT IS THE DESIGN. The carried accrual is four scalar fields, so cycle 40 costs exactly what cycle 3 costs -- unlike a conversation kit, whose input grows with every turn. What it breaks on is four other things. FIRST, THE CHAIN IS SERIAL. Cycles within one account cannot be parallelised, because cycle 3's prompt contains an accrual produced by cycle 2. This run went 50 chains wide with 12 workers and took 345.3 seconds of wall clock for 150 calls; an account with forty cycles of history would take forty serial calls and no width would help. SECOND, THE STATE STORE IS A FILE. src/state.py writes data/state.json atomically, which is correct for one process and is not a concurrency model. Two schedulers advancing the same account would race, and the loser's cycle would silently vanish from the accrual -- and because the carried quantity is a RUNNING TOTAL, a lost cycle is not a lost row, it is a permanently wrong number. THIRD, NOTHING HERE SCHEDULES ANYTHING. This is a monitor in the sense that it carries state between cycles; it is still invoked, not woken. FOURTH, A REAL UNBILLED BOOK IS NOT 50 ACCOUNTS. A large retailer's is six figures, and at one call per account per cycle the arithmetic that matters becomes the bill rather than the wall clock: see Cost.cost_at_10x.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
ACC-0007-C3, the whole argument on one screen. A suspension opened inside cycle 2 and closed on 2026-07-23 inside cycle 3, so this statement carries a CLOSED entry with no OPENED to pair it with and nothing on the page says how long the stop ran. The free one-page floor therefore counts every day since the oldest unbilled read and reports BREACHED at 396 days -- write the balance off. The carried accrual says 347 days were already counted and a suspension was open; the answer is AT_RISK at 360, five days of headroom against a fourteen-day lead time. The right-hand column is the floor, shown every time, so a reader can see which rows the model bought nothing on. ⚠︎ THIS FRAME IS A SEPARATE LIVE CALL FROM THE SCORED RUN, on a prompt verified byte-identical to it, and the two disagree by one day: the scored run answered 359 and this call answered the correct 360. It is not a staged frame -- it is the whole of this kit's evidence that its day-count error may be noise rather than a convention it has misread, and it is recorded in Eval.could_not_verify rather than being quietly re-shot.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page with NO API_KEY configured. It does not error and it does not go blank: the statement, the carried accrual, the parsed dates, the withheld-section list and the entire free one-page floor are all computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called. This is the honest failure state, not a staged one.failureOpen full size →Before anything is checked. The two things this page has to get right are already on it: the carried accrual, verbatim as it goes into the prompt, and the list of which sections left the machine and which did not -- Customer Contact, the household's name and supply address, marked WITHHELD rather than silently absent. A page that simply does not mention them cannot be told apart from one that quietly sent them.failureOpen full size →
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150cycle statements
0.53 MiBjson 1 · jsonl 1 · txt 150
50accounts (3 cycles each) · p50 3 chars
$0.00setup · 0.0s
How it is cutWhat one accounts (3 cycles each) is
No split, and no chunking. The unit is an ACCOUNT -- a run of three consecutive monthly cycle statements processed strictly in order, because cycle 3's prompt contains an accrual produced by cycle 2. Each statement goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step. tools/build_corpus.py writes 150 documents and the answer key from a fixed seed with no clock read and no model called; nothing is embedded, ranked or cached.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every account, tariff, meter, reading, exception entry, household name and address is invented here. Verified against the repository's own LICENSE file on 2026-08-23.
Bring your ownBring your own cycle statements
Replace src/backbill.py FIRST, and not as an option. It is your regulator's window and your operation's lead times, the shipped ones are invented, and data/gold.jsonl is literally its output -- change the rules and re-run tools/build_corpus.py and the answer key follows. Then point the corpus generator at your own accounts, or delete it and drop real cycle statements into data/corpus/ named <ACCOUNT>-C<n>.txt with the same seven section headings. src/segment.py, src/select.py, src/prompt.py and src/state.py read sections by name and do not care where the statements came from.
⚠︎ And what stops being true when you do: Every figure here is a property of a three-cycle corpus of 50 invented households, measured once. Three of the denominators are small: 11 AT_RISK cells, 25 memory-dependent cells and ONE false alarm -- one row is not a rate. And the headline hides its own shape: 99.33 pct band accuracy is 149 correct cells and ONE mis-mapped blocker, so a book with a different mix of recorded exception wording moves that figure without the model having changed at all. ⚠︎ The thing most likely not to survive contact with a real book is the blocker taxonomy, which is the field this run is WORST on and the one the catalogue row already records as unconfirmed.
What breaks it
⚠︎ A HOUSEHOLD DOES NOT USE 6,834 kWh OF ELECTRICITY IN ELEVEN MONTHS -- FOUND BY READING, FIXED, AND GATED. The first build drew ONE daily-consumption range (7-26 kWh) and ONE unit-rate range for BOTH fuels. Gas and electricity differ by roughly four times in volume and four times in price, in opposite directions, so a single range is wrong for both: it billed homes about two and a half times the electricity they use, at an electricity price, and produced unbilled estimates near 1,975 USD where 900 would be right. EVERY BAND, EVERY GOLD LABEL AND EVERY ASSERTION IN THE REPOSITORY WAS PERFECTLY HAPPY WITH IT -- the arithmetic is internally consistent whatever the numbers are, which is exactly why nothing caught it. It was found by opening a generated document and reading it. DAILY_KWH, UNIT_RATE and STANDING are now per fuel, NON_DOMESTIC_MULTIPLIER covers the micro-business accounts, and evals/check_labels.py carries a permanent per-fuel range assertion. What remains untested is the class, not the instance: a quantity whose plausible range this corpus does not know about would reproduce the same defect and the same silence.
⚠︎ THE RULE TEXT CONTRADICTED THE ANSWER KEY, AND THE KEY WOULD HAVE WON. Rule 3's first draft read 'the window is SUSPENDED for any period in which the delay is attributable to the customer'. On the 32 cycles whose standing block reason is 'no access to the meter' or 'the customer undertook to send a reading' with NO suspension entry raised against it, that sentence says the clock has been stopped the whole time -- while src/backbill.py counts every one of those days. A model reading the rule as written would have been marked wrong for reading it correctly. The rule now names the DATED ENTRY as the thing that suspends the window and says explicitly that a customer-sounding standing reason does not. Those 32 cycles are a trap for a reader, not an ambiguity in the key.
Three cycles per account. Long enough to hide a suspension behind a boundary; not long enough to test an accrual that has been running for years, a window that closes mid-history, or a suspension that opens and closes twice.
A hidden suspension on a LONG-LEAD blocker. The two spanning-suspension patterns draw their standing blocker from the short-lead codes on purpose, because a hidden stop of thirty-odd days only moves the BAND when thirty days is comparable with the lead time. On a 60-day NO_ACCESS block the same hidden stop moves the day count and usually leaves the band where it was. That is a corpus decision, stated in the generator, and it means the published band separation is a property of the mix.
A dispute that is later withdrawn. src/backbill.step zeroes the day count when an account leaves the protection, so an account that came back would restart its clock from nothing. Nothing here withdraws a dispute, so the defect is unmeasured rather than absent.
A real read history. Every account has exactly one oldest-unbilled-actual-read date and it never moves. A real book has partial reads, disputed reads, reads later withdrawn, and meter exchanges that reset the register.
One back-billing rule, reproduced identically on every page. A real supplier's rules differ by segment, carry exceptions and get amended; none of that is here.
Free-text agent prose. The Operational Notes section is this kit's injection surface and it is deliberately sent, but on this corpus it is one of three fixed sentences chosen by a seeded generator. A book whose notes are written freely by agents is the real surface and this corpus does not exercise it.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
134
not measured
instruction
2,357
not measured
carried state
186
not measured
Synthetic Record
256
not measured
Account
332
not measured
Back-Billing Rule
1,686
not measured
Billing Position
516
not measured
Exception Log
510
not measured
Operational Notes
205
not measured
Total
1,486
This is the cost lesson as arithmetic: of the 6,182 characters assembled, 2,491 are instructions — 40% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.py::build on ACC-0001-C1 with the exact carried accrual r001-unbilled-watch recorded for that call ({"days_at_risk": 0, "prev_band": null, "protected": true, "suspension_open": false}) -- the identical code path the run used, not a paraphrase. THE SPLIT IS IN CHARACTERS AND THE PAGE SAYS SO: per-part TOKEN counts were not measured, because measuring them means sending nested prefixes of the prompt to the provider and this kit spent no calls on it. The billed totals (222,929 in / 306,959 out over 150 calls) are measured; nothing here apportions them across the parts. The stateless control's build of the same statement differs only inside the carried-state block, and evals/check_labels.py asserts every other byte is identical -- that is the entire experiment.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You apply a written back-billing rule to one unbilled energy account for one cycle. You answer with one JSON object and no other text.
You are reviewing one energy account that is unbilled beyond its billing cycle, for one cycle, and
deciding how close it is to the back-billing limit written in the statement below.
The back-billing rule, its window length and the remediation lead time for every blocker are
reproduced in the statement. Apply them exactly as written. The rule depends on how many days of
this account's unbilled period have ALREADY been counted towards the window in EARLIER cycles, which
you cannot see; what is known is stated under "Carried state" and is the only history available to
you. Do not assume anything about earlier cycles beyond it.
How to count:
- Days counted in this cycle are stated in the Billing Position.
- Subtract from them every day inside a customer-attributable suspension the Exception Log records.
A suspension runs from the day it is OPENED up to but not including the day it is CLOSED. A
standing block whose reason sounds customer-attributable does NOT suspend the window unless a
suspension was opened for it. The Exception Log shows only the entries raised in the period this
statement covers; a suspension the carried state says was already open has been running since
before this cycle and stops on the day this statement records it closing.
- Add what is left to the days the carried state says were already counted.
- An account outside the protection has no window at all: answer band NO_LIMIT, clock_status
NOT_PROTECTED, and days_at_risk 0.
Answer with a single JSON object and nothing else:
{"band": "WITHIN_LIMIT|APPROACHING|AT_RISK|BREACHED|NO_LIMIT",
"clock_status": "RUNNING|SUSPENDED|NOT_PROTECTED",
"days_at_risk": <whole number of days now counted towards the window>,
"blocker_cause": "EXCEPTION_HOLD|STUCK_ORDER|METER_PROBLEM|NO_ACCESS|READ_DISPUTE|CAUSE_UNKNOWN",
"rationale": "one sentence, naming the rule you applied and the days you counted"}
"band" is the band AFTER the suspension rule has been applied, which is not the band the elapsed
time since the last actual read alone would imply. "clock_status" is SUSPENDED only if a
customer-attributable suspension is still open at this cycle end. "blocker_cause" is the standing
billing block mapped to one of the six codes above; use CAUSE_UNKNOWN when the recorded reason maps
to none of the other five rather than choosing the nearest.
Carried state
----------------------------------------------------------------
No earlier cycle has been recorded for this account. This is its first appearance on the watchlist, so the days already counted towards the window are only those shown on this statement.
Cycle statement
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes no real
customer, household, address, meter, supplier or charge. Account ACC-0001, cycle C1.
Account
----------------------------------------------------------------
Account : ACC-0001
Segment : domestic (inside the back-billing protection)
Product : Prepayment converted to credit, electricity
Meter type : smart, non-communicating
Supply since : 2020-05-14
Billing cycle : monthly, cycle end on day 20
Back-Billing Rule
----------------------------------------------------------------
Back-billing window : 365 days
Remediation lead time published for each blocker:
EXCEPTION_HOLD 14 days
STUCK_ORDER 7 days
METER_PROBLEM 45 days
NO_ACCESS 60 days
READ_DISPUTE 30 days
CAUSE_UNKNOWN 30 days
1. A charge for energy consumed more than 365 days before the date a valid bill is issued cannot
be recovered from the customer and must be written off.
2. The 365 days run from the date of the OLDEST UNBILLED ACTUAL METER READ. They do not run from
the last bill issued, and they do not run from the date the billing exception was raised.
3. The window is SUSPENDED only for a period the Exception Log records as a customer-attributable
suspension -- access to the meter refused or not arranged, or a reading the customer undertook
to provide and did not. A suspension runs from the day it is OPENED up to but not including the
day it is CLOSED, and days inside it do not count towards the window. A standing block whose
recorded reason sounds customer-attributable does NOT suspend the window unless a suspension
was opened for it.
4. The protection does not apply to an account outside the protected segment, nor to an account
where a valid bill was issued and the customer has disputed it. Those accounts have NO LIMIT
and are not on this watchlist.
5. An account is AT RISK when the days left before the window closes are fewer than the published
remediation lead time for its blocker; APPROACHING when they are fewer than twice that lead
time; BREACHED when the window has already closed.
Billing Position
----------------------------------------------------------------
Cycle : C1
Cycle end : 2026-06-20
Previous cycle end : none - this is the first cycle this account appears on the watchlist
Days counted in this cycle : 153
Oldest unbilled actual read : 2026-01-18
Last bill issued : 2026-04-27 (estimated - does not close the unbilled period)
Unbilled consumption estimate : 1,637 kWh
Unbilled value estimate : 579.20 USD
Exception Log
----------------------------------------------------------------
Standing billing block : BLK-71935, raised 2026-02-15, still open
reason as recorded : "meter room locked and no arrangement has been made to open it"
Entries raised in the period this statement covers. A suspension is a separate, dated episode and
does not change the standing block reason above. Entries from earlier statements are not reproduced
here.
(no exception entries were raised in the period this statement covers)
Operational Notes
----------------------------------------------------------------
This account was picked up in the unbilled sweep for this cycle. Nothing has been billed and
nothing has been written off.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"band":"WITHIN_LIMIT","clock_status":"RUNNING","days_at_risk":153,"blocker_cause":"NO_ACCESS","rationale":"Applying the 365-day back-billing window with no customer-attributable suspension recorded, the 153 days counted this cycle leave 212 days remaining, which is within the 60-day NO_ACCESS remediation lead time."}
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch unbilled power accounts before the charge expires — 150 cycle statements. One model answered, and every answer was then graded Five different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
There is no LLM-as-judge in this kit and that is not a shortcut. Three of the four answered fields are closed sets -- five bands, three clock states, six blocker codes -- and the fourth is a whole number of days. A judge is for answers whose correctness is a matter of reading; these are matters of equality and arithmetic, and asking a model to grade 'AT_RISK == AT_RISK' would add cost, variance and a second thing to be wrong. evals/scoring.py compares each cell exactly against a gold set that src/backbill.replay computed, and scores the stateless control and the free floor through the same function.
150cycle statements
150source documents
1model tier
5grading methods
MeasurementsWhat was measured
COUNTED149 · 103 / 150band accuracy pct — cycles, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED54 · 13 / 54band on watchlist pct — cycles whose correct band is not WITHIN_LIMIT or NO_LIMIT, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED147 · 65 / 150days at risk accuracy pct — cycles, exact whole-day match, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 · 138 / 150clock accuracy pct — cycles, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED137 · 133 / 150blocker cause accuracy pct — cycles, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED25 · 9 / 25memory band accuracy pct — cycles whose answer is NOT derivable from their own page, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED1 · 1 / 96false alarm rate pct — quiet cycles, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 1 / 1exposure error pct — portfolio total in USD reported as beyond the window, with the carried accrualDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 150 / 150unparsed replies — calls -- every reply parsed; largest 12,664 tokens under a 24,000 capDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The comparison cannot be wrong about itself; the risk is in the labels. Those come from src/backbill.replay, which the runtime never calls -- src/ asks the model for the band and never computes it -- so the answer key is arithmetic done independently of the code under test rather than the code under test marking its own homework. evals/check_labels.py re-derives all 150 gold rows from the planted inputs and refuses to let a run spend if any disagrees. ⚠︎ WHAT THE KEY CANNOT SETTLE is whether one of its six blocker codes is NAMED for what it describes; see could_not_verify and results/ambiguity-r001-unbilled-watch.json.
2,046.39output tokens · the fast tier, with the carried accrual · 12,343 ms p50
4,738.99output tokens · the same tier, memory removed (THE CONTROL) · 22,557 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 1.8× as long, and lands one row apart on 150. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One cycle statement
1,000 cycle statements
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.006882
$6.88
11%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.002753
$2.75
11%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.117181
$117.18
13%
Same work, 43× the bill
The same cycle statements, the same tokens — only the rate card changed. And across all 3 cards between 11% and 13% of what you pay is the prompt this pipeline sends, not the answer it writes.
provider-side reasoning -- 95.3 pct of this run's output tokens, and output is 89.2 pct of the projected bill. Nothing else is close, and this kit has measured neither what disabling it costs in accuracy nor what a lower ceiling would have saved.
Rates checked 2026-08-18. The provider that actually ran all 309 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. A measured $0.00, not an unpriced one -- and evals/ambiguity.py re-scores the whole run under the other reading of one blocker code for $0.00 as well, because it reuses the recorded replies.
The gradersFive ways to grade
READ THIS ROW BEFORE THE STATELESS CONTROL, BECAUSE IT IS THE STRONGER NO-MEMORY ARM AND ON TWO FIELDS IT BEATS THE MODEL. It is not a straw man: it already avoids the mistake rule 2 exists to warn about (it anchors on the oldest unbilled ACTUAL read, not the last bill), it subtracts every suspended day the page shows, and its keyword table tests 'reason not recorded' BEFORE the topical words, which is the most generous reading of a keyword matcher available. It gets the blocker code right on 98.0 pct of cells against the model's 91.33, and it raises no missed watchlist cycles at all. What it cannot do is the three shapes that need memory: a CLOSED entry with no OPENED, an empty exception log hiding a still-open suspension, and a dispute raised in an earlier cycle. ⚑ AND ITS EXPOSURE ERROR RUNS THE OPPOSITE WAY TO THE STATELESS MODEL'S. The floor over-states the balance beyond the window by 49.39 pct ($43,358.99 against a true $29,024.35); the stateless model UNDER-states it by 65.98 pct ($9,873.05). Two arms with the same information, erring in opposite directions -- which means 'no memory' is not a bias anybody could correct for with a constant.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key For each of the 150 cycles and each of the four answered fields, did the reply equal the computed answer key? Band, clock and blocker are compared exactly; the day count is compared as a whole number, and a reply that is not a whole number of days is a MISS rather than a zero -- 0 is a meaningful answer here (it is what an unprotected account gets) and a parse failure must not be scored as one.
$0.00
no
yes
the fast tier, with the carried accrual 99.3% band accuracy · the fast tier, memory removed (THE CONTROL) 68.7% band accuracy · the free one-page floor, no model 86.7% band accuracy · 3 more measured on each run
Band accuracy restricted to the cycles that need somebody to touch them Of the 54 cycles whose correct band is not WITHIN_LIMIT and not NO_LIMIT, how many were banded correctly? This is the number the kit is actually about; the 150-cycle average is 64 pct quiet and something that answers WITHIN_LIMIT to everything scores well on it.
$0.00
no
yes
the fast tier, with the carried accrual 100.0% accuracy · the fast tier, memory removed (THE CONTROL) 24.1% accuracy · the free one-page floor, no model 81.5% accuracy
False alarms on the cycles where nothing needs doing Of the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT, how many were put on the watchlist anyway? This is the other direction, and it costs somebody an hour where the first costs the balance.
$0.00
no
yes
the fast tier, with the carried accrual 1.0% false alarm rate · the fast tier, memory removed (THE CONTROL) 1.0% false alarm rate · the free one-page floor, no model 9.4% false alarm rate
The cycles whose answer cannot be derived from their own page On the 25 cycles where a suspension began, or a dispute was raised, in a statement the model is not holding, was the band right and was the day count right? This is the subset the whole monitor argument rests on, and it is marked in the answer key by the generator's own structure -- never by what any arm answered.
$0.00
no
yes
the fast tier, with the carried accrual 100.0% band accuracy · the fast tier, memory removed (THE CONTROL) 36.0% band accuracy · the free one-page floor, no model 20.0% band accuracy · 1 more measured on each run
The portfolio total reported as beyond the back-billing window Sum the unbilled value of every cycle the arm reported BREACHED, and compare it with the truth. This is the number that lands in a month-end pack: how much of the unbilled book is already unrecoverable. It is a DIFFERENT question from per-account accuracy and it is the question the catalogue row asks -- 'reconcile that against this month's unbilled-revenue estimate'.
$0.00
no
yes
the fast tier, with the carried accrual 0.0% error · the fast tier, memory removed (THE CONTROL) 66.0% error · the free one-page floor, no model 49.4% error
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Decisively in one direction, and the kit says which comparison is doing the work. THE STRONGER NO-MEMORY ARM IS THE FREE FLOOR, NOT THE STATELESS MODEL, and the headline is measured against it: band 99.33 pct against 86.67 (+12.66), watchlist 100.0 pct against 81.48 (+18.52), day count 98.0 pct against 83.33, and the memory-dependent subset 100.0 pct against 20.0. The stateless control separates further still (68.67 pct band, 24.07 pct watchlist) but ⚠︎ ITS DAY-COUNT FIGURE IS CONFOUNDED and the page says so: 61 of its cycle-2 and cycle-3 answers report only the current cycle's days, because the instruction tells the model to ADD to a carried figure and the control removes the figure without rewriting the instruction. That is partly a measurement of a prompt written for a stateful arm. WHAT THIS LABELLED SET CANNOT SEPARATE: the stateful arm from a perfect answer on the band (99.33 pct is one miss, and that miss is downstream of the taxonomy field, so the band grader has saturated); a monitor that understands the half-open interval convention from one that is consistently a day out in the safe direction; and the two readings of one blocker code, which it can only PRICE -- evals/ambiguity.py values the disagreement at 4.67 points of blocker accuracy and zero points of band accuracy.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Sorting an unbilled queue where nothing was ever suspended and every account's history is on its own page
the free one-page floor, at $0.00
The floor gets the band right on 86.67 pct of all cycles and the blocker code on 98.0 pct -- BETTER than the model's 91.33. On the clean-drift, suspension-inside-one-cycle and not-protected-by-segment patterns it is exactly right, because the elapsed calendar IS the risk clock there.
Paying a model to do date subtraction a regex already does perfectly.
An account whose suspension began in a cycle you are not holding -- a CLOSED entry with no OPENED, or an empty exception log
the fast tier WITH the carried accrual
100.0 pct of the bands and 92.0 pct of the day counts on the 25 memory-dependent cells, against the floor's 20.0 pct and 0 pct. This is the half the accrual plainly buys, and it is the difference between AT_RISK at 360 days and BREACHED at 396 on the same statement.
Nothing here. It is also cheaper than the same model without the accrual.
Producing the month-end unbilled-exposure figure for finance
the fast tier WITH the carried accrual, with the per-account list published beside the total
The total is exact ($29,024.35 of $29,024.35, 0.0 pct error) where the floor is 49.39 pct OVER and the stateless arm 65.98 pct UNDER. ⚠︎ But publish the list too: this grader is where errors cancel, and a total can be right while rows are wrong.
Quoting the total on its own. It is the one figure on this page that can be right for the wrong reason.
Deciding which exception code an account's recorded reason belongs to
neither, as built -- free code beats the model and both are worse than a curated list
The floor scores 98.0 pct and the model 91.33 pct (96.0 pct under the most generous reading of its own answers). The model invented ten CAUSE_UNKNOWNs on reasons that map cleanly, while getting 18 of 18 genuinely unrecorded ones right. And the catalogue row already records the real taxonomy as unconfirmed, so this field is measuring agreement with an invented list.
Reading 91.33 pct as the model being bad at classification. Two recorded phrasings account for all 13 misses, and one of the two is a defect in the code's NAME.
Deciding to write a balance off
a person, reading the band, the day count and the carried accrual the kit printed
The highest-consequence output here is BREACHED, and both no-memory arms get the total wrong in opposite directions -- the floor by 49.39 pct over, the stateless arm by 65.98 pct under. Nothing in this kit writes anything off and nothing should be added that does.
Wiring any of this to a write-off. There is no such code path and that is the guardrail.
An account with years of history rather than three cycles
nothing here yet -- measure it first
The accrual is four scalar fields, so the COST does not grow. The accuracy is unmeasured: every rule in this corpus resolves inside three cycles, and the one error shape this run has (the half-open boundary, one day out) is exactly the shape a longer history gives more chances to accumulate -- on a RUNNING TOTAL, one day out per cycle is forty days out after forty cycles.
Assuming any figure here transfers. The corpus mix decides the headline.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
blocker_read_omission_unmapped
A reading the customer never sent, not mapped to the read code
10
Ten cycles across four accounts carry the recorded reason "the customer undertook to send a reading and has not done so". Seven answered CAUSE_UNKNOWN and three answered NO_ACCESS; the key says READ_DISPUTE. ⚠︎ THIS ONE IS ARGUABLY THE KIT'S FAULT…
blocker_hold_unmapped
A plain exception hold answered as unrecorded
3
ACC-0029-C1, ACC-0029-C2 and ACC-0030-C1 carry "cycle run completed with this account held back for checking" and answered CAUSE_UNKNOWN where the key says EXCEPTION_HOLD. Unlike the row above there is no defensible second reading: the sentence says the cycle…
day_count_off_by_one
The half-open suspension interval, counted one day out
3
ACC-0007-C3 (359 against 360), ACC-0012-C2 (323 against 324) and ACC-0024-C1 (307 against 306). All three are spanning-suspension accounts and all three are the boundary convention -- a suspension runs from the day it is OPENED up to but NOT including the day…
lead_time_from_the_wrong_blocker
The right day count, banded against the wrong lead time
1
ACC-0034-C2 -- the only band miss in 150 cycles. Its own rationale: "...giving 304 days at risk and 61 days remaining, which is fewer than twice the 60-day NO_ACCESS lead time, so APPROACHING." Every step is right except which blocker it is: the key says…
no_verdict
No reply at all
0
None on the scored run. All 150 replies parsed, finish_reason was never 'length' and the largest reply reached 12,664 of the 24,000-token ceiling (53 pct). ⚠︎ The STATELESS control at the identical ceiling lost one -- ACC-0024-C2 spent the whole budget and…
correct
Every cell correct
134
134 of 150 cycles answered all four fields correctly, including every BREACHED cycle (18 of 18), every NOT_PROTECTED cycle (25 of 25), all 54 cycles on the watchlist and 25 of the 25 memory-dependent cells on the band.
What we could NOT verify
⚠︎ THE STATELESS CONTROL IS CONFOUNDED ON THE DAY COUNT AND THE PAGE PUBLISHES IT ANYWAY RATHER THAN QUIETLY DROPPING IT. 61 of its cycle-2 and cycle-3 replies answer only the days in the current cycle (30 or 31), not the days since the oldest unbilled read. That is a defensible reading of an instruction which says 'add what is left to the days the carried state says were already counted' when there is no carried state to add to. So 43.33 pct is a measurement of memory-removal AND of a prompt written for a stateful arm, mixed, and nothing here separates them. The clean experiment is a third prompt that tells a memory-less arm to anchor on the oldest read; it costs one more 150-call run and has not been fired. THE FREE FLOOR IS THE HONEST NO-MEMORY COMPARATOR FOR THIS FIELD (83.33 pct) and it is what the headline is measured against.
⚠︎ WHETHER ONE OF THE SIX BLOCKER CODES IS NAMED FOR WHAT IT DESCRIBES. READ_DISPUTE is named for a disagreement and the key applies it to "the customer undertook to send a reading and has not done so", which is an omission. evals/ambiguity.py publishes what every figure becomes under the other reading (blocker cause 91.33 pct -> 96.0 pct over the 12 cycles carrying that sentence; the model gains 7 and the band does not move, because both codes carry the same 30-day lead time). The wording was NOT edited after the answers came back and the LOWER figure is published everywhere. What is not settled is which reading a billing operation would actually want -- and the catalogue row already records the blocker taxonomy itself as unconfirmed.
⚠︎ THE BACK-BILLING WINDOW IS INVENTED. 365 days, six remediation lead times, a segment threshold and a dispute rule, none of them confirmed against any source. Every band on this page is agreement with that invented rule, not with anybody's regulator. The catalogue row records the real value as an open question shared with a sibling row and this kit ships BLOCKED-PENDING-ANCHOR on it.
One run per arm, each fired once. Nothing here is a distribution: whether 99.33 pct band accuracy repeats on a second identical run is unmeasured, and the small denominators (54 watchlist cycles, 25 memory-dependent cells, 11 AT_RISK cells, 1 false alarm) make any repeat noisy.
One tier. No second model was run, so nothing on this page compares two models -- the three arms differ by MEMORY and by whether a model was called at all, never by which model.
⚠︎ WHETHER THE THREE OFF-BY-ONE DAY COUNTS ARE AN ERROR AT ALL, OR SAMPLING NOISE. ⚠︎ AND ONE UNPLANNED REPEAT SAYS THIS IS NOT A STABLE MISREADING. The UI screenshot fires one live call on ACC-0007-C3, on a prompt verified byte-identical to the scored run's for that cycle. The scored run answered 359 days; that call answered 360, which is right. A single repeat is evidence and not a measurement -- but it is enough to withdraw 'the model is a day out on boundary suspensions' as a finding. What is measured is that 3 of 150 day counts were one day out ON THIS RUN; whether that is a convention error or sampling noise is UNRESOLVED, and resolving it costs one repeat run. Two experiments are outstanding and neither has been fired: a repeat run of the same 150 cycles, which would say whether the rate is stable, and a prompt that spells the half-open convention out with a worked example, which would say whether it can be taught. Until the first of those, this run's 98.0 pct exact figure is a measurement and its EXPLANATION is not.
Whether the exposure total's 0.00 pct error survives an arm that makes offsetting errors. This arm made none to offset. The figure has never been exercised in the regime it is most likely to mislead in.
Whether rendering the carried accrual as JSON rather than English matters. src/state.describe argues for English and says plainly that the argument is unmeasured; scoring the same 150 cycles with the accrual rendered both ways would cost one more full run and has not been paid for.
Whether disabling provider-side reasoning changes the answers. 95.3 pct of this run's output tokens were reasoning, left at the provider's default. Turning it off would change the price by an order of magnitude and this kit has not measured what it does to the bands.
Whether the 24,000-token ceiling is enough. This run's largest reply reached 12,664 (53 pct of it) with 0 failures -- but the stateless control at the identical ceiling lost one call outright, so the headroom is adequate for the published configuration and demonstrably not for a harder one.
Whether any quantity in this corpus other than consumption carries a plausibility bound the generator does not know about. The per-fuel range defect (Data.breaks_on) is fixed and gated for the quantity it was found on; the gate asserts the instance, and nothing asserts the class.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried accrual
1,486.19
2,046.39
12,343 ms
$0.006882
$0.002753
$0.117181
the same tier, memory removed (THE CONTROL)
1,461.13
4,738.99
22,557 ms
$0.014948
$0.005979
$0.251561
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same cycle, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
309 live calls are on the shared ledger for this kit and every one of them is behind a figure on this page: 9 calibration (c000, at a 32,000-token ceiling on the three hardest accounts), 150 scored, 150 the stateless control. NOTHING WAS DISCARDED and no run was re-fired -- the corpus defects this kit found (see Data.breaks_on) were both caught by reading generated documents BEFORE any call was made, which is the whole argument for reading your own corpus first. Everything else -- the free one-page floor, the wiring stub, the scorer, the pre-flight and the ambiguity re-score -- is pure code and costs $0.00. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING, 95.3 pct of output tokens (292,625 of 306,959) and therefore most of the bill. It was left at the provider's default and this kit has not measured what turning it off does to the answers.
HOW HARD THE INTERVAL ARITHMETIC IS, NOT HOW LONG THE PAGE IS. Every statement is between 3,641 and 3,869 bytes -- within 6 pct of each other -- and output tokens ran from 282 to 12,664, a 44.9x spread on a corpus of one shape. The long replies are the boundary-suspension cycles.
WHETHER THE PROMPT CARRIES THE ACCRUAL -- AND IT IS A SAVING, NOT A COST. One short sentence adds a handful of input tokens a cycle and CUTS the output by 2.32x, because a model with the accrual has an addition to do and a model without one has a puzzle it cannot solve.
THE FIXED HEAD OF THE PROMPT. The system message and the instruction are 2,491 of 6,356 characters -- 39 pct, byte-identical on all 150 calls. Add the back-billing rule block, which varies in nothing but nothing, and the account's own numbers are the smaller half of what is sent.
Your volumeWhat it costs at your volume
Linear in CYCLES, and flat in history length, which is the unusual half. Each cycle is one call whose input is one statement plus four carried fields, so an account on its fortieth cycle costs exactly what one on its third costs. 150 cycles project to $1.0323 on the shared rate card, so ten times the set is about $10.32. ⚠︎ AND A REAL UNBILLED BOOK IS THE PLACE THIS ARITHMETIC STOPS BEING COMFORTABLE: at $0.006882 a cycle, a hundred thousand unbilled accounts checked monthly is about $688 a month, every month, and the free one-page floor already gets 86.67 pct of the bands right for nothing. The honest reading is that this belongs on the subset the floor cannot answer, not on the whole book. What does NOT scale is wall clock: cycles within one account are strictly serial.
Where pricing changes shape
THE OUTPUT CEILING IS A CLIFF AND THIS KIT HAS BEEN OVER IT -- ON THE CONTROL, NOT ON THE SCORED RUN. At the published 24,000 the scored run lost none of 150, and its largest reply reached 12,664 tokens (53 pct of the cap). The STATELESS control, at the identical ceiling, lost ONE: ACC-0024-C2 spent the whole 24,000-token budget and returned nothing. A ceiling is not billed (the provider charges tokens produced, not tokens allowed), so the calls that finish in a few hundred tokens pay nothing for the headroom -- but a call that hits it is billed in full and returns nothing at all.
THE MEMORY STEP IS A DISCOUNT, NOT A SURCHARGE. Adding the carried accrual moves the price per cycle from $0.0150479 to $0.0068823 -- a 2.19x SAVING for one sentence, effectively all of it in output tokens.
Your return, with your numbers
Volumeaccount-cycles per billing run -- this run judged 150 (50 accounts x 3 cycles) in one pass per arm
What it replacessomebody opening each unbilled account, finding the oldest read that was actually billed, reconstructing how much of the delay was the customer's own from earlier cycles' exception entries, and totalling what is already lost
Time saved per itemnot measured here -- depends on how long reconstructing a suspension history takes in the reader's own billing system
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start a question whose honest answer might be 'free code already does most of this', which on the blocker-code field it does, better than the model.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
222,929input tokens · this run
306,959output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 150 cycles, one completion call each, one tier. The 150-call stateless control and the 9 calibration calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.413
$0.413
$2.75
2026-09-12
gemini-3-flash
Google
$1.032
$1.032
$6.88
2026-09-18
gemini-3-8-flash
Google
$1.318
$1.318
$8.79
2026-09-18
llama-5
Meta
$1.583
$1.583
$10.55
2026-09-18
claude-haiku-4-5
Anthropic
$1.758
$1.758
$11.72
2026-09-12
grok-4-5
xAI
$2.288
$2.288
$15.25
2026-09-18
grok-4-6
xAI
$2.288
$2.288
$15.25
2026-09-18
claude-sonnet-5
Anthropic
$3.515
$3.515
$23.44
2026-09-12
gemini-3-1-pro
Google
$4.129
$4.129
$27.53
2026-09-18
gpt-5-6-terra
OpenAI
$4.129
$4.129
$27.53
2026-09-12
gpt-5-6-sol
OpenAI
$7.031
$7.031
$46.87
2026-09-12
claude-opus-4-8
Anthropic
$8.789
$8.789
$58.59
2026-09-12
claude-opus-5
Anthropic
$8.789
$8.789
$58.59
2026-09-12
claude-fable-5
Anthropic
$17.577
$17.577
$117.18
2026-09-18
claude-fable-5-1
Anthropic
$17.577
$17.577
$117.18
2026-09-18
gpt-6-astra
OpenAI
$17.577
$17.577
$117.18
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 95.3 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (292,625 of 306,959), left at the provider's default, so every row below prices a reasoning-on workload. Output is 89.2 pct of the projected bill on the shared card, which means most of what these rows charge for is the model doing interval arithmetic. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
⚑ EVERY ROW PRICES THE ARM WITH THE CARRIED ACCRUAL, AND IT IS THE CHEAPER ONE. The stateless control costs $0.0150479 a cycle on the same card -- 2.19x the stateful price -- because removing the memory more than doubled the output tokens. A cheaper row is not a cheaper kit, and the cheapest option on this page is the free one-page floor at $0.00, whose scores are published beside every model row in Eval.baseline -- including the blocker-code field, where it BEATS the model.
Output length here is driven by how hard the INTERVAL ARITHMETIC is, not by document size -- 282 to 12,664 tokens on statements that are all within 6 pct of each other. A book with a different mix of suspension shapes would move every row below, and a book with histories longer than three cycles has not been priced at all.
⚠︎ AT SCALE THESE ROWS ARGUE AGAINST THEMSELVES. At $0.006882 a cycle on the shared card, a hundred thousand unbilled accounts checked monthly is about $688 a month -- and the free floor already gets 86.67 pct of the bands right for nothing. The honest deployment is the model on the subset the floor cannot answer, which on this corpus is 25 cycles of 150.
Nothing here includes retries, and this run had none that reached the harness. It also includes no tokens billed for an answer that never arrived: all 150 replies parsed. The stateless control's one lost call WAS billed in full and returned nothing, and it is not in these rows.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
14 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 50 accounts x 3 consecutive monthly cycles = 150 statements from a fixed seed (SEED = 20260823), across two fuels and both segments. It plants eight patterns, of which 20 accounts carry a suspension and 9 of those span a cycle boundary -- the trap. Consumption, unit rates and standing charges are drawn PER FUEL; see Data.breaks_on for the defect that rule was written from. The gold labels are src/backbill.replay's output, never typed.
You change it to: Point it at your own accounts, tariffs and exception vocabulary, or delete it and drop real cycle statements into data/corpus/ named <ACCOUNT>-C<n>.txt with the same seven section headings. Everything downstream reads statements by account and cycle and does not care where they came from.
tools/build_corpus.py
# Generate the synthetic unbilled-account corpus and its computed answer key.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
RULE = "-" * 64
CYCLE_MONTHS = [(2026, 6), (2026, 7), (2026, 8)]
PATTERNS = [("CD", 8), ("SE", 9), ("SO", 6), ("SL", 5),
PRODUCTS = [
src/backbill.pythe back-billing arithmetic — a swap seam
The window rule as pure code: days accrued, minus dated customer-attributable suspensions, against the remediation lead time for the standing blocker. No model, no judgement. ⚠︎ The window length and every lead time in it are INVENTED and the kit ships that as a stated blocker -- see Data.why_this_corpus.
You change it to: BACKBILL_WINDOW_DAYS, LEAD_DAYS and step() are your regulator's window and your operation's lead times. Change them there and re-run tools/build_corpus.py, which recomputes the whole answer key from the same function -- the gold cannot drift from the rules because it is their output. ⚠︎ THIS IS THE FIRST FILE TO REPLACE, not an optional one: the shipped window is invented.
src/backbill.py
# The back-billing window, and the arithmetic it implies. Pure code, no model.
BACKBILL_WINDOW_DAYS = 365
WITHIN_LIMIT = "WITHIN_LIMIT"
APPROACHING = "APPROACHING"
AT_RISK = "AT_RISK"
BREACHED = "BREACHED"
NO_LIMIT = "NO_LIMIT"
BANDS = (WITHIN_LIMIT, APPROACHING, AT_RISK, BREACHED, NO_LIMIT)
RUNNING = "RUNNING"
SUSPENDED = "SUSPENDED"
src/state.pythe carried accrual — a swap seam
SEAM 2 -- the thing that makes this a monitor. Four fields per account (days already counted, whether a suspension is open, the band last reported, whether it is still protected), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on cycle 40 as on cycle 3.
You change it to: load()/save() are a JSON file today. Point them at a table, a key-value store or a billing data mart and nothing else in the kit changes -- for_account() and advance() are the whole interface, and describe() is the only thing the prompt sees.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_account(store, account_id):
def advance(store, account_id, elapsed, suspended, susp_open_end, cause, protected):
def describe(state):
src/segment.pythe section splitter
Splits a cycle statement into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 documents before a run may spend.
src/segment.py
# Split a cycle statement into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Account", "Back-Billing Rule", "Billing Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector — a swap seam
Decides which sections reach the model. Customer Contact -- the household's name, supply address, phone and email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
You change it to: SECTION_HINTS maps a field to the sections it needs; NEVER_SENT names what is withheld whatever happens. ⚠︎ On a domestic book this is the load-bearing seam: check NEVER_SENT covers every section your statements carry that names a person, before anything spends.
src/select.py
# Pick which sections of a cycle statement are sent. Pure code -- the last deterministic step
BANNER = "Synthetic Record"
ACCOUNT = "Account"
RULEBLOCK = "Back-Billing Rule"
POSITION = "Billing Position"
EXCEPTIONS = "Exception Log"
CONTACT = "Customer Contact"
NOTES = "Operational Notes"
NEVER_SENT = (CONTACT,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-accrual sentence, and the selected sections in document order. The stateless build replaces the accrual sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost document.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/monitor.pythe watch
One cycle, one call. Parses the cycle dates and the unbilled value off the page with a regex (the model is never asked to read a date), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
src/monitor.py
# One cycle, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You apply a written back-billing rule to one unbilled energy account for one cycle. "
MAX_TOKENS = 24000
FIELDS = ("band", "clock_status", "days_at_risk", "blocker_cause")
def documents():
def accounts():
def load_doc(doc_id):
def position_of(text):
src/app.pythe local UI
One account, one cycle, its carried state and its proximity, on 127.0.0.1:9004. Renders with no key. It shows the carried sentence verbatim and the free one-page floor's answer beside the model's, so a reader can see on which rows the model bought nothing.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9004"))
def _gold():
GOLD_ROWS = _gold()
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe free one-page floor
Everything a reader of ONE statement can compute: the oldest unbilled actual read, the suspended days this page shows, an ordered keyword map of the standing reason, and the segment line. 0 calls, $0.00, scored through the identical scorer.
evals/baseline.py
# The free, no-model floor: the whole answer, computed from ONE statement, with no memory.
KEYWORDS = [
def _date(s):
def classify(reason):
def read_page(text):
def review(text):
evals/scoring.pythe scorer
Exact match per cell against the computed gold, split five ways an average would hide: the four fields, the cycles on the watchlist, the quiet cycles, the memory-dependent subset, and the portfolio exposure in dollars. No judge model.
evals/scoring.py
# Score a run against the computed gold. Pure code, no model, no judge.
FIELDS = ("band", "clock_status", "days_at_risk", "blocker_cause")
QUIET_BANDS = ("WITHIN_LIMIT", "NO_LIMIT")
def _pct(n, d):
def _norm(v):
def _as_int(v):
def score(records, golds):
def compare(stateful, stateless):
evals/check_labels.pythe pre-flight
Everything that must be true before a run may spend: 150 documents, 50 gap-free accounts, every document declaring itself synthetic on line 1, gold reproducible from backbill.replay, no statement naming another cycle, the two prompt builds differing in exactly one block, the privacy guard red-proven in both directions, per-fuel consumption inside a household's range, and no code path anywhere that releases a bill or writes a charge off.
evals/check_labels.py
# Everything that must be true before a run may spend. Free, no key, no model.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAILS = []
def check(label, ok, detail=""):
def _date(s):
def main():
evals/ambiguity.pythe ambiguity re-score
Re-scores the recorded run against the OTHER defensible reading of one blocker code. Free, no new calls. It exists because ten of the thirteen blocker misses are one recorded sentence whose code name does not describe it, and the honest response is to publish both numbers rather than reword the taxonomy after seeing the answers.
evals/ambiguity.py
# Re-score the recorded run under the OTHER defensible reading of one blocker code. Free, no calls.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
DISPUTED_REASON = "the customer undertook to send a reading and has not done so"
ALSO_ACCEPTED = B.CAUSE_UNKNOWN
def reason_of(doc_id):
def main(run_id="r001-unbilled-watch"):
evals/run.pythe run harness
50 account chains, three strictly-ordered cycles each, 12 concurrent workers. The accrual advances even for a cycle whose CALL failed, so one transport error cannot turn into three scored failures.
evals/run.py
# Run the watch over the 50 accounts and write one result file. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def load_gold():
def stub_complete(cfg, system, user, max_tokens=1024):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 50 accounts x 3 consecutive monthly cycles = 150 statements from a fixed seed (SEED = 20260823), across two fuels and both segments. It plants eight patterns, of which 20 accounts carry a suspension and 9 of those span a cycle boundary -- the trap. Consumption, unit rates and standing charges are drawn PER FUEL; see Data.breaks_on for the defect that rule was written from. The gold labels are src/backbill.replay's output, never typed. A swap seam.
src/backbill.pyThe window rule as pure code: days accrued, minus dated customer-attributable suspensions, against the remediation lead time for the standing blocker. No model, no judgement. ⚠︎ The window length and every lead time in it are INVENTED and the kit ships that as a stated blocker -- see Data.why_this_corpus. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Four fields per account (days already counted, whether a suspension is open, the band last reported, whether it is still protected), written from the arithmetic and never from the model's reply, and rendered into one English sentence for the prompt. It costs the same on cycle 40 as on cycle 3. A swap seam.
src/segment.pySplits a cycle statement into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 documents before a run may spend.
src/select.pyDecides which sections reach the model. Customer Contact -- the household's name, supply address, phone and email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing. A swap seam.
src/prompt.pyThree parts: the fixed instruction, the carried-accrual sentence, and the selected sections in document order. The stateless build replaces the accrual sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost document. A swap seam.
src/monitor.pyOne cycle, one call. Parses the cycle dates and the unbilled value off the page with a regex (the model is never asked to read a date), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling.
evals/baseline.pyEverything a reader of ONE statement can compute: the oldest unbilled actual read, the suspended days this page shows, an ordered keyword map of the standing reason, and the segment line. 0 calls, $0.00, scored through the identical scorer.
evals/scoring.pyExact match per cell against the computed gold, split five ways an average would hide: the four fields, the cycles on the watchlist, the quiet cycles, the memory-dependent subset, and the portfolio exposure in dollars. No judge model.
evals/check_labels.pyEverything that must be true before a run may spend: 150 documents, 50 gap-free accounts, every document declaring itself synthetic on line 1, gold reproducible from backbill.replay, no statement naming another cycle, the two prompt builds differing in exactly one block, the privacy guard red-proven in both directions, per-fuel consumption inside a household's range, and no code path anywhere that releases a bill or writes a charge off.
evals/ambiguity.pyRe-scores the recorded run against the OTHER defensible reading of one blocker code. Free, no new calls. It exists because ten of the thirteen blocker misses are one recorded sentence whose code name does not describe it, and the honest response is to publish both numbers rather than reword the taxonomy after seeing the answers.
evals/run.py50 account chains, three strictly-ordered cycles each, 12 concurrent workers. The accrual advances even for a cycle whose CALL failed, so one transport error cannot turn into three scored failures.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1486 input and 2046 output tokens per cycle, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Cycles/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per cycle directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Operational Notes, which are one of three fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- prose a billing agent types into an exception queue -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
The experimentWe did not attack it -- and the boundary that matters most on this vertical was proven by breaking it on purpose, twice
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus has none. So the five rows above are boundaries, not payloads. One of them was red-proven: the privacy guard was replaced with the naive fallback every sibling kit once shipped and all 150 statements leaked the household's name, supply address, phone number and email, then it was restored and all 150 held. ⚠︎ It took two attempts to do that honestly. The first red-proof swapped the guard and measured ZERO leaks -- which looked like the guard working and was the test never firing, because the fallback is unreachable on today's corpus. A guard that has never been MADE to fail is a guard nobody has tested, and a red-proof that cannot fail is worse, because it reports a green. Confirmed by assertion and by reading the recorded run, not by an attack trial, on 2026-08-23 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a household's name and address ever leave the machine?
Every cycle statement carries a Customer Contact section -- the account holder, their supply address, a contact number and an email. It is the only place in this corpus whose subject is a PERSON, and on a domestic energy book it is exactly what a supplier would most mind sending anywhere. A selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the statement MINUS that section rather than to the statement. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py over all 150 documents. ⚠︎ AND THE FIRST VERSION OF THAT PROOF MEASURED ZERO AND WAS THE TEST NOT FIRING. Swapping the guard for the naive or list(secs) changes nothing on today's corpus, because every hint names a section all 150 documents carry -- so the fallback is not on any live code path. The hole is CONDITIONAL, so the proof now reproduces the condition: a statement with no Exception Log, which is ordinary. With the guard, 0 of 150 leak; with or list(secs), 150 of 150.
Can a wrong answer poison the next cycle?
A monitor that fed its own verdict back into its next prompt would compound one bad cycle into every cycle after it -- and here the carried quantity is a RUNNING TOTAL, so one over-count in March would still be in the number in December.
It cannot. src/state.advance() calls src/backbill.step() on the dates parsed from the page; the model's reply is scored and never written. ⚠︎ Read that honestly: the property holds by construction, and this run cannot demonstrate it. The three day-count misses are one day out each and every account they sit on was answered correctly at its next cycle anyway, because the accrual the next cycle got came from the arithmetic and not from the reply.
Can a cycle statement tell the model what happened in an earlier cycle?
If any statement named an earlier cycle's band, accrual or suspension, the stateless control could read the history off the page and the gap this kit publishes would be measuring nothing.
evals/check_labels.py scans all 150 documents for any cycle token other than their own and refuses to let a run spend if it finds one. 0 leaks. It also asserts that the stateful and stateless prompt builds are byte-identical outside the carried-state block, and that none of the 25 memory-dependent cells is answerable from its own page.
Can this kit release a bill or write a charge off?
A watchlist that grew an apply path would be a system that decides, unsupervised, that a household's charge is unrecoverable -- and the guardrail on the catalogue row is 'no forced release, no write-off, non-configurable'.
There is no endpoint, no function, no flag and no button anywhere in src/, evals/, tools/ or ui/ that issues, releases, adjusts, suppresses or writes off a charge. The only things this kit writes are results/*.json and data/state.json. evals/check_labels.py asserts it mechanically, by looking for what would have to be present -- crude and grep-shaped, and the only mechanical form a guarantee about ABSENCE can take.
Can a figure measured under a non-published output ceiling be mistaken for a scored one?
A calibration fired at a higher cap produces better-looking numbers under a configuration the page does not name.
evals/run.py refuses --max-tokens unless the run id begins with 'c'. The calibration is c000, fired at 32,000; the scored run and its control carry the published MAX_TOKENS = 24,000 from src/monitor.py.
Each boundary above was checked by running an assertion or by reading the recorded run, not by an attack trial -- there is no untrusted field on this corpus to construct a payload against. One of the five is red-proven, meaning the guard was removed and the failure observed rather than merely asserted while passing, and its first red-proof was itself wrong and is recorded as such.
The result0 attack trials, five boundaries checked -- and the privacy boundary red-proven by removing the guard and watching all 150 statements leak a household's name and supply address.
0untrusted input fields on this corpus
0 of 0attack trials run
1 of 5boundaries red-proven, not just asserted
0 of 150documents leak a household's name or address
The Operational Notes ARE the field an outside party would influence in a real deployment, and this kit sends them rather than hiding them -- but on this corpus they are one of three sentences chosen by a seeded generator, so there is nothing adversarial in them to catch. A version pointed at real agent prose reopens the question and should be attacked before it ships. What IS measured is the privacy boundary, in both directions, before every run.
Read this twice
The accrual is written from the arithmetic, never from the model's reply, and on this kit that matters more than usual because the carried quantity is a RUNNING TOTAL. A monitor that fed its own day count forward would not lose one cycle to a bad answer, it would carry the error for the life of the account -- one day out in March is still one day out in December, and forty cycles of it is forty days. src/backbill.step() never reads the reply. ⚠︎ AND THIS RUN CANNOT PROVE IT, WHICH THE PAGE SAYS RATHER THAN LETTING THE NUMBER IMPLY OTHERWISE: the guarantee is in the code path, and the run is consistent with it. What the boundary gives you is narrower than correctness -- it guarantees only that a wrong answer stays where it is.
HonestyWhat this does not prove
Whether a real deployment's Operational Notes -- prose a billing agent writes freely -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at a real book.
Whether the state file is safe under concurrency. src/state.save() replaces atomically, which is correct for one writer and is not a concurrency model; two schedulers advancing the same account have never been run and would race. On a RUNNING TOTAL a lost cycle is a permanently wrong number rather than a missing row.
Whether a truncated reply can ever be partially trusted. The stateless control lost one call to the ceiling and src/monitor._parse rejected it outright, which is the safe behaviour and is not the same as having measured how often a fragment would have been right.
Whether the grep-shaped no-write-off assertion would catch a write-off path written under a name it does not know. It asserts the absence of four function names and two method calls; a path called something else would pass it.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No forced release and no write-off, non-configurable. This kit produces a proximity band, a clock status, a day count and a blocker code for a person to read. It never issues, releases, adjusts, suppresses or writes off a charge, and there is no setting that makes it. Separately: the carried accrual is written by code from the parsed dates and the previous accrual; the model's four answers are scored and are never written back into the history the next cycle is judged against.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). src/state.advance() -> src/backbill.step() is the only thing that touches the accrual, and it is called by evals/run.py after every cycle including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit releases a bill or writes anything off
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for release/write-off function names and method calls and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
A wrong answer does not propagate into the next cycle's accrual
17 wrong cells on r001-unbilled-watch across 3 fields, and every account carrying one was answered correctly at its next cycle. ⚠︎ The property is guaranteed by the code path (backbill.step() never reads the reply); this run is CONSISTENT with it rather than a demonstration of it.
A cycle whose CALL failed still advances the accrual
evals/run.py's exception branch calls backbill.step() before continuing, so one transport error cannot turn into three scored failures. The scored run had 0 failures so the branch was not exercised; the STATELESS control had 1 and did exercise it -- ACC-0024-C2's cycle 3 was still judged against a correct accrual.
The household's name and supply address never reach the provider
0 of 150 statements leak the Customer Contact section, and 150 of 150 leak it when the guard is removed AND the condition that reaches the fallback is reproduced. Both directions asserted by evals/check_labels.py before any run may spend.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The accrual is correct whatever the model says, which means a wrong band is reported wrongly for that cycle and nothing catches it -- see Eval.taxonomy's single band miss, which this guardrail does nothing to prevent.
IT IS NOT AN APPROVAL GATE, BECAUSE THERE IS NOTHING TO APPROVE. There is no apply path to gate: nothing here releases or writes off anything, so there is no 'undo' because there is nothing that could 'do'. A deployment that adds one adds the gate at the same time or it has no guardrail at all.
⚠︎ IT IS NOT A GUARANTEE ABOUT THE WINDOW ITSELF. The back-billing window and every lead time in src/backbill.py are invented. A watchlist that is arithmetically perfect against an invented rule is arithmetically perfect against nothing.
It does not make the state file safe to share. One writer, one file, replaced atomically -- two schedulers on one account would race, and on a running total the loser's cycle is a permanently wrong number rather than a missing row.
It does not validate the blocker taxonomy. Six codes, invented, and the run's worst field. One of the six is arguably named for something other than what it covers -- see Eval.could_not_verify.
WatchedWhat is watched, and why that one
3runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 33 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
11 measured by the latest run22 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
alarm
band_accuracy_pct; days_at_risk_accuracy_pct; blocker_cause_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A cycle that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. The scored run had 0 of 150 -- but the stateless control at the identical ceiling had 1, so 0 is this arm's result and not a property of the design.
band-on-the-watchlist
Band accuracy restricted to the cycles that need somebody to touch them
alarm
band_on_watchlist_pct; missed_watchlist; watchlist_cells — alarm on any cycle reported quiet whose correct band is not. This arm has none; one would be a new failure mode rather than a worse rate, because the cost is the whole unbilled balance rather than an hour of somebody's time.
false-alarms-on-quiet-cycles
False alarms on the cycles where nothing needs doing
alarm
false_alarm_rate_pct; false_alarms; unparsed_replies — alarm on any false alarm whose day count is CORRECT. This run's one is exactly that shape -- the arithmetic was right and the lead time was wrong -- which points at the blocker taxonomy rather than at the model's reasoning, and those need different fixes.
memory-dependent-subset
The cycles whose answer cannot be derived from their own page
alarm
memory_band_accuracy_pct; memory_days_accuracy_pct; memory_cells — alarm on any drop at all on the band. Every arm's number here is either 100 or structurally low, so a first movement on the stateful arm is a new failure rather than a worse rate.
exposure-reconciliation
The portfolio total reported as beyond the back-billing window
alarm
exposure_reported_usd; exposure_signed_error_usd; exposure_error_pct — alarm on a small error_pct beside a falling band figure. That combination is the signature of cancellation and is the one reading of this grader that is worse than useless.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
559,640
cycle statements edited — the count held, the bytes did not
split.count
50
the accounts count moved — a different set was scored
split.size_p50
3
the median size of one account moved
split.size_p95
3
the 95th-percentile size of one account moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cycles_scored 150, watchlist_cells 54) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
proximity band accuracy
not yet known
150 cycles
A band is the spread between runs of the SAME model at the same settings, and there has been no repeat. r001 and s001 are one run each of two DIFFERENT prompts, which is a comparison and not a band.
band on the watchlist
0 -- saturated at 100.0 pct, and that is a statement about the corpus
54 cycles on the watchlist
54 of 54, with 0 reported quiet. A grader an arm aces has stopped discriminating; the honest reading is that this corpus's watchlist cycles are within reach of the configuration, not that the configuration cannot be beaten.
blocker code accuracy
91.33 to 96.0 pct -- the two readings of one code name, not a spread between runs
150 cycles
evals/ambiguity.py, re-scoring the recorded replies for $0.00. The lower figure is published. This is the only field where the model loses to free code (98.0 pct).
day count, exact
98.0 pct -- 147 of 150, all three misses one day, and one of the three did NOT reproduce
150 cycles
evals/scoring.py::score, r001-unbilled-watch. Mean absolute error 0.02 days, worst single error 1 day, against the free floor's 17.43 and 424. This grader DOES discriminate: 98.0 pct against 83.33.
watchlist cycles reported quiet -- the expensive direction
0 of 54 on this run, and it is the number to watch rather than the accuracy
54 cycles on the watchlist
evals/scoring.py::score, r001-unbilled-watch. This is the half of the band figure that costs money: a quiet account put on the list costs somebody an hour, an account near the limit reported quiet costs the whole unbilled balance, and an average over both prices them the same. The stateless control reported 39 of 54 watchlist cycles as quiet on the identical corpus, and the free one-page floor reported 0 -- it errs the other way, by counting days it cannot see suspended.
clock status
0 spread -- saturated at 100.0 pct, and saying so is the point of banding it
150 cycles
evals/scoring.py::score, r001-unbilled-watch: 150 of 150. ⚠︎ A GRADER AN ARM ACES HAS STOPPED DISCRIMINATING, and this one nearly does for the other arms too -- the free one-page floor scores 92.67 pct and the stateless control 92.0 pct, so the whole spread across three arms is under eight points. RUNNING is the correct answer on 104 of the 150 cycles, so an arm that answered RUNNING to everything would score well here; what the field actually catches is the 21 SUSPENDED cycles and the 25 NOT_PROTECTED ones, and those are where every arm's misses are.
false alarms
not yet known
96 quiet cycles
One row of 96. 1.04 pct is what a single positive works out to, not a rate -- there is nothing here to estimate a spread from.
the memory-dependent subset
band 0 spread (saturated at 100.0 pct); day count 92.0 pct
25 memory-dependent cycles
25 of 25 bands and 23 of 25 day counts on r001-unbilled-watch, against the free floor's 20.0 pct and 0 pct. ⚠︎ The corpus draws these patterns from SHORT-lead blockers on purpose, so the subset is easier to separate than a random one would be.
the portfolio exposure total
0.00 pct on one run, and it is the figure most likely to be right for the wrong reason
one number per run, over 150 cycles
$29,024.35 reported against $29,024.35. The arm made no offsetting errors, so cancellation was never exercised; the floor is 49.39 pct OVER and the stateless control 65.98 pct UNDER.
the denominators themselves
0 -- constants of the corpus, not results
150 cycles
54 on the watchlist, 96 quiet, 25 memory-dependent, 18 BREACHED, 25 NOT_PROTECTED on every arm. Which cycles land where is decided by tools/build_corpus.py before any model sees anything.
input volume
1.7 pct between the two arms, and it is a measurement rather than a spread
150 calls
222,929 input tokens with the carried accrual, 219,169 without -- the difference is the accrual sentence. Prompt assembly is pure code, so this figure is model-independent.
output volume
not yet known
150 calls
⚑ 306,959 tokens WITH the carried accrual against 710,849 WITHOUT -- the memory-less arm produced 2.32x MORE output, which is the opposite of what this series has measured before. Two different prompts, not two runs of one, so it is a comparison and not a band.
provider-side reasoning share
not yet known
150 calls
95.3 pct of output on the scored run (292,625 of 306,959) and 98.1 pct on the control (697,178 of 710,849). A distribution within one run, which is not the same thing as a band between runs.
latency
not yet known
150 calls
p50 12,343ms and p95 55,907ms with the carried accrual; 22,557ms and 125,161ms without. One recorded run per configuration, and the tail is 4.5 times the median.
replies that did not parse
0 of 150 with the accrual, 1 of 150 without
150 calls
Two points at two configurations at the SAME ceiling, which is a reason to keep the ceiling high rather than a band. The lost call spent the whole 24,000-token budget.
HistoryRun history
3 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
b000-unbilled-watch-onepage 2026-08-23
r001-unbilled-watch 2026-08-23
s001-unbilled-watch-stateless 2026-08-23
band accuracy, %
86.67
99.33
68.67
band on watchlist, %
81.48
100.00
24.07
blocker cause accuracy, %
98.00
91.33
88.67
clock accuracy, %
92.67
100.00
92.00
days at risk accuracy, %
83.33
98.00
43.33
false alarm rate, %
9.38
1.04
1.04
input tokens, whole run
0
222929
219169
model latency p50 ms
0.00
12343.00
22557.00
model latency p95 ms
0.00
55907.00
125161.00
missed watchlist, %
0.00
0.00
72.22
output tokens, whole run
0
306959
710849
not a time series No two of these 3 runs measured the same system — they differ on documents, max_tokens, provider, stateless, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 3 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the carried accrual
band 68.67 pct -> 99.33 pct, watchlist 24.07 pct -> 100.0 pct, day count 43.33 pct -> 98.0 pct, memory-dependent band 36.0 pct -> 100.0 pct, exposure error 65.98 pct under -> 0.0 pct, cost per cycle $0.0150479 -> $0.0068823 (CHEAPER)
measured
r001-unbilled-watch against s001-unbilled-watch-stateless. Same corpus, same model, same grader, same 150 cycles; the prompts differ in exactly one block and evals/check_labels.py asserts they are byte-identical everywhere else. ⚠︎ The day-count half of this row is confounded -- see Eval.could_not_verify.
r001-unbilled-watch against b000-unbilled-watch-onepage over the same 150 cycles and the same scorer. The blocker field moves the WRONG way and the kit publishes it: an ordered keyword table beats the model by 6.67 points.
whether the account's answer depends on a cycle you are not holding
band 100.0 pct on the 25 memory-dependent cells against 99.2 pct on the other 125; for the free floor, 20.0 pct against 100.0 pct
measured
results/eval-r001-unbilled-watch.json and results/eval-b000-unbilled-watch-onepage.json, cells grouped by the memory_dependent flag the generator set. EVERY one of the free floor's band misses is a memory-dependent cell; it is exactly right on the other 125.
which blocker code the answer key applies
blocker code 91.33 pct -> 96.0 pct. Seven cycles move; the band does NOT move at all
measured
results/ambiguity-r001-unbilled-watch.json, produced by re-scoring the recorded replies under the other reading of READ_DISPUTE. No new calls. The band is unmoved because READ_DISPUTE and CAUSE_UNKNOWN carry the same 30-day lead time.
which blocker the model NAMES, when it names the wrong one
one band cell, from WITHIN_LIMIT to APPROACHING -- the only band miss in the run
measured
ACC-0034-C2 on r001-unbilled-watch. Day count 304 and 61 remaining, both exactly right; the model applied NO_ACCESS's 60-day lead time where the key says READ_DISPUTE's 30. The lead time is what turns a day count into a band, so a mis-mapped blocker moves the band silently.
the published output ceiling
0 of 150 replies lost at 24,000 with the accrual; 1 of 150 lost at the SAME ceiling without it; largest replies 12,664 and 24,000
measured
eval-c000 (cap 32,000 on the three hardest accounts, largest reply 5,934), eval-r001-unbilled-watch (cap 24,000, largest 12,664, 0 failures), eval-s001-unbilled-watch-stateless (same cap, 1 failure at the cap).
rendering the carried accrual as English rather than as JSON
unknown -- the experiment would cost one more 150-call run and has not been paid for
reasoning
src/state.describe's own docstring argues for English on the grounds that the rule text says 'days do not count towards the window' in English, and states plainly that this is a design choice rather than a measurement.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
proximity band accuracy
nothing yet
band on the watchlist
any cycle at all reported quiet whose correct band is not. The cost of that direction is the whole unbilled balance
blocker code accuracy
any answer other than READ_DISPUTE or CAUSE_UNKNOWN on the 12 cycles carrying the disputed sentence -- three of this run's answers are NO_ACCESS and neither reading recovers them
day count, exact
any error larger than one day, or any error on a cycle with no suspension at all. Both would be a different failure from the boundary convention this run shows
watchlist cycles reported quiet -- the expensive direction
any cycle at all whose correct band is APPROACHING, AT_RISK or BREACHED and which is reported WITHIN_LIMIT or NO_LIMIT. One is enough: this arm has none, so a first movement is a new failure mode and not a worse rate
clock status
any drop at all, and read it against the SUSPENDED and NOT_PROTECTED cycles rather than against the total -- a miss on a RUNNING cycle would be a different failure from a miss on a suspended one
false alarms
nothing yet
the memory-dependent subset
any band miss at all on this subset
the portfolio exposure total
a small error_pct beside a falling band figure -- the signature of cancellation
the denominators themselves
any change at all, on any run
input volume
any change without a corresponding change to the prompt or the corpus
output volume
nothing yet
provider-side reasoning share
nothing yet
latency
nothing yet
replies that did not parse
any unparsed reply on the published configuration
NextThe three you would add first
A human step in front of anything that writes a balance offBREACHED is the highest-consequence output this kit has and it is a claim that a household's charge is unrecoverable. Both no-memory arms get the portfolio total wrong in opposite directions (49.39 pct over and 65.98 pct under), and the stateful arm's exactness is measured on one run of an invented rule.
The real back-billing window, before anything elseBACKBILL_WINDOW_DAYS = 365 is invented and the catalogue row records the real value as an open question shared with a sibling row. Every band on these pages is agreement with an invented number.
A refusal state for a reply that did not parseThe scored run had none, but the stateless control lost one call to the ceiling and it scored as four wrong cells. In a deployment that cycle needs to be visibly UNANSWERED rather than silently wrong, and the kit has no vocabulary for that -- the scorer's 'absent' verdict exists but nothing downstream reads it.
A state store with a concurrency modeldata/state.json is one file replaced atomically. That is right for one process and is not a design for a scheduler, and the carried quantity is a running total, so a lost write is permanent rather than transient.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/backbill.py, src/segment.py or src/select.py -- it refuses to let a run spend if the corpus, the answer key, the experimental control, the privacy guard or the no-write-off property has moved. Re-run the scored eval AND its stateless control together (paid, 300 calls) on any change to src/prompt.py or src/state.py: the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. evals/ambiguity.py is free and should be re-run whenever the blocker taxonomy is touched.
What this cannot tell you
One run per arm. Whether the 12.66-point band gap against the free floor is stable across repeats is not measured, and the small denominators (54 watchlist cycles, 25 memory-dependent cells, 1 false alarm) would make a repeat noisy.
Whether the no-propagation property would hold under a wrong answer mid-chain. It is true by construction -- backbill.step() never reads the reply -- and this run is only consistent with it.
Whether the no-write-off guarantee holds against a path named something the assertion does not know. It greps for four function names and two method calls; the guarantee is about ABSENCE and that is the only mechanical form it can take.
Whether telling the model which lead time to use, or spelling out the half-open interval with a worked example, would remove the run's four remaining error shapes. Neither has been tried; each costs one more 150-call run.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried accrual
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries an ACCRUAL -- four scalar fields written by arithmetic. The whole reason the cost is flat in history length is that nothing here remembers what was said, only what was counted, and a memory layer would give back the growth this design exists to avoid. It would also give back the failure this design exists to avoid: a running total fed from its own model output compounds forever
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the account chain
evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect)
this is the seam where a framework would genuinely earn its place, and the kit says so: 50 independent chains of 3 strictly-ordered steps, with cadence, retries and late-arriving cycles all outside the kit. A ThreadPoolExecutor is the right size for an eval and the wrong size for a monitor that actually runs on a billing calendar
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
four exact-match comparisons, three slices of the same cells and one dollar sum is a dict comprehension, not a platform
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each account is a chain of three steps -- cycle 1, cycle 2, cycle 3 -- with no branching and exactly one edge between consecutive cycles, carrying four fields. Different accounts never touch. A framework would add an orchestrator to a for-loop that already runs 50 chains wide.
The other sideWhat a framework costs you
No scheduler. This is a monitor that carries state and is still INVOKED, not woken -- the cadence, the retry policy and the question of what to do when a cycle does not arrive are all outside the kit.
No state store worth the name. data/state.json is one file replaced atomically; a framework would bring a checkpointer with a concurrency model, which this does not have -- and on a running total that gap is worse than on a per-request kit.
No observability beyond what evals/run.py prints and writes to results/. There is no tracing and no dashboard integration.
No retry/backoff beyond src/adapters' own bounded retry, which covers a busy provider and a dropped connection and nothing else.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-unbilled-watch on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
12,343 ms
not yet known
nothing yet
Model, p95
55,907 ms
not yet known
nothing yet
Input tokens
222,929
1.7 pct between the two arms, and it is a measurement rather than a spread
any change without a corresponding change to the prompt or the corpus
Output tokens
306,959
not yet known
nothing yet
No movement column. Not one of the 2 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-unbilled-watch12,343 ms
s001-unbilled-watch-stateless22,557 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
1 run not plotted. b000-unbilled-watch-onepage recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — one period report goes whole into one call, and the prior period's carried state goes with it.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 3 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
cycle statements
data/corpus/<ACCOUNT>-C<n>.txt -- 150 files, 559,640 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Customer Contact -- the household's name, supply address, phone and email -- never does, by src/select.NEVER_SENT
the answer key
data/gold.jsonl -- 150 rows, the output of src/backbill.replay over the planted inputs, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the carried accrual
data/state.json in a deployment (src/state.py, written atomically). ⚠︎ IN AN EVAL IT IS SCOPED TO THE RUN AND NEVER TOUCHES DISK -- evals/run.py builds a fresh in-memory store per account, because a run that started from the previous run's memory could not be re-run or compared with its own control
one short English sentence per call, produced by state.describe -- four scalar fields, never a prior statement and never a prior reply
the recorded runs
results/eval-*.json -- the scored run, the stateless control, the free one-page floor, the calibration and the wiring stub, plus results/ambiguity-*.json, all committed
never -- they are read by the page and by evals/ambiguity.py, which re-scores them for $0.00
the key
.env or the shared repo-root .env -- never committed, 0600
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py). src/app.py redacts it out of any error it renders
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
windowing
the scoreable unit, made before anything judges it. 150 documents become 50 chains of three consecutive monthly cycle statements, processed strictly in order by evals/run.py because cycle 3's prompt contains an accrual produced by cycle 2. Different accounts are independent and run concurrently, 50 chains wide; cycles inside one never can be. ⚠︎ NOTHING IS CUT FROM A STREAM HERE -- the corpus arrives as discrete cycle statements and the 'window' is simply which three belong to one account. The decision keeps the variant's name because the contract requires it; what it actually decides is the chain, and the thing carried across the boundary has its own row below.
50 chains x 3 cycles = 150 calls, 0 gaps and 0 out-of-order cycles; evals/check_labels.py asserts all 50 sequences are complete before a run may spend. Wall clock 345.3s at 12 workers against 150 serial calls (r001-unbilled-watch, evals/check_labels.py)
three cycles. Long enough to hide a suspension behind a boundary and not long enough to test an accrual that has been running for years -- and the one error shape this run has (one day out on a half-open interval) is exactly the shape a longer chain accumulates, because the carried quantity is a running total
a chain longer than three cycles invalidates the day-count headline rather than extending it: one day out per cycle is forty days out after forty cycles, and nothing here measures that
state
four scalar fields per account -- days already counted towards the window, whether a suspension is open, the band last reported, whether it is still protected -- written by src/backbill.step from the dates PARSED off the page and never from the model's reply, and rendered by src/state.describe into one English sentence. That single sentence is the entire route from one cycle to the next, and it is what both scored arms differ by.
186 characters on the worked example. Input tokens 222,929 with it against 219,169 without -- and OUTPUT 306,959 against 710,849, so the accrual CUTS the bill: $0.0068823 a cycle with it against $0.0150479 without. Band 99.33 pct against 68.67 pct. (r001-unbilled-watch against s001-unbilled-watch-stateless, lenses.LLM.prompt_parts)
the accrual is four scalars, so cycle 40 costs what cycle 3 costs -- the OPPOSITE curve to an intake kit, whose input grows with the square of the turns. What is NOT bounded is accuracy over a longer history, which is unmeasured, and the store itself: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model
lose the carried accrual and the day count collapses -- 43.33 pct against 98.0 -- with every reply still well-formed and nothing raising an error. See environment.signatures' traceless row
model
one call per CYCLE carrying the instruction, the carried-accrual sentence and six of the statement's seven sections, behind src/adapters/__init__.py, at the published MAX_TOKENS = 24,000 with provider-side reasoning left at the default -- the configuration both scored arms ship
0 of 150 replies failed to parse on the scored run; the largest was 12,664 output tokens (53 pct of the cap). The stateless control at the IDENTICAL ceiling lost one call outright, which spent the whole budget and returned nothing. (lenses.Eval.scores, r001-unbilled-watch and s001-unbilled-watch-stateless)
the 24,000-token ceiling was set at 4.0 times the largest reply the calibration had seen (5,934 on 9 calls at a 32,000 cap), deliberately high because a ceiling is not billed. The scored run came within 53 pct of it anyway, and the control went over
one tier was run, so nothing here compares two models; the three arms differ by MEMORY and by whether a model was called at all. Point .env at your own model and the free scorer re-runs on your numbers
labels
data/gold.jsonl, 150 rows, computed by src/backbill.replay over the planted inputs at generation time -- a function the runtime never calls, so the key is arithmetic done independently of the code under test rather than the code under test marking its own homework
71 WITHIN_LIMIT / 25 APPROACHING / 11 AT_RISK / 18 BREACHED / 25 NO_LIMIT over 150 cycles; 54 on the watchlist, 96 quiet, 25 memory-dependent, 32 cycles carrying a customer-sounding standing reason with no suspension open against it (lenses.Eval.dataset, unbilled-watch-2026-08-23-50accounts-150cycles)
⚠︎ THE KEY IS ONLY AS GOOD AS AN INVENTED RULE. BACKBILL_WINDOW_DAYS = 365 and all six lead times were written for this kit and confirmed against nothing; the catalogue row records the real value as an open question shared with a sibling row. And one of the six blocker codes is arguably named for something other than what it covers -- evals/ambiguity.py prices that at 12 cycles and 4.67 points
your own book: hand-label the gold, which is the real work. This kit's key is a luxury of controlling both the generator and the rules, and a hand-labelled key has an error rate nothing here has measured
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
an account reported BREACHED whose exception log shows a CLOSED suspension with no OPENED entry
something read the statement without the carried accrual and counted every day since the oldest unbilled read, including the stretch the customer was responsible for
check the carried accrual before disputing the answer. On ACC-0007-C3 the free one-page floor reports BREACHED at 396 days and the truth is AT_RISK at 360 -- a 36-day difference that is entirely a suspension the page cannot see (results/eval-b000-unbilled-watch-onepage.json against results/eval-r001-unbilled-watch.json, ACC-0007-C3)
an account on the watchlist whose day count is right and whose band looks one step too severe
the lead time came from the wrong blocker. The band is the days LEFT compared with the remediation lead time for the standing block, so a mis-mapped blocker moves the band with the arithmetic untouched
read the rationale for the lead time it names, not the day count. ACC-0034-C2 says '61 days remaining, which is fewer than twice the 60-day NO_ACCESS lead time' -- 61 and 304 are both correct and READ_DISPUTE's lead time is 30, which makes it WITHIN_LIMIT. It is the only band miss in the run (results/eval-r001-unbilled-watch.json, ACC-0034-C2)
a blocker code of CAUSE_UNKNOWN on an account whose recorded reason reads perfectly clearly
the bucket is being used as a refuge rather than as a bucket. The model got 18 of 18 genuinely unrecorded reasons right on this run and invented ten more, on two phrasings
group the CAUSE_UNKNOWN rows by their recorded reason before working them. Ten of this run's are one sentence about a reading the customer never sent, and the free one-page floor maps all of them without a model (results/eval-r001-unbilled-watch.json, blocker_cause cells, and results/ambiguity-r001-unbilled-watch.json)
No machine symptom — this failure leaves no trace in any output.
There is none, and this is the failure with no machine symptom: a deployment that quietly lost its carried accrual would keep answering, keep parsing, keep costing money -- MORE money, 2.32x the output tokens -- and keep reporting a portfolio exposure 65.98 pct BELOW the truth, with every reply well-formed and nothing anywhere raising an error. No check in this kit compares the two arms at run time. The only thing that would catch it is noticing that the exposure total fell, which on a quiet month is exactly what everybody hopes to see.
⚠︎ THIS KIT'S FLOW VARIANT IS 'triage' AND IT IS THE WRONG SHAPE, INHERITED RATHER THAN CHOSEN. monitor is a pattern in build/domain/readiness.py with no flow variant of its own, so a cadence kit borrows triage and inherits station labels that do not describe it -- 'Cut into windows' over a thing that is not a window, 'Page threshold' over arithmetic that pages nobody. They are used unchanged rather than reworded, because inventing station words silently breaks lens reachability while a wrong-sounding label with an honest caption under it is merely visible. ⚠︎ AND index_answer IS REQUIRED VERBATIM FROM A CLOSED LIST WHOSE NEAREST MEMBER SAYS 'period report' WHERE THIS KIT SAYS 'cycle statement'. The closest accepted wording is shipped unchanged rather than reworded, and the mismatch is reported here instead of being papered over; the proposed addition is recorded with the kit rather than edited into the standard. WHAT IS ALSO NOT MEASURED: concurrency and worker sizing -- 50 account chains, 12 workers, nothing measured past 150 calls. Whether cycles within an account can ever be run out of order -- they cannot, by construction, and no throughput figure accounts for that serial constraint at longer histories. Provider-side retention, training use and log residency -- provider-dependent, a third state. Whether the provider's prompt cache was hit on any call -- every figure prices the cache-miss rate for exactly this reason, and the fixed head of this prompt (2,491 of 6,356 characters, byte-identical on all 150 calls) is precisely what a cache would have discounted. Whether either arm repeats -- every configuration ran once.
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every account, tariff, meter, reading, exception entry, household name and address is invented here. Verified against the repository's own LICENSE file on 2026-08-23. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineThe band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
For each of the 150 cycles and each of the four answered fields, did the reply equal the computed answer key? Band, clock and blocker are compared exactly; the day count is compared as a whole number, and a reply that is not a whole number of days is a MISS rather than a zero -- 0 is a meaningful answer here (it is what an unprotected account gets) and a parse failure must not be scored as one.
$0.00per 1,000 cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/ambiguity.py are scored through.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The cycle
ACC-0007-C3 -- a spanning suspension (pattern SE), cycle 3 of 3
Dates on the statement
Oldest unbilled actual read 2025-07-04; cycle end 2026-08-04; 31 days counted this cycle. Last bill issued 2026-07-01 -- an ESTIMATE, and the distractor.
What the exception log carries
one exception entry: 'suspension CLOSED' on 2026-07-23. There is no OPENED entry anywhere on this statement to pair it with.
Carried accrual, in the prompt
As at the previous cycle, 347 days of this account's unbilled period already counted towards the back-billing window. A customer-attributable suspension was OPEN at the end of that cycle, so the window has been stopped since then and stays stopped until this statement records it closing. It was reported APPROACHING.
The free one-page floor
BREACHED at 396 days -- it counts every day since the oldest read, because it cannot know when the stop began
The same model, memory removed
WITHIN_LIMIT at 31 days -- told there is no history, it counted only the days this cycle
The model, with the carried accrual
AT_RISK at 359 days, clock RUNNING, blocker EXCEPTION_HOLD
Ground truth
AT_RISK at 360 days, clock RUNNING, blocker EXCEPTION_HOLD
Scored as
3 of 4 cells correct -- the day count is ONE day out, the only kind of day-count error in the whole run, and the band, clock and blocker are all right. The floor would have told finance to write this balance off.
Grader
Verdict
Why
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
band hit, clock hit, blocker hit, day count MISS by one
ACC-0007-C3, cycle 3 of a spanning-suspension account. The statement carries one exception entry -- 'suspension CLOSED' on 2026-07-23 -- with no OPENED anywhere to pair it with. The carried accrual said 347 days were already counted and a suspension was open. The model answered AT_RISK at 359 days against a key of 360: one day out, on the half-open interval convention (a suspension runs from the day it opens up to but NOT including the day it closes). Band, clock and blocker all match.
Band accuracy restricted to the cycles that need somebody to touch them
hit -- one of the 54
This cycle's correct band is AT_RISK, so it is inside the 54 cycles this grader scores. The stateful arm got 100.0 pct of them; the free one-page floor got 81.48 pct and the stateless control 24.07 pct. On THIS cycle the floor answered BREACHED -- it would have told finance the balance was already unrecoverable.
False alarms on the cycles where nothing needs doing
not in scope
This grader scores only the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT. This one is AT_RISK, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The cycles whose answer cannot be derived from their own page
band hit, day count miss
This cycle is one of the 25 memory-dependent cells: 36 days of suspension happened in a cycle whose statement is not in the prompt. The stateful arm got the BAND right on 100.0 pct of these and the day count on 92.0 pct; the free floor got 20.0 pct and 0.0 pct, and the stateless control 36.0 pct and 0.0 pct. This is the subset the whole monitor argument rests on.
The portfolio total reported as beyond the back-billing window
in scope, and correctly excluded
This grader sums the unbilled value of every cycle reported BREACHED. This account is AT_RISK, so the correct behaviour is to leave its $1,185.92 OUT of the total -- which the stateful arm did and the free floor did not. Across all 150 cycles the stateful arm reported $29,024.35 against a true $29,024.35 (0.0 pct error); the floor reported $43,358.99 (49.39 pct over) and the stateless control $9,873.05 (65.98 pct under).
The formulaWhat it computes
accuracy = hits / 150 per field. A cycle whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried accrual
99.3% band accuracy · 3 more measured on this row
the fast tier, memory removed (THE CONTROL)
68.7% band accuracy · 3 more measured on this row
the free one-page floor, no model
86.7% band accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/backbill.replay over the planted inputs at generation time. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known to be arguable -- the NAME of one of its six blocker codes. READ_DISPUTE is named for a disagreement and covers a sentence describing an omission; evals/ambiguity.py prices that at 12 cycles and 4.67 points of blocker accuracy. The other three fields have no known ambiguity.
Watch these
band_accuracy_pct
days_at_risk_accuracy_pct
blocker_cause_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A cycle that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. The scored run had 0 of 150 -- but the stateless control at the identical ceiling had 1, so 0 is this arm's result and not a property of the design.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because two of them are small: 54 watchlist cycles and 25 memory-dependent cycles, so one row moves the first by 1.9 points and the second by 4.0.
Cadence: Re-run evals/check_labels.py (free) on any change to the corpus generator, src/backbill.py, src/segment.py or src/select.py. Re-run the scored eval (paid, one call per cycle) on any change to src/prompt.py or src/state.py -- both of those change what the model is told.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real unbilled book, where whether a delay was the customer's own is decided by a person reading a note. That is why this corpus is generated rather than captured.
Band accuracy restricted to the cycles that need somebody to touch them
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineBand accuracy restricted to the cycles that need somebody to touch them
Of the 54 cycles whose correct band is not WITHIN_LIMIT and not NO_LIMIT, how many were banded correctly? This is the number the kit is actually about; the 150-cycle average is 64 pct quiet and something that answers WITHIN_LIMIT to everything scores well on it.
$0.00per 1,000 cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader -- a slice of the same cells, not a second comparison.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The cycle
ACC-0007-C3 -- a spanning suspension (pattern SE), cycle 3 of 3
Dates on the statement
Oldest unbilled actual read 2025-07-04; cycle end 2026-08-04; 31 days counted this cycle. Last bill issued 2026-07-01 -- an ESTIMATE, and the distractor.
What the exception log carries
one exception entry: 'suspension CLOSED' on 2026-07-23. There is no OPENED entry anywhere on this statement to pair it with.
Carried accrual, in the prompt
As at the previous cycle, 347 days of this account's unbilled period already counted towards the back-billing window. A customer-attributable suspension was OPEN at the end of that cycle, so the window has been stopped since then and stays stopped until this statement records it closing. It was reported APPROACHING.
The free one-page floor
BREACHED at 396 days -- it counts every day since the oldest read, because it cannot know when the stop began
The same model, memory removed
WITHIN_LIMIT at 31 days -- told there is no history, it counted only the days this cycle
The model, with the carried accrual
AT_RISK at 359 days, clock RUNNING, blocker EXCEPTION_HOLD
Ground truth
AT_RISK at 360 days, clock RUNNING, blocker EXCEPTION_HOLD
Scored as
3 of 4 cells correct -- the day count is ONE day out, the only kind of day-count error in the whole run, and the band, clock and blocker are all right. The floor would have told finance to write this balance off.
Grader
Verdict
Why
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
band hit, clock hit, blocker hit, day count MISS by one
ACC-0007-C3, cycle 3 of a spanning-suspension account. The statement carries one exception entry -- 'suspension CLOSED' on 2026-07-23 -- with no OPENED anywhere to pair it with. The carried accrual said 347 days were already counted and a suspension was open. The model answered AT_RISK at 359 days against a key of 360: one day out, on the half-open interval convention (a suspension runs from the day it opens up to but NOT including the day it closes). Band, clock and blocker all match.
Band accuracy restricted to the cycles that need somebody to touch them
hit -- one of the 54
This cycle's correct band is AT_RISK, so it is inside the 54 cycles this grader scores. The stateful arm got 100.0 pct of them; the free one-page floor got 81.48 pct and the stateless control 24.07 pct. On THIS cycle the floor answered BREACHED -- it would have told finance the balance was already unrecoverable.
False alarms on the cycles where nothing needs doing
not in scope
This grader scores only the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT. This one is AT_RISK, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The cycles whose answer cannot be derived from their own page
band hit, day count miss
This cycle is one of the 25 memory-dependent cells: 36 days of suspension happened in a cycle whose statement is not in the prompt. The stateful arm got the BAND right on 100.0 pct of these and the day count on 92.0 pct; the free floor got 20.0 pct and 0.0 pct, and the stateless control 36.0 pct and 0.0 pct. This is the subset the whole monitor argument rests on.
The portfolio total reported as beyond the back-billing window
in scope, and correctly excluded
This grader sums the unbilled value of every cycle reported BREACHED. This account is AT_RISK, so the correct behaviour is to leave its $1,185.92 OUT of the total -- which the stateful arm did and the free floor did not. Across all 150 cycles the stateful arm reported $29,024.35 against a true $29,024.35 (0.0 pct error); the floor reported $43,358.99 (49.39 pct over) and the stateless control $9,873.05 (65.98 pct under).
The formulaWhat it computes
hits among band cells where gold is APPROACHING, AT_RISK or BREACHED, divided by 54.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried accrual
100.0% accuracy
the fast tier, memory removed (THE CONTROL)
24.1% accuracy
the free one-page floor, no model
81.5% accuracy
In operationWhat to monitor
Reference standard: the same data/gold.jsonl the reference grader uses, restricted to cells whose band is not WITHIN_LIMIT or NO_LIMIT. It measures the reference grader's own answers on a subset, so it has no independent standard.
These rates are UNKNOWN, on purpose
Whether 100.0 pct is a ceiling or a corpus. This arm got every watchlist cycle right, so nothing here bounds how it behaves on a harder mix -- a saturated grader tells you the corpus stopped discriminating, not that the model cannot be beaten. The denominator is also a property of the generator's band mixture, which is stated in tools/build_corpus.py and is a decision rather than a finding.
Watch these
band_on_watchlist_pct
missed_watchlist
watchlist_cells
Alarm on
any cycle reported quiet whose correct band is not. This arm has none; one would be a new failure mode rather than a worse rate, because the cost is the whole unbilled balance rather than an hour of somebody's time.
How tight can the band be? 54 rows means the finest honest band is about 1.9 points. Nothing is tuned.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
The book is mostly comfortable, which every real one is.
Do not use it
As a headline on its own. It ignores the 96 quiet cycles entirely, so something that escalated everything would score 100 pct here and be useless.
False alarms on the cycles where nothing needs doing
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineFalse alarms on the cycles where nothing needs doing
Of the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT, how many were put on the watchlist anyway? This is the other direction, and it costs somebody an hour where the first costs the balance.
$0.00per 1,000 cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The cycle
ACC-0007-C3 -- a spanning suspension (pattern SE), cycle 3 of 3
Dates on the statement
Oldest unbilled actual read 2025-07-04; cycle end 2026-08-04; 31 days counted this cycle. Last bill issued 2026-07-01 -- an ESTIMATE, and the distractor.
What the exception log carries
one exception entry: 'suspension CLOSED' on 2026-07-23. There is no OPENED entry anywhere on this statement to pair it with.
Carried accrual, in the prompt
As at the previous cycle, 347 days of this account's unbilled period already counted towards the back-billing window. A customer-attributable suspension was OPEN at the end of that cycle, so the window has been stopped since then and stays stopped until this statement records it closing. It was reported APPROACHING.
The free one-page floor
BREACHED at 396 days -- it counts every day since the oldest read, because it cannot know when the stop began
The same model, memory removed
WITHIN_LIMIT at 31 days -- told there is no history, it counted only the days this cycle
The model, with the carried accrual
AT_RISK at 359 days, clock RUNNING, blocker EXCEPTION_HOLD
Ground truth
AT_RISK at 360 days, clock RUNNING, blocker EXCEPTION_HOLD
Scored as
3 of 4 cells correct -- the day count is ONE day out, the only kind of day-count error in the whole run, and the band, clock and blocker are all right. The floor would have told finance to write this balance off.
Grader
Verdict
Why
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
band hit, clock hit, blocker hit, day count MISS by one
ACC-0007-C3, cycle 3 of a spanning-suspension account. The statement carries one exception entry -- 'suspension CLOSED' on 2026-07-23 -- with no OPENED anywhere to pair it with. The carried accrual said 347 days were already counted and a suspension was open. The model answered AT_RISK at 359 days against a key of 360: one day out, on the half-open interval convention (a suspension runs from the day it opens up to but NOT including the day it closes). Band, clock and blocker all match.
Band accuracy restricted to the cycles that need somebody to touch them
hit -- one of the 54
This cycle's correct band is AT_RISK, so it is inside the 54 cycles this grader scores. The stateful arm got 100.0 pct of them; the free one-page floor got 81.48 pct and the stateless control 24.07 pct. On THIS cycle the floor answered BREACHED -- it would have told finance the balance was already unrecoverable.
False alarms on the cycles where nothing needs doing
not in scope
This grader scores only the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT. This one is AT_RISK, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The cycles whose answer cannot be derived from their own page
band hit, day count miss
This cycle is one of the 25 memory-dependent cells: 36 days of suspension happened in a cycle whose statement is not in the prompt. The stateful arm got the BAND right on 100.0 pct of these and the day count on 92.0 pct; the free floor got 20.0 pct and 0.0 pct, and the stateless control 36.0 pct and 0.0 pct. This is the subset the whole monitor argument rests on.
The portfolio total reported as beyond the back-billing window
in scope, and correctly excluded
This grader sums the unbilled value of every cycle reported BREACHED. This account is AT_RISK, so the correct behaviour is to leave its $1,185.92 OUT of the total -- which the stateful arm did and the free floor did not. Across all 150 cycles the stateful arm reported $29,024.35 against a true $29,024.35 (0.0 pct error); the floor reported $43,358.99 (49.39 pct over) and the stateless control $9,873.05 (65.98 pct under).
The formulaWhat it computes
quiet cells whose reply names a watchlist band, divided by 96.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried accrual
1.0% false alarm rate
the fast tier, memory removed (THE CONTROL)
1.0% false alarm rate
the free one-page floor, no model
9.4% false alarm rate
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to the 96 cells whose band is WITHIN_LIMIT or NO_LIMIT.
These rates are UNKNOWN, on purpose
Its own rate cannot be separated from the reference grader's, and with a single positive there is nothing here to estimate a rate FROM. One row of 96 is consistent with a true rate anywhere from a fraction of a percent to several percent, and no repeat run exists to narrow it.
Watch these
false_alarm_rate_pct
false_alarms
unparsed_replies
Alarm on
any false alarm whose day count is CORRECT. This run's one is exactly that shape -- the arithmetic was right and the lead time was wrong -- which points at the blocker taxonomy rather than at the model's reasoning, and those need different fixes.
How tight can the band be? 96 rows, so one row is 1.04 points. No threshold is tuned.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Always, and BESIDE the watchlist rate rather than under it. The two move in opposite directions and a single accuracy hides the trade entirely.
Do not use it
As a measure of caution on an arm that under-alarms by construction -- on the stateless control it is 1.04 pct and means almost nothing, because that arm reports 39 of 54 watchlist cycles as quiet.
The cycles whose answer cannot be derived from their own page
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineThe cycles whose answer cannot be derived from their own page
On the 25 cycles where a suspension began, or a dispute was raised, in a statement the model is not holding, was the band right and was the day count right? This is the subset the whole monitor argument rests on, and it is marked in the answer key by the generator's own structure -- never by what any arm answered.
$0.00per 1,000 cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The cycle
ACC-0007-C3 -- a spanning suspension (pattern SE), cycle 3 of 3
Dates on the statement
Oldest unbilled actual read 2025-07-04; cycle end 2026-08-04; 31 days counted this cycle. Last bill issued 2026-07-01 -- an ESTIMATE, and the distractor.
What the exception log carries
one exception entry: 'suspension CLOSED' on 2026-07-23. There is no OPENED entry anywhere on this statement to pair it with.
Carried accrual, in the prompt
As at the previous cycle, 347 days of this account's unbilled period already counted towards the back-billing window. A customer-attributable suspension was OPEN at the end of that cycle, so the window has been stopped since then and stays stopped until this statement records it closing. It was reported APPROACHING.
The free one-page floor
BREACHED at 396 days -- it counts every day since the oldest read, because it cannot know when the stop began
The same model, memory removed
WITHIN_LIMIT at 31 days -- told there is no history, it counted only the days this cycle
The model, with the carried accrual
AT_RISK at 359 days, clock RUNNING, blocker EXCEPTION_HOLD
Ground truth
AT_RISK at 360 days, clock RUNNING, blocker EXCEPTION_HOLD
Scored as
3 of 4 cells correct -- the day count is ONE day out, the only kind of day-count error in the whole run, and the band, clock and blocker are all right. The floor would have told finance to write this balance off.
Grader
Verdict
Why
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
band hit, clock hit, blocker hit, day count MISS by one
ACC-0007-C3, cycle 3 of a spanning-suspension account. The statement carries one exception entry -- 'suspension CLOSED' on 2026-07-23 -- with no OPENED anywhere to pair it with. The carried accrual said 347 days were already counted and a suspension was open. The model answered AT_RISK at 359 days against a key of 360: one day out, on the half-open interval convention (a suspension runs from the day it opens up to but NOT including the day it closes). Band, clock and blocker all match.
Band accuracy restricted to the cycles that need somebody to touch them
hit -- one of the 54
This cycle's correct band is AT_RISK, so it is inside the 54 cycles this grader scores. The stateful arm got 100.0 pct of them; the free one-page floor got 81.48 pct and the stateless control 24.07 pct. On THIS cycle the floor answered BREACHED -- it would have told finance the balance was already unrecoverable.
False alarms on the cycles where nothing needs doing
not in scope
This grader scores only the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT. This one is AT_RISK, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The cycles whose answer cannot be derived from their own page
band hit, day count miss
This cycle is one of the 25 memory-dependent cells: 36 days of suspension happened in a cycle whose statement is not in the prompt. The stateful arm got the BAND right on 100.0 pct of these and the day count on 92.0 pct; the free floor got 20.0 pct and 0.0 pct, and the stateless control 36.0 pct and 0.0 pct. This is the subset the whole monitor argument rests on.
The portfolio total reported as beyond the back-billing window
in scope, and correctly excluded
This grader sums the unbilled value of every cycle reported BREACHED. This account is AT_RISK, so the correct behaviour is to leave its $1,185.92 OUT of the total -- which the stateful arm did and the free floor did not. Across all 150 cycles the stateful arm reported $29,024.35 against a true $29,024.35 (0.0 pct error); the floor reported $43,358.99 (49.39 pct over) and the stateless control $9,873.05 (65.98 pct under).
The formulaWhat it computes
hits among cells where gold.memory_dependent is true, divided by 25, reported separately for the band and for the day count.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried accrual
100.0% band accuracy · 1 more measured on this row
the fast tier, memory removed (THE CONTROL)
36.0% band accuracy · 1 more measured on this row
the free one-page floor, no model
20.0% band accuracy · 1 more measured on this row
In operationWhat to monitor
Reference standard: the same data/gold.jsonl, restricted to the 25 cells the generator marked memory_dependent. evals/check_labels.py asserts that count and asserts that none of those cells is answerable from its own page.
These rates are UNKNOWN, on purpose
⚠︎ THIS GRADER IS SATURATED ON THE BAND FOR THE STATEFUL ARM (100.0 pct) AND IT SAYS SO RATHER THAN BANKING IT. 25 of 25 is a strong result and it is also 25 rows; what it cannot tell you is how the arm behaves on a hidden suspension long enough, or a lead time short enough, to be harder than these. The corpus draws the spanning-suspension patterns from SHORT-lead blockers on purpose (see Data.breaks_on), so this subset is easier to separate than a random one would be, and the separation is a property of that choice.
Watch these
memory_band_accuracy_pct
memory_days_accuracy_pct
memory_cells
Alarm on
any drop at all on the band. Every arm's number here is either 100 or structurally low, so a first movement on the stateful arm is a new failure rather than a worse rate.
How tight can the band be? 25 rows. One row is 4.0 points, so no band finer than about four points means anything on this grader.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Reading whether the memory is worth its price, rather than reading a headline. The 25 rows are invisible in a 150-cycle average.
Do not use it
As the kit's headline. It double-counts cells already inside the band figure, and a headline that summed them would report the same 25 answers twice.
The portfolio total reported as beyond the back-billing window
Catch unbilled power accounts before the charge expires
PresenterOpens the private repo. Visible to admins only.
In one lineThe portfolio total reported as beyond the back-billing window
Sum the unbilled value of every cycle the arm reported BREACHED, and compare it with the truth. This is the number that lands in a month-end pack: how much of the unbilled book is already unrecoverable. It is a DIFFERENT question from per-account accuracy and it is the question the catalogue row asks -- 'reconcile that against this month's unbilled-revenue estimate'.
$0.00per 1,000 cycle statements
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the reference grader.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The cycle
ACC-0007-C3 -- a spanning suspension (pattern SE), cycle 3 of 3
Dates on the statement
Oldest unbilled actual read 2025-07-04; cycle end 2026-08-04; 31 days counted this cycle. Last bill issued 2026-07-01 -- an ESTIMATE, and the distractor.
What the exception log carries
one exception entry: 'suspension CLOSED' on 2026-07-23. There is no OPENED entry anywhere on this statement to pair it with.
Carried accrual, in the prompt
As at the previous cycle, 347 days of this account's unbilled period already counted towards the back-billing window. A customer-attributable suspension was OPEN at the end of that cycle, so the window has been stopped since then and stays stopped until this statement records it closing. It was reported APPROACHING.
The free one-page floor
BREACHED at 396 days -- it counts every day since the oldest read, because it cannot know when the stop began
The same model, memory removed
WITHIN_LIMIT at 31 days -- told there is no history, it counted only the days this cycle
The model, with the carried accrual
AT_RISK at 359 days, clock RUNNING, blocker EXCEPTION_HOLD
Ground truth
AT_RISK at 360 days, clock RUNNING, blocker EXCEPTION_HOLD
Scored as
3 of 4 cells correct -- the day count is ONE day out, the only kind of day-count error in the whole run, and the band, clock and blocker are all right. The floor would have told finance to write this balance off.
Grader
Verdict
Why
The band, the clock, the day count and the blocker code, per cycle, exact match against the computed answer key
band hit, clock hit, blocker hit, day count MISS by one
ACC-0007-C3, cycle 3 of a spanning-suspension account. The statement carries one exception entry -- 'suspension CLOSED' on 2026-07-23 -- with no OPENED anywhere to pair it with. The carried accrual said 347 days were already counted and a suspension was open. The model answered AT_RISK at 359 days against a key of 360: one day out, on the half-open interval convention (a suspension runs from the day it opens up to but NOT including the day it closes). Band, clock and blocker all match.
Band accuracy restricted to the cycles that need somebody to touch them
hit -- one of the 54
This cycle's correct band is AT_RISK, so it is inside the 54 cycles this grader scores. The stateful arm got 100.0 pct of them; the free one-page floor got 81.48 pct and the stateless control 24.07 pct. On THIS cycle the floor answered BREACHED -- it would have told finance the balance was already unrecoverable.
False alarms on the cycles where nothing needs doing
not in scope
This grader scores only the 96 cycles whose correct band is WITHIN_LIMIT or NO_LIMIT. This one is AT_RISK, so the row is outside its denominator entirely -- shown rather than omitted, because a blank cell and an out-of-scope cell look identical and mean different things.
The cycles whose answer cannot be derived from their own page
band hit, day count miss
This cycle is one of the 25 memory-dependent cells: 36 days of suspension happened in a cycle whose statement is not in the prompt. The stateful arm got the BAND right on 100.0 pct of these and the day count on 92.0 pct; the free floor got 20.0 pct and 0.0 pct, and the stateless control 36.0 pct and 0.0 pct. This is the subset the whole monitor argument rests on.
The portfolio total reported as beyond the back-billing window
in scope, and correctly excluded
This grader sums the unbilled value of every cycle reported BREACHED. This account is AT_RISK, so the correct behaviour is to leave its $1,185.92 OUT of the total -- which the stateful arm did and the free floor did not. Across all 150 cycles the stateful arm reported $29,024.35 against a true $29,024.35 (0.0 pct error); the floor reported $43,358.99 (49.39 pct over) and the stateless control $9,873.05 (65.98 pct under).
The formulaWhat it computes
abs(sum(value where reported band == BREACHED) - sum(value where gold band == BREACHED)) / the gold total. The gold total is $29,024.35 of $158,520.97 unbilled in the whole book.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried accrual
0.0% error
the fast tier, memory removed (THE CONTROL)
66.0% error
the free one-page floor, no model
49.4% error
In operationWhat to monitor
Reference standard: the unbilled_value_usd field of data/gold.jsonl, summed over cells whose gold band is BREACHED. The values themselves are generated and asserted against a per-fuel household range by evals/check_labels.py.
These rates are UNKNOWN, on purpose
Whether 0.00 pct survives an arm that makes offsetting errors. This arm made none to offset, so the figure has never been exercised in the regime it is most likely to mislead in -- a run that misses two accounts and adds two others would report a small error and a badly wrong list, and nothing in this grader would say so.
Watch these
exposure_reported_usd
exposure_signed_error_usd
exposure_error_pct
Alarm on
a small error_pct beside a falling band figure. That combination is the signature of cancellation and is the one reading of this grader that is worse than useless.
How tight can the band be? No threshold. The gold total is dominated by 18 cycles across 6 accounts, so one account moves it by several percent -- the denominator is small in accounts even though it is large in dollars.
Cadence: Whenever the reference grader runs -- it is the same pass.
The decisionWhen to reach for it
Use it
Whenever a number leaves the operation and goes to finance. Per-account accuracy does not answer 'is the total right'.
Do not use it
⚠︎ AS A SUBSTITUTE FOR THE PER-ACCOUNT FIGURES, WHICH IS THE WHOLE HAZARD. This is a place errors CANCEL: two accounts wrong in opposite directions net to nothing here, so a report can be exactly right in aggregate and wrong on rows somebody then works. Both are published and neither is allowed to stand in for the other.
A living map of modern AI — kept current every morning