Every shift, you check whether a meter's missing readings are catching up or stuck, and guess which ones are gone for good. This app watches each meter-day and tells you, in plain words, whether today's run is the one that has to chase it.
PresenterOpens the private repo. Visible to admins only.
For the meter-data analystEnergy & Utilities · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A meter-data operations analyst at a utility, checking each meter-day's completeness every shift.
✕Today's manual process
1Open each meter-day still inside its settlement window, one at a time.
2Compare the gap count against yesterday's export, in a spreadsheet, to see if it moved.
3Read the exception notes to guess which quarantined reads are coming back.
4Miss the trend and a re-poll goes out too late, or not at all, before the window closes.
Every meter-day checked manually, shift after shift.
✓With the app
1Every meter-day is read the moment its settlement window opens.
2The gap count is compared to last time automatically, so nothing is guessed.
3Each quarantine reason is judged on its own words, not a keyword list.
4The re-poll call is made the moment it's needed, before the window closes.
The app flags only the ones that moved.
See it work
One real case: what the app read, step by step
One meter-day has been stuck at 10 missing intervals for two runs, and 5 of them are gone for good.
Catch a utility's missing meter readings in timeReference appBuilt to be shaped to your process
4
1What it remembers 10 intervals were missing then, no re-poll sent yet.
2This run's status Still stalled at 10 missing intervals, same as last time.
3What won't come back 5 missing reads are gone for good, not just late.
4The call Nothing moved, so this run raises the re-poll now.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"How many intervals are missing on this meter-day" is a count, and a twelve-line regex over the quality flags gets it exactly right for nothing -- this kit measures that and publishes it. The question a completeness watch actually exists for is the next one, and it has two terms neither the register nor a count can answer. Has the gap count MOVED since the last scheduled run? A feed filling in normally needs nothing; a feed that stopped needs a re-poll now, and the two look identical on one snapshot. And of the reads that are quarantined, which are coming back at all? That turns on a sentence a validation analyst typed, not on a code. Behind both there is a deadline that does not reopen: at the settlement close every interval still absent is substituted by an estimate and the true read is never put back. Somebody working the head-end's completeness report at shift change: opening each meter-day still inside its window, comparing the gap count against yesterday's export to see whether it has moved, reading the exception queue to decide which quarantined reads are coming back, and remembering which service points already have an on-demand read outstanding.
Audience
A meter-data operations analyst deciding what to chase this shift, and the settlement lead who has to report what share of the day's energy was settled against an estimate rather than a read. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual interval reads
The corpus is 120 interval reads, 0.72 MB (json 2 · jsonl 1 · txt 120). A UTILITY'S INTERVAL REGISTER IS THE HOUSEHOLD. A 15-minute load curve says when the occupants woke up, when the house emptied, whether anybody was home on Tuesday and whether they charge a car overnight. No utility publishes one, and the two alternatives were both worse than generating it: a scrubbed real extract is not de-identified, because the curve IS the identifying thing and taking the name off the top is the step that makes people forget that; and a kit with no corpus is a description of a kit. So the corpus is generated and the generator is committed beside it.
The corpus
The 120 interval readsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your interval reads. That is the whole change — there is no database to migrate.
One interval read, as the model receives itSP-0001-R1.txt · 1 of 120
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated interval-data completeness snapshot
for an AI use-case kit; it reproduces no utility's head-end extract, no market's
settlement timetable and no real premises. The settlement rules are ILLUSTRATIVE.
Service Point
----------------------------------------------------------------
Service point : SP-0001
Meter : MTR-456696
Channel : B1 -- kWh delivered, net site
Interval length : 15 minutes
Intervals expected : 96
Flow date (meter-day read) : 2026-03-01
Network region : NORTH-11
Market segment : large commercial, half-hourly settled
Settlement Terms
----------------------------------------------------------------
Settlement close : 2026-03-03 12:00 local
Re-poll recovery lag on file : 8.00 hours
Estimation policy : any interval still absent when the window shuts is
substituted by the estimation engine and the true read
is never put back
Rule G-1 A meter-day is COMPLETE when every EXPECTED interval on the channel carries a validated
actual read. The expected count is on the service point's own record (96 at 15 minutes,
48 at 30 minutes) and is never assumed.
Rule G-2 THE FLAG DECIDES, NOT THE VALUE. An interval reported 0.000 with flag A is a real
reading -- a vacant premises, a disconnected service, on-site generation offsetting
load -- and is NOT a gap. A head-end zero-fills an interval it never received and writes
Abridged — the file continues.
The outcomeWhat a good result looks like
Every meter-day inside its settlement window carries a status, a gap count, a completeness percentage and an explicit call on whether this run raises the re-poll -- with the ones that have stopped moving separated from the ones that are filling in. On this corpus that separation is worth 51 intervals: the arm with the carried state raises all 20 re-polls the rule requires and loses 83 intervals at close, and the identical arm with the memory removed raises 3 and loses 134.
And when it cannot
Two ways, and they cost different things. A MISSED re-poll is data nobody can recover: the window shuts, the estimation engine substitutes, and the true read is never put back. A DUPLICATE re-poll is an on-demand read queued at the head-end for data already on its way, and at the scale a utility runs this at, enough of them is the reason nobody reads the exception queue any more. On this corpus the model made 0 of each. The control made 18 missed and 1 duplicate; the memory-blind floor made 0 missed and 61 duplicates.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
You only need to know how many intervals are missing on each meter-day — a regex over the quality flags, or the b002 floor in this repo Counting M and Q in a 96-cell grid is exact and free. On this corpus b002 gets 100.00 pct of the gap counts and 100.00 pct of the completeness figures for nothing at all, which is exactly what the model gets.
Your head-end's exception queue carries coded reasons from a closed list — the b002 floor with a code lookup instead of its keyword table The model's entire margin on this corpus is reading free prose -- 80.83 against 72.50 on unrecoverable, 22 of 23 against 15 of 23 where one sentence decides. Take the prose away and there is nothing left to buy.
Your exception queue is prose a validation analyst types, and it names the meter whether or not the meter is the problem — this kit, with the carried state That is exactly the corpus these figures were measured on, and it is where the whole margin comes from. It is also where the model still fails: 20 of its 23 misses over-call a read as dead.
You want the fewest intervals lost and you can absorb the noise — the b000 floor -- chase every gap on every run Measured, and it is the uncomfortable answer: chasing everything loses 49 intervals at close against the correct rule's 83, because the re-poll lands earlier. Rule G-6 minimises alarms, not loss.
You want the watch to actually wake up on a schedule — your own scheduler -- cron, Airflow, Temporal, whatever already runs evals/run.py is INVOKED, not woken. Nothing in this kit detects a missed run, back-fills it or marks its readings late -- and on this task that absence is priced: one skipped run costs 6 intervals permanently, three cost 41.
At a glanceHow the whole thing runs
100%interval gap completeness pct
8,411 msp50, end to end
$5.43per 1,000 interval reads · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a utility's missing meter readings in time14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own head-end, or replace data/corpus/*.txt and data/gold.jsonl directly. The measured figures on this page do not travel with your corpus, and four of them are corpus properties rather than model properties.Corpus lens →
When is this the wrong choice?
Avoid: AVOID this kit entirely for that job. Paying a model per meter-day per run to count four letters is the most expensive way to get an answer a regex already has. That is the case against the best-fitting scenario (“You only need to know how many intervals are missing on each meter-day”). 5 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A HEAD-END THAT DOES NOT ZERO-FILL. This corpus prints an absent interval as 0.000 M, which is what makes the zero trap a trap. 6 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER THE INSTRUCTION-SHAPED OPERATIONAL NOTE MOVES THE ANSWER. It reaches the model verbatim on a third of the corpus and the raise/hold call was 100.00 pct correct anyway, which is consistent with the note having no effect and is not a measurement of it. 9 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-interval-gap. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces: the 120 snapshots and the answer key (regenerate them with python3 tools/build_corpus.py and diff), the fourteen pre-flight checks (python3 -m evals.check_labels), the whole cadence and per-arm loss table (python3 tools/cadence.py), all three free floors, and the UI with every locally computed panel plus a replay of what r001 actually answered. What needs a key is one thing: a new live reading.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
8,411 msp50, end to end
35,415 msp95
1 minclone to first result
What the clock covers. One completion call for one meter-day at one scheduled run, wall clock from the request leaving src/adapters to the reply parsing, measured over the 120 readings of r001-interval-gap at 12 concurrent meter-day chains. It does NOT include the corpus build, the register parse or the scorer, all of which are pure code. ⚠︎ The p95 is 4.2x the p50 because a reading is a 96-cell count and the model narrates its way through the hard grids; the same readings without the carried state ran to a 90.5s p95.
Current processWhat it replaces
Somebody working the head-end's completeness report at shift change: opening each meter-day still inside its window, comparing the gap count against yesterday's export to see whether it has moved, reading the exception queue to decide which quarantined reads are coming back, and remembering which service points already have an on-demand read outstanding.
Where it is not good enough
⚑ FOUR OF THE FIVE BANDS ARE A DEAD TIE WITH FREE CODE AND THE PAGE LEADS WITH THAT. The completeness figure, the status, the gap count and the raise/hold call all score 100.00 for the model and 100.00 for the b002 floor, which costs $0.00. Counting four letters in a ninety-six cell grid is what a regex is for and this kit will not pretend otherwise. The model's entire margin is one band: unrecoverable, 80.83 against 72.50, which is whether a quarantine reason means the read is coming back. On the 23 readings where exactly ONE reason line decides the answer, the model gets 22 and the keyword table gets 15.
⚑ AND THE FLOOR THAT LOSES EVERY ACCURACY BAND RECOVERS THE MOST DATA. b000 chases every gap on every run -- no memory, no trend, the alarm-fatigue rule everybody criticises -- and because it raises the re-poll at the first run instead of waiting for the trend, it loses 49 intervals at close against the correct rule's 83. Rule G-6 is the alarm-MINIMISING rule, not the loss-minimising one, and this kit measured that rather than assuming its own rule was right. The price of those 34 recovered intervals is 61 duplicate re-polls in 100 quiet readings. Which of those two a utility wants is a decision about its head-end's capacity and its analysts' attention, and nothing on this page can make it.
⚑ THE UNRECOVERABLE BAND FAILS IN THE MORE EXPENSIVE DIRECTION. 20 of the model's 23 misses OVER-call a read as dead when it was coming back -- which sends a substitution request for data that would have arrived. Only 3 under-call. The keyword floor splits 21 over and 12 under. Neither is safe and the page publishes both counts rather than one accuracy figure.
⚠︎ ONE MODEL, ONE CORPUS, ONE CADENCE. Every figure here is one tier on a synthetic corpus at an eight-hour cadence with a one-cadence recovery lag. No cross-tier claim is made anywhere, and the cadence table is arithmetic over the same 24 meter-days rather than twenty-four different feeds.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt120jsonl1json2
24 service points, every meter-day inside its settlement window re-read whole
Recorded failureskip one run and 6 intervals are settled on an estimate for good; skip three and 41 are — measured, and no model is called to measure it
flags + memory 100.00% on four bands, 72.50 on the fifth
Recorded failurethe strongest free floor loses only where a SENTENCE decides — 15 of 23 against the model's 22, and its table holds 'failed', not 'failure'
100.00% completeness figure · 0 missed re-polls of 20
80.83% unrecoverable against the free floor's 72.50
$0.00free flag-count floor did four of five bands for
6one skipped run costs intervals, permanently
2026-08-23as of
It produces a completeness position and a re-poll request for a meter-data analyst to action. It never writes an estimate, substitutes an interval value, closes a settlement window or submits anything to a market, and there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is three scalars written by src/completeness.step() from the parsed position and never from the model's answer — the reply is not an argument to it, which is why a wrong reading is a wrong worklist row and never a corrupted history. The clock is the second: at the settlement close the estimation engine substitutes whatever is still absent and the true read is never put back, so a run that did not happen is not a late fix.
⚠︎ THE FREE FLOOR IS THE STORY AND THE PAGE LEADS WITH IT. The completeness figure, the status, the gap count and the raise/hold call are 100.00 pct for the model and 100.00 pct for a floor that costs nothing; counting four letters in a 96-cell grid is what a regex is for. The model's margin is one band — whether a quarantined read is coming back — and it is real: 22 of 23 against 15 of 23 where a single sentence decides, because two sentences break the keyword table completely ('phase imbalance alarm on the site's own switchboard rather than on the metering', and 'failed the day-on-day tolerance check ... queued for manual release').
⚠︎ AND THE CONTROL IS A COLLAPSE, NOT A CONFOUND: with the carried-state sentence replaced and nothing else changed, status falls from 100.00 to 57.50 pct, to 3.77 on the 53 memory-dependent readings, 18 of 20 re-polls are missed, and the SAME model spends 2.15x the output tokens reasoning about a trend it has no way to see. Memory is a discount here, not an overhead.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the settlement rule
src/completeness.py
RULE_TEXT and step(). The rules ship illustrative; a market's real timetable, estimation policy and quality-flag vocabulary all differ. Change them and regenerate the corpus -- the answer key comes from this module.
the cadence and the recovery lag
src/completeness.py
CADENCE_HOURS and RECOVERY_LAG_HOURS. Both are corpus properties AND answer properties: tools/cadence.py shows a 24-hour cadence loses 46 more intervals than the shipped 8-hour one on the same population.
what is sent
src/select.py
SECTION_HINTS and NEVER_SENT. Adding a field to the answer means naming the sections it needs; the fallback subtracts NEVER_SENT unconditionally.
the free floor
evals/baseline.py
METER_SIDE_WORDS and the three MODES. A team with coded exception reasons rather than prose should replace the keyword table with a code lookup and re-run b002 -- that is where the model's whole margin lives.
the corpus
tools/build_corpus.py
SEED, POINTS, RUN_HOURS, CLOSE_OFFSETS and the two reason pools. Keep the seven section headings and the value flag cell shape and everything downstream still parses.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 24 service points x 5 scheduled runs = 120 completeness snapshots from a fixed seed (20260823), and computes data/gold.jsonl from the meter-day MODEL rather than from the rendered page. 756,083 bytes, seven fixed sections each.
the section parser
src/segment.py
Splits a snapshot into its seven named sections. Pure code. All seven are asserted present in all 120 documents before a run may spend.
the section selector
src/select.py
Decides which sections go on the wire. Customer Contact is mapped by nothing and subtracted unconditionally by the fallback, so it never leaves the machine.
the register and clock reader
src/watch.py
One regex over the grid returning (value, flag) per interval. Reads the FLAG and never the value, which is the whole of Rule G-2. Used by the floors, the UI and the pre-flight -- never to answer for the model. Also pulls the run number, the snapshot time, the expected interval count, the settlement close and whether the window has shut, in pure code. The model is never asked to read a timestamp.
the rule
src/completeness.py
Eight illustrative rules and the arithmetic behind them. step() owns the carried position and is the ONLY thing that advances it; the model's reply never reaches it.
the carried state
src/state.py
Three scalars per meter-day -- the previous gap count, whether a re-poll is out, and the previous status -- rendered as one English sentence. Flat in history length.
the prompt
src/prompt.py
Three parts, string-concatenated: the question, the carried-state sentence, the snapshot. --stateless replaces the middle block and changes nothing else, which check_labels asserts byte for byte.
the model call
src/adapters/__init__.py
One completion over urllib. Transient and terminal HTTP failures separated, four bounded retries, the daily call cap checked here so every caller is covered.
the free floors
evals/baseline.py
Three of them, all $0.00, all scored through the identical scorer. b002 is the column the model has to beat and it wins four of the five bands.
the scorer
evals/scoring.py
Exact match per cell, no model. Missed and duplicate re-polls counted apart; the completeness figure derived from the reported gap count rather than asked for.
the cadence experiment
tools/cadence.py
Walks the same 24 meter-days through eight schedules and each scored arm's own raise/hold answers, and reports what is permanently estimated at close. Calls nothing; costs $0.00.
the pre-flight
evals/check_labels.py
Fourteen claims the page makes, checked against the corpus that will be sent, before any run may spend. Includes both directions of the privacy guard and an assertion that the free keyword table does NOT separate the quarantine reasons perfectly.
the local UI
src/app.py
http.server and hand-written HTML/JS on port 9013. Renders with no key; the second button replays what r001 actually answered off the committed result file. HTML/JS lives beside it under ui/.
Where it breaks at scale
⚑ THE CADENCE IS THE SCALING VARIABLE AND IT IS NOT THE CORPUS SIZE. One run of this watch is one call per meter-day still inside its settlement window; the bill is meter-days x runs, and this watch wakes three times a day. A utility with 50,000 service points on a five-day settlement window is 750,000 calls a day, which is a different kind of system from this one and should be: at that scale the register scan is free and only the quarantine attribution is worth a model, so the shape that survives is code for the count and a call only for the exception queue -- which on this corpus is 41 reason lines across 120 readings.
⚠︎ AND THE INPUT DOES NOT SHRINK. Every reading sends the whole 96-cell register because the count is the task, so input is 2,468 tokens a reading and flat. What is NOT flat is output: 1,400 tokens a reading on the shipped arm, 93.1 pct of it provider-side reasoning, because the model counts by narrating. A 15-minute channel costs roughly twice a 30-minute one for the same day.
⚠︎ THE STATE STORE IS ONE FILE. data/state.json is replaced atomically, correct for one writer. Two watches on one head-end is two writers and nothing here has tested it.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
SP-0024-R2, the whole argument on one screen, replayed off the committed r001 result file rather than called live. 22 of this meter-day's 48 intervals read exactly 0.000 with flag A -- a real premises drawing nothing -- so a value-counting reader is wrong before it starts. The gap count has not moved since the 06:00 run, which is stated only in the carried-state panel at the top and appears nowhere on the register. Both columns agree on the status, the count, the completeness figure and the re-poll call; they differ on ONE row. The model says 5 intervals are unrecoverable and the free keyword table says 3, because one reason line reads “battery-backed clock failure ... there is no read to recover” and the table holds “failed”, not “failure”. The key says 5.successOpen full size →The same reading before anything is read. Everything below the buttons -- the carried state, the free floor, the flag counts, the zero-consumption count and the quarantine reasons in full -- is computed locally and needs no key. The model column is the only empty one, and it says “not checked yet” rather than showing a blank.emptyOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
THE SAME PANEL WITH THE MODEL WRONG AND THE FREE FLOOR RIGHT. SP-0006-R2, also replayed off r001. Four intervals are quarantined because “the mapping to this channel was wrong at the time of measurement and the raw pulses were discarded at the collector” -- the pulses are gone, so nothing is recoverable. The model reads “collector” and “mapping” as a platform problem and answers 0 unrecoverable; its own rationale, printed on the page, says so. The keyword table happens to hold “discarded” and answers 4, which is right. This is the direction that costs: four reads somebody would wait for, and they were never going to arrive.failureOpen full size →The read button pressed with no API_KEY configured. It returns 200 with a sentence rather than an error, and the page still carries every locally-computed panel. A forker's first ten minutes after a clone look like this, which is why the second button exists.failureOpen full size →
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
120interval reads
0.72 MiBjson 2 · jsonl 1 · txt 120
p50 6,414chars per bytes per snapshot
$0.00setup · 0s
How it is cutWhat one bytes per snapshot is
no split. Every one of the 120 readings is scored, and the same 120 are scored for the stateless control and all three free floors.
SetupWhat the setup figure measured
There is no index. Every meter-day still inside its settlement window is re-read WHOLE on every scheduled run -- that is what a monitor is -- so there is nothing to embed, nothing to chunk and no retrieval step. What makes a reading a CHANGE is the carried state, which is three integers and costs nothing to compute.
LicenceLicence
MIT, same as the repository. There is no third-party data in this kit.
Bring your ownBring your own interval reads
Point tools/build_corpus.py at your own head-end, or replace data/corpus/*.txt and data/gold.jsonl directly. Keep the seven section headings (src/segment.py::SECTIONS -- the parser asserts all seven in every document before a run may spend) and keep the register's value flag cell shape, and the scanner, the floors, the UI and the scorer all keep working. Your gold rows need the same fields data/gold.jsonl carries; evals/check_labels.py will tell you which are missing before you spend anything.
⚠︎ And what stops being true when you do: The measured figures on this page do not travel with your corpus, and four of them are corpus properties rather than model properties. The 100.00 pct on the completeness figure is a property of a register with an unambiguous flag column. The 80.83 pct on unrecoverable is a property of THESE sixteen quarantine sentences, written to collide in both directions. The whole cadence table is a property of an eight-hour watch against a one-cadence recovery lag and the settlement closes this generator chose. And the 969 zero-consumption intervals are why the zero trap scores what it scores. Re-measure all four before quoting any of them.
What breaks it
⚠︎ A CORPUS WHOSE QUARANTINE REASONS ARE SEPARABLE BY KEYWORD MEASURES THE TEMPLATES, NOT THE WORK -- FOUND BEFORE ANY CALL WAS MADE AND IT COST A REWRITE. The first pool had transient reasons that all named a queue or the head-end and meter-side reasons that all named the meter, and a fourteen-word keyword table separated all sixteen perfectly. That is a fact about templates. The pool now carries transient entries full of fault vocabulary ("a tamper flag was raised on this installation earlier in the day and cleared by field services") and meter-side entries full of platform vocabulary ("the collection queue holds no entry for this interval and never will"), and evals/check_labels.py asserts the table misattributes at least 5 readings so the band can never silently become a template match again. Measured: 33 of 120.
A HEAD-END THAT DOES NOT ZERO-FILL. This corpus prints an absent interval as 0.000 M, which is what makes the zero trap a trap. A feed that prints an empty cell, or omits the row entirely, changes what the register scanner has to do and makes the b000 floor look better than it is.
CODED EXCEPTION REASONS. The model's entire margin on this kit is reading free prose. A head-end whose exception queue carries a closed list of reason codes takes that margin away completely and should use the b002 floor with a code lookup instead of a keyword table.
A REVISION CYCLE. The register FREEZES at the settlement close in this corpus, because the estimation pass takes ownership of the meter-day. A market that reopens for a revision cycle behaves differently and this kit does not model one -- every permanent-loss figure on this page assumes the close is final.
A POPULATION THAT CHANGES BETWEEN RUNS. All 24 meter-days are present at all 5 runs. A real watch picks up new flow dates every day and drops settled ones, and nothing here tests a meter-day appearing mid-window or leaving the watch.
A LONGER WINDOW. Five scheduled runs is long enough to hide a trend behind a run boundary and short enough that no carried count has been advanced more than four times. A real settlement window is days; half of this corpus's own closes fall after the fifth run, which is why tools/cadence.py walks the whole arc and the eval does not.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
195
not measured
the question and the JSON shape
3,494
not measured
Carried state — the experiment
147
not measured
Synthetic Record
249
not measured
Service Point
372
not measured
Settlement Terms
2,704
not measured
Completeness Position
283
not measured
Interval Register
1,526
not measured
Operational Notes
87
not measured
Total
2,468
This is the cost lesson as arithmetic: of the 9,057 characters assembled, 5,221 are documents — 58% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
TWO MESSAGES, JOINED IN SEND ORDER. The provider receives a system message and a user message; prompt_verbatim is the system message, a blank line, then the user message exactly as src/prompt.build assembled it for SP-0024-R2 on run r001-interval-gap. Nothing is elided and nothing is re-ordered. The register is 48 lines here because SP-0024 is a 30-minute channel; a 15-minute channel sends 24 rows of four cells instead of two.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled interval-data completeness watch. You apply a written settlement rule to one service point's meter-day at one scheduled run. You answer with one JSON object and no other text.
You are the scheduled interval-data completeness watch for one utility's meter data operations. It
wakes every eight hours, re-reads every meter-day still inside its settlement window, and reports
each one. You are reading ONE service point's meter-day at ONE scheduled run.
The settlement rules, the expected interval count and the settlement close are reproduced in the
snapshot below. Apply them exactly as written. You cannot see the earlier scheduled runs; what is
known about them is stated under "Carried state" and is the only history available to you. Do not
assume anything about earlier runs beyond it.
How to count:
- The Interval Register lists every expected interval with a VALUE and a one-letter QUALITY FLAG.
THE FLAG DECIDES AND THE VALUE NEVER DOES (Rule G-2). An interval reading "0.000 A" is a real
validated read of zero consumption and is NOT missing. An interval reading "0.000 M" was never
received and IS missing. They differ by one letter.
- "missing_intervals" is the number of expected intervals flagged M or Q (Rule G-3). Count them.
Flags A and E are not missing.
- Under the register, each quarantine reason names the interval it covers, or the first and last
interval of the run it covers ("05:15-05:45" is three intervals on a 15-minute channel, two on
a 30-minute one). Every Q on the grid is covered by exactly one of those lines.
- "unrecoverable" is how many of those missing intervals can no longer be recovered:
* while the settlement window is OPEN, it is the number of Q intervals whose quarantine reason
in the register is a METER-SIDE fault -- the measurement element, the register, the
installation, the device clock, a tamper or alarm condition, a decommissioned or swapped
device. Those never re-validate (Rule G-4). A Q whose reason is a TRANSIENT condition in the
collection or validation chain -- a queue, a re-sync, a pending deduplication, a manual
release, a profile correction -- is recoverable and does NOT count here.
* once the window has CLOSED, every missing interval is unrecoverable, so "unrecoverable"
equals "missing_intervals" (Rule G-5).
- If the expected interval count or the settlement close is not on file, the meter-day is
CONTEXT_INCOMPLETE: status CONTEXT_INCOMPLETE, missing_intervals 0, unrecoverable 0,
raise_chase NO. Never assume a default interval count (Rule G-7).
Answer with a single JSON object and nothing else:
{"status": "COMPLETE|BASELINE|HEALING|STALLED|SETTLED|CONTEXT_INCOMPLETE",
"missing_intervals": <integer>,
"unrecoverable": <integer, not greater than missing_intervals>,
"raise_chase": "YES|NO",
"rationale": "one sentence, naming the rule you applied and the counts you reached"}
Precedence for "status", applied in this order: CONTEXT_INCOMPLETE beats everything; then SETTLED if
the settlement window has closed, whether or not anything is missing; then COMPLETE if nothing is
missing; then BASELINE if this is the first scheduled run of this meter-day, because there is no
earlier count to compare against; then HEALING if fewer intervals are missing now than the carried state
says were missing at the previous run; otherwise STALLED.
"raise_chase" is YES only on the run that raises the re-poll -- the status is STALLED and the
carried state does not already say a re-poll was raised. On every later run it is NO, and it is
never YES on a first run, a healing run or a settled one (Rule G-6, Rule G-8).
Carried state
----------------------------------------------------------------
At the previous scheduled run this meter-day had 10 intervals missing. No re-poll has been raised for this meter-day yet. It was reported BASELINE.
Completeness snapshot
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated interval-data completeness snapshot
for an AI use-case kit; it reproduces no utility's head-end extract, no market's
settlement timetable and no real premises. The settlement rules are ILLUSTRATIVE.
Service Point
----------------------------------------------------------------
Service point : SP-0024
Meter : MTR-601546
Channel : E1 -- kWh delivered
Interval length : 30 minutes
Intervals expected : 48
Flow date (meter-day read) : 2026-03-01
Network region : CENTRAL-02
Market segment : small commercial, profile class 4
Settlement Terms
----------------------------------------------------------------
Settlement close : 2026-03-02 20:00 local
Re-poll recovery lag on file : 8.00 hours
Estimation policy : any interval still absent when the window shuts is
substituted by the estimation engine and the true read
is never put back
Rule G-1 A meter-day is COMPLETE when every EXPECTED interval on the channel carries a validated
actual read. The expected count is on the service point's own record (96 at 15 minutes,
48 at 30 minutes) and is never assumed.
Rule G-2 THE FLAG DECIDES, NOT THE VALUE. An interval reported 0.000 with flag A is a real
reading -- a vacant premises, a disconnected service, on-site generation offsetting
load -- and is NOT a gap. A head-end zero-fills an interval it never received and writes
the same 0.000 with flag M. The two are one letter apart on the page.
Rule G-3 A GAP is any expected interval flagged M (no read received) or Q (received, failed
validation, quarantined). Both count against completeness. They differ in what closes
them, not in whether they count.
Rule G-4 A Q interval is RECOVERABLE where its quarantine reason is a TRANSIENT condition in the
collection or validation chain -- it re-validates once the condition clears. A Q whose
reason is a METER-SIDE fault -- the measurement element, the register, the installation,
the device clock -- will never re-validate and is UNRECOVERABLE from this run onward: it
must be substituted now rather than chased.
Rule G-5 An M interval is recoverable while the settlement window is open and unrecoverable once
it has closed. At close, EVERY interval still missing becomes an estimate and the true
reading is never substituted back in.
Rule G-6 A re-poll is raised ONCE per meter-day, on the first scheduled run at which the gap count
has NOT fallen since the previous run. The first run of a meter-day establishes the
baseline and raises nothing -- with no earlier count there is no trend, and chasing every
gap on sight is the alarm-fatigue failure this watch exists to avoid.
Rule G-7 The expected interval count and the settlement close are OPERATOR-SUPPLIED, read from the
service point record. Where either is not on file the meter-day is CONTEXT_INCOMPLETE:
nothing is counted, nothing is chased, and no default interval count may be assumed.
Rule G-8 Once the settlement window has closed the meter-day is SETTLED. Its missing intervals are
final, all of them unrecoverable, and no re-poll is raised against it.
Completeness Position
----------------------------------------------------------------
Snapshot taken : 2026-03-02 14:00 (scheduled run 2 of this meter-day)
Watch cadence : every 8 hours (06:00 / 14:00 / 22:00 local)
Previous scheduled run : 2026-03-02 06:00
Settlement window : IN WINDOW -- 6.00 hours until close
Interval Register
----------------------------------------------------------------
Every expected interval on this channel, with its value in kWh and its quality flag.
Flags: A actual, telemetered and validated M no read received Q quarantined, reason below E substituted at a prior revision
Hour :00 :30
00 0.441 A 0.000 M
01 0.366 A 0.432 A
02 0.411 A 0.000 A
03 0.000 A 0.000 A
04 0.000 A 0.000 A
05 0.000 A 0.000 A
06 0.000 A 0.000 A
07 0.000 A 0.000 A
08 0.000 A 0.000 A
09 0.000 A 0.000 A
10 0.000 A 0.000 A
11 0.000 Q 0.000 Q
12 0.000 A 0.000 A
13 0.000 A 0.000 A
14 0.000 Q 0.000 Q
15 0.000 Q 0.000 A
16 0.837 A 0.809 A
17 0.916 A 0.813 A
18 1.042 A 1.039 Q
19 1.145 A 1.461 A
20 1.452 A 0.892 A
21 0.000 M 0.000 M
22 0.000 M 0.475 A
23 0.303 A 0.271 A
Quarantine reasons raised against this meter-day
11:00-11:30 Q battery-backed clock failure means the device cannot attribute anything to this timestamp; there is no read to recover
14:00-15:00 Q the head-end has been accepting reads from a decommissioned register since the exchange; nothing measured this interval
18:30 Q the device's comms module re-registered on the mesh mid-interval; the read is buffered on the meter and comes back on the next poll
Operational Notes
----------------------------------------------------------------
Nothing outstanding on this service point's exception queue at the last shift handover.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"status": "STALLED",
"missing_intervals": 10,
"unrecoverable": 5,
"raise_chase": "YES",
"rationale": "Applying G-3 and G-4, I count 10 missing intervals (4 M and 6 Q) with 5 Q intervals unrecoverable due to meter-side clock/register faults, and because the 10 missing equals the carried 10 with no re-poll already raised, status is STALLED and chase is YES."
}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a utility's missing meter readings in time — 120 interval reads. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the meter-day model -- the gap blocks, their arrival times, their attribution and the settlement close -- and never from the rendered page. evals/check_labels.py re-derives the whole key from src/completeness.step before any run may spend, and separately asserts that the RENDERED register agrees with the key on all 120 documents, so a rendering that disagreed with the model behind it would refuse the run rather than quietly score against the wrong truth. The two integer fields are compared as integers; the status and the raise/hold call are compared as strings from a closed list; the completeness figure is DERIVED from the reported gap count rather than asked for, so a right count with a slipped decimal cannot look like a wrong count.
120interval reads
120source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED110 · 108 / 110interval gap completeness pct — readings whose channel record carries an interval count, completeness derived from the reported gap count, exact matchDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED120 · 69 / 120status accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED120 · 118 / 120missing intervals accuracy pct — readings, exact integer match, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED97 · 98 / 120unrecoverable accuracy pct — readings, exact integer match, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED120 · 101 / 120raise chase accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 18 / 20missed chase pct — readings whose correct answer is to RAISE the re-poll, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 100 / 100duplicate chase rate pct — readings that must NOT raise a re-poll, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED53 · 2 / 53memory status accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED53 · 34 / 53memory chase accuracy pct — readings whose answer is NOT derivable from their own snapshot, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED45 / 45zero trap accuracy pct — readings carrying at least one interval read 0.000 with flag A, gap count exactDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED20 / 20context incomplete recall pct — readings with no interval count or no settlement close on file, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED120 / 120answered pct — readings, replies that parsed, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is asserted: evals/check_labels.py runs fourteen checks before any run may spend, and all fourteen pass. They include the answer key replaying from src/completeness.step (0 mismatches), the rendered register agreeing with the key (0 disagreements), every Q on the grid being covered by exactly one reason line (0 mismatches), the privacy guard in BOTH directions (0 of 120 leak with the guard, 120 of 120 without it under a reproduced schema change), the stateful and stateless prompts differing on exactly one line, no snapshot stating an earlier gap count or status or outstanding re-poll, and two negative controls: counting zeros instead of flags must get a different answer (it does, on 78 of 120) and the free keyword table must NOT separate the quarantine reasons perfectly (it misattributes 33 of 120).
1,399.92output tokens · the fast tier, with the carried state · 8,411 ms p50
3,005.65output tokens · the same tier, memory removed (THE CONTROL) · 13,220 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 1.6× as long, and lands one row apart on 120. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One interval read
1,000 interval reads
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.005434
$5.43
23%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.002174
$2.17
23%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.094677
$94.68
26%
Same work, 44× the bill
The same interval reads, the same tokens — only the rate card changed. And across all 3 cards between 23% and 26% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE SPLIT. The only lever on this page that changes the bill by an order of magnitude is not a model setting -- it is sending the model less. Four of the five bands are already 100.00 pct free, so a deployment that counted the register in code and called the model ONLY for the quarantine reasons would pay for 41 reason lines across 120 readings instead of 120 full registers. That is a design this kit measures the case for and does not ship, because shipping it would remove the comparison that makes the case.
Rates checked 2026-08-18. The provider that actually ran all 255 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors, the fourteen pre-flight checks or the whole cadence and per-arm loss table. The only money on this kit is the readings themselves.
The gradersThree ways to grade
⚑ THE STRONGEST FREE FLOOR TIES THE MODEL ON FOUR OF THE FIVE BANDS AND THE KIT LEADS WITH THAT. b002 -- carried state, the register counted on the flag column, attribution by keyword table -- scores 100.00 on the completeness figure, the status, the gap count and the raise/hold call, exactly what the model scores, for $0.00. The model's whole margin is unrecoverable: 80.83 against 72.50.
⚑ AND ON THE 23 READINGS WHERE ONE SENTENCE DECIDES THE ANSWER, THE MARGIN IS 22 AGAINST 15. Restricting to open-window readings that carry exactly one quarantine reason line makes the attribution unambiguous per sentence. Two sentences break the table completely: "phase imbalance alarm on the site's own switchboard rather than on the metering" (0 of 4 -- the table holds "phase" and "alarm" and the sentence says explicitly that it is not on the metering) and "failed the day-on-day tolerance check ... queued for manual release" (0 of 4 -- the table holds "failed"). The model reads both correctly. The model's own single miss in that subset is "the meter's clock had drifted against the collector", which it confuses with the battery-clock failure sentence.
⚑ WHAT A MONITOR BUYS IS MEMORY, AND THE MODEL IS WHAT YOU ADD ON TOP. The largest movement anywhere on this kit is b000 -> b001: the identical zero-counting rule, given three carried scalars, drops duplicate re-polls from 61 in 100 quiet readings to 1, and raise/hold accuracy goes 49.17 -> 90.83. There is no model in that at all.
⚠︎ ONE ASYMMETRY FLATTERS b002 AND IT IS THE SAME ONE EVERY KIT IN THIS SERIES HAS: the floor's counting half and the answer key are two expressions of ONE implementation, so the floor cannot misread the counting rule -- it IS the rule -- while the model has only the prose. The attribution half is NOT shared: the keyword table was written from utility vocabulary and deliberately never tuned against the corpus, which is why it loses 8 of 23 on the sentences.
⚑ AND THE WEAKEST FLOOR RECOVERS THE MOST DATA. b000 has no memory and chases every gap on every run. Because it therefore raises the re-poll at the FIRST run instead of waiting for the trend, tools/cadence.py measures it losing 49 intervals at the settlement close against the correct rule's 83. That is the uncomfortable finding on this kit and it is published rather than buried: Rule G-6 is the alarm-minimising rule, not the loss-minimising one.
the fast tier, with the carried state 100.0% interval gap completeness · the same tier, memory removed (THE CONTROL) 98.2% interval gap completeness · 4 more measured on each run
The two re-poll directions, counted apart A MISSED re-poll is a reading whose correct answer was to raise one and did not -- data that the settlement close then substitutes permanently. A DUPLICATE re-poll is one raised against a meter-day that already has one outstanding -- an on-demand read queued at the head-end for data already on its way. Different costs, different owners, never averaged.
$0.00
no
yes
the fast tier, with the carried state 0.0% missed chase · the same tier, memory removed (THE CONTROL) 90.0% missed chase · 1 more measured on each run
What each arm's own raise/hold answers cost in permanently estimated intervals Reads each arm's committed result file, takes the FIRST run at which it said YES as that meter-day's re-poll time, holds the recovery lag and the settlement close exactly as the key has them, and counts what is still absent when the window shuts. It is the only place on this kit where a percentage of readings becomes the thing the reader actually loses.
$0.00
no
yes
no headline metric on any of its 4 runs — they record intervals estimated
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set CAN tell the arms apart, and the evidence is that it does: five arms scored through the identical scorer span 20.00 to 100.00 on the discriminator, 40.83 to 100.00 on status, and 48.33 to 80.83 on unrecoverable. It also separates them in the right PLACES -- the two arms without a flag-aware count are the two that fail the zero trap at 0.00 pct, the two arms without memory are the two that fail the memory-dependent status slice, and the one band where the shipped arm beats the strongest floor is the one band that turns on prose. A corpus where every arm scored alike would be measuring nothing.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
You only need to know how many intervals are missing on each meter-day
a regex over the quality flags, or the b002 floor in this repo
Counting M and Q in a 96-cell grid is exact and free. On this corpus b002 gets 100.00 pct of the gap counts and 100.00 pct of the completeness figures for nothing at all, which is exactly what the model gets.
AVOID this kit entirely for that job. Paying a model per meter-day per run to count four letters is the most expensive way to get an answer a regex already has.
Your head-end's exception queue carries coded reasons from a closed list
the b002 floor with a code lookup instead of its keyword table
The model's entire margin on this corpus is reading free prose -- 80.83 against 72.50 on unrecoverable, 22 of 23 against 15 of 23 where one sentence decides. Take the prose away and there is nothing left to buy.
AVOID assuming the measured gap transfers. It was measured against a keyword table over sixteen sentences written to collide, not against your reason codes.
Your exception queue is prose a validation analyst types, and it names the meter whether or not the meter is the problem
this kit, with the carried state
That is exactly the corpus these figures were measured on, and it is where the whole margin comes from. It is also where the model still fails: 20 of its 23 misses over-call a read as dead.
AVOID running it without the carried state. The control shows what that costs: 18 of 20 re-polls missed, status accuracy 3.77 pct on the memory-dependent readings, 51 more intervals permanently estimated, and 2.15x the output tokens.
You want the fewest intervals lost and you can absorb the noise
the b000 floor -- chase every gap on every run
Measured, and it is the uncomfortable answer: chasing everything loses 49 intervals at close against the correct rule's 83, because the re-poll lands earlier. Rule G-6 minimises alarms, not loss.
AVOID it if your analysts already ignore the exception queue. It raises 61 duplicate re-polls in 100 quiet readings, and an alert nobody reads recovers nothing.
You want the watch to actually wake up on a schedule
your own scheduler -- cron, Airflow, Temporal, whatever already runs
evals/run.py is INVOKED, not woken. Nothing in this kit detects a missed run, back-fills it or marks its readings late -- and on this task that absence is priced: one skipped run costs 6 intervals permanently, three cost 41.
AVOID reading this kit's figures as a claim about a running deployment. They are measured over five runs that all happened.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
FAULT_NOUN_OVERRIDES_THE_CLAUSE_THAT_CANCELS_IT
a quarantine reason names a fault and then says it was cleared
11
"a tamper flag was raised on this installation earlier in the day and cleared by field services; the interval is held while the flags are re-applied". The read exists and will validate, so the interval is RECOVERABLE. On the 11 readings that carry this…
TWO_CLOCK_SENTENCES
a drifted clock and a failed clock read the same at a glance
9
"the meter's clock had drifted against the collector and the read is held pending a time re-sync" is RECOVERABLE; "battery-backed clock failure means the device cannot attribute anything to this timestamp; there is no read to recover" is not. The model gets…
PLATFORM_WORDS_ON_A_DEAD_READ
the sentence names the collector, and the read is gone anyway
4
SP-0006-R2: "the mapping to this channel was wrong at the time of measurement and the raw pulses were discarded at the collector", four intervals. The pulses are gone, so nothing is recoverable and the key says 4. The model answers 0 and its own rationale…
RULE_RAISES_A_REPOLL_THAT_CANNOT_LAND
not a model error at all -- the shipped rule's own limitation
3
SP-0024-R2 among them: the re-poll is raised at 14:00, the recovery lag on file is 8 hours and the settlement window shuts at 20:00. Rule G-6 raises it because the gap count has stopped falling, and it cannot possibly land in time. The correct move is to…
What we could NOT verify
WHETHER THE INSTRUCTION-SHAPED OPERATIONAL NOTE MOVES THE ANSWER. It reaches the model verbatim on a third of the corpus and the raise/hold call was 100.00 pct correct anyway, which is consistent with the note having no effect and is not a measurement of it. The experiment -- score the same 120 readings with the note removed -- costs one more 120-call run and was not paid for.
WHETHER THE CARRIED STATE READS BETTER AS ENGLISH OR AS JSON. src/state.describe renders three scalars as a sentence because that is the form Rule G-6 is written about. It is a design choice and it is unmeasured; the obvious experiment is another 120-call run.
WHAT THE STATELESS ARM WOULD HAVE DONE ON A SIXTH RUN. The per-arm permanent-loss table treats an arm that never raised a re-poll inside the five shipped runs as never having chased. Half the settlement closes in this corpus fall after run 5, so for those meter-days the control's 134 is an upper bound rather than a measurement.
A SECOND MODEL. One tier was run. Every figure on this page is one provider's fast tier and no cross-tier claim is made anywhere.
WHAT DISABLING PROVIDER-SIDE REASONING DOES. 93.1 pct of this run's output tokens were reasoning left at the provider's default. src/adapters accepts a thinking field and this kit's harness never sent one, so the lever exists and was never pulled.
WHETHER THE ATTRIBUTION OF THE 23 UNRECOVERABLE MISSES IS EXACT. It is exact on the 23 readings carrying exactly one quarantine reason line, and heuristic on the rest -- the model answers a TOTAL, not a verdict per line, so on a reading with three reasons there is no way to say which one it got wrong. The taxonomy counts say "likely driver" and mean it.
CONCURRENCY. data/state.json is replaced atomically, which is correct for one writer. Two watches over one head-end is two writers and nothing here has tested it.
A CADENCE OTHER THAN EIGHT HOURS, AS AN ACCURACY CLAIM. tools/cadence.py measures what other schedules LOSE, which is arithmetic. It does not re-run the model on them, so no accuracy figure on this page is known at any other cadence.
PROVIDER-SIDE RETENTION. The prompt carries no customer identifiers by construction, but it does carry a full day's load curve, and what a provider keeps of a request is outside this repository and nobody here has verified it.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
2,468.12
1,399.92
8,411 ms
$0.005434
$0.002174
$0.094677
the same tier, memory removed (THE CONTROL)
2,444.46
3,005.65
13,220 ms
$0.010239
$0.004096
$0.174727
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same reading (one meter-day, at one scheduled run), the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
255 live calls were made for this kit and 255 returned: 10 for c000 (an 8,000-token ceiling probe on the two hardest meter-day chains, of which 3 replies were cut off and returned nothing -- which is exactly what a calibration run is for), 5 for c001 (the same chain at 24,000, all 5 correct), 120 scored, and 120 for the stateless control. NOTHING WAS DISCARDED BY CHOICE and no run was re-fired. The corpus defect this kit found -- a quarantine pool a keyword table separated perfectly -- was caught by writing the free floor and reading its score BEFORE any call was made, which is the whole argument for running your floors first. Every screenshot on this page is free: the shoot script blanks API_KEY and the two answered frames replay the committed r001 file. Everything else -- all three floors, the wiring stub, the scorer, the pre-flight and the entire cadence and permanent-loss table -- is pure code and costs $0.00. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, and almost all of them are reasoning. 93.1 pct of this run's output (156,477 of 167,990) was provider-side reasoning left at the default, and output is 77.3 pct of the projected bill on the shared card. The model is counting a grid by narrating it.
THE REGISTER, which is the input floor and does not move. Every reading sends all 96 (or 48) cells because counting them is the task -- 2,468 input tokens, flat. A 15-minute channel costs about twice a 30-minute one for the same meter-day.
THE CADENCE, which multiplies everything. Three runs a day over a five-day settlement window is fifteen readings per meter-day, not one.
THE CARRIED STATE, which is a DISCOUNT and not a cost. It is one sentence -- 24 tokens a reading more on input -- and it halves the output, because a model that can see the trend stops reasoning about what the trend might be.
Your volumeWhat it costs at your volume
LINEAR IN METER-DAYS x RUNS x WINDOW LENGTH, AND THAT IS THE WHOLE WARNING. Ten times the service points is ten times the bill; nothing here batches, caches or shares context between readings, and nothing should -- each reading is one meter-day at one instant. What does NOT scale linearly is the value: the cadence table shows the recovered intervals falling off sharply as runs are removed, so halving the bill by halving the cadence costs 14 permanently estimated intervals on this population and cutting to daily costs 46.
Where pricing changes shape
The output ceiling. MAX_TOKENS = 24000 is set from calibration. c000 fired the two hardest meter-day chains at an 8,000-token cap and 3 of the 10 replies were cut off and returned nothing at all -- billed in full, worth nothing. c001 re-fired the harder chain at 24,000 and topped out at 8,867. On the scored run the largest was 12,513. A ceiling below about 16,000 on this corpus loses readings.
The interval length. A 15-minute channel is 96 cells and a 30-minute one is 48. The input roughly doubles and the reasoning more than doubles, because the count is longer. Half this corpus is each; a utility on 5-minute intervals would be reading 288 cells a call and should not be using this shape at all.
Provider-side reasoning defaults. 93.1 pct of the output is reasoning nobody asked for. A provider that changes its default reprices this kit without anything a reader can see changing.
Your return, with your numbers
Volumemeter-days still inside their settlement window, per scheduled run -- this run judged 120 (24 meter-days x 5 runs) per arm, on an eight-hourly watch
What it replacessomebody working the head-end's completeness report at shift change: opening each meter-day still inside its window, comparing the gap count against yesterday's export, reading the exception queue to decide which quarantined reads are coming back, and remembering which service points already have an on-demand read outstanding
Time saved per itemnot measured here -- depends on how long an analyst takes to compare two exports and read an exception queue. What IS measured is the other side of the ledger: 60 of the 109 recoverable intervals in this population are saved by running the watch at all, and 34 more would be saved by chasing everything instead of waiting for the trend.
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared connection was already pointed at, and the kit's whole claim is that swapping it is .env plus one more run. No cross-tier comparison is made anywhere on this page because only one tier was run, and the projection table below is arithmetic on this run's own token counts rather than a second measurement.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
296,175input tokens · this run
167,990output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 120 readings, one completion call each, one tier. The 120-call stateless control and the 15 calibration calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here. The whole cadence and permanent-loss table cost nothing and is not in any of these figures.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.261
$0.261
$2.17
2026-09-12
gemini-3-flash
Google
$0.652
$0.652
$5.43
2026-09-18
gemini-3-8-flash
Google
$0.852
$0.852
$7.10
2026-09-18
llama-5
Meta
$1.084
$1.084
$9.03
2026-09-18
claude-haiku-4-5
Anthropic
$1.136
$1.136
$9.47
2026-09-12
grok-4-5
xAI
$1.600
$1.600
$13.34
2026-09-18
grok-4-6
xAI
$1.600
$1.600
$13.34
2026-09-18
claude-sonnet-5
Anthropic
$2.272
$2.272
$18.94
2026-09-12
gemini-3-1-pro
Google
$2.608
$2.608
$21.74
2026-09-18
gpt-5-6-terra
OpenAI
$2.608
$2.608
$21.74
2026-09-12
gpt-5-6-sol
OpenAI
$4.545
$4.545
$37.87
2026-09-12
claude-opus-4-8
Anthropic
$5.681
$5.681
$47.34
2026-09-12
claude-opus-5
Anthropic
$5.681
$5.681
$47.34
2026-09-12
claude-fable-5
Anthropic
$11.361
$11.361
$94.68
2026-09-18
claude-fable-5-1
Anthropic
$11.361
$11.361
$94.68
2026-09-18
gpt-6-astra
OpenAI
$11.361
$11.361
$94.68
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of date with nothing checking it against the vendor since.
⚑ 93.1 PCT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (156,477 of 167,990), left at the provider's default, so every row below prices a reasoning-on workload. Output is 77.3 pct of the projected bill on the shared card, which means most of what these rows charge for is the model counting a grid by narrating it. A vendor whose default differs, or a caller who turns it off, would see a materially different figure, and this kit has not measured what turning it off does to the answers.
⚑ EVERY ROW PRICES ONE READING, AND A DEPLOYMENT DOES NOT BUY ONE READING. This watch wakes 3 times a day and re-reads every meter-day still inside its settlement window each time, so the bill is meter-days x runs x the length of the window in days. Multiply any row below by your own open-meter-day count and by 3 before comparing it with anything.
⚑ AND EVERY ROW PRICES THE ARM WITH THE CARRIED STATE, WHICH IS THE CHEAPER ONE. The stateless control cost 1.88 times as much per reading on the same card, because removing the memory made the model reason for longer -- 360,678 output tokens against 167,990 on the same 120 readings. Memory is not an overhead here; it is a discount.
⚠︎ AND NONE OF THESE ROWS IS THE CHEAPEST OPTION FOR FOUR OF THE FIVE BANDS. The b002 floor scores exactly what the model scores on the completeness figure, the status, the gap count and the raise/hold call, for $0.00. Every row above is the price of a 8.33-point margin on one field.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
13 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 24 service points x 5 scheduled runs = 120 completeness snapshots from a fixed seed (20260823), and computes data/gold.jsonl from the meter-day MODEL rather than from the rendered page. 756,083 bytes, seven fixed sections each.
You change it to: SEED, POINTS, RUN_HOURS, CLOSE_OFFSETS and the two reason pools. Keep the seven section headings and the value flag cell shape and everything downstream still parses.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
POINTS = 24
FLOW_DATE = "2026-03-01"
BASE_DAY = 2
RUN_HOURS = (6.0, 14.0, 22.0, 30.0, 38.0)
src/segment.pythe section parser
Splits a snapshot into its seven named sections. Pure code. All seven are asserted present in all 120 documents before a run may spend.
src/segment.py
# Split a completeness snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Service Point", "Settlement Terms", "Completeness Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector — a swap seam
Decides which sections go on the wire. Customer Contact is mapped by nothing and subtracted unconditionally by the fallback, so it never leaves the machine.
You change it to: SECTION_HINTS and NEVER_SENT. Adding a field to the answer means naming the sections it needs; the fallback subtracts NEVER_SENT unconditionally.
src/select.py
# Pick which sections of a snapshot are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
POINT = "Service Point"
TERMS = "Settlement Terms"
POSITION = "Completeness Position"
REGISTER = "Interval Register"
CUSTOMER = "Customer Contact"
NOTES = "Operational Notes"
NEVER_SENT = (CUSTOMER,)
SECTION_HINTS = {
src/watch.pythe register and clock reader
One regex over the grid returning (value, flag) per interval. Reads the FLAG and never the value, which is the whole of Rule G-2. Used by the floors, the UI and the pre-flight -- never to answer for the model. Also pulls the run number, the snapshot time, the expected interval count, the settlement close and whether the window has shut, in pure code. The model is never asked to read a timestamp.
src/watch.py
# One meter-day, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled interval-data completeness watch. You apply a written settlement "
MAX_TOKENS = 24000
FIELDS = ("status", "missing_intervals", "unrecoverable", "raise_chase")
def documents():
def points():
def load_doc(doc_id):
CELL = re.compile(r"(\d+\.\d{3})\s+([AMQE])\b")
src/completeness.pythe rule — a swap seam
Eight illustrative rules and the arithmetic behind them. step() owns the carried position and is the ONLY thing that advances it; the model's reply never reaches it.
You change it to: CADENCE_HOURS and RECOVERY_LAG_HOURS. Both are corpus properties AND answer properties: tools/cadence.py shows a 24-hour cadence loses 46 more intervals than the shipped 8-hour one on the same population.
src/completeness.py
# The completeness rule as arithmetic. Pure code, no model, no dependency.
COMPLETE = "COMPLETE"
BASELINE = "BASELINE"
HEALING = "HEALING"
STALLED = "STALLED"
SETTLED = "SETTLED"
CONTEXT_INCOMPLETE = "CONTEXT_INCOMPLETE"
STATUSES = (COMPLETE, BASELINE, HEALING, STALLED, SETTLED, CONTEXT_INCOMPLETE)
YES = "YES"
NO = "NO"
src/state.pythe carried state
Three scalars per meter-day -- the previous gap count, whether a re-poll is out, and the previous status -- rendered as one English sentence. Flat in history length.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot counter.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_point(store, point_id):
def describe(state):
src/prompt.pythe prompt
Three parts, string-concatenated: the question, the carried-state sentence, the snapshot. --stateless replaces the middle block and changes nothing else, which check_labels asserts byte for byte.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model call — a swap seam
One completion over urllib. Transient and terminal HTTP failures separated, four bounded retries, the daily call cap checked here so every caller is covered.
You change it to: PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
evals/baseline.pythe free floors — a swap seam
Three of them, all $0.00, all scored through the identical scorer. b002 is the column the model has to beat and it wins four of the five bands.
You change it to: METER_SIDE_WORDS and the three MODES. A team with coded exception reasons rather than prose should replace the keyword table with a code lookup and re-run b002 -- that is where the model's whole margin lives.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("zerocount", "zerocount-mem", "flagcount-mem")
ASSUMED_EXPECTED = 96
METER_SIDE_WORDS = ("tamper", "phase", "measurement element", "ct ratio", "decommissioned",
def _num(pat, text, cast=float):
def _is_meter_side(reason):
def _zero_count(text):
def review(text, carried=None, mode="flagcount-mem"):
evals/scoring.pythe scorer
Exact match per cell, no model. Missed and duplicate re-polls counted apart; the completeness figure derived from the reported gap count rather than asked for.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("status", "missing_intervals", "unrecoverable", "raise_chase")
INT_FIELDS = ("missing_intervals", "unrecoverable")
def _pct(n, d):
def _int(v):
def _completeness(expected, missing):
def score(records, golds):
tools/cadence.pythe cadence experiment
Walks the same 24 meter-days through eight schedules and each scored arm's own raise/hold answers, and reports what is permanently estimated at close. Calls nothing; costs $0.00.
tools/cadence.py
# What a missed run costs, in intervals that can never be read again. FREE -- no model, no key.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
OUT = os.path.join(HERE, "data", "cadence.json")
HORIZON_HOURS = 96.0
def runs_every(period, start=6.0, skip=()):
SCHEDULES = [
def points():
def evaluate(ps, run_hours):
ARMS = [
def arm_losses(ps):
evals/check_labels.pythe pre-flight
Fourteen claims the page makes, checked against the corpus that will be sent, before any run may spend. Includes both directions of the privacy guard and an assertion that the free keyword table does NOT separate the quarantine reasons perfectly.
evals/check_labels.py
# Everything that must be true BEFORE a run is allowed to spend. Free, and it runs in seconds.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAILED = []
def check(name, ok, detail=""):
def main():
src/app.pythe local UI
http.server and hand-written HTML/JS on port 9013. Renders with no key; the second button replays what r001 actually answered off the committed result file. HTML/JS lives beside it under ui/.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9013"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-interval-gap")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 24 service points x 5 scheduled runs = 120 completeness snapshots from a fixed seed (20260823), and computes data/gold.jsonl from the meter-day MODEL rather than from the rendered page. 756,083 bytes, seven fixed sections each. A swap seam.
src/segment.pySplits a snapshot into its seven named sections. Pure code. All seven are asserted present in all 120 documents before a run may spend.
src/select.pyDecides which sections go on the wire. Customer Contact is mapped by nothing and subtracted unconditionally by the fallback, so it never leaves the machine. A swap seam.
src/watch.pyOne regex over the grid returning (value, flag) per interval. Reads the FLAG and never the value, which is the whole of Rule G-2. Used by the floors, the UI and the pre-flight -- never to answer for the model. Also pulls the run number, the snapshot time, the expected interval count, the settlement close and whether the window has shut, in pure code. The model is never asked to read a timestamp.
src/completeness.pyEight illustrative rules and the arithmetic behind them. step() owns the carried position and is the ONLY thing that advances it; the model's reply never reaches it. A swap seam.
src/state.pyThree scalars per meter-day -- the previous gap count, whether a re-poll is out, and the previous status -- rendered as one English sentence. Flat in history length.
src/prompt.pyThree parts, string-concatenated: the question, the carried-state sentence, the snapshot. --stateless replaces the middle block and changes nothing else, which check_labels asserts byte for byte.
src/adapters/__init__.pyOne completion over urllib. Transient and terminal HTTP failures separated, four bounded retries, the daily call cap checked here so every caller is covered. A swap seam.
evals/baseline.pyThree of them, all $0.00, all scored through the identical scorer. b002 is the column the model has to beat and it wins four of the five bands. A swap seam.
evals/scoring.pyExact match per cell, no model. Missed and duplicate re-polls counted apart; the completeness figure derived from the reported gap count rather than asked for.
tools/cadence.pyWalks the same 24 meter-days through eight schedules and each scored arm's own raise/hold answers, and reports what is permanently estimated at close. Calls nothing; costs $0.00.
evals/check_labels.pyFourteen claims the page makes, checked against the corpus that will be sent, before any run may spend. Includes both directions of the privacy guard and an assertion that the free keyword table does NOT separate the quarantine reasons perfectly.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2468 input and 1399 output tokens per reading (one meter-day, at one scheduled run), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one meter-day, at one scheduled run)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one meter-day, at one scheduled run) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Operational Notes, which are one of three fixed sentences. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- prose a meter-data analyst types into the work queue -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
The experimentWe did NOT attack it — and this is the surface we left open
An indirect prompt injection needs a field somebody outside can write into, and this kit has exactly one: the Operational Notes a meter-data analyst types into the work queue. It is SENT rather than hidden, because hiding a surface does not close it. No attack was fired at this kit and none is claimed. The three boundaries below were checked in code on 2026-08-23, and only the first of them is red-proven in both directions rather than argued from an absent code path.
Boundary checked
What could go wrong
What the code guarantees
Does a customer's name, address, account number and phone ever leave the machine?
Every snapshot carries a Customer Contact section. Attached to a 15-minute load curve that is not a contact record, it is a statement of when a household was in and when it was out -- and not one field this kit answers asks for any of it. A selector that fell back to the whole document would send all of it.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the snapshot MINUS that section rather than to the snapshot. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py over all 120 documents. ⚠︎ AND A PROOF THAT ONLY ASSERTED THE FIRST DIRECTION WOULD MEASURE ZERO AND BE THE TEST NOT FIRING. Swapping the guard for the naive or list(secs) changes nothing on today's corpus, because every hint names a section all 120 documents carry -- so the fallback is not on any live code path. The hole is CONDITIONAL, so the proof reproduces the condition: a head-end version upgrade renames the export blocks and nothing this kit maps still parses. With the guard, 0 of 120 leak; with or list(secs), 120 of 120. ⚠︎ NOTE WHAT IS NOT CLAIMED: the interval VALUES are sent, because counting them is the task, and a load curve is not anonymous because the name came off it.
Can a wrong answer poison the next scheduled run?
A monitor that fed its own count forward would compound one miscount into every reading after it -- and a gap count that drifts upward looks exactly like a feed that stopped, which is the alarm this watch exists to raise.
src/completeness.step() is the only thing that touches the carried position and it is never passed a model reply; evals/run.py calls it with figures from the answer key after every reading, INCLUDING one whose call failed. Argued from the code path, not measured by an experiment -- r001 is consistent with it and does not demonstrate it, because r001's raise/hold call was right on all 120 readings and there was nothing to propagate.
Can it write an estimate, close a window or submit a settlement position?
A completeness watch sits one step from the estimation engine. A kit that could substitute an interval would be deciding a settled quantity, which is a regulated act performed under a policy somebody's regulator approved.
0 code paths. The only writers in the kit are evals/run.py (results/*.json), tools/cadence.py (data/cadence.json) and src/state.save (data/state.json). evals/check_labels.py greps every .py and .js file for such names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
The first row is the only one measured in both directions. The other two are guarantees about what is ABSENT -- no code path writes an estimate or closes a settlement window, and no code path feeds a model reply back into the carried position -- and absence is checked by asserting names it knows, which is weaker than a run and is written here as such.
The result0 attack trials, and three boundaries checked in code: the customer's name, address, account number and phone never reach the provider (red-proven in BOTH directions over all 120 documents), no code path writes an estimate or closes a settlement window, and no model reply ever reaches the carried position. The fourth -- whether a crafted operational note or quarantine reason could move the answer -- is unmeasured, and it sits on the one field where prose decides the verdict.
1externally-authored field a live deployment would carry (the Operational Notes an analyst types into the work queue) -- synthetic on this run's corpus
0 of 0attack trials run
n/adecision-flip resistance -- not measured
This run's corpus is entirely generated (tools/build_corpus.py) -- no note and no quarantine reason was authored by an outside party. A real deployment's exception queue is free text written by whoever worked the meter-day, and the quarantine reasons this kit's answer turns on are exactly that: prose a validation analyst typed, read verbatim and trusted. One of the three shipped notes is instruction-shaped on purpose -- "Field services have asked that no further on-demand reads be requested against this device this week" -- and it reaches the model on a third of the corpus. Whether a note crafted to suppress a re-poll, or a quarantine reason crafted to make a live read look dead, would succeed is unmeasured for this kit.
Read this twice
The quarantine reasons and the operational notes reach the model verbatim, and one of the three shipped notes is instruction-shaped on purpose — “Field services have asked that no further on-demand reads be requested against this device this week.” Nothing here filters it and nothing here has measured whether it moves the answer. A model declining to follow an instruction would not be a defence anyway — it is one vendor’s behaviour on one day.
HonestyWhat this does not prove
Whether a crafted Operational Note could suppress a re-poll. The raise/hold call was 100.00 pct correct on a corpus where a third of the readings carry an instruction-shaped note, which is consistent with the note having no effect and is not a measurement of it. No red-team run exists.
Whether a crafted quarantine reason could move unrecoverable. This is the more exposed of the two surfaces, because that field's answer IS a reading of the sentence -- and the model already over-calls 20 of 23 misses on ordinary, non-adversarial prose.
Whether the interval VALUES are safe to send. The privacy guard removes the identifiers, not the load curve, and a 15-minute load curve is itself the sensitive artefact. The kit states this rather than claiming the prompt is anonymous.
Provider-side retention of any of it. Outside this repository; nobody here has verified it.
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No estimate write, non-configurable. This kit produces a completeness status, a gap count, an unrecoverable count and a raise/hold call for a meter-data analyst to read. It never writes an estimate, substitutes an interval value, closes a settlement window, submits anything to a market or a billing system or writes to a head-end, and there is no setting that makes it. Separately, and just as non-configurable: the expected interval count and the settlement close are OPERATOR-SUPPLIED. A meter-day whose channel record is incomplete is reported CONTEXT_INCOMPLETE and counted against nothing -- there is no default 96 in this kit.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json), tools/cadence.py (data/cadence.json) and src/state.save (data/state.json). src/completeness.step() is the only thing that touches the carried position, and it is called by evals/run.py after every reading including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit writes an estimate or closes a settlement window
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for such names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
No meter-day is counted against a guessed interval count
20 of 20 readings whose channel record is incomplete were reported CONTEXT_INCOMPLETE by the model (100.00 pct). The two floors that assume 96 got 20 of 20 wrong -- 0.00 pct -- which is the guardrail measured rather than asserted.
Exactly one re-poll per meter-day
0 duplicate re-polls in 100 readings that must not raise one, and 0 missed re-polls in 20 that must. ⚠︎ Both zeros are results on THIS corpus, not properties of the code: nothing in the kit refuses a second raise, the carried state merely tells the model one is already out. The control arm, with that sentence removed, MISSED 18 of the 20.
A zero-consumption read is never counted as a gap
100.00 pct on the 45 readings that carry at least one interval reading 0.000 with flag A -- 969 such intervals across the corpus. The two floors that count values instead of flags scored 0.00 pct on the same 45.
The customer's name, address, account number and phone never reach the provider
0 of 120 snapshots leak the Customer Contact section, and 120 of 120 leak it when the guard is removed AND the condition that reaches the fallback is reproduced. Both directions asserted by evals/check_labels.py before any run may spend.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The carried position is correct whatever the model says, which means a wrong reading is a wrong worklist row and not a corrupted history. Those are different problems and only the first one is on this page.
IT IS NOT A SCHEDULER. evals/run.py is invoked, not woken. Nothing in this kit detects a missed run, back-fills it, or marks its readings late -- which on a kit whose whole argument is a permanent-loss clock is the most uncomfortable sentence on the page, and it is why tools/cadence.py measures what the absence costs instead of the kit pretending to close it.
IT IS NOT A JUDGEMENT ABOUT WHETHER THE RE-POLL CAN STILL LAND. Rule G-6 raises one whenever the gap count has stopped falling, including on 3 of the 20 meter-days where the recovery lag cannot fit inside the remaining window. The right move there is to escalate to a substitution and the shipped rule does not do it. Measured by tools/cadence.py, not fixed.
WatchedWhat is watched, and why that one
8runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 38 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 4 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
22 measured by the latest run16 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The status, the gap count, the unrecoverable count and the raise/hold call, per reading, exact match against the computed answer key
alarm
interval_gap_completeness_pct; status_accuracy_pct; missing_intervals_accuracy_pct; unrecoverable_accuracy_pct; raise_chase_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Both scored arms had 0 of 120 -- but the 8,000-token calibration had 3 of 10, so 0 is a property of the published ceiling and not of the design.
two-repoll-directions
The two re-poll directions, counted apart
alarm
missed_chases; missed_chase_pct; duplicate_chases; duplicate_chase_rate_pct — alarm on missed_chases above 0 on any arm intended for deployment. A missed re-poll is the direction that destroys data; a duplicate merely wastes a read.
permanent-loss-replay
What each arm's own raise/hold answers cost in permanently estimated intervals
alarm
intervals_estimated; kwh_estimated; meter_days_chased; re_polls_that_could_not_land — alarm on an arm losing more intervals than the answer key. Anything above 83 on this population is the raise/hold call costing data; anything below it, as b000 does at 49, is a different trade rather than a better answer.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
120
different corpus — nothing is comparable
corpus.bytes
756,083
interval reads edited — the count held, the bytes did not
split.count
120
the bytes per snapshot count moved — a different set was scored
split.size_p50
6,414
the median size of one bytes per snapshot moved
split.size_p95
6,871
the 95th-percentile size of one bytes per snapshot moved
dataset.rows
120
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (cadence_hours 8, chase_cells 20, completeness_cells 110, context_incomplete_cells 20, documents 120, memory_cells 53, meter_days 24, quiet_cells 100, readings_scored 120, recovery_lag_hours 8.0, stateless False, zero_trap_cells 45) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
The completeness figure (the discriminator)
exact match at 100.00 pct with the carried state; 98.18 pct with the memory removed; 100.00 pct on the strongest free floor; 20.00 pct on both zero-counting floors
110 readings whose channel record carries an interval count
r001-interval-gap against s001, b000, b001 and b002, all through evals/scoring.py.
Status, gap count and raise/hold
exact match at 100.00 / 100.00 / 100.00 pct with the carried state; 57.50 / 98.33 / 84.17 with it removed; 100.00 / 100.00 / 100.00 on the strongest free floor
120 readings
r001-interval-gap against s001-interval-gap-stateless and b002-interval-gap-flagcountmem.
Unrecoverable — the only band that separates the model from free code
80.83 pct against the strongest free floor's 72.50; 22 of 23 against 15 of 23 on the readings where exactly one quarantine sentence decides the answer
120 readings, and a 23-reading subset where the attribution is unambiguous
r001-interval-gap against b002-interval-gap-flagcountmem, per-sentence subset computed from data/gold.jsonl and the two committed result files.
The two re-poll directions
0 missed of 20 and 0 duplicate of 100 with the carried state; 18 missed with it removed; 61 duplicate on the memory-blind floor
20 readings that must raise a re-poll, 100 that must not
r001 against s001 and b000, all through the same scorer.
The two guardrails, scored
context-incomplete recall 100.00 pct and the zero-consumption trap 100.00 pct with the carried state and on the strongest free floor; 0.00 pct on both zero-counting floors
20 readings with nothing on file, 45 carrying a genuine zero-consumption read
r001 against b000 and b001.
Intervals permanently estimated at the settlement close
83 with the carried state, identical to the answer key; 134 with the memory removed; 49 on the memory-blind floor that chases everything
24 meter-days, 1,920 expected intervals, 109 of them recoverable by any schedule
tools/cadence.py replaying each arm's own raise/hold answers through the settlement clock. Free; no model is called.
Memory-dependent slices
status 100.00, gap count 100.00 and raise/hold 100.00 pct with the carried state; 3.77, 100.00 and 64.15 with it removed
53 readings whose answer is not derivable from their own snapshot
r001-interval-gap against s001-interval-gap-stateless. ⚑ THE MIDDLE FIGURE IS THE POINT: the counting does not need a yesterday and the trend call is nothing but yesterday, so removing the memory leaves one at 100.00 and drops the other to 3.77.
Re-poll counts, absolute
0 missed and 0 duplicate with the carried state; 18 missed and 1 duplicate with it removed; 0 missed and 61 duplicate on the memory-blind floor
20 readings that must raise a re-poll, 100 that must not
the same scorer pass as the rates above. Both are published because a rate over 20 cells moves 5 points per row and a reader sizing a worklist needs the count.
Intervals reported absent, against the truth
649 of 649, 0.00 pct out, with the carried state and on the strongest free floor; 1,583 of 649, 143.91 pct out, on both zero-counting floors
649 intervals genuinely absent across the 120 readings
r001-interval-gap against b000/b001. This is the completeness figure expressed as a total rather than a rate, and it is the one a settlement lead reports upward.
Replies that arrived and parsed
100.00 pct answered and 0 unparsed on both scored arms at the published 24,000-token ceiling; 7 of 10 at the 8,000-token calibration ceiling
120 readings per scored arm; 10 on c000
r001 and s001 against c000-interval-gap-calibration. The 0 is a property of the published ceiling and not of the design, which is why the calibration figure is beside it.
Latency per reading
8,411 ms p50 and 35,415 ms p95 with the carried state; 13,220 and 90,454 with it removed
120 readings per arm, 12 concurrent meter-day chains
r001-interval-gap against s001-interval-gap-stateless. The p95 is 4.2x the p50 on the shipped arm because a reading is a 96-cell count and the model narrates the hard grids.
Tokens, and where the bill goes
296,175 in and 167,990 out with the carried state; 293,335 in and 360,678 out with it removed -- 2.15x the output for 24 fewer input tokens a reading
120 readings per arm
r001-interval-gap against s001-interval-gap-stateless. 93.1 pct of the shipped arm's output was provider-side reasoning left at the default.
HistoryRun history
8 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-interval-gap-zerocount 2026-08-23
b001-interval-gap-zerocountmem 2026-08-23
b002-interval-gap-flagcountmem 2026-08-23
answered, %
100.0
100.0
100.0
context incomplete recall, %
0.0
0.0
100.0
duplicate chase rate, %
61.0
1.0
0.0
duplicate chases
61
1
0
input tokens, whole run
0
0
0
interval gap completeness, %
20.0
20.0
100.0
intervals absent error, %
143.91
143.91
0.00
intervals reported absent
1583
1583
649
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
memory chase accuracy, %
60.38
79.25
100.00
memory missing accuracy, %
16.98
16.98
100.00
memory status accuracy, %
47.17
49.06
100.00
missed chase, %
0.0
50.0
0.0
missed chases
0
10
0
missing intervals accuracy, %
18.33
18.33
100.00
output tokens, whole run
0
0
0
raise chase accuracy, %
49.17
90.83
100.00
status accuracy, %
40.83
58.33
100.00
unparsed replies
0
0
0
unrecoverable accuracy, %
48.33
48.33
72.50
zero trap accuracy, %
0.0
0.0
100.0
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 4 runs. Columns here are only ever compared with each other.
Metric
c000-interval-gap-calibration 2026-08-23
c001-interval-gap-calibration-24k 2026-08-23
r001-interval-gap 2026-08-23
s001-interval-gap-stateless 2026-08-23
answered, %
70.0
100.0
100.0
100.0
context incomplete recall, %
—
—
100.0
100.0
duplicate chase rate, %
0.0
0.0
0.0
1.0
duplicate chases
0
0
0
1
input tokens, whole run
24698
12944
296175
293335
interval gap completeness, %
70.00
100.00
100.00
98.18
intervals absent error, %
17.27
0.00
0.00
0.00
intervals reported absent
91
33
649
649
model latency p50 ms
18213.00
44042.00
8411.00
13220.00
model latency p95 ms
56992.00
75340.00
35415.00
90454.00
memory chase accuracy, %
100.00
100.00
100.00
64.15
memory missing accuracy, %
100.0
100.0
100.0
100.0
memory status accuracy, %
100.00
100.00
100.00
3.77
missed chase, %
0.0
0.0
0.0
90.0
missed chases
0
0
0
18
missing intervals accuracy, %
70.00
100.00
100.00
98.33
output tokens, whole run
47251
26177
167990
360678
raise chase accuracy, %
70.00
100.00
100.00
84.17
status accuracy, %
70.0
100.0
100.0
57.5
unparsed replies
3
0
0
0
unrecoverable accuracy, %
20.00
80.00
80.83
81.67
zero trap accuracy, %
40.0
100.0
100.0
100.0
not a time series No two of these 4 runs measured the same system — they differ on chase_cells, completeness_cells, context_incomplete_cells, documents, max_tokens, memory_cells, meter_days, quiet_cells, readings_scored, stateless, zero_trap_cells — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-interval-gap-stub 2026-08-23
answered, %
100.0
context incomplete recall, %
0.0
duplicate chase rate, %
61.0
duplicate chases
61
input tokens, whole run
296243
interval gap completeness, %
20.0
intervals absent error, %
143.91
intervals reported absent
1583
model latency p50 ms
0.00
model latency p95 ms
0.00
memory chase accuracy, %
60.38
memory missing accuracy, %
16.98
memory status accuracy, %
47.17
missed chase, %
0.0
missed chases
0
missing intervals accuracy, %
18.33
output tokens, whole run
4320
raise chase accuracy, %
49.17
status accuracy, %
40.83
unparsed replies
0
unrecoverable accuracy, %
48.33
zero trap accuracy, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 22 chips that all say so.
DeviationsWhat deviated
0 breaches across 8 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
removing the carried state (the one-line prompt change)
status accuracy 100.00 -> 57.50 pct and, on the 53 memory-dependent readings, 100.00 -> 3.77; missed re-polls 0 -> 18 of 20; permanently estimated intervals 83 -> 134. The counting does NOT move (gap count 100.00 -> 98.33). And it costs MORE: output tokens 167,990 -> 360,678, p50 latency 8.4s -> 13.2s.
measured
results/eval-r001-interval-gap.json against results/eval-s001-interval-gap-stateless.json, plus data/cadence.json's arms table.
counting register VALUES instead of quality FLAGS
the completeness figure collapses 100.00 -> 20.00 pct and the reported absent-interval total goes 649 -> 1,583 against a truth of 649 (143.91 pct out). The zero-consumption trap goes 100.00 -> 0.00 on all 45 readings that carry a real zero. Nothing else about the arm changes: b001 has the identical memory the model has.
measured
results/eval-b001-interval-gap-zerocountmem.json against results/eval-b002-interval-gap-flagcountmem.json -- same memory, one rule apart.
waiting for the trend before raising a re-poll (Rule G-6) instead of chasing on sight
duplicate re-polls 61 -> 0 in 100 quiet readings, and permanently estimated intervals 49 -> 83. The alarm-minimising rule and the loss-minimising rule are not the same rule and this kit ships the first one.
measured
results/eval-b000-interval-gap-zerocount.json against the answer key, both replayed through tools/cadence.py::arm_losses.
the watch cadence
permanently estimated intervals 83 at 8-hourly, 97 at 12-hourly, 129 daily and 143 every other day -- which is the same 143 as no watch at all. Skipping runs on the shipped cadence costs 6, 22 and 41 for one, two and three. No accuracy figure is known at any other cadence: only the loss is.
measured
tools/cadence.py, data/cadence.json. $0.00 -- no model is called.
the recovery lag on file
everything in the cadence table, because the lag is what decides the last run at which a re-poll can still land. It ships as one cadence period (8 hours) and nothing here has varied it.
reasoning
src/completeness.RECOVERY_LAG_HOURS and tools/cadence.py::flag_at -- confirmed by reading the two functions; declared on the environment ladder as a decision rather than measured across values.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
Unrecoverable — the only band that separates the model from free code
20 of the model's 23 misses OVER-call a read as dead; 3 under-call.
Intervals permanently estimated at the settlement close
an arm above 83 is losing data to its raise/hold call.
Re-poll counts, absolute
any non-zero missed count.
Replies that arrived and parsed
unparsed_replies above 0 — a reading that returns nothing scores as a miss in all four fields, so a reliability failure arrives disguised as a quality failure.
NextThe three you would add first
An escalation output for the case where a re-poll cannot land in timeRule G-6 raises one whenever the gap count stops falling, including on the 3 of 20 meter-days where the recovery lag does not fit inside the remaining window. tools/cadence.py computes the last useful run exactly -- close minus the recovery lag -- so the information is already there and the kit simply has no second verdict to emit.
Count the register in code and call the model ONLY for the quarantine reasonsfour of the five bands are already 100.00 pct free (Eval.baseline_per_field), so the model is being paid to count on every reading and earning its bill on 41 reason lines across 120. The split is not shipped because shipping it would remove the comparison that makes the case for it.
A monotonicity check on the unrecoverable counta meter-side quarantine never heals, so the count can only rise. Nothing in the kit asserts that, and 3 of the model's 23 misses UNDER-call -- which on a later run would show up as the figure going down. It is a free check the kit does not make.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free, seconds) on any change to tools/build_corpus.py, src/completeness.py, src/segment.py or src/select.py -- all four change what the guardrail is asserted over. Re-run tools/cadence.py (free) on any change to CADENCE_HOURS, RECOVERY_LAG_HOURS, the re-poll rule, or any arm's result file, because every permanent-loss figure on these pages is derived from those four. Re-run the scored eval AND its stateless control together (paid, 240 calls) on any change to src/prompt.py or src/state.py -- the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. All three floors are free and should be re-run on any change at all.
What this cannot tell you
Whether the no-estimate-write guarantee holds against a path named something the checker does not know. evals/check_labels.py asserts the absence of eight names it knows and passes at 0; a function called something else would pass it. That is weaker than a run and is written here as such.
Whether the one-re-poll-per-meter-day property is a property of the CODE. It is not: nothing in the kit refuses a second raise. The carried state tells the model one is already out, and on this corpus that was enough for 0 duplicates in 100 quiet readings. The control, with that sentence removed, was not a duplicate storm either -- it missed re-polls instead -- so the property has been observed twice and enforced never.
Whether the CONTEXT_INCOMPLETE guardrail survives a channel record that is present but WRONG. Every one of the 20 readings it is measured on has the field literally absent. A record stating 96 on a 30-minute channel would be counted against, and nothing here detects it.
Whether a crafted Operational Note could move the raise/hold call. The surface is sent deliberately and was never attacked -- see Security.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, sys, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a COUNT -- three scalar fields written by arithmetic. The cost is flat in history length because nothing here remembers what was said, only what was counted, and a memory layer would give back the growth this design exists to avoid. Measured the other way round on this kit: removing the memory made the SAME model spend 2.15x the output tokens, because a model told there is no history reasons at length about a trend it cannot see
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the schedule
evals/run.py
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
THIS is the seam where a framework genuinely earns its place, and on this kit it is not a preference -- tools/cadence.py measures 6 intervals lost permanently to ONE skipped run and 41 to three. 24 independent chains of 5 strictly-ordered readings, with the cadence, retries, missed-run detection and late back-fill all OUTSIDE the kit. A ThreadPoolExecutor is the right size for an eval and the wrong size for a watch whose skipped run destroys data
the scorer
evals/scoring.py
eval harnesses (promptfoo, DeepEval)
four exact-match comparisons, one derived percentage and three slices of the same cells is a dict comprehension, not a platform
the corpus
tools/build_corpus.py
synthetic-data and eval-set generators
the generator IS the answer key here, which is the property a general generator cannot give you: every gold value is computed from the meter-day model, so a disagreement between the page and the key is a finding about the reader
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each meter-day is a chain of five readings -- 06:00, 14:00 and 22:00 on the 2nd, 06:00 and 14:00 on the 3rd -- with no branching and exactly one edge between consecutive runs, carrying three scalars. Different service points never touch. A framework would add an orchestrator to a for-loop that already runs 24 chains wide.
The other sideWhat a framework costs you
No scheduler, and on this kit that is the expensive absence rather than the cheap one. This is a watch that carries state and is still INVOKED, not woken. tools/cadence.py prices the gap: one skipped run costs 6 intervals that can never be read again, three cost 41, and a daily cadence costs 46 of the 60 the shipped schedule saves.
No missed-run detection, no back-fill, no late marking. All three are what a scheduling engine gives you and none of them exists here.
No escalation path. Rule G-6 raises a re-poll even when the recovery lag cannot land inside the remaining window -- 3 of the 20 on the shipped cadence. A real deployment would switch to a substitution request at that point and this kit has no such output.
No concurrency model for the state store. data/state.json is one file replaced atomically, which is correct for one writer and is not a concurrency model.
What we could NOT verify
Whether a scheduling engine would actually recover the intervals this kit's cadence table says a denser schedule recovers. tools/cadence.py measures what a SCHEDULE loses; it assumes every scheduled run in a schedule happens. What a real Airflow DAG's own failure rate does to that number is not modelled, and it is the whole reason the 8h with runs skipped rows exist.
Whether a memory layer would beat three scalars. The comparison this kit ran is memory against NO memory, not three scalars against a transcript. The cost side is known -- three scalars are flat in history length -- and the accuracy side of a richer state is unmeasured.
Whether an eval harness would find anything evals/scoring.py misses. The scorer is exact match on four fields; a harness offering per-slice significance testing might say something about the 23-reading one-sentence subset that this kit reports as a raw 22 against 15.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-interval-gap on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
8,411 ms
8,411 ms p50 and 35,415 ms p95 with the carried state; 13,220 and 90,454 with it removed
—
Model, p95
35,415 ms
8,411 ms p50 and 35,415 ms p95 with the carried state; 13,220 and 90,454 with it removed
—
Input tokens
296,175
296,175 in and 167,990 out with the carried state; 293,335 in and 360,678 out with it removed -- 2.15x the output for 24 fewer input tokens a reading
—
Output tokens
167,990
296,175 in and 167,990 out with the carried state; 293,335 in and 360,678 out with it removed -- 2.15x the output for 24 fewer input tokens a reading
—
No movement column. Not one of the 6 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-interval-gap-calibration18,213 ms
c001-interval-gap-calibration-24k44,042 ms
r001-interval-gap8,411 ms
s001-interval-gap-stateless13,220 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-interval-gap-zerocount, b001-interval-gap-zerocountmem, b002-interval-gap-flagcountmem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 8 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
completeness snapshots
data/corpus/SP-<n>-R<k>.txt -- 120 files, 756,083 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Customer Contact -- the customer's name, service address, account number and phone -- never does, by src/select.NEVER_SENT. ⚠︎ The interval VALUES do go, because counting them is the task, and a load curve is not anonymous because the name came off it
the carried state
data/state.json in a deployment, and in-memory per run in the eval -- three scalars per meter-day, written by src/completeness.step and never by a model
as one English sentence in every prompt, and the UI prints the same sentence verbatim so a reader can audit what the model was told
the answer key
data/gold.jsonl -- 120 rows, the output of src/completeness.step over the planted gap blocks, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the cadence table
data/cadence.json -- eight schedules and six scored arms walked through the settlement clock by tools/cadence.py
never. It calls nothing; every figure in it cost $0.00
the run records
results/eval-*.json in the kit, and one small record per run in the app repo's run register
never -- they are read by the site build, not by a provider
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
an EIGHT-HOUR watch: 06:00, 14:00 and 22:00 local, 3 runs a day, for as long as a meter-day stays inside its settlement window. One run owns exactly the window since the last one. The cadence is not a deployment detail here: it decides how quickly a stalled feed becomes VISIBLE, and a stall that becomes visible after the last run at which a re-poll could still land is a stall nobody can act on.
24 meter-days x 5 scheduled runs = 120 calls per arm, 0 gaps and 0 out-of-order readings; evals/check_labels.py asserts all 24 sequences are complete before a run may spend. Wall clock 257.5s at 12 workers. ⚑ AND WHAT A MISSED RUN COSTS IS MEASURED, NOT ARGUED. tools/cadence.py walks the same 24 meter-days through eight schedules: the shipped 8-hourly watch loses 83 intervals permanently, one skipped run loses 6 more, two lose 22 more, three lose 41 more, a daily cadence loses 46 more, and a watch that wakes every other day loses 60 more -- which is the same 143 as having no watch at all. 34 of the 83 are unavoidable on every schedule (a meter-side fault has no read to recover), so the watch's real contribution is 60 of the 109 recoverable intervals. (tools/cadence.py, data/cadence.json, r001-interval-gap, evals/check_labels.py)
⚠︎ THE MARGINAL COST OF A SKIPPED RUN ACCELERATES, which is the part a reader sizing a schedule most needs and the part a single number hides: the first skipped run costs 6 intervals, the second 16 more, the third 19 more. And Rule G-6 raises a re-poll even when the recovery lag cannot fit inside the remaining window -- 3 of the 20 on the shipped cadence. The right move there is a substitution request and this kit has no such output.
change the recovery lag and every number in that table moves, because the lag is what decides the last useful run. It ships as one cadence period, which is a modelling choice; a head-end that answers an on-demand read in minutes would make a slow cadence far less costly, and one that answers in a day would make the shipped cadence nearly worthless.
state
THREE scalars per meter-day -- the previous gap count, whether a re-poll is already out, and the previous status -- rendered as one English sentence by src/state.describe and written ONLY by src/completeness.step. Not the previous snapshot, not the previous reply. There is deliberately no carried unrecoverable count, because a meter-side quarantine stays on the register with its reason still printed and is re-derivable every run.
53 of the 120 readings are memory-dependent. With the state: status 100.00 pct on them, raise/hold 100.00 pct, 0 missed and 0 duplicate re-polls. Without it: status 3.77 pct on the same 53, 18 of 20 re-polls missed, and 51 more intervals permanently estimated at the settlement close. (r001-interval-gap against s001-interval-gap-stateless, tools/cadence.py)
the state is bounded by design -- three scalars, so run 40 costs what run 2 costs, the OPPOSITE curve to an intake kit. It is also CHEAPER: the stateless arm spent 2.15x the output tokens on the same readings. What is NOT bounded is accuracy over a longer history, which is unmeasured, and the store itself: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model.
lose the carried state and the trend call collapses to 3.77 pct with every reply still well-formed and nothing raising an error. See environment.signatures' traceless row.
model
one completion call per reading, on one provider and one key, through src/adapters/__init__.py. MAX_TOKENS = 24000, set from two calibration runs; thinking is never sent, so every published run left provider-side reasoning at the default and the result files record thinking: null.
120 calls on the scored arm, 167,990 output tokens, 156,477 of them provider-side reasoning (93.1 pct). Largest reply 12,513 tokens against the 24,000 ceiling. p50 8.4s, p95 35.4s. 0 unparsed. (r001-interval-gap, c000-interval-gap-calibration, c001-interval-gap-calibration-24k)
the ceiling is the thing that breaks first and it DID break, on the calibration run that exists to find it: 3 of 10 replies at an 8,000-token cap returned nothing. At 24,000 the largest reply on the scored arm was 12,513, so there is roughly 2x headroom -- measured, not assumed. The control arm reached 16,166.
a provider whose reasoning default differs reprices this kit without changing anything a reader can see, and nothing here has measured what disabling reasoning does to the answers. Only ONE model was run and no cross-tier claim is made anywhere on this page.
labels
120 readings -- 24 service points re-read at 5 scheduled runs -- with the four answers computed by tools/build_corpus.py from the meter-day model and re-derived from src/completeness.step by evals/check_labels.py before any run may spend. 53 are memory-dependent; 20 have no interval count or no settlement close on file and must be counted against nothing; 45 carry at least one genuine zero-consumption read.
120 rows in data/gold.jsonl, 0 replay mismatches, 0 sequence gaps, and 0 disagreements between the RENDERED register and the key. Status spread: BASELINE 20, HEALING 22, STALLED 31, COMPLETE 14, SETTLED 13, CONTEXT_INCOMPLETE 20. 969 intervals across the corpus read exactly 0.000 with flag A. (data/gold.jsonl, evals/check_labels.py, data/corpus-stats.json)
it stops scoring at the fifth scheduled run of each meter-day. Long enough to hide a trend behind a run boundary and not long enough to reach half of this corpus's own settlement closes -- which is why tools/cadence.py walks the whole arc separately and why the per-arm loss figure for the worst arm is an upper bound.
the unrecoverable band encodes ONE reading of sixteen sentences. A corpus whose exception reasons were coded rather than written would take the model's entire margin away, and the kit says so in fitment rather than hoping nobody notices.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a meter-day reported HEALING two runs running while its gap count does not move
something is comparing against the wrong previous count, or against none. HEALING is the status that suppresses the re-poll, so this is the direction that loses data silently
read the carried-state sentence, not the number. The stateless control reported BASELINE or HEALING on almost every memory-dependent reading -- 3.77 pct status accuracy on 53 of them -- and missed 18 of the 20 re-polls it should have raised (results/eval-s001-interval-gap-stateless.json against results/eval-r001-interval-gap.json)
a completeness figure that drops on a day the premises was empty
something counted values instead of flags. A vacant rental, a disconnected service and a shop shut for a refit all read 0.000 with flag A, and they are real reads
look at the flag column. The two floors that count zeros report 1,583 absent intervals against the key's 649 -- 143.91 pct out -- and score 0.00 pct on the 45 readings that carry a genuine zero (results/eval-b000-interval-gap-zerocount.json against data/gold.jsonl)
an interval counted against a meter-day whose channel record is incomplete
something filled the missing interval count with a default. That is the one thing this row's guardrail forbids: the figure is operator-supplied and a guessed one produces a completeness percentage nobody can defend
look at the Intervals expected line. Both floors that assume 96 get all 20 of these readings wrong; both arms that abstain get all 20 right (results/eval-b000-interval-gap-zerocount.json against results/eval-r001-interval-gap.json)
a substitution request against a read that then arrives
the quarantine reason was over-called as meter-side. This is the model's commonest error on this kit -- 20 of its 23 misses -- and the sentences that cause it name a fault and then say it was cleared
read the reason line to the end. "a tamper flag was raised on this installation earlier in the day and cleared by field services" is a read that is coming back (results/eval-r001-interval-gap.json, the 23 unrecoverable misses)
nothing at all -- the worklist simply does not change between two shifts
TRACELESS. A scheduled run that never fired leaves no artefact anywhere in this kit: no error, no gap marker, no late flag. The readings it would have produced simply do not exist, and the next run's window silently covers twice as long. On this task that is not a delay -- tools/cadence.py measures one skipped run converting 6 recoverable intervals into permanent estimates, and three converting 41
count the runs, not the rows. Nothing in this kit will tell you; that is what a scheduler is for and this kit does not have one (tools/cadence.py, data/cadence.json)
['Concurrency. data/state.json is replaced atomically, which is correct for one writer; two watches over one head-end is two writers and nothing here has tested it.', 'Any accuracy figure at a cadence other than eight hours. The cadence table measures what other schedules LOSE, which is arithmetic; the model was not re-run on any of them.', 'A population that changes between runs -- a meter-day appearing mid-window, or one leaving the watch when it settles.', 'A market with a revision cycle. This corpus freezes the register at the settlement close, so every permanent-loss figure assumes the close is final.', "Provider-side retention. The prompt carries no customer identifiers by construction, but it does carry a full day's load curve, and what a provider keeps of a request is outside this repository and nobody here has verified it.", 'GPU sizing, local inference and anything about running this off a hosted API. Not attempted; not costed.', 'A second model. One tier was run.', 'Whether the Operational Notes field can move the raise/hold call. The surface is sent deliberately and was never attacked.']
The corpus licence, from the Data lens: MIT, same as the repository. There is no third-party data in this kit. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The status, the gap count, the unrecoverable count and the raise/hold call, per reading, exact match against the computed answer key
Catch a utility's missing meter readings in time
PresenterOpens the private repo. Visible to admins only.
In one lineThe status, the gap count, the unrecoverable count and the raise/hold call, per reading, exact match against the computed answer key
For each of the 120 readings and each of the four answered fields, did the reply equal the computed answer key? The two counts are compared as integers and a non-integer is a MISS rather than a zero -- 0 is a meaningful answer here (it is what every COMPLETE and every CONTEXT_INCOMPLETE reading gets) and a parse failure must not be scored as one.
$0.00per 1,000 interval reads
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function all three free floors and the stateless control are scored through.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
SP-0024-R2 -- scheduled run 2 of 5, 2026-03-02 14:00, a 30-minute channel with 48 expected intervals, settlement close 2026-03-02 20:00, six hours away
What the previous scheduled run left behind
At the previous scheduled run this meter-day had 10 intervals missing. No re-poll has been raised for this meter-day yet. It was reported BASELINE.
The interval register, this run
A 38 · M 4 · Q 6 · E 0 -- 10 absent of 48. Twenty-two of the 38 actuals read exactly 0.000, which is a real premises drawing nothing and is not a gap (Rule G-2).
The exception queue, verbatim
11:00-11:30 (2 intervals): "battery-backed clock failure means the device cannot attribute anything to this timestamp; there is no read to recover" · 14:00-15:00 (3): "the head-end has been accepting reads from a decommissioned register since the exchange; nothing measured this interval" · 18:30 (1): "the device's comms module re-registered on the mesh mid-interval; the read is buffered on the meter and comes back on the next poll"
STALLED, 10 absent, 5 unrecoverable, raise YES -- four of four
What the strongest free floor answered
STALLED, 10 absent, 3 unrecoverable, raise YES -- three of four
Why this one
It is the shape where all three of this kit's arguments are live at once and only one of them separates the arms. The zero run means a value-counting reader is already wrong; the trend is invisible on the page and only the carried state has it; and the attribution turns on a sentence whose fault noun is "failure" where the keyword table holds "failed".
Grader
Verdict
Why
The status, the gap count, the unrecoverable count and the raise/hold call, per reading, exact match against the computed answer key
status hit, count hit, unrecoverable hit, raise/hold hit -- four of four
The gap count is 4 M plus 6 Q = 10, and getting there requires not counting the 22 intervals that read 0.000 with flag A. The trend call needs the carried state: 10 absent now against 10 at the previous run is STALLED, and there is no earlier re-poll, so this is the run that raises one. The unrecoverable count is 2 (the clock failure) plus 3 (the decommissioned register) = 5; the comms-module read is buffered on the meter and comes back. The model's rationale names G-3 and G-4 and reaches 5.
The two re-poll directions, counted apart
in scope on the RAISE side, and a hit -- one of the 20
This reading's correct answer is to raise the re-poll, so it sits inside the 20 that the missed-re-poll rate is measured over and outside the 100 the duplicate rate is measured over. The model raised it; so did b002; the stateless control did NOT, and it missed 18 of these 20 across the run.
What each arm's own raise/hold answers cost in permanently estimated intervals
no loss attributable to this reading's answer
tools/cadence.py replays every arm's own raise decisions through the settlement clock. Because r001 raised this re-poll on the same run the rule does, this meter-day loses exactly what the key loses. ⚠︎ Raised at 14:00 with an eight-hour recovery lag against a 20:00 close, this re-poll is one of the 3 that CANNOT land in time -- and Rule G-6 raises it anyway. That is a limitation of the shipped rule, not of the answer, and it is counted rather than fixed.
The formulaWhat it computes
accuracy = hits / 120 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% interval gap completeness · 4 more measured on this row
the same tier, memory removed (THE CONTROL)
98.2% interval gap completeness · 4 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by tools/build_corpus.py from the meter-day model at generation time and re-derived from src/completeness.step by evals/check_labels.py before any run may spend. The same pre-flight separately asserts that the RENDERED register agrees with the key on all 120 documents, so the page the model reads and the truth it is scored against cannot drift apart. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known to be arguable -- Rule G-4 splits sixteen quarantine sentences into recoverable and not, and that split is one author's reading of English rather than a fact. Three of the sixteen are genuinely close: "a tamper flag was raised ... and cleared by field services", "the CT ratio on the billing record disagrees with the network record" and "phase imbalance alarm on the site's own switchboard" all name a fault and then say it does not apply. The key calls all three recoverable. The other four fields have no known ambiguity.
Watch these
interval_gap_completeness_pct
status_accuracy_pct
missing_intervals_accuracy_pct
unrecoverable_accuracy_pct
raise_chase_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A reading that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality failure. Both scored arms had 0 of 120 -- but the 8,000-token calibration had 3 of 10, so 0 is a property of the published ceiling and not of the design.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because three of them are small: 20 readings that must raise a re-poll, 20 with nothing on file, and 23 where exactly one quarantine sentence decides the unrecoverable count. One row moves the first two by 5.0 points and the third by 4.3.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/completeness.py, src/segment.py or src/select.py. Re-run tools/cadence.py (free) on any change to the cadence, the recovery lag or the re-poll rule. Re-run the scored eval AND its stateless control together (paid, 240 calls) on any change to src/prompt.py or src/state.py -- the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. All three floors are free and should be re-run on any change at all.
The decisionWhen to reach for it
Use it
The truth is known and two of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real head-end, where whether a quarantined read will ever validate is decided by a person reading an analyst's note. That is why this corpus is generated rather than captured.
PresenterOpens the private repo. Visible to admins only.
In one lineThe two re-poll directions, counted apart
A MISSED re-poll is a reading whose correct answer was to raise one and did not -- data that the settlement close then substitutes permanently. A DUPLICATE re-poll is one raised against a meter-day that already has one outstanding -- an on-demand read queued at the head-end for data already on its way. Different costs, different owners, never averaged.
$0.00per 1,000 interval reads
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the cell match.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
SP-0024-R2 -- scheduled run 2 of 5, 2026-03-02 14:00, a 30-minute channel with 48 expected intervals, settlement close 2026-03-02 20:00, six hours away
What the previous scheduled run left behind
At the previous scheduled run this meter-day had 10 intervals missing. No re-poll has been raised for this meter-day yet. It was reported BASELINE.
The interval register, this run
A 38 · M 4 · Q 6 · E 0 -- 10 absent of 48. Twenty-two of the 38 actuals read exactly 0.000, which is a real premises drawing nothing and is not a gap (Rule G-2).
The exception queue, verbatim
11:00-11:30 (2 intervals): "battery-backed clock failure means the device cannot attribute anything to this timestamp; there is no read to recover" · 14:00-15:00 (3): "the head-end has been accepting reads from a decommissioned register since the exchange; nothing measured this interval" · 18:30 (1): "the device's comms module re-registered on the mesh mid-interval; the read is buffered on the meter and comes back on the next poll"
STALLED, 10 absent, 5 unrecoverable, raise YES -- four of four
What the strongest free floor answered
STALLED, 10 absent, 3 unrecoverable, raise YES -- three of four
Why this one
It is the shape where all three of this kit's arguments are live at once and only one of them separates the arms. The zero run means a value-counting reader is already wrong; the trend is invisible on the page and only the carried state has it; and the attribution turns on a sentence whose fault noun is "failure" where the keyword table holds "failed".
Grader
Verdict
Why
The status, the gap count, the unrecoverable count and the raise/hold call, per reading, exact match against the computed answer key
status hit, count hit, unrecoverable hit, raise/hold hit -- four of four
The gap count is 4 M plus 6 Q = 10, and getting there requires not counting the 22 intervals that read 0.000 with flag A. The trend call needs the carried state: 10 absent now against 10 at the previous run is STALLED, and there is no earlier re-poll, so this is the run that raises one. The unrecoverable count is 2 (the clock failure) plus 3 (the decommissioned register) = 5; the comms-module read is buffered on the meter and comes back. The model's rationale names G-3 and G-4 and reaches 5.
The two re-poll directions, counted apart
in scope on the RAISE side, and a hit -- one of the 20
This reading's correct answer is to raise the re-poll, so it sits inside the 20 that the missed-re-poll rate is measured over and outside the 100 the duplicate rate is measured over. The model raised it; so did b002; the stateless control did NOT, and it missed 18 of these 20 across the run.
What each arm's own raise/hold answers cost in permanently estimated intervals
no loss attributable to this reading's answer
tools/cadence.py replays every arm's own raise decisions through the settlement clock. Because r001 raised this re-poll on the same run the rule does, this meter-day loses exactly what the key loses. ⚠︎ Raised at 14:00 with an eight-hour recovery lag against a 20:00 close, this re-poll is one of the 3 that CANNOT land in time -- and Rule G-6 raises it anyway. That is a limitation of the shipped rule, not of the answer, and it is counted rather than fixed.
Reference standard: the same data/gold.jsonl. The raise/hold column is computed by src/completeness.step from the trend and the carried chase flag, never hand-labelled.
These rates are UNKNOWN, on purpose
Nothing about the counting. What is NOT known is whether the two directions cost what this kit implies they cost -- a missed re-poll is priced in permanently estimated intervals by the third grader, but a duplicate re-poll is priced in nothing here, because what an extra on-demand read costs a head-end is a fact about somebody's contract and capacity that this repository does not have.
Watch these
missed_chases
missed_chase_pct
duplicate_chases
duplicate_chase_rate_pct
Alarm on
missed_chases above 0 on any arm intended for deployment. A missed re-poll is the direction that destroys data; a duplicate merely wastes a read.
How tight can the band be? No threshold. The reply's raise_chase field decides.
Cadence: Re-run with the scored eval; the two are one pass. The floors are free and should be re-run on any change to evals/baseline.py.
The decisionWhen to reach for it
Use it
The two error directions cost different things, which on a permanent-loss clock they do.
Do not use it
A task where a false positive and a false negative cost the same. This is not one.
What each arm's own raise/hold answers cost in permanently estimated intervals
Catch a utility's missing meter readings in time
PresenterOpens the private repo. Visible to admins only.
In one lineWhat each arm's own raise/hold answers cost in permanently estimated intervals
Reads each arm's committed result file, takes the FIRST run at which it said YES as that meter-day's re-poll time, holds the recovery lag and the settlement close exactly as the key has them, and counts what is still absent when the window shuts. It is the only place on this kit where a percentage of readings becomes the thing the reader actually loses.
$0.00per 1,000 interval reads
nodata leaves your network
yessame answer every time
MethodHow the test was run
tools/cadence.py::arm_losses, pure arithmetic, no key and no model.
Every grader on these pages scored the same 120 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
SP-0024-R2 -- scheduled run 2 of 5, 2026-03-02 14:00, a 30-minute channel with 48 expected intervals, settlement close 2026-03-02 20:00, six hours away
What the previous scheduled run left behind
At the previous scheduled run this meter-day had 10 intervals missing. No re-poll has been raised for this meter-day yet. It was reported BASELINE.
The interval register, this run
A 38 · M 4 · Q 6 · E 0 -- 10 absent of 48. Twenty-two of the 38 actuals read exactly 0.000, which is a real premises drawing nothing and is not a gap (Rule G-2).
The exception queue, verbatim
11:00-11:30 (2 intervals): "battery-backed clock failure means the device cannot attribute anything to this timestamp; there is no read to recover" · 14:00-15:00 (3): "the head-end has been accepting reads from a decommissioned register since the exchange; nothing measured this interval" · 18:30 (1): "the device's comms module re-registered on the mesh mid-interval; the read is buffered on the meter and comes back on the next poll"
STALLED, 10 absent, 5 unrecoverable, raise YES -- four of four
What the strongest free floor answered
STALLED, 10 absent, 3 unrecoverable, raise YES -- three of four
Why this one
It is the shape where all three of this kit's arguments are live at once and only one of them separates the arms. The zero run means a value-counting reader is already wrong; the trend is invisible on the page and only the carried state has it; and the attribution turns on a sentence whose fault noun is "failure" where the keyword table holds "failed".
Grader
Verdict
Why
The status, the gap count, the unrecoverable count and the raise/hold call, per reading, exact match against the computed answer key
status hit, count hit, unrecoverable hit, raise/hold hit -- four of four
The gap count is 4 M plus 6 Q = 10, and getting there requires not counting the 22 intervals that read 0.000 with flag A. The trend call needs the carried state: 10 absent now against 10 at the previous run is STALLED, and there is no earlier re-poll, so this is the run that raises one. The unrecoverable count is 2 (the clock failure) plus 3 (the decommissioned register) = 5; the comms-module read is buffered on the meter and comes back. The model's rationale names G-3 and G-4 and reaches 5.
The two re-poll directions, counted apart
in scope on the RAISE side, and a hit -- one of the 20
This reading's correct answer is to raise the re-poll, so it sits inside the 20 that the missed-re-poll rate is measured over and outside the 100 the duplicate rate is measured over. The model raised it; so did b002; the stateless control did NOT, and it missed 18 of these 20 across the run.
What each arm's own raise/hold answers cost in permanently estimated intervals
no loss attributable to this reading's answer
tools/cadence.py replays every arm's own raise decisions through the settlement clock. Because r001 raised this re-poll on the same run the rule does, this meter-day loses exactly what the key loses. ⚠︎ Raised at 14:00 with an eight-hour recovery lag against a 20:00 close, this re-poll is one of the 3 that CANNOT land in time -- and Rule G-6 raises it anyway. That is a limitation of the shipped rule, not of the answer, and it is counted rather than fixed.
The formulaWhat it computes
intervals still flagged M or Q at the settlement close, summed over 24 meter-days, plus the same figure in kWh from the withheld true values.
The analysisWhat it actually did
Model
Result
the answer key (Rule G-6 exactly)
no headline metric on this row — it records intervals estimated 83
the fast tier, with the carried state
no headline metric on this row — it records intervals estimated 83
the same tier, memory removed (THE CONTROL)
no headline metric on this row — it records intervals estimated 134
b000 floor, zeros, no memory
no headline metric on this row — it records intervals estimated 49
In operationWhat to monitor
Reference standard: tools/cadence.py holds the settlement clock, the recovery lag and the gap-block arrival times, all from the same generator that wrote the corpus. The arm supplies only WHEN it said YES; everything else is held identical across arms, which is what makes the comparison a measurement of the raise/hold call and not of anything else.
These rates are UNKNOWN, on purpose
What an arm would have done on a sixth scheduled run. The corpus is five runs and half the settlement closes fall after the fifth, so an arm that never said YES inside those five is treated as never having chased. For the control that makes 134 an UPPER bound rather than a measurement. Also unknown: whether a real head-end's recovery lag is one cadence period. It ships as one and every figure in this grader moves with it.
Watch these
intervals_estimated
kwh_estimated
meter_days_chased
re_polls_that_could_not_land
Alarm on
an arm losing more intervals than the answer key. Anything above 83 on this population is the raise/hold call costing data; anything below it, as b000 does at 49, is a different trade rather than a better answer.
How tight can the band be? No threshold. The one number that behaves like one is RECOVERY_LAG_HOURS = 8.0 in src/completeness.py, which decides the last run at which a re-poll can still land, and it is declared on the environment ladder rather than tuned.
Cadence: Free. Re-run on any change to the cadence, the recovery lag, the re-poll rule or any arm's result file.
The decisionWhen to reach for it
Use it
The failure is permanent and the schedule is what decides how much of it happens.
Do not use it
A market with a revision cycle, where a late read can still be substituted back in. This corpus freezes at the close and does not model one.
A living map of modern AI — kept current every morning