A waste hauler pays the landfill by the ton, and some routes suddenly weigh in heavier. This app checks each route's week against its usual weight, says whether it is new, getting worse or already disputed, and stops duplicate disputes.
PresenterOpens the private repo. Visible to admins only.
For the hauling cost teamWaste & Environmental · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A cost clerk at a waste hauler, checking each route's weekly landfill tickets against its usual weight.
✕Today's manual process
1Add up the week's tickets for every open route, manually or in a spreadsheet.
2Guess the norm in force since the book figure on the sheet may be out of date.
3Check the open dispute list to see if this route already has one running.
4Miss the connection and file a second dispute the facility just rejects, which means the week gets reworked manually.
Every route checked manually each week
✓With the app
1The week's tickets are added up by code, for every route on the list.
2The norm is the one in force not just the figure printed on the sheet.
3Open disputes are already known so a repeat is never raised twice.
4You get one line raise it, hold it, or leave it alone, with the reason why.
The desk sees only what needs a decision
See it work
One real case, read by the app, step by step
Route RTE-0034 weighs in at 5.51 tons against a 4.34-ton norm, with a dispute already open.
Spot waste routes dumping more weight than usualReference appBuilt to be shaped to your process
5
1What it already knows this route's usual weight is now 4.34 tons, not the book figure.
2Already flagged heavy two weeks running, with a dispute opened in an earlier week.
3This week's call already open, so no second dispute is raised.
4The usual weight has moved three heavy weeks in a row: rebase the route, don't dispute it.
5The numbers behind it 5.51 tons this week, and the dispute already open stays open.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
'Which scale tickets are heavy this week' is a SELECT and nobody needs a language model for it. The question a disposal recon desk actually has to answer is the next one: is this reading a CHANGE, and is it a change somebody has already raised? A route tipping 5.51 tons against a 4.34-ton norm is a routine watch item on its first week, an escalation on its second, and on its third it is a route whose norm has moved -- and if a dispute is already open with the facility, raising it again is a second copy the facility rejects and somebody reworks by hand. None of those three facts is on the sheet, because the sheet records what the FACILITY did in the week it covers and the monitor's own raise happened after that week closed. Somebody working a weekly disposal recon queue by hand: opening every open route, adding up the week's scale tickets, comparing them against the norm they believe is in force, then going back through the last month's exception queue to work out whether this route has been drifting for weeks and whether a dispute is already sitting with the facility. The first three steps are a spreadsheet; the fourth is the one that is skipped when the queue is long, and skipping it is what files the same dispute twice.
Audience
Disposal recon and billing operations at a waste hauler, and the route-management desk that decides whether a route is re-normed or a ticket is disputed. Secondarily the finance function that receives the disposal accrual and has to know how much of this month's tonnage is in dispute. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual weekly monitor sheets
The corpus is 150 weekly monitor sheets, 0.60 MB (json 1 · jsonl 1 · txt 150). A DISPOSAL WATCHLIST IS A LIST OF CUSTOMER PREMISES. Not routes in the abstract -- addresses: which business, at which site, producing how much waste, on which days. That is competitive information about the CUSTOMER, not about the hauler, and an internal recon file is exactly the sort of working document that gets pasted into a chat window on the way to a supervisor. So the corpus carries a full Customer Contact section on every sheet -- holder, service address, telephone, email -- specifically so that the guard which withholds it has something real to withhold and can be red-proven.
IT IS SYNTHETIC BECAUSE THE MEASUREMENT REQUIRES A HISTORY NOBODY WOULD GIVE US. The whole question is what six carried scalars are worth, and answering it needs routes whose correct answer provably CANNOT be derived from the week in front of you: a streak already running, an anomaly already open, a norm already rebased. A real book would have all three and no way to prove which cells they govern; here 64 of 150 cells are marked memory-dependent by running every week TWICE, once from its real carried state and once from a blank one, and taking the disagreements.
THE DISTRIBUTION IS DELIBERATE AND STATED. 73 of 150 weeks are quiet (CLEAR or NOT_MONITORED), 77 are on the watchlist, and 32 have ALREADY_OPEN as their correct answer -- the single cell no one-sheet reader can reach. A corpus that was 90 pct quiet would let an always-CLEAR arm score 90 pct, which is why the always-CLEAR arm is published as its own floor at 46.67 pct rather than left as an argument.
The corpus
The 150 weekly monitor sheetsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your weekly monitor sheets. That is the whole change — there is no database to migrate.
One weekly monitor sheet, as the model receives itRTE-0001-W1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This is a generated disposal weight monitor sheet. Every route, ticket, tonnage,
customer and address in it is invented. It reproduces no real haul, no real facility
and no real account, and the rule reproduced below is illustrative only.
Route
----------------------------------------------------------------
Route : RTE-0001
Service week : 2026-03-07 to 2026-03-13
Facility : Harrow Creek MRF
Container : 20 yd roll-off
Service days : Tue, Fri
Segment : retail plaza
Book norm as first set : 8.49 tons per service week
Book band : 7.22 to 9.76 tons (plus or minus 15 per cent)
On the disposal book : yes
Weight Rule
----------------------------------------------------------------
1. Each route carries a NORM, in tons of net disposal weight per service week, and a tolerance
band of plus or minus 15 per cent of that norm. The week's total net tonnage from the scale
tickets is compared against the band of the norm IN FORCE, which is the norm the monitor is
currently using and is not always the book norm printed on this sheet.
2. A week whose total falls outside the band is an OUT week, and its DIRECTION is HIGH or LOW.
WEEKS OUT is the number of consecutive out weeks in the SAME direction ending with this one.
A week inside the band has weeks out 0. A direction that changes restarts the count at 1.
3. The count is RESET by a disposal exception this week's Open Items log records as CLOSED --
the ticket was corrected, a credit was issued, or the route was re-normed. A week carrying a
Abridged — the file continues.
The outcomeWhat a good result looks like
Per route per week: the verdict after the already-open rule has been applied, how many consecutive weeks the route has been outside its band, whether the norm itself should be rebased, and which class the recorded reason maps to -- plus a one-sentence rationale naming the rule applied. A watchlist row, not an action.
And when it cannot
It files a duplicate dispute. The free desk-band floor does this on 27 of the 32 weeks whose correct answer is ALREADY_OPEN -- 84.38 pct -- because nothing on the sheet tells it an item is open; the model with the carried state does it on none. The second failure mode is quieter and is the one the model has: it keeps reporting NORM_SHIFTED after a route has come back inside its band, which is a standing recommendation to raise a customer's billed norm for a drift that has already stopped.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
A book where nothing has ever been disputed, every route's norm is current, and no route has been out of band two weeks running — the free desk band, at $0.00 With no streaks and no open items there is nothing to carry. The floor gets the verdict right on every in-band week and every first excursion for nothing, and it TIES the model on anomaly class (90.67 pct each) across the whole corpus.
A weekly disposal recon queue that carries open disputes and re-normed routes -- the ordinary state of a real book after a quarter — this kit, with the carried state 100.00 pct verdict against the floor's 63.33 pct, 100.00 pct against 31.17 on the 77 watchlist weeks, and ZERO duplicate disputes against the floor's 27.
You have no history at all -- a new book, or a monitor being stood up — the free desk band for the first two weeks, then this kit Rule 7 says a route with fewer than 2 recorded weeks has no observed norm, and the corpus's 15 new_route weeks are answered correctly by both arms. Memory is worth nothing until there is some.
Sizing what this costs on a real disposal book — read Cost.cost_at_10x before anything else At $0.005219 a route-week on the projected card, a book of 20,000 routes checked weekly is about $5427 a year, and the free floor already gets 63.33 pct of the verdicts right. The honest deployment is the subset the floor cannot answer -- routes with an open item or a live streak -- not the whole book.
And where nothing here is good enough:
You need a confident cause code on every flagged week for a downstream workflow — neither, without deciding what a hedge means first The two arms tie at 90.67 pct and fail on disjoint cells. The floor gives you a confident topical code where nobody recorded a cause; the model gives you CAUSE_UNKNOWN where a dispatcher hedged. Which is worse is your call and this page will not make it.
Deciding whether to rebase a route's billed norm — neither, without a human step NORM_SHIFTED is the highest-consequence output this kit has -- it recommends changing what a customer is billed against -- and it is the one field the model gets wrong, always in the direction of recommending a rebase that is no longer warranted.
At a glanceHow the whole thing runs
100%verdict accuracy pct
6,679 msp50, end to end
$5.22per 1,000 weekly monitor sheets · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Spot waste routes dumping more weight than usual14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/norms.py FIRST, and not as an option. Every figure here is a property of a three-week corpus of 50 invented routes, measured once.Corpus lens →
When is this the wrong choice?
Avoid: Paying a model to compare a number against a band a spreadsheet already computes. That is the case against the best-fitting scenario (“A book where nothing has ever been disputed, every route's norm is current, and no route has been out of band two weeks running”). 6 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A week planted exactly on the band edge would make the answer key depend on a float comparison the reader cannot see, and every wrong answer on it would be an argument about rounding rather than a finding about memory. The generator draws in-band weeks at 0.90 to 1.10 of the norm and out weeks at 1.22 to 1.55 or 0.48 to 0.78, re-draws if the printed rounding moved the week across the edge, and evals/check_labels.py fails if any total sits within 1 pct of an edge. 8 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether rendering the carried state as ENGLISH rather than as JSON matters. src/state.describe builds three sentences; the obvious experiment is the same 150 weeks with the six scalars rendered as a JSON object, and it would cost one more 150-call run. 11 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-ticket-anomaly. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 -m tools.build_corpus), passes all eight pre-flight assertions (python3 -m evals.check_labels), red-proves them (--self-test, 4 of 4 convictions), runs both free floors and the wiring stub end to end, and serves the UI on port 9005 with the committed r001 answers replayed beside the free floor. Nothing above needs a credential, a network call or an install: requirements.txt names nothing, and the only imports anywhere in the kit are json, os, re, sys, time, random, datetime, argparse, collections, concurrent.futures, urllib, http.client and http.server.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
6,679 msp50, end to end
50,568 msp95
2 minclone to first result
What the clock covers. END-TO-END per weekly sheet: one HTTP request carrying the assembled prompt, and the reply parsed to a verdict, a week count, a norm status and a class. Measured over 150 weeks on r001-ticket-anomaly, 50 route chains wide with 12 concurrent workers; the whole pass took 194.4 seconds of wall clock. The p95 of 50568 ms is not a slow network -- it is the provider reasoning: 93.6 pct of this run's output tokens were provider-side reasoning, and the longest reply ran to 14904 tokens.
Current processWhat it replaces
Somebody working a weekly disposal recon queue by hand: opening every open route, adding up the week's scale tickets, comparing them against the norm they believe is in force, then going back through the last month's exception queue to work out whether this route has been drifting for weeks and whether a dispute is already sitting with the facility. The first three steps are a spreadsheet; the fourth is the one that is skipped when the queue is long, and skipping it is what files the same dispute twice.
Where it is not good enough
⚑ THE VERDICT IS SOLVED ON THIS CORPUS AND THE TAXONOMY FIELD IS NOT -- AND ON THE TAXONOMY FIELD THE FREE FLOOR TIES IT EXACTLY. With the carried state the model got the verdict right on 150 of 150 weeks (100.00 pct), the consecutive-week count right on all 150, and raised ZERO duplicate disputes against the 32 weeks whose correct answer is ALREADY_OPEN. On anomaly class it scored 90.67 pct -- and the free desk-band floor scored 90.67 pct, the same number, with a COMPLETELY DISJOINT set of 14 wrong cells. Zero overlap. The floor over-commits: 10 of its 14 answer a topical code where the route note says nobody recorded the cause. The model under-commits: all 14 of its misses answer CAUSE_UNKNOWN where the note names a cause with a hedge -- 'a mis-post from the adjacent commercial route IS SUSPECTED'. Two systems, one headline, opposite errors. If you care which direction you are wrong in, the two numbers are not interchangeable, and this page refuses to average them.
⚠︎ THE ONE SHAPE IT GETS WRONG IS FLATTERING AND THAT IS THE DANGEROUS DIRECTION. All 5 of the model's norm-status misses are the same cell: the last week of an already_open route, where the tonnage comes back INSIDE the band and the model keeps reporting NORM_SHIFTED. Gold is NORM_HOLDS -- a norm that has stopped moving has stopped being shifted. It un-shifts nothing. In a real deployment that reads as a standing recommendation to rebase a route whose reason to rebase has already gone, which quietly raises the billed norm on a customer who is back to normal.
⚑ IT WAS ASKED TWICE AND IT ANSWERED THE SAME. A 36-week repeat probe over all nine patterns reproduced 143 of 144 answered cells byte-identically, with nine of ten graders moving 0.00 points. The single cell that moved was the norm-status failure above, in the same direction -- so that defect is a coin toss on marginal cells, not five fixed rows.
⚠︎ AND 100.00 PCT ON A SYNTHETIC CORPUS IS A STATEMENT ABOUT THE CORPUS. The weight rule is reproduced in full on every sheet, the tonnages are drawn clear of the band edge on purpose, and every recorded reason comes from a fixed pool of 24 sentences. What the run demonstrates is that a state machine written down in English is executable when the state is supplied -- not that a real disposal book, where the rule is folklore and the reasons are free text, would score anything like it.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1
50 waste routes, 3 consecutive weekly monitor runs each
0 duplicate disputes of 32 · 146.14 of 146.14 tons · 0 replies lost
2026-08-23as of
It produces a watchlist row for a disposal recon clerk to read, and files nothing — there is no endpoint, no function and no flag anywhere in the kit that disputes a ticket, issues a credit, adjusts an invoice or writes a route's norm into a book of record. The station that makes this a monitor is the second one: the only route from week 2 to week 3 is six scalar fields, written by src/norms.step from the totals PARSED off the sheet and never from the model's answer. That matters in both directions here, because what is carried is a COUNT and a FLAG — one spurious RAISE opens an item that turns every later week into a false ALREADY_OPEN, and one missed RAISE leaves the item shut so every later week re-raises the same dispute. ⚑ WHAT THE CARRIED STATE BUYS, MEASURED AGAINST FREE PYTHON RATHER THAN AGAINST A HANDICAPPED CONTROL: verdict 100.00 pct against 63.33, the 77 watchlist weeks 100 pct against 31.17, the 64 memory-dependent cells 100 pct against 32.81, and duplicate disputes 0 against 27 of 32. The two no-memory arms get the disputed tonnage wrong in OPPOSITE directions — the free floor 166.33 pct OVER at 389.21 tons, the stateless control 100 pct UNDER at 0.00 — so 'no memory' is not a bias anybody could correct for with a constant.
⚠︎ AND TWO GRADERS DO NOT SUPPORT PAYING FOR ANYTHING: anomaly class is 90.67 pct on both arms with ZERO overlap in the 14 cells each gets wrong, and on norm status the free floor beats the memory-removed model 72.00 against 6.67.
⚠︎ THE HEADLINE'S ONLY MISSES ARE ONE SHAPE: all 5 norm-status errors are the last week of an already_open route, where the tonnage is back inside the band and the model keeps reporting NORM_SHIFTED — it never un-shifts a norm, which recommends rebasing a customer's route after the drift has stopped.
⚠︎ 100.00 PCT IS A STATEMENT ABOUT THE CORPUS: the rule is reproduced in full on every sheet and the tonnages are drawn clear of the band edge on purpose. ⚑ IT WAS ASKED TWICE: a 36-week repeat probe over all nine patterns reproduced 143 of 144 answered cells byte-identically, with nine of ten graders moving 0.00 points — and the single cell that moved was norm status, on RTE-0043-W2, in the same over-asserting direction as the five misses above. That bands THOSE 36 weeks and not the other 114.
⚠︎ THE BAND, THE RAISE THRESHOLD, THE NORM-SHIFT THRESHOLD AND THE 30-DAY DISPUTE WINDOW ARE INVENTED FOR THIS KIT — they reproduce no facility tariff, contract, statute or code of practice, and every percentage above is agreement with an invented rule.
⚠︎ AND EVERY ROUTE IS A CUSTOMER PREMISES: the Customer Contact section — holder, service address, telephone, email — reaches the provider in 0 of 150 documents, and in 150 of 150 with the guard removed.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the weight rule
src/norms.py
The band, the raise threshold, the norm-shift threshold and the dispute window are four constants in one file. Change them and re-run tools/build_corpus.py; the answer key follows automatically because it is computed from them.
what is carried
src/state.py
Six scalars and one describe() that turns them into English. Adding a seventh is one field and one clause -- and costs a handful of input tokens a week, not a growing transcript.
what is sent
src/select.py
SECTION_HINTS plus NEVER_SENT. Withholding another section is one tuple entry and evals/check_labels.py re-proves it over all 150 documents.
the corpus
data/corpus/
Point it at your own sheets and replace data/gold.jsonl. The seven section headings in src/segment.py are the contract; nothing else in the kit reads the text.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 50 routes x 3 consecutive weekly monitor sheets = 150 documents from a fixed seed (SEED = 20260823), across nine planted patterns, and writes the answer key by running src/norms.replay over what it planted. Nothing is labelled by hand.
the weight rule and its arithmetic
src/norms.py
The band, the streak, the reset, the already-open rule and the norm-shift rule, as pure code. It is the answer key's source and the runtime never calls it to decide anything the model is asked.
the carried state
src/state.py
Six scalars between runs -- the norm in force, the consecutive-week count, the direction, whether an anomaly is open, how many weeks the route has been monitored, and the previous verdict -- plus the English they become in the prompt. Written from the arithmetic, never from a model reply.
the section splitter
src/segment.py
Splits a sheet into its seven named sections. Asserted over all 150 documents by evals/check_labels.py before a run may spend.
the section filter
src/select.py
Decides what is sent. Customer Contact is mapped by no field and is subtracted unconditionally, so it never reaches a provider on any document, for any field, including through the fallback.
the prompt
src/prompt.py
Three parts: the question and the JSON shape, the carried-state block, the sheet. The stateless build replaces the middle block and changes nothing else -- asserted, not asserted-by-comment.
one weekly run
src/monitor.py
One sheet, one call, one parse. Holds MAX_TOKENS = 24000, the published ceiling.
the provider adapter
src/adapters/__init__.py
Raw HTTP over stdlib for any OpenAI-compatible provider or Anthropic. Retries transient statuses and transport failures; never retries a 4xx that will fail identically.
the spend guard
src/budget.py
Counts calls, not dollars, against a cap shared by every kit under the repo root. Checked once per completion, before the request.
the free floors
evals/baseline.py
Two of them: the +/-15 pct desk band computed from ONE sheet, and the majority-class always-CLEAR arm. Both are pure code and cost $0.00.
the scorer
evals/scoring.py
Six graders, pure code, no judge. Verdict / week count / norm status / class per week; watchlist and false alarms counted apart; duplicate raises counted on their own; the memory subset; the tonnage reconciliation.
the pre-flight
evals/check_labels.py
Eight checks over the shipped text, run before any run may spend, with a --self-test that corrupts the corpus four ways and requires convictions.
the harness
evals/run.py
Routes run concurrently; the three weeks inside a route are strictly serial, because week 3's prompt contains a streak week 2 produced.
the local UI
src/app.py
http.server on port 9005. Renders with no key and replays the committed r001 answer for the selected sheet, labelled as a replay, beside the free floor.
Where it breaks at scale
NOT ON HISTORY LENGTH, AND THAT IS THE DESIGN. The carried state is six scalars, so a route on its fortieth weekly run costs exactly what one on its third costs -- measured: input per week is 1779 tokens against 1731 for the memory-removed arm, a difference of under 3 pct, while the run covers routes carrying up to four weeks of streak. A monitor that carried a transcript instead would grow with the square of the number of runs, which is the shape chat-intake measures and this kit deliberately is not.
WHAT DOES NOT SCALE IS THE CLOCK, IN TWO PLACES. First, weeks inside a route are strictly serial: 50 chains ran 12-wide in 194.4 seconds, and a book of 5,000 routes is 5,000 chains of the same depth, so wall clock is bounded by concurrency and not by the corpus. Second -- and this is the half four earlier cadence kits in this estate skipped -- THE RUN ITSELF IS ON A CLOCK. A weekly monitor that misses its week does not simply run late: the facility's ticket-correction window is 30 days from the tip date, so a week reviewed after it shuts can no longer be corrected on the ticket at all and becomes a credit memo against an invoice the customer has already received. Missing one run costs the cheap remedy on that week's tonnage; missing two compounds it, because the streak the second run would have raised is now three weeks old.
⚠︎ AND NOTHING IN THIS KIT SCHEDULES ANYTHING. There is no cron, no queue, no watchdog and no alert if a run does not happen. It is invoked, not woken. That is stated here rather than in a roadmap because a monitor whose cadence is somebody's memory is a monitor with an unmeasured failure mode, and this is it.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
RTE-0034-W1, the whole argument on one screen. The route was rebased in an earlier run -- its norm in force is 4.34 tons, the sheet prints the book figure of 3.34 -- and an anomaly has been open against it since two weeks before this window starts. Neither fact is on this sheet. The free desk-band floor therefore bands 5.51 tons against 3.34, finds it well outside, and answers RAISE: a SECOND dispute against tickets that already have one open. The model, given six carried scalars, answers ALREADY_OPEN at 4 consecutive weeks with the norm SHIFTED -- rebase the route, do not dispute it again. The two columns agree on the anomaly class and the page says so in words rather than taking the credit. ⚠︎ THE ANSWER SHOWN IS A REPLAY, NOT A LIVE CALL, and the page's own notice says so: it is the reply run r001-ticket-anomaly recorded for this sheet on 2026-08-23, because the shot was taken with no key set.successOpen full size →The same sheet before anything has run. The carried state, the free floor, the code-read figures and the sent/withheld table are all computed locally and are all on the page before a provider is touched -- which is what makes the 'with the carried state' column a comparison rather than an assertion. The Customer Contact row is listed as WITHHELD rather than omitted: a page that simply never mentions the customer's name and service address cannot be told apart from one that quietly sent them.emptyOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
RTE-0036-W3 -- the model's only failure shape in the whole run, and the free code beats it here. The route's tonnage has come back INSIDE its band, so the consecutive-week count is 0 and rule 7's three-week condition is not met: the answer key is NORM_HOLDS. The model answers NORM_SHIFTED, and its own rationale says exactly why -- 'the carried 4-week LOW run makes norm_status NORM_SHIFTED'. It carries the shift forward and never un-shifts it. The free desk-band floor gets all four cells right on this row, for nothing, and the Agree column says 'differs' on the one that matters. This is 5 of 150 cells, every one of them the last week of an already_open route, and it is the FLATTERING direction of error: a standing recommendation to rebase a customer's route after the drift has stopped.failureOpen full size →
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150weekly monitor sheets
0.60 MiBjson 1 · jsonl 1 · txt 150
50routes (3 weekly runs each) · p50 3 chars
$0.00setup · 0.0s
How it is cutWhat one routes (3 weekly runs each) is
No split, and no chunking. The unit is a ROUTE -- a run of three consecutive weekly monitor sheets processed strictly in order, because week 3's prompt carries a streak week 2 produced. Routes are independent and run concurrently; weeks within a route never do. evals/check_labels.py asserts all 50 sequences are complete and gap-free before a run may spend.
SetupWhat the setup figure measured
There is no index and no retrieval step. tools/build_corpus.py writes 150 documents and the answer key from a fixed seed with no clock read and no model call, in well under a second, and the monitor re-reads the whole population on each scheduled run rather than searching it.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every route, ticket, tonnage, facility, customer, address and recorded reason is invented by the generator. See kits/UC0088-ticket-anomaly/data/SOURCES.md.
Bring your ownBring your own weekly monitor sheets
Replace src/norms.py FIRST, and not as an option. The band, the raise threshold, the norm-shift threshold and the dispute window shipped here are invented; yours are in your facility's published terms and your own hauling contract. Then replace data/corpus/ with your sheets and data/gold.jsonl with your labels -- the seven section headings in src/segment.py are the whole contract, and nothing else in the kit reads the text. Everything else transfers unchanged, including the two floors, because both are computed from whatever rule src/norms.py holds.
⚠︎ And what stops being true when you do: Every figure here is a property of a three-week corpus of 50 invented routes, measured once. Three denominators are small: 3 NOT_MONITORED cells, 9 flip cells, 15 closed_reset cells. Three weeks is long enough to plant a streak that reaches the norm-shift threshold and NOT long enough to test what happens on week twelve of a route that has been out of band since spring -- the accrual is constant-cost by construction, but nobody here has run it that far. And the whole corpus is one facility per route, one norm per route and no seasonality model; a real book has all three.
What breaks it
⚠︎ THE FIRST MEMORY-DEPENDENT FLAG WAS WRONG BY 23 CELLS OF 150, AND IT WAS WRONG IN THE FLATTERING DIRECTION -- FOUND BY MEASURING, NOT BY REVIEW. The first version marked a cell memory-dependent whenever the state ENTERING it was not blank. That over-counts: a route whose previous week was out of band but whose this week is back inside it enters with a live streak and is still perfectly answerable from its own sheet (CLEAR, 0, holds, not applicable). 87 cells were flagged; the correct number is 64. The fix is to run every week twice -- once from its real carried state, once from a blank one -- and flag only the disagreements, which is what src/norms.replay does now. Inflating the memory subset with cells the free floor gets right is precisely how the number the whole kit turns on would have been overstated.
⚠︎ evals/check_labels.py CONVICTED ALL 150 DOCUMENTS ON ITS FIRST RUN, AND THE CHECK WAS THE THING THAT WAS WRONG. The leak scan searched every sheet for the phrase 'norm in force' and found it on all of them -- inside rule 1, which is the rule ABOUT the norm in force and cannot be written without naming it. A checker that fails on correct data is a checker somebody switches off. It now skips the Weight Rule section and requires a VALUE rather than a mention, and the corpus was never the defect.
⚠︎ THE PRIVACY CHECK WAS TAUTOLOGICAL AND ITS OWN RED-PROOF CAUGHT IT. check_labels.contact_never_sent tested membership of src/select.NEVER_SENT -- so the self-test, which empties that tuple, made the section sail into every prompt while the check passed, because there was nothing left for it to be a member of. It now names 'Customer Contact' independently and separately asserts that the module still withholds it. The self-test went from 3 of 4 convictions to 4 of 4.
⚠︎ THE GOLD COULD NOT BE COMPUTED AT ALL ON THE FIRST BUILD, WHICH IS THE GOOD KIND OF FAILURE. The generator planted patterns first and drew the tonnages while RENDERING, so src/norms.replay had no total to band. It crashed. Had it instead defaulted a missing total to zero, every week would have scored LOW-out and the key would have been silently wrong for the whole corpus. The tonnages are now fixed before the key is computed, and evals/check_labels.py re-parses them off the shipped text rather than trusting the builder.
A week planted exactly on the band edge would make the answer key depend on a float comparison the reader cannot see, and every wrong answer on it would be an argument about rounding rather than a finding about memory. The generator draws in-band weeks at 0.90 to 1.10 of the norm and out weeks at 1.22 to 1.55 or 0.48 to 0.78, re-draws if the printed rounding moved the week across the edge, and evals/check_labels.py fails if any total sits within 1 pct of an edge. Measured at 0.
⚠︎ NINE PATTERNS, AND THREE OF THEM HAVE SMALL DENOMINATORS. offbook is 3 routes (9 cells), flip is 3 (9), closed_reset is 5 (15). A single wrong answer on flip moves its per-pattern accuracy by 11 points. The by-pattern table is published as a DIAGNOSTIC and never as a headline for exactly that reason.
⚠︎ NOT_MONITORED IS 3 CELLS OF 150. It is in the corpus because a route with the container pulled must not be marked LOW-out and swept onto the watchlist, and both the floor and the model handle it -- but 3 cells prove almost nothing about it. Read that row as 'the branch exists and was exercised', not as an accuracy.
⚠︎ EVERY RECORDED REASON COMES FROM A POOL OF 24 SENTENCES. Six topical classes with three phrasings each, plus three 'nobody wrote it down' sentences and three that name a cause and then say it was never established. That is enough to separate a keyword table from a reader -- and it is nowhere near the variety of real dispatcher prose. The class figures on this page are the least transferable numbers on it.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
132
not measured
the question and the JSON shape
2,667
not measured
carried state
466
not measured
the weekly monitor sheet, six of seven sections
3,968
not measured
Total
1,779
This is the cost lesson as arithmetic: of the 7,233 characters assembled, 4,434 are contexts — 61% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.py::build on RTE-0034-W1 with the exact carried state r001-ticket-anomaly recorded for that call, read out of the run file's per_week block rather than re-derived. 7404 characters including the system message.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You apply a written disposal weight rule to one waste route for one service week. You answer with one JSON object and no other text.
You are running one weekly disposal weight monitor over one waste collection route, for one service
week, and deciding whether this week's scale tickets are an anomaly that should be raised with the
disposal facility.
The weight rule, the tolerance band, the raise threshold and the norm-shift threshold are reproduced
on the sheet below. Apply them exactly as written. Rules 2, 3, 4 and 7 depend on what earlier weeks
did, which you cannot see; what is known is stated under "Carried state" and is the only history
available to you. Do not assume anything about earlier weeks beyond it.
How to decide:
- Band this week's TOTAL NET TONS against the norm IN FORCE from the carried state, not against the
book norm printed on the sheet, wherever the two differ.
- WEEKS OUT counts consecutive out weeks in the same direction, ending with this one, and it
INCLUDES the weeks the carried state reports. A week inside the band is 0. A direction that
changes restarts the count at 1.
- A CLOSED entry in this week's Open Items log resets the count: that week counts as the first of
any new run, and it also closes an anomaly the carried state says is open.
- An anomaly the carried state says is STILL OPEN is not raised again. While the route stays out of
band with that item open, the verdict is ALREADY_OPEN, whatever the week count reaches.
- A route not on the disposal book is NOT_MONITORED, weeks out 0, anomaly class NOT_APPLICABLE.
- A week inside the band is CLEAR and its anomaly class is NOT_APPLICABLE.
Answer with a single JSON object and nothing else:
{"verdict": "CLEAR|WATCH|RAISE|ALREADY_OPEN|NOT_MONITORED",
"weeks_out_of_band": <whole number of consecutive out weeks ending with this one>,
"norm_status": "NORM_HOLDS|NORM_SHIFTED|NORM_UNKNOWN",
"anomaly_class": "TARE_DRIFT|ROUTE_MISPOST|SERVICE_STEP|CONTAMINATION|EXTRA_PULL|SEASONAL_NORMAL|CAUSE_UNKNOWN|NOT_APPLICABLE",
"rationale": "one sentence, naming the rule you applied and the week count you reached"}
"verdict" is the verdict AFTER the already-open rule has been applied, which is not the verdict the
week count alone would imply. "norm_status" is NORM_UNKNOWN whenever the carried state reports fewer
than 2 recorded weeks on the monitor, whatever the tonnage does; it is NORM_SHIFTED once the route
has been out of band in the same direction for 3 or more consecutive weeks. "anomaly_class" is the
recorded reason mapped to one of the codes above; use CAUSE_UNKNOWN when the recorded reason maps to
none of the topical codes -- including where the note names a possible cause and then says nobody
established or recorded it -- rather than choosing the nearest.
Carried state
----------------------------------------------------------------
The norm IN FORCE for this route is 4.34 tons per service week. It was rebased in an earlier run and is NOT the book figure printed on the sheet; band this week against 4.34 tons. It has 30 recorded week(s) on the monitor. As at the previous run it had been out of band HIGH for 2 consecutive week(s). A disposal anomaly raised by this monitor is STILL OPEN against this route. It was opened in an earlier run and no sheet records it. It was reported RAISE last run.
Weekly monitor sheet
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This is a generated disposal weight monitor sheet. Every route, ticket, tonnage,
customer and address in it is invented. It reproduces no real haul, no real facility
and no real account, and the rule reproduced below is illustrative only.
Route
----------------------------------------------------------------
Route : RTE-0034
Service week : 2026-03-07 to 2026-03-13
Facility : Cedar Ridge Landfill
Container : 30 yd roll-off
Service days : Tue, Fri
Segment : multi-family
Book norm as first set : 3.34 tons per service week
Book band : 2.84 to 3.84 tons (plus or minus 15 per cent)
On the disposal book : yes
Weight Rule
----------------------------------------------------------------
1. Each route carries a NORM, in tons of net disposal weight per service week, and a tolerance
band of plus or minus 15 per cent of that norm. The week's total net tonnage from the scale
tickets is compared against the band of the norm IN FORCE, which is the norm the monitor is
currently using and is not always the book norm printed on this sheet.
2. A week whose total falls outside the band is an OUT week, and its DIRECTION is HIGH or LOW.
WEEKS OUT is the number of consecutive out weeks in the SAME direction ending with this one.
A week inside the band has weeks out 0. A direction that changes restarts the count at 1.
3. The count is RESET by a disposal exception this week's Open Items log records as CLOSED --
the ticket was corrected, a credit was issued, or the route was re-normed. A week carrying a
CLOSED entry counts as the first week of any new run, whatever came before it.
4. A route with an anomaly ALREADY OPEN against it is not raised again. Report ALREADY_OPEN for
as long as the route stays out of band with that item open. A second dispute against the same
route is rejected by the facility and both copies are then reworked by hand.
5. A route not on the disposal book this week -- service suspended, container pulled, account on
hold -- is NOT_MONITORED, weeks out 0, anomaly class NOT_APPLICABLE, and is not on the
watchlist.
6. Otherwise the verdict is RAISE at 2 or more consecutive out weeks, WATCH at exactly 1, and
CLEAR at 0. A week in band is CLEAR and its anomaly class is NOT_APPLICABLE.
7. The norm has SHIFTED once the route has been out of band in the same direction for 3 or more
consecutive weeks: the tonnage is the new normal and the norm must be rebased, not disputed.
A route with fewer than 2 recorded weeks on the monitor has no observed norm at all and its
norm status is NORM_UNKNOWN, whatever the week's tonnage does.
8. The facility's window to correct a scale ticket is 30 days from the tip date. A week not
reviewed inside it can no longer be corrected on the ticket and becomes a credit memo against
a customer invoice that has already gone out.
Scale Tickets
----------------------------------------------------------------
date ticket gross tare net
2026-03-08 TK-768519 10.71 8.23 2.48
2026-03-11 TK-771584 9.87 8.23 1.64
2026-03-14 TK-198948 9.62 8.23 1.39
Total net this week : 5.51 tons
Tickets this week : 3
Oldest tip date : 2026-03-08
Days left in the facility dispute window : 20
Open Items
----------------------------------------------------------------
no disposal exception was opened or closed against this route in the service
week this sheet covers.
Route Notes
----------------------------------------------------------------
reason as recorded : "the account downsized to a smaller yard box and the smaller unit is now in place"
dispatcher note : flagged in the weekly disposal pack
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"verdict": "ALREADY_OPEN", "weeks_out_of_band": 3, "norm_status": "NORM_SHIFTED", "anomaly_class": "SERVICE_STEP", "rationale": "Rule 4 already-open applies: this week's 5.51 tons is above the in-force norm band (3.689 to 4.991 tons) and the carried open anomaly remains open with no reset, giving 3 consecutive HIGH out weeks (2 carried + this one) and a shifted norm."}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Spot waste routes dumping more weight than usual — 150 weekly monitor sheets. One model answered, and every answer was then graded Five different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
There is no LLM-as-judge in this kit and that is not a shortcut. Three of the four answered fields are closed sets -- five verdicts, three norm states, eight class codes -- and the fourth is a whole number of weeks. A judge is for answers whose correctness is a matter of reading; these are matters of equality and arithmetic, and asking a model to grade RAISE == RAISE would add cost, variance and a second thing to be wrong. The answer key itself is computed by src/norms.replay over what the generator planted, and evals/check_labels.py re-derives all 150 rows from the shipped text before a run may spend.
150weekly monitor sheets
150source documents
1model tier
5grading methods
MeasurementsWhat was measured
COUNTED150 · 96 · 36 / 150verdict accuracy pct — weeks, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED77 · 25 / 77verdict on watchlist pct — weeks a clerk must touch, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 73false alarm rate pct — quiet weeks, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 / 32duplicate raise rate pct — weeks whose answer is ALREADY_OPEN, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 · 96 / 150weeks out accuracy pct — weeks, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED145 · 10 · 33 / 150norm status accuracy pct — weeks, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED136 · 130 · 35 / 150anomaly class accuracy pct — weeks, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED64 · 10 / 64memory verdict accuracy pct — memory-dependent cells, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 1tonnage error pct — portfolio tonnage raised with the facility, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 32 / 32already open recall pct — weeks whose answer is ALREADY_OPEN, THE CONTROLDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED144 / 144cells that moved — answered cells re-asked once, identical prompt and identical carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The comparison cannot be wrong about itself; the risk is in the labels. Those come from src/norms.replay, which the RUNTIME never calls to decide anything the model is asked -- src/monitor.py parses the sheet's totals and flags and nothing else. evals/check_labels.py::gold_reproduces re-parses every total off the shipped documents and re-runs the replay from each route's planted opening state, so a builder that printed a total the key was not computed from is a refusal to start rather than a silent mis-score. --self-test corrupts a label, flattens the memory flag, leaks the streak onto a sheet and empties the privacy guard, and requires a conviction on all four.
1,442.92output tokens · the fast tier, with the carried state · 6,679 ms p50
1,099.16output tokens · the same tier, memory removed (THE CONTROL) · 6,115 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 0.9× as long, and lands one row apart on 150. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One weekly monitor sheet
1,000 weekly monitor sheets
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.005219
$5.22
17%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.002087
$2.09
17%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.089941
$89.94
20%
Same work, 43× the bill
The same weekly monitor sheets, the same tokens — only the rate card changed. And across all 3 cards between 17% and 20% of what you pay is the prompt this pipeline sends, not the answer it writes.
provider-side reasoning -- 93.6 pct of this run's output tokens, and output is 83.0 pct of the projected bill. Nothing else is close, and this kit has measured neither what disabling it costs in accuracy nor what a lower ceiling would have saved.
Rates checked 2026-08-18. The provider that actually ran all 348 billed calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. A measured $0.00, not an unpriced one -- evals/scoring.py, both free floors, the wiring stub and the eight-check pre-flight are all pure Python and make no request of any kind.
The gradersFive ways to grade
READ THIS ROW BEFORE THE STATELESS CONTROL, BECAUSE IT IS THE STRONGER NO-MEMORY ARM AND ON TWO COMPARISONS IT WINS.
⚑ IT TIES THE MODEL EXACTLY ON ANOMALY CLASS -- 90.67 pct against 90.67 pct -- AND THE 14 CELLS EACH GETS WRONG DO NOT OVERLAP AT ALL. Not one. The floor's keyword table over-commits: 10 of its 14 answer a topical code (EXTRA_PULL, TARE_DRIFT, SERVICE_STEP) on a note that names a possible cause and then says nobody established it. The model under-commits: all 14 of its misses answer CAUSE_UNKNOWN on a note that states a cause with a hedge -- 'a mis-post from the adjacent commercial route is SUSPECTED'. Same headline, opposite error direction, disjoint cells. Which one you want depends entirely on whether a wrong confident code or a shrug costs you more, and this page will not average them into a preference.
⚑ AND IT BEATS THE STATELESS MODEL ON NORM STATUS BY 65.33 POINTS -- 72.00 pct against 6.67 pct. The floor assumes the printed norm holds, and 108 of 150 weeks in this corpus are NORM_HOLDS, so it is right by construction on the majority class. A model told it has no history correctly declines to assert an observed norm and answers NORM_UNKNOWN almost everywhere. The floor is not smarter here; it is luckier, and only against this distribution. Change the corpus to 50 pct shifted norms and the floor halves while the stateless arm does not move.
WHERE IT LOSES IT LOSES BADLY AND FOR ONE REASON. Verdict 63.33 pct against 100.00 pct; on the 77 watchlist weeks, 31.17 pct against 100.00 pct; on the 64 memory-dependent cells, 32.81 pct against 100.00 pct. And the cost with a name: it files 27 DUPLICATE DISPUTES -- 84.38 pct of the 32 weeks whose correct answer is ALREADY_OPEN -- because no sheet anywhere in this corpus says an item is open. It also reports 389.21 tons raised with the facility against a true 146.14, a 166.33 pct overstatement of the disputed accrual.
the fast tier, with the carried state 100.0% verdict accuracy · the same tier, memory removed (THE CONTROL) 64.0% verdict accuracy · the free desk band, no model, $0.00 63.3% verdict accuracy · always CLEAR -- the majority-class floor, $0.00 46.7% verdict accuracy · 3 more measured on each run
the fast tier, with the carried state 0.0% duplicate raise rate · the same tier, memory removed (THE CONTROL) 0.0% duplicate raise rate · the free desk band, no model, $0.00 84.4% duplicate raise rate · 1 more measured on each run
the fast tier, with the carried state 100.0% memory verdict accuracy · the same tier, memory removed (THE CONTROL) 15.6% memory verdict accuracy · the free desk band, no model, $0.00 32.8% memory verdict accuracy · 1 more measured on each run
the fast tier, with the carried state 100.0% verdict on watchlist · the same tier, memory removed (THE CONTROL) 32.5% verdict on watchlist · the free desk band, no model, $0.00 31.2% verdict on watchlist · always CLEAR -- the majority-class floor, $0.00 0.0% verdict on watchlist · 1 more measured on each run
the fast tier, with the carried state 0.0% tonnage error · the same tier, memory removed (THE CONTROL) 100.0% tonnage error · the free desk band, no model, $0.00: no headline metric, 1 measurement
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Decisively in one direction, and the kit says which comparison is doing the work.
THE STRONGER NO-MEMORY ARM IS THE FREE FLOOR, NOT THE STATELESS MODEL, on three of the six graders -- verdict 63.33 against 64.00, the memory subset 32.81 against 15.62, norm status 72.00 against 6.67. So the headline this kit publishes is measured against the FLOOR: verdict 100.00 pct against 63.33, a gap of 36.67 points; the 77 watchlist weeks 100.00 against 31.17, a gap of 68.83; the 64 memory-dependent cells 100.00 against 32.81, a gap of 67.19. The stateless control is published beside it because it isolates the SENTENCE rather than the SYSTEM, and the two answer different questions.
WHERE IT IS NOT SEPARABLE, SAID PLAINLY. On anomaly class the model and the free floor are EQUAL to two decimal places, 90.67 pct each, over 150 weeks. There is no separation at all on that grader and the only difference is which cells each gets wrong -- and those do not overlap on a single row. On norm status the floor BEATS the memory-removed model by 65.33 points. Two of six graders therefore do not support paying for anything.
⚑ ONE RUN PER ARM, AND ONE REPEAT PROBE BEHIND THE SCORED ARM. p002-ticket-anomaly-repeat re-ran 36 of the 150 weeks -- 12 routes spanning all nine patterns -- through the identical prompt and the identical carried state. 143 of 144 answered cells came back byte-identical, and 9 of the 10 graders reproduced to the digit: verdict, watchlist, week count, anomaly class, duplicate raises, ALREADY_OPEN recall, the memory subset, false alarms and the tonnage total all moved 0.00 points. ⚠︎ WHAT THAT IS AND IS NOT: it is two passes over 36 weeks, so it bands THOSE 36 and says nothing about the other 114, and one cell is the resolution limit -- on this subset a single cell is 2.78 points of verdict accuracy and 2.78 of norm status. A zero here means 'nothing moved on these weeks, asked twice', never 'the kit is stable'.
⚠︎ AND THE ONE CELL THAT DID MOVE IS THE KIT'S OWN KNOWN DEFECT, WHICH MAKES IT THE MOST USEFUL RESULT ON THIS PAGE. RTE-0043-W2, norm status, NORM_HOLDS on the first pass and NORM_SHIFTED on the second, from an identical prompt and an identical carried state. Both passes apply rule 3's CLOSED reset correctly to the verdict and the week count -- WATCH at 1 -- and the second one fails to carry that reset into the fourth field, reasoning instead that "the carried 3 prior LOW weeks mean the norm has shifted". So the un-shifting failure described in Eval.taxonomy is not a deterministic quirk of five particular rows: it is a coin toss on marginal cells, and r001 happened to win this one. Norm status is the only field that moved and it fell 2.77 points on the subset.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
A book where nothing has ever been disputed, every route's norm is current, and no route has been out of band two weeks running
the free desk band, at $0.00
With no streaks and no open items there is nothing to carry. The floor gets the verdict right on every in-band week and every first excursion for nothing, and it TIES the model on anomaly class (90.67 pct each) across the whole corpus.
Paying a model to compare a number against a band a spreadsheet already computes.
A weekly disposal recon queue that carries open disputes and re-normed routes -- the ordinary state of a real book after a quarter
this kit, with the carried state
100.00 pct verdict against the floor's 63.33 pct, 100.00 pct against 31.17 on the 77 watchlist weeks, and ZERO duplicate disputes against the floor's 27.
Reading the 100.00 pct verdict figure as a property of your book. It is a property of a synthetic corpus whose rule is printed on every page.
You need a confident cause code on every flagged week for a downstream workflow
neither, without deciding what a hedge means first
The two arms tie at 90.67 pct and fail on disjoint cells. The floor gives you a confident topical code where nobody recorded a cause; the model gives you CAUSE_UNKNOWN where a dispatcher hedged. Which is worse is your call and this page will not make it.
Treating the two 90.67 pct figures as interchangeable because they are equal.
Deciding whether to rebase a route's billed norm
neither, without a human step
NORM_SHIFTED is the highest-consequence output this kit has -- it recommends changing what a customer is billed against -- and it is the one field the model gets wrong, always in the direction of recommending a rebase that is no longer warranted.
Wiring norm_status to anything that writes a rate.
You have no history at all -- a new book, or a monitor being stood up
the free desk band for the first two weeks, then this kit
Rule 7 says a route with fewer than 2 recorded weeks has no observed norm, and the corpus's 15 new_route weeks are answered correctly by both arms. Memory is worth nothing until there is some.
Paying for a carried state that is empty.
Sizing what this costs on a real disposal book
read Cost.cost_at_10x before anything else
At $0.005219 a route-week on the projected card, a book of 20,000 routes checked weekly is about $5427 a year, and the free floor already gets 63.33 pct of the verdicts right. The honest deployment is the subset the floor cannot answer -- routes with an open item or a live streak -- not the whole book.
Running it over every route every week because the per-call price looks small.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
class_hedge_declined
A recorded reason that names a cause with a hedge, answered CAUSE_UNKNOWN
14
All 14 of the model's anomaly-class misses are this shape, and 9 of them are one sentence: 'a mis-post from the adjacent commercial route is suspected on the second ticket'. The model's own rationale on RTE-0031-W1 says it outright -- "the route note's…
norm_never_unshifts
A route back inside its band still reported NORM_SHIFTED
5
Every one of the model's 5 norm-status misses is the same cell: the final week of an already_open route, where the tonnage comes back INSIDE the band. Gold is NORM_HOLDS -- weeks out is 0, so rule 7's three-consecutive-week condition is not met. The model…
floor_keyword_overcommits
The FREE FLOOR's failure, published beside the model's because they are opposite
10
10 of the free desk band's 14 class misses answer a topical code where the key says CAUSE_UNKNOWN -- EXTRA_PULL x6, TARE_DRIFT x2, SERVICE_STEP x2. The corpus plants three reasons that name a cause and then say it was never established ('an additional pull is…
floor_duplicate_dispute
The free floor re-raises a dispute that is already open
27
27 of the 32 ALREADY_OPEN weeks. It is not a scoring artefact: each one is a second dispute filed against tickets that already carry one, which the facility rejects and somebody reworks. It also inflates the tonnage the desk reports as in dispute from 146.14…
stateless_undercalls
The memory-removed model answers WATCH where the key says RAISE or ALREADY_OPEN
34
The stateless arm files ZERO duplicates -- which reads as a clean sheet until you notice its ALREADY_OPEN recall is 0.00 pct. It cannot count past one week, so every escalation collapses to WATCH. Its per-pattern figures show it exactly: carried_streak 0.00…
stateless_norm_collapse
The memory-removed model refuses to assert a norm status and loses to a spreadsheet
140
Norm status falls from 96.67 pct to 6.67 pct when the carried state is removed -- BELOW the free desk band's 72.00 pct. Told it has no history, the model correctly answers NORM_UNKNOWN, which is the honest answer and the wrong one on the 108 of 150 weeks…
What we could NOT verify
⚠︎ RUN-TO-RUN STABILITY ON THE OTHER 114 WEEKS. p002-ticket-anomaly-repeat asked 36 of the 150 weeks a second time and 143 of 144 cells came back identical, so the nine graders that moved 0.00 points are banded ON THOSE 36 AND NOWHERE ELSE. Two passes is also n=2: it cannot separate a stable answer from an unstable one that happened to agree twice, and the single cell that did move (see Eval.taxonomy) is the evidence that at least one field is the second kind. ⚑ THE FIRST ATTEMPT AT THIS PROBE FAILED AND ITS RECORD IS KEPT: all 36 calls returned 402 Insufficient Balance from the shared provider account earlier the same day. It lives at kits/UC0088-ticket-anomaly/docs/attempted-p002-ticket-anomaly-repeat-402.json, deliberately NOT in results/ and NOT in the run register -- a result file of 36 transport failures carries a full set of zero scores with real denominators and would commit a record of zeros that a board reads as a clean run. The successful probe was written to its own path rather than over it, so the discarded run keeps its own token totals.
⚠︎ THE ANOMALY-CLASS KEY IS CONTESTED ON UP TO 14 OF 150 CELLS AND HAS NOT BEEN RE-SCORED. Every one of the model's class misses answers CAUSE_UNKNOWN on a note that names a cause with a hedge ('is suspected', 'carry the same tare as the 30-yard unit'). The prompt tells it to answer CAUSE_UNKNOWN where a note names a possible cause and then says nobody established it; the key treats a hedged attribution as an attribution. The obvious experiment -- re-score the committed run under the other reading, which costs nothing because the replies are on disk -- was not run, and the figure on this page is the strict-key one.
Whether rendering the carried state as ENGLISH rather than as JSON matters. src/state.describe builds three sentences; the obvious experiment is the same 150 weeks with the six scalars rendered as a JSON object, and it would cost one more 150-call run. Not paid for.
What disabling provider-side reasoning does. 93.6 pct of this run's output tokens were reasoning and output is the overwhelming majority of the bill, so the lever is obvious -- and src/adapters carries a documented THINKING_OFF shape it has never sent. Every published run left reasoning at the provider's default and the result files record thinking: null so a future run that DOES send it cannot be quietly compared against one that did not.
Whether a longer chain behaves. Three weeks is long enough to reach the norm-shift threshold and not long enough to test a route that has been out of band since spring. The carried state is constant-size by construction, so the cost curve is known; the ACCURACY curve on week twelve is not.
What a duplicate dispute actually costs a hauler in rework, and what a ton is worth at any real facility. The kit counts duplicates and reports tons; it prices neither.
Whether the four invented thresholds are anywhere near a real facility's. The band, the raise threshold, the norm-shift threshold and the 30-day dispute window were chosen for this kit and are the four things a forker must replace first.
How the kit behaves when a route disappears from the population entirely rather than going off the book for one week. Not in the corpus, and the carried state has no expiry -- a route pulled permanently keeps its open item for ever.
Whether the model would hold up on real dispatcher prose. Every recorded reason here is one of 24 generated sentences; a real route sheet carries abbreviations, half-sentences and internal codes.
What happens under an indirect prompt injection in the Route Notes. No attack was fired; see redteam_page.
Whether any of the six graders would survive a corpus with a different class balance. Three denominators are small (3 NOT_MONITORED cells, 9 flip, 15 closed_reset) and the whole distribution was chosen by the generator, not observed.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,779.48
1,442.92
6,679 ms
$0.005219
$0.002087
$0.089941
the same tier, memory removed (THE CONTROL)
1,730.75
1,099.16
6,115 ms
$0.004163
$0.001665
$0.072265
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same route-week, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
386 lines are on the shared ledger for this kit and 348 of them reached a provider: 6 calibration at an 8,000-token ceiling (c000, which lost one reply and is committed anyway because a ceiling that failed is a measurement), 6 more at 24,000 (c001), 150 scored, 150 the stateless control, and 36 the repeat probe. A further 38 lines -- a first attempt at that same probe, plus two single-call checks -- returned 402 Insufficient Balance from the shared account and were billed NOTHING; they are counted as discarded_calls at $0.00 because the ledger records a line before each request, and the successful probe was written to its own path rather than over the failed one so the discarded run keeps its own totals. Everything else -- both free floors, the wiring stub, the scorer and the eight-check pre-flight -- is pure code and costs $0.00. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING, 93.6 pct of output tokens (202520 of 216438) and therefore most of the bill. It was left at the provider's default and this kit has not measured what turning it off does to the answers.
HOW MANY RULES THE WEEK TRIPS, NOT HOW LONG THE SHEET IS. Every sheet is between 4008 and 4291 bytes -- within 7 pct of each other -- and output tokens ran from 268 to 14904, a 55.6x spread on a corpus of one shape. The long replies are the routes where a rebased norm, a live streak, a direction change and a CLOSED entry all have to be resolved before a verdict falls out.
WHETHER THE PROMPT CARRIES THE STATE -- AND HERE IT IS A COST, NOT A SAVING. Three sentences add 49 input tokens a week and 1.31x the output, because a model with a streak and an open item has four rules to reconcile and a model without them has one comparison to make. It costs 1.25x to be right.
THE FIXED HEAD OF THE PROMPT. The system message and the instruction are 2799 of 7404 characters -- 38 pct, byte-identical on all 150 calls. Add the weight-rule block reproduced on every sheet and the route's own numbers are much the smaller half of what is sent.
Your volumeWhat it costs at your volume
Linear in ROUTE-WEEKS, and flat in history length, which is the unusual half. Each week is one call whose input is one sheet plus six carried scalars, so a route on its fortieth weekly run costs exactly what one on its third costs. 150 route-weeks project to $0.7828 on the shared rate card, so ten times the set is about $7.83.
⚠︎ AND A REAL DISPOSAL BOOK IS WHERE THIS ARITHMETIC STOPS BEING COMFORTABLE. At $0.005219 a route-week, 20,000 routes checked every week is about $104 a week and $5427 a year -- every year -- and the free desk band already gets 63.33 pct of the verdicts right for nothing. The honest reading is that this belongs on the SUBSET the floor cannot answer: routes carrying an open item or a live streak. In this corpus that is 64 of 150 weeks, which would cut the bill by about 57 pct.
WHAT DOES NOT SCALE IS WALL CLOCK WITHIN A ROUTE: weeks are strictly serial. 50 chains ran 12-wide in 194.4 seconds; widening the pool is the only lever, and the depth is fixed at however many periods you replay.
Where pricing changes shape
THE OUTPUT CEILING IS A CLIFF AND THIS KIT HAS BEEN OVER IT -- ON A CALIBRATION RUN, WHICH IS WHERE YOU WANT TO FIND IT. c000-ticket-anomaly-calibration fired the two hardest routes at a 8,000-token ceiling and LOST ONE OF SIX: RTE-0041-W2 spent the whole budget and returned nothing. c001 re-fired the identical six at 24,000 and got them all through at a 6651-token maximum. The published ceiling is 24000 and the scored run's largest reply reached 14904, 62 pct of it. A ceiling is not billed -- the provider charges tokens produced, not tokens allowed -- so the weeks that finish in a few hundred tokens pay nothing for the headroom; a call that HITS it is billed in full and returns nothing at all.
THE MEMORY STEP IS A SURCHARGE ON THIS TASK, NOT A DISCOUNT. Adding the carried state moves the price per route-week from $0.0041629 to $0.0052185 -- 1.25x -- effectively all of it in output tokens. It buys 36.00 points of verdict accuracy and 27 avoided duplicate disputes, and the kit states both halves rather than only the flattering one.
Your return, with your numbers
Volumeroute-weeks per scheduled run -- this run judged 150 (50 routes x 3 weekly runs) in one pass per arm
What it replacessomebody opening every open route, adding up the week's scale tickets, comparing them against the norm they believe is in force, and then going back through the exception queue to work out whether the route has been drifting and whether a dispute is already sitting with the facility
Time saved per itemnot measured here -- depends on how long reconstructing a route's dispute history takes in the reader's own disposal system
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start a question whose honest answer might be 'free code already does most of this', which on the anomaly-class field it does, exactly as well as the model.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
266,922input tokens · this run
216,438output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 150 route-weeks, one completion call each, one tier. The 150-call stateless control and the 12 calibration calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.313
$0.313
$2.09
2026-09-12
gemini-3-flash
Google
$0.783
$0.783
$5.22
2026-09-18
gemini-3-8-flash
Google
$1.012
$1.012
$6.75
2026-09-18
llama-5
Meta
$1.254
$1.254
$8.36
2026-09-18
claude-haiku-4-5
Anthropic
$1.349
$1.349
$8.99
2026-09-12
grok-4-5
xAI
$1.832
$1.832
$12.22
2026-09-18
grok-4-6
xAI
$1.832
$1.832
$12.22
2026-09-18
claude-sonnet-5
Anthropic
$2.698
$2.698
$17.99
2026-09-12
gemini-3-1-pro
Google
$3.131
$3.131
$20.87
2026-09-18
gpt-5-6-terra
OpenAI
$3.131
$3.131
$20.87
2026-09-12
gpt-5-6-sol
OpenAI
$5.396
$5.396
$35.98
2026-09-12
claude-opus-4-8
Anthropic
$6.746
$6.746
$44.97
2026-09-12
claude-opus-5
Anthropic
$6.746
$6.746
$44.97
2026-09-12
claude-fable-5
Anthropic
$13.491
$13.491
$89.94
2026-09-18
claude-fable-5-1
Anthropic
$13.491
$13.491
$89.94
2026-09-18
gpt-6-astra
OpenAI
$13.491
$13.491
$89.94
2026-09-17
Read this against the numbers above
A projection onto published rate cards, not a bill. Nobody paid any of these figures.
Rates checked on 2026-08-18. A vendor moving a price makes every row here stale and nothing in this kit notices.
Token counts are this model's tokenizer on this corpus. Another vendor's tokenizer will not agree exactly, and the difference is largest on the ticket tables.
Output is 83 pct of the projected bill, and 93.6 pct of the output was provider-side reasoning. A vendor that bills reasoning differently changes every row.
Nothing here says another model would ANSWER as well. Only one model was run; the rows price the same token counts elsewhere and price nothing about quality.
The provider that actually ran the calls is not in the table, so no row is the real spend.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
14 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator
Generates 50 routes x 3 consecutive weekly monitor sheets = 150 documents from a fixed seed (SEED = 20260823), across nine planted patterns, and writes the answer key by running src/norms.replay over what it planted. Nothing is labelled by hand.
tools/build_corpus.py
# Generate the 50-route, 150-week disposal weight corpus and its computed answer key.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
RULE = "-" * 64
FACILITIES = ["Northbay Transfer Station", "Cedar Ridge Landfill", "Pike Street Transfer",
CONTAINERS = ["4 yd front-load", "6 yd front-load", "8 yd front-load", "20 yd roll-off",
DAYSETS = [("Mon", "Thu"), ("Tue", "Fri"), ("Wed",), ("Mon", "Wed", "Fri"), ("Tue", "Thu")]
src/norms.pythe weight rule and its arithmetic — a swap seam
The band, the streak, the reset, the already-open rule and the norm-shift rule, as pure code. It is the answer key's source and the runtime never calls it to decide anything the model is asked.
You change it to: The band, the raise threshold, the norm-shift threshold and the dispute window are four constants in one file. Change them and re-run tools/build_corpus.py; the answer key follows automatically because it is computed from them.
src/norms.py
# The disposal weight rule, and the arithmetic it implies. Pure code, no model.
BAND_PCT = 15
RAISE_AT = 2
SHIFT_AT = 3
HISTORY_FOR_NORM = 2
DISPUTE_WINDOW_DAYS = 30
CLEAR = "CLEAR"
WATCH = "WATCH"
RAISE = "RAISE"
ALREADY_OPEN = "ALREADY_OPEN"
src/state.pythe carried state — a swap seam
Six scalars between runs -- the norm in force, the consecutive-week count, the direction, whether an anomaly is open, how many weeks the route has been monitored, and the previous verdict -- plus the English they become in the prompt. Written from the arithmetic, never from a model reply.
You change it to: Six scalars and one describe() that turns them into English. Adding a seventh is one field and one clause -- and costs a handful of input tokens a week, not a growing transcript.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_route(store, route_id, book_norm):
def describe(state, book_norm=None):
src/segment.pythe section splitter
Splits a sheet into its seven named sections. Asserted over all 150 documents by evals/check_labels.py before a run may spend.
src/segment.py
# Split a weekly monitor sheet into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Route", "Weight Rule", "Scale Tickets", "Open Items",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section filter — a swap seam
Decides what is sent. Customer Contact is mapped by no field and is subtracted unconditionally, so it never reaches a provider on any document, for any field, including through the fallback.
You change it to: SECTION_HINTS plus NEVER_SENT. Withholding another section is one tuple entry and evals/check_labels.py re-proves it over all 150 documents.
src/select.py
# Pick which sections of a weekly monitor sheet are sent. Pure code -- the last deterministic step
BANNER = "Synthetic Record"
ROUTE = "Route"
RULEBLOCK = "Weight Rule"
TICKETS = "Scale Tickets"
OPENITEMS = "Open Items"
NOTES = "Route Notes"
CONTACT = "Customer Contact"
NEVER_SENT = (CONTACT,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts: the question and the JSON shape, the carried-state block, the sheet. The stateless build replaces the middle block and changes nothing else -- asserted, not asserted-by-comment.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, book_norm=None, stateless=False):
src/monitor.pyone weekly run
One sheet, one call, one parse. Holds MAX_TOKENS = 24000, the published ceiling.
src/monitor.py
# One weekly run over one route, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You apply a written disposal weight rule to one waste route for one service week. "
MAX_TOKENS = 24000
FIELDS = ("verdict", "weeks_out_of_band", "norm_status", "anomaly_class")
def documents():
def routes():
def load_doc(doc_id):
def position_of(text):
src/adapters/__init__.pythe provider adapter — a swap seam
Raw HTTP over stdlib for any OpenAI-compatible provider or Anthropic. Retries transient statuses and transport failures; never retries a 4xx that will fail identically.
You change it to: .env plus the same run again. Any OpenAI-compatible provider or Anthropic.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/budget.pythe spend guard
Counts calls, not dollars, against a cap shared by every kit under the repo root. Checked once per completion, before the request.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
evals/baseline.pythe free floors
Two of them: the +/-15 pct desk band computed from ONE sheet, and the majority-class always-CLEAR arm. Both are pure code and cost $0.00.
evals/baseline.py
# The free, no-model floors: the whole answer computed from ONE sheet, with no memory.
DESK_RAISE_DEVIATION_PCT = 30
def read_sheet(text):
def deskband(text):
def allclear(_text):
MODES = {"deskband": deskband, "allclear": allclear}
def review(text, mode="deskband"):
evals/scoring.pythe scorer
Six graders, pure code, no judge. Verdict / week count / norm status / class per week; watchlist and false alarms counted apart; duplicate raises counted on their own; the memory subset; the tonnage reconciliation.
evals/scoring.py
# Score a run against the computed gold. Pure code, no model, no judge.
FIELDS = ("verdict", "weeks_out_of_band", "norm_status", "anomaly_class")
QUIET = tuple(v.upper() for v in N.QUIET_VERDICTS)
def _pct(n, d):
def _norm(v):
def _as_int(v):
def score(records, golds):
def _by_pattern(vrows):
def compare(stateful, stateless):
evals/check_labels.pythe pre-flight
Eight checks over the shipped text, run before any run may spend, with a --self-test that corrupts the corpus four ways and requires convictions.
evals/check_labels.py
# Everything this kit ASSERTS about its own corpus, checked before a run is allowed to spend.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def gold_rows():
def sections_present(docs, rows):
WITHHELD = ("Customer Contact",)
def contact_never_sent(docs, rows):
def streak_not_on_the_sheet(docs, rows):
def gold_reproduces(docs, rows):
def sequences_complete(docs, rows):
evals/run.pythe harness
Routes run concurrently; the three weeks inside a route are strictly serial, because week 3's prompt contains a streak week 2 produced.
evals/run.py
# Run the weekly monitor over the 50 routes and write one result file. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def load_gold():
def observation(g):
def opening_state(route_rows):
def stub_complete(cfg, system, user, max_tokens=1024):
def main():
src/app.pythe local UI
http.server on port 9005. Renders with no key and replays the committed r001 answer for the selected sheet, labelled as a replay, beside the free floor.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9005"))
def _gold():
GOLD_ROWS = _gold()
RECORDED_RUN = "r001-ticket-anomaly"
def _recorded():
RECORDED = _recorded()
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 50 routes x 3 consecutive weekly monitor sheets = 150 documents from a fixed seed (SEED = 20260823), across nine planted patterns, and writes the answer key by running src/norms.replay over what it planted. Nothing is labelled by hand.
src/norms.pyThe band, the streak, the reset, the already-open rule and the norm-shift rule, as pure code. It is the answer key's source and the runtime never calls it to decide anything the model is asked. A swap seam.
src/state.pySix scalars between runs -- the norm in force, the consecutive-week count, the direction, whether an anomaly is open, how many weeks the route has been monitored, and the previous verdict -- plus the English they become in the prompt. Written from the arithmetic, never from a model reply. A swap seam.
src/segment.pySplits a sheet into its seven named sections. Asserted over all 150 documents by evals/check_labels.py before a run may spend.
src/select.pyDecides what is sent. Customer Contact is mapped by no field and is subtracted unconditionally, so it never reaches a provider on any document, for any field, including through the fallback. A swap seam.
src/prompt.pyThree parts: the question and the JSON shape, the carried-state block, the sheet. The stateless build replaces the middle block and changes nothing else -- asserted, not asserted-by-comment.
src/monitor.pyOne sheet, one call, one parse. Holds MAX_TOKENS = 24000, the published ceiling.
src/adapters/__init__.pyRaw HTTP over stdlib for any OpenAI-compatible provider or Anthropic. Retries transient statuses and transport failures; never retries a 4xx that will fail identically. A swap seam.
src/budget.pyCounts calls, not dollars, against a cap shared by every kit under the repo root. Checked once per completion, before the request.
evals/baseline.pyTwo of them: the +/-15 pct desk band computed from ONE sheet, and the majority-class always-CLEAR arm. Both are pure code and cost $0.00.
evals/scoring.pySix graders, pure code, no judge. Verdict / week count / norm status / class per week; watchlist and false alarms counted apart; duplicate raises counted on their own; the memory subset; the tonnage reconciliation.
evals/check_labels.pyEight checks over the shipped text, run before any run may spend, with a --self-test that corrupts the corpus four ways and requires convictions.
evals/run.pyRoutes run concurrently; the three weeks inside a route are strictly serial, because week 3's prompt contains a streak week 2 produced.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1779 input and 1442 output tokens per route-week, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Route-weeks/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per route-week directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py and a fixed pool of 24 generated sentences. The Route Notes are the field an outside party WOULD influence in a real deployment -- a dispatcher types them, and on a commercial book a customer can dictate what goes in them -- and this kit sends them deliberately rather than hiding its own injection surface. The one section that is never sent is Customer Contact, and that is a subtraction in code rather than a hint that happens not to match.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never written to a result file, a log line, a screenshot or a page. src/app.py redacts the key and the base URL out of any error string before returning it. src/config.save writes .env at 0600 with os.open BEFORE the file exists, so it is never world-readable even for an instant. No credential has ever been in this repository.
The experimentWe did not attack it -- and the boundary that matters most on this vertical was proven by breaking it on purpose
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus has none: 24 generated sentences, one generator, one seed. So instead of an attack trial the kit red-proves the boundary it CAN prove. evals/check_labels.py --self-test empties src/select.NEVER_SENT and requires the privacy check to convict; it leaks a week count and an open-item line onto a sheet and requires the leak scan to convict; it corrupts a gold label and flattens the memory flag and requires the replay check to convict. Four of four. The first version of that self-test scored three of four, and the one that failed was the privacy check -- because it tested membership of the very tuple the test empties, so emptying it made the check vacuously pass while the section sailed into every prompt. The red-proof found the checker, not the corpus. Confirmed by assertion and by reading the recorded run, not by an attack trial, on 2026-08-23 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does a customer's name and service address ever leave the machine?
Every weekly sheet carries a Customer Contact section -- the account holder, the service address, a telephone number and an email. On a commercial disposal book that is which premises produce what tonnage on which days, which is competitive information about the customer.
src/select.NEVER_SENT holds it and _fallback() subtracts it unconditionally, so a hint that matches nothing falls back to the sheet MINUS that section rather than to the whole sheet. evals/check_labels.py checks the union AND every per-field fallback AND a deliberately unmatchable field name, over all 150 documents, and separately asserts the account holder's own name does not appear anywhere in the assembled prompt.
Can the kit dispute a ticket, issue a credit or change a customer's billed norm?
A monitor that recommends RAISE and NORM_SHIFTED is one integration away from filing the dispute and re-norming the route, and the integration is the tempting part.
There is no such code path and no configuration flag that adds one. The only writers anywhere in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). It is a property of what is ABSENT, which is why it is stated at the call site in src/monitor.py as well as here.
Can a wrong model reply corrupt the next week's answer?
A monitor that fed its own verdict back into its next prompt would compound one bad week into every week after it -- and on this task both directions have teeth: a spurious RAISE opens an item that turns every later week into a false ALREADY_OPEN, and a missed RAISE leaves the item shut so every later week re-raises the same dispute.
src/norms.step writes the next state from the ARITHMETIC over the sheet's own totals and flags. evals/run.py advances the state from norms.step on every week including the ones whose CALL failed, and never from r['answer']. The model's reply is scored and discarded.
Can a run spend more than expected?
One key funds every kit in this repository. A loop with a bad exit condition, a --limit left off, or a run started twice in two terminals, and the first sign is the provider's dashboard tomorrow.
src/budget.py counts calls against a cap read from the shared .env, checked ONCE per completion inside src/adapters.complete -- the single line every caller goes through, including the app and any script a forker writes. The harness prints what it is about to spend before it spends it. When no cap is configured it says so out loud rather than being silent.
Can a corpus edit quietly deflate the finding?
Every claim on this page about memory is a claim about what 150 generated sheets do NOT say. Leak a streak onto a sheet and the stateless arm reads it, the gap closes, and no gate anywhere notices.
evals/check_labels.py searches every non-rule section of all 150 documents for a stated week count, a stated open item, a stated norm in force or a stated monitor age, and evals/run.py refuses to spend if it is red. --self-test plants each leak and requires a conviction.
Each boundary above was checked by running an assertion or by reading the recorded run, not by an attack trial. The privacy boundary was red-proven by emptying the guard and watching the check convict -- and that red-proof caught the check itself being tautological, which is the whole reason to red-prove one.
The result0 attack trials, five boundaries checked -- and the privacy boundary red-proven by removing the guard and watching the check convict, which is also how the check's own tautology was found.
0untrusted input fields on this corpus
0attack trials fired
5boundaries checked by assertion or by reading a run
4corpus corruptions the pre-flight convicts, red-proven
The Route Notes ARE the field an outside party would influence in a real deployment and this kit sends them, deliberately, rather than hiding its own injection surface. On the shipped corpus they are drawn from 24 generated sentences, so there is nothing adversarial to measure; what that surface does under attack is unmeasured and is the first thing to fire if this is pointed at a real book.
Read this twice
The carried state is written from the arithmetic, never from the model's reply, and on this kit that matters in both directions. The quantity being carried is a COUNT and a FLAG. One spurious RAISE opens an item that turns every following week into a false ALREADY_OPEN; one missed RAISE leaves the item shut so every following week re-raises the same dispute. Feed the model's own verdict forward and a single bad week becomes a permanent one. src/norms.step advances the state from the sheet's own totals and flags, and evals/run.py advances it even on weeks whose call FAILED -- so a transport error costs one week's answer and never three.
HonestyWhat this does not prove
Whether a real deployment's Route Notes -- prose a dispatcher writes freely, and on a commercial book prose a customer can dictate -- would carry an instruction the model follows. Not applicable to the shipped corpus, and the first thing to attack if this is pointed at a real one.
Whether the provider retains prompts, and for how long. Outside the kit and unmeasured.
Whether a forker's own sheets carry personal data in a section this kit DOES send. The guard names one section; a real sheet might put a contact name in the Route Notes.
What the model does with a sheet whose sections are missing or reordered. evals/check_labels.py asserts all seven are present in this corpus, so the degraded path is guarded rather than measured.
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No dispute, no credit and no re-norm, non-configurable. This kit produces a verdict, a consecutive-week count, a norm status and a class code, and stops. Nothing in src/, evals/, tools/ or ui/ files a dispute with a facility, issues a credit memo, adjusts an invoice or writes a route's norm into a book of record.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json), tools/build_corpus.py (data/) and src/state.save (data/state.json). There is no HTTP client anywhere except src/adapters, which speaks to one completions endpoint.
EvidenceDoes it hold?
What
Measured
Nothing in this kit disputes a ticket, credits an invoice or re-norms a route
0 code paths. The kit has no client for any system of record and no endpoint that takes a write; src/app.py serves four GETs and one POST, and the POST returns a verdict.
The carried state never contains a model reply
0 assignments from r['answer'] into the state. evals/run.py advances the state only from norms.step over the sheet's own totals and flags, on every week including the 150 whose call could have failed and did not.
The withheld section never reaches a provider
0 of 150 documents, checked in three directions -- the union, every per-field fallback, and a field name that matches nothing -- plus a search for the account holder's name in the assembled prompt.
A run cannot start over a corpus that has drifted from its key
8 pre-flight checks, all green on the shipped corpus, run inside evals/run.py before any spend. 4 of 4 red-proven by --self-test.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The carried state is correct whatever the model says, which means a wrong verdict is reported wrongly for that week and nothing catches it -- see Eval.taxonomy's norm_never_unshifts row, where five weeks recommend a rebase that is no longer warranted and the state machine is entirely happy.
IT IS NOT AN APPROVAL WORKFLOW. There is no reviewer, no queue and no sign-off. The output is a watchlist row; who reads it and what they do next is outside the kit.
IT IS NOT A SCHEDULER. The kit is invoked, not woken. Nothing here fires the weekly run, nothing notices if it does not happen, and nothing alerts on a missed week -- which is a real gap on a monitor whose whole value proposition is a cadence. See Architecture.breaks_at_scale.
IT IS NOT A DEFENCE AGAINST A POISONED SHEET. If the Open Items log falsely records an anomaly CLOSED, the streak resets, the open item clears, and the kit reports WATCH on a route that should be ALREADY_OPEN. Nothing detects it and there is no second source.
IT IS NOT A LIMIT ON SPEND BEYOND CALL COUNT. src/budget.py counts calls, not dollars, because the kit does not know the rate card a forker is pointed at.
WatchedWhat is watched, and why that one
8runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 37 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
16 measured by the latest run21 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
alarm
verdict_accuracy_pct; duplicate_raise_rate_pct; anomaly_class_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0, and duplicate_raise_rate_pct above 0. A week that returns nothing scores as a miss in all four fields, so a reliability failure arrives disguised as a quality one.
duplicate-raise
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
alarm
duplicate_raise_rate_pct; already_open_recall_pct — alarm on duplicate_raise_rate_pct above 0 on any arm that carries state -- it means the open-item flag stopped reaching the prompt.
memory-subset
The cells whose answer cannot be derived from the week in front of you
alarm
memory_verdict_accuracy_pct; memory_weeks_accuracy_pct — alarm on the subset shrinking below 20 cells -- evals/check_labels.py fails the build rather than publishing a rate with no denominator.
two-directions
Watchlist recall and false alarms, counted apart and never summed
alarm
verdict_on_watchlist_pct; false_alarm_rate_pct; missed_watchlist — alarm on missed_watchlist above 0.
tonnage-reconciliation
The tonnage this run would put in dispute with the facility
alarm
tonnage_error_pct; tonnage_reported_raised — alarm on tonnage_error_pct moving while the per-week figures do not -- that is cancellation, not accuracy.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
629,038
weekly monitor sheets edited — the count held, the bytes did not
split.count
50
the routes count moved — a different set was scored
split.size_p50
3
the median size of one route moved
split.size_p95
3
the 95th-percentile size of one route moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (already_open_cells 32, documents 150, memory_cells 64, routes 50, stateless False, watchlist_cells 77, weeks_scored 150) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
verdict accuracy
0.00 points over 2 passes on a 36-week subset
150 route-weeks; the repeat covered 36
MEASURED, on a subset. p002-ticket-anomaly-repeat re-asked 36 of the 150 weeks with an identical prompt and an identical carried state; r001 and p002 both score 100.00 pct on those 36. ⚠︎ IT BANDS THOSE 36 AND NOT THE OTHER 114, and n=2 cannot separate a stable answer from an unstable one that agreed twice. One cell is the resolution limit: 2.78 points on this subset.
watchlist recall
0.00 points over 2 passes on a 36-week subset
77 weeks a clerk must touch; the repeat covered 23 of them
MEASURED on the subset: 100.00 pct on both passes.
duplicate raises
0.00 points over 2 passes on a 36-week subset
32 ALREADY_OPEN weeks
MEASURED on the subset: 0 duplicates on both passes, over 11 ALREADY_OPEN weeks of the 32. The free floor is 84.38 pct on the full set. ⚠︎ TWO PASSES IS NOT PROOF THE ZERO HOLDS -- it is proof it held twice on these weeks.
anomaly class
0.00 points over 2 passes on a 36-week subset
150 route-weeks
MEASURED on the subset: 97.22 pct on both passes. The key on this field is separately contested on up to 14 cells -- a reproducible number can still be scored against a debatable label.
norm status
2.77 points over 2 passes on a 36-week subset -- THE ONLY FIELD THAT MOVED
150 route-weeks; the repeat covered 36
MEASURED, and it is the one band on this page that is not zero. r001 scored 94.44 pct on the 36 re-asked weeks and p002 scored 91.67. The whole gap is ONE CELL: RTE-0043-W2 answered NORM_HOLDS on the first pass and NORM_SHIFTED on the second, from an identical prompt and an identical carried state. It is the same un-shifting defect Eval.taxonomy names, which means that defect is a coin toss on marginal cells rather than five fixed rows.
consecutive weeks out of band
0.00 points over 2 passes on a 36-week subset
150 route-weeks; the repeat covered 36
MEASURED on the subset: 100.00 pct on both passes, MAE 0.00 weeks and a worst error of 0 on each. It is the field the whole carried state exists to make answerable -- the free floor and the memory-removed arm both cap out at 1 week by construction.
the memory subset
0.00 points over 2 passes on a 36-week subset
64 memory-dependent cells
MEASURED on the subset: 100.00 pct on both passes over 20 memory-dependent cells of the 64. The subset DEFINITION was separately corrected once -- see Data.breaks_on.
false alarms
0.00 points over 2 passes on a 36-week subset
73 quiet weeks
MEASURED on the subset: 0 of 13 quiet weeks on both passes. The full run is 0 of 73.
tonnage reconciliation
0.00 points over 2 passes on a 36-week subset
one portfolio total over 150 weeks
MEASURED on the subset: 0.00 pct error on both passes. Errors cancel in this grader, so this is the least informative band on the page and is published only so that no row here reads as unmeasured when it is not.
latency
p50 8650 to 8784 ms, p95 62904 to 121524 ms over 2 passes on the SAME 36 weeks
150 calls
MEASURED, and it is wide. Both sides restricted to the same 36 weeks, because latency is per call and differencing a 150-week run against a 36-week probe would band the SUBSET rather than the model. The spread is provider reasoning, not network: the answers were identical on 143 of 144 cells over those weeks while the clock was not.
tokens in
0 tokens over 2 passes
150 calls
MEASURED: deterministic given the corpus and the prompt, and it did not move. The input is assembled by src/prompt.py from committed text and a computed state, so a change here would mean the corpus or the builder moved.
tokens out
60588 to 83604 tokens over 2 passes on the SAME 36 weeks -- a 38 pct swing
150 calls
MEASURED, and it is the widest band on the page. 93.6 pct of output is provider-side reasoning, and the two passes produced near-identical ANSWERS at materially different token cost -- which is the honest reason this kit's price per route-week is a projection and not a quote.
unparsed replies
0 across every scored call at the published ceiling
336 calls -- r001, s001 and the p002 repeat
Measured: 0 failures on r001, 0 on s001 and 0 on the p002 repeat at the published 24000-token ceiling, over 336 calls. c000 lost 1 of 6 at 8,000.
the free floor
exact
150 route-weeks
Pure code over a fixed corpus. It is deterministic and would reproduce to the digit on any machine.
the majority-class floor
exact
150 route-weeks
One line of code over a fixed corpus: 46.67 pct.
the pre-flight
8 of 8
8 checks over 150 documents
Deterministic. 4 of them red-proven by --self-test.
HistoryRun history
8 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 2 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · with the model in the path — 7 runs. Columns here are only ever compared with each other.
Metric
b000-ticket-anomaly-deskband 2026-08-23
b001-ticket-anomaly-allclear 2026-08-23
c000-ticket-anomaly-calibration 2026-08-23
c001-ticket-anomaly-calibration 2026-08-23
p002-ticket-anomaly-repeat 2026-08-23
r001-ticket-anomaly 2026-08-23
s001-ticket-anomaly-stateless 2026-08-23
already open recall, %
0.0
0.0
100.0
100.0
100.0
100.0
0.0
anomaly class accuracy, %
90.67
48.67
83.33
100.00
97.22
90.67
86.67
duplicate raise rate, %
84.38
0.00
0.00
0.00
0.00
0.00
0.00
false alarm rate, %
2.74
0.00
—
—
0.00
0.00
2.74
input tokens, whole run
0
0
10842
10842
64244
266922
259612
model latency p50 ms
0.00
0.00
8672.00
12605.00
8650.00
6679.00
6115.00
model latency p95 ms
0.00
0.00
25164.00
52414.00
121524.00
50568.00
24530.00
memory verdict accuracy, %
32.81
10.94
100.00
100.00
100.00
100.00
15.62
memory weeks accuracy, %
15.62
10.94
100.00
100.00
100.00
100.00
15.62
missed watchlist, %
2.6
100.0
0.0
0.0
0.0
0.0
2.6
norm status accuracy, %
72.00
72.00
83.33
83.33
91.67
96.67
6.67
output tokens, whole run
0
0
15803
15083
83604
216438
164874
tonnage error, %
166.33
100.00
0.00
0.00
0.00
0.00
100.00
verdict accuracy, %
63.33
46.67
83.33
100.00
100.00
100.00
64.00
verdict on watchlist, %
31.17
0.00
83.33
100.00
100.00
100.00
32.47
weeks out accuracy, %
64.00
48.67
83.33
100.00
100.00
100.00
64.00
not a time series No two of these 7 runs measured the same system — they differ on already_open_cells, baseline, documents, max_tokens, memory_cells, provider, routes, stateless, thinking, watchlist_cells, weeks_scored, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-ticket-anomaly-stub 2026-08-23
already open recall, %
0.0
anomaly class accuracy, %
90.67
duplicate raise rate, %
84.38
false alarm rate, %
2.74
input tokens, whole run
266717
model latency p50 ms
0.00
model latency p95 ms
0.00
memory verdict accuracy, %
32.81
memory weeks accuracy, %
15.62
missed watchlist, %
2.6
norm status accuracy, %
72.0
output tokens, whole run
6252
tonnage error, %
166.33
verdict accuracy, %
63.33
verdict on watchlist, %
31.17
weeks out accuracy, %
64.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 16 chips that all say so.
DeviationsWhat deviated
0 breaches across 8 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
r001-ticket-anomaly against s001-ticket-anomaly-stateless. Same corpus, same model, same grader, same 150 weeks, prompts byte-identical except the carried-state block -- asserted by evals/check_labels.py, not assumed.
b000-ticket-anomaly-deskband against r001-ticket-anomaly, same 150 weeks, same scorer. The class row is a TIE and is listed as one.
answering CLEAR to everything
verdict 46.67 pct, watchlist 0.00 pct, false alarms 0.00 pct, norm status 72.00 pct -- the same as the desk band
measured
b001-ticket-anomaly-allclear. It is on the page so that nobody reads a verdict percentage without knowing what a single line of code scores.
the output ceiling
8,000 -> 24,000 turns 1 lost reply of 6 into 0 of 6 on the two hardest routes; the scored run's largest reply is 14904, 62 pct of the published cap
measured
c000-ticket-anomaly-calibration against c001-ticket-anomaly-calibration, identical routes, identical prompt, ceiling the only difference.
provider-side reasoning
unknown -- 93.6 pct of output tokens and therefore most of the bill, and no run has turned it off. Listed as REASONING rather than measured: the token share is a measurement, the effect of removing it is not.
reasoning
src/adapters carries a documented THINKING_OFF shape it has never sent. Every result file records thinking: null so a future run cannot be quietly compared against these.
rendering the carried state as JSON instead of English
unknown. Listed as REASONING: nothing here has run the other rendering.
reasoning
src/state.describe builds three sentences. The experiment is one more 150-call run and was not paid for.
chain length
unknown beyond three weeks. The COST half IS measured -- input per week is 1779 tokens with the state against 1731 without, flat across routes carrying up to four weeks of streak -- but the ACCURACY curve on week twelve is not, and the edge is typed for the half that is unknown.
reasoning
r001 and s001 token totals over 150 weeks bound the cost. No chain longer than three weeks exists anywhere in this kit, so the accuracy claim is reasoning from a constant-size state, not a measurement.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
verdict accuracy
any movement at all on a third pass -- the two passes so far are identical
watchlist recall
any movement on a third pass
duplicate raises
any duplicate at all on any pass
anomaly class
any movement on a third pass
norm status
this one already has -- treat norm_status as the unstable field and never wire it to anything that writes a rate
consecutive weeks out of band
any movement on a third pass
the memory subset
any movement on a third pass
false alarms
any false alarm on any pass
tonnage reconciliation
movement here while the per-week figures hold -- that is cancellation, not accuracy
latency
nothing -- latency is not gated here, and a page that alarmed on it would alarm on the provider's mood
tokens in
any change at all
tokens out
nothing -- but it is why cost_at_10x is stated as a range in prose rather than as a single figure anybody should budget against
unparsed replies
anything above 0 -- a week that returns nothing is scored as a miss in all four fields, so a reliability failure arrives disguised as a quality one
the free floor
any change at all -- it would mean the corpus or evals/baseline.py moved
the majority-class floor
any change at all -- it is the corpus's own class balance
the pre-flight
any red check -- evals/run.py refuses to spend
NextThe three you would add first
A human step in front of anything that changes a customer's billed normNORM_SHIFTED is the highest-consequence output this kit has -- it recommends changing what a customer is billed against -- and it is the one field the model gets wrong, always in the direction of recommending a rebase after the drift has stopped. 5 of 150 weeks.
A scheduler, and an alarm for a run that did not happenThe facility's ticket-correction window is 30 days from the tip date. A weekly monitor that silently misses two weeks has lost the cheap remedy on that tonnage, and nothing in this kit would say so.
A second source for the open-item flagThe whole ALREADY_OPEN answer rests on one boolean this kit carries itself. If it drifts from the facility's own dispute queue, the kit is confidently wrong in whichever direction it drifted, and neither the model nor the scorer can tell.
An expiry on carried stateA route pulled permanently keeps its open item for ever. Not in the corpus, not handled, and it would surface as a route that can never be raised again.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/norms.py, src/segment.py or src/select.py -- it is eight checks over 150 documents and takes under a second. Re-run the scored eval (150 calls) only when the prompt, the model or the corpus changes, and re-run the stateless control WITH it: the headline on this page is a difference between two arms, and half a difference is not a number. The two free floors cost nothing and should be re-run with every corpus change, because they are what tells you whether the corpus got easier.
⚠︎ AND THE KIT'S OWN CADENCE IS NOT THIS ONE. The thing being described above is how often to re-MEASURE. How often the monitor RUNS is weekly, because the unit is a service week and the facility's correction window is 30 days from the tip date -- a week reviewed after that shuts becomes a credit memo against an invoice the customer already has. Nothing in this kit schedules that run or notices a missed one.
What this cannot tell you
Whether the 36.67-point verdict gap against the free floor is stable across repeats OF THE FULL SET. p002 re-asked 36 of the 150 weeks and the verdict figure did not move; the other 114 have been answered once.
Whether the 0 duplicate-raise rate holds beyond two passes. It held on both, over 11 of the 32 ALREADY_OPEN weeks. Two is not many.
Whether the FLOORS or the CONTROL are stable. Neither was repeated: the free floors are pure code and would reproduce to the digit by construction, but s001 is a model arm with one pass behind it, so the memory gap is banded on only one of its two sides.
Whether the guardrail holds on a forked kit. It is a property of absent code, and a forker who adds a facility client removes it without touching anything this kit checks.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, sys, time, random, datetime, argparse, collections, concurrent.futures, urllib, http.client and http.server, every one of them in the standard library. requirements.txt names nothing and says why.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior runs)
those carry a TRANSCRIPT; this carries six SCALARS written by arithmetic. The whole reason the cost curve is flat in history length is that nothing here grows with the number of runs -- measured: 1779 input tokens a week with the state against 1731 without, on routes carrying up to four weeks of streak.
the provider call
src/adapters/__init__.py
an LLM client library or a router (LiteLLM, OpenRouter, a vendor SDK)
one function per wire format, 40 lines each, plus a retry loop that distinguishes a transient status from a terminal one and a transport failure from both. A router would add a dependency and a second place for the key to live.
the scheduled run
-- there is no file
a scheduler or orchestrator (Airflow, Dagster, Temporal, cron)
⚠︎ THIS IS THE ONE PLACE THE KIT IS GENUINELY INCOMPLETE FOR ITS PATTERN. A monitor is defined by its cadence and this one is INVOKED, not woken. There is no scheduler, no retry policy for a run that did not happen, and no alarm on a missed week -- and a missed week costs the facility's 30-day correction window on that tonnage. Adding one is outside a kit by design (no deployment machinery), which means the gap is real and is stated rather than closed.
the scorer
evals/scoring.py
an eval framework (Braintrust, promptfoo, DeepEval, an LLM-judge harness)
six graders over closed sets and one integer. There is nothing here a framework would do better, and a judge would add cost, variance and a second thing to be wrong on comparisons that are string equality.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each route is a chain of three weekly runs; each run is parse, assemble, call, parse, score, advance. There is no branching, no tool selection and no retry-with-a-different-prompt. Drawing it as a graph would suggest choices the code does not make.
The other sideWhat a framework costs you
No scheduler. This is a monitor that carries state and is still INVOKED, not woken -- the cadence, the retry policy and the question of what to do when a week does not arrive are all outside the kit, and the last of those has a 30-day deadline behind it.
No state store. src/state.py writes one JSON file atomically. Two runs against the same file would race, and nothing here prevents it.
No observability beyond a result file. There is no trace, no span and no dashboard; the run file records every call's tokens, latency, finish reason and carried state, and that is the whole of it.
No queue and no backpressure. evals/run.py runs 12 route chains concurrently by default and a provider under load is handled by a bounded backoff inside src/adapters, not by a rate limiter.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-ticket-anomaly on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
6,679 ms
p50 8650 to 8784 ms, p95 62904 to 121524 ms over 2 passes on the SAME 36 weeks
nothing -- latency is not gated here, and a page that alarmed on it would alarm on the provider's mood
Model, p95
50,568 ms
p50 8650 to 8784 ms, p95 62904 to 121524 ms over 2 passes on the SAME 36 weeks
nothing -- latency is not gated here, and a page that alarmed on it would alarm on the provider's mood
Input tokens
266,922
0 tokens over 2 passes
any change at all
Output tokens
216,438
60588 to 83604 tokens over 2 passes on the SAME 36 weeks -- a 38 pct swing
nothing -- but it is why cost_at_10x is stated as a range in prose rather than as a single figure anybody should budget against
No movement column. Not one of the 6 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-ticket-anomaly-calibration8,672 ms
c001-ticket-anomaly-calibration12,605 ms
p002-ticket-anomaly-repeat8,650 ms
r001-ticket-anomaly6,679 ms
s001-ticket-anomaly-stateless6,115 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
2 runs not plotted. b000-ticket-anomaly-deskband, b001-ticket-anomaly-allclear recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 8 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
weekly monitor sheets
data/corpus/<ROUTE>-W<n>.txt -- 150 files, 629038 bytes, generated once from a fixed seed and committed. No database, no object store, no fetch.
six of the seven sections go to the provider in the prompt; Customer Contact -- the account holder, the service address, a telephone number and an email -- is mapped by no field and subtracted unconditionally, so it never leaves the machine on any document.
the carried state
data/state.json in a deployment, written atomically by src/state.save. In an EVAL it is scoped to the run and never touches disk -- a run that started from the previous run's memory could not be re-run and could not be compared with its own control.
yes, and it is the point: three sentences built by src/state.describe go into every stateful prompt. Six scalars, no transcript, no prior reply.
the answer key
data/gold.jsonl -- one line per route-week carrying the four labels, the route's opening state, the week's total and the memory-dependent flag. Computed by src/norms.replay, never typed.
no. The scorer reads it in-process; nothing about it is ever sent.
run records
results/eval-<run-id>.json -- every call's tokens, latency, finish reason, carried state and parsed answer, plus the six graders. Committed to the kit repository.
no. They are written locally and read by the site's own extractor; nothing posts them anywhere.
the provider key
.env beside the kit or at the repository root, gitignored, written at 0600 by src/config.save with the mode set before the file exists.
only as an Authorization header to the one completions endpoint. It is never written to a result file, a log line, a page or an error string -- src/app.py redacts it and the base URL out of any exception before returning.
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never written to a result file, a log line, a screenshot or a page. src/app.py redacts the key and the base URL out of any error string before returning it. src/config.save writes .env at 0600 with os.open BEFORE the file exists, so it is never world-readable even for an instant. No credential has ever been in this repository.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
one run per SERVICE WEEK over the whole open population, and each run owns exactly the tickets that posted since the last one. The unit is a route-week because the scale tickets are settled weekly at the facility; a run does not sample, it re-reads every open route.
3 consecutive weekly runs per route over 50 routes = 150 route-weeks in one pass; the whole population takes 194.4 seconds of wall clock at 12 concurrent route chains, so the run itself is never what makes a week late. A MISSED run costs the facility correction window on that week: 30 days from the tip date, and every sheet prints how many are left -- the corpus ranges from 19 to 28 days remaining at the moment of review. Miss one weekly run and the cheap remedy on that tonnage is gone; the ticket can no longer be corrected and the money moves to a credit memo against an invoice the customer already holds. Miss two and the streak the second run would have raised is three weeks old, which is past this kit's own norm-shift threshold. (r001-ticket-anomaly, data/corpus-stats.json, src/norms.DISPUTE_WINDOW_DAYS)
the cadence stops being weekly. A daily monitor over the same book multiplies the bill by five and measures noise -- a single day's tickets are not a comparable unit against a weekly norm. A monthly one is inside the 30-day window only if it lands early in it, and a route that drifts in week one is not looked at until week four.
moving to a different period invalidates the norm, the band and every threshold on this page at once, because all of them are expressed per service week. It does NOT invalidate the cost figures, which are per call and flat in history length.
state
six scalars per route, carried between runs and nothing else: the norm in force, the consecutive-week count, the direction, whether an anomaly is open, how many weeks the route has been monitored, and the previous verdict. Written from the arithmetic in src/norms.step, never from a model reply.
input per route-week is 1779 tokens with the state and 1731 without -- a difference of under 3 pct, on routes carrying up to four weeks of streak, so the cost is flat in history length by measurement and not only by design. The state is worth 36.00 points of verdict accuracy, 67.53 points on the 77 watchlist weeks and 84.38 points on the 64 memory-dependent cells, and it takes ALREADY_OPEN recall from 0.00 pct to 100.00 pct. (r001-ticket-anomaly against s001-ticket-anomaly-stateless)
one JSON file written atomically by one process. Two scheduled runs against the same file would race, and nothing in the kit prevents it. That is the point at which this needs a real store, and the seam is src/state.py -- load() and save() are the whole interface.
a state store that is not atomic invalidates every ALREADY_OPEN answer at once, in the worst direction: a truncated file reads as no history, which closes every open item and makes the next run re-raise every live dispute in the book on the same morning.
model
one completion call per route-week against an OpenAI-compatible endpoint or Anthropic, chosen by PROVIDER in .env. One system message, one user message, no tools, no streaming, temperature left at the provider's default.
150 calls on the scored run, 150 on the stateless control and 12 on the two calibrations. p50 6679 ms, p95 50568 ms. Output ceiling published at 24000 and the largest reply reached 14904; a calibration at 8,000 lost 1 of 6. 93.6 pct of output tokens were provider-side reasoning. (r001-ticket-anomaly, s001-ticket-anomaly-stateless, c000-ticket-anomaly-calibration, c001-ticket-anomaly-calibration)
the published 24000-token ceiling. Above it the reply is cut off and billed in full, and the week is scored as a miss in all four fields. c000 is the committed evidence of what that looks like.
changing the model invalidates every accuracy figure on this page and none of the corpus figures. Changing the ceiling invalidates the failure count only.
labels
150 route-weeks over 50 invented routes, four labelled fields each, computed by src/norms.replay from a planted opening state and three observations. Nine patterns, chosen so that 64 of the cells cannot be answered from the week in front of you.
77 watchlist weeks, 73 quiet, 32 ALREADY_OPEN, 64 memory-dependent, 3 NOT_MONITORED. 8 pre-flight checks re-derive every label from the shipped text before a run may spend, and 4 of them are red-proven. (data/corpus-stats.json, evals/check_labels.py)
the labels stop being credible where the corpus stops being representative: three weeks per route, 24 generated reason sentences, one facility per route and no seasonality. The anomaly-class figures are the least transferable numbers on the page, and the key on that field is contested on up to 14 cells -- see Eval.taxonomy.
a longer chain invalidates the week-count headline rather than extending it, because nothing here has been run past week three. Real dispatcher prose invalidates the class figures outright.
corpus refresh
the population is re-read WHOLE on every scheduled run -- there is no index to rebuild, no cache to invalidate and no incremental load. A route that appears, disappears or changes container between runs is simply read as it is on the day.
regenerating all 150 sheets and the answer key from the fixed seed takes under a second and costs $0.00 -- no clock is read and no model is called, which is why the corpus in the repository is byte-identical on any machine. (tools/build_corpus.py, data/corpus-stats.json)
a route that leaves the book PERMANENTLY. Off-the-book for one week is handled (NOT_MONITORED, streak reset, open item preserved); a route pulled for good keeps its open item for ever, because the carried state has no expiry. Not in the corpus and not handled.
nothing on this page. Re-reading the population whole is what makes a reading comparable to the last one, and it is the reason there is no index answer to give.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
the same route raised week after week, and the facility rejecting each dispute as a duplicate
the open-item flag is not reaching the prompt -- either the carried state was lost or the run is being made without it. A one-sheet reader cannot know an item is open, so it raises every week the route is out of band.
compare the run's per_week[].carried_in against data/state.json before disputing the answer. On RTE-0034-W1 the free desk-band floor answers RAISE and the stateful arm answers ALREADY_OPEN, on identical text. (results/eval-b000-ticket-anomaly-deskband.json against results/eval-r001-ticket-anomaly.json, RTE-0034-W1)
a whole book of live disputes re-raised on the same morning
the state file was truncated and read as an empty history, which closes every open item at once. It is the quietest way this kit can fail and nothing downstream would notice, because every individual answer looks reasonable.
check data/state.json is non-empty and parses before the next run. src/state.save writes to a temp file and os.replace()s it precisely so a kill mid-write leaves the PREVIOUS state rather than half of this one. (src/state.py -- the atomic write and the comment stating this failure mode; no run has produced it)
NORM_SHIFTED on a route whose tonnage is back inside its band
the norm status is being carried forward rather than recomputed. Gold is NORM_HOLDS whenever the week count is 0, because rule 7's three-consecutive-week condition is not met.
read the week count in the same reply. If it says 0 and the norm status says shifted, the reply is internally inconsistent and the norm status is the wrong half. This is the model's only failure shape on this run: 5 cells, all of them the last week of an already_open route. (results/eval-r001-ticket-anomaly.json, the norm_status misses)
a confident topical cause code on a note that says nobody recorded the cause
something is keyword-matching the reason rather than reading it. The free desk band does this on 10 weeks; the model does the opposite and answers CAUSE_UNKNOWN on notes that hedge.
read the note. If it names a cause and then says it was never established, CAUSE_UNKNOWN is the intended answer and a topical code is the failure. The two arms get 14 and 14 class cells wrong with ZERO overlap. (results/eval-b000-ticket-anomaly-deskband.json against results/eval-r001-ticket-anomaly.json, anomaly_class cells)
No machine symptom — this failure leaves no trace in any output.
the defence is outside the kit and has to be: a scheduler with an alarm on a run that did not start, plus the days-left figure the sheets already print (19 to 28 in this corpus) read as a deadline rather than as decoration. Architecture.breaks_at_scale and guardrails.add_first both name it, and frameworks.mapping records it as the one place this kit is genuinely incomplete for its pattern.
⚠︎ THE CADENCE HALF OF THIS LADDER IS DECLARED AND THE MACHINERY BEHIND IT IS NOT BUILT. This kit runs weekly because the unit is a service week and the facility's correction window is 30 days; that is measured and stated. What is NOT measured is anything about the schedule actually holding: there is no scheduler, no retry for a run that did not happen, no alarm on a missed week, and no figure anywhere for how often a real disposal desk actually manages a weekly pass. A monitor whose cadence is somebody's memory has an unmeasured failure mode and this is it.
⚑ RUN-TO-RUN VARIANCE IS NOW PARTLY MEASURED, AND ONE FIELD MOVED. p002-ticket-anomaly-repeat re-asked 36 of the 150 weeks with an identical prompt and an identical carried state; 143 of 144 answered cells came back byte-identical and nine of ten graders reproduced to the digit. The exception is norm status, where a single cell (RTE-0043-W2) answered NORM_HOLDS on the first pass and NORM_SHIFTED on the second -- the kit's own un-shifting defect, showing itself to be a coin toss on marginal cells rather than five fixed rows. ⚠︎ WHAT IS STILL NOT MEASURED: the other 114 weeks, ever; the floors and the stateless control, at all; and anything beyond n=2, which cannot separate a stable answer from an unstable one that agreed twice. The 100.00 pct headline is banded on 36 weeks and asserted on 150.
⚠︎ CONCURRENCY. The eval ran 12 route chains at once and nothing measured what happens at 50 or at 500. src/state.py writes one JSON file and two scheduled runs against it would race; the kit does not prevent it and no run has provoked it.
⚠︎ PROVIDER-SIDE RETENTION. Whether the provider keeps the prompts, and for how long, is outside the kit and unmeasured. What IS known is what is sent: six of seven sections, with the customer's name, service address, telephone and email in the section that never leaves.
⚠︎ GPU SIZING AND SELF-HOSTING. Nothing here has been run against a local model. src/adapters speaks to any OpenAI-compatible endpoint including a local server, so the seam exists; no measurement does.
⚠︎ WHAT A DUPLICATE DISPUTE COSTS. The kit counts them (27 on the free floor, 0 with the carried state) and prices neither the rework nor a ton at any real facility.
⚠︎ WHETHER THE FOUR SHIPPED THRESHOLDS ARE ANYWHERE NEAR A REAL FACILITY'S. The band, the raise threshold, the norm-shift threshold and the dispute window were invented for this kit and are the first four things a forker must replace.
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every route, ticket, tonnage, facility, customer, address and recorded reason is invented by the generator. See kits/UC0088-ticket-anomaly/data/SOURCES.md. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
Spot waste routes dumping more weight than usual
PresenterOpens the private repo. Visible to admins only.
In one lineThe verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
For each of the 150 weeks and each of the four answered fields, did the reply equal the computed answer key? Verdict, norm status and class are compared case-insensitively as strings; the week count is compared as an integer, and a value that is not a whole number is scored absent rather than coerced to 0 -- 0 is a meaningful answer here.
$0.00per 1,000 weekly monitor sheets
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function both free floors and the stateless control are scored through.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The week
RTE-0034-W1 -- an already_open route whose norm was rebased in an earlier run (pattern already_open, week 1 of 3)
Read off the sheet, in code
Total net this week 5.51 tons over 3 tickets. Book norm printed on the sheet 3.34 tons, book band 2.84 to 3.84. Norm actually in force 4.34 tons, band 3.69 to 4.99. On the disposal book. 20 days left in the facility dispute window.
What the Open Items log carries
no disposal exception was opened or closed against this route in the service week this sheet covers. There is nothing anywhere on it that says a dispute is already open, or that the route has been out of band for two weeks, or that the norm was rebased.
Carried state, in the prompt
The norm IN FORCE for this route is 4.34 tons per service week. It was rebased in an earlier run and is NOT the book figure printed on the sheet. It has 30 recorded weeks on the monitor. As at the previous run it had been out of band HIGH for 2 consecutive weeks. A disposal anomaly raised by this monitor is STILL OPEN against this route. It was reported RAISE last run.
The free desk-band floor
RAISE at 1 week out, NORM_HOLDS -- a second dispute against tickets that already have one open
The same model, memory removed
WATCH at 1 week out, NORM_UNKNOWN -- told there is no history, it can only see this week
The model, with the carried state
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Ground truth
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Scored as
4 of 4 cells correct. The row matters because the two no-memory arms are wrong in opposite and both-expensive ways: the floor files a duplicate the facility rejects, the stateless model under-calls a route that has been drifting for a month.
Grader
Verdict
Why
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
verdict hit, week count hit, norm status hit, anomaly class hit -- 4 of 4
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, SERVICE_STEP, against a key of the same four. The free desk-band floor scores 1 of 4 on this row (RAISE / 1 week / NORM_HOLDS / SERVICE_STEP) and the memory-removed arm scores 1 of 4 as well, on a different cell.
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
not a duplicate
This is one of the 32 weeks whose key is ALREADY_OPEN. The model answered ALREADY_OPEN. The free desk-band floor answered RAISE, which makes this row one of its 27 duplicate disputes -- a second copy filed against tickets that already carry one.
The cells whose answer cannot be derived from the week in front of you
in the subset, and hit on both fields
Running this week from a blank state gives RAISE at 1 week with the norm holding; running it from the route's real carried state gives ALREADY_OPEN at 4 with the norm shifted. The two disagree, so the cell is one of the 64 memory-dependent ones, and the model got both fields right.
Watchlist recall and false alarms, counted apart and never summed
on the watchlist, and reported on it
ALREADY_OPEN is not a quiet verdict, so this row sits in the 77-week watchlist denominator rather than the 73-week quiet one. It is neither a miss nor a false alarm.
The tonnage this run would put in dispute with the facility
contributes nothing, correctly
Only RAISE counts towards the disputed total -- ALREADY_OPEN tonnage is already with the facility under an existing dispute. The model adds 0 tons from this row. The free floor's RAISE adds this route's 5.51 tons to a portfolio figure that is already 166.33 pct over.
The formulaWhat it computes
accuracy = hits / 150 per field. A week whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% verdict accuracy · 3 more measured on this row
the same tier, memory removed (THE CONTROL)
64.0% verdict accuracy · 3 more measured on this row
the free desk band, no model, $0.00
63.3% verdict accuracy · 3 more measured on this row
always CLEAR -- the majority-class floor, $0.00
46.7% verdict accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/norms.replay over the planted opening state and observations at generation time. This grader IS the reference, so its own TPR and TNR are not measurable.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one thing about it is known and contested -- see Eval.taxonomy's first row, where 9 cells turn on whether 'a mis-post is SUSPECTED' names a cause or declines to.
Watch these
verdict_accuracy_pct
duplicate_raise_rate_pct
anomaly_class_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0, and duplicate_raise_rate_pct above 0. A week that returns nothing scores as a miss in all four fields, so a reliability failure arrives disguised as a quality one.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because three are small: 3 NOT_MONITORED cells, 9 flip cells, 15 closed_reset cells.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/norms.py, src/segment.py or src/select.py. Re-run the scored eval (150 calls) only when the prompt, the model or the corpus changes -- and re-run the stateless control WITH it, because the headline on this page is a difference between the two and half a difference is not a number.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real disposal book, where whether a week is 'really' out of band depends on a norm nobody has agreed and a reason a dispatcher typed in a hurry.
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
Spot waste routes dumping more weight than usual
PresenterOpens the private repo. Visible to admins only.
In one lineDuplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
Of the 32 weeks where an anomaly raised by an earlier run is still open, how many were raised again? Every one is a second dispute against tickets that already carry one, which the facility rejects and somebody then reworks by hand.
$0.00per 1,000 weekly monitor sheets
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The week
RTE-0034-W1 -- an already_open route whose norm was rebased in an earlier run (pattern already_open, week 1 of 3)
Read off the sheet, in code
Total net this week 5.51 tons over 3 tickets. Book norm printed on the sheet 3.34 tons, book band 2.84 to 3.84. Norm actually in force 4.34 tons, band 3.69 to 4.99. On the disposal book. 20 days left in the facility dispute window.
What the Open Items log carries
no disposal exception was opened or closed against this route in the service week this sheet covers. There is nothing anywhere on it that says a dispute is already open, or that the route has been out of band for two weeks, or that the norm was rebased.
Carried state, in the prompt
The norm IN FORCE for this route is 4.34 tons per service week. It was rebased in an earlier run and is NOT the book figure printed on the sheet. It has 30 recorded weeks on the monitor. As at the previous run it had been out of band HIGH for 2 consecutive weeks. A disposal anomaly raised by this monitor is STILL OPEN against this route. It was reported RAISE last run.
The free desk-band floor
RAISE at 1 week out, NORM_HOLDS -- a second dispute against tickets that already have one open
The same model, memory removed
WATCH at 1 week out, NORM_UNKNOWN -- told there is no history, it can only see this week
The model, with the carried state
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Ground truth
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Scored as
4 of 4 cells correct. The row matters because the two no-memory arms are wrong in opposite and both-expensive ways: the floor files a duplicate the facility rejects, the stateless model under-calls a route that has been drifting for a month.
Grader
Verdict
Why
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
verdict hit, week count hit, norm status hit, anomaly class hit -- 4 of 4
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, SERVICE_STEP, against a key of the same four. The free desk-band floor scores 1 of 4 on this row (RAISE / 1 week / NORM_HOLDS / SERVICE_STEP) and the memory-removed arm scores 1 of 4 as well, on a different cell.
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
not a duplicate
This is one of the 32 weeks whose key is ALREADY_OPEN. The model answered ALREADY_OPEN. The free desk-band floor answered RAISE, which makes this row one of its 27 duplicate disputes -- a second copy filed against tickets that already carry one.
The cells whose answer cannot be derived from the week in front of you
in the subset, and hit on both fields
Running this week from a blank state gives RAISE at 1 week with the norm holding; running it from the route's real carried state gives ALREADY_OPEN at 4 with the norm shifted. The two disagree, so the cell is one of the 64 memory-dependent ones, and the model got both fields right.
Watchlist recall and false alarms, counted apart and never summed
on the watchlist, and reported on it
ALREADY_OPEN is not a quiet verdict, so this row sits in the 77-week watchlist denominator rather than the 73-week quiet one. It is neither a miss nor a false alarm.
The tonnage this run would put in dispute with the facility
contributes nothing, correctly
Only RAISE counts towards the disputed total -- ALREADY_OPEN tonnage is already with the facility under an existing dispute. The model adds 0 tons from this row. The free floor's RAISE adds this route's 5.51 tons to a portfolio figure that is already 166.33 pct over.
The formulaWhat it computes
duplicates / 32. It is a COST WITH A NAME, not a slice of an accuracy figure, which is why it is counted on its own.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
0.0% duplicate raise rate · 1 more measured on this row
the same tier, memory removed (THE CONTROL)
0.0% duplicate raise rate · 1 more measured on this row
the free desk band, no model, $0.00
84.4% duplicate raise rate · 1 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl. A week is ALREADY_OPEN when the carried state says an item is open and the week is out of band -- both facts computed by src/norms.step, neither on any sheet.
These rates are UNKNOWN, on purpose
What a duplicate dispute actually costs a hauler. Nobody here has measured the rework, and the kit states the count rather than pricing it.
Watch these
duplicate_raise_rate_pct
already_open_recall_pct
Alarm on
duplicate_raise_rate_pct above 0 on any arm that carries state -- it means the open-item flag stopped reaching the prompt.
How tight can the band be? No threshold. The verdict string decides.
Cadence: With every scored run, and with the stateless control beside it.
The decisionWhen to reach for it
Use it
The kit carries an open-item flag between runs. It is unreachable without one: no sheet in this corpus states that an item is open.
Do not use it
A one-shot classifier. It has no open-item flag, so this grader would be measuring the absence of a feature rather than a failure.
The cells whose answer cannot be derived from the week in front of you
Spot waste routes dumping more weight than usual
PresenterOpens the private repo. Visible to admins only.
In one lineThe cells whose answer cannot be derived from the week in front of you
On the 64 of 150 cells where running the week from a BLANK state gives a different answer from running it from its real carried state, how often was the verdict and the week count right?
$0.00per 1,000 weekly monitor sheets
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The week
RTE-0034-W1 -- an already_open route whose norm was rebased in an earlier run (pattern already_open, week 1 of 3)
Read off the sheet, in code
Total net this week 5.51 tons over 3 tickets. Book norm printed on the sheet 3.34 tons, book band 2.84 to 3.84. Norm actually in force 4.34 tons, band 3.69 to 4.99. On the disposal book. 20 days left in the facility dispute window.
What the Open Items log carries
no disposal exception was opened or closed against this route in the service week this sheet covers. There is nothing anywhere on it that says a dispute is already open, or that the route has been out of band for two weeks, or that the norm was rebased.
Carried state, in the prompt
The norm IN FORCE for this route is 4.34 tons per service week. It was rebased in an earlier run and is NOT the book figure printed on the sheet. It has 30 recorded weeks on the monitor. As at the previous run it had been out of band HIGH for 2 consecutive weeks. A disposal anomaly raised by this monitor is STILL OPEN against this route. It was reported RAISE last run.
The free desk-band floor
RAISE at 1 week out, NORM_HOLDS -- a second dispute against tickets that already have one open
The same model, memory removed
WATCH at 1 week out, NORM_UNKNOWN -- told there is no history, it can only see this week
The model, with the carried state
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Ground truth
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Scored as
4 of 4 cells correct. The row matters because the two no-memory arms are wrong in opposite and both-expensive ways: the floor files a duplicate the facility rejects, the stateless model under-calls a route that has been drifting for a month.
Grader
Verdict
Why
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
verdict hit, week count hit, norm status hit, anomaly class hit -- 4 of 4
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, SERVICE_STEP, against a key of the same four. The free desk-band floor scores 1 of 4 on this row (RAISE / 1 week / NORM_HOLDS / SERVICE_STEP) and the memory-removed arm scores 1 of 4 as well, on a different cell.
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
not a duplicate
This is one of the 32 weeks whose key is ALREADY_OPEN. The model answered ALREADY_OPEN. The free desk-band floor answered RAISE, which makes this row one of its 27 duplicate disputes -- a second copy filed against tickets that already carry one.
The cells whose answer cannot be derived from the week in front of you
in the subset, and hit on both fields
Running this week from a blank state gives RAISE at 1 week with the norm holding; running it from the route's real carried state gives ALREADY_OPEN at 4 with the norm shifted. The two disagree, so the cell is one of the 64 memory-dependent ones, and the model got both fields right.
Watchlist recall and false alarms, counted apart and never summed
on the watchlist, and reported on it
ALREADY_OPEN is not a quiet verdict, so this row sits in the 77-week watchlist denominator rather than the 73-week quiet one. It is neither a miss nor a false alarm.
The tonnage this run would put in dispute with the facility
contributes nothing, correctly
Only RAISE counts towards the disputed total -- ALREADY_OPEN tonnage is already with the facility under an existing dispute. The model adds 0 tons from this row. The free floor's RAISE adds this route's 5.51 tons to a portfolio figure that is already 166.33 pct over.
The formulaWhat it computes
hits / 64, separately for verdict and for the week count. The subset itself is computed by src/norms.replay running every week twice, not labelled by hand.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% memory verdict accuracy · 1 more measured on this row
the same tier, memory removed (THE CONTROL)
15.6% memory verdict accuracy · 1 more measured on this row
the free desk band, no model, $0.00
32.8% memory verdict accuracy · 1 more measured on this row
In operationWhat to monitor
Reference standard: src/norms.replay, run twice per week -- once from the real carried state, once from the blank state a one-sheet reader would have.
These rates are UNKNOWN, on purpose
Whether the subset is the RIGHT 64 cells is a property of the corpus generator, not of the model. An earlier definition over-counted it by 23; see Data.breaks_on.
Watch these
memory_verdict_accuracy_pct
memory_weeks_accuracy_pct
Alarm on
the subset shrinking below 20 cells -- evals/check_labels.py fails the build rather than publishing a rate with no denominator.
How tight can the band be? No threshold.
Cadence: With every scored run.
The decisionWhen to reach for it
Use it
The kit claims memory is worth something and the claim needs a denominator.
Do not use it
As a headline. It double-counts cells already inside the verdict figure, and a headline that summed them would report the same answers twice.
Watchlist recall and false alarms, counted apart and never summed
Spot waste routes dumping more weight than usual
PresenterOpens the private repo. Visible to admins only.
In one lineWatchlist recall and false alarms, counted apart and never summed
Of the 77 weeks a clerk must touch, how many were reported as something to touch? And of the 73 quiet weeks, how many were put on the watchlist anyway?
$0.00per 1,000 weekly monitor sheets
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The week
RTE-0034-W1 -- an already_open route whose norm was rebased in an earlier run (pattern already_open, week 1 of 3)
Read off the sheet, in code
Total net this week 5.51 tons over 3 tickets. Book norm printed on the sheet 3.34 tons, book band 2.84 to 3.84. Norm actually in force 4.34 tons, band 3.69 to 4.99. On the disposal book. 20 days left in the facility dispute window.
What the Open Items log carries
no disposal exception was opened or closed against this route in the service week this sheet covers. There is nothing anywhere on it that says a dispute is already open, or that the route has been out of band for two weeks, or that the norm was rebased.
Carried state, in the prompt
The norm IN FORCE for this route is 4.34 tons per service week. It was rebased in an earlier run and is NOT the book figure printed on the sheet. It has 30 recorded weeks on the monitor. As at the previous run it had been out of band HIGH for 2 consecutive weeks. A disposal anomaly raised by this monitor is STILL OPEN against this route. It was reported RAISE last run.
The free desk-band floor
RAISE at 1 week out, NORM_HOLDS -- a second dispute against tickets that already have one open
The same model, memory removed
WATCH at 1 week out, NORM_UNKNOWN -- told there is no history, it can only see this week
The model, with the carried state
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Ground truth
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Scored as
4 of 4 cells correct. The row matters because the two no-memory arms are wrong in opposite and both-expensive ways: the floor files a duplicate the facility rejects, the stateless model under-calls a route that has been drifting for a month.
Grader
Verdict
Why
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
verdict hit, week count hit, norm status hit, anomaly class hit -- 4 of 4
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, SERVICE_STEP, against a key of the same four. The free desk-band floor scores 1 of 4 on this row (RAISE / 1 week / NORM_HOLDS / SERVICE_STEP) and the memory-removed arm scores 1 of 4 as well, on a different cell.
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
not a duplicate
This is one of the 32 weeks whose key is ALREADY_OPEN. The model answered ALREADY_OPEN. The free desk-band floor answered RAISE, which makes this row one of its 27 duplicate disputes -- a second copy filed against tickets that already carry one.
The cells whose answer cannot be derived from the week in front of you
in the subset, and hit on both fields
Running this week from a blank state gives RAISE at 1 week with the norm holding; running it from the route's real carried state gives ALREADY_OPEN at 4 with the norm shifted. The two disagree, so the cell is one of the 64 memory-dependent ones, and the model got both fields right.
Watchlist recall and false alarms, counted apart and never summed
on the watchlist, and reported on it
ALREADY_OPEN is not a quiet verdict, so this row sits in the 77-week watchlist denominator rather than the 73-week quiet one. It is neither a miss nor a false alarm.
The tonnage this run would put in dispute with the facility
contributes nothing, correctly
Only RAISE counts towards the disputed total -- ALREADY_OPEN tonnage is already with the facility under an existing dispute. The model adds 0 tons from this row. The free floor's RAISE adds this route's 5.51 tons to a portfolio figure that is already 166.33 pct over.
The formulaWhat it computes
recall = hits / 77 on the non-quiet weeks; false alarm rate = quiet weeks reported non-quiet / 73.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% verdict on watchlist · 1 more measured on this row
the same tier, memory removed (THE CONTROL)
32.5% verdict on watchlist · 1 more measured on this row
the free desk band, no model, $0.00
31.2% verdict on watchlist · 1 more measured on this row
always CLEAR -- the majority-class floor, $0.00
0.0% verdict on watchlist · 1 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl. Quiet is CLEAR or NOT_MONITORED; everything else is a week somebody has to look at.
These rates are UNKNOWN, on purpose
What a false alarm actually costs this desk. Not measured.
Watch these
verdict_on_watchlist_pct
false_alarm_rate_pct
missed_watchlist
Alarm on
missed_watchlist above 0.
How tight can the band be? No threshold.
Cadence: With every scored run.
The decisionWhen to reach for it
Use it
Always. The two errors cost differently: a quiet route on the watchlist costs a clerk twenty minutes, a route drifting for a month reported quiet costs the month's over-tipped tonnage, uncorrectable once the 30-day facility window shuts.
Do not use it
As one number. Averaging them prices the two errors the same.
The tonnage this run would put in dispute with the facility
Spot waste routes dumping more weight than usual
PresenterOpens the private repo. Visible to admins only.
In one lineThe tonnage this run would put in dispute with the facility
Add up the week totals of every route reported RAISE. Is the portfolio number the disposal desk would take to the facility on Monday right?
$0.00per 1,000 weekly monitor sheets
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model.
Every grader on these pages scored the same 150 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The week
RTE-0034-W1 -- an already_open route whose norm was rebased in an earlier run (pattern already_open, week 1 of 3)
Read off the sheet, in code
Total net this week 5.51 tons over 3 tickets. Book norm printed on the sheet 3.34 tons, book band 2.84 to 3.84. Norm actually in force 4.34 tons, band 3.69 to 4.99. On the disposal book. 20 days left in the facility dispute window.
What the Open Items log carries
no disposal exception was opened or closed against this route in the service week this sheet covers. There is nothing anywhere on it that says a dispute is already open, or that the route has been out of band for two weeks, or that the norm was rebased.
Carried state, in the prompt
The norm IN FORCE for this route is 4.34 tons per service week. It was rebased in an earlier run and is NOT the book figure printed on the sheet. It has 30 recorded weeks on the monitor. As at the previous run it had been out of band HIGH for 2 consecutive weeks. A disposal anomaly raised by this monitor is STILL OPEN against this route. It was reported RAISE last run.
The free desk-band floor
RAISE at 1 week out, NORM_HOLDS -- a second dispute against tickets that already have one open
The same model, memory removed
WATCH at 1 week out, NORM_UNKNOWN -- told there is no history, it can only see this week
The model, with the carried state
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Ground truth
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, class SERVICE_STEP
Scored as
4 of 4 cells correct. The row matters because the two no-memory arms are wrong in opposite and both-expensive ways: the floor files a duplicate the facility rejects, the stateless model under-calls a route that has been drifting for a month.
Grader
Verdict
Why
The verdict, the week count, the norm status and the anomaly class, per week, exact match against the computed answer key
verdict hit, week count hit, norm status hit, anomaly class hit -- 4 of 4
ALREADY_OPEN at 4 weeks out, NORM_SHIFTED, SERVICE_STEP, against a key of the same four. The free desk-band floor scores 1 of 4 on this row (RAISE / 1 week / NORM_HOLDS / SERVICE_STEP) and the memory-removed arm scores 1 of 4 as well, on a different cell.
Duplicate disputes: weeks whose correct answer is ALREADY_OPEN that were answered RAISE
not a duplicate
This is one of the 32 weeks whose key is ALREADY_OPEN. The model answered ALREADY_OPEN. The free desk-band floor answered RAISE, which makes this row one of its 27 duplicate disputes -- a second copy filed against tickets that already carry one.
The cells whose answer cannot be derived from the week in front of you
in the subset, and hit on both fields
Running this week from a blank state gives RAISE at 1 week with the norm holding; running it from the route's real carried state gives ALREADY_OPEN at 4 with the norm shifted. The two disagree, so the cell is one of the 64 memory-dependent ones, and the model got both fields right.
Watchlist recall and false alarms, counted apart and never summed
on the watchlist, and reported on it
ALREADY_OPEN is not a quiet verdict, so this row sits in the 77-week watchlist denominator rather than the 73-week quiet one. It is neither a miss nor a false alarm.
The tonnage this run would put in dispute with the facility
contributes nothing, correctly
Only RAISE counts towards the disputed total -- ALREADY_OPEN tonnage is already with the facility under an existing dispute. The model adds 0 tons from this row. The free floor's RAISE adds this route's 5.51 tons to a portfolio figure that is already 166.33 pct over.
The formulaWhat it computes
|reported - true| / true, over the 150 weeks. Only RAISE counts: ALREADY_OPEN tonnage is already with the facility under an existing dispute and counting it again would double the number rule 4 exists to stop.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
0.0% tonnage error
the same tier, memory removed (THE CONTROL)
100.0% tonnage error
the free desk band, no model, $0.00
no headline metric on this row — it records tonnage error pct 166.33
In operationWhat to monitor
Reference standard: data/gold.jsonl week totals, summed over the RAISE weeks.
These rates are UNKNOWN, on purpose
The dollar value of a ton at any real facility. The kit reports tons and prices nothing.
Watch these
tonnage_error_pct
tonnage_reported_raised
Alarm on
tonnage_error_pct moving while the per-week figures do not -- that is cancellation, not accuracy.
How tight can the band be? No threshold.
Cadence: With every scored run.
The decisionWhen to reach for it
Use it
Somebody upstream reads one aggregate rather than the rows.
Do not use it
As a substitute for per-route accuracy. Errors in opposite directions cancel here, so this can be right while every row is wrong.
A living map of modern AI — kept current every morning