Catch new royalty problems at your franchise stores
Every store's royalty numbers get flagged the same way every month, and most of them were already explained last time. This app checks each store's report against last month's and only raises the ones that are new.
PresenterOpens the private repo. Visible to admins only.
For the royalty analystRestaurants & QSR · Cross-domain
Why it matters
Today's manual process, and the same job with the app
A franchise finance team at a multi-store restaurant brand, reviewing every store's royalty report each month.
✕Today's manual process
1Pull every store's report and recompute four indicators against that store's own history.
2Remember what was already flagged or dig through last month's file to check.
3Read the store's operating note and judge by eye whether it covers the problem.
4Miss a real one and a legitimate under-reporting case goes unnoticed for a month.
Every store re-read manually each month
✓With the app
1The app reads every store and recomputes the same four indicators from the printed figures.
2It remembers last run and only raises what's new since then.
3It reads the note itself checking the cause against the approved list.
4Nothing new stays quiet and only a genuine change reaches the analyst.
Only new, unexplained crossings reach anyone
See it work
One real case, read by the app, step by step
Store STR-0006's average ticket dipped during an approved remodel closure, and its own note already explains why.
Catch new royalty problems at your franchise storesReference appBuilt to be shaped to your process
5
1One number moved average ticket ran 6.10% below its own average this period.
2What's outside its normal range the average ticket indicator, flagged by the arithmetic.
3The store explains why closed for an approved remodel, confirmed by the area coach.
4Nothing new this period no indicator crossed for the first time.
5The call: no change a correct no-change, so no false alarm reaches anyone.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch new royalty problems at your franchise stores
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A franchisor's royalty function gets one reporting file per store per period, and the arithmetic in it is easy: cash share against the store's own baseline, deposits against reported net sales, discounts against the segment median, average ticket against the store's own baseline. A spreadsheet flags all four perfectly. What the spreadsheet produces is a flag column that lights up the same forty stores every month, most of them for reasons somebody already accepted, and the column that says 'this is new' does not exist. So the file gets read by a person, and it gets read late. Somebody re-reading every open store's royalty file each month, working out which indicators are outside their bands, remembering which of those were already outside last month, and deciding whether the store's own operating note excuses the ones that are new.
Audience
Franchise finance and internal audit teams who already compute these indicators and cannot triage them, and anyone building a monitor of any kind -- the cadence half of this kit is the transferable part, and it is the half four earlier monitors in this series left out. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual store-period readings
The corpus is 117 store-period readings, 0.43 MB (jsonl 1 · txt 117). A real royalty file names a real operator, their bank deposits and their sales mix, and publishing one would be publishing the exact thing the kit exists to look at. So the SHAPE is taken from what a monthly royalty reporting pack contains and everything in it is invented -- every franchisee email is on the reserved .invalid top-level domain, which by RFC 2606 can never resolve.
What the corpus is FOR is the six operating-note classes, and they were designed before anything ran: coded_stale 4, coded_valid 6, coded_wrong_indicator 2, none 89, prose_unlisted 4, prose_valid 12. Three of the six are readable by a free regex and three are not, in both directions -- an approved cause written in words with no code (which must excuse) and a plausible cause that is not on the list (which must not). A corpus whose exemptions were all machine-readable would have decided the result before the run.
The POPULATION MOVES on purpose: two stores open between run 1 and run 2, one closes after run 2. A fixed rectangle of N stores x M periods cannot test what a monitor does when it meets a row it has never seen, which is the commonest thing that happens to one.
The corpus
The 117 store-period readingsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your store-period readings. That is the whole change — there is no database to migrate.
One store-period reading, as the model receives itSTR-0001-P1.txt · 1 of 117
Synthetic Record
----------------------------------------------------------------
Generated by tools/build_corpus.py from a fixed seed. Every store, franchisee, dollar
and deposit in this file is invented. It reproduces no real franchise system's data.
Store
----------------------------------------------------------------
Store : STR-0001
Franchisee code : FR-888
Trade segment : INLINE
Reporting period : 2026-01 (2026-01-01 to 2026-01-31)
Royalty reporting status : OPEN
Reporting Rule
----------------------------------------------------------------
Four under-reporting indicators are assembled for every open store each period. An
indicator is OUTSIDE ITS BAND when the figure below crosses the stated limit. All four
bands are one-sided; a movement in the other direction is not an indicator.
CASH_MIX cash share of gross sales, below the store's own 12-period baseline
by more than 6.0 percentage points
DEPOSIT_GAP net sales less recorded bank deposits and documented non-deposit items,
more than 3.0 per cent of net sales
VOID_RATE voids, comps and discounts as a share of gross sales, more than 2.0
percentage points above the segment median for the same period
TICKET_DRIFT average ticket below the store's own 12-period baseline by more than
5.0 per cent
An indicator that is outside its band is EXCUSED only by an operating note that names an
approved cause, whose dates fall inside this reporting period, and whose cause covers
that indicator. The approved causes and what each one covers:
POS_OUTAGE point-of-sale terminals unavailable covers CASH_MIX, TICKET_DRIFT
Abridged — the file continues.
The outcomeWhat a good result looks like
Per reading: one of five movements, the indicators outside band on the arithmetic, the ones that are NEW this run, the ones an operating note excuses, and one descriptive sentence. A run raises a store only on NEW_CROSSING; every other movement is described and left standing.
And when it cannot
There is no recorded wrong answer on this corpus, so the honest statement of failure is structural rather than empirical: the two directions that cost different things are a MISSED CROSSING (a period nobody ever looked at, and if the indicator returns to band before the next run, nobody ever will) and a FALSE RAISE (an operator asked to explain a reading that was already explained, which is how a monitor gets switched off). Both are measured and reported apart, never averaged. The free floors fail in the second direction at up to 63.49 per cent, and the cadence ablation shows what the first one costs.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Your field notes are entered against a cause code, always, with dates. Most systems with a structured field-operations queue are here. — the free regex floor, b002, at $0.00 It scored 95.73 per cent on movement and 89.74 per cent on the exemption set with no key, no model and nothing leaving your network. The model's margin over it on this corpus is 4.27 points, all of it on notes with no code.
Your operating notes are prose written by area coaches, and the cause code field is optional or unused. — this kit, one call per store per period This is the only case the measurement supports. The free floor's movement accuracy on prose-only exemptions is 58.33 per cent against the model's 100, over 12 readings.
You want to know WHICH stores are outside their bands this period, and nothing about change. — the memoryless threshold floor, b001, at $0.00 It scores 100.0 per cent on the outside-band set -- the arithmetic is exact and free. Every arm in this kit, model and floor alike, scores the same 100.0 there.
You cannot run on a reliable schedule -- the extract lands late, or a run is skipped when somebody is on leave. — fix the schedule before buying anything Measured, not asserted: skipping one monthly run out of three lost 4 of the 15 crossings in the skipped period PERMANENTLY -- no later run can recover them -- and reported the other 11 in the wrong period. The model was still 100 per cent accurate on the degraded schedule, which is the point: the loss is a property of the cadence and no model quality touches it.
And where nothing here is good enough:
You need the OUTPUT to carry a conclusion -- a recommendation, a severity, an action. — not this kit, and not a variant of it The cap is absolute and has no expansion path: the pack describes indicators and every reply is checked against a closed forbidden-language list. Removing that is not a configuration change, it is a different kit with a different risk posture and different research owed.
At a glanceHow the whole thing runs
100%movement accuracy pct and the four answered fields, per reading
7,201 msp50, end to end
$3.87per 1,000 store-period readings · Google Gemini 3 Flash
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch new royalty problems at your franchise stores14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/rules.py FIRST -- it is your franchise system's reporting rule, not a utility. Every figure on these pages is a property of THIS corpus: 117 readings, 40 stores, four indicators, five approved causes, one note per period at most, at most three runs per store, and prose exemptions that use the approved cause's wording verbatim.Corpus lens →
When is this the wrong choice?
Avoid: Paying per store per period for a judgement a regex over your own cause-code field already makes correctly. That is the case against the best-fitting scenario (“Your field notes are entered against a cause code, always, with dates. Most systems with a structured field-operations queue are here.”). 5 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
More than one operating note in a period. Every document here carries exactly one note or none, and evals/check_labels.py asserts that shape rather than assuming it -- so the kit has never been measured on a period where a store files three notes from three people, one of which contradicts another. 7 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER src/rules.py IS THE RIGHT RULE. It is both the policy the prompt paraphrases and the code the answer key applies, so 100 per cent means the model reproduces this file. 8 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
6 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-royalty-watch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes all twelve pre-flight checks, runs all three free floors and renders the UI. Nothing is fetched and nothing is downloaded. What it cannot do is reproduce the model arm: that needs a provider key, and the committed result files are the evidence for those figures.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
7,201 msp50, end to end
17,504 msp95
3 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed. There is no retrieval step and no index, so the only thing outside the provider call is reading a 3.8 kB file off disk and four float comparisons. The figures are over the 117 readings of r001-royalty-watch, run 12 wide -- a whole scheduled run took 98.9 seconds of wall clock.
Current processWhat it replaces
Somebody re-reading every open store's royalty file each month, working out which indicators are outside their bands, remembering which of those were already outside last month, and deciding whether the store's own operating note excuses the ones that are new.
Where it is not good enough
⚑ IT SCORED 100 PER CENT ON EVERY FIGURE AND THAT IS A STATEMENT ABOUT THE CORPUS, NOT ABOUT THE MODEL. 117 of 117 readings correct on movement, 23 of 23 crossings raised, 0 false raises, 0 conclusory replies. A perfect score means the labelled set CANNOT SEPARATE this model from a better one, and it cannot tell you where the model would break -- the taxonomy of failures below is empty because there were none to put in it, which is the least useful state an eval can be in. The strongest free floor already reached 95.73 per cent, so the whole measured margin is 4.27 points and it lives in exactly one place: 12 readings whose operating note describes an approved cause in plain words and enters no cause code. Twelve readings is a thin denominator for the only claim this kit makes. A corpus with ambiguous prose -- a note that half-describes an approved cause, or names two -- would separate these arms and this one does not have any.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt117jsonl1
40 franchised stores, re-read on every scheduled run
Recorded failurethe strongest free floor still raises 5 stores it should not — every one a note that names an approved cause in words and enters no code
Recorded failure0 model failures on this corpus — which is a criticism of the corpus: an empty failure taxonomy cannot separate this tier from a better one
It produces a change column beside a flag column a spreadsheet already fills, for a franchise auditor to read. It opens no case, issues no notice and contacts nobody, and every reply is checked against a closed list of conclusory language — 0 of 117 on this run, with the checker red-proven in both directions before the run was allowed to spend. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is four fields written by src/state.advance() from the parsed figures and never from the model's answer — the reply is not an argument to it, which is why the whole schedule's state is computable before a single call and why every reading in a run fires concurrently. The clock is the second: one run owns one reading of every open store for exactly one period, and evals/check_labels.py asserts every (store, period) is owned exactly once.
⚠︎ THE FREE FLOOR IS THE STORY ON COST AND THE PAGE LEADS WITH IT. Threshold arithmetic alone scores 32.48 pct on movement; add last period's flag column and it is 92.31; add a regex over the cause codes, their dates and the coverage map and it is 95.73, all three for $0.00. The model's 100.0 is a 4.27-point margin and it lives entirely in 12 readings whose operating note describes an approved cause in plain English and enters no cause code — the free floor gets 58.33 pct of those and the model gets all of them.
⚠︎ AND THE STATELESS CONTROL IS A CONFOUND, NOT A COLLAPSE: it scores 0.0 pct because all 77 of its readings were answered FIRST_READING, which is the correct answer under the instruction's own first precedence test when no earlier run is available — and that arm removed it. It proves the state block is the only route between runs and that the model invents no history it was not given (0 false raises in 77); the memory figure comes from the two free floors instead. The four indicator bands and five approved exemption causes are INVENTED for this kit and reproduce no franchise agreement, disclosure document or state relationship statute; anchor research on default and termination process is owed and has not been done. No red-team run exists for this kit; this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the reporting rule
src/rules.py
Your bands, your approved causes, your coverage map and your own forbidden-language list. Replace this file FIRST -- it is a policy somebody wrote down, not a utility, and every figure on these pages is a property of the one shipped here.
the population and its readings
tools/build_corpus.py
Point it at your own extract instead of generating one. The only contract downstream is the seven section headings and the printed figure labels src/figures.py reads.
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic; adding a third is one function and one entry in PROVIDERS.
the schedule
src/cadence.py
FREQUENCY and RUNS_AFTER_PERIOD_CLOSE_DAYS. Five days is this corpus's settlement-lag assumption; a franchisor whose depository posts same-day would set it to 1, and a weekly cycle changes what one run owns.
what is never sent
src/select.py
NEVER_SENT and SECTION_HINTS. Adding a section to a document does not add it to the prompt -- a field has to ask for it.
the carried state
src/state.py
What persists between runs. It is four fields today; adding a fifth changes the prompt by one clause and the cost by nothing measurable, because the state is constant-size by construction.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Writes 117 period files across 40 stores from a fixed seed, solving each store's printed figures backwards from a scenario table so that exactly the intended indicators breach. It also plants the six operating-note classes, which are the experiment. Free and offline.
the printed figures, parsed
src/figures.py
Reads the eleven printed numbers off a period file and derives the four indicator distances from the ROUNDED values the page shows. The model is never asked to divide one dollar figure by another.
the approved cause list
src/exemptions.py
Five causes, their plain-English wording and the indicators each can excuse, as data. Printed in full inside every document, and check_labels reads it back off the page to prove the two copies have not drifted.
the note reader, twice
src/notes.py
claims() finds cause CODE tokens and the dates beside them -- what a free regex can see. claims_full() also matches an approved cause written in plain words. The gap between the two functions IS the measurement, which is why they are deliberately not one function.
the reporting rule
src/rules.py
The four bands, the three-part exemption test, the five-step movement precedence, and FORBIDDEN -- the closed list of language that would turn a description into an allegation. The answer key and the first file a forker replaces.
the carried state
src/state.py
Four fields written from the ARITHMETIC and never from a model reply, rendered as one English sentence. Constant size whatever the history length, so the twelfth run of a store costs what the second did.
the clock
src/cadence.py
The schedule as code: monthly, five days after period close; one run owns one reading of every open store for exactly one period; gap() and missed() say how many scheduled runs this one is standing in for.
the section split
src/segment.py
Splits a period file into its seven named sections. The shape is a property of the corpus builder and evals/check_labels.py asserts it on every document, so a parser drift shows up as a refusal to start rather than as a quietly truncated prompt.
what is withheld
src/select.py
Sends five of the seven sections. Franchisee Contact (a named individual, address, telephone, email) and Audit File Access are mapped by no field, and the subtraction is UNCONDITIONAL so the fallback path -- the one that fails quietly -- cannot reintroduce them.
the prompt
src/prompt.py
Four parts in assembly order: the question, the schedule, the carried state, the reading. The stateless arm REPLACES the third block rather than deleting it, so the control does not also measure a missing heading.
one reading, one call
src/monitor.py
Assembles, calls, parses. Holds MAX_TOKENS = 20000, the published ceiling a scored run must use.
the answer key
evals/answers.py
Generates the key by WALKING A SCHEDULE. The correct answer to a March reading depends on which runs happened, so a static gold file would be silently wrong for the one experiment this kit exists to run.
the three free floors
evals/baseline.py
Threshold only; threshold plus memory; threshold plus memory plus a regex note reader. All three scored by the same grader on the same readings.
the grader
evals/scoring.py
Six figures, never averaged, plus conclusory() -- the negative assertion, checked on every reply against src/rules.FORBIDDEN. No model, no key, $0.00.
the pre-flight
evals/check_labels.py
Twelve checks that must pass before a run may spend, including a red proof of the language checker in both directions and a proof that the missed-run key actually differs from the full one.
the harness
evals/run.py
Fires a schedule. Because the state is arithmetic, every reading is independent and they run concurrently -- the guardrail visible in the shape of the loop.
the local UI
src/app.py
http.server on port 9003, serving ui/index.html, ui/app.js and ui/app.css -- hand-written, no build step. Shows the carried state verbatim, the free floor beside the answer, the access log, and which sections were withheld. Renders with no key: only /api/assemble calls a provider, and without one it returns 200 with a sentence rather than an error.
Where it breaks at scale
NOT ON HISTORY LENGTH, WHICH IS THE UNUSUAL HALF. The carried state is one boolean, a set of at most four codes, an integer and a period label, so the prompt for a store's twelfth run is the same size as its second and the bill is FLAT in history. That is the opposite curve from a conversation kit, where the prefix is re-sent whole every turn.
It breaks on POPULATION and on CADENCE, and they multiply. Cost is exactly linear in (open stores x runs per year): 117 readings cost $0.4528 projected on the shared card, so a 2,000-store system read monthly is 24,000 readings a year and about $93. A weekly cadence on the same book is four times that for the same number of stores, which is the decision this kit exists to make legible -- and a run that does not finish before the next one is due is a missed run, whose cost is measured in a000-royalty-watch-missedrun.
The second ceiling is the note reader. One operating note per period is this corpus's shape and it is asserted by the pre-flight, not assumed. A real store files several notes a period, from several people, some of them contradicting each other -- and nothing here has been measured on that. The third is the ONE-RULEBOOK assumption: every store in this population is judged against the same four bands and the same five causes, printed on every page. A franchise system with different terms by cohort, region or agreement vintage has N rulebooks, and the prompt carries one.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
STR-0006-P2, chosen by reading the answer key rather than by looking for a flattering frame -- it is this kit's entire margin over free code, on one screen. The average ticket is 6.10 per cent below the store's own baseline, so TICKET_DRIFT is outside its band on the arithmetic and BOTH columns agree about that. The operating note says the dining room was closed for an approved remodel from 2026-02-03 to 2026-02-21 and enters NO CAUSE CODE. The free regex floor sees no code, excuses nothing and raises the store: NEW CROSSING. The model reads the sentence, applies the approved-cause list printed on the same page, and answers NO CHANGE. The row beneath says the arithmetic was 'same -- free code got here too', which is the page refusing to take credit it did not earn.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same store with NO API_KEY configured. Nothing was called and the page says so in a sentence rather than throwing -- but look at what is left: the free regex floor's answer, NEW CROSSING, sitting in the only populated column. That is the wrong answer, and it is the answer a reader gets for free. A kit that rendered an empty page here would be hiding the fact that its free alternative is confidently wrong on exactly the readings the kit exists for.failureOpen full size →The stateless control arm, drawn from its committed result file by tools/shoot_terminal.py. 0.0 per cent movement accuracy over 77 readings, 0 of 23 crossings raised -- and it is a CONFOUND, not a score. Every one of the 77 was answered FIRST_READING, which is what the instruction's own first precedence test says to answer when no earlier run is available, and this arm removed the earlier run. What it does establish is worth having: the carried state is the ONLY route from one run to the next, and the model did not invent a history it was not given -- 0 false raises in 77. What it CANNOT tell you is how much worse the kit is without memory; that number comes from the free floors instead.failureOpen full size →
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
117store-period readings
0.43 MiBjsonl 1 · txt 117
p50 1chars per store-period readings (one file, one call)
$0.00setup · 0.0s
How it is cutWhat one store-period readings (one file, one call) is
No split and no chunking. The unit is one STORE-PERIOD READING -- a whole period file goes into one call, with the previous run's carried state beside it. What varies is not the size of the input but the SCHEDULE it sits in: 40 readings are one store's first sight by this monitor and cannot be a change, and 77 have a previous run to be compared against.
SetupWhat the setup figure measured
There is no index and no retrieval step. The population is re-read WHOLE on every scheduled run -- that is what a monitor is -- so there is nothing to build, nothing to warm and nothing to invalidate. tools/build_corpus.py writes 117 files and the harness reads all of the ones a run owns.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved, because there is no third-party material here to grant. The pack reproduces no real franchise agreement, no franchise disclosure document and no state franchise-relationship statute, and quotes from none.
Bring your ownBring your own store-period readings
Replace src/rules.py FIRST -- it is your franchise system's reporting rule, not a utility. Its four bands, its five approved causes and the indicators each cause covers are a policy somebody wrote down and yours will differ in all three. Then point tools/build_corpus.py at your own extract, keeping the seven section headings and the printed figure labels src/figures.py reads. Then set src/cadence.py's FREQUENCY and settlement lag. Everything above those three files -- the state, the prompt, the answer key, the floors and the grader -- is written against the rule rather than against these numbers, and needs no edit.
⚠︎ And what stops being true when you do: Every figure on these pages is a property of THIS corpus: 117 readings, 40 stores, four indicators, five approved causes, one note per period at most, at most three runs per store, and prose exemptions that use the approved cause's wording verbatim. The 100 per cent scores in particular do not travel -- they say the labelled set could not separate this model from a better one, and a corpus with paraphrased or contradictory notes would separate them. The measured margin over free code is 4.27 points and rests on 12 readings.
What breaks it
More than one operating note in a period. Every document here carries exactly one note or none, and evals/check_labels.py asserts that shape rather than assuming it -- so the kit has never been measured on a period where a store files three notes from three people, one of which contradicts another. src/notes.py returns a LIST and the rule takes the first claim that covers an indicator, so a later note withdrawing an earlier one is silently ignored. That is a real deployment case and it is unmeasured here.
A note that half-describes an approved cause. claims_full() matches the approved cause's wording as a SUBSTRING, so 'the point-of-sale terminals were unavailable' hits and 'the tills were down' does not. Every prose exemption in this corpus uses the approved wording verbatim, which is the easiest version of the problem and the reason the model scores 100 per cent on it. Paraphrase is where this kit's only claim would actually be tested and there is no paraphrase in the corpus.
More than one rulebook. Every store is judged against the same four bands and the same five causes, printed on every page. A franchise system with different terms by cohort or agreement vintage has several, and nothing here selects between them.
Indicators that move together for ordinary trading reasons. The generator pushes indicators outside band independently. In real trading a promotion moves the void rate, the average ticket and the cash mix at once, and a monitor that raises three indicators for one cause is three times as annoying as one that raises one. Not produced, not measured.
⚠︎ A KNIFE-EDGE READING. The pre-flight REFUSES a corpus where any derived figure lands within 0.15 of its threshold, so nothing here tests what happens at 6.01 points against a 6.0 band. That is a deliberate exclusion -- an answer key that turns on a rounding decision measures the rounding -- and it means the kit has no evidence about the boundary, which is where a real dispute starts.
⚠︎ A HISTORY LONGER THAN THREE RUNS. Every store is read at most three times. The carried state is constant-size so nothing should change, but 'should' is not a measurement: there is no run here in which an indicator has been standing for eight periods, and nothing has been measured about whether the model keeps calling it STILL_STANDING or eventually re-raises it.
⚠︎ AN OUT-OF-ORDER OR DUPLICATED RUN. cadence.owned() enumerates one reading per open store per period and check_labels asserts each is owned exactly once, so a scheduler that fires twice for the same period would re-read readings already carried into the state and report every standing indicator as fresh. The assertion is on the SCHEDULE, not on a deployment's scheduler, and no deployment scheduler has been tested.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
the question and the JSON shape
3,057
not measured
the schedule, and which period this run owns
220
not measured
what the last run left behind
106
not measured
this period's file, minus what is never sent
3,265
not measured
Total
1,647
This is the cost lesson as arithmetic: of the 6,648 characters assembled, 3,265 are documents — 49% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from the kit's own src/prompt.py::parts_of on STR-0001-P1 with the exact carried state the answer key gives that reading, and VERIFIED against what the run wrote down: results/eval-r001-royalty-watch.json records the four part sizes (instruction 3057 / schedule 220 / carried_state 106 / reading 3265) and the replay reproduces them exactly. The published raw_response is that same reading's recorded reply, so the prompt and the answer on this page are two halves of one call.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are assembling under-reporting INDICATORS for one franchised store, for one reporting period,
as one scheduled run of a monitor that reads every open store each period.
The four indicators, their bands and the approved exemption causes are reproduced in the reading
below. Apply them exactly as written. Every figure you need is printed; do not estimate, and do not
use any figure that is not on the page.
What makes this a monitor rather than a classification: you are not asked whether this store looks
bad, you are asked WHAT MOVED since the last scheduled run. What the last run saw is stated under
"Carried state" and is the only history available to you. Do not assume anything about earlier runs
beyond it.
How to decide:
- An indicator is OUTSIDE ITS BAND on the arithmetic alone. List every one that is, whether or not
an operating note excuses it.
- An operating note EXCUSES an outside-band indicator only if all three hold: the cause is on the
approved list in the reading, its dates fall inside this reporting period, and the cause covers
that indicator. A cause not on the list excuses nothing however reasonable it reads; a cause
whose dates closed before this period opened excuses nothing; a cause that covers a different
indicator excuses nothing. A note may state an approved cause in plain words without naming its
code -- that still counts, provided the words match one of the approved causes.
- Then choose exactly one movement, testing in this order and stopping at the first that holds:
FIRST_READING the carried state says no earlier run has read this store.
NEW_CROSSING an indicator that is outside its band, is NOT excused, and was NOT outside
at the last run.
STILL_STANDING an indicator outside its band now that was ALSO outside at the last run.
RETURNED_TO_BAND an indicator that was outside at the last run and is inside now.
NO_CHANGE none of the above.
The carried state records what was outside on the ARITHMETIC at the last run, excused or not --
so an excuse expiring is not a new crossing.
Answer with a single JSON object and nothing else:
{"movement": "FIRST_READING|NEW_CROSSING|STILL_STANDING|RETURNED_TO_BAND|NO_CHANGE",
"indicators_outside": [indicator codes outside their band on the arithmetic, alphabetical],
"indicators_new": [the codes that make this a NEW_CROSSING; empty for every other movement],
"exemptions_applied": [the codes of indicators an operating note excuses this period, alphabetical],
"description": "one sentence describing what the figures show and what moved"}
"description" is DESCRIPTIVE ONLY. State what the figures are and what changed since the last run.
Do not state or imply a cause, a conclusion, wrongdoing, an allegation, or any action that should
follow. Do not use words such as under-reporting, fraud, concealment, misstatement, breach, default,
termination or demand. This pack describes indicators for a human auditor to decide about; it does
not decide anything.
Schedule
----------------------------------------------------------------
Runs monthly, 5 days after each period closes. One run owns one reading of every open store for exactly one period; the previous run's carried state is what makes that reading a change. This run owns reporting period P1.
Carried state
----------------------------------------------------------------
This store has not been read by any earlier run. There is no previous reading to compare this one against.
This period's reading
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Generated by tools/build_corpus.py from a fixed seed. Every store, franchisee, dollar
and deposit in this file is invented. It reproduces no real franchise system's data.
Store
----------------------------------------------------------------
Store : STR-0001
Franchisee code : FR-888
Trade segment : INLINE
Reporting period : 2026-01 (2026-01-01 to 2026-01-31)
Royalty reporting status : OPEN
Reporting Rule
----------------------------------------------------------------
Four under-reporting indicators are assembled for every open store each period. An
indicator is OUTSIDE ITS BAND when the figure below crosses the stated limit. All four
bands are one-sided; a movement in the other direction is not an indicator.
CASH_MIX cash share of gross sales, below the store's own 12-period baseline
by more than 6.0 percentage points
DEPOSIT_GAP net sales less recorded bank deposits and documented non-deposit items,
more than 3.0 per cent of net sales
VOID_RATE voids, comps and discounts as a share of gross sales, more than 2.0
percentage points above the segment median for the same period
TICKET_DRIFT average ticket below the store's own 12-period baseline by more than
5.0 per cent
An indicator that is outside its band is EXCUSED only by an operating note that names an
approved cause, whose dates fall inside this reporting period, and whose cause covers
that indicator. The approved causes and what each one covers:
POS_OUTAGE point-of-sale terminals unavailable covers CASH_MIX, TICKET_DRIFT
BANK_CUTOFF bank deposit cut-off moved by the depository covers DEPOSIT_GAP
REMODEL_CLOSURE dining room closed for approved remodel covers TICKET_DRIFT
SYSTEM_PROMO system-wide promotion priced by the franchisor covers VOID_RATE, TICKET_DRIFT
DELIVERY_ONBOARD third-party delivery channel onboarding covers CASH_MIX, TICKET_DRIFT
A cause not on this list excuses nothing, however reasonable it reads. A cause on the
list whose dates closed before this period opened excuses nothing. A cause on the list
that does not cover the indicator in question excuses nothing.
This pack DESCRIBES indicators. It states no conclusion about the store, the franchisee
or the cause of any reading, and no output of this pack may be used toward a formal
demand or notice.
Period Figures
----------------------------------------------------------------
Gross sales : 166,018.40 USD
Net sales : 150,360.39 USD
Cash share of gross sales : 15.4 %
Cash share, 12-period baseline : 13.9 %
Bank deposits recorded : 148,435.43 USD
Documented non-deposit items : 992.73 USD
Voids, comps and discounts : 2.4 % of gross sales
Segment median, same period : 2.6 %
POS transaction count : 13,271
Average ticket : 11.33 USD
Average ticket, 12-period base : 11.50 USD
Operating Notes
----------------------------------------------------------------
No operating note was filed for this period.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"movement": "FIRST_READING","indicators_outside": [],"indicators_new": [],"exemptions_applied": [],"description": "This period all four indicators are within their bands and no operating note was filed; this is the store's first reading, so there is no prior state to compare against."}
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch new royalty problems at your franchise stores — 117 store-period readings. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
There is no LLM-as-judge in this kit and that is not a shortcut. Every field the model answers is a word from a closed list or a set of codes from a closed list, so 'correct' is exact match and a judge would be an expensive way to compare two strings. The one free-text field -- the description -- is not graded for quality either; it is checked against a closed list of forbidden phrases, which is again exact matching. Where judgement genuinely lives is in src/rules.py, and that is stated as a risk in validated rather than dressed up as a measurement.
117store-period readings
117source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED117 · 39 · 0 · 24 / 117movement accuracy pct — store-period readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED77 / 77change movement accuracy pct — readings that HAVE a previous run -- where memory can matter at allDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED23 · 15 / 23change detection pct — readings whose answer key is a genuine new crossingDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 23false raise rate pct — readings the model raisedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED117 / 117outside set exact pct — readings, exact set match on the arithmetic the free floors also get rightDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED117 / 117exemption set exact pct — readings, exact set match on what an operating note excusesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 117conclusory language rate pct — replies checked against src/rules.FORBIDDEN -- the negative assertion; target 0.0Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 117unparsed replies — calls -- every reply parsed; largest 3469 output tokens under a 20000 capDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
⚠︎ THE COMPARISON CANNOT BE WRONG ABOUT ITSELF; THE RISK IS ENTIRELY IN THE ANSWER KEY, AND THE KEY IS CODE THAT ALSO DEFINES THE TASK. src/rules.py holds both the policy the prompt describes in English and the arithmetic the key applies, so a kit that scores 100 per cent has demonstrated that the model reproduces src/rules.py -- not that src/rules.py is the right rule. Two things narrow that, neither of them eliminates it. First, tools/build_corpus.py states independently what it PLANTED in each document (planted_outside, note_class) and evals/check_labels.py asserts rules.outside() agrees with it on all 117 readings -- two statements written for different purposes that have to match. Second, the five-step precedence is written out in both places, in code in src/rules.py and in English in the prompt, and they were reconciled by reading rather than by testing. What is NOT validated is whether a franchise auditor would read the written rule the way the code does; nobody outside this repository has been asked.
Run it twiceThe same set, run again
THE STRUCTURED VERDICT IS STABLE AND THE PROSE IS NOT, and a reader needs both halves. Over the 24 readings both runs share, 0 changed on ANY of the four answered fields -- movement, outside set, new set, exemption set. Over the same 24 readings, 0 descriptions were byte-identical. So the thing this kit is scored on reproduces exactly, and the sentence beside it is rewritten every time: do not diff two runs' descriptions and read the difference as a change of view.
Run date
the fast tier, first pass
the fast tier, the same 24 readings asked again
2026-08-23
100.0% r001-royalty-watch
100.0% p002-royalty-watch-repeat
movement_accuracy_pct and the four answered fields, per reading — There is nothing to average -- both passes scored 100.0 per cent on every field. Averaging identical numbers would imply a spread that was measured and found to be zero on 24 readings, which is a much stronger claim than one probe supports.
What did not move
Everything deterministic reproduced to the digit: the answer key, all three free floors, the arithmetic and the section split are byte-identical across the two runs, because none of them calls anything.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One store-period reading
1,000 store-period readings
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.003870
$3.87
21%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.001548
$1.55
21%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.067251
$67.25
24%
Same work, 43× the bill
The same store-period readings, the same tokens — only the rate card changed. And across all 3 cards between 21% and 24% of what you pay is the prompt this pipeline sends, not the answer it writes.
The cadence. Everything else here is a few per cent; this one is a multiple -- weekly instead of monthly is 4x the same book, and it is the one knob a buyer actually sets. The second lever is provider-side reasoning at 90.6 per cent of output tokens, unmeasured with it disabled.
Rates checked 2026-08-18. The provider that actually ran all 263 calls behind this kit is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. A MEASURED $0.00, not an unpriced one: evals/scoring.py imports src.rules and compares dictionaries, the three floors it is compared against import src.notes and compare dictionaries, and evals/check_labels.py runs twelve assertions. Not one of them constructs an adapter. A forker re-running the whole evaluation pays for the pipeline and nothing on top.
The gradersThree ways to grade
⚑ READ THE FLOOR FIRST AND READ ALL THREE. The strongest free floor scores 95.73 per cent on movement against the model's 100.0 -- a margin of 4.27 points -- and it costs nothing, needs no key and sends nothing anywhere. If your field notes are always entered against a cause code, run b002 and do not pay for this kit. That is the honest headline and it is printed above the model's numbers on purpose.
THE THREE FLOORS ISOLATE THE TWO THINGS SEPARATELY, WHICH THE MODEL ARMS COULD NOT.
b001 threshold only, NO MEMORY movement 32.48 pct, false raises 63.49 pct (40 of 63)
b000 threshold PLUS memory movement 92.31 pct, false raises 28.12 pct (9 of 32)
b002 the above PLUS a regex note reader movement 95.73 pct, false raises 17.86 pct (5 of 28)
b001 to b000 is MEMORY and nothing else -- both are blind to the notes -- and it is worth 59.83 points of movement accuracy. b000 to b002 is NOTE READING and nothing else, worth a further 3.42. The model's margin over b002 is the third thing: 12 readings whose operating note describes an approved cause in plain words and enters no code, which b002 scores 58.33 per cent on and the model scores 100 on.
⚠︎ AND THIS IS WHY THE STATELESS CONTROL IS NOT THE MEMORY MEASUREMENT. s001-royalty-watch-stateless scores 0.0 per cent, which looks like a 100-point collapse and is not one: all 77 of its readings were answered FIRST_READING, which is the correct answer under the instruction's own first precedence test when no earlier run is available -- and that arm removed the earlier run. The arm proves the carried state is the only route between runs and that the model does not invent a history it was not given (0 false raises in 77). It does not measure what memory is worth, and the free floors do, so the memory figure on these pages comes from b001 against b000.
the fast tier, with the carried state 100.0% · the fast tier, one scheduled run MISSED 100.0% · the strongest free floor, no model 95.7% · the threshold floor with no memory, no model 32.5%
the fast tier, with the carried state 0.0% conclusory language rate · the fast tier, memory removed (THE CONTROL) 0.0% conclusory language rate · the fast tier, one scheduled run MISSED 0.0% conclusory language rate · the strongest free floor, no model 0.0% conclusory language rate
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
NOT AT ALL, IN THE DIRECTION THAT MATTERS MOST, AND THE KIT SAYS SO RATHER THAN CELEBRATING THE SCORE. Every model arm scored 100 per cent on every field, so this labelled set cannot distinguish the shipped model from a better one, cannot rank two models, and produced an empty failure taxonomy. A set that separates nothing at the top is a set that will not notice a regression until it is large.
It separates decisively in three OTHER directions, all of them measured on the same 117 readings by the same grader. MEMORY: b001 to b000, 32.48 to 92.31 per cent on movement, with note reading held constant at blind. NOTE READING: b000 to b002, 92.31 to 95.73 per cent, with memory held constant. PROSE: b002 to the model, 95.73 to 100.0, and the by-note-class table localises the whole of it to 12 readings.
And it separates decisively on CADENCE, which is not a model comparison at all: the same model on a schedule with one run missing is still 100 per cent accurate and still loses 4 crossings forever.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Your field notes are entered against a cause code, always, with dates. Most systems with a structured field-operations queue are here.
the free regex floor, b002, at $0.00
It scored 95.73 per cent on movement and 89.74 per cent on the exemption set with no key, no model and nothing leaving your network. The model's margin over it on this corpus is 4.27 points, all of it on notes with no code.
Paying per store per period for a judgement a regex over your own cause-code field already makes correctly.
Your operating notes are prose written by area coaches, and the cause code field is optional or unused.
this kit, one call per store per period
This is the only case the measurement supports. The free floor's movement accuracy on prose-only exemptions is 58.33 per cent against the model's 100, over 12 readings.
Reading the headline as a general result. It rests on 12 readings whose prose uses the approved cause's wording verbatim; paraphrase is untested.
You want to know WHICH stores are outside their bands this period, and nothing about change.
the memoryless threshold floor, b001, at $0.00
It scores 100.0 per cent on the outside-band set -- the arithmetic is exact and free. Every arm in this kit, model and floor alike, scores the same 100.0 there.
Reading its 32.48 per cent movement accuracy as a failure of the idea. It is not trying to answer that question and it raises 63 of 117 readings doing so.
You need the OUTPUT to carry a conclusion -- a recommendation, a severity, an action.
not this kit, and not a variant of it
The cap is absolute and has no expansion path: the pack describes indicators and every reply is checked against a closed forbidden-language list. Removing that is not a configuration change, it is a different kit with a different risk posture and different research owed.
Reading conclusory_language_rate_pct 0.0 as a quality score. It is a constraint being met, and the floors meet it too by being unable to write a sentence.
You cannot run on a reliable schedule -- the extract lands late, or a run is skipped when somebody is on leave.
fix the schedule before buying anything
Measured, not asserted: skipping one monthly run out of three lost 4 of the 15 crossings in the skipped period PERMANENTLY -- no later run can recover them -- and reported the other 11 in the wrong period. The model was still 100 per cent accurate on the degraded schedule, which is the point: the loss is a property of the cadence and no model quality touches it.
Buying accuracy to fix a reliability problem. On this corpus a missed run costs 27 per cent of that period's crossings and the best model in the world costs 0 per cent of them.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
none_observed
No model failure was recorded on this corpus
0
There is nothing to quote. All 117 readings of r001-royalty-watch matched the answer key on all four fields, all 39 of the missed-run ablation matched, and all 24 of the repeat probe matched. An empty taxonomy is a fact about the corpus's difficulty, not…
floor_false_raise_no_code
THE FREE FLOOR's failure, which is the one this kit exists to measure
5
STR-0006-P2: "The dining room closed for approved remodel from 2026-02-03 to 2026-02-21 ... No cause code was entered against the period." b002 finds no cause code token, excuses nothing and answers NEW_CROSSING. The key says NO_CHANGE. 5 of b002's 28 raises…
floor_memoryless_overraise
The memoryless floor raising a standing indicator as if it were new
40
Any store on the archetype where VOID_RATE is outside band in both P2 and P3 -- b001 answers NEW_CROSSING both times, because it has nothing to compare against. 40 of its 63 raises are false.
cadence_permanent_loss
A crossing that appeared in a skipped run and returned to band before the next one
4
STR-0003, STR-0009, STR-0023, STR-0029. Their P2 crossing is invisible to a schedule that skipped P2, because the reading it would have been compared against was never taken. No model quality recovers this and no later run can.
cadence_misdated
A crossing that started in a skipped run and is reported as new in the next one
11
STR-0002, STR-0007, STR-0008, STR-0011, STR-0015, ... -- raised, so nothing is missed, but dated to P3 when the indicator actually crossed in P2. The period is the date an auditor works back from.
What we could NOT verify
WHETHER src/rules.py IS THE RIGHT RULE. It is both the policy the prompt paraphrases and the code the answer key applies, so 100 per cent means the model reproduces this file. Nobody outside this repository -- no franchise auditor, no franchisor's finance team -- has been asked whether the five-step precedence or the three-part exemption test matches how the question is actually decided. The corpus generator's independent statement of what it planted narrows the risk to the RULE and does not touch it.
WHETHER THE MODEL WOULD HOLD UP ON PARAPHRASE. Every prose exemption in this corpus states the approved cause in the approved wording. The one claim this kit makes rests on 12 readings of verbatim prose, and a note reading 'the tills were down all week' was never put in front of it.
WHETHER THE 100 PER CENT WOULD SURVIVE A SECOND, HARDER CORPUS. One repeat probe over 24 readings found zero movement, which says the answers are stable and says nothing about whether they are robust. Stability under repetition and robustness under distribution shift are different properties and only the first was measured.
WHAT PROVIDER-SIDE REASONING IS BUYING. 90.6 per cent of this run's output tokens were reasoning, left at the provider's default and never disabled. The adapter can send a disabled-thinking field and this kit never has, so nobody knows whether the same answers come back at a fraction of the output bill.
WHAT A LONGER HISTORY DOES. No store here is read more than three times. Whether an indicator standing for eight consecutive runs keeps being reported STILL_STANDING, or whether the model eventually re-raises it, is unmeasured -- and it is the single most likely source of alert fatigue in a real deployment.
WHETHER A SECOND OPERATING NOTE CHANGES THE ANSWER. Every document carries at most one, asserted by the pre-flight. src/notes.py returns a list and the rule takes the first covering claim, so a later note withdrawing an earlier one is silently ignored. Nothing was run on that shape.
WHETHER THE CADENCE COST GENERALISES. 4 of 15 crossings lost is one number from one skipped period on one corpus, and it depends entirely on how many crossings return to band within one period -- a property of the population, not of the monitor. A book whose indicators stay outside band for months would lose almost nothing to a missed run.
WHETHER ANY OF THIS IS LAWFUL TO ACT ON. Anchor research on state franchise-relationship law bearing on default and termination process is OWED and has not been done. The monitoring runs; whether a franchisor may take any step on the strength of it is outside everything measured here, and is named as a precondition on the kit's own pages.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,646.67
1,015.68
7,201 ms
$0.003870
$0.001548
$0.067251
the same tier, one scheduled run MISSED (the ablation)
1,653.26
1,159.87
7,215 ms
$0.004306
$0.001722
$0.074526
the same tier, memory removed (THE CONTROL)
1,627.17
999.94
7,429 ms
$0.003813
$0.001525
$0.066269
the strongest free floor, no provider
0.0
0.0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
⚑ GRADING IS FREE AND THE EVIDENCE AROUND THE HEADLINE IS NOT. The scored run is $0.4528 projected; everything that makes it believable -- the control, the cadence ablation, the repeat probe and the calibration -- is another $0.5651, so the headline is under half the model spend on this kit. The three free floors and the grader add exactly $0.00, measured: no code path in evals/ constructs an adapter. Two further calls were spent on the published screenshot, one of them discarded when the frame was retaken at a shorter height.
Cost driversWhat actually moves the bill
PROVIDER-SIDE REASONING, 90.6 per cent of output tokens (107,628 of 118,834) and therefore most of the bill. It was left at the provider's default on every published run. Output is 78.7 per cent of the projected cost on the shared card, so most of what this kit charges for is the model reasoning about four bands and one note.
THE NUMBER OF SCHEDULED RUNS, which is a decision and not a fact. Cost is exactly (open stores x runs per year) x $0.003870. Moving from monthly to weekly quadruples the bill on an unchanged book, and it is the only lever here that changes the bill by a multiple.
THE REPORTING RULE PRINTED IN EVERY DOCUMENT. The Reporting Rule section is reproduced on every page and goes into every prompt -- roughly 1780 of the 3265 characters of the reading block. A deployment could hoist it into the system prompt once; this kit does not, because a page that carries its own rule can be read on its own, and a forker diffing two documents can see the rulebook they were judged against.
NOT THE CARRIED STATE. It is 106 characters and constant in history length, so the twelfth run of a store costs what its second did. On this corpus the stateful and stateless arms differ by 19 input tokens a reading.
Your volumeWhat it costs at your volume
LINEAR IN READINGS, AND FLAT IN HISTORY. A reading is one call whatever the store's history, so 10x the stores is 10x the bill and 10x the runs is 10x the bill, with no term that grows faster. 117 readings cost $0.4528 projected on the shared card; a 2,000-store system read monthly is 24,000 readings a year, about $93.
⚠︎ THE THING THAT DOES NOT SCALE LINEARLY IS THE WALL CLOCK AGAINST THE SCHEDULE. This run took 98.9 seconds for 117 readings at 12 workers. 24,000 readings at the same width is about 6 hours, which is comfortably inside a month and not inside a night -- and a run that does not finish before the next one is due IS a missed run, whose cost is measured in a000-royalty-watch-missedrun rather than guessed at.
Where pricing changes shape
THE OUTPUT CEILING IS A CLIFF AND THIS KIT MEASURED HOW CLOSE IT SITS. c000-royalty-watch-calibration ran four readings at max_tokens=2000 and the largest reply used 1752 -- 88 per cent of the cap, on four calls. The published ceiling is 20000 and the scored run's largest reply was 3469, so the headroom is real, but a forker who sets a 2,000 cap to save money will lose replies to truncation on the hard readings and an unparsed reply counts as a miss in every field.
A CEILING IS NOT A COST. The provider bills tokens produced, not tokens allowed, so the 20000 cap costs nothing on the 117 readings that finish in a few hundred tokens. Setting it low is a false economy in both directions.
PROVIDER-SIDE REASONING AT THE DEFAULT. If a vendor's default changes, or a forker disables it, the output bill moves by up to 90.6 per cent and this kit has NOT measured what that does to the answers.
Your return, with your numbers
Volumestore-period readings per scheduled run -- this run owned 117 over 40 stores and three monthly periods. Yours is (open stores) x (runs per year).
What it replacessomebody re-reading every open store's royalty file each period, computing four indicators, remembering which were already outside band last period, and deciding whether the store's own operating note excuses the new ones
Time saved per itemnot measured here -- it depends entirely on whether your notes carry cause codes. If they do, the free floor already does this at $0.00 and the time saved by the model is zero; the fitment table says which case you are in.
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared key is pointed at, and it is the cheap one -- the right place to start on a task whose free floor already scores 95.73 per cent. A more expensive tier cannot be justified from this corpus in either direction: the cheap one already scored 100 on every field, so there is nothing above it to buy, and there is no evidence here about what a cheaper or a local model would do.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
192,660input tokens · this run
118,834output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every model figure on these pages: 117 readings, one completion call each, one tier, one healthy schedule. The 77-call stateless control, the 39-call cadence ablation, the 24-call repeat probe, the 4 calibration calls and the 2 screenshot calls are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.181
$0.181
$1.55
2026-09-12
gemini-3-flash
Google
$0.453
$0.453
$3.87
2026-09-18
gemini-3-8-flash
Google
$0.590
$0.590
$5.04
2026-09-18
llama-5
Meta
$0.746
$0.746
$6.37
2026-09-18
claude-haiku-4-5
Anthropic
$0.787
$0.787
$6.73
2026-09-12
grok-4-5
xAI
$1.098
$1.098
$9.39
2026-09-18
grok-4-6
xAI
$1.098
$1.098
$9.39
2026-09-18
claude-sonnet-5
Anthropic
$1.574
$1.574
$13.45
2026-09-12
gemini-3-1-pro
Google
$1.811
$1.811
$15.48
2026-09-18
gpt-5-6-terra
OpenAI
$1.811
$1.811
$15.48
2026-09-12
gpt-5-6-sol
OpenAI
$3.147
$3.147
$26.90
2026-09-12
claude-opus-4-8
Anthropic
$3.934
$3.934
$33.63
2026-09-12
claude-opus-5
Anthropic
$3.934
$3.934
$33.63
2026-09-12
claude-fable-5
Anthropic
$7.868
$7.868
$67.25
2026-09-18
claude-fable-5-1
Anthropic
$7.868
$7.868
$67.25
2026-09-18
gpt-6-astra
OpenAI
$7.868
$7.868
$67.25
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, as of each row's own rates_as_of date, with nothing checking it against the vendor since.
⚑ 90.6 PER CENT OF THIS RUN'S OUTPUT TOKENS WERE PROVIDER-SIDE REASONING (107,628 of 118,834), left at the provider's default, so every row below prices a reasoning-ON workload. Output is 78.7 per cent of the projected bill on the shared card. A vendor whose default differs, or a caller who turns it off, would see a materially different figure -- and this kit has not measured what turning it off does to the answers.
⚠︎ EVERY ROW PRICES ONE MONTHLY RUN OF 117 READINGS, WHICH IS THE WRONG UNIT FOR A BUYER AND THE RIGHT ONE FOR A COMPARISON. A monitor's bill is (open stores x runs per year), so multiply the per-query column by your own two numbers rather than reading the per-run total. On the shared card, a 2,000-store book read monthly is about $93 a year and read weekly is about $403.
THE FREE FLOOR IS NOT ON THIS TABLE AND IT SHOULD BE READ FIRST. b002 costs $0.00 and scored 95.73 per cent on movement against the model's 100.0. Every row here is the price of a 4.27-point margin on this corpus.
LATENCY IS NOT PROJECTED. p50 7201 ms and p95 17504 ms are properties of the provider that actually ran, at 12 concurrent workers, and say nothing about any row here.
THE PROVIDER THAT ACTUALLY RAN IS NOT ON THIS TABLE, per this estate's naming rule. Nothing here is what was paid.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
17 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Writes 117 period files across 40 stores from a fixed seed, solving each store's printed figures backwards from a scenario table so that exactly the intended indicators breach. It also plants the six operating-note classes, which are the experiment. Free and offline.
You change it to: Point it at your own extract instead of generating one. The only contract downstream is the seven section headings and the printed figure labels src/figures.py reads.
tools/build_corpus.py
# Generate the store population and its answer key. Deterministic, offline, free.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
SEED = 20260823
PERIODS = ("P1", "P2", "P3")
PERIOD_LABEL = {"P1": "2026-01", "P2": "2026-02", "P3": "2026-03"}
PERIOD_RANGE = {"P1": ("2026-01-01", "2026-01-31"),
PRIOR_RANGE = {"P1": ("2025-12-02", "2025-12-09"),
SEGMENTS = ("FREESTANDING", "INLINE", "MALL_FOOD_COURT", "TRAVEL_PLAZA")
src/figures.pythe printed figures, parsed
Reads the eleven printed numbers off a period file and derives the four indicator distances from the ROUNDED values the page shows. The model is never asked to divide one dollar figure by another.
src/figures.py
# Read the printed figures off one period document. Pure code, no model.
def _num(text, pat):
def read(text):
def store_and_period(text):
src/exemptions.pythe approved cause list
Five causes, their plain-English wording and the indicators each can excuse, as data. Printed in full inside every document, and check_labels reads it back off the page to prove the two copies have not drifted.
src/exemptions.py
# The approved cause list, in one place.
APPROVED = {
COVERS = {k: set(v[1]) for k, v in APPROVED.items()}
CAUSE = {k: v[0] for k, v in APPROVED.items()}
src/notes.pythe note reader, twice
claims() finds cause CODE tokens and the dates beside them -- what a free regex can see. claims_full() also matches an approved cause written in plain words. The gap between the two functions IS the measurement, which is why they are deliberately not one function.
src/notes.py
# Read an operating note as an exemption CLAIM. Pure code -- and deliberately only as good as a
DATE = r"(\d{4}-\d{2}-\d{2})"
def period_range(text):
def note_text(text):
def _claim(code, body, at, lo, hi, how):
def claims(text):
def claims_full(text):
src/rules.pythe reporting rule — a swap seam
The four bands, the three-part exemption test, the five-step movement precedence, and FORBIDDEN -- the closed list of language that would turn a description into an allegation. The answer key and the first file a forker replaces.
You change it to: Your bands, your approved causes, your coverage map and your own forbidden-language list. Replace this file FIRST -- it is a policy somebody wrote down, not a utility, and every figure on these pages is a property of the one shipped here.
src/rules.py
# The written reporting rule, as code: bands, exemptions, and what counts as a CHANGE.
INDICATORS = ("CASH_MIX", "DEPOSIT_GAP", "VOID_RATE", "TICKET_DRIFT")
BANDS = {
MOVEMENTS = ("FIRST_READING", "NEW_CROSSING", "STILL_STANDING", "RETURNED_TO_BAND", "NO_CHANGE")
FORBIDDEN = (
def outside(figures):
def margins(figures):
def excused(raw_outside, text):
def classify(figures, text, carried):
def raises(answer):
src/state.pythe carried state — a swap seam
Four fields written from the ARITHMETIC and never from a model reply, rendered as one English sentence. Constant size whatever the history length, so the twelfth run of a store costs what the second did.
You change it to: What persists between runs. It is four fields today; adding a fifth changes the prompt by one clause and the cost by nothing measurable, because the state is constant-size by construction.
src/state.py
# The carried state -- what makes a reading a CHANGE rather than a classification.
def initial():
def advance(carried, figures, text, run_label, movement=None):
def describe(carried):
src/cadence.pythe clock — a swap seam
The schedule as code: monthly, five days after period close; one run owns one reading of every open store for exactly one period; gap() and missed() say how many scheduled runs this one is standing in for.
You change it to: FREQUENCY and RUNS_AFTER_PERIOD_CLOSE_DAYS. Five days is this corpus's settlement-lag assumption; a franchisor whose depository posts same-day would set it to 1, and a weekly cycle changes what one run owns.
src/cadence.py
# The clock. When this monitor wakes, what one run owns that the last did not, and what a missed
FREQUENCY = "monthly"
RUNS_AFTER_PERIOD_CLOSE_DAYS = 5
UNIT = "(store, period)"
PERIOD_ORDER = ("P1", "P2", "P3")
def owned(population, period):
def gap(prev_period, this_period):
def missed(prev_period, this_period):
def describe():
src/segment.pythe section split
Splits a period file into its seven named sections. The shape is a property of the corpus builder and evals/check_labels.py asserts it on every document, so a parser drift shows up as a refusal to start rather than as a quietly truncated prompt.
src/segment.py
# Split one period document into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Store", "Reporting Rule", "Period Figures",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pywhat is withheld — a swap seam
Sends five of the seven sections. Franchisee Contact (a named individual, address, telephone, email) and Audit File Access are mapped by no field, and the subtraction is UNCONDITIONAL so the fallback path -- the one that fails quietly -- cannot reintroduce them.
You change it to: NEVER_SENT and SECTION_HINTS. Adding a section to a document does not add it to the prompt -- a field has to ask for it.
src/select.py
# Pick which sections of a period document are sent. Pure code -- the last deterministic step
BANNER = "Synthetic Record"
STORE = "Store"
RULEBLOCK = "Reporting Rule"
FIGURES = "Period Figures"
NOTES = "Operating Notes"
CONTACT = "Franchisee Contact"
ACCESS = "Audit File Access"
NEVER_SENT = (CONTACT, ACCESS)
SECTION_HINTS = {
src/prompt.pythe prompt
Four parts in assembly order: the question, the schedule, the carried state, the reading. The stateless arm REPLACES the third block rather than deleting it, so the control does not also measure a missing heading.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, run_label, stateless=False):
def parts_of(text, carried, run_label, stateless=False):
src/monitor.pyone reading, one call
Assembles, calls, parses. Holds MAX_TOKENS = 20000, the published ceiling a scored run must use.
src/monitor.py
# One store, one period, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You assemble under-reporting indicators for one franchised store for one reporting "
MAX_TOKENS = 20000
FIELDS = ("movement", "indicators_outside", "indicators_new", "exemptions_applied")
def documents():
def population():
def load_doc(doc_id):
def _parse(txt):
evals/answers.pythe answer key
Generates the key by WALKING A SCHEDULE. The correct answer to a March reading depends on which runs happened, so a static gold file would be silently wrong for the one experiment this kit exists to run.
evals/answers.py
# Build the answer key by WALKING THE SCHEDULE, not by reading a file.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
def plan_rows():
def walk(periods, population=None, texts=None):
def full():
def missed(skip):
evals/baseline.pythe three free floors
Threshold only; threshold plus memory; threshold plus memory plus a regex note reader. All three scored by the same grader on the same readings.
evals/baseline.py
# THE FREE FLOORS. Three of them, all pure Python, all $0.00, all scored by the same grader the
def _describe(raw, new, exc, movement):
def _movement(raw, reportable, carried):
def review(text, carried, kind="b002"):
evals/scoring.pythe grader
Six figures, never averaged, plus conclusory() -- the negative assertion, checked on every reply against src/rules.FORBIDDEN. No model, no key, $0.00.
evals/scoring.py
# Grade a run. Pure code, no key, no model, no cost -- a measured $0.00, not an unpriced one.
MOVEMENTS = R.MOVEMENTS
def conclusory(description):
def _set(x):
def score(records, golds):
def cadence_cost(missed_records, full_golds, missed_golds):
evals/check_labels.pythe pre-flight
Twelve checks that must pass before a run may spend, including a red proof of the language checker in both directions and a proof that the missed-run key actually differs from the full one.
evals/check_labels.py
# THE PRE-FLIGHT. Free, offline, and it must pass before any run is allowed to spend.
MIN_PER_MOVEMENT = {"FIRST_READING": 20, "NEW_CROSSING": 15, "STILL_STANDING": 8,
MIN_PER_NOTE_CLASS = {"none": 40, "coded_valid": 4, "prose_valid": 8, "coded_stale": 4,
KNIFE_EDGE = 0.15
SEEDED_BAD = ("The deposit gap and the cash-mix drop together indicate under-reporting by the "
SEEDED_OK = ("Cash share is 8.1 points below the store's 12-period baseline and the deposit gap "
def main():
evals/run.pythe harness
Fires a schedule. Because the state is arithmetic, every reading is independent and they run concurrently -- the guardrail visible in the shape of the loop.
evals/run.py
# Fire the monitor over the store population and write one result file. THIS SPENDS MONEY.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
def stub_complete(cfg, system, user, max_tokens=1024):
def main():
src/app.pythe local UI
http.server on port 9003, serving ui/index.html, ui/app.js and ui/app.css -- hand-written, no build step. Shows the carried state verbatim, the free floor beside the answer, the access log, and which sections were withheld. Renders with no key: only /api/assemble calls a provider, and without one it returns 200 with a sentence rather than an error.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "9003"))
KEY = A.full()
def _access_log(text):
class H(BaseHTTPRequestHandler):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyWrites 117 period files across 40 stores from a fixed seed, solving each store's printed figures backwards from a scenario table so that exactly the intended indicators breach. It also plants the six operating-note classes, which are the experiment. Free and offline. A swap seam.
src/figures.pyReads the eleven printed numbers off a period file and derives the four indicator distances from the ROUNDED values the page shows. The model is never asked to divide one dollar figure by another.
src/exemptions.pyFive causes, their plain-English wording and the indicators each can excuse, as data. Printed in full inside every document, and check_labels reads it back off the page to prove the two copies have not drifted.
src/notes.pyclaims() finds cause CODE tokens and the dates beside them -- what a free regex can see. claims_full() also matches an approved cause written in plain words. The gap between the two functions IS the measurement, which is why they are deliberately not one function.
src/rules.pyThe four bands, the three-part exemption test, the five-step movement precedence, and FORBIDDEN -- the closed list of language that would turn a description into an allegation. The answer key and the first file a forker replaces. A swap seam.
src/state.pyFour fields written from the ARITHMETIC and never from a model reply, rendered as one English sentence. Constant size whatever the history length, so the twelfth run of a store costs what the second did. A swap seam.
src/cadence.pyThe schedule as code: monthly, five days after period close; one run owns one reading of every open store for exactly one period; gap() and missed() say how many scheduled runs this one is standing in for. A swap seam.
src/segment.pySplits a period file into its seven named sections. The shape is a property of the corpus builder and evals/check_labels.py asserts it on every document, so a parser drift shows up as a refusal to start rather than as a quietly truncated prompt.
src/select.pySends five of the seven sections. Franchisee Contact (a named individual, address, telephone, email) and Audit File Access are mapped by no field, and the subtraction is UNCONDITIONAL so the fallback path -- the one that fails quietly -- cannot reintroduce them. A swap seam.
src/prompt.pyFour parts in assembly order: the question, the schedule, the carried state, the reading. The stateless arm REPLACES the third block rather than deleting it, so the control does not also measure a missing heading.
src/monitor.pyAssembles, calls, parses. Holds MAX_TOKENS = 20000, the published ceiling a scored run must use.
evals/answers.pyGenerates the key by WALKING A SCHEDULE. The correct answer to a March reading depends on which runs happened, so a static gold file would be silently wrong for the one experiment this kit exists to run.
evals/baseline.pyThreshold only; threshold plus memory; threshold plus memory plus a regex note reader. All three scored by the same grader on the same readings.
evals/scoring.pySix figures, never averaged, plus conclusory() -- the negative assertion, checked on every reply against src/rules.FORBIDDEN. No model, no key, $0.00.
evals/check_labels.pyTwelve checks that must pass before a run may spend, including a red proof of the language checker in both directions and a proof that the missed-run key actually differs from the full one.
evals/run.pyFires a schedule. Because the state is arithmetic, every reading is independent and they run concurrently -- the guardrail visible in the shape of the loop.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1646 input and 1015 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version, and one field that WOULD be the surface in a real deployment. Every byte that reaches the prompt on this corpus is generated by tools/build_corpus.py from a fixed seed. In production the Operating Notes section is free prose an area coach types, and it is deliberately SENT -- hiding this kit's own injection surface from the run that is supposed to measure it would be worse than having one. Two sections are never sent at all: Franchisee Contact (a named individual, address, telephone, email) and Audit File Access.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and merged key by key so a kit-local file holding only MODEL keeps the shared key. It is never written to a result file, never printed, and src/app.py substitutes [API_KEY] and [BASE_URL] into any provider error before it reaches the browser. No reader of this kit is ever asked for a key: the corpus, the answer key, all three free floors and every committed result ship in the repository.
The experimentWe did not attack it -- and the two boundaries that matter most were proven by breaking them on purpose
An indirect prompt injection needs a field an outside party controls that reaches the prompt. This kit has exactly one and names it: the Operating Notes body. On this corpus it is generated from a fixed seed, so the trial would be self-attacking. What WAS done instead is the thing an attack run cannot do -- the two guarantees were removed and the checks were confirmed to fail. Hide a mapped section and the fallback test convicts; seed a conclusory sentence and the language checker convicts; seed a descriptive one and it acquits. Both are in evals/check_labels.py and both run before a scored run may spend. Confirmed by assertion and by reading the recorded runs, not by an attack trial, on 2026-08-23 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Can a wrong answer poison the next run?
A monitor that carried its own verdict forward would compound one wrong reading into every reading after it, and the compounding would be invisible: each later answer would be locally consistent with the state it was given.
It cannot. src/state.advance() takes (carried, figures, text, run_label) and never the reply -- the answer does not appear in the signature. evals/answers.py computes the whole schedule's state before a single call is made, which is why evals/run.py can fire every reading concurrently. If the state depended on the model, that loop could not be written that way.
Can a franchisee's name or address reach the provider?
Sending the whole document, or falling back to it when a mapped section is missing -- the or list(secs) fallback every sibling kit once shipped.
Franchisee Contact and Audit File Access are mapped by no field, and src/select._fallback subtracts NEVER_SENT UNCONDITIONALLY, so the fallback path cannot reintroduce them. evals/check_labels.py asserts it on every one of the 117 documents AND red-proves the fallback specifically, by stripping every mapped section from a document and confirming the withheld ones still do not appear.
Can the pack state a conclusion?
Trusting the instruction. A prompt that says 'do not conclude' is a request, and a request that nothing checks is a claim about a run nobody looked at.
evals/scoring.conclusory() checks every reply against src/rules.FORBIDDEN, the local UI prints the result beside each answer, and evals/check_labels.py RED-PROVES the checker in both directions before any run may spend -- it must convict a seeded conclusory sentence and acquit a seeded descriptive one. Measured 0 of 257 replies across four model runs.
Can the kit do anything to a store or a franchisee?
A monitor that grows a notify step, a case-opening step or a 'send to legal' flag, one feature at a time, each one reasonable on its own.
There is no writer anywhere outside results/*.json and docs/shots/. No endpoint in src/app.py opens a case; the only POST is /api/assemble and it returns JSON. This is a property of what is ABSENT, which is why it is stated where the calls are made (src/monitor.py's docstring) rather than only on a page.
Can a run silently read a period twice?
A scheduler that fires twice would re-read readings already carried into the state, and every standing indicator would be reported as a fresh crossing -- a monitor generating its own false raises.
PARTLY, AND THE HONEST ANSWER IS THAT THE GUARANTEE IS ON THE SCHEDULE AND NOT ON A SCHEDULER. cadence.owned() enumerates one reading per open store per period and evals/check_labels.py asserts every (store, period) is owned exactly once across the whole schedule. Nothing here tests a real cron that double-fires, because this kit ships no cron.
Each boundary above was checked by running an assertion or by reading the recorded run, never by reasoning about the code. Two of the five are red-proven -- the withheld sections and the language checker -- meaning the check was shown to FAIL when the guarantee is removed, which is the only way to tell a working check from one that has never fired.
The result0 attack trials, five boundaries checked -- two of them red-proven by removing the guarantee and watching the check fail -- and one injection surface named rather than hidden.
0attack trials fired
1prompt field an outside party would control in a real deployment
2sections withheld from the model on every reading
The Operating Notes body IS the field an outside party would control in a real deployment -- prose typed by whoever files the note -- and it is sent. On THIS corpus every byte of it is generated from a fixed seed, so there is nothing adversarial to find and firing an attack run would measure the attacker's imagination rather than the system. The honest state is: the surface is identified, it is not hidden, and it has not been attacked.
Read this twice
The carried state is written from the arithmetic, never from the model's reply. src/state.advance() does not take the answer -- it is not in the signature. That is what stops one wrong reading compounding through every later run of the same store, and it is also why the whole schedule's state is computable before any call is made, which is why every reading in a run is independent. A monitor that fed its own verdicts forward would be cheaper to write and impossible to audit.
HonestyWhat this does not prove
Whether a real operating note -- prose an area coach writes freely -- would carry an instruction the model follows. That is the live question for this kit and it has not been tested, because a synthetic corpus cannot answer it honestly.
Whether the withheld sections stay withheld on a FORKER's documents. The guarantee is that no field maps to them and the fallback subtracts them; it is not that a section named something else in your corpus will be recognised as sensitive. Adding a section adds it to SECTIONS, and a section nobody adds to NEVER_SENT is sent.
Whether the forbidden-language list is the right list. It has 22 entries chosen here, and it cannot catch an implication that uses none of them -- 'the deposit pattern speaks for itself' passes. A phrase list is a floor on this constraint, not a ceiling.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
The pack DESCRIBES indicators and states no conclusion, and the carried state is written by code from the arithmetic rather than from the model's reply. Two rules, both enforced in code, neither of them configurable.
src/rules.FORBIDDEN checked by evals/scoring.conclusory() on every reply and rendered beside every answer in src/app.py; src/state.advance() called by evals/answers.walk() with the figures and the document, never the answer.
EvidenceDoes it hold?
What
Measured
No reply states a conclusion
0 of 257 replies across four model runs (r001 117, s001 77, a000 39, p002 24) contained any of the 22 forbidden phrases.
The checker that measures it is not asleep
Red-proven in both directions on every pre-flight: it convicts a seeded conclusory sentence and acquits a seeded descriptive one. Twelve checks pass; a scored run cannot start until they do.
A wrong answer cannot propagate to the next run
Structural, not statistical: the model's reply is not an argument to src/state.advance(). There were no wrong answers on r001 to propagate, so this is a property of the call signature rather than a rate.
A named individual's contact details never leave the machine
117 of 117 documents send exactly 5 of 7 sections. Red-proven through the fallback path, which is the one that fails quietly.
One scheduled run owns one reading of each open store, exactly once
117 readings owned across three periods, 0 duplicates, 0 gaps within a store's own membership -- asserted by evals/check_labels.py over the schedule.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The state is correct whatever the model says, and the language check passes on a confidently wrong reading as long as it is phrased descriptively. Correctness is the Eval lens's job and the two must not be read as one another.
IT IS NOT A GUARANTEE THE PACK IS SAFE TO ACT ON. The forbidden-language list stops the OUTPUT stating a conclusion; it does nothing about a reader who draws one. The real control on that is the cap -- no output of this kit may be used toward a formal demand -- and a cap is a policy, enforced by people.
IT IS NOT A COMPLETE PHRASE LIST. 22 entries chosen here, matched on word boundaries. An implication that uses none of them passes, and 'the deposit pattern speaks for itself' is the example that does. A phrase list is a floor on this constraint, never a ceiling.
IT IS NOT PROTECTION AGAINST A MISSED RUN. Every guarantee here is about what one run does. Nothing in the code notices that a run did not happen, and a000-royalty-watch-missedrun measured what that costs: 4 of 15 crossings lost permanently. The guard against it is a scheduler this kit does not ship.
WatchedWhat is watched, and why that one
9runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 32 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
13 measured by the latest run19 need the model half
Metric
Owner
Role
Why this one
movement-exact-match
The movement, the outside-band set, the new set and the exemption set, per reading, exact match against the generated answer key
alarm
movement_accuracy_pct; change_movement_accuracy_pct; change_detection_pct; false_raise_rate_pct; outside_set_exact_pct; exemption_set_exact_pct; conclusory_language_rate_pct; unparsed_replies — alarm on conclusory_language_rate_pct FIRST, on any value above 0.0 -- it is the one figure with no acceptable trade-off, because a pack that states a conclusion has done the thing this kit's cap forbids whatever else it got right. Then change_detection_pct, because a missed crossing can become permanently unrecoverable at the next run and a false raise cannot.
two-directions-apart
Missed crossings and false raises, counted separately against the key
alarm
change_detection_pct; false_raise_rate_pct; crossings_missed; crossings_lost — on the missed-run arm only — alarm on crossings_missed, on any increase at all. It is a COUNT and not a rate because the denominator is 23 and a rate over 23 rows hides the first one.
conclusory-language
The negative assertion: does the description state a conclusion rather than describe a reading
alarm
conclusory_language_rate_pct; conclusory_replies; the phrases each hit matched — results[].language_hits[].phrases — alarm on any hit at all. One reply out of 117 is 100 per cent of this kit's acceptable count.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
117
different corpus — nothing is comparable
corpus.bytes
446,941
store-period readings edited — the count held, the bytes did not
split.count
117
the store-period readings count moved — a different set was scored
split.size_p50
1
the median size of one store-period reading moved
split.size_p95
1
the 95th-percentile size of one store-period reading moved
dataset.rows
117
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
movement accuracy
not yet known
117 readings (77 of them memory-dependent)
A band is the spread between runs of the SAME model at the same settings. There have been two -- r001 and the p002 repeat probe over a 24-reading subset -- and they agreed on every field of every shared reading, so the measured spread is 0 over 24 readings. One agreement is not a band.
the two directions
not yet known
23 crossings, 23 raises
Same reason. And the denominators are small enough that a band would be wide even with more runs: one row is 4.35 points of detection.
the arithmetic
zero tolerance
117 readings
The outside-band set is arithmetic on printed numbers and all three free floors score 100.0 per cent on it. Any model value below that is the model failing at something a regex does for nothing, which is a defect and not a trade-off.
the negative assertion
one reply
every reply, every run
A COUNT, not a rate. 257 replies across four model runs have produced none, and one would be 100 per cent of this kit's acceptable count -- the cap has no expansion path, so there is no value of this figure that is a trade-off.
the cadence
one missed run
15 crossings in the skipped period
Measured on one ablation, not a spread: skipping one of three monthly runs lost 4 of the 15 crossings in the skipped period permanently and misdated 11. It is a property of the population -- how fast indicators return to band -- and would differ on another book.
reply integrity
zero tolerance
117 calls
0 unparsed replies over 117 calls on the scored run, with the largest reply at 3469 output tokens under a 20000 cap. An unparsed reply counts as a miss in every field, so this figure has to be read before any accuracy above it.
cost and latency
not yet known
117 readings
One scored run. p50 7201 ms and p95 17504 ms over 117 readings at 12 workers; the p95 is 2.4x the p50, which is a concurrency artefact as much as a model one and has not been separated.
HistoryRun history
9 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 2 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · with the model in the path — 8 runs. Columns here are only ever compared with each other.
Metric
a000-royalty-watch-missedrun 2026-08-23
b000-royalty-watch-noteblind 2026-08-23
b001-royalty-watch-flagonly 2026-08-23
b002-royalty-watch-regexnotes 2026-08-23
c000-royalty-watch-calibration 2026-08-23
p002-royalty-watch-repeat 2026-08-23
r001-royalty-watch 2026-08-23
s001-royalty-watch-stateless 2026-08-23
answered, %
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
change detection, %
100.0
100.0
100.0
100.0
100.0
100.0
100.0
0.0
change movement accuracy, %
100.00
88.31
49.35
93.51
100.00
100.00
100.00
0.00
conclusory language rate, %
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
crossings lost
4
—
—
—
—
—
—
—
crossings misdated
11
—
—
—
—
—
—
—
exemption set exact, %
100.00
84.62
84.62
89.74
100.00
100.00
100.00
100.00
false raise rate, %
0.00
28.12
63.49
17.86
0.00
0.00
0.00
—
input tokens, whole run
64477
0
0
0
6602
39456
192660
125292
model latency p50 ms
7215.00
0.00
0.00
0.00
8804.00
7259.00
7201.00
7429.00
model latency p95 ms
25071.00
0.00
0.00
0.00
13700.00
13098.00
17504.00
15166.00
movement accuracy, %
100.00
92.31
32.48
95.73
100.00
100.00
100.00
0.00
output tokens, whole run
45235
0
0
0
4312
22510
118834
76995
outside set exact, %
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
unparsed replies
0
0
0
0
0
0
0
0
not a time series No two of these 8 runs measured the same system — they differ on change_cells, crossings_in_set, floor, max_tokens, periods_run, provider, raised, readings, skipped_period, stateless, stores, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-royalty-watch-stub 2026-08-23
answered, %
100.0
change detection, %
100.0
change movement accuracy, %
49.35
conclusory language rate, %
0.0
exemption set exact, %
84.62
false raise rate, %
63.49
input tokens, whole run
204561
model latency p50 ms
0.00
model latency p95 ms
0.00
movement accuracy, %
32.48
output tokens, whole run
5609
outside set exact, %
100.0
unparsed replies
0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 13 chips that all say so.
DeviationsWhat deviated
0 breaches across 9 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the floor can read the operating notes at all
b000-royalty-watch-noteblind against b002-royalty-watch-regexnotes. Same corpus, same schedule, same grader, same 117 readings; the only difference is that b002 runs src/notes.claims() and b000 does not.
b001-royalty-watch-flagonly against b000-royalty-watch-noteblind. Both are blind to the notes, so this pair isolates MEMORY and nothing else -- which is why the memory figure on these pages comes from here rather than from the stateless model arm.
whether the prompt carries the carried state
movement accuracy 100.0 pct -> 0.0 pct, and every reading answered FIRST_READING
measured
r001-royalty-watch against s001-royalty-watch-stateless, 77 shared readings. ⚠︎ CONFOUNDED: FIRST_READING is the CORRECT answer under the instruction's own precedence when no earlier run exists, so this edge measures that the state block is the only route between runs, not what memory is worth.
whether a scheduled run actually happens
4 of 15 crossings in the skipped period lost permanently, 11 more reported in the wrong period, and the surviving readings cost 11.3 pct more each
measured
a000-royalty-watch-missedrun -- period P3 replayed with the state period P1 left behind, 39 readings, against the full schedule's answer key. Model accuracy was 100 pct on both, which is the finding: this lever does not go through the model at all.
the output ceiling
largest reply 1752 tokens against a 2,000 cap on the calibration; 3469 against the published 20000
measured
c000-royalty-watch-calibration, four readings at max_tokens=2000. No reply was truncated, but the largest used 88 pct of that cap on four calls, which is why the published ceiling is not 2,000.
asking the same readings again
no structured field moved on any of 24 shared readings; 0 of 24 descriptions were byte-identical
measured
p002-royalty-watch-repeat against r001-royalty-watch on the same 24 readings, same settings, same day.
disabling provider-side reasoning
up to 90.6 pct of output tokens, and therefore most of the bill -- and an unknown amount of accuracy
reasoning
Not run. src/adapters/__init__.py can send the disabled-thinking field and this kit's harness never has; every published run left it at the provider's default and the result files record thinking: null. The token share is measured; the effect on answers is not.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
movement accuracy
nothing yet
the two directions
nothing yet
the arithmetic
nothing automatically -- the kit ships no alerting. It is the first figure to read when a run looks wrong, because it separates a reading failure from a judgement failure.
the negative assertion
nothing automatically. The local UI prints the matched phrases beside the answer, which is the only place this fires today.
the cadence
nothing yet, and this is the band most worth wiring: a run that did not START is invisible to everything in this repository.
reply integrity
nothing automatically. c000-royalty-watch-calibration is the evidence for where the ceiling sits.
cost and latency
nothing yet
NextThe three you would add first
A human step in front of anything that leaves the buildingThe kit produces indicator readings for an auditor. The row's own cap is absolute -- no output toward a franchise default notice, with no expansion path -- and the anchor research that would say what is lawful has not been done.
A scheduler with a MISSED-RUN ALARM, not just a scheduleThis is the cheapest large win available and this kit does not ship it. A missed monthly run cost 4 of 15 crossings permanently in the measured ablation, against a model that was 100 per cent accurate on the degraded schedule. Alarm on a run that did not start, not only on one that failed.
Access control on the output, not only on the inputAudit File Access is in the corpus because who has read an indicator file is part of the record. The kit renders it and enforces nothing -- a real deployment's restricted-by-default tier is a guardrail, not a data label, and it belongs outside this repository.
A second reader for any store raised more than N runs runningNothing here has been measured beyond three runs per store, and the most likely real failure is alert fatigue on a store that stays outside band for months.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
⚑ THIS KIT HAS TWO CADENCES AND THEY ARE DIFFERENT THINGS. The MONITOR runs monthly, five days after each period closes (src/cadence.py) -- that is the product, and what a missed one costs is measured in a000-royalty-watch-missedrun.
The EVAL cadence is this: re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/rules.py, src/notes.py, src/exemptions.py, src/state.py or src/cadence.py, and re-run the three free floors with it (also free) -- they take seconds and they are the comparison every model figure is read against. Re-run the SCORED arm only when the prompt, the model or the settings change, because that one spends. Nothing here re-runs automatically.
What this cannot tell you
One run per arm on the model side. Whether the 100 per cent scores are stable across repeats is measured only on a 24-reading subset (p002), which found zero movement -- strong evidence of stability on those readings and no evidence at all about the other 93.
Whether the guardrail holds on prose this repository did not write. Every reply checked was produced from a synthetic corpus, and the forbidden-language list has never been tested against a model asked to write about a real store.
Whether the cadence cost figures generalise. 4 of 15 crossings lost is one number from one skipped period on one population, and it depends entirely on how fast indicators return to band -- a property of the book, not of the monitor.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, argparse, urllib, http.server and concurrent.futures, every one of them standard library. requirements.txt names nothing, and it says so in a sentence rather than being empty.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores -- LangGraph checkpointers, LangChain memory, a vector store used as a scratchpad
those carry a TRANSCRIPT and grow with it; this carries four scalars written by arithmetic and is constant-size. Same word, opposite cost curve, and the difference is measurable: the stateful and stateless arms here differ by 19 input tokens a reading.
the schedule
src/cadence.py
Airflow, Dagster, Prefect, Temporal, or a cron line -- anything that owns 'this ran, that did not'
this file is the CONTRACT, not the runner: what a run owns, how many runs it is standing in for, and which periods nobody read. A scheduler tells you a job failed; cadence.missed() tells you which readings were never taken, which is the thing that costs money. Wire the two together and alarm on a run that did not START.
the model
src/adapters/__init__.py
any vendor SDK, LiteLLM, or an OpenAI-compatible gateway
one function per provider returning one normalised shape including token counts and finish_reason. Lens 05 publishes those and lens 07 prices them, so a provider that returns no usage cannot be published from this kit.
the rulebook
src/rules.py, src/exemptions.py
a rules engine -- Drools, json-logic, a decision table in a spreadsheet, or your existing GL/POS analytics rules
the four bands and the coverage map genuinely are a decision table and a rules engine would hold them well. What a rules engine cannot do is the one thing the model is here for -- decide whether a sentence of prose names an approved cause. b002 is the rules-engine arm, measured at 95.73 per cent.
the grader
evals/scoring.py
an eval framework -- Ragas, DeepEval, promptfoo, Braintrust
every field here is a word or a code from a closed list, so grading is dictionary comparison and a framework would be a dependency to compare two strings. Where a framework earns its place is trend storage across runs, which this kit does not do at all.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the place it looks most like one is a straight line. A monitor reads a store, carries four fields to the next run, and reads it again -- which invites a state machine, a checkpointer and a workflow engine. All three would be modelling something that is already a for-loop over a schedule, because the state transition is arithmetic and does not depend on the model's reply. That is why evals/answers.py can compute the WHOLE schedule's state before any call is made, and why every reading in a run is independent. If the state depended on the reply, this would be a graph and the loop could not be written this way.
The other sideWhat a framework costs you
NO SCHEDULER, AND THIS ONE IS A REAL GAP RATHER THAN A STYLE CHOICE. src/cadence.py is the contract and there is no runner behind it: nothing in this repository notices that a run did not happen. The ablation measured what that costs -- 4 of 15 crossings lost permanently -- so the missing piece has a price on it.
NO STATE STORE. The carried state is scoped to the run, deliberately: an eval that kept state in a file could not be re-run and could not be compared against its own control, because the second run would start from the first one's memory. A deployment needs a real store, and choosing one is out of scope here.
NO TREND STORAGE. Each result file is a snapshot; nothing in the kit compares this month's run with last month's beyond the four carried fields. The site's own run register does that for the kit's METRICS, which is a different question from a franchisor asking how store STR-0006 has trended.
NO RETRY OR DEAD-LETTER POLICY ABOVE THE CALL. The adapter retries a transient provider failure four times with backoff; a reading that still fails is recorded as a failure and the run continues. What a deployment does with three failed readings out of two thousand -- re-run them, hold the whole run, or accept the gap -- is a policy this kit does not have, and a gap in a monitor's population is not the same as a gap in a batch job's output.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative. The one exception is the rules-engine row: b002 IS that arm, implemented and scored at 95.73 per cent, which is why it is the only row here with a number in it.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-royalty-watch on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
7,201 ms
not yet known
nothing yet
Model, p95
17,504 ms
not yet known
nothing yet
Input tokens
192,660
not yet known
nothing yet
Output tokens
118,834
not yet known
nothing yet
No movement column. Not one of the 7 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
a000-royalty-watch-missedrun7,215 ms
c000-royalty-watch-calibration8,804 ms
p002-royalty-watch-repeat7,259 ms
r001-royalty-watch7,201 ms
s001-royalty-watch-stateless7,429 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-royalty-watch-noteblind, b001-royalty-watch-flagonly, b002-royalty-watch-regexnotes recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 9 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
period files
data/corpus/<STORE>-<PERIOD>.txt -- 117 files, 446,941 bytes, regenerated from a fixed seed by tools/build_corpus.py
five of the seven sections go to the provider in the prompt. Franchisee Contact never does -- a named individual, address, telephone and email -- and neither does Audit File Access. In your deployment this artifact is the royalty reporting extract you already produce.
the carried state
in memory for the length of a run; a deployment persists it. Four fields per store, written by src/state.advance() from the arithmetic
yes, as one English sentence of about 106 characters inside the prompt. It is the smallest thing in the prompt and the reason the kit exists.
the schedule
src/cadence.py -- frequency, settlement lag, and what one run owns
one line of it does: the prompt tells the model which period this run owns. Nothing else about the schedule reaches the provider.
the reporting rule
src/rules.py and src/exemptions.py, and printed in full inside every period file
yes -- the Reporting Rule section is in every prompt, roughly 1780 characters of it. That is deliberate: a document that carries its own rulebook can be read on its own, and a forker diffing two runs can see which rule each was judged against.
run results and the answer key
results/eval-*.json, committed; the key is GENERATED at run time by evals/answers.py and is never a file
nothing. Grading calls nothing -- a measured $0.00, not an unpriced one.
the provider key
.env (gitignored) or the shared repo-root .env, read by src/config.py and merged key by key
only as the Authorization header on the completion call. It is never written to a result file, never printed, and substituted out of any provider error before it reaches the browser.
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and merged key by key so a kit-local file holding only MODEL keeps the shared key. It is never written to a result file, never printed, and src/app.py substitutes [API_KEY] and [BASE_URL] into any provider error before it reaches the browser. No reader of this kit is ever asked for a key: the corpus, the answer key, all three free floors and every committed result ship in the repository.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
monthly, five days after each period closes. ONE RUN OWNS ONE READING OF EVERY OPEN STORE FOR EXACTLY ONE PERIOD -- cadence.owned() enumerates it and the pre-flight asserts every (store, period) is owned exactly once. What one run owns that the last did not is one period per open store, plus any store that opened since.
117 readings owned across 3 periods, 0 duplicates; a whole run took 98.9 seconds of wall clock at 12 workers. A MISSED RUN COSTS 4 of the 15 crossings in the skipped period PERMANENTLY -- no later run recovers them -- and reports 11 more in the wrong period, while the surviving readings each cost 11.3 pct more. (r001-royalty-watch, a000-royalty-watch-missedrun, evals/check_labels.py)
the run has to finish before the next one is due. At this width, 117 readings took 98.9 seconds; 24,000 readings a year is about 6 hours of wall clock per annual cycle, which fits a monthly window and does not fit a nightly one at this concurrency. Past that ceiling you widen the pool or you shorten the population, and a run that overruns IS a missed run.
changing the frequency invalidates the headline rather than scaling it. Every figure here is measured on ONE reading per store per period; on a weekly cadence the same indicators are read four times as often, most readings become STILL_STANDING, and both the detection denominator and the false-raise rate are over a different population of events.
state
four fields per store, carried between runs: whether the store has been seen, the indicators outside band on the ARITHMETIC at the last run, how many runs it has been read in, and whether the last run raised. Written by src/state.advance() from the figures and the document -- the model's reply is not an argument to it.
106 characters in the prompt on the published reading; the stateful and stateless arms differ by 19 input tokens a reading, so the state is 1.2 pct of input. Constant in history length by construction: a store's twelfth run carries the same four fields as its second. (r001-royalty-watch against s001-royalty-watch-stateless)
the run's own memory. State is scoped to a run here, deliberately, so the eval can be re-run and compared against its own control. A deployment needs a real store keyed by (store, period), and this kit does not ship one -- seam: null, that is the point the kit is outgrown.
carrying anything the MODEL produced would invalidate every figure on these pages, because a wrong reading would then propagate into later prompts and the run would be measuring compounding rather than judgement. The signature of advance() is the guarantee.
model
one completion call per store-period reading, one system message and one user message, max_tokens 20000, temperature at the provider default, reasoning left at the provider default. Any OpenAI-compatible provider or Anthropic via src/adapters/__init__.py.
1647 input and 1016 output tokens a reading over 117 readings; $0.003870 a reading projected on the shared card; p50 7201 ms, p95 17504 ms end to end. 90.6 pct of output was provider-side reasoning. Largest reply 3469 tokens under a 20000 cap; 0 unparsed replies. (r001-royalty-watch, c000-royalty-watch-calibration)
the output ceiling. The calibration ran four readings at max_tokens 2,000 and the largest used 1752 of it -- 88 pct on four calls -- so a forker who caps low to save money will lose replies to truncation, and an unparsed reply is a miss in every field. A ceiling is not a cost: the provider bills tokens produced.
changing the model, the ceiling or the reasoning setting invalidates every model figure on these pages. The three free floors are unaffected -- they call nothing -- which is why the comparison survives a model change and the headline does not.
labels
the answer key is GENERATED, not stored: evals/answers.py walks the schedule applying src/rules.py, carrying state exactly as a deployment would. data/gold.jsonl holds what the corpus GENERATOR planted and is not the key. The documents are this repository's own synthetic pack, MIT.
117 labelled readings over 40 stores: FIRST_READING 40, NEW_CROSSING 23, STILL_STANDING 22, RETURNED_TO_BAND 8, NO_CHANGE 24. Six operating-note classes: coded_stale 4, coded_valid 6, coded_wrong_indicator 2, none 89, prose_unlisted 4, prose_valid 12. The pre-flight asserts rules.outside() agrees with the generator's independent statement on all 117, and refuses any reading whose derived figure sits within 0.15 of its threshold. (evals/check_labels.py, tools/build_corpus.py)
the corpus stops scoring when the model stops being wrong, and it already has: every model arm scored 100 pct on every field, so this set cannot rank two models or notice a small regression. The next corpus needs paraphrased causes, more than one note per period, and readings near the band edge -- all three deliberately absent here.
swapping in your own documents invalidates every rate on these pages and keeps every mechanism. src/rules.py is the file to replace first; the key regenerates itself from it.
corpus refresh
the population is re-read WHOLE each run and membership is read off disk, so a store that opened since the last run appears with no carried state and a store that closed simply stops being owned. A first sight is FIRST_READING and never a change, whatever the figures say.
40 of the 117 readings are a store's first sight by this monitor: 38 opening readings in period P1, plus 2 stores that opened between run 1 and run 2, and 1 store that closed after run 2. The model answered all 40 correctly. (r001-royalty-watch, evals/check_labels.py)
a store that leaves the population and comes back. Nothing here tests it: a re-opened store would still carry state from before it closed, and whether that state is meaningful after a gap is a policy decision the kit does not make.
nothing measured, as long as membership is contiguous -- the pre-flight asserts it. A non-contiguous membership would make cadence.gap() misreport how many runs a reading stands in for.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a store raised NEW_CROSSING for an indicator that was outside band last run too
the carried state did not reach the prompt, or reached it empty. The state is the only route between runs, and without it every outside-band reading looks new.
check the prompt's Carried state block before doubting the answer. The stateless arm is what this looks like at scale: 77 of 77 readings answered FIRST_READING. (results/eval-s001-royalty-watch-stateless.json, movement_confusion)
a store raised for an indicator whose operating note plainly names an approved cause
the note carries no cause CODE token and something in the path is reading notes with src/notes.claims() rather than claims_full() -- which is exactly what the free floor does, by design.
look at whether the reading came from a floor run or a model run. results/eval-b002-royalty-watch-regexnotes.json makes this error 5 times out of 28 raises and r001 makes it 0 times. (results/eval-b002-royalty-watch-regexnotes.json misses, e.g. STR-0006-P2)
a crossing dated to a period in which the indicator did not actually cross
a scheduled run did not happen. The reading was compared against the last reading anybody took, which is older than the last period.
read cadence.gap() for that store's carried last_run against the period being reported. A gap greater than 1 means this run is standing in for a run that never fired, and 11 of 15 such crossings in the measured ablation were misdated rather than missed. (results/eval-a000-royalty-watch-missedrun.json, cadence.cost)
a reply that did not parse
almost always the output ceiling. finish_reason is recorded on every failure row alongside the output token count and the cap.
read failures[].at_ceiling in the result file. 0 of 117 on the scored run at a 20000 cap; the calibration at 2,000 came within 88 pct of it on four calls. (results/eval-c000-royalty-watch-calibration.json)
No machine symptom — this failure leaves no trace in any output.
There is no machine symptom for this and there cannot be one inside the kit: a run that never starts writes no result file, produces no failure row and leaves no trace anywhere in this repository. The defence is entirely outside it -- a scheduler that alarms on a run that did not START, not only on one that failed. It is named first in guardrails.add_first and it is the single most expensive gap in this kit, priced at 4 of 15 crossings lost permanently per missed monthly run.
⚠︎ THE SCHEDULER. This kit ships a schedule and no runner: src/cadence.py is a contract and nothing executes it. Every cadence figure here comes from replaying a schedule inside an eval, not from a monitor that actually woke up on a clock for three months. What a real cron does when a run overruns, double-fires or half-completes is unmeasured, and the traceless signature above is the shape of the risk.
⚠︎ CONCURRENCY BEYOND 12 WORKERS. Everything here ran 12 wide. Whether a 2,000-store population at 64 wide hits a provider rate limit, and what the retry policy does to the wall clock when it does, has not been tried -- and the wall clock is what decides whether a run finishes before the next one is due.
⚠︎ THE STATE STORE. State is scoped to a run so the eval is reproducible. No persistence layer was chosen, sized or measured, and the ceiling row above says seam:null rather than inventing a join.
⚠︎ PROVIDER-SIDE RETENTION. What the provider keeps of 192,660 input and 118,834 output tokens, and for how long, is the provider's policy and is not measured here. Two sections are withheld from every prompt precisely because that question has no answer this kit can give.
⚠︎ REASONING DISABLED. 90.6 pct of output tokens were provider-side reasoning at the default. The adapter can send the disabled-thinking field; this kit never has, so the cheaper configuration's accuracy is unknown.
⚠︎ GPU OR SELF-HOSTED SIZING. Nothing was run on a local model. src/adapters points at any OpenAI-compatible endpoint including a local server, and no figure here says what hardware that would need.
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved, because there is no third-party material here to grant. The pack reproduces no real franchise agreement, no franchise disclosure document and no state franchise-relationship statute, and quotes from none. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The movement, the outside-band set, the new set and the exemption set, per reading, exact match against the generated answer key
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineThe movement, the outside-band set, the new set and the exemption set, per reading, exact match against the generated answer key
For each of the 117 readings and each of the four answered fields, did the reply equal the key. A set is compared as a sorted tuple, so order in the reply is irrelevant and membership is everything.
$0.00per 1,000 store-period readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in process, no key and no model. The same function scores the three free floors, which is what makes the comparison a comparison.
Every grader on these pages scored the same 117 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
STR-0006-P2 -- a MALL_FOOD_COURT store, second scheduled run
The arithmetic, in code
Average ticket 6.10 per cent below the store's own 12-period baseline, against a 5.0 per cent band. The other three indicators are inside: cash share 0.40 points down against 6.0, deposit gap 0.34 per cent against 3.0, voids 0.90 points above the segment median against 2.0.
Carried state, in the prompt
The last run read this store in period P1 and no indicator was outside its band. That run did not raise a new crossing.
The operating note
"The dining room closed for approved remodel from 2026-02-03 to 2026-02-21, which the area coach visited and confirmed. No cause code was entered against the period." REMODEL_CLOSURE is on the approved list, its dates are inside the period, and it covers TICKET_DRIFT.
The strongest free floor
NEW_CROSSING, indicators_new TICKET_DRIFT, exemptions none. The regex finds no cause code token in the note, so it excuses nothing and raises the store.
NO_CHANGE / outside TICKET_DRIFT / new none / excused TICKET_DRIFT
Scored as
4 of 4 fields correct, and the free floor is wrong on three of them. This single reading is the whole of this kit's margin over free code, which is why it is also the published screenshot.
Grader
Verdict
Why
The movement, the outside-band set, the new set and the exemption set, per reading, exact match against the generated answer key
movement hit, outside-band hit, new-set hit, exemption-set hit
All four answered fields equal the key: NO_CHANGE, outside [TICKET_DRIFT], new [], excused [TICKET_DRIFT]. The strongest free floor matches on one of the four.
Missed crossings and false raises, counted separately against the key
not a crossing, and not raised -- neither direction fires
The key movement is NO_CHANGE, so this reading is not in the detection denominator; the model did not raise it, so it is not in the false-raise denominator either. The free floor DID raise it, which is one of that arm's 5 false raises.
The negative assertion: does the description state a conclusion rather than describe a reading
clean -- 0 forbidden phrases
The reply describes the ticket figure against the baseline, names the three indicators inside their bands, and says the operating note covers TICKET_DRIFT and falls within the period. It states no cause and no consequence, which is the whole constraint.
The formulaWhat it computes
accuracy = hits / answered, per field, reported apart and never summed. A reply that did not parse counts as a miss in every field.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
scored 100.0%
the fast tier, one scheduled run MISSED
scored 100.0%
the strongest free floor, no model
scored 95.7%
the threshold floor with no memory, no model
scored 32.5%
In operationWhat to monitor
Reference standard: The answer key evals/answers.py generates by walking the schedule -- src/rules.py's five ordered tests applied to the printed figures and the operating note, carrying state exactly as a deployment would. It is CODE, not a person and not a model, so it cannot be wrong about itself; it can only be wrong about the policy, and that risk is stated in Eval.validated rather than hidden in a rate.
These rates are UNKNOWN, on purpose
This grader IS the reference standard, so its own TPR and TNR are not measured and cannot be -- scoring it against itself is circular and would read as evidence. What CAN be wrong is the key it applies, and Eval.validated names exactly what: src/rules.py holds both the policy the prompt paraphrases and the arithmetic the key applies, so a 100 per cent score says the model reproduces that file and not that the file is right.
Watch these
movement_accuracy_pct
change_movement_accuracy_pct
change_detection_pct
false_raise_rate_pct
outside_set_exact_pct
exemption_set_exact_pct
conclusory_language_rate_pct
unparsed_replies
Alarm on
conclusory_language_rate_pct FIRST, on any value above 0.0 -- it is the one figure with no acceptable trade-off, because a pack that states a conclusion has done the thing this kit's cap forbids whatever else it got right. Then change_detection_pct, because a missed crossing can become permanently unrecoverable at the next run and a false raise cannot.
How tight can the band be? The denominators are small and they differ per figure, which is why they travel beside every rate: movement is over 117 readings, the memory-dependent figure over 77, detection over 23 crossings and the false-raise rate over 23 raises. One crossing is 4.35 points of detection, so no alarm band on that figure can be finer than about 4 points without alarming on a single row. The language check is the exception and it is a COUNT, not a rate: the band is one reply.
Cadence: Free, so re-run it on every change to tools/build_corpus.py, src/rules.py, src/notes.py, src/state.py or src/cadence.py. evals/check_labels.py runs first and refuses to let a scored run spend if the corpus and the rulebook have drifted apart.
The decisionWhen to reach for it
Use it
The truth is known and every answer is a word or a code from a closed list.
Do not use it
The truth is not known -- the normal state of a real reporting rule, whose exemptions are argued rather than computed.
Missed crossings and false raises, counted separately against the key
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineMissed crossings and false raises, counted separately against the key
Of the readings that genuinely were a new crossing, how many were raised (detection); and of the readings that were raised, how many were not crossings (false raise).
$0.00per 1,000 store-period readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, the same pass as the grader above.
Every grader on these pages scored the same 117 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
STR-0006-P2 -- a MALL_FOOD_COURT store, second scheduled run
The arithmetic, in code
Average ticket 6.10 per cent below the store's own 12-period baseline, against a 5.0 per cent band. The other three indicators are inside: cash share 0.40 points down against 6.0, deposit gap 0.34 per cent against 3.0, voids 0.90 points above the segment median against 2.0.
Carried state, in the prompt
The last run read this store in period P1 and no indicator was outside its band. That run did not raise a new crossing.
The operating note
"The dining room closed for approved remodel from 2026-02-03 to 2026-02-21, which the area coach visited and confirmed. No cause code was entered against the period." REMODEL_CLOSURE is on the approved list, its dates are inside the period, and it covers TICKET_DRIFT.
The strongest free floor
NEW_CROSSING, indicators_new TICKET_DRIFT, exemptions none. The regex finds no cause code token in the note, so it excuses nothing and raises the store.
NO_CHANGE / outside TICKET_DRIFT / new none / excused TICKET_DRIFT
Scored as
4 of 4 fields correct, and the free floor is wrong on three of them. This single reading is the whole of this kit's margin over free code, which is why it is also the published screenshot.
Grader
Verdict
Why
The movement, the outside-band set, the new set and the exemption set, per reading, exact match against the generated answer key
movement hit, outside-band hit, new-set hit, exemption-set hit
All four answered fields equal the key: NO_CHANGE, outside [TICKET_DRIFT], new [], excused [TICKET_DRIFT]. The strongest free floor matches on one of the four.
Missed crossings and false raises, counted separately against the key
not a crossing, and not raised -- neither direction fires
The key movement is NO_CHANGE, so this reading is not in the detection denominator; the model did not raise it, so it is not in the false-raise denominator either. The free floor DID raise it, which is one of that arm's 5 false raises.
The negative assertion: does the description state a conclusion rather than describe a reading
clean -- 0 forbidden phrases
The reply describes the ticket figure against the baseline, names the three indicators inside their bands, and says the operating note covers TICKET_DRIFT and falls within the period. It states no cause and no consequence, which is the whole constraint.
The formulaWhat it computes
detection = raised-and-true / true; false raise = raised-and-false / raised. Two denominators, deliberately different, never combined into an F-score.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
scored 100.0%
the strongest free floor, no model
scored 100.0%
the threshold floor with no memory, no model
scored 100.0%
In operationWhat to monitor
Reference standard: The same generated answer key, restricted to the readings whose key movement is NEW_CROSSING (detection) and to the readings the arm actually raised (false raise).
These rates are UNKNOWN, on purpose
This grader has no rates of its own to be blank. It is arithmetic over the reference grader's per-reading verdicts -- two counts and two denominators -- so there is no second opinion it could be scored against. What is unmeasured is whether the DENOMINATORS are big enough to carry a rate at all: 23 crossings and 23 raises is a small set, and every figure it produces is wide.
Watch these
change_detection_pct
false_raise_rate_pct
crossings_missed
crossings_lost — on the missed-run arm only
Alarm on
crossings_missed, on any increase at all. It is a COUNT and not a rate because the denominator is 23 and a rate over 23 rows hides the first one.
How tight can the band be? 23 crossings and 23 raises. One row is 4.35 points of detection and 4.35 of the false-raise rate, so neither band can be finer than about four points.
Cadence: Free, so re-run it on every change to tools/build_corpus.py, src/rules.py, src/notes.py, src/state.py or src/cadence.py. evals/check_labels.py runs first and refuses to let a scored run spend if the corpus and the rulebook have drifted apart.
The decisionWhen to reach for it
Use it
The two errors cost different things. Here they cost very different things: a false raise wastes an operator's morning, and a missed crossing can become permanently unrecoverable at the next run.
Do not use it
The classes are balanced and both errors cost the same, in which case one accuracy says it all and two figures are noise.
The negative assertion: does the description state a conclusion rather than describe a reading
Catch new royalty problems at your franchise stores
PresenterOpens the private repo. Visible to admins only.
In one lineThe negative assertion: does the description state a conclusion rather than describe a reading
Whether the reply's one free-text field contains any phrase from src/rules.FORBIDDEN -- 22 of them, matched case-insensitively on word boundaries with hyphen and space treated alike.
$0.00per 1,000 store-period readings
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::conclusory, in process. RED-PROVEN IN BOTH DIRECTIONS by evals/check_labels.py before any run may spend: it must convict a seeded conclusory sentence and acquit a seeded descriptive one. A checker that has only ever returned zero is indistinguishable from one that is not wired up, and this kit's whole cap rests on this function.
Every grader on these pages scored the same 117 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The reading
STR-0006-P2 -- a MALL_FOOD_COURT store, second scheduled run
The arithmetic, in code
Average ticket 6.10 per cent below the store's own 12-period baseline, against a 5.0 per cent band. The other three indicators are inside: cash share 0.40 points down against 6.0, deposit gap 0.34 per cent against 3.0, voids 0.90 points above the segment median against 2.0.
Carried state, in the prompt
The last run read this store in period P1 and no indicator was outside its band. That run did not raise a new crossing.
The operating note
"The dining room closed for approved remodel from 2026-02-03 to 2026-02-21, which the area coach visited and confirmed. No cause code was entered against the period." REMODEL_CLOSURE is on the approved list, its dates are inside the period, and it covers TICKET_DRIFT.
The strongest free floor
NEW_CROSSING, indicators_new TICKET_DRIFT, exemptions none. The regex finds no cause code token in the note, so it excuses nothing and raises the store.
NO_CHANGE / outside TICKET_DRIFT / new none / excused TICKET_DRIFT
Scored as
4 of 4 fields correct, and the free floor is wrong on three of them. This single reading is the whole of this kit's margin over free code, which is why it is also the published screenshot.
Grader
Verdict
Why
The movement, the outside-band set, the new set and the exemption set, per reading, exact match against the generated answer key
movement hit, outside-band hit, new-set hit, exemption-set hit
All four answered fields equal the key: NO_CHANGE, outside [TICKET_DRIFT], new [], excused [TICKET_DRIFT]. The strongest free floor matches on one of the four.
Missed crossings and false raises, counted separately against the key
not a crossing, and not raised -- neither direction fires
The key movement is NO_CHANGE, so this reading is not in the detection denominator; the model did not raise it, so it is not in the false-raise denominator either. The free floor DID raise it, which is one of that arm's 5 false raises.
The negative assertion: does the description state a conclusion rather than describe a reading
clean -- 0 forbidden phrases
The reply describes the ticket figure against the baseline, names the three indicators inside their bands, and says the operating note covers TICKET_DRIFT and falls within the period. It states no cause and no consequence, which is the whole constraint.
The formulaWhat it computes
rate = replies containing at least one forbidden phrase / replies. The only acceptable value is 0.0; this is a constraint, not a trade-off.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
0 false pass, 0 false fail
the fast tier, memory removed (THE CONTROL)
0 false pass, 0 false fail
the fast tier, one scheduled run MISSED
0 false pass, 0 false fail
the strongest free floor, no model
0 false pass, 0 false fail
In operationWhat to monitor
Reference standard: src/rules.FORBIDDEN, a closed list of phrases. There is no truth to compare against here -- the list IS the standard, which is why it is written down in the kit rather than described on a page.
These rates are UNKNOWN, on purpose
There is no rate of its own to publish because there is no truth to compare against: the forbidden-phrase list IS the standard here, so this grader cannot be wrong about a match. What it can be wrong about is COVERAGE -- an implication that uses none of the 22 phrases passes, and 'the deposit pattern speaks for itself' is the example. Nobody has measured how often that happens, because measuring it needs a second, independent judgement about tone that this kit deliberately does not have.
Watch these
conclusory_language_rate_pct
conclusory_replies
the phrases each hit matched — results[].language_hits[].phrases
Alarm on
any hit at all. One reply out of 117 is 100 per cent of this kit's acceptable count.
How tight can the band be? A count, not a rate. The band is ONE REPLY, across every run, and 257 replies over four model runs have produced none.
Cadence: Free, so re-run it on every change to tools/build_corpus.py, src/rules.py, src/notes.py, src/state.py or src/cadence.py. evals/check_labels.py runs first and refuses to let a scored run spend if the corpus and the rulebook have drifted apart.
The decisionWhen to reach for it
Use it
The constraint on the output is about what it may not SAY, and the forbidden language is a closed list somebody is willing to write down.
Do not use it
The constraint is about tone or implication rather than vocabulary. A phrase list cannot catch 'the deposit pattern speaks for itself', and this grader would pass it.
A living map of modern AI — kept current every morning