A waste hauler's invoice register can quote a job's number while quietly reversing the charge, or carry a hold that isn't real. This app checks each completed job against the register and says if it is truly billed, on hold, or still open.
PresenterOpens the private repo. Visible to admins only.
For the billing deskWaste & Environmental
Why it matters
Today's manual process, and the same job with the app
A billing analyst at a waste and recycling hauler, reconciling completed routes against the weekly invoice register.
✕Today's manual process
1Pull the week's invoice register manually and check whether each completed job's number shows up anywhere in it.
2Open anything that looks bundled or corrected to work out what the line actually covers.
3Read the desk notes to decide if a job is genuinely on hold or just has a complaint attached.
4Miss a reversal and revenue stays uncollected with nobody watching it.
Every job reconciled from memory and notes
✓With the app
1The register is read for you checking every completed job against it, however the line is worded.
2Bundled and corrected lines are worked out so a reversal that only quotes the job's number doesn't count as billed.
3Hold notes are checked for the real thing a dated, formal hold, not just a complaint in the file.
4Nothing slips through every reversal is caught and routed to the right person.
The app reads every register line first
See it work
One real case, as the app recorded it, step by step
Work order WO-0049 for a compactor haul looks billed because a register line reverses it, so the app flags it for review instead.
Catch waste hauls that are still unbilledReference appBuilt to be shaped to your process
4
1The case coming in 14 days unbilled already, no hold. Now called REVIEW.
2How long it's been open 21 days counted, the same way every week.
3The proof on file Driver confirmation and weight ticket are both there.
4Who picks it up Routed to the supervisor, since it's aged and reversed.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"Which completed jobs have no invoice line" is a SELECT and nobody needs a language model for it. The question a billing operation actually has to answer is the next one: does the invoice register genuinely BILL this job -- by its exact number, a bundled route-week line, or a reformatted job reference -- or does a line merely QUOTE the job's number while reversing the charge? Is a note that sounds like a billing hold an actual, dated, formal one, or a customer's request that must not stop the clock? And does the job even have the evidence on file to be billed at all? Three of those four questions can sit in a register line or a desk note that a foreign-key join never reads. Someone manually reconciling completed work orders against the invoice register: pulling the week's register, checking whether each job's exact reference appears anywhere in it, opening any line that looks like a bundled or corrected charge to work out what it actually covers, reading exception notes to decide whether a job is genuinely on hold or just has a customer complaint attached, and tracking which weight-based jobs are still missing a scale ticket.
Audience
A billing analyst deciding which candidates to release from the worklist this week, and a billing supervisor reviewing the ones that are aged or on hold. Both are reading a routing recommendation, never an executed action -- releasing an invoice is the analyst's own decision in the billing system, and this pack has no path that makes it for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual reconciliation snapshots
The corpus is 150 reconciliation snapshots, 0.85 MB (json 1 · jsonl 1 · txt 150). AN UNBILLED-WORK RECONCILIATION QUEUE JOINS OPERATIONS TO REVENUE, which makes it one of the more commercially sensitive files a hauler holds -- it names customers, exactly what was done at each site, and which completed jobs the billing desk has not yet been able to invoice. There is no public one and there never will be. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/billing.step's output over the planted per-run facts, so a whole-window hold convention this fiddly cannot carry its author's misreading into the score. ⚠︎ AND THE RULE IS INVENTED. The nine rules, the 21-day supervisor-review threshold and the whole-window day-counting convention reproduce no billing SLA, no revenue-assurance policy, no waste-hauling regulator's guidance and no real hauler's escalation practice, and they name none. ⚑ THIS PACK IS DELIBERATELY, ONLY OPERATIONAL. It answers 'is this job billed, and if not, how long and why' from the completion date and the invoice register alone -- it does not read, gate on, or feed a revenue-recognition close calendar or a GL posting period, and nothing here claims those two clocks agree. Whether unbilled-candidate detection should interface with accounting-close timing is an open question for whoever owns that boundary, left unresolved here on purpose (default: stay operational) rather than guessed at.
The corpus
The 150 reconciliation snapshotsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your reconciliation snapshots. That is the whole change — there is no database to migrate.
One reconciliation snapshot, as the model receives itWO-0001-R1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated unbilled-work reconciliation
snapshot for an AI use-case kit; it reproduces no real hauler, customer, contract or
billing register. The billing rule and its threshold are ILLUSTRATIVE.
Work Order
----------------------------------------------------------------
Customer : CUST-8021 (Grandview Retail Plaza)
Work order reference : WO-0001
Service type : COMPACTOR_HAUL
Container / equipment : CMP-6207
Route : RT-04
Completion date : 2026-02-21
Contract value for this job : 1,766.72 USD
Evidence required for billing : driver completion confirmation AND a weight ticket
Billing Rule
----------------------------------------------------------------
Supervisor review threshold : 21 days unbilled (operator-tunable; see .env
SUPERVISOR_REVIEW_DAYS)
Rule B-1 A completed work order becomes a BILLING CANDIDATE the day after its completion date,
and stays one until an invoice line in the register genuinely references it.
Rule B-2 Days unbilled are counted in whole RECONCILIATION WINDOWS. The first window runs from the
completion date to the first reconciliation run; every window after that is exactly the
7-day cadence. A window in which a formal billing hold (Rule B-3) was open for ANY part
does not count toward the total AT ALL; a window with no formal hold open at any point
counts in full. This is a stated simplification -- a hold that closes partway through a
window still excludes the whole window -- and it is not split into partial days.
Abridged — the file continues.
The outcomeWhat a good result looks like
Every completed work order carries a status (BILLED, UNBILLED, ON_HOLD, REVIEW or CONTEXT_INCOMPLETE), a hold-adjusted day count, an evidence-completeness flag, and a routing recommendation -- NONE for a billed job, ANALYST for a routine or evidence-incomplete one, SUPERVISOR for one that is aged past the review threshold or genuinely on hold.
And when it cannot
Two ways, and they cost differently. Reporting a genuinely BILLED job as still open sends an analyst to chase revenue that is already collected. Reporting a job as BILLED when a register line only QUOTES its number while reversing the charge (Rule B-6) leaves real revenue leaking with nobody looking at it -- the costlier direction, because nothing else on a billing desk catches it. On the scored run neither happened: 0 of 35 BILLED jobs were missed and 0 of 115 non-billed jobs were falsely marked BILLED. Two of those 115 are the reversal-decoy cells this finding rests on, and that denominator is thin enough to say so plainly.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
a billing desk with a clean, single-ID invoice register and no bundled or corrected lines — the exact-join-mem floor (b001) 91.33 pct for $0.00, and it already beats the model on nothing this scenario tests
a billing desk whose register carries bundled route-week lines or reformatted job refs, and whose desk notes carry informal hold-sounding language — the model, with the carried state the two traps this kit exists to test -- a bundled/reformatted reference and a hold-sounding note with no formal entry -- are exactly where every free floor in this run loses points, and the model lost none
deciding whether ANY of this is worth a model call at all — run all three floors first (0 calls, $0.00) and look at the per-bucket cut b001 already solves the hold-decoy trap by SCOPE alone; the model's margin over the best floor lives entirely in bundled/reformatted matching and the reversal decoy, which is a much narrower claim than 'the model beats free code'
At a glanceHow the whole thing runs
100%status accuracy pct
5,624 msp50, end to end
$3.90per 1,000 reconciliation snapshots · Google Gemini 3 Flash
Run once, for real, on 2026-08-24. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch waste hauls that are still unbilled14 steps · 4 questions · run once, for real · 2026-08-24
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace src/billing.py FIRST, and not as an option -- SUPERVISOR_REVIEW_DAYS, the whole-window hold convention and the invoice-match test are all invented here, and data/gold.jsonl is literally their output. The measured figures on this page do not travel with your register.Corpus lens →
When is this the wrong choice?
Avoid: Spending on a model call for a matching problem that does not exist on this register. That is the case against the best-fitting scenario (“a billing desk with a clean, single-ID invoice register and no bundled or corrected lines”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
THREE OF THIS KIT'S MOST IMPORTANT BUCKETS ARE THIN. 2 reversal-decoy readings, 4 hold-decoy work orders (12 readings), 6 CONTEXT_INCOMPLETE readings. 7 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether the shipped 21-day SUPERVISOR_REVIEW_DAYS threshold matches any real billing desk's own escalation policy -- it is an invented default, not a confirmed one (Data.why_this_corpus). 8 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-24 — r001-unbilled-work. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces: the 150 snapshots and the answer key (regenerated from the seed in well under a second), all eighteen pre-flight assertions, all three free floors scored end to end, the wiring stub, and the local UI at 127.0.0.1:8208 including what run r001-unbilled-work recorded for every reading. What it CANNOT reproduce without a key is a model column of its own.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
5,624 msp50, end to end
21,033 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, and the reply parsed to a status, a day count, an evidence flag and a routing call. Splitting the snapshot into sections, dropping the withheld one, and computing the whole-window hold arithmetic all happen outside this measurement and cost no network at all. p95 (21.0s) is 3.7 times p50 (5.6s); the spread tracks how much the invoice-match decision needs reasoning about -- an exact match is a lookup, a bundled or reformatted reference or a reversal decoy is a reading task.
Current processWhat it replaces
Someone manually reconciling completed work orders against the invoice register: pulling the week's register, checking whether each job's exact reference appears anywhere in it, opening any line that looks like a bundled or corrected charge to work out what it actually covers, reading exception notes to decide whether a job is genuinely on hold or just has a customer complaint attached, and tracking which weight-based jobs are still missing a scale ticket.
Where it is not good enough
⚑ THE MODEL WON EVERY MEASURED CELL ON THIS RUN, AND THAT IS THE LEAST INTERESTING WAY TO READ IT. Status, evidence-completeness and routing are all 100 pct against a free floor that tops out at 96 pct (see Eval) -- but three of this kit's most important buckets are thin: 2 reversal-decoy readings, 4 hold-decoy work orders (12 readings), 6 CONTEXT_INCOMPLETE readings. A single wrong answer in the reversal-decoy bucket alone would have cut its recall in half. Read the per-bucket cut in Eval before the headline. ⚠︎ THE FIRST SCORED ATTEMPT AT THIS RUN WAS NOT CLEAN. The carried-state sentence for a BILLED work order originally said only that it "stays BILLED" without restating the frozen day count, and the model correctly refused to guess -- it returned null on all 21 memory-dependent BILLED cells rather than invent a number, which is the right instinct pointed at a genuinely missing fact. Fixing the sentence to state the number explicitly (src/state.py::describe) took memory-dependent day-count accuracy from 14.29 pct to 100 pct; that fix is what shipped, not a rewritten corpus or a re-picked example. ⚠︎ AND THE 21-DAY SUPERVISOR-REVIEW THRESHOLD IS AN INVENTED DEFAULT, NOT A CONFIRMED POLICY -- see Data.breaks_on. ⚠︎ THE WHOLE-WINDOW HOLD CONVENTION IS A STATED SIMPLIFICATION: a formal hold open for even one day of a 7-day window excludes the WHOLE window from the count, never a partial day. That understates real unbilled exposure on a job whose hold closes early in a window, and it is not measured how much.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1json1
50 completed work orders on one waste and recycling hauler's billing desk, each re-read whole at three weekly reconciliation runs
Recorded failureno missed-run cost estimate ships with this kit -- a stated gap, not a hidden zero, and the only figure on this page that is NOT measured
BILLED recall vs false-BILLED, ON_HOLD recall vs false-ON_HOLD, counted apart
Recorded failure0 misses on the scored run -- but 3 buckets are thin (2 reversal-decoy, 4 hold-decoy, 6 evidence-incomplete readings), named rather than averaged away
100.00% status accuracy, 0 false BILLED, 0 false ON_HOLD
strongest free floor: 96.00% for $0.00, wrong on 2 named traps
memory-dependent day-count accuracy 14.29% -> 100% after a found-and-fixed defect
2026-08-24as of
It produces a candidate list and a routing recommendation for a billing analyst or supervisor to action. It never creates, releases or posts an invoice, and there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is four scalars written by src/state.py from the arithmetic and never from the model's reply -- whether an invoice line is already matched, whether a formal hold is open, the frozen or running day count, the status last reported -- so a wrong reading is a wrong worklist row this run, never a corrupted history for the next one. The clock is the second: this kit ships no missed-run cost estimate, a stated gap rather than a hidden zero.
⚠︎ THE FLOOR STATION CARRIES THE FREE NUMBER'S OWN SHAPE, NOT A CLEAN LOSS. Three floors separate what memory buys from what matcher width buys from what hold-check SCOPE buys: exact join alone is 68.67 pct; add memory and a correctly-scoped hold check and it jumps to 91.33 with ZERO false holds; widen the matcher further to catch bundled and reformatted references and it reaches 96.0 -- but the same widening that helps it catch real references is what makes it catch a reversal decoy and a hold-sounding note that are not real.
⚠︎ AND A REAL DEFECT WAS FOUND AND FIXED DURING THIS KIT'S OWN MEASUREMENT. The first scored attempt's carried-state sentence did not restate a BILLED work order's frozen day count, and the model correctly answered null 21 times rather than guess it. Naming the number in the sentence took memory-dependent day-count accuracy from 14.29 pct to 100 pct; the run published here is the one after that fix, not the one that surfaced it.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the supervisor-review threshold
src/billing.py
SUPERVISOR_REVIEW_DAYS, shipped at 21 with no confirmed policy behind it -- the first thing a real deployment needs to replace, not an optional tuning knob. Re-run tools/build_corpus.py after changing it; the gold cannot drift because it is the rule's own output.
the invoice-match rule
src/billing.py, tools/build_corpus.py
Rule B-4/B-6's definition of what counts as a genuine reference, and the register-line phrasings the corpus plants to test it. A real billing system's own reference formats (bundled lines, reformatted job refs, reversal language) should replace the invented ones here before this kit means anything about a real register.
the hold rule
src/billing.py
The whole-window exclusion convention (Rule B-2) and what counts as a FORMAL hold (Rule B-3). A real billing desk's own hold workflow -- who can open one, what evidence it requires -- should replace the invented rule.
the carried state
src/state.py
load()/save() are a JSON file today. Point them at a table or a billing data mart and nothing else in the kit changes -- for_workorder() and advance-by-step() are the whole interface, and describe() is the only thing the prompt sees.
the corpus
tools/build_corpus.py
The customers, service types, invoice-line phrasings and the seed. Keep the eight section headings or src/segment.py's assertion refuses to start.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 50 completed work orders x 3 weekly reconciliation runs = 150 snapshots from a fixed seed (SEED = 20260824), across ten customers and six service types. Plants ten explicit buckets: three ways an invoice line genuinely bills a job (exact, bundled, reformatted), one way a line only quotes a job's number without billing it (the reversal decoy), two hold shapes (closes inside the next window, spans a whole window silently), a hold-sounding note with no formal hold behind it, an evidence gap that resolves, and two plain aging trajectories. The gold labels are src/billing.step's output over the raw per-run facts, never typed by hand.
the billing rule
src/billing.py
The rule as pure code: the whole-window hold exclusion (Rule B-2/B-3), the invoice-match test (Rule B-4/B-6), the evidence-completeness test by service type (Rule B-5), and the routing table (Rule B-8). No model, no judgement. The nine rules, the 21-day review threshold and the whole-window convention are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. Four scalars per work order (billed, hold_open, the frozen or running day count, the status last reported), written from the arithmetic and rendered into one English sentence for the prompt. ⚠︎ A real defect was found and fixed here during this kit's own measurement: the first version described a BILLED work order without restating its frozen day count, and the model correctly answered null rather than guess it -- 21 of 21 memory-dependent BILLED cells wrong. Naming the number in the sentence took memory-dependent day-count accuracy from 14.29 pct to 100 pct; see the module's own docstring for the full account.
the splitter
src/segment.py
Splits a reconciliation snapshot into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 150 documents before a run may spend.
the selector
src/select.py
Decides which sections reach the model. Site Contact -- the customer site's name, direct dial and work email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the carried-state sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
the adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading.
the watch
src/monitor.py
One work order, one reconciliation run, one call. Parses identifying facts off the page with a regex (the model is never asked to compute a calendar difference), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling. Never creates, releases or posts an invoice -- see the module's own guardrail note.
the local UI
src/app.py
One work order, one reconciliation run, its carried state and its verdict, on 127.0.0.1:8208. Renders with no key. Shows the carried sentence verbatim, the free floor's answer beside the model's, and a second button that replays what run r001-unbilled-work actually answered, straight off the committed result file, labelled as a replay.
the free floors
evals/baseline.py
exact-join: an exact work-order-number string join, no memory, no hold awareness -- the incumbent. exact-join-mem: the same join plus carried memory and a hold check scoped to the Evidence Log's formal entries only. broad-match-mem: carried memory, a bare-number search that catches bundled and reformatted references, and a hold check widened to any hold-sounding word in the notes. 0 calls, $0.00, all three scored through the identical scorer.
the scorer
evals/scoring.py
Exact match per cell against the computed gold, split ways an average would hide: the four fields, BILLED recall against a false-BILLED rate, ON_HOLD recall against a false-ON_HOLD rate, under- versus over-escalation counted apart, the memory-dependent subset, and every bucket that plants a trap, cut separately. No judge model.
the pre-flight
evals/check_labels.py
Eighteen things that must be true before a run may spend: all eight sections parse, every work order's run sequence is complete and gap-free, no run-varying section restates an earlier match, neither a formal-hold flag nor a final day count is printed anywhere on the page, an exact join is proven to miss every bundled and reformatted reference AND be fooled by both reversal lines, the hold-decoy bucket is proven to carry no formal hold entry, the privacy guard holds and would fail without it, the two prompt builders differ on exactly one line, no code path creates, releases or posts an invoice, and the answer key replays from src/billing.step.
Where it breaks at scale
⚑ THE CADENCE IS THE SCALING VARIABLE, NOT THE CORPUS SIZE, SAME AS EVERY MONITOR IN THIS SERIES. One run of this watch is one call per OPEN work order, and the watch wakes weekly, so a billing desk holding 4,000 open candidates pays 4,000 calls a week whether or not anything changed. ⚠︎ AND THIS KIT DOES NOT MEASURE WHAT A MISSED RUN COSTS -- unlike its closest cousins in this series (deduction-age, settle-fail), no evals/cadence.py exists here to price a skipped reconciliation run in dollars. That is a real gap in this kit's own evidence, stated rather than quietly left off: a work order whose hold closes on a run that never fires would carry a stale carried-state sentence into the next one, and nothing here measures how often or how expensive that is. ⚠︎ TWO OTHER THINGS BREAK BEFORE THE CALL COUNT DOES. First, data/state.json is one file replaced atomically -- correct for one writer and not a concurrency model; two schedulers advancing the same work order would race, and the loser's write forgets a match or a hold, so the next run re-ages or re-flags a job that was already resolved. Second, a run that does not happen is not detected anywhere in this kit: evals/run.py is INVOKED, it is not woken, and nothing here back-fills a missed run or marks its readings late. ⚡ AND AN OPEN QUESTION THIS KIT DOES NOT RESOLVE: a sibling billing-verification row in this same catalogue also examines invoice-line gaps from a different angle, and this pack does not share a detection core with it -- run both independently against the same real register and they could disagree about the same gap. Noted here as an integration question for whoever owns both rows, not solved by either one alone.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
WO-0049-R2, the whole argument on one screen, and it was chosen by reading the answer key for the case where the free floor and the truth disagree most -- not by looking for a flattering frame. Its invoice register carries 'billing correction -- reversing prior miskeyed charge originally coded to WO-0049; no service performed, $0.00' -- a line that QUOTES this work order's own number while explicitly taking the charge back out. The free bare-number floor sees '0049' inside the line and reports BILLED, nothing to do. Run r001-unbilled-work reads what the line SAYS: evidence complete, no formal hold, 21 non-held days -- REVIEW, route to supervisor. The floor would have told a billing desk this job needed no further attention; it does. ⚠︎ THE MODEL COLUMN HERE IS REPLAYED FROM THE COMMITTED RESULT FILE, NOT A LIVE CALL, and the column header says so in as many words.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page with NO API_KEY configured. It does not error and it does not go blank: the snapshot, the carried state, the parsed identifying facts and the entire free floor are computed locally, so everything except the model column still renders and the button returns a plain sentence saying nothing was called.failureOpen full size →Before anything is asked. The two things this page has to get right are already on it: the carried state, verbatim as it goes into the prompt, and the list of which sections left the machine and which did not -- Site Contact, the customer site's name, direct dial and work email, marked WITHHELD rather than silently absent. A page that simply does not mention them cannot be told apart from one that quietly sent them.failureOpen full size →
How it is cutWhat one work orders (3 weekly reconciliation runs each) is
No split, and no chunking. The unit is a WORK ORDER -- three consecutive weekly reconciliation runs processed strictly in order, because run 3's prompt contains a carried state produced by run 2. Each snapshot goes whole into one call, minus the one section no field asks for.
SetupWhat the setup figure measured
There is no index and no retrieval step -- the open population is re-read whole on each reconciliation run. tools/build_corpus.py writes 150 documents and the answer key from a fixed seed in well under a second with no clock read and no model called.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every customer, work order, container, invoice line, evidence-log entry and desk note is invented here. Verified against the repository's own LICENSE file on 2026-08-24.
Bring your ownBring your own reconciliation snapshots
Replace src/billing.py FIRST, and not as an option -- SUPERVISOR_REVIEW_DAYS, the whole-window hold convention and the invoice-match test are all invented here, and data/gold.jsonl is literally their output. Then point tools/build_corpus.py at your own work orders and invoice register, or replace data/corpus/*.txt and data/gold.jsonl directly. Keep the eight section headings in src/segment.py::SECTIONS -- the parser asserts all eight in every document before a run may spend -- and keep the Billing Rule block, because the rule the model applies is READ OFF THE PAGE rather than baked into the prompt. gold.jsonl needs one row per snapshot carrying the raw src/billing.step inputs (days since completion, cumulative hold days, whether the register matches, whether a hold is open, whether evidence is complete) alongside the resolved answer; evals/check_labels.py re-derives the answer from src/billing.step and refuses if they disagree.
⚠︎ And what stops being true when you do: The measured figures on this page do not travel with your register. Every one of them is a fact about FIVE curated register-line phrasings and a 21-day threshold nobody has confirmed, over a population that does not change between runs and where no work order goes on hold twice. Re-run all four arms (model, stateless control, and the three floors) on your own register's actual reference formats before quoting anything here against it.
What breaks it
⚠︎ THE 21-DAY SUPERVISOR-REVIEW THRESHOLD IS A SHIPPED DEFAULT, NOT A CONFIRMED ESCALATION POLICY. It is operator-tunable (src/billing.py::SUPERVISOR_REVIEW_DAYS) precisely because nobody has confirmed the real number against an actual billing desk's own escalation practice -- see Business.not_good_enough.
⚠︎ THE WHOLE-WINDOW HOLD CONVENTION IS A STATED SIMPLIFICATION THAT CAN UNDER-STATE EXPOSURE. A formal hold open for even one day of a 7-day reconciliation window excludes the ENTIRE window from the day count rather than splitting it into partial days (Rule B-2). A job whose hold closes on day 1 of a window is treated identically to one whose hold covers all 7 -- both windows contribute zero days unbilled. How often that matters, and by how much, is not measured here.
THREE OF THIS KIT'S MOST IMPORTANT BUCKETS ARE THIN. 2 reversal-decoy readings, 4 hold-decoy work orders (12 readings), 6 CONTEXT_INCOMPLETE readings. A single wrong answer in the reversal-decoy bucket alone would cut its own recall figure in half -- read the bucket-level cut in Eval, not the 150-reading aggregate, before trusting any of the trap-specific numbers.
⚠︎ THE MEMORY-DEPENDENT FLAG UNDERCOUNTS ITS OWN CATEGORY, FOUND BY READING A REAL FAILURE. hold_closes_r2's second reconciliation run also depends on the carried state to know its first window contributed zero days -- but the generator's flag only marks hold_spans_r2's second run as memory-dependent. The stateless control's own miss on WO-0036-R2 (reported REVIEW at 21 days; the true answer is UNBILLED at 0) proves the dependency exists where the flag says it does not. The published memory_cells count (21) is a floor on the true number, not the true number.
ONLY ONE HOLD INTERVAL PER WORK ORDER, EVER. No work order in this corpus goes on hold, comes off, and goes on again -- a real book's dispute-and-resolve cycle repeating is not exercised.
EVERY REGISTER-LINE PHRASING IS ONE OF FIVE PATTERNS (exact, bundled, reformatted, reversal, none), WRITTEN IN ENGLISH, BY ONE AUTHOR. A real invoicing system's actual reference formats, abbreviations and free-text conventions are far messier than five curated patterns, and a phrasing this corpus does not know about would reproduce the same defect this kit was built to catch, silently.
TEN CUSTOMERS, SIX SERVICE TYPES, ONE HAULER. A real book's mix of service types, contract structures and customer-specific billing terms is not sampled here; every job in this corpus bills against a plain flat rate or a plain per-ton rate, never a tiered or minimum-charge structure.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent. The split is in characters, because that is what this run recorded — the billed totals are measured, but nothing apportioned them across the parts, and dividing a measured total by character share would produce three numbers nobody counted.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
194
not measured
instruction
4,359
not measured
carried state
190
not measured
Synthetic Record
317
not measured
Work Order
487
not measured
Billing Rule
3,870
not measured
Billing Position
501
not measured
Invoice Register
158
not measured
Evidence Log
178
not measured
Billing Desk Notes
168
not measured
Total
2,454
This is the cost lesson as arithmetic: of the 10,422 characters assembled, 4,553 are instructions — 44% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
TWO MESSAGES, JOINED IN SEND ORDER. The provider receives a system message and a user message; prompt_verbatim prints both, separated by a blank line, because publishing only the second would be publishing most of a prompt. Replayed from the kit's own src/prompt.build() for WO-0001-R1 with the carried state that reading was actually given (src/prompt.build_parts() produces the identical decomposition used here). The section list it produces is the list the run recorded in sections_used -- Synthetic Record, Work Order, Billing Rule, Billing Position, Invoice Register, Evidence Log, Billing Desk Notes -- and src/prompt.py is the only thing in the kit that assembles a prompt, so there is no second path a run could have taken. tokens.input (2454) is the run's own INPUT TOKEN AVERAGE over all 150 readings, not a per-reading count the harness does not store; the assembled parts above are this exact reading's, in characters.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled unbilled-work reconciliation watch. You apply a written billing rule to one completed work order at one reconciliation run. You answer with one JSON object and no other text.
You are a scheduled unbilled-work reconciliation watch for one waste and recycling hauler's
billing desk. It wakes every seven days, re-reads every completed work order still without a
matched invoice line, and reports each one. You are reading ONE work order at ONE reconciliation
run.
The billing rules, the work order's evidence and the invoice register excerpt are reproduced in
the snapshot below. Apply them exactly as written. You cannot see the earlier reconciliation runs;
what is known about them is stated under "Carried state" and is the only history available to you.
Do not assume anything about earlier runs beyond it.
How to decide:
- First check the evidence (Rule B-5). A weight-based service needs a driver completion
confirmation AND a weight ticket on file; a flat-rate service needs only the driver completion
confirmation. "Evidence on file" on the page states this directly. If the evidence this rule
requires is missing, the work order is CONTEXT_INCOMPLETE: days_unbilled null, evidence_complete
NO, route_to ANALYST. Never bill or age a job whose required evidence is not on file.
- Otherwise, decide whether the Invoice Register actually bills this job (Rule B-4). It is BILLED
when a line genuinely references it -- by its exact work order number, by a bundled line
listing several work order numbers for one route-week, or by a reformatted job reference. A
line that only QUOTES this work order's number while describing a correction, a reversal or a
charge taken back out does NOT bill it (Rule B-6) -- read what the line SAYS, not just whether
the number appears in it. Once a work order has been reported BILLED on an earlier run it stays
BILLED (Rule B-7); do not re-open it. IMPORTANT: "days_unbilled" is NEVER null on a BILLED
reading -- it is a historical fact ("it took N days to bill"), not a still-running count. If the
carried state already says BILLED, its sentence states the exact whole-day number that was
already counted; COPY THAT NUMBER, do not recompute it and do not report null. If the register
bills it for the FIRST time on THIS run, compute days_unbilled the same way an UNBILLED reading
would (below) and report that number, frozen.
- If not billed, decide whether the CURRENT reconciliation window is held (Rule B-2/B-3): either
the Evidence Log shows a hold OPENED in this window with no matching close in it, or the carried
state already says a hold was open coming in and this window's Evidence Log shows no CLOSE entry
for it. A standing note that only sounds like a hold -- a customer disputing a charge, a contact
asking to delay invoicing -- does NOT count, however it reads.
- "Days since completion, as of this run" on the page is the total elapsed days, already computed
for you; do not re-derive it from the calendar. The FIRST reconciliation window runs from
completion to the first run; every window after that is exactly 7 days. Days unbilled = days
since completion, MINUS the full length of every window (this one and any earlier one) that was
held for any part -- a window with the hold open even briefly does not count at all (Rule B-2).
Add that to what the carried state already established for earlier windows. Report REVIEW if the
total is 21 or more; otherwise UNBILLED. If the current window is held, report ON_HOLD instead
(days_unbilled may still be reported, reflecting only the non-held windows).
Answer with a single JSON object and nothing else:
{"status": "BILLED|UNBILLED|ON_HOLD|REVIEW|CONTEXT_INCOMPLETE",
"days_unbilled": <whole days counted as unbilled, hold-adjusted; null ONLY for CONTEXT_INCOMPLETE>,
"evidence_complete": "YES|NO",
"route_to": "NONE|ANALYST|SUPERVISOR",
"rationale": "one sentence, naming the rule you applied"}
Precedence for "status", applied in this order: CONTEXT_INCOMPLETE beats everything; then BILLED
if the carried state already says one, or if the register bills it this run; then ON_HOLD if a
formal hold is open; then REVIEW if the day count is 21 or more; otherwise UNBILLED.
"route_to": NONE for BILLED, ANALYST for CONTEXT_INCOMPLETE and UNBILLED, SUPERVISOR for ON_HOLD
and REVIEW (Rule B-8). This pack never creates, releases or posts an invoice -- "route_to" is a
recommendation for a person, not an action this reading takes (Rule B-9).
Carried state
----------------------------------------------------------------
No earlier reconciliation run has been recorded for this work order. This is its first appearance on the watch, so no invoice line has been matched yet and no earlier day count is available.
Unbilled-work reconciliation snapshot
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated unbilled-work reconciliation
snapshot for an AI use-case kit; it reproduces no real hauler, customer, contract or
billing register. The billing rule and its threshold are ILLUSTRATIVE.
Work Order
----------------------------------------------------------------
Customer : CUST-8021 (Grandview Retail Plaza)
Work order reference : WO-0001
Service type : COMPACTOR_HAUL
Container / equipment : CMP-6207
Route : RT-04
Completion date : 2026-02-21
Contract value for this job : 1,766.72 USD
Evidence required for billing : driver completion confirmation AND a weight ticket
Billing Rule
----------------------------------------------------------------
Supervisor review threshold : 21 days unbilled (operator-tunable; see .env
SUPERVISOR_REVIEW_DAYS)
Rule B-1 A completed work order becomes a BILLING CANDIDATE the day after its completion date,
and stays one until an invoice line in the register genuinely references it.
Rule B-2 Days unbilled are counted in whole RECONCILIATION WINDOWS. The first window runs from the
completion date to the first reconciliation run; every window after that is exactly the
7-day cadence. A window in which a formal billing hold (Rule B-3) was open for ANY part
does not count toward the total AT ALL; a window with no formal hold open at any point
counts in full. This is a stated simplification -- a hold that closes partway through a
window still excludes the whole window -- and it is not split into partial days.
Rule B-3 A billing hold suspends a window's count only where the Evidence Log records a DATED,
FORMAL hold entry — opened and, once resolved, closed — or where the carried state says a
formal hold was already open coming into this window. A standing note that SOUNDS like a
hold (a customer disputing a weight, an account contact asking to delay invoicing) does
NOT suspend anything unless a formal hold was actually opened for it.
Rule B-4 A work order is BILLED when the Invoice Register carries a line that genuinely references
it: by its exact work order number, by a BUNDLED line listing several work order numbers
for one route-week, or by a REFORMATTED job reference the billing system appends its own
suffix to. A line that merely QUOTES this work order's number while describing a
correction, a reversal or a charge taken back out does NOT bill it — quoting the number is
not billing the job, and Rule B-6 is what such a line is for.
Rule B-5 Evidence completeness depends on the service type. A WEIGHT-BASED service (compactor haul,
roll-off pull, disposal trip) needs a driver completion confirmation AND a weight ticket on
file. A FLAT-RATE service (container swap, roll-off delivery, extra pickup) needs only the
driver completion confirmation. Where the evidence this rule requires is not on file, the
work order is CONTEXT_INCOMPLETE: no status band is reported and it is routed to the
analyst to chase the evidence, never billed on a guess.
Rule B-6 A correction or reversal line in the register that names this work order while stating no
service was performed, or that the charge was miskeyed, does not change this work order's
status in either direction. It is evidence about the LEDGER, not about the JOB.
Rule B-7 Once a work order is reported BILLED, it stays BILLED on every later reconciliation run —
this pack does not re-open, re-flag or re-age a billed work order, even where a later
register shows that invoice voided or credited. Voiding is out of scope for this pack.
Rule B-8 Routing: a CONTEXT_INCOMPLETE work order routes to the ANALYST to gather evidence. A BILLED
work order routes to NONE — nothing to do. An ON_HOLD work order routes to the SUPERVISOR,
because a hold means a person is already deciding something. An UNBILLED work order younger
than SUPERVISOR_REVIEW_DAYS routes to the ANALYST; at or past that age it is a REVIEW and
routes to the SUPERVISOR.
Rule B-9 This pack never creates, releases or posts an invoice, and never marks a work order billed
on its own authority. "BILLED" is read off the register, never asserted; every reading's
output is a routing recommendation for a person to act on.
Billing Position
----------------------------------------------------------------
Reconciliation run : 2026-03-02 (scheduled run 1 of this work order)
Watch cadence : every 7 days (Mondays, 08:00 local)
Previous reconciliation run : -- this is the first scheduled run of this work order
Completion date : 2026-02-21
Days since completion, as of this run : 9
Evidence on file : driver completion confirmation ON FILE; weight ticket ON FILE
Invoice Register
----------------------------------------------------------------
Invoice INV-713778, line 2: WO-0001, compactor haul, weight-based, $1,766.72
Evidence Log
----------------------------------------------------------------
Entries raised in this window only.
2026-02-25 job re-assigned to a different analyst in the queue
Billing Desk Notes
----------------------------------------------------------------
Customer's site manager asked for an itemized breakdown before the invoice goes out.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"status": "BILLED", "days_unbilled": 9, "evidence_complete": "YES", "route_to": "NONE", "rationale": "Invoice register line INV-713778 line 2 genuinely references WO-0001 for the compactor haul, so Rule B-4 marks it BILLED with 9 days unbilled and no formal hold applies."}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch waste hauls that are still unbilled — 150 reconciliation snapshots. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the billing model -- the resolved status, the hold-adjusted day count, the evidence flag and the routing call -- and re-derived from src/billing.step by evals/check_labels.py before any run may spend. A null days_unbilled MATCHES a null one, because 'no count can be computed' is the correct answer on the 6 CONTEXT_INCOMPLETE readings rather than a missing value. No model grades anything, here or anywhere in this kit.
150reconciliation snapshots
150source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED150 · 123 · 144 · 137 · 103 / 150status accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 / 150days unbilled accuracy pct — readings, exact whole-day match, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 / 150evidence complete accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 / 150route to accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED35 · 17 · 35 / 35billed recall pct — readings truly BILLED, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 115false billed pct — readings NOT billed, including the reversal-decoy trap, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED9 · 6 · 9 / 9hold recall pct — readings truly ON_HOLD, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 4 / 141false hold pct — readings NOT on hold, including the hold-decoy trap, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED6 / 6context incomplete recall pct — readings with required evidence missing, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED21 · 0 / 21memory status accuracy pct — readings whose answer is NOT derivable from their own page, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED21 / 21memory days unbilled accuracy pct — readings whose answer is NOT derivable from their own page, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 · 0 / 1exposure usd error pct — portfolio total in USD reported as still open, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 · 150 / 150answered pct — calls, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is asserted: evals/check_labels.py replays all 50 work-order chains through src/billing.step and requires the committed gold to match on all four fields, so the key and the arithmetic cannot drift apart. It also proves the two traps are real rather than assumed: an exact work-order-number join is shown to miss every bundled and reformatted reference and to be fooled by both reversal lines, and the hold-decoy bucket is shown to carry no formal hold entry at all.
Grading costWhat it costs
Every dollar on this page is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One reconciliation snapshot
1,000 reconciliation snapshots
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.003902
$3.90
31%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.001561
$1.56
31%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.069118
$69.12
36%
Same work, 44× the bill
The same reconciliation snapshots, the same tokens — only the rate card changed. And across all 3 cards between 31% and 36% of what you pay is the prompt this pipeline sends, not the answer it writes.
Rates checked 2026-08-24. The provider that actually ran all 480 calls made while building this kit is kept out of these projection tables per this estate's naming rule for rendered pages, so nothing here is what was paid. The real spend for this kit's whole build is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors it is compared against or the pre-flight. The cost of the RUN is a different figure and lives in lens 07.
The gradersOne way to grade, and why it is the only one
⚑ THE MODEL WINS EVERY MEASURED CELL, AND THE FLOOR IS NOT A STRAW MAN. b002 (broad-match-mem) reaches 96.0 pct status accuracy for $0.00 by widening its matcher to bare numbers (catching the bundled and reformatted references b000/b001 miss) and its hold check to any hold-sounding word anywhere in the notes. Both widenings that make it strong are also what fool it: 2.84 pct false ON_HOLD (the hold-decoy notes it cannot tell from a formal entry) and it is fooled by both reversal-decoy lines the same way the two weaker floors are. ⚑ THREE FLOORS, BECAUSE ONE CANNOT SEPARATE THE MATCH FROM THE MEMORY FROM THE HOLD SCOPE. b000 is the incumbent -- an exact work-order-number join with no memory and no hold awareness: 68.67 pct, missing every bundled and reformatted match and every ON_HOLD reading. b001 adds carried memory and a hold check scoped correctly to the Evidence Log's own formal entries: 91.33 pct, +22.66 points, and already reaches 100 pct hold recall with ZERO false holds -- proving the hold-decoy trap is beaten by SCOPE, not by width. b002 then widens the matcher to catch bundled/reformatted (+33 points of billed recall, 68.57 -> 100.0) but widens the hold check too far and starts catching decoys. ⚠︎ READ THE PER-BUCKET CUT, NOT THE AGGREGATE: b002's 96.0 pct aggregate hides that it is wrong on 4 of 12 hold-decoy readings and both reversal-decoy readings -- exactly the two buckets this kit was built to test, and exactly where the model's 100 pct earns its keep.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
The status, day count, evidence flag and routing call, per reading, exact match against the computed answer key For each of the 150 readings and each of the four answered fields, did the reply equal the computed answer key? Status, evidence-complete and route_to are compared exactly against closed lists; days_unbilled is compared as a whole number, and a null MATCHES a null because 'no count can be computed' is the correct answer on 6 readings and must not be scored as a parse failure.
$0.00
no
yes
the fast tier, with the carried state 100.0% status accuracy · the fast tier, memory removed (THE CONTROL) 82.0% status accuracy · the strongest free floor, no model 96.0% status accuracy · 2 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set separates the four arms cleanly: 68.67 pct (no memory, no hold awareness) / 91.33 pct (memory, correctly-scoped hold) / 96.0 pct (memory, broad matching, over-broad hold) / 100.0 pct (the model). Only one grader is used (exact match against a computed key), so grader-versus-grader separability does not apply the way it would with an LLM-judge; what IS separable, and measured, is which capability each point of the gap buys.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
a billing desk with a clean, single-ID invoice register and no bundled or corrected lines
the exact-join-mem floor (b001)
91.33 pct for $0.00, and it already beats the model on nothing this scenario tests
spending on a model call for a matching problem that does not exist on this register
a billing desk whose register carries bundled route-week lines or reformatted job refs, and whose desk notes carry informal hold-sounding language
the model, with the carried state
the two traps this kit exists to test -- a bundled/reformatted reference and a hold-sounding note with no formal entry -- are exactly where every free floor in this run loses points, and the model lost none
the broad-match-mem floor alone: its own widening is what makes it wrong on hold-decoy and reversal-decoy rows
deciding whether ANY of this is worth a model call at all
run all three floors first (0 calls, $0.00) and look at the per-bucket cut
b001 already solves the hold-decoy trap by SCOPE alone; the model's margin over the best floor lives entirely in bundled/reformatted matching and the reversal decoy, which is a much narrower claim than 'the model beats free code'
quoting the 150-reading aggregate as the reason to spend -- the aggregate is exactly what a thin-denominator bucket hides
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
MEMORY_BLIND_BILLED_RECALL
Without the carried state, an already-billed work order is re-aged as if still open
21
WO-0001-R3 (stateless control, s001-unbilled-work-stateless): "Applying Rules B-4, B-2 and B-8: no register line bills WO-0001, no formal hold is open, and its 23 days unbilled meet the 21-day review threshold, so it routes to the supervisor." -- WO-0001 was…
MEMORY_BLIND_HOLD_HISTORY
Without the carried state, a hold that already closed loses its own history and the job ages from zero as if it had never been held
4
WO-0036-R2 (stateless control): "Flat-rate service evidence is complete (Rule B-5), the register has no billing line for WO-0036 (Rule B-4), no formal hold is open because only a close entry appears with no opening entry or carried hold state (Rule B-3), and…
What we could NOT verify
Whether the shipped 21-day SUPERVISOR_REVIEW_DAYS threshold matches any real billing desk's own escalation policy -- it is an invented default, not a confirmed one (Data.why_this_corpus).
Whether this pack's detection should share one join-and-hold-check core with a sibling billing-verification row in this catalogue that also examines invoice-line gaps, rather than the two independently risking disagreement about the same gap. Noted as an integration question for whoever owns both rows; not solved by this kit alone.
Whether unbilled-candidate detection should interface with revenue-recognition close timing or a GL posting calendar. This pack is deliberately, only operational -- its clock is the completion date and the reconciliation cadence, not any accounting period -- and that boundary is a stated design choice here, not a verified non-issue.
The true size of the memory-dependent category. The corpus generator's memory_dependent flag undercounts (see Data.breaks_on and the taxonomy above); a corrected flag would move memory_cells from 21 to at least 25, and this run does not re-derive the correct figure.
Whether disabling provider-side reasoning (the thinking setting, never sent on any published run here) changes any of these figures.
How this kit performs against a real billing register's actual reference formats, abbreviations and free-text conventions -- only five curated phrasings (exact, bundled, reformatted, reversal, none) are tested here.
What a missed reconciliation run costs. Unlike its closest cousins in this series, this kit ships no evals/cadence.py, so the dollar cost of a skipped weekly run is a real gap in this kit's own evidence rather than a measured figure.
Whether the second modelled arm needed to satisfy this standard's two-distinct-models rule should be a second REAL model rather than the free floor. Only one real provider was run; rules-baseline stands in as the second entry in scores, same convention as every monitor kit in this series.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
2,454.23
891.51
5,624 ms
$0.003902
$0.001561
$0.069118
the same tier, memory removed (THE CONTROL)
2,431.13
1,280.29
7,135 ms
$0.005056
$0.002023
$0.088326
the three free floors, no model
0
0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-24. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same reading (one work order, at one reconciliation run), the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
480 live calls were attempted building this kit; all 480 returned something. 309 are the committed measurement set (9 final calibration, 150 scored, 150 stateless control) at $1.3692 projected. 171 were DISCARDED, not because a call failed but because an exploratory r001 attempt (150 calls) surfaced a genuine prompt-clarity defect -- the carried-state sentence for a BILLED work order did not restate its frozen day count, so the model correctly answered null rather than guess -- which was fixed in src/state.py and the run re-fired clean; three earlier calibration probes (21 calls total) were superseded the same way while finding the right token ceiling. Nothing here was re-fired to chase a better SCORE; the defect that was fixed is documented in Business.not_good_enough and Architecture.components, and the discarded runs' own numbers are not quoted anywhere as findings. Grading itself costs $0.00 on every arm -- evals/scoring.py is pure code.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, priced at 6x input on the shared projection card. 68.5 pct of this run's projected bill is output, even though output tokens are only 27 pct of the total token count -- what you are paying for is the model reasoning about which of five register-line shapes applies, not reading the page.
THE CADENCE, which multiplies everything else. One reading is one call; the watch wakes weekly; the bill is rows x runs, not rows.
THE OPEN QUEUE, not the desk's total job volume. A work order that gets billed leaves the effective watch the run after it is matched (Rule B-7 freezes it, but this kit still re-reads and re-reports it every later run rather than dropping it from the corpus -- a real deployment would retire a billed job from the watch list entirely, which this kit's own corpus does not model).
Your volumeWhat it costs at your volume
LINEAR IN WORK ORDERS x RECONCILIATION RUNS, AND THAT IS THE WHOLE WARNING. Ten times the open work orders is ten times the calls at the same cadence -- there is no batching, no cache and no early exit, because every open candidate is re-read whole on every run by design. A billing desk holding 500 open work orders on this weekly watch is 500 calls a week, about $0.39 a week on the shared projection card. Nothing about the per-call price changes; the multiplier is the schedule and the queue size.
Where pricing changes shape
Your return, with your numbers
Volumecompleted work orders per reconciliation run -- this run judged 150 (50 work orders x 3 weekly runs) per arm, on a weekly watch
What it replacesa billing analyst manually reconciling completed jobs against the invoice register each week: checking whether each job's exact reference appears, opening any bundled or corrected line to work out what it covers, and reading exception notes to decide whether a hold is real
Time saved per itemnot measured here -- depends on how long an analyst takes to open and read one register line, which is the whole task this kit automates
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
2,454input tokens · this run
892output tokens
$0.004what it actually cost
per-query average, this run
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.234
$0.234
$1.56
2026-09-12
gemini-3-flash
Google
$0.585
$0.585
$3.90
2026-09-18
gemini-3-8-flash
Google
$0.778
$0.778
$5.18
2026-09-18
llama-5
Meta
$1.029
$1.029
$6.86
2026-09-18
claude-haiku-4-5
Anthropic
$1.037
$1.037
$6.91
2026-09-12
grok-4-5
xAI
$1.539
$1.539
$10.26
2026-09-18
grok-4-6
xAI
$1.539
$1.539
$10.26
2026-09-18
claude-sonnet-5
Anthropic
$2.074
$2.074
$13.82
2026-09-12
gemini-3-1-pro
Google
$2.341
$2.341
$15.61
2026-09-18
gpt-5-6-terra
OpenAI
$2.341
$2.341
$15.61
2026-09-12
gpt-5-6-sol
OpenAI
$4.147
$4.147
$27.65
2026-09-12
claude-opus-4-8
Anthropic
$5.184
$5.184
$34.56
2026-09-12
claude-opus-5
Anthropic
$5.184
$5.184
$34.56
2026-09-12
claude-fable-5
Anthropic
$10.368
$10.368
$69.12
2026-09-18
claude-fable-5-1
Anthropic
$10.368
$10.368
$69.12
2026-09-18
gpt-6-astra
OpenAI
$10.368
$10.368
$69.12
2026-09-17
Read this against the numbers above
Projection only -- no other model was actually called against this corpus.
The reasoning-token share for the fast tier is not separately broken out here; a different model's reasoning behaviour is unmeasured and could move these projections substantially.
Accuracy is NOT projected, only cost -- a cheaper or pricier model is not implied to score the same 100.0 pct status accuracy.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
12 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Six of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 50 completed work orders x 3 weekly reconciliation runs = 150 snapshots from a fixed seed (SEED = 20260824), across ten customers and six service types. Plants ten explicit buckets: three ways an invoice line genuinely bills a job (exact, bundled, reformatted), one way a line only quotes a job's number without billing it (the reversal decoy), two hold shapes (closes inside the next window, spans a whole window silently), a hold-sounding note with no formal hold behind it, an evidence gap that resolves, and two plain aging trajectories. The gold labels are src/billing.step's output over the raw per-run facts, never typed by hand.
You change it to: The customers, service types, invoice-line phrasings and the seed. Keep the eight section headings or src/segment.py's assertion refuses to start.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260824
COUNT = 50
RULE = "-" * 64
RUN_DATES = (dt.date(2026, 3, 2), dt.date(2026, 3, 9), dt.date(2026, 3, 16))
CUSTOMERS = [
src/billing.pythe billing rule — a swap seam
The rule as pure code: the whole-window hold exclusion (Rule B-2/B-3), the invoice-match test (Rule B-4/B-6), the evidence-completeness test by service type (Rule B-5), and the routing table (Rule B-8). No model, no judgement. The nine rules, the 21-day review threshold and the whole-window convention are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
You change it to: The whole-window exclusion convention (Rule B-2) and what counts as a FORMAL hold (Rule B-3). A real billing desk's own hold workflow -- who can open one, what evidence it requires -- should replace the invented rule.
src/billing.py
# The unbilled-work rule as arithmetic. Pure code, no model, standard library only.
BILLED = "BILLED"
UNBILLED = "UNBILLED"
ON_HOLD = "ON_HOLD"
REVIEW = "REVIEW"
CONTEXT_INCOMPLETE = "CONTEXT_INCOMPLETE"
STATUSES = (BILLED, UNBILLED, ON_HOLD, REVIEW, CONTEXT_INCOMPLETE)
YES = "YES"
NO = "NO"
EVIDENCE = (YES, NO)
src/state.pythe carried state — a swap seam
SEAM 2 -- the thing that makes this a monitor. Four scalars per work order (billed, hold_open, the frozen or running day count, the status last reported), written from the arithmetic and rendered into one English sentence for the prompt. ⚠︎ A real defect was found and fixed here during this kit's own measurement: the first version described a BILLED work order without restating its frozen day count, and the model correctly answered null rather than guess it -- 21 of 21 memory-dependent BILLED cells wrong. Naming the number in the sentence took memory-dependent day-count accuracy from 14.29 pct to 100 pct; see the module's own docstring for the full account.
You change it to: load()/save() are a JSON file today. Point them at a table or a billing data mart and nothing else in the kit changes -- for_workorder() and advance-by-step() are the whole interface, and describe() is the only thing the prompt sees.
src/state.py
# The carried state -- the thing that makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_workorder(store, workorder_id):
def describe(state):
src/segment.pythe splitter
Splits a reconciliation snapshot into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 150 documents before a run may spend.
src/segment.py
# Split a reconciliation snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Work Order", "Billing Rule", "Billing Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe selector
Decides which sections reach the model. Site Contact -- the customer site's name, direct dial and work email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/select.py
# Pick which sections of a reconciliation snapshot are sent. Pure code -- the last deterministic
BANNER = "Synthetic Record"
WORKORDER = "Work Order"
RULE = "Billing Rule"
POSITION = "Billing Position"
REGISTER = "Invoice Register"
EVIDENCE = "Evidence Log"
CONTACT = "Site Contact"
NOTES = "Billing Desk Notes"
NEVER_SENT = (CONTACT,)
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the carried-state sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string
INSTRUCTION = """\
def build(text, carried, stateless=False):
SYSTEM = ("You are a scheduled unbilled-work reconciliation watch. You apply a written billing "
def build_parts(text, carried, stateless=False):
src/adapters/__init__.pythe adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading.
You change it to: PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/monitor.pythe watch
One work order, one reconciliation run, one call. Parses identifying facts off the page with a regex (the model is never asked to compute a calendar difference), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling. Never creates, releases or posts an invoice -- see the module's own guardrail note.
src/monitor.py
# One work order, one reconciliation run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = P.SYSTEM
MAX_TOKENS = 24000
FIELDS = ("status", "days_unbilled", "evidence_complete", "route_to")
def documents():
def chains():
def load_doc(doc_id):
def position_of(text):
src/app.pythe local UI
One work order, one reconciliation run, its carried state and its verdict, on 127.0.0.1:8208. Renders with no key. Shows the carried sentence verbatim, the free floor's answer beside the model's, and a second button that replays what run r001-unbilled-work actually answered, straight off the committed result file, labelled as a replay.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "8208"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-unbilled-work")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe free floors
exact-join: an exact work-order-number string join, no memory, no hold awareness -- the incumbent. exact-join-mem: the same join plus carried memory and a hold check scoped to the Evidence Log's formal entries only. broad-match-mem: carried memory, a bare-number search that catches bundled and reformatted references, and a hold check widened to any hold-sounding word in the notes. 0 calls, $0.00, all three scored through the identical scorer.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("exact-join", "exact-join-mem", "broad-match-mem")
REVIEW_DAYS = B.SUPERVISOR_REVIEW_DAYS
HOLD_WORDS = ("hold the invoice", "disputes the", "asked us to hold", "budget resets",
SECTION_NAMES = ("Synthetic Record", "Work Order", "Billing Rule", "Billing Position",
def _num(pat, text, cast=int):
def _section(text, name):
def review(text, carried=None, mode="broad-match-mem"):
evals/scoring.pythe scorer
Exact match per cell against the computed gold, split ways an average would hide: the four fields, BILLED recall against a false-BILLED rate, ON_HOLD recall against a false-ON_HOLD rate, under- versus over-escalation counted apart, the memory-dependent subset, and every bucket that plants a trap, cut separately. No judge model.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("status", "days_unbilled", "evidence_complete", "route_to")
BUCKET_CUTS = ("billed_exact", "billed_bundled", "billed_reformatted", "unbilled_watch",
def _pct(n, d):
def _days(v):
def score(records, golds):
evals/check_labels.pythe pre-flight
Eighteen things that must be true before a run may spend: all eight sections parse, every work order's run sequence is complete and gap-free, no run-varying section restates an earlier match, neither a formal-hold flag nor a final day count is printed anywhere on the page, an exact join is proven to miss every bundled and reformatted reference AND be fooled by both reversal lines, the hold-decoy bucket is proven to carry no formal hold entry, the privacy guard holds and would fail without it, the two prompt builders differ on exactly one line, no code path creates, releases or posts an invoice, and the answer key replays from src/billing.step.
evals/check_labels.py
# Everything that must be true BEFORE a run is allowed to spend. Free, and it runs in seconds.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAILED = []
def check(name, ok, detail=""):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 50 completed work orders x 3 weekly reconciliation runs = 150 snapshots from a fixed seed (SEED = 20260824), across ten customers and six service types. Plants ten explicit buckets: three ways an invoice line genuinely bills a job (exact, bundled, reformatted), one way a line only quotes a job's number without billing it (the reversal decoy), two hold shapes (closes inside the next window, spans a whole window silently), a hold-sounding note with no formal hold behind it, an evidence gap that resolves, and two plain aging trajectories. The gold labels are src/billing.step's output over the raw per-run facts, never typed by hand. A swap seam.
src/billing.pyThe rule as pure code: the whole-window hold exclusion (Rule B-2/B-3), the invoice-match test (Rule B-4/B-6), the evidence-completeness test by service type (Rule B-5), and the routing table (Rule B-8). No model, no judgement. The nine rules, the 21-day review threshold and the whole-window convention are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Four scalars per work order (billed, hold_open, the frozen or running day count, the status last reported), written from the arithmetic and rendered into one English sentence for the prompt. ⚠︎ A real defect was found and fixed here during this kit's own measurement: the first version described a BILLED work order without restating its frozen day count, and the model correctly answered null rather than guess it -- 21 of 21 memory-dependent BILLED cells wrong. Naming the number in the sentence took memory-dependent day-count accuracy from 14.29 pct to 100 pct; see the module's own docstring for the full account. A swap seam.
src/segment.pySplits a reconciliation snapshot into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 150 documents before a run may spend.
src/select.pyDecides which sections reach the model. Site Contact -- the customer site's name, direct dial and work email -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/prompt.pyThree parts: the fixed instruction, the carried-state sentence, and the selected sections in document order. The stateless build replaces the carried-state sentence with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost reading. A swap seam.
src/monitor.pyOne work order, one reconciliation run, one call. Parses identifying facts off the page with a regex (the model is never asked to compute a calendar difference), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 24000, the published ceiling. Never creates, releases or posts an invoice -- see the module's own guardrail note.
evals/baseline.pyexact-join: an exact work-order-number string join, no memory, no hold awareness -- the incumbent. exact-join-mem: the same join plus carried memory and a hold check scoped to the Evidence Log's formal entries only. broad-match-mem: carried memory, a bare-number search that catches bundled and reformatted references, and a hold check widened to any hold-sounding word in the notes. 0 calls, $0.00, all three scored through the identical scorer.
evals/scoring.pyExact match per cell against the computed gold, split ways an average would hide: the four fields, BILLED recall against a false-BILLED rate, ON_HOLD recall against a false-ON_HOLD rate, under- versus over-escalation counted apart, the memory-dependent subset, and every bucket that plants a trap, cut separately. No judge model.
evals/check_labels.pyEighteen things that must be true before a run may spend: all eight sections parse, every work order's run sequence is complete and gap-free, no run-varying section restates an earlier match, neither a formal-hold flag nor a final day count is printed anywhere on the page, an exact join is proven to miss every bundled and reformatted reference AND be fooled by both reversal lines, the hold-decoy bucket is proven to carry no formal hold entry, the privacy guard holds and would fail without it, the two prompt builders differ on exactly one line, no code path creates, releases or posts an invoice, and the answer key replays from src/billing.step.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2454 input and 891 output tokens per reading (one work order, at one reconciliation run), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one work order, at one reconciliation run)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one work order, at one reconciliation run) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No injection or red-team run was fired against this kit. Billing Desk Notes is this kit's one injection surface -- free text a billing analyst types, sent to the model verbatim -- and one of the shipped notes is instruction-shaped by design ('Customer's accounts-payable contact asked us to hold off invoicing until next quarter's budget resets.'), matching the convention sibling kits in this series use, but its effect on a live model's status or route_to call was never measured.
API_KEY is read from .env (gitignored) or the real environment; never logged, never placed in a prompt, and redacted from any provider error message src/app.py passes to the browser (exact-string replace on api_key and base_url before the message is returned).
The experimentWe did NOT attack it -- and this is the surface we left open
An indirect prompt injection needs a field somebody outside the process can write into, and this kit's is the Billing Desk Notes -- free text a billing analyst types, sent to the model verbatim. One of the shipped notes is instruction-shaped by design, but no adversarial trial was fired against it. Both were run for real on 2026-08-24.
Boundary checked
What could go wrong
What the code guarantees
Whether an instruction embedded in a corpus field can move this kit's own routing or status decision
No adversarial trial was fired against this kit's corpus.
evals/check_labels.py's banned-code-path scan (0 violations) -- the only automated gate this kit runs before spending.
No boundary here is measured in both directions. The one gate below reports the banned-code-path scan, which is a static check on what the code CAN do, not a measurement of what the model did under attack.
The resultNot measured. No adversarial trial was run against the Billing Desk Notes injection surface, including the instruction-shaped note already shipped in the corpus.
0attack trials fired
No red-team run was attempted for this kit. evals/check_labels.py's banned-code-path scan (0 violations) -- the only automated gate this kit runs before spending. Cited for reference only, same corpus and model as r001-unbilled-work.
Read this twice
The Billing Desk Notes reach the model verbatim — there is no filter between a desk note and the prompt, and one shipped note is instruction-shaped ('asked us to hold off invoicing until next quarter's budget resets'). This pack never creates, releases, posts or voids an invoice, so the worst a followed instruction can do here is route a job to the wrong queue.
HonestyWhat this does not prove
Whether the instruction-shaped note ('asked us to hold off invoicing until next quarter's budget resets') actually suppresses or changes route_to on a live model -- named as a surface, never probed.
Whether the model would follow an injected instruction to falsely report BILLED, or a different status, on a job that is neither -- never tested.
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
Never create, release, post or void an invoice under any authority, and never present the 21-day threshold or hold convention as any real hauler's actual policy.
Stated to the model on every call in src/prompt.py's INSTRUCTION block (Rule B-9), and enforced mechanically by evals/check_labels.py's banned-code-path scan, which greps every .py/.js file in the kit for the names of a release/post/create-invoice path before any run may spend.
EvidenceDoes it hold?
What
Measured
The banned-code-path scan
Measured at 0 banned code paths across the whole kit, on every run of check_labels.py, including the one immediately before r001-unbilled-work, s001-unbilled-work-stateless and every calibration run spent.
The limitWhat a guardrail is not
This is a PROMPT RULE PLUS A STATIC-ANALYSIS SCAN, not a runtime enforcement layer. Nothing in the code stops a forker from adding a release_invoice() function tomorrow; the scan only catches it the next time someone runs evals/check_labels.py, a manual step, not a hook that fires on every commit.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 39 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
20 measured by the latest run19 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The status, day count, evidence flag and routing call, per reading, exact match against the computed answer key
alarm
status_accuracy_pct; billed_recall_pct; false_hold_pct; days_unbilled_accuracy_pct — alarm on any reply that does not parse. The scored run had 0 of 150; the calibration run at an under-sized 8,000-token ceiling had 1 of 9, which is why the published ceiling is 24,000.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
891,503
reconciliation snapshots edited — the count held, the bytes did not
split.count
50
the work orders count moved — a different set was scored
split.size_p50
3
the median size of one work order moved
split.size_p95
3
the 95th-percentile size of one work order moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (analyst_enough_cells 120, billed_cells 35, cadence_days 7, context_incomplete_cells 6, documents 150, exposure_usd_gold 83399.01, hold_cells 9, memory_cells 21, needs_supervisor_cells 30, not_billed_cells 115, not_hold_cells 141, readings_scored 150, stateless False, supervisor_review_days 21, workorders 50) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Status
100.00 pct
150 readings
r001-unbilled-work exact match against computed gold status
Days unbilled
100.00 pct
150 readings
r001-unbilled-work exact match against computed gold days_unbilled
Route-to routing
100.00 pct
150 readings
r001-unbilled-work exact match against computed gold route_to
Answered
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Billed recall
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Context incomplete recall
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Evidence complete accuracy
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Exposure, USD error
0.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Exposure, USD reported
$83,399.01 on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
False billed
0.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
False hold
0.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Hold recall
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Input tokens, whole run
368,135 on r001-unbilled-work
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Latency, median ms
5,624 ms on r001-unbilled-work
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Latency, 95th percentile ms
21,033 ms on r001-unbilled-work
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Memory-dependent days unbilled accuracy
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Memory-dependent status accuracy
100.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Output tokens, whole run
133,726 on r001-unbilled-work
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Over escalation
0.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
Under escalation
0.00 pct on r001-unbilled-work
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-unbilled-work's own, re-derived from its result file by build/measured/runlog.py.
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-unbilled-work-exactjoin 2026-08-24
b001-unbilled-work-exactjoinmem 2026-08-24
b002-unbilled-work-broadmatch 2026-08-24
answered, %
100.0
100.0
100.0
billed recall, %
17.14
68.57
100.00
context incomplete recall, %
100.0
100.0
100.0
days unbilled accuracy, %
76.00
98.00
95.33
evidence complete accuracy, %
100.0
100.0
100.0
exposure usd error, %
18.14
4.68
3.59
exposure usd reported
98531.37
87302.76
80408.13
false billed, %
1.74
1.74
1.74
false hold, %
0.00
0.00
2.84
hold recall, %
0.0
100.0
100.0
input tokens, whole run
0
0
0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
memory days unbilled accuracy, %
0.0
100.0
100.0
memory status accuracy, %
0.0
100.0
100.0
output tokens, whole run
0
0
0
over escalation, %
12.50
1.67
3.33
route to accuracy, %
68.67
91.33
96.00
status accuracy, %
68.67
91.33
96.00
under escalation, %
33.33
3.33
3.33
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 3 runs. Columns here are only ever compared with each other.
Metric
c000-unbilled-work-calibration 2026-08-24
r001-unbilled-work 2026-08-24
s001-unbilled-work-stateless 2026-08-24
answered, %
100.0
100.0
100.0
billed recall, %
100.00
100.00
48.57
context incomplete recall, %
—
100.0
100.0
days unbilled accuracy, %
100.0
100.0
80.0
evidence complete accuracy, %
100.0
100.0
100.0
exposure usd error, %
0.00
0.00
13.46
exposure usd reported
2406.90
83399.01
94627.62
false billed, %
0.0
0.0
0.0
false hold, %
0.0
0.0
0.0
hold recall, %
—
100.00
66.67
input tokens, whole run
22143
368135
364669
model latency p50 ms
3708.00
5624.00
7135.00
model latency p95 ms
8180.00
21033.00
23850.00
memory days unbilled accuracy, %
100.0
100.0
0.0
memory status accuracy, %
100.0
100.0
0.0
output tokens, whole run
4822
133726
192043
over escalation, %
0.0
0.0
10.0
route to accuracy, %
100.0
100.0
82.0
status accuracy, %
100.0
100.0
82.0
under escalation, %
—
0.0
10.0
not a time series No two of these 3 runs measured the same system — they differ on analyst_enough_cells, billed_cells, context_incomplete_cells, documents, exposure_usd_gold, hold_cells, memory_cells, needs_supervisor_cells, not_billed_cells, not_hold_cells, readings_scored, stateless, workorders — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-unbilled-work-stub 2026-08-24
answered, %
100.0
billed recall, %
17.14
context incomplete recall, %
100.0
days unbilled accuracy, %
76.0
evidence complete accuracy, %
100.0
exposure usd error, %
18.14
exposure usd reported
98531.37
false billed, %
1.74
false hold, %
0.0
hold recall, %
0.0
input tokens, whole run
375259
model latency p50 ms
0.00
model latency p95 ms
0.00
memory days unbilled accuracy, %
0.0
memory status accuracy, %
0.0
output tokens, whole run
6417
over escalation, %
12.5
route to accuracy, %
68.67
status accuracy, %
68.67
under escalation, %
33.33
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 20 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
the tier/threshold or routing constants in src/*.py
If SUPERVISOR_REVIEW_DAYS or the whole-window hold convention in src/billing.py change, every published accuracy figure and the corpus's own gold labels need regenerating together -- src/billing.step is the single place both the corpus builder and the scorer call, so a threshold edit ripples to data/gold.jsonl, every results/*.json and this spec at once, never silently in only one.
reasoning
not independently re-measured -- the module is read, not re-run, to reach this claim
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
Answered
nothing yet — a second scored run is what would give this column a spread to fire on.
Billed recall
nothing yet — a second scored run is what would give this column a spread to fire on.
Context incomplete recall
nothing yet — a second scored run is what would give this column a spread to fire on.
Evidence complete accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Exposure, USD error
nothing yet — a second scored run is what would give this column a spread to fire on.
Exposure, USD reported
nothing yet — a second scored run is what would give this column a spread to fire on.
False billed
nothing yet — a second scored run is what would give this column a spread to fire on.
False hold
nothing yet — a second scored run is what would give this column a spread to fire on.
Hold recall
nothing yet — a second scored run is what would give this column a spread to fire on.
Memory-dependent days unbilled accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Memory-dependent status accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Over escalation
nothing yet — a second scored run is what would give this column a spread to fire on.
Under escalation
nothing yet — a second scored run is what would give this column a spread to fire on.
NextThe three you would add first
Automate the manual scan stepA CI hook (or pre-commit) that runs evals/check_labels.py's banned-path scan automatically, since today it is a manual step a developer has to remember to run.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Checked on every run via evals/check_labels.py's banned-path scan, run manually immediately before any paid call -- not on a schedule and not on a git hook.
What this cannot tell you
Whether a forker who added their own invoice-release function under a different name elsewhere in src/ would be caught -- the scan matches known function-name patterns, not behaviour.
Whether the illustrative 21-day threshold and whole-window hold convention would need retuning against a real hauler's actual billing policy once waste:BCR-UNB's operator supplies one -- genuinely unresolved.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
NO FRAMEWORK. The same position as every kit in this series: a folder of readable Python, no framework dependency, and a prompt anyone can read end to end in src/prompt.py. LangChain/LlamaIndex-style abstractions would own a retrieval step -- there is none here, the whole snapshot goes into one prompt -- and the memory/checkpoint layer, which is already the entire surface of src/state.py's four scalars.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the model call
src/adapters/__init__.py
an LLM wrapper / Runnable
a wrapper buys a swappable provider interface; src/adapters/__init__.py is one dict, one function, and the seam this kit measures is what the model returns, not how it is called.
the carried state
src/state.py
a memory or checkpoint object
a checkpoint object buys persistence and concurrency; this kit carries four-to-a-few scalars written by arithmetic, never by the model, so a framework would add machinery around a fact that fits in one line.
the corpus
tools/build_corpus.py
a document loader
a document loader buys format handling across many source types; this kit reads one flat synthetic format it fully controls, so the loader would abstract nothing that varies.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Linear: src/monitor.py -> src/prompt.py -> src/adapters -> evals/scoring.py. No branching, no tool-calling and no agent loop -- a framework's graph/DAG abstraction has nothing to route here.
The other sideWhat a framework costs you
The cost of not using a framework is that swapping providers means editing src/adapters/__init__.py's PROVIDERS dict by hand (one function, one dict entry) rather than swapping a framework's provider string. In exchange the prompt has no hidden templating layer and requirements.txt stays empty.
What we could NOT verify
Whether a framework's built-in memory abstraction would have caught the billed/hold-open distinctions src/state.py encodes by hand -- no port to a framework was built to compare.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-unbilled-work on the same tier, memory removed (THE CONTROL), 2026-08-24. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
5,624 ms
5,624 ms on r001-unbilled-work
—
Model, p95
21,033 ms
21,033 ms on r001-unbilled-work
—
Input tokens
368,135
368,135 on r001-unbilled-work
—
Output tokens
133,726
133,726 on r001-unbilled-work
—
No movement column. Not one of the 5 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-unbilled-work-calibration3,708 ms
r001-unbilled-work5,624 ms
s001-unbilled-work-stateless7,135 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-unbilled-work-exactjoin, b001-unbilled-work-exactjoinmem, b002-unbilled-work-broadmatch recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-24, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
reconciliation snapshots
data/corpus/WO-<n>-R<k>.txt -- 150 files, 891503 bytes, generated once from a fixed seed by tools/build_corpus.py
seven of the eight sections go to the provider in the prompt; Site Contact -- the customer site's name, direct dial and work email -- never does, by src/select.NEVER_SENT
the carried state
data/state.json in a deployment, and in-memory per run in the eval -- four scalars per work order, written by src/billing.step and never by a model
as one English sentence in every prompt, and the UI prints the same sentence verbatim so a reader can audit what the model was told
the answer key
data/gold.jsonl -- 150 rows, the output of src/billing.step over the planted per-run facts, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the run records
results/eval-*.json in the kit
never -- they are read by the site build, not by a provider
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
API_KEY is read from .env (gitignored) or the real environment; never logged, never placed in a prompt, and redacted from any provider error message src/app.py passes to the browser (exact-string replace on api_key and base_url before the message is returned).
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a SEVEN-DAY watch: every Monday at 08:00 local. One run owns exactly the window since the last one, and the reconciliation-window day count is DEFINED as this cadence (Rule B-2), so changing the schedule changes what the answers mean as well as what they cost.
50 work orders x 3 scheduled runs = 150 calls per arm, 0 gaps and 0 out-of-order readings; evals/check_labels.py asserts every sequence is complete before a run may spend. A desk of 500 open work orders on this cadence is 500 calls a week, about $0.39 on the shared projection card. (r001-unbilled-work, evals/check_labels.py, src/billing.CADENCE_DAYS)
⚠︎ UNLIKE ITS CLOSEST COUSINS IN THIS SERIES, THIS KIT DOES NOT MEASURE WHAT A MISSED RUN COSTS. No evals/cadence.py exists here to sweep other cadences or price a skipped reconciliation run in dollars -- a real gap in this kit's own evidence, stated rather than quietly left off (see Architecture.breaks_at_scale).
every accuracy figure on this page is per a seven-day watch, and the whole-window hold exclusion is DEFINED as this cadence. A daily or monthly watch is not a comparable run of the same question, and no arm has been run at one.
state
four scalar fields per work order -- whether an invoice line is already matched, whether a formal hold is open, the frozen or running day count, and the status last reported -- written by src/billing.step from figures PARSED off the page and never from the model's reply, and rendered by src/state.describe into one English sentence. That sentence is the entire route from one reconciliation run to the next.
190 characters on the worked first-run example, growing to a full sentence naming a frozen day count once a work order is billed. Status accuracy with it: 100.0 pct. Status accuracy without it (the stateless control): 82.0 pct, and 0.0 pct on the 21 cells flagged memory-dependent, by construction. (r001-unbilled-work against s001-unbilled-work-stateless, lenses.LLM.prompt_parts)
four scalars, so run 40 costs what run 2 costs. What is NOT bounded is accuracy over a longer history, which is unmeasured (only 3 runs per work order), and the store itself: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model. ⚠︎ A REAL DEFECT WAS FOUND AND FIXED HERE DURING THIS KIT'S OWN MEASUREMENT: the first version of describe() did not restate a BILLED work order's frozen day count, and the model correctly answered null 21 times rather than guess it. Naming the number took memory-dependent day-count accuracy from 14.29 pct to 100 pct.
lose the carried state and BILLED recall collapses to 48.57 pct and ON_HOLD recall to 66.67 pct, while every reply stays well-formed and nothing raises an error -- see environment.signatures' traceless row.
model
one completion call per reading, on one provider and one key, through src/adapters/__init__.py. MAX_TOKENS = 24000; thinking is never sent, so every published run left provider-side reasoning at the default and the result files record thinking: null.
150 calls on the scored arm, largest reply 4308 output tokens against the 24000 ceiling. p50 5.6s, p95 21.0s. 0 unparsed. (r001-unbilled-work, c000-unbilled-work-calibration)
⚠︎ THE FIRST CALIBRATION PROBE UNDER-ESTIMATED THIS: an 8,000-token cap on the three hardest work orders lost one reading to the ceiling (WO-0007-R2, a bundled-match case). Re-run at the full 24,000-token ceiling it parsed at 8,330 output tokens. The published ceiling is set from that measurement, not from the smaller probe, and a ceiling is not a cost -- the provider bills tokens produced, not tokens allowed.
a provider whose reasoning default differs reprices this kit without changing anything a reader can see. Only ONE real model was run; rules-baseline stands in as the second scores entry, and no cross-tier claim is made anywhere on this page.
labels
150 readings -- 50 completed work orders re-read at 3 weekly reconciliation runs -- with the four answers computed by tools/build_corpus.py from the billing model and re-derived from src/billing.step by evals/check_labels.py before any run may spend. 21 are memory-dependent by the corpus generator's own flag (at least 25 by a corrected count, see Data.breaks_on); 6 have required evidence missing and must not be billed on a guess; 2 are reversal-decoy lines and 4 work orders (12 readings) carry a hold-sounding note with no formal hold behind it.
150 rows in data/gold.jsonl, 0 replay mismatches, 0 sequence gaps. Status spread: BILLED 35, UNBILLED 79, REVIEW 21, ON_HOLD 9, CONTEXT_INCOMPLETE 6. evals/check_labels.py asserts an exact work-order-number join misses every bundled and reformatted reference AND is fooled by both reversal lines, and that the hold-decoy bucket carries no formal hold entry at all -- without those two proofs the corpus's two central traps would be assumed rather than measured. (data/gold.jsonl, evals/check_labels.py, data/corpus-stats.json)
it stops scoring at the third weekly reconciliation run of each work order. Long enough for a match or a hold opened in run 1 to have to be remembered in run 3, and not long enough to test a work order that has sat unbilled for a quarter, or a hold that opens, closes, and opens again.
the labels encode ONE author's five register-line phrasings and one whole-window hold convention. A real billing register's actual reference formats are far messier, and a phrasing this corpus does not know about would reproduce the same defect this kit was built to catch, silently.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a work order reported UNBILLED or REVIEW that was in fact already billed on an earlier run
the reading was made without the carried state. Nothing on a later snapshot restates an earlier invoice match, so a reader who cannot see the previous run re-ages a job that revenue has already been collected on
check what the carried state said before chasing this job. The stateless control makes this exact confusion 21 times in 150 readings; the stateful arm makes it 0 (results/eval-s001-unbilled-work-stateless.json against results/eval-r001-unbilled-work.json)
a work order marked BILLED whose invoice register line explicitly describes a correction, reversal or miskeyed charge
something matched on the work order's NUMBER alone without reading what the line SAYS. A correction that quotes a job's reference while taking the charge back out is not a bill, and Rule B-6 exists because a bare-number search cannot tell the two apart
read the Invoice Register line in full, not just whether the work order's number appears in it. The broadest free floor in this kit gets this wrong on both reversal-decoy readings; the model gets it wrong on none (results/eval-b002-unbilled-work-broadmatch.json against results/eval-r001-unbilled-work.json)
a work order marked ON_HOLD with no dated 'opened' entry anywhere in its Evidence Log and no carried hold history
something treated a Billing Desk Note that merely sounds like a hold -- a customer dispute, a request to delay invoicing -- as a formal suspension. Rule B-3 is exactly the guardrail this is supposed to trip
look for a dated 'Billing hold ... opened' entry. If there is none, and the carried state does not already say a hold was open, ON_HOLD is wrong regardless of what the desk notes say (results/eval-b002-unbilled-work-broadmatch.json, the hold_decoy bucket)
nothing at all -- the worklist simply does not change between two weekly runs
TRACELESS. A scheduled run that never fired leaves no artefact anywhere in this kit: no error, no gap marker, no late flag. The readings it would have produced simply do not exist
there is nothing on the worklist to look at, so look at the SCHEDULER instead: compare the number of readings this kit produced against the number the cadence says it should have (open work orders x runs). On this corpus that is 50 x 3 = 150 and evals/check_labels.py asserts it; in a deployment nothing does. Unlike its closest cousins in this series, this kit has not priced what a missed run costs in dollars -- see Architecture.breaks_at_scale. (evals/check_labels.py's sequence-completeness assertion)
['Concurrency. data/state.json is replaced atomically, which is correct for one writer; two watches on one queue is two writers and nothing here has tested it. A lost write here forgets a match or a hold, so the next run re-ages or re-flags a job that was already resolved.', 'What a missed reconciliation run costs, in dollars. Unlike deduction-age and settle-fail, this kit ships no evals/cadence.py -- a stated gap, not a zero.', 'A cadence other than seven days AS A SCORING QUESTION. No arm has been scored at any other interval, and the whole-window hold convention is defined as this cadence.', "A population that changes between runs -- a work order completed mid-window, or one whose hold opens, closes, and opens again. This kit's vocabulary carries only one hold interval per work order, ever.", 'Provider-side retention. The prompt carries no personal data by construction, but what the provider keeps of a request is outside this repository and nobody here has verified it.', 'GPU sizing, local inference and anything about running this off a hosted API. Not attempted; not costed.', 'A second REAL model. One tier was run; rules-baseline stands in as the second scores entry.', "Whether the Billing Desk Notes field can move route_to or status. The surface is sent deliberately and one shipped note is worded like an instruction ('hold off invoicing until next quarter's budget resets'), but no adversarial run against it has been fired -- see Eval.redteam_page.", "The true size of the memory-dependent category. The corpus generator's own flag undercounts it (see Data.breaks_on); this run does not re-derive the corrected figure.", "Whether this pack's detection should share a core with a sibling billing-verification row in this catalogue. Not resolved here -- see Eval.could_not_verify."]
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every customer, work order, container, invoice line, evidence-log entry and desk note is invented here. Verified against the repository's own LICENSE file on 2026-08-24. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The status, day count, evidence flag and routing call, per reading, exact match against the computed answer key
Catch waste hauls that are still unbilled
PresenterOpens the private repo. Visible to admins only.
In one lineThe status, day count, evidence flag and routing call, per reading, exact match against the computed answer key
For each of the 150 readings and each of the four answered fields, did the reply equal the computed answer key? Status, evidence-complete and route_to are compared exactly against closed lists; days_unbilled is compared as a whole number, and a null MATCHES a null because 'no count can be computed' is the correct answer on 6 readings and must not be scored as a parse failure.
$0.00per 1,000 reconciliation snapshots
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py is scored through.
The inputOne real row, seen by every grader
The reading
WO-0049-R2 -- a reversal decoy, the register quoting the job's own number while reversing the charge
What the Invoice Register carries
"Invoice INV-727042, line 5: billing correction -- reversing prior miskeyed charge originally coded to WO-0049; no service performed, $0.00" -- the work order's exact reference appears in the line, and the line explicitly says no service was performed.
The answer key
REVIEW, evidence complete, no formal hold, 21 days unbilled, route SUPERVISOR (Rule B-6: a reversal that quotes the number does not bill the job)
What run r001-unbilled-work answered
REVIEW, evidence complete, no formal hold, 21 days unbilled, route SUPERVISOR -- rationale: "Evidence is complete (Rule B-5), the register line merely reverses a miskeyed charge so it does not bill (Rule B-6), no formal hold is open, and the hold-adjusted unbilled total is 21 days, reaching the supervisor review threshold (Rule B-2/B-8)."
The strongest free floor
BILLED, route NONE -- its bare-number search finds "0049" inside the line and stops reading
Why this one
It is the shape where a wrong match does not merely change a status label -- it tells a billing desk a job needs no further attention when $1,901.60 of contract value is genuinely still unbilled and now aged past the review threshold, and it is the one row where the free floor's own widening (the thing that makes it beat the two weaker floors) is exactly what defeats it.
Grader
Verdict
Why
The status, day count, evidence flag and routing call, per reading, exact match against the computed answer key
status hit, days hit, evidence hit, route hit -- four of four
The model read the reversal language and correctly withheld BILLED; all four fields match the key. On billed recall versus a false-billed rate, counted apart: in scope on the NOT-BILLED side, and a hit -- one of 115 — This reading's correct status is REVIEW, so it sits inside the 115 not-billed cells the false-BILLED rate is measured over. The model correctly withheld BILLED; the free floor did not, and is one of two false-BILLED cells in the whole run. On under- versus over-escalation, counted apart: in scope on the under-escalation side, and a hit -- one of 30 — The correct routing is SUPERVISOR (a REVIEW case), so this reading is one of 30 cells where under-escalation is possible. The model routed SUPERVISOR; the free floor routed NONE, which is the failure direction that leaves an aged, unbilled job with nobody looking at it. On the portfolio exposure reported as still open, in usd: in scope, and correctly included — This grader sums the contract value of every reading NOT reported BILLED. The model correctly counted this job's $1,901.60 as open exposure; the free floor excluded it. Across all 150 readings the model's total matched the gold exactly (0.0 pct error); the free floor's total was 3.59 pct off.
The formulaWhat it computes
accuracy = hits / 150 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% status accuracy · 2 more measured on this row
the fast tier, memory removed (THE CONTROL)
82.0% status accuracy · 2 more measured on this row
the strongest free floor, no model
96.0% status accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/billing.step over the planted inputs at generation time. This grader IS the reference, so its own TPR/TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate cannot be measured: it IS the reference. What can be wrong is the RULE it encodes -- see Data.breaks_on for the invented threshold and the whole-window simplification.
Watch these
status_accuracy_pct
billed_recall_pct
false_hold_pct
days_unbilled_accuracy_pct
Alarm on
any reply that does not parse. The scored run had 0 of 150; the calibration run at an under-sized 8,000-token ceiling had 1 of 9, which is why the published ceiling is 24,000.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. Denominators are printed beside every rate because three are thin: 35 BILLED cells, 9 ON_HOLD cells, 6 CONTEXT_INCOMPLETE cells, 2 reversal-decoy readings.
Cadence: Re-run evals/check_labels.py (free) on any change to the corpus generator, src/billing.py, src/segment.py or src/select.py. Re-run the scored eval (paid) on any change to src/prompt.py or src/state.py.
The decisionWhen to reach for it
Use it
The truth is known and three of the four answers are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real unbilled-work queue, where whether a register line genuinely bills a job is decided by a person reading it. That is why this corpus is generated rather than captured.
A living map of modern AI — kept current every morning