Decide what to do about a client's broken payment promise
Customers promise to pay by a date, then some of those promises get broken without anyone noticing which ones, or for how long. This app checks every promise each week and tells the desk exactly which step to take today.
PresenterOpens the private repo. Visible to admins only.
For the collections deskTechnology & SaaS · Cross-domain
Why it matters
Today's manual process, and the same job with the app
Accounts receivable at a software company, reviewing every payment promise on a weekly cycle.
✕Today's manual process
1Open each arrangement and reread the notes to see if the promise on screen still matches what the customer agreed.
2Work out how many weeks it's been broken, since nothing on the screen counts that automatically.
3Remember last week's call and decide today's step from memory, case by case.
4Guess wrong and either a broken promise sits untouched another week, or a paying customer gets an unwanted call.
Every promise reviewed from memory
✓With the app
1Each arrangement is checked every week, against the promise, the payments and the notes on file.
2How many weeks it's been broken carries forward on its own, written from the numbers, not memory.
3Today's step is recommended with the reason spelled out in plain words.
4A recent payment holds the line so a customer who just paid isn't pushed to the harshest step by mistake.
Every promise reviewed with its full history
See it work
One real case, read by the app, step by step
PTP-0045: a client promised $7,108 by July 20, paid $3,199 partway, and is now three weekly checks broken.
Decide what to do about a client's broken payment promiseReference appBuilt to be shaped to your process
4
1What it found broken: $7,108.02 was promised by July 20.
2What the desk does collector call, since $3,198.61 already came in against it.
3How long it's run three weekly checks broken, 28 days past the date.
4What's driving it a short partial payment, no dispute on file. The collector calls today.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Decide what to do about a client's broken payment promise
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
'Which promises are past their date' is a SELECT and nobody needs a language model for it -- the days past the date are printed on the page. The question a collections desk actually has to answer is the next one: what does somebody DO about this arrangement today. The ladder that answers it counts SCHEDULED OBSERVATIONS rather than days, so three of the terms it needs are not on the snapshot in front of you: how many consecutive weekly runs this promise has already been broken at, what was recommended at the last one, and how much had been credited then. A ledger that shows '28 days past' shows none of them. Somebody working a promise-to-pay queue by hand every Monday: opening each arrangement, working out from the collections notes whether the promise the ledger shows is still the promise the customer agreed to, remembering how many weeks in a row it has already been broken and what was done about it last week, spotting the part payment that landed on a Wednesday and appears in no dated entry, and only then deciding who picks up the phone.
Audience
Collections and AR operations in a software company, and the credit function that owns the escalation policy -- anyone running a promise-to-pay book on a weekly cycle, and anyone deciding whether a language model has any business near one. This kit's own answer to that last question is NO on the corpus it ships, and it publishes the numbers that say so. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual watch snapshots
The corpus is 150 watch snapshots, 0.70 MB (jsonl 1 · txt 150). It is generated because the alternative is somebody's real collections book, and a collections book is a list of named people at named companies who owe money. Generating it also buys the one thing a monitor most needs and cannot buy any other way: a CHAIN. Every arrangement is observed three times a week apart with events planted in the gaps, so the kit can ask 'what changed since the last reading' rather than 'what does this page say'. The nine planted patterns cover a promise kept, a promise part-paid then kept, a promise broken at three consecutive runs, a promise broken then renegotiated, a promise renegotiated then broken again, a dispute that caps the ladder, cash arriving in good faith between runs, and a promise not yet due -- and 14 observations carry a renegotiation or a dispute recorded only in a collector's sentence, because that is where the ledger and the human record actually disagree.
The corpus
The 150 watch snapshotsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your watch snapshots. That is the whole change — there is no database to migrate.
One watch snapshot, as the model receives itPTP-0001-R1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This file is generated for a public evaluation kit. Every customer, contact, invoice
number, amount and collector's note in it is invented, and the collections policy
reproduced below is illustrative: it is not any company's credit-management standard.
Arrangement
----------------------------------------------------------------
Arrangement : PTP-0001
Account : ACCT-4007 (Mid-Market, Team plan)
Customer : Northgate Lab
Scheduled run : 2026-08-03 07:00 UTC (run 1 of the weekly watch)
Promise recorded on : 2026-07-24
Promise generation : 1
Promised amount : 14,247.35 USD
Promised by : 2026-08-06
Invoices covered : INV-20013, INV-20014
Collections Policy
----------------------------------------------------------------
The promise-to-pay watch runs once a week, every Monday at 07:00 UTC, over every open
arrangement in the book. One run is one observation of an arrangement. The rungs below count
OBSERVATIONS, not days: an arrangement seen broken at three consecutive weekly runs is on rung 5
whether that took 15 days or 25.
1. Nothing is due and nothing has been missed -> NONE
2. The promise falls due within 2 days of this run and less than
half the promised amount has been credited against it -> REMINDER
3. The promise is broken -- the promised date has passed and the
credited amount is short -- and this is the FIRST scheduled run at
which it has been broken -> COLLECTOR_CALL
Abridged — the file continues.
The outcomeWhat a good result looks like
Per observation: the band the promise is in, the rung of the written ladder the desk should be on, how many consecutive scheduled runs the promise has been broken at, and what is holding the money -- plus a five-field carried state the next run is judged against, written by code from the parsed figures and never from the model's reply.
And when it cannot
It recommends the wrong rung, and the two directions cost very differently. Recommending too little leaves a broken promise unworked for another week; recommending too much means a phone call somebody did not deserve, and at the top rung it means proposing a credit-hold review of a customer who is either paying in instalments or in a genuine dispute. This run over-escalated 8 observations and under-escalated 0, and made 3 credit-hold recommendations the policy does not support.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Your ladder is arithmetic on figures your ledger already prints — a rule engine, and no model at all Measured here: free rules carrying five scalars scored 100.0 / 100.0 / 100.0 on band, rung and count against the model's 99.33 / 94.0 / 99.33, for $0.00 and no key.
Your collectors' notes carry facts your ledger does not — the model, on the prose fields only Rewrite the notes into unseen phrasings and the free keyword table falls to 16.67 pct on the driver and 83.33 pct on the band; the model on the identical set scored 76.67 pct and 100.0 pct.
You want the model to write the state as well as read it — no -- keep the state in code The carried quantity here is a counter that only climbs. A monitor that fed its own verdict into next week's prompt would turn one bad week into a credit-hold review two weeks later, and nothing downstream would notice.
And where nothing here is good enough:
You want to run the watch fortnightly to halve the bill — do not, or change the ladder first Measured for $0.00 by evals/cadence.py: skipping one weekly run leaves 9 surviving observations on a LOWER rung than they should be, never reaches 24 rungs at all, delays 10 by a run, loses 9 credit-hold recommendations, and makes 12 ledger events invisible to every later snapshot -- because the activity window is 3 days and the gap would be 14.
Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Decide what to do about a client's broken payment promise14 steps · 4 questions · run once, for real · 2026-08-23
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace data/ entirely. The measured result does NOT travel with the code.Corpus lens →
When is this the wrong choice?
Avoid: Do not buy a model to do arithmetic you can write down. If the rule fits on a page, write the page. That is the case against the best-fitting scenario (“Your ladder is arithmetic on figures your ledger already prints”). 4 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A snapshot missing any of the seven named sections. segment.split() returns what it finds and the prompt is assembled from it, so a parser drift produces a shorter prompt rather than an error -- evals/check_labels.py refuses to let a run spend until all seven are present in all 150. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
WHETHER A WRONG WEEK PROPAGATES. evals/run.py advances the carried state from the ledger's own planted facts, never from the model's reply -- which is the correct design and is also a boundary on what was measured. 6 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-23 — r001-ptp-watch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured reproduces everything except the model's answers: python3 tools/build_corpus.py rebuilds all 150 snapshots and the answer key byte for byte, python3 -m evals.check_labels runs the whole pre-flight, both free floors score in under a second, python3 -m evals.cadence prices a missed run, and python3 -m src.app renders the UI with every panel filled in. What it cannot do is call a provider, and the button says so in a sentence rather than an error.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
99.33%rows answered
14,901 msp50, end to end
50,865 msp95
2 minclone to first result
What the clock covers. END-TO-END per observation: one HTTP request carrying the assembled prompt, and the reply parsed to a band, a rung, a count and a driver code. Splitting the snapshot into sections, dropping the one section no field asks for, parsing the figures and advancing the carried state all happen outside this clock and cost no network at all. p95 (50,865 ms) is 3.4x p50 (14,901 ms) -- output length here is driven by how many of the nine rules interact on a given arrangement, not by page length: every snapshot is within a few hundred bytes of every other and replies ran from 534 to 9770 tokens.
Current processWhat it replaces
Somebody working a promise-to-pay queue by hand every Monday: opening each arrangement, working out from the collections notes whether the promise the ledger shows is still the promise the customer agreed to, remembering how many weeks in a row it has already been broken and what was done about it last week, spotting the part payment that landed on a Wednesday and appears in no dated entry, and only then deciding who picks up the phone.
Where it is not good enough
⚑ THE FREE RULE ENGINE WINS AND THIS KIT'S OWN PAGE LEADS WITH IT. A pure-Python floor carrying the same five scalars the model is given scores 100.0 pct on the band, 100.0 pct on the escalation rung and 100.0 pct on the consecutive count over all 150 observations -- exact, for $0.00, in under a second, with no key. The model scored 99.33 / 94.0 / 99.33 on the identical set and made 3 credit-hold recommendations the policy does not support against the floor's 0. On the corpus this kit ships, buying a model for this job buys nothing and costs $0.007904 an observation.
THAT IS NOT A WEAK MODEL, IT IS THE SHAPE OF THE TASK. Given the amounts and dates printed on the page and five scalars carried from last week, the ladder is arithmetic: a comparison, a counter and two overrides. There is nothing in it to read.
⚠︎ WHERE THE FLOOR ACTUALLY BREAKS, AND IT COST 30 CALLS TO FIND OUT. The 100 pct is partly a property of the fixture: the corpus draws every collector note from a fixed pool and the keyword table was written knowing what is in it. Rewrite the notes on the 10 prose-dependent arrangements into phrasings the table has never seen, change nothing else, and the free floor falls to 83.33 pct on the band, 63.33 pct on the rung and 16.67 pct on the driver, and starts recommending credit holds it should not (4 of 30). The model on the identical rewritten set scored 100.0 / 100.0 / 76.67. THAT is the comparison a buyer should read, and it is the only place on this page where compute has anything to sell.
⚠︎ THE MODEL'S SINGLE LARGEST FAILURE IS A REFUGE BUCKET, AND IT INVERTS. On the shipped corpus it maps the driver on 68.0 pct of observations against free code's 98.0 -- 30.0 points WORSE -- and 34 of the 48 it got wrong were filed under NO_REASON_RECORDED when a mapped code was available. A bucket for 'nobody wrote the reason down' is the honest way to answer an open question about a taxonomy, and it is also, measurably, where a model puts anything it does not want to commit on. ⚑ ON THE REWRITTEN NOTES THE SAME BUCKET IS USED 0 TIMES AND THE MODEL SCORES 76.67 PCT AGAINST FREE CODE'S 16.67. The model is not worse at reading than a keyword table; it is more willing to decline when the table is right by construction, and far better when it is not.
⚑ WHAT THE CARRIED STATE ACTUALLY BOUGHT, AND IT IS NOT THE BAND. The stateless control scored 100.0 pct on the band -- higher than the stateful run's 99.33 -- because the band is derivable from figures printed on the page. What memory bought is the COUNTER (99.33 pct against 70.67) and the RUNG (94.0 against 84.67), and on the 48 observations whose answer is not derivable from their own page the count goes 97.92 pct against 37.5. ⚠︎ AND REMOVING IT MADE THE RUN DEARER, NOT CHEAPER: 425,696 output tokens without the carried paragraph against 347,975 with it, $0.009383 an observation against $0.007904, and one reply spent 18,005 of its 20,000-token budget. A model told there is no history still has to answer 'how many consecutive runs', and it reasons about a number it cannot derive until the budget runs out.
⚠︎ AND THE CADENCE IS A CEILING NOBODY PRICES. Skipping one weekly run of three puts 9 surviving observations on a lower rung, never reaches 24 rungs at all, loses 9 credit-hold recommendations, and makes 12 ledger events invisible to every later snapshot. That is not a latency cost, it is a different answer -- and it is measured for $0.00 by evals/cadence.py, which is the half of a monitor this estate keeps skipping.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1
50 open promise-to-pay arrangements, re-read on 3 weekly runs
0.007904 per observation · the fast tierUnit cost ↗
Recorded failure1 of 150 replies lost: 1 returned an EMPTY body with finish_reason 'stop' after billing 7,427 output tokens -- budget spent on hidden reasoning, nowhere near the ceiling
free rules with the same memory: 100.0 / 100.0 / 100.0
on rewritten prose: model 100.0 pct, free 83.33 pct
2026-08-23as of
It produces a band, a rung RECOMMENDATION, a consecutive count and a driver code for a collector to read, and acts on none of it — there is no endpoint, no function and no flag anywhere in the kit that suspends an account, places a hold, cancels a subscription, writes anything off or emails anybody, and evals/check_labels.py asserts it before any run may spend. ⚑ THE STATION THAT MAKES THIS A MONITOR IS THE SECOND, AND THE ONE THAT MAKES IT A CADENCE KIT IS THE THIRD. The only route from last week is five scalars, written by src/promise.py from the figures PARSED off the page and never from the model's answer — which matters here because the carried quantity is a COUNTER THAT ONLY CLIMBS, so one bad week fed back would be a credit-hold review two weeks later. And the clock is a step rather than an instrument: the rungs count observations, so a missed run cannot be reconstructed by arithmetic.
⚠︎ THE HONEST HEADLINE IS ON THE SCORING STATION: free Python carrying the same five scalars is EXACT on band, rung and count and the model is not, so this kit's own recommendation is to write the rules down rather than buy compute for them. The one axis where that flips is prose the ledger never structured, and the prose probe measures it.
⚠︎ EVERY THRESHOLD IN THE POLICY — the weekly cadence, the five rungs, the two-day reminder lead, the half-the-amount test, the $25,000 shift, the dispute cap and the good-faith hold — IS INVENTED for this kit and reproduces no company's credit-management standard.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, MODEL in .env. Nothing else. The kit-local .env overrides the shared repo-root one key by key, so pointing one kit at a different model is one line.
the carried state
src/state.py
which scalars cross the week. Five here; a book with an instalment plan would carry the plan's next due date as a sixth. Change describe() and the prompt changes with it -- there is no template engine in between.
the ladder and the cadence
src/promise.py
RUN_INTERVAL_DAYS, the five rungs, the reminder lead, the half-the-amount test, LARGE_BALANCE_USD, the dispute cap and the good-faith hold. All in one file, all invented, all meant to be replaced by your own book's.
the corpus
tools/build_corpus.py
the whole of data/. Point src/watch.py's CORPUS at your own snapshots in the same seven-section shape and hand-label a gold set.
what never leaves the machine
src/select.py
NEVER_SENT. One tuple decides which sections are subtracted before anything is sent, and _fallback() subtracts them unconditionally.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 50 arrangements x 3 consecutive WEEKLY scheduled runs = 150 snapshots from a fixed seed (SEED = 20260823) across three customer segments. It plants nine patterns, of which 14 observations carry a renegotiation or a dispute recorded ONLY in the collector's prose and 9 carry a payment that moved outside the snapshot's activity window. The gold labels are src/promise.replay's output, never typed.
the collections ladder
src/promise.py
The nine rules as pure code: band, rung, the consecutive-run counter, the dispute cap, the good-faith hold and the large-balance shift. No model, no judgement. ⚠︎ The cadence, every rung and every threshold in it are INVENTED and the kit ships that as a stated blocker -- see Data.why_this_corpus.
the cadence
src/promise.py
RUN_INTERVAL_DAYS = 7, every Monday at 07:00 UTC, over the whole open book. It is not a setting hanging off the run -- it IS the run's trigger, and the rungs count observations rather than days, so a missed run cannot be reconstructed by arithmetic. evals/cadence.py prices one for $0.00.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. Five scalars per arrangement (the previous run's date, its band, its rung, the consecutive-broken count and the cash credited as at that run), written by the arithmetic and never from the model's reply, and rendered into one short paragraph for the prompt. It costs the same at run 40 as at run 3.
the section splitter
src/segment.py
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 snapshots before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Account Contact -- a named accounts-payable clerk, their work email, their direct line and a billing address -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state paragraph, and the selected sections in document order. The stateless build replaces the paragraph with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost observation.
the watch
src/watch.py
One arrangement, one scheduled run, one call. Parses the amounts, the dates and the ledger dispute flag off the page with regexes (the model is never asked to read a figure), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 20000, the published ceiling.
the local UI
src/app.py
One arrangement, one scheduled run, its carried state and both free floors beside the model's answer, on 127.0.0.1:9004. Renders with no key. Where the model agrees with the free memory floor, the page says the model bought nothing on that row.
the free floors
evals/baseline.py
TWO of them. onesnapshot is the ageing report a desk already runs: parsed figures, the ledger flag, an ordered keyword table over the note, and the consecutive count INFERRED from the days past the date. ledgermem is the same rules carrying the same five scalars the model is given. 0 calls, $0.00, scored through the identical scorer.
the cadence cost
evals/cadence.py
Replays every arrangement against a schedule that skips one run and diffs the two answer keys. 0 calls. It is the only place in this kit that measures the CLOCK rather than the reading.
the prose probe
evals/paraphrase.py
Rewrites the collector note on the 10 prose-dependent arrangements into phrasings the keyword table has never seen, so the free floor's score is tested off its own training set. Written in one pass and neither side edited afterwards.
the scorer
evals/scoring.py
Exact match per cell against the computed gold, split six ways an average would hide: the four fields, the observations where a rung fires, the quiet ones, the counter-reset subset, the memory-dependent subset, and the credit-hold recommendations the policy does not support. No judge model.
the pre-flight
evals/check_labels.py
Refuses to let a run spend until the parser, the schedule, the privacy guard, the stateless control's one-place difference, the answer key's provenance and the absent-guardrail grep all pass. Every one of them fails SILENTLY at run time.
Where it breaks at scale
Three ceilings, in the order they arrive. (1) THE STATE STORE. data/state.json is one file replaced atomically -- correct for one writer and not a concurrency model. A book big enough to want two workers writing state wants a database, and that is the point at which this kit is outgrown. (2) THE SCHEDULE. evals/run.py walks the runs serially and the arrangements inside a run concurrently, 50 wide; the wall clock for one scored run of 150 observations was 504.0 seconds at 12 workers. A book of a hundred thousand arrangements is a queue and a worker pool, not a ThreadPoolExecutor. (3) THE BILL. At $0.007904 an observation, a hundred thousand arrangements watched weekly is about $790 a week -- against a free rule engine that was EXACT on three of the four fields over this corpus.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
PTP-0045 at run 3, before anything is called. The two free columns are already filled in and they DISAGREE: reading this snapshot alone gives 5 consecutive broken runs and CREDIT HOLD REVIEW, because the page says 28 days past the promised date and the policy says the watch runs weekly. The truth is 3 -- the watch first saw this arrangement when it was already three weeks overdue -- and cash was credited between run 2 and run 3 on a day outside the snapshot's three-day activity window, so rule 8 holds the rung at COLLECTOR CALL. Nothing on the page distinguishes the two columns except numbers that were carried.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same page with no API_KEY configured, after pressing Run this observation. It returns a plain sentence saying nothing was called rather than an error, and every other panel -- the carried state, both free floors, the parsed figures and the list of what left the machine -- still renders. A kit that cannot be looked at without a credential is a kit nobody evaluates.failureOpen full size →PTP-0045 at run 3, answered live. The model's column sits beside the two free ones and the right-hand column says, row by row, whether it agreed with free code. This is a FAILURE shot in the sense the standard means: it is the page on which the model demonstrably bought nothing that free arithmetic had not already produced, and the kit publishes it rather than a row where it looked clever.failureOpen full size →
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
150watch snapshots
0.70 MiBjsonl 1 · txt 150
p50 4,880chars per watch snapshot
$0.00setup · 0.0s
How it is cutWhat one watch snapshot is
one snapshot per arrangement per scheduled run -- 50 arrangements x 3 weekly runs, never chunked. The whole open book is re-read on every scheduled run, so a 'split' here is the population divided by time rather than a document divided by length.
SetupWhat the setup figure measured
There is no index and nothing was built, so both figures are a MEASURED zero rather than an unmeasured one: no embedding call, no vector store, no chunk file, nothing on disk under data/index/. A monitor re-reads its whole population on every scheduled run and the finding lives in the difference between two readings, so there is nothing an index could be built over. Regenerating the corpus and the answer key takes under a second and needs no key -- that is corpus GENERATION, not index building, and it is counted in Data.cold_clone rather than here.
LicenceLicence
MIT, the same as the repository. Nothing in data/ is derived from anyone else's work, so there is nothing here to attribute.
Bring your ownBring your own watch snapshots
Replace data/ entirely. Point src/watch.py's CORPUS at your own snapshots in the same seven-section shape (Synthetic Record, Arrangement, Collections Policy, Cash Position, Activity Log, Account Contact, Collector Notes), put your own book's rules and cadence in src/promise.py, and hand-label a gold set. The scorer, the pre-flight, both free floors, the cadence measurement and the prose probe all work unchanged on any corpus in that shape.
⚠︎ And what stops being true when you do: The measured result does NOT travel with the code. This kit's answer key is COMPUTED by the same nine rules the kit publishes, from inputs the generator planted -- which is a luxury of controlling both the generator and the rules, and is why a free rule engine scores 100 pct on three of the four fields. On a real book the figures are parsed from documents somebody else wrote, the rules are your credit policy rather than these nine, and the key is hand-labelled with an error rate nothing here has measured. Bring the harness; re-measure everything.
What breaks it
A snapshot missing any of the seven named sections. segment.split() returns what it finds and the prompt is assembled from it, so a parser drift produces a shorter prompt rather than an error -- evals/check_labels.py refuses to let a run spend until all seven are present in all 150.
An arrangement whose scheduled runs are out of order or have a gap. The consecutive counter is a running total, so run 3 processed before run 2 is silently wrong rather than loudly broken. The pre-flight asserts all 50 sequences are complete and gap-free.
A collector note phrased outside the fixed pool. Measured, not asserted: on the shipped corpus the free keyword table maps the driver on 98.0 pct of observations, and on the same arrangements with the notes rewritten it maps 16.67 pct. That is a property of a fixed pool, not of keyword matching, and this kit paid 30 calls to find out.
A promise whose amount CHANGES on renegotiation. Every renegotiation planted here is a date extension for the same money, so the arithmetic never has to reconcile a reduced balance against cash already credited. A payment-plan renegotiation would, and nothing here has measured it.
An arrangement observed for the first time when it is already many weeks overdue. This is not a failure -- it is the corpus's main trap and 44 of 150 observations carry it -- but it is the shape that defeats every ageing inference: the page says 28 days past and the count is 1, because the watch had never seen it before.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
the question
2,496
604
the carried state
323
78
the snapshot
4,602
1,113
Total
1,795
This is the cost lesson as arithmetic: of the 1,795 tokens assembled, 1,113 are snapshots — 62% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
REPLAYED from the kit's own src/prompt.build() for PTP-0045-R3, not logged by the run. That is only trustworthy because the replayed section list matches what the run recorded (sections_used), and because the carried state it is built from is recomputed by the same src/promise.step the run used. The per-part token counts are apportioned by character share of the run's measured mean input (1794.94 tokens) -- the provider bills one number for the whole prompt and never breaks it down, so the split is arithmetic on a measured total rather than a second measurement.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are running one scheduled weekly observation of one promise-to-pay arrangement on a software
company's collections book, and deciding what the collections desk should do about it TODAY.
The collections policy, its rungs and its cadence are reproduced in the snapshot below. Apply them
exactly as written.
The policy counts SCHEDULED RUNS, not days. How many consecutive runs this arrangement has already
been broken at, what rung was recommended at the previous run, and how much had been credited as at
that run are NOT on this snapshot. What is known about earlier runs is stated under "Carried state"
and is the only history available to you. Do not assume anything about earlier runs beyond it.
Two things to read carefully, because the ledger and the collector disagree on purpose:
- The Activity Log covers only the three days up to and including this run, while the watch runs
every seven. Cash that arrived earlier in the week is in the Cash Position total and in no dated
entry anywhere on this page. Rule 8 asks whether cash MOVED since the previous run, which you can
only answer from the credited figure in the carried state.
- A renegotiation (rule 6) and an open dispute (rule 7) count when the Activity Log records them OR
when the collector's note for this run records them. A customer ASKING for new terms is not a
renegotiation; a customer who has agreed new terms that replace the promise above is.
Answer with a single JSON object and nothing else:
{"band": "STANDING|AT_RISK|BROKEN|KEPT|RENEGOTIATED",
"escalation": "NONE|REMINDER|COLLECTOR_CALL|ACCOUNT_MANAGER|CREDIT_HOLD_REVIEW",
"consecutive_broken_runs": <whole number of consecutive scheduled runs, INCLUDING this one, at
which this promise has been broken; 0 if it is not broken at this run>,
"top_driver": "DISPUTE_OPEN|AWAITING_PO|BANK_IN_FLIGHT|PART_PAID_SHORT|RENEGOTIATION_REQUESTED|NO_REASON_RECORDED",
"rationale": "one sentence, naming the rule you applied and the count you carried"}
"band" is the state of the promise at this run after rules 6 and 7 have been read. "escalation" is
the rung after the cap in rule 7 and the hold in rule 8 have been applied. "top_driver" is why the
money has not arrived, mapped to one of the six codes above; use NO_REASON_RECORDED when nothing on
the page gives a reason that maps to the other five, rather than choosing the nearest.
Carried state
----------------------------------------------------------------
The previous scheduled run was on 2026-08-10. It reported this arrangement BROKEN and recommended COLLECTOR_CALL. Counting that run, the promise has now been broken at 2 consecutive scheduled runs before this one. As at that run, 1,777.01 USD had been credited against the promise. The arrangement was then at generation 1.
Watch snapshot
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This file is generated for a public evaluation kit. Every customer, contact, invoice
number, amount and collector's note in it is invented, and the collections policy
reproduced below is illustrative: it is not any company's credit-management standard.
Arrangement
----------------------------------------------------------------
Arrangement : PTP-0045
Account : ACCT-4315 (SMB, Team plan)
Customer : Denholm Cranes
Scheduled run : 2026-08-17 07:00 UTC (run 3 of the weekly watch)
Promise recorded on : 2026-07-10
Promise generation : 1
Promised amount : 7,108.02 USD
Promised by : 2026-07-20
Invoices covered : INV-20585
Collections Policy
----------------------------------------------------------------
The promise-to-pay watch runs once a week, every Monday at 07:00 UTC, over every open
arrangement in the book. One run is one observation of an arrangement. The rungs below count
OBSERVATIONS, not days: an arrangement seen broken at three consecutive weekly runs is on rung 5
whether that took 15 days or 25.
1. Nothing is due and nothing has been missed -> NONE
2. The promise falls due within 2 days of this run and less than
half the promised amount has been credited against it -> REMINDER
3. The promise is broken -- the promised date has passed and the
credited amount is short -- and this is the FIRST scheduled run at
which it has been broken -> COLLECTOR_CALL
4. Broken at the SECOND consecutive scheduled run -> ACCOUNT_MANAGER
5. Broken at the THIRD consecutive scheduled run or later -> CREDIT_HOLD_REVIEW
6. A RENEGOTIATED promise starts again: the band is RENEGOTIATED, the rung returns to NONE and
the consecutive count returns to zero at that run. A promise is renegotiated when the Activity
Log records a RENEGOTIATED entry, OR when the collector's note for this run records the
customer agreeing new terms that supersede the arrangement above. A note in which the customer
merely ASKS for new terms is not a renegotiation.
7. CAP: while a dispute is open against any invoice this promise covers, the rung is capped at
COLLECTOR_CALL however many consecutive runs have passed. A dispute is open when the billing
ledger's dispute flag is set, OR when the collector's note for this run records the customer
disputing a covered invoice.
8. HOLD: if cash was credited against this promise since the PREVIOUS scheduled run, AND that
run recommended COLLECTOR_CALL or higher, the rung does not climb above the rung recommended
at that run. The consecutive count still rises; the action does not. A first breach is not
held by this rule, however much cash arrived.
9. A broken promise of 25,000.00 USD or more climbs one rung early: ACCOUNT_MANAGER at the
first broken run and CREDIT_HOLD_REVIEW at the second. Rules 7 and 8 are applied after it.
10. The driver is the single thing most responsible for the money not having arrived at this run.
NOTHING_OUTSTANDING is used when, and only when, the promise has been kept in full.
TERMS_RENEGOTIATION covers both a renegotiation agreed and one the customer has only asked
for; the band, not the driver, is what distinguishes them.
This watch produces a recommendation for a collector to read. It suspends nothing, holds no
orders, writes nothing off and sends nothing to a customer.
Cash Position
----------------------------------------------------------------
As at this scheduled run:
Promised amount : 7,108.02 USD
Credited against the promise : 3,198.61 USD
Outstanding on the promise : 3,909.41 USD
Days past the promised date : 28
Dispute flag on the billing ledger: not set
Activity Log
----------------------------------------------------------------
Entries dated in the 3 days up to and including this run. Anything older is reflected
in the Cash Position above and appears nowhere else on this page.
no entries in the 3 days up to this run
Collector Notes
----------------------------------------------------------------
2026-08-17 "They paid what they could this month and went quiet on the balance."
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"band": "STANDING", "escalation": "NONE", "consecutive_broken_runs": 0, "top_driver": "BANK_IN_FLIGHT", "rationale": "Rule 1 applies because the promised date is 2026-08-06, three days after this first weekly run, so nothing is due or missed; the consecutive broken count is 0, and the collector note that the bank shows payment sent but not yet arrived maps to BANK_IN_FLIGHT."}
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Decide what to do about a client's broken payment promise — 150 watch snapshots. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
Three of the four answered fields are closed sets -- five bands, five rungs, seven driver codes -- and the fourth is a whole number of scheduled runs. A judge is for answers whose correctness is a matter of reading; these are matters of equality, and asking a model to grade COLLECTOR_CALL == COLLECTOR_CALL would add cost, variance and a second thing to be wrong.
150watch snapshots
150source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED149 · 150 · 30 / 150band accuracy pct — observations, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED141 · 127 · 30 · 45 / 150escalation accuracy pct — observations, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED90 / 99escalation when a rule fires pct — observations whose correct rung is not NONE, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 51false alarm rate pct — observations whose correct rung IS NONE, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED149 · 106 · 48 / 150count accuracy pct — observations, exact whole-number match, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED43 · 18 / 44counter reset accuracy pct — observations where the page's own ageing inference disagrees with the truthDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED102 · 23 / 150top driver accuracy pct — observations, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED102 / 128driver when something outstanding pct — observations where the promise is NOT already kept in fullDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED47 / 48memory band accuracy pct — observations whose answer is NOT derivable from their own pageDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED42 / 48memory escalation accuracy pct — observations whose answer is NOT derivable from their own pageDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED12 / 12rule8 escalation accuracy pct — observations where the good-faith hold in rule 8 changes the rungDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED4 · 0 / 150bad credit holds — CREDIT_HOLD_REVIEW recommendations the policy does not support -- THE COSTLY ERRORDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and by mutation. By construction: the answer key is src/promise.replay's output over the planted inputs, and evals/check_labels.py asserts the corpus builder's own walk and the published replay() agree on all 150 observations. By mutation: the same pre-flight changes the credited figure in a rendered snapshot and asserts the free memory floor's band flips with it, which red-proves that the floor parses the page rather than reading the fixture. What is NOT validated is the ladder itself -- it is invented, and no amount of internal consistency makes an invented policy correct.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid. ⚠︎ THE THREE CARDS WERE NOT ALL VERIFIED ON THE SAME DAY: each row carries its own rates_as_of read from build/facts/models.json (google/gemini-3-flash 2026-08-18, openai/gpt-5-6-luna 2026-08-18, anthropic/claude-fable-5 2026-08-23), and this block's date is the newest of them rather than a single date asserted over all three.
Priced at
Per 1M in / out
One watch snapshot
1,000 watch snapshots
Share that is the prompt
google/gemini-3-flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than price lists
$0.50 / $3.00
$0.007904
$7.90
11%
openai/gpt-5-6-luna the cheapest row on the published set, for the low end
$0.20 / $1.20
$0.003161
$3.16
11%
anthropic/claude-fable-5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.134719
$134.72
13%
Same work, 43× the bill
The same watch snapshots, the same tokens — only the rate card changed. And across all 3 cards between 11% and 13% of what you pay is the prompt this pipeline sends, not the answer it writes.
The cadence. Everything else on this page is a per-call cost and the cadence is the number of calls -- but it is NOT a free knob: evals/cadence.py measures that skipping one weekly run puts 9 surviving observations on a lower rung, never reaches 24 rungs at all and hides 12 ledger events for good. Halving the bill halves the answer.
Rates checked 2026-08-23. The provider that actually ran all 389 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is recorded in the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER is free. evals/scoring.py makes no model call, and neither do either of the free floors it compares against, nor evals/cadence.py, nor the pre-flight. The unit that costs money is the OBSERVATION, and it is priced in the Cost lens; pricing the ruler in the units of the thing it measures is how a $0.00 grader gets reported as expensive.
The gradersOne way to grade, and why it is the only one
⚑ READ THIS BEFORE ANY OTHER NUMBER ON THIS PAGE. The free memory floor scores 100.0 pct on the band, 100.0 pct on the escalation rung and 100.0 pct on the consecutive count over all 150 observations. It is exact. It costs $0.00, it needs no key, it runs in under a second, and the model does not beat it -- 99.33 pct, 94.0 pct and 99.33 pct on the same three fields. On the corpus this kit ships, THE HONEST DEPLOYMENT OF THIS USE CASE IS A RULE ENGINE.
That is not an accident of a weak model and this kit says why: given the figures printed on the page and five scalars carried from last week, the ladder is arithmetic. There is nothing in it for a language model to read.
⚠︎ AND THE 100 PCT IS PARTLY AN ARTEFACT OF THE FIXTURE, WHICH IS WHY THE PROSE PROBE EXISTS. The corpus draws every collector note from a fixed pool and the keyword table was written knowing what is in it -- a matcher tested on the sentences it was written for. Rewrite the notes on the 10 prose-dependent arrangements into phrasings it has never seen and the free floor falls to 83.33 pct on the band, 63.33 pct on the rung and 16.67 pct on the driver, and starts recommending credit-hold reviews it should not (4 of 30). The model was run on the identical rewritten set and scored 100.0 / 100.0 / 76.67. THAT comparison, not the headline one, is what a buyer should read.
⚑ AND THE TWO FLOORS FAIL IN OPPOSITE, READABLE WAYS. The memoryless one NEVER under-escalates -- 0 of 150 observations, zero rungs below the key -- it only climbs too far, and its 25 unsupported credit-hold recommendations land almost entirely on the three patterns memory exists for: goodfaith_hold (12), broken_climbing_stale (9), broken_then_renegotiated (4). The memory floor's only weakness is prose: on the rewritten notes it files 29 of 30 drivers under NO_REASON_RECORDED, which is a keyword table meeting sentences it was not written for and answering 'nobody wrote it down' to almost all of them.
no headline metric on any of its 6 runs — they record band accuracy · escalation accuracy · consecutive count accuracy · top driver accuracy
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set can separate the two FLOORS cleanly -- 100.0 pct against 80.0 pct on the escalation rung, over 150 observations -- and it can separate memory from no memory. What it cannot do is separate the model from the free memory floor on the three closed fields, because the floor is at the ceiling: there is no headroom above 100.0 pct for anything to be better in. The only axis on this set with room to move is the driver, and the prose probe is the only arm that opens it.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Your ladder is arithmetic on figures your ledger already prints
a rule engine, and no model at all
Measured here: free rules carrying five scalars scored 100.0 / 100.0 / 100.0 on band, rung and count against the model's 99.33 / 94.0 / 99.33, for $0.00 and no key.
do not buy a model to do arithmetic you can write down. If the rule fits on a page, write the page.
Your collectors' notes carry facts your ledger does not
the model, on the prose fields only
Rewrite the notes into unseen phrasings and the free keyword table falls to 16.67 pct on the driver and 83.33 pct on the band; the model on the identical set scored 76.67 pct and 100.0 pct.
do not let it near the arithmetic. Parse the figures, carry the state in code, and hand the model only the sentence nobody structured.
You want to run the watch fortnightly to halve the bill
do not, or change the ladder first
Measured for $0.00 by evals/cadence.py: skipping one weekly run leaves 9 surviving observations on a LOWER rung than they should be, never reaches 24 rungs at all, delays 10 by a run, loses 9 credit-hold recommendations, and makes 12 ledger events invisible to every later snapshot -- because the activity window is 3 days and the gap would be 14.
do not read a cadence change as a cost decision alone. The rungs count observations, so halving the cadence changes the ANSWER, not just the latency.
You want the model to write the state as well as read it
no -- keep the state in code
The carried quantity here is a counter that only climbs. A monitor that fed its own verdict into next week's prompt would turn one bad week into a credit-hold review two weeks later, and nothing downstream would notice.
avoid any design where the model's answer becomes next period's input. It is the one architectural mistake this shape makes that a score cannot see.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
ESC-OVER
Recommended a rung ABOVE the one the ladder allows
8
PTP-0034-R2 (dispute_cap, run R2). The key says COLLECTOR_CALL; the model answered ACCOUNT_MANAGER. The page showed 9 days past the promised date, 0.00 USD credited of 8,861.97 promised, ledger dispute flag set; the carried state said the promise had been…
CNT-WRONG
Consecutive-run count wrong
1
PTP-0036-R1 (dispute_cap, run R1). The key says 1; the model answered None. The page showed 8 days past the promised date, 0.00 USD credited of 66,661.94 promised, ledger dispute flag set; the carried state said the promise had been broken at 0 consecutive…
BAND-WRONG
Band wrong
1
PTP-0036-R1 (dispute_cap, run R1). The key says BROKEN; the model answered None. The page showed 8 days past the promised date, 0.00 USD credited of 66,661.94 promised, ledger dispute flag set; the carried state said the promise had been broken at 0…
DRV-WRONG
Driver wrong
48
PTP-0001-R2 (kept_clean, run R2). The key says NOTHING_OUTSTANDING; the model answered NO_REASON_RECORDED. The page showed 4 days past the promised date, 14,247.35 USD credited of 14,247.35 promised, ledger dispute flag not set; the carried state said the…
What we could NOT verify
WHETHER A WRONG WEEK PROPAGATES. evals/run.py advances the carried state from the ledger's own planted facts, never from the model's reply -- which is the correct design and is also a boundary on what was measured. The model is never handed a carried state that is itself wrong, so nothing here measures what happens on the run after a bad one in a deployment whose renegotiation detection is imperfect.
WHETHER ENGLISH IS THE RIGHT WAY TO CARRY STATE. src/state.describe() renders five scalars as a paragraph rather than as JSON. The obvious experiment -- score the same 150 observations with the state rendered both ways -- costs one more full run and has not been paid for.
WHETHER THE MODEL IS STABLE. Every model figure here comes from ONE run of each arm. No observation was fired twice, so nothing on this page distinguishes a finding from sampling noise, and a single-run difference of a point or two should be read as no difference.
WHETHER A SECOND MODEL WOULD DO BETTER. One tier was run. This kit compares MEMORY against no memory and CODE against a model; it does not compare two models, and nothing here ranks any.
WHETHER THE INVENTED LADDER RESEMBLES ANY REAL ONE. The cadence, the five rungs, the two-day reminder lead, the half-the-amount test, the 25,000 USD large-balance shift, the dispute cap and the good-faith hold were all written for this kit and confirmed against nothing. Every score on this page is an accuracy against an invented rulebook.
WHETHER A HAND-LABELLED KEY WOULD AGREE. This key is computed, which is a luxury of controlling both the generator and the rules. A real book's key is hand-labelled and has an error rate nothing here has measured.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
google/gemini-3-flash
openai/gpt-5-6-luna
anthropic/claude-fable-5
the fast tier, with the carried state
1,794.94
2,335.4
14,901 ms
$0.007904
$0.003161
$0.134719
the same tier, memory removed (THE CONTROL)
1,738.31
2,837.97
17,968 ms
$0.009383
$0.003753
$0.159282
free rules carrying the same five scalars
0
0
0 ms
$0.000000
$0.000000
$0.000000
free rules, one snapshot, no memory
0
0
0 ms
$0.000000
$0.000000
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-23. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same observation, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
389 live calls are on the shared ledger for this kit and every one of them is behind a figure on this page. GRADING ITSELF COSTS NOTHING: evals/scoring.py makes no model call, and neither do the two free floors, the cadence measurement, the pre-flight or the prose probe's free arm. A forker re-running the eval pays the pipeline bill and no second bill on top of it. Figures are projected onto the Google Gemini 3 Flash card, not the real spend.
Cost driversWhat actually moves the bill
Output tokens. 2335.4 per observation against 1794.94 in, so output is 88.6 pct of the projected bill on the shared card.
The collections policy, reproduced on every snapshot. It is roughly 2774 characters of the 7421-character prompt -- about 37 pct of every call's input -- and one scheduled run sends it once per arrangement, so a book of 50 costs 50 copies of it a week, forever. That is the price of not asking a model to apply nine rules it cannot see, and it is paid on every observation rather than once.
MEMORY, AND IT PUSHES THE BILL DOWN RATHER THAN UP. Removing the five-scalar paragraph cost $0.009383 an observation against $0.007904 with it -- 1.19x -- because output grew from 2335.4 tokens to 2837.97. On this shape a carried counter is not an overhead you pay for accuracy; it is the cheaper configuration AND the accurate one.
The CADENCE, which is the multiplier nobody prices. One observation is one call; a weekly watch over a book of N arrangements is N calls a week forever, not N calls once. That is the cost shape of a monitor and it is the one a per-document price hides.
Provider-side reasoning, left at the default. 95.9 pct of this run's output was reasoning tokens that never reach the parsed answer and still count against max_tokens.
Your volumeWhat it costs at your volume
Linear in arrangements and linear in runs, with no threshold anywhere in the range tested. Ten times the book is ten times the bill: $0.007904 an observation, so a book of 500 watched weekly is about $3.95 a week and 100,000 is about $790. ⚠︎ THE ARITHMETIC ARGUES AGAINST THE KIT AT EVERY SCALE, because the free rule engine on the row above costs $0.00 and was exact on three of the four fields.
Where pricing changes shape
max_tokens = 20000 is a ceiling, not a cost -- the provider bills tokens produced, not tokens allowed. It was set from results/eval-c000-ptp-watch-calibration.json, whose largest reply was 10684 output tokens at a 32,000-token ceiling; the published ceiling is 1.9x that. The scored run's largest reply was 9770, 48.9 pct of the cap.
Your return, with your numbers
Volumearrangement-observations per scheduled run -- this run judged 150 (50 arrangements x 3 weekly runs) per arm. A real book multiplies by the number of open arrangements AND by 52 runs a year.
What it replacessomebody working a promise-to-pay queue by hand every Monday: reconstructing how many weeks in a row each promise has been broken, what was done last week, and whether the collector's note supersedes what the ledger says
Time saved per itemnot measured here -- and on this corpus the honest answer is that a free rule engine saves the same time for nothing, so the return to compute is zero on three of the four fields
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
One tier was run, and the comparison this kit is built for is not between two models -- it is between a model and free code. Point .env at your own model and every arm re-runs on your numbers.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
267,446input tokens · this run
347,975output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 149 observations, one completion call each, one tier. The stateless control, the prose probe and the cadence ablation are accounted for separately in Cost.cost_of_evaluation_usd and are not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.471
$1.347
$3.16
2026-09-12
gemini-3-flash
Google
$1.178
$3.367
$7.90
2026-09-18
gemini-3-8-flash
Google
$1.505
$4.304
$10.10
2026-09-18
llama-5
Meta
$1.813
$5.184
$12.17
2026-09-18
claude-haiku-4-5
Anthropic
$2.007
$5.739
$13.47
2026-09-12
grok-4-5
xAI
$2.623
$7.499
$17.60
2026-09-18
grok-4-6
xAI
$2.623
$7.499
$17.60
2026-09-18
claude-sonnet-5
Anthropic
$4.015
$11.478
$26.94
2026-09-12
gemini-3-1-pro
Google
$4.711
$13.468
$31.61
2026-09-18
gpt-5-6-terra
OpenAI
$4.711
$13.468
$31.61
2026-09-12
gpt-5-6-sol
OpenAI
$8.029
$22.956
$53.89
2026-09-12
claude-opus-4-8
Anthropic
$10.037
$28.695
$67.36
2026-09-12
claude-opus-5
Anthropic
$10.037
$28.695
$67.36
2026-09-12
claude-fable-5
Anthropic
$20.073
$57.390
$134.72
2026-09-18
claude-fable-5-1
Anthropic
$20.073
$57.390
$134.72
2026-09-18
gpt-6-astra
OpenAI
$20.073
$57.390
$134.72
2026-09-17
Read this against the numbers above
Nobody paid any of these. They are measured token counts times published list prices checked on 2026-08-18, and a vendor reprices without telling you.
usd_reproduce_this_whole_run is FOUR live arms plus the calibration, not one eval pass: the scored run, the stateless control, the prose probe and the cadence ablation are all behind figures on these pages, and a forker who re-measures the argument pays for all of them.
The spread across the three cards is a property of the PRICE LISTS, not of the work. Do not read a ratio between two of these rows as a fact about models.
Every row here is above $0.00, and the free rule engine on this kit's own scoreboard is at $0.00 and was exact on three of the four fields. The right column to compare these against is that one.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
15 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 50 arrangements x 3 consecutive WEEKLY scheduled runs = 150 snapshots from a fixed seed (SEED = 20260823) across three customer segments. It plants nine patterns, of which 14 observations carry a renegotiation or a dispute recorded ONLY in the collector's prose and 9 carry a payment that moved outside the snapshot's activity window. The gold labels are src/promise.replay's output, never typed.
You change it to: the whole of data/. Point src/watch.py's CORPUS at your own snapshots in the same seven-section shape and hand-label a gold set.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic from one seed; no network, no model.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260823
RULE = "-" * 64
RUN_DATES = [dt.date(2026, 8, 3), dt.date(2026, 8, 10), dt.date(2026, 8, 17)]
CUSTOMERS = [
FIRST = ["Dana", "Marcus", "Priya", "Tomas", "Aoife", "Kenji", "Lena", "Ruben", "Sofia", "Idris",
src/promise.pythe collections ladder — a swap seam
The nine rules as pure code: band, rung, the consecutive-run counter, the dispute cap, the good-faith hold and the large-balance shift. No model, no judgement. ⚠︎ The cadence, every rung and every threshold in it are INVENTED and the kit ships that as a stated blocker -- see Data.why_this_corpus.
You change it to: RUN_INTERVAL_DAYS, the five rungs, the reminder lead, the half-the-amount test, LARGE_BALANCE_USD, the dispute cap and the good-faith hold. All in one file, all invented, all meant to be replaced by your own book's.
src/promise.py
# The promise-to-pay policy, and the arithmetic it implies. Pure code, no model.
RUN_INTERVAL_DAYS = 7
RUN_TIME_UTC = "07:00"
RUN_DAY_NAME = "Monday"
ACTIVITY_WINDOW_DAYS = 3
STANDING = "STANDING"
AT_RISK = "AT_RISK"
BROKEN = "BROKEN"
KEPT = "KEPT"
RENEGOTIATED = "RENEGOTIATED"
src/promise.pythe cadence — a swap seam
RUN_INTERVAL_DAYS = 7, every Monday at 07:00 UTC, over the whole open book. It is not a setting hanging off the run -- it IS the run's trigger, and the rungs count observations rather than days, so a missed run cannot be reconstructed by arithmetic. evals/cadence.py prices one for $0.00.
You change it to: RUN_INTERVAL_DAYS, the five rungs, the reminder lead, the half-the-amount test, LARGE_BALANCE_USD, the dispute cap and the good-faith hold. All in one file, all invented, all meant to be replaced by your own book's.
src/promise.py
# The promise-to-pay policy, and the arithmetic it implies. Pure code, no model.
RUN_INTERVAL_DAYS = 7
RUN_TIME_UTC = "07:00"
RUN_DAY_NAME = "Monday"
ACTIVITY_WINDOW_DAYS = 3
STANDING = "STANDING"
AT_RISK = "AT_RISK"
BROKEN = "BROKEN"
KEPT = "KEPT"
RENEGOTIATED = "RENEGOTIATED"
src/state.pythe carried state — a swap seam
SEAM 2 -- the thing that makes this a monitor. Five scalars per arrangement (the previous run's date, its band, its rung, the consecutive-broken count and the cash credited as at that run), written by the arithmetic and never from the model's reply, and rendered into one short paragraph for the prompt. It costs the same at run 40 as at run 3.
You change it to: which scalars cross the week. Five here; a book with an instalment plan would carry the plan's next due date as a sixth. Change describe() and the prompt changes with it -- there is no template engine in between.
src/state.py
# The carried state -- what makes this a monitor and not another one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_arrangement(store, arrangement_id):
def advance(store, arrangement_id, days_past, credited, promised, renegotiated, dispute_open,
def _usd(v):
def describe(state):
src/segment.pythe section splitter
Splits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 snapshots before a run may spend.
src/segment.py
# Split a watch snapshot into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Arrangement", "Collections Policy", "Cash Position",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector — a swap seam
Decides which sections reach the model. Account Contact -- a named accounts-payable clerk, their work email, their direct line and a billing address -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing.
You change it to: NEVER_SENT. One tuple decides which sections are subtracted before anything is sent, and _fallback() subtracts them unconditionally.
src/select.py
# Pick which sections of a snapshot are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
ARRANGEMENT = "Arrangement"
POLICY = "Collections Policy"
CASH = "Cash Position"
ACTIVITY = "Activity Log"
CONTACT = "Account Contact"
NOTES = "Collector Notes"
NEVER_SENT = (CONTACT,)
SECTION_HINTS = {
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state paragraph, and the selected sections in document order. The stateless build replaces the paragraph with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string concatenation
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost observation.
You change it to: PROVIDER, BASE_URL, MODEL in .env. Nothing else. The kit-local .env overrides the shared repo-root one key by key, so pointing one kit at a different model is one line.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe watch
One arrangement, one scheduled run, one call. Parses the amounts, the dates and the ledger dispute flag off the page with regexes (the model is never asked to read a figure), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 20000, the published ceiling.
src/watch.py
# One arrangement, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You apply a written collections policy to one promise-to-pay arrangement at one "
MAX_TOKENS = 20000
FIELDS = ("band", "escalation", "consecutive_broken_runs", "top_driver")
def documents():
def arrangements():
def scheduled_runs():
def load_doc(doc_id):
src/app.pythe local UI
One arrangement, one scheduled run, its carried state and both free floors beside the model's answer, on 127.0.0.1:9004. Renders with no key. Where the model agrees with the free memory floor, the page says the model bought nothing on that row.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "9004"))
def _gold():
GOLD_ROWS = _gold()
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe free floors
TWO of them. onesnapshot is the ageing report a desk already runs: parsed figures, the ledger flag, an ordered keyword table over the note, and the consecutive count INFERRED from the days past the date. ledgermem is the same rules carrying the same five scalars the model is given. 0 calls, $0.00, scored through the identical scorer.
evals/baseline.py
# The free, no-model floors. TWO of them, because one would have flattered the model.
DRIVER_KEYWORDS = [
ASKED_PAT = r"has asked to|they want to re-cut|nothing agreed|have not accepted|would come back"
AGREED_PAT = (r"agreed with their|new date agreed|settled on the call|we have re-cut|"
DISPUTE_NOTE_PAT = r"disputes? |contested|not paying|will not pay|query on|is wrong and"
def classify_driver(note, credited, promised):
def read_page(text):
def _answer(band, rung, n, driver, why):
def review_onesnapshot(text, _state=None):
def review_ledgermem(text, state):
evals/cadence.pythe cadence cost
Replays every arrangement against a schedule that skips one run and diffs the two answer keys. 0 calls. It is the only place in this kit that measures the CLOCK rather than the reading.
evals/cadence.py
# WHAT A MISSED SCHEDULED RUN COSTS. Pure code over the answer key. Zero calls, zero dollars.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
OUT = os.path.join(HERE, "results", "cadence-cost.json")
def main():
evals/paraphrase.pythe prose probe
Rewrites the collector note on the 10 prose-dependent arrangements into phrasings the keyword table has never seen, so the free floor's score is tested off its own training set. Written in one pass and neither side edited afterwards.
evals/paraphrase.py
# THE PROBE THAT MAKES THE FREE FLOOR'S SCORE HONEST.
PARA = {
NOTE_RE = re.compile(r"(Collector Notes\n-+\n \S+ \")(.*)(\")", re.S)
def is_prose_dependent(gold_row):
def _key(gold_row):
def rewrite(text, gold_row, seq):
evals/scoring.pythe scorer
Exact match per cell against the computed gold, split six ways an average would hide: the four fields, the observations where a rung fires, the quiet ones, the counter-reset subset, the memory-dependent subset, and the credit-hold recommendations the policy does not support. No judge model.
evals/scoring.py
# Score a run against the computed gold. Pure code, no model, no judge.
FIELDS = ("band", "escalation", "consecutive_broken_runs", "top_driver")
RUNGS = ("NONE", "REMINDER", "COLLECTOR_CALL", "ACCOUNT_MANAGER", "CREDIT_HOLD_REVIEW")
QUIET_RUNG = "NONE"
COSTLY_RUNG = "CREDIT_HOLD_REVIEW"
NO_REASON = "NO_REASON_RECORDED"
NOTHING = "NOTHING_OUTSTANDING"
def _pct(n, d):
def _norm(v):
def _as_int(v):
evals/check_labels.pythe pre-flight
Refuses to let a run spend until the parser, the schedule, the privacy guard, the stateless control's one-place difference, the answer key's provenance and the absent-guardrail grep all pass. Every one of them fails SILENTLY at run time.
evals/check_labels.py
# Pre-flight. Everything that must be true BEFORE a run is allowed to spend.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
FAILURES = []
def check(name, ok, detail=""):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 50 arrangements x 3 consecutive WEEKLY scheduled runs = 150 snapshots from a fixed seed (SEED = 20260823) across three customer segments. It plants nine patterns, of which 14 observations carry a renegotiation or a dispute recorded ONLY in the collector's prose and 9 carry a payment that moved outside the snapshot's activity window. The gold labels are src/promise.replay's output, never typed. A swap seam.
src/promise.pyThe nine rules as pure code: band, rung, the consecutive-run counter, the dispute cap, the good-faith hold and the large-balance shift. No model, no judgement. ⚠︎ The cadence, every rung and every threshold in it are INVENTED and the kit ships that as a stated blocker -- see Data.why_this_corpus. A swap seam.
src/promise.pyRUN_INTERVAL_DAYS = 7, every Monday at 07:00 UTC, over the whole open book. It is not a setting hanging off the run -- it IS the run's trigger, and the rungs count observations rather than days, so a missed run cannot be reconstructed by arithmetic. evals/cadence.py prices one for $0.00. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Five scalars per arrangement (the previous run's date, its band, its rung, the consecutive-broken count and the cash credited as at that run), written by the arithmetic and never from the model's reply, and rendered into one short paragraph for the prompt. It costs the same at run 40 as at run 3. A swap seam.
src/segment.pySplits a snapshot into its seven named sections. Pure code; evals/check_labels.py asserts all seven are present in all 150 snapshots before a run may spend.
src/select.pyDecides which sections reach the model. Account Contact -- a named accounts-payable clerk, their work email, their direct line and a billing address -- is mapped by no field and is subtracted unconditionally by _fallback(), so the one section whose subject is a PERSON never leaves the machine even when a hint matches nothing. A swap seam.
src/prompt.pyThree parts: the fixed instruction, the carried-state paragraph, and the selected sections in document order. The stateless build replaces the paragraph with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text, and separates a transport failure from an HTTP status so a dropped connection is retried rather than recorded as a lost observation. A swap seam.
src/watch.pyOne arrangement, one scheduled run, one call. Parses the amounts, the dates and the ledger dispute flag off the page with regexes (the model is never asked to read a figure), assembles the prompt, calls the model, parses the reply. Holds MAX_TOKENS = 20000, the published ceiling.
evals/baseline.pyTWO of them. onesnapshot is the ageing report a desk already runs: parsed figures, the ledger flag, an ordered keyword table over the note, and the consecutive count INFERRED from the days past the date. ledgermem is the same rules carrying the same five scalars the model is given. 0 calls, $0.00, scored through the identical scorer.
evals/cadence.pyReplays every arrangement against a schedule that skips one run and diffs the two answer keys. 0 calls. It is the only place in this kit that measures the CLOCK rather than the reading.
evals/paraphrase.pyRewrites the collector note on the 10 prose-dependent arrangements into phrasings the keyword table has never seen, so the free floor's score is tested off its own training set. Written in one pass and neither side edited afterwards.
evals/scoring.pyExact match per cell against the computed gold, split six ways an average would hide: the four fields, the observations where a rung fires, the quiet ones, the counter-reset subset, the memory-dependent subset, and the credit-hold recommendations the policy does not support. No judge model.
evals/check_labels.pyRefuses to let a run spend until the parser, the schedule, the privacy guard, the stateless control's one-place difference, the answer key's provenance and the absent-guardrail grep all pass. Every one of them fails SILENTLY at run time.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1794 input and 2335 output tokens per observation, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Observations/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per observation directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py's fixed seed, including the Collector Notes. src/select.py names the notes as the field an outside party WOULD influence in a real deployment -- a customer's own words, quoted by a collector -- and sends them deliberately, so the surface is visible rather than hidden. On this corpus nobody outside the repository wrote a byte of it.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
The experimentWe did not attack it -- and the boundary that matters most on this vertical was proven by trying to break it
An indirect prompt injection needs a field an outside party controls that reaches the prompt, and this corpus has none. So the three rows above are boundaries, not payloads. One of them is red-proven in both directions before every run: evals/check_labels.py strips the sections that Account Contact's siblings map to, checks that the fallback STILL subtracts the contact, and then checks that the same stripped document WOULD carry it if the subtraction were removed -- so the test is known to be able to fail. That second half is the point. A guard that has never been made to fail is a guard nobody has tested, and this estate has already paid once for a red-proof that measured zero because the path it exercised was unreachable. Confirmed by assertion and by reading the recorded run, not by an attack trial, on 2026-08-23 -- no attack run was fired; see Eval.redteam_page.
Boundary checked
What could go wrong
What the code guarantees
Does a named accounts-payable clerk's email and direct line ever leave the machine?
Every snapshot carries an Account Contact section -- a named person, their work email, their direct line and a billing address. It is the only place in this corpus whose subject is a PERSON, and on a collections book the person named is usually NOT the debtor: they are an AP clerk whose employer is late. A watchlist of arrangement ids is an operations artefact; the same list with names and direct lines on it is a disclosure about individuals.
src/select.SECTION_HINTS maps no field to Account Contact, and _fallback() subtracts it UNCONDITIONALLY rather than falling back to the whole document. evals/check_labels.py asserts it is absent from every one of the 150 prompts AND red-proves the guard by stripping the mapped sections and checking the fallback still subtracts it.
Can this kit suspend an account, place a hold or write a balance off?
A collections tool that recommends a credit-hold review is one refactor away from performing one, and the refactor is usually somebody adding a convenience endpoint.
There is no such code path. The only writers anywhere in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). evals/check_labels.py greps every .py, .js and .mjs file for suspension, hold, write-off, cancellation and email function names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
Can the model's own answer change what the next run is told?
The carried quantity here is a counter that only climbs. If the model's reply wrote the state, one wrong week would become a credit-hold review two weeks later and nothing downstream would notice.
src/promise.step() takes parsed figures and the previous state and never sees the reply. evals/run.py advances the state from the ledger's own facts in the same branch whether the call succeeded or failed.
Three boundaries, all of them properties of what is ABSENT rather than of something that runs. That is the honest shape for a kit whose whole guarantee is 'it recommends and does not act', and it is also the shape that is easiest to lose in a refactor -- which is why two of the three are asserted by a gate that runs before any money is spent.
The result0 attack trials, three boundaries checked -- and the privacy boundary red-proven in both directions by evals/check_labels.py before any run may spend.
0untrusted input fields on this corpus
0 of 0attack trials run
1 of 3boundaries red-proven, not just asserted
0 of 150prompts carrying a named person's email or direct line
The Collector Notes ARE the field an outside party would influence in a real deployment, and this kit sends them rather than hiding them -- rules 6 and 7 both do their work there, so a version that withheld them would be measuring a different task. On this corpus they are drawn from a fixed pool by a seeded generator, so there is nothing adversarial in them to catch. What IS measured is the privacy boundary, in both directions, before every run.
Read this twice
The carried state is written from the arithmetic, never from the model's reply, and on this kit that matters more than usual because the carried quantity is a COUNTER THAT ONLY CLIMBS. A monitor that fed its own rung forward would not lose one week to a bad answer, it would escalate on it: one wrong BROKEN in week 2 is an ACCOUNT_MANAGER in week 3 and a CREDIT_HOLD_REVIEW in week 4, with nothing downstream able to tell. src/promise.step() never reads the reply. ⚠︎ AND THIS RUN CANNOT PROVE IT, which the page says rather than letting the number imply otherwise: the guarantee is in the code path and the run is consistent with it. No arm was fired in which the model wrote the state, so what is guaranteed is narrower than correctness -- only that a wrong answer stays where it is.
HonestyWhat this does not prove
No injection attempt was fired at the collector's note, which is this kit's own stated surface. Nothing here measures whether a customer's quoted sentence can move the band.
The privacy guard is asserted over the corpus that ships. A snapshot in a different shape -- a contact block inside the Activity Log, say -- would not be caught by a section-name rule.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No action of any kind, non-configurable. This kit produces a band, a rung RECOMMENDATION, a count and a driver code for a collector to read. It never suspends an account, places a credit hold, cancels a subscription, adjusts, writes off or emails anybody, and there is no setting that makes it. Separately: the carried state is written by code from the parsed figures and the previous state; the model's four answers are scored and are never written back into the history the next run is judged against.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json) and src/state.save (data/state.json). src/state.advance() -> src/promise.step() is the only thing that touches the carried state, and evals/run.py calls it after every observation including one whose call FAILED.
EvidenceDoes it hold?
What
Measured
Nothing in this kit suspends, holds, cancels, writes off or emails
0 code paths. evals/check_labels.py greps every .py, .js and .mjs file for nine such names and passes at 0. ⚠︎ It asserts the absence of names it knows; a path called something else would pass it.
A wrong answer does not propagate into the next run's state
59 wrong cells on r001-ptp-watch across 150 observations, and the carried state for every following observation was byte-identical to what it would have been had the model been right -- because src/promise.step() never reads the reply. ⚠︎ The property is guaranteed by the code rather than measured by an experiment: no run was fired in which the model wrote the state.
The top rung is a RECOMMENDATION to review, never a hold
CREDIT_HOLD_REVIEW appears 11 times in the answer key and 14 times in the model's answers on r001. In both cases it is a string in a JSON object that a person reads. Nothing consumes it.
The limitWhat a guardrail is not
Not a credit-control system. It answers one question about one arrangement at one point on a clock and hands the answer to a person.
Not a dunning engine. It sends nothing to a customer and there is no template, no address book and no outbound path anywhere in the kit.
Not a collections workflow. It opens no case, assigns no owner, sets no follow-up date and writes to nothing but results/*.json and data/state.json.
Not a decision. The top rung is a proposal to REVIEW whether a credit hold is warranted, and the kit's own scoreboard says free code makes fewer unsupported ones than the model does.
Not a scheduler. It is invoked rather than woken, so nothing in it can notice that a scheduled run did not happen -- which evals/cadence.py prices at 9 lost credit-hold recommendations and 12 permanently hidden ledger events per missed weekly run.
WatchedWhat is watched, and why that one
10runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 41 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
14 measured by the latest run27 need the model half
Metric
Owner
Role
Why this one
cell-exact
Cell-exact match against the computed key
alarm
escalation_when_a_rule_fires_pct; false_alarm_rate_pct; bad_credit_holds; counter_reset_accuracy_pct — alarm on bad_credit_holds. Every other rung is somebody picking up a phone; this one is a proposal to review whether a paying customer's service should be suspended, and it is wrong most expensively on exactly the rows rules 7 and 8 exist to protect.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
732,034
watch snapshots edited — the count held, the bytes did not
split.count
150
the watch snapshots count moved — a different set was scored
split.size_p50
4,880
the median size of one watch snapshot moved
split.size_p95
4,880
the 95th-percentile size of one watch snapshot moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (stateless False) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
band accuracy
not yet known
150 observations
A band is the spread between repeats of the SAME arm at the same settings, and there has been no repeat: every model figure on these pages comes from one run. r001, s001, a001 and a002 differ by prompt or by schedule, so differencing them measures those changes and not this instrument. ⚠︎ AND THIS ONE IS NEARLY AT CEILING ANYWAY: the scored run is 99.33 pct and free code is 100.0, so the 0.67-point gap is a single cell and should be read as no difference.
escalation rung accuracy
not yet known
150 observations; 99 of them where a rung fires
A band is the spread between repeats of the SAME arm at the same settings, and there has been no repeat: every model figure on these pages comes from one run. r001, s001, a001 and a002 differ by prompt or by schedule, so differencing them measures those changes and not this instrument.
consecutive-run count on the counter-reset subset
not yet known
44 observations where the ageing inference disagrees with the truth
A band is the spread between repeats of the SAME arm at the same settings, and there has been no repeat: every model figure on these pages comes from one run. r001, s001, a001 and a002 differ by prompt or by schedule, so differencing them measures those changes and not this instrument.
top driver accuracy
not yet known
150 observations
A band is the spread between repeats of the SAME arm at the same settings, and there has been no repeat: every model figure on these pages comes from one run. r001, s001, a001 and a002 differ by prompt or by schedule, so differencing them measures those changes and not this instrument. ⚠︎ DO NOT READ THE SHIPPED-CORPUS AND REWRITTEN-PROSE FIGURES AS A BAND EITHER (68.0 pct against 76.67): those are two different question sets, not two readings of one.
false alarms on the quiet observations
not yet known
51 observations whose correct rung is NONE
A band is the spread between repeats of the SAME arm at the same settings, and there has been no repeat: every model figure on these pages comes from one run. r001, s001, a001 and a002 differ by prompt or by schedule, so differencing them measures those changes and not this instrument. ⚠︎ AND NO BAND FINER THAN ONE ROW OF THE NEGATIVE CLASS IS MEASURABLE HERE: 51 quiet observations means one false alarm is 1.96 pct, so any band tighter than that would be alarming on arithmetic.
unsupported credit-hold recommendations
not yet known
150 observations
A band is the spread between repeats of the SAME arm at the same settings, and there has been no repeat: every model figure on these pages comes from one run. r001, s001, a001 and a002 differ by prompt or by schedule, so differencing them measures those changes and not this instrument. ⚠︎ THIS IS THE METRIC AN OPERATOR SHOULD WATCH and the one with no band, which is the honest state and not a reassuring one: the scored run made 3, free rules with the same memory made 0, and free rules with none made 25.
end-to-end latency
14,901 to 50,865 ms (p50 to p95)
149 answered observations on r001-ptp-watch
MEASURED, and unlike the accuracy rows this one does not need a repeat: it is the spread WITHIN the single run, not between runs. Full range 5,417 to 83,653 ms -- a 15.4x tail, driven by how many of the nine rules interact on a given arrangement rather than by page length, since every snapshot is within a few hundred bytes of every other.
output tokens
1,745 to 5,848 tokens an observation (p50 to p95)
149 answered observations on r001-ptp-watch
MEASURED within the run: 534 to 9,770 tokens, a 18.3x spread, on a corpus whose snapshots are all nearly the same size. 95.9 pct of the output is provider-side reasoning, so this band is mostly a band on how hard the arithmetic was.
input tokens
not yet known
149 answered observations on r001-ptp-watch
⚠︎ NOT FOR THE USUAL REASON. This one is unknown because evals/run.py records input tokens as a RUN TOTAL and never per observation, so there is no distribution to take a spread from -- only a mean of 1794.94. The corpus makes a narrow spread likely (every snapshot is within a few hundred bytes and the collections policy is identical on all of them), and 'likely' is not a measurement. Recording per-call input would close this for $0.00 on the next run.
the outstanding count, not just the band
0 — the scored run is at 99.33 pct with a mean absolute error of 0.0 and a worst case of 0
149 observations scored
Banding a cell right and counting it right are different abilities and only one of them is visible in band accuracy. The one-snapshot floor and the stub both sit at 70.67 pct with an MAE of 0.6 and a worst case of 4 — they band well and cannot count — while the ledger-memory floor reaches 100.00 pct, so the count is what the carried ledger buys. Removing memory costs 28.66 points: the stateless arm is 70.67 pct, MAE 0.44, worst case 2.
driver attribution where something IS outstanding
79.69 pct on the scored run, and the free floors beat it — this band is a loss, printed
128 cells with something outstanding, out of 150
The one-snapshot floor and the stub both reach 97.66 pct here against the scored run's 79.69, and removing memory barely moves it (78.91 pct stateless). A floor that wins by 17.97 points on a sub-population is the honest reading of this metric: naming the largest open driver is close to a lookup on these rows, and the model is spending its attention on the count and the escalation rung instead. It is banded rather than dropped because a reader comparing kits should see where the arithmetic is enough.
HistoryRun history
10 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 2 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · with the model in the path — 9 runs. Columns here are only ever compared with each other.
Metric
a000-ptp-watch-paraphrase-free 2026-08-23
a001-ptp-watch-paraphrase 2026-08-23
a002-ptp-watch-skiprun 2026-08-23
a002f-ptp-watch-skiprun-free 2026-08-23
b000-ptp-watch-onesnapshot 2026-08-23
b001-ptp-watch-ledgermem 2026-08-23
c000-ptp-watch-calibration 2026-08-23
r001-ptp-watch 2026-08-23
s001-ptp-watch-stateless 2026-08-23
band accuracy, %
83.33
100.00
98.00
100.00
100.00
100.00
100.00
99.33
100.00
count accuracy, %
83.33
100.00
98.00
100.00
70.67
100.00
100.00
99.33
70.67
count mae
0.27
0.00
0.00
0.00
0.60
0.00
0.00
0.00
0.44
count worst error
2
0
0
0
4
0
0
0
2
counter reset accuracy, %
100.00
100.00
100.00
100.00
0.00
100.00
100.00
97.73
40.91
driver when something outstanding, %
16.67
76.67
72.09
100.00
97.66
97.66
66.67
79.69
78.91
escalation accuracy, %
63.33
100.00
92.00
100.00
80.00
100.00
77.78
94.00
84.67
escalation when a rule fires, %
72.73
100.00
90.32
100.00
69.70
100.00
71.43
90.91
76.77
false alarm rate, %
62.5
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
input tokens, whole run
—
53590
88876
—
—
—
16096
267446
260746
model latency p50 ms
0.00
15173.00
14554.00
0.00
0.00
0.00
18615.00
14901.00
17968.00
model latency p95 ms
0.00
26449.00
67411.00
0.00
0.00
0.00
84035.00
50865.00
65831.00
output tokens, whole run
—
56873
145215
—
—
—
32080
347975
425696
top driver accuracy, %
16.67
76.67
62.00
100.00
98.00
98.00
66.67
68.00
67.33
not a time series No two of these 9 runs measured the same system — they differ on counter_reset_units, documents, driver_open_cells, escalations_that_fire, max_tokens, provider, stateless, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-ptp-watch-stub 2026-08-23
band accuracy, %
100.0
count accuracy, %
70.67
count mae
0.6
count worst error
4
counter reset accuracy, %
0.0
driver when something outstanding, %
97.66
escalation accuracy, %
80.0
escalation when a rule fires, %
69.7
false alarm rate, %
0.0
input tokens, whole run
277915
model latency p50 ms
0.00
model latency p95 ms
0.00
output tokens, whole run
10300
top driver accuracy, %
98.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 14 chips that all say so.
DeviationsWhat deviated
0 breaches across 10 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the five-scalar state paragraph
consecutive count 70.67 pct -> 99.33 pct, counter-reset subset 40.91 -> 97.73, rung 84.67 -> 94.0, memory-dependent count 37.5 -> 97.92, cost per observation $0.009383 -> $0.007904 (CHEAPER with the state). The BAND barely moves: 100.0 -> 99.33.
measured
r001-ptp-watch against s001-ptp-watch-stateless. Same corpus, same model, same grader, same 150 observations; the prompts differ in exactly one block and evals/check_labels.py asserts they are byte-identical everywhere else.
whether the collector's notes use phrasings the free keyword table was written for
free floor band 100.0 pct -> 83.33, rung 100.0 -> 63.33, driver 98.0 -> 16.67, unsupported credit holds 0 -> 4. The MODEL on the same rewritten set: band 100.0, rung 100.0, driver 76.67.
measured
a000-ptp-watch-paraphrase-free against b001-ptp-watch-ledgermem, and a001-ptp-watch-paraphrase against both, over the 10 prose-dependent arrangements. ⚠︎ 30 observations is a small denominator and a difference of a few points on it is not a finding.
whether the scheduled run happens
150 observations become 100; 9 surviving observations land on a LOWER rung; 24 rungs are never reached and 10 arrive a run late; 9 CREDIT_HOLD_REVIEW recommendations never happen; 12 ledger events appear on no later snapshot
measured
results/cadence-cost.json (evals/cadence.py, 0 calls) diffs the two answer keys; a002-ptp-watch-skiprun re-fires the model on the gapped schedule and is scored against THAT schedule's key, because the correct answer moved when the schedule did.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
band accuracy
nothing yet
escalation rung accuracy
nothing yet
consecutive-run count on the counter-reset subset
nothing yet
top driver accuracy
nothing yet
false alarms on the quiet observations
nothing yet
unsupported credit-hold recommendations
nothing yet
end-to-end latency
nothing automatically. A p95 above the ceiling's own timeout would start losing observations, and src/adapters retries a transport failure rather than recording a lost one -- but nothing in the kit watches this number.
output tokens
nothing automatically. This is the number that sets the bill -- output is the larger half of every projected dollar on the Cost page.
input tokens
nothing yet
the outstanding count, not just the band
a worst case above 0, or MAE above 0.0. The worst case moves first — accuracy can stay high while one cell is out by four.
driver attribution where something IS outstanding
below 79.69, the figure the scored run actually reached. This band is not aspirational and must not be re-read as one — the target here is free code, and free code is ahead.
NextThe three you would add first
A human step in front of anything that acts on CREDIT_HOLD_REVIEWIt is the highest-consequence output this kit has -- a proposal to review whether a paying customer's service should be suspended -- and it is wrong most expensively on exactly the rows rules 7 and 8 exist to protect. The scored run made 3 that the policy does not support; the memoryless floor made 25.
An alert when a scheduled run does not fireThis is the traceless failure of a monitor: no error, no wrong answer, no artifact. Measured for $0.00, one missed weekly run of three loses 9 credit-hold recommendations and hides 12 ledger events for good. src/state.py carries last_run_date so a gap is VISIBLE in the transcript, and nothing in the kit can make a missed run happen.
A rule engine, before a modelThis kit's own scoreboard says free Python carrying the same five scalars is exact on band, rung and count. Buy compute for the collector's prose, which is the one axis where it measurably wins, and not for the arithmetic.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
The guardrail is checked by evals/check_labels.py, which runs before any spend, and by reading src/watch.py's module docstring, which states it where the calls are made rather than only on a page.
What this cannot tell you
Nothing here has measured what a downstream consumer would do with a CREDIT_HOLD_REVIEW string. The guarantee is that this kit does not act; it is not a guarantee about whatever a forker wires to its output.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures from the standard library. requirements.txt names nothing and says in a comment that the emptiness is load-bearing: a line there means something under src/ or evals/ imports it.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior turns)
those carry a TRANSCRIPT; this carries a LEDGER LINE -- five scalars written by arithmetic. The cost is flat in history length because nothing here remembers what was said, only what was counted, and a memory layer would give back the growth this design exists to avoid. It would also give back the failure this design exists to avoid: a counter fed from its own model output only ever climbs.
the model call
src/adapters/__init__.py
LLM client wrappers and provider routers (LangChain, LiteLLM)
one completion call in one shape is small enough that stdlib is honestly the right size, and it keeps the fork test to one install
the schedule
evals/run.py and src/promise.RUN_INTERVAL_DAYS
workflow and scheduling engines (Airflow, Temporal, Prefect, plain cron)
this is the seam where a framework genuinely earns its place, and it is the ONLY one. A weekly watch over a real book needs a scheduler that knows a run did not happen -- evals/cadence.py measures what that costs and the answer is not small. What a scheduler must NOT be allowed to do is retry a missed run as though it had happened on time: the rungs count observations, so a late run and an on-time run are different answers to the same question.
the state store
src/state.py
any database
data/state.json is one file replaced atomically. Correct for one writer; not a concurrency model. This is the first ceiling a real book hits.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each arrangement is a chain of 3 observations -- run 1, run 2, run 3 -- with no branching and exactly one edge between consecutive runs, carrying five scalars. Different arrangements never touch. What IS a real structure is the other axis: one scheduled run fans out across the whole open book at once, which is a for-loop 50 wide, not an orchestration.
The other sideWhat a framework costs you
No scheduler, and on a cadence kit that is the biggest thing missing. This monitor is INVOKED, not woken: the clock, the retry policy and the question of what to do when a run does not happen are all outside the kit. evals/cadence.py measures what a missed run costs ($0.00) and can do nothing to prevent one.
No state store worth the name. data/state.json is one file replaced atomically; a framework would bring a checkpointer with a concurrency model, which this does not have -- and on a counter that only climbs that gap is worse than on a per-request kit.
No observability beyond what evals/run.py prints and writes to results/. No tracing, no dashboard integration, no alert on a run that did not fire.
No retry or backoff beyond src/adapters' own bounded retry, which covers a busy provider and a dropped connection and nothing else. It did not cover the one reply this run lost, which returned HTTP 200 with an empty body.
What we could NOT verify
No framework was benchmarked against this kit. The claim is that stdlib was sufficient here, not that any named framework is slower or worse.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-ptp-watch on the same tier, memory removed (THE CONTROL), 2026-08-23. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
14,901 ms
14,901 to 50,865 ms (p50 to p95)
nothing automatically. A p95 above the ceiling's own timeout would start losing observations, and src/adapters retries a transport failure rather than recording a lost one -- but nothing in the kit watches this number.
Model, p95
50,865 ms
14,901 to 50,865 ms (p50 to p95)
nothing automatically. A p95 above the ceiling's own timeout would start losing observations, and src/adapters retries a transport failure rather than recording a lost one -- but nothing in the kit watches this number.
Input tokens
267,446
not yet known
nothing yet
Output tokens
347,975
1,745 to 5,848 tokens an observation (p50 to p95)
nothing automatically. This is the number that sets the bill -- output is the larger half of every projected dollar on the Cost page.
No movement column. Not one of the 8 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
a001-ptp-watch-paraphrase15,173 ms
a002-ptp-watch-skiprun14,554 ms
c000-ptp-watch-calibration18,615 ms
r001-ptp-watch14,901 ms
s001-ptp-watch-stateless17,968 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
4 runs not plotted. a000-ptp-watch-paraphrase-free, a002f-ptp-watch-skiprun-free, b000-ptp-watch-onesnapshot, b001-ptp-watch-ledgermem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Decide what to do about a client's broken payment promise
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-23, across 10 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
watch snapshots
data/corpus/<PTP-####>-R<n>.txt -- 150 files, 732,034 bytes, generated once from a fixed seed by tools/build_corpus.py
six of the seven sections go to the provider in the prompt; Account Contact -- a named AP clerk, their work email, their direct line and a billing address -- never does, by src/select.NEVER_SENT
the answer key
data/gold.jsonl -- 150 rows, the output of src/promise.replay over the planted inputs, never hand-authored
never -- evals/scoring.py is pure code, no model and no key
the carried state
data/state.json in a deployment (src/state.py, written atomically). ⚠︎ IN AN EVAL IT IS SCOPED TO THE RUN AND NEVER TOUCHES DISK -- evals/run.py builds a fresh in-memory store per arrangement, because a run that started from the previous run's memory could not be re-run or compared with its own control
one short paragraph per call, produced by state.describe -- five scalar fields, never a prior snapshot and never a prior reply
the recorded runs
results/eval-*.json -- the scored run, the stateless control, two free floors, the prose probe's two arms, the cadence ablation's two arms, the wiring stub and the calibration
never -- they are written locally and committed
the provider key
.env at the repo root or beside the kit, gitignored, read by src/config.py
only as an Authorization header on the one HTTPS request per observation
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 56
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged. config.sources() reports WHICH file a value came from and never the value. config.save() opens the file 0600 before the write rather than chmod-ing it after, so the key never exists at the default umask even for an instant. src/app.py redacts the key and the base URL out of any provider error before it reaches the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
RUN_INTERVAL_DAYS = 7. One scheduled run every Monday at 07:00 UTC, over EVERY open arrangement in the book -- 50 here, and the whole population is re-read each time rather than a delta being applied. What ONE run owns that the last did not: seven days of cash application, disputes, renegotiations and collector notes, and exactly one increment of the consecutive-broken counter. The rungs count OBSERVATIONS, not days, which is why a run cannot be made up later.
Skipping ONE run of three: 150 observations become 100; 9 surviving observations land on a LOWER rung than they should; 24 rungs are never reached at all and 10 arrive a run late; 9 CREDIT_HOLD_REVIEW recommendations never happen; and 12 ledger events -- renegotiations and payments dated inside the skipped week -- appear on NO later snapshot, because the Activity Log window is 3 days and the gap would be 14. (results/cadence-cost.json (evals/cadence.py, 0 calls), and the model arm a002-ptp-watch-skiprun against r001-ptp-watch)
three scheduled runs. Long enough to reach the third rung and to put a renegotiation between two observations; not long enough to test a book where an arrangement has been on the watch for a year. The two windows are deliberately misaligned -- a 3-day activity log against a 7-day cadence -- and that gap is the whole reason the carried credited figure exists.
change the cadence and every escalation figure on this page is measuring a different question, because the ladder counts runs. A fortnightly watch is not this kit run half as often; it is a different policy, and the honest move is to rewrite the rungs in days before changing the clock.
state
five scalars per arrangement -- the previous run's date, the band it reported, the rung it recommended, the consecutive-broken count and the cash credited as at that run -- written by src/promise.step from the figures PARSED off the page and never from the model's reply, and rendered by src/state.describe into one short paragraph. That paragraph is the entire route from one week to the next, and it is what the two scored arms differ by.
323 characters on the worked example. Band 99.33 pct with it against 100.0 pct without; rung 94.0 pct against 84.67 pct; consecutive count 99.33 pct against 70.67 pct. Input tokens 267,446 with it against 260,746 without, and output 347,975 against 425,696. (r001-ptp-watch against s001-ptp-watch-stateless, lenses.LLM.prompt_parts)
five scalars, so run 40 costs what run 3 costs -- the OPPOSITE curve to an intake kit, whose input grows with the square of the turns. What is NOT bounded is the store: data/state.json is one file replaced atomically, correct for one writer and not a concurrency model.
remove the carried state and the kit is an ordinary one-shot classifier over a single snapshot -- which is exactly what the b000 floor is, for $0.00. ⚠︎ AND THE FREE MEMORY FLOOR HAS THE SAME FIVE SCALARS AND IS EXACT, so the state is load-bearing for the TASK and not for the MODEL.
model
one call per OBSERVATION carrying the instruction, the carried-state paragraph and six of the snapshot's seven sections, behind src/adapters/__init__.py, at the published MAX_TOKENS = 20,000 with provider-side reasoning left at the default -- the configuration every scored arm ships.
1 of 150 replies failed to parse on the scored run; the largest reply that DID parse was 9,770 output tokens (48.9 pct of the cap). p50 14,901 ms, p95 50,865 ms. Of those: 1 returned an EMPTY body with finish_reason 'stop' after billing 7,427 output tokens -- the whole budget went to provider-side reasoning that never reached the text, at 37.1 pct of the cap, so RAISING THE CEILING WOULD NOT HAVE SAVED IT. (lenses.Eval.scores, r001-ptp-watch and c000-ptp-watch-calibration)
the 20,000-token ceiling was set at 1.9x the largest reply the calibration had seen (10,684 on 9 calls at a 32,000 cap), deliberately high because a ceiling is not billed.
one tier was run, so nothing here compares two models. The comparisons this kit does make are memory against no memory, and a model against free code. Point .env at your own model and the free scorer re-runs on your numbers.
labels
data/gold.jsonl, 150 rows, computed by src/promise.replay over the planted inputs at generation time -- a function the runtime never calls, so the key is arithmetic done independently of the code under test rather than the code under test marking its own homework.
{"AT_RISK": 15, "BROKEN": 84, "KEPT": 22, "RENEGOTIATED": 11, "STANDING": 18} over 150 observations; 99 rungs fire and 51 are quiet; 48 observations are memory-dependent, 44 carry a counter the page's own ageing inference gets wrong, 12 are held by the good-faith rule, and 14 carry a renegotiation or a dispute recorded only in a collector's sentence. (lenses.Eval.dataset, ptp-watch-2026-08-23-50arrangements-150observations)
⚠︎ THE KEY IS ONLY AS GOOD AS AN INVENTED POLICY. The cadence, the five rungs, the reminder lead, the half-the-amount test, LARGE_BALANCE_USD, the dispute cap and the good-faith hold were all written for this kit and confirmed against nothing. And the collector notes come from a fixed pool the free keyword table was written against -- evals/paraphrase.py prices that at 16.67 pct against 98.0 pct on the driver.
your own book: hand-label the gold, which is the real work. This kit's key is a luxury of controlling both the generator and the rules, and a hand-labelled key has an error rate nothing here has measured.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a CREDIT_HOLD_REVIEW recommended on an arrangement whose Cash Position shows more credited than it did last week
something read the snapshot without the carried credited figure and could not apply the good-faith hold in rule 8 -- the payment landed mid-week and appears in no dated entry, because the activity window is 3 days and the cadence is 7
check the carried state before disputing the answer. On PTP-0045 at run 3 the memoryless floor recommends CREDIT HOLD REVIEW at an inferred 5 consecutive runs; the truth is 3 and the rung is COLLECTOR_CALL. (results/eval-b000-ptp-watch-onesnapshot.json)
a consecutive-run count that equals the days past the promised date divided by seven, plus one
the ageing inference was used instead of a carried counter. It is wrong on 44 of 150 observations here -- every arrangement the watch first saw when it was already weeks overdue, and every one whose count was returned to zero by a renegotiation.
compare against the previous run's recorded count. There is no arithmetic that recovers it from the page. (results/eval-b000-ptp-watch-onesnapshot.json)
a rung that climbed while the ledger's dispute flag was set
rule 7's cap was not applied, or the dispute exists only in the collector's note and whatever read the page did not read the sentence
read the Collector Notes section for that run. 9 observations in this corpus carry a dispute that the ledger flag does not. (data/corpus-stats.json)
No machine symptom — this failure leaves no trace in any output.
the scheduler, not the kit. src/state.py carries last_run_date precisely so that a gap is VISIBLE in the transcript, and state.describe() prints the date out loud -- but nothing in this kit can make a missed run happen. That is a deployment control and evals/cadence.py exists to price it.
Concurrency. data/state.json is one file replaced atomically and no run has ever had two writers. Nothing here measures what two workers on one book would do to it. Provider-side retention. The prompt carries no personal data by construction, but what a provider keeps and for how long is a contract term nothing in this kit can check. GPU sizing and self-hosting. Every figure here is against a hosted, OpenAI-compatible endpoint. Nothing has been run locally. Whether the flow variant's station labels fit. monitor was added to the standard on 2026-08-23 and this is among the first kits to use it; every station reads correctly for this kit, which is recorded here so that a later kit for which one does not can say so. Whether English or JSON is the better rendering for the carried state. One rendering was run. Whether a real collections book's notes resemble either the shipped pool or the rewritten one. Both are invented; the probe measures sensitivity to phrasing, not realism.
The corpus licence, from the Data lens: MIT, the same as the repository. Nothing in data/ is derived from anyone else's work, so there is nothing here to attribute. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
a reader of this snapshot alone infers 2 consecutive runs from 9 days past the date and a 7-day cadence; the truth is 2
Grader
Verdict
Why
Cell-exact match against the computed key
band HIT (key BROKEN, answered BROKEN) · rung MISS (key COLLECTOR_CALL, answered ACCOUNT_MANAGER) · count HIT (key 2, answered 2) · driver HIT (key DISPUTE_OPEN, answered DISPUTE_OPEN)
PTP-0034-R2, run R2 of the dispute_cap pattern. The page showed 9 days past the promised date and 0.00 USD credited of 8,861.97 promised. THE FOUR CUTS THIS GRADER REPORTS, on this one observation: (1) per-cell exactness, above. (2) The counter-reset cut -- this is one of the 44 observations where the page's own ageing arithmetic disagrees with the truth: a reader with only this snapshot infers 2 consecutive runs from 9 days past the date and a 7-day cadence, and the watch has seen it 2 time(s). (3) The driver cut -- the key says DISPUTE_OPEN. This is the field the model loses to free code on the shipped corpus (68.0 pct against 98.0) and wins on once the notes are rewritten (76.67 against 16.67). (4) The costly-error cut -- unsupported CREDIT_HOLD_REVIEW recommendations: not one of them here, and across the run 3 from the model, 0 from free rules with the same memory, 25 from free rules with none. On this observation rule 8 holds the rung at COLLECTOR_CALL because cash was credited since the previous run.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
no headline metric on this row — it records band accuracy 0.99 · escalation accuracy 0.94 · consecutive count accuracy 0.99 · top driver accuracy 0.68
the fast tier, memory removed (THE CONTROL)
no headline metric on this row — it records band accuracy 1 · escalation accuracy 0.85 · consecutive count accuracy 0.71 · top driver accuracy 0.67
free rules carrying the same five scalars
no headline metric on this row — it records band accuracy 1 · escalation accuracy 1 · consecutive count accuracy 1 · top driver accuracy 0.98
free rules, one snapshot, no memory
no headline metric on this row — it records band accuracy 1 · escalation accuracy 0.8 · consecutive count accuracy 0.71 · top driver accuracy 0.98
the fast tier, collector notes REWRITTEN
no headline metric on this row — it records band accuracy 1 · escalation accuracy 1 · consecutive count accuracy 1 · top driver accuracy 0.77
free rules + same state, notes REWRITTEN
no headline metric on this row — it records band accuracy 0.83 · escalation accuracy 0.63 · consecutive count accuracy 0.83 · top driver accuracy 0.17
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by src/promise.replay over the planted inputs -- never hand-authored.
These rates are UNKNOWN, on purpose
It cannot be scored against itself. This grader IS the reference standard here: the key is computed by src/promise.replay and equality against it is the definition of correct, so a TPR/TNR for it would be circular.
Watch these
escalation_when_a_rule_fires_pct
false_alarm_rate_pct
bad_credit_holds
counter_reset_accuracy_pct
Alarm on
bad_credit_holds. Every other rung is somebody picking up a phone; this one is a proposal to review whether a paying customer's service should be suspended, and it is wrong most expensively on exactly the rows rules 7 and 8 exist to protect.
How tight can the band be? The alarm band cannot be finer than one row of the negative class. There are 51 quiet observations, so one false alarm is 1.96 pct and no band tighter than that is measurable on this set.
Cadence: re-run on every corpus change and on every model change. It is free, so there is no reason to run it less often than that.
The decisionWhen to reach for it
Use it
always. Every field here is a closed set or a whole number, so equality IS the grade and there is nothing for a judge to add.
Do not use it
the moment a field becomes free text. The rationale the model returns is not scored at all by this grader, and nothing in this kit scores it.
A living map of modern AI — kept current every morning