The business caseThe problem this solves
An invoice fell due and was not paid, and nobody has written down why. A collections desk has the open-item row, the settlement lines behind it, every evidence record on file and an account notes block the desks have written to each other, and has to say which of nine causes it is and which records prove it — then hand it to the right desk before it ages. Today that is a person reading every record against the 270-day window, the subscription and the FINAL test, then reading every note for a closure and joining each note's own date to the record it names. reading one overdue invoice end to end — the subscription, 270-day window and FINAL tests on every evidence record, every account note read for a closure and joined to the record's own date, and the predecessor rule that makes a note dated before the record it names inert — before the nine-rung ladder is applied and the desk chosen.
Audience
A collections or AR operations desk coding overdue invoices on a subscription ledger, and the manager deciding whether a model call is worth buying for it. This report's own answer for this corpus is NO — the call is significantly WORSE than free code — and the numbers are laid out so that answer can be checked rather than taken. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual overdue invoices
The corpus is 64 overdue invoices, 0.10 MB (txt 64). 64 invoices = 32 fact patterns x 2 mirror twins, generated from seed 20260918, because the question is whether reading an account note beats not reading one and that needs a measured ratio of both. 38 invoices have a note that really closes a record and 26 have one that names a record and closes nothing; 24 carry a second true cause, 28 a date trap, 26 a near-miss subscription, 34 a window edge, 24 an aged review and 16 a cap probe. Every invoice carries three or four notes, one of them decoy-shaped and one filler, asserted in code. And the decoy pool is MEASURED: a candidate qualifies only if removing the record it names actually moves that invoice's answer, and it prefers an out-mover on 52 of 64. Before that fix a naive word list scored 34 of 64 against a 26 floor — a word list BEATING the rule that reads no note, because a word rule only ever CLOSES records; after it, 12.
The corpus
- The 64 overdue invoicesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your overdue invoices. That is the whole change — there is no database to migrate.
OVERDUE INVOICE OIV-0001
LEDGER EXTRACTED 2026-07-10 STANDARD OICS-2026
ACCOUNT ACT-16653 SUBSCRIPTION SUB-673 PLAN PLN-4
INVOICE DATE 2026-02-15 DUE DATE 2026-04-01 AS OF 2026-06-30
ORIGINAL AMOUNT 9445.84 OPEN AMOUNT 8396.31
DECLARED SCOPE EVIDENCE WINDOW 270 DAYS AGEING BUCKETS 30/60/90/120
[1] SETTLEMENT ACTIVITY
LINE ENTRY POSTED AMOUNT
1 ENT-59757 2026-04-10 1049.53
[2] EVIDENCE ON FILE
ID KIND SUBSCRIPTION DATED STATUS REFERENCE AMOUNT
EVD-30001 SLA-CLAIM SUB-673 2026-07-06 FINAL SLA-90552 3390.80
EVD-30002 PO-AUTHORITY SUB-673 2026-02-06 FINAL POA-61812 8431.74
EVD-30003 UPLIFT-QUERY SUB-763 2026-05-15 FINAL QRY-54623 32.28
EVD-30004 CREDIT-COMMIT SUB-673 2026-03-04 FINAL CRQ-20073 35.43
[3] NOTICES SENT
LINE NOTICE SENT CHANNEL
1 NTC-50201 2026-06-25 post
[4] ACCOUNT NOTES
Billing ops: the evidence index for this invoice was regenerated on 2026-07-10; no evidence content changed.
Dispute desk: a query on 2026-05-26 asked whether EVD-30004 still stands for SUB-673; it does not stand - EVD-30004 was closed on 2026-05-26 because the line behind it was rebilled.
Dispute desk: a query on 2026-07-06 asked whether EVD-30001 still stands for SUB-673; it was never closed - EVD-30001 was only rechecked on 2026-07-06 because the line behind it was rebilled.
The outcomeWhat a good result looks like
One invoice in, one coded invoice out: the PRIMARY cause, one of the nine rungs of OICS-2026; every lower rung that also fires, as a ranked secondary set; and the exact set of qualifying evidence record ids. Then — in code, for every arm alike — the days past due, the ageing bucket, the aged-review flag, the owning desk, the route and the records in front of the desk. It never contacts a customer, never sends, holds or releases a notice, never changes an account's treatment, stage, terms or balance, and emits no verdict, no total and no pay-by date.
And when it cannot
And when nothing on file settles it: no-cause-found, with one short unresolved line saying what is missing, no cause asserted and nothing cited. 8 of the 64 invoices are residual in the key; the paid call reaches 0 of them — it names a cause on all 8 — and a free constant reaches all 8, which is why the paid arm's residual figures may never be quoted as a result.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You are coding overdue invoices on a subscription ledger and want a defensible answer today. — the free columns bar, and do not buy a call at all
26 of 64 invoices whole for $0.00 against the paid call's 7, exact McNemar p = 0.0008779. This is not a close call and the kit does not dress it as one. - Your causes really do live in free-text account notes rather than in a dispute flag. — measure a constant and a word list on YOUR corpus first — both are free and both are in this kit
on the 38 invoices here where a note closes a record, the paid call takes 6 and a fixed residual reply and a word list each take 4: p = 0.7539062. The reading win is +2 invoices and not separable. - You need the ageing bucket, the owning desk or the aged-review flag. — src/policy.py::rollup, for nothing
the bucket is 64 of 64 for every arm — a subtraction of two printed dates — and the desk is a lookup on the primary. None of it is ever bought. - Your hard invoices are the ones with two true causes, where the PRIMARY names the desk. — free code, and treat the call as no opinion at all
on the 24 two-cause invoices free code is 24 of 24 on the primary and the call is 10, p = 0.0001221 against. The kit's single most common named failure isdemoted_the_primary_to_a_secondary— 26 of 64 — it finds the right cause and files it second. - You want a defensible number for a buying decision on your own corpus. — clone, replace data/corpus and data/gold.jsonl, run
evals.baseline --allandevals.check_labels, then decide
every floor, every grader and every paired test is pure code and costs $0.00, so the comparison is reproducible before a single call is bought.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own invoices in the same plain-text shape — open-item row, settlement lines, evidence register, dunning notices, account notes — and write data/gold.jsonl with one {invoice_id, cause, secondary, cited} per invoice. EVERY QUALITY FIGURE ON THIS PAGE STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per invoice for a reading that is significantly worse than a rule which reads no note. That is the case against the best-fitting scenario (“You are coding overdue invoices on a subscription ledger and want a defensible answer today.”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | AN INVOICE WHOSE CAUSE IS A PRINTED FLAG rather than a sentence in an account note. The whole contest here is reading notes; on the 26 invoices where a note names a record but closes nothing, free code that reads no note is 26 of 26 and the call is 1. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NOTHING. THE CALL DOES NOT BEAT FREE CODE AND THE TEST SAYS SO — 7 of 64 whole against 26 for a rule that reads no note, 6 discordant to 25, p = 0.0008779 on 64 invoices. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-18 — r001-overdue-cause. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board and scores all eight free arms offline, and the scored run replays from the committed results/ rather than being re-bought; --selfcheck re-derives 576 of 576 arm-invoices and exits 0. Measured on a scratch copy with results/cache-*.jsonl deleted: the corpus panels still read, the paid arm's single-invoice view NAMES the missing file rather than rendering dashes, and --selfcheck exits 1 reporting 512 of 576. That is why the reply caches ship — the repository root already declares kits/*/results/ checked in, and this kit adds no .gitignore of its own.

















