The business caseThe problem this solves
A credit lands in a suspense account and nobody knows why. A settlement desk has the item's printed header, the postings behind it, the evidence records on file and a notes block the desks have written to each other, and has to say which of nine causes it is and which records prove it — then hand it to the right queue before it ages. Today that is a person reading every record against the window, the corridor and the FINAL test, then reading every note for a void and joining each note's own date to the record it names. reading one suspense item end to end — the window, corridor and FINAL tests on every record, every note read for a void and joined to the record's own date and, for a standing instruction, the AS OF date — before the nine-rung ladder is applied and the queue chosen.
Audience
A payment operations or settlement reconciliation desk coding suspense items, and the manager deciding whether a model call is worth buying for it. This report's own answer for this corpus is NO — the call does not separate from free code — and the numbers are laid out so that answer can be checked rather than taken. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual settlement suspense items
The corpus is 64 settlement suspense items, 0.09 MB (txt 64). 64 items = 32 fact patterns x 2 mirror twins, generated from seed 20260917, because the question is whether reading a note beats not reading one and that needs a measured ratio of both. 42 items have a note that really voids a record and 22 have none that does; 24 carry a predecessor-dated note, 22 an inert void, 18 a near-miss corridor, 32 a window edge and 16 a cap probe. Every item carries exactly one void-shaped note, one decoy and one filler, asserted in code — an earlier build leaked the answer through the note COUNT. And the decoy pool is MEASURED: a candidate qualifies only if removing the record it names actually moves that item's answer. Before that fix a naive word list scored 22 of 64; after it, 6. 0 of 256 domain-filter combinations beat reading no note at all, on every metric.
The corpus
- The 64 settlement suspense itemsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your settlement suspense items. That is the whole change — there is no database to migrate.
SETTLEMENT SUSPENSE ITEM SUS-0001
REGISTER EXTRACTED 2026-07-05 STANDARD SSCS-2026
SUSPENSE ACCOUNT SA-919 CORRIDOR COR-168 CHANNEL CHN-3
CREDITED ACCOUNT ACC-99160 VALUE DATE 2026-06-21 AS OF 2026-06-30
AMOUNT IN SUSPENSE 11586.63
DECLARED SCOPE EVIDENCE WINDOW 180 DAYS AGE BANDS 30/60/90/180
[1] LEDGER POSTINGS
LINE ENTRY BOOKED AMOUNT
1 ENT-86442 2026-06-24 8255.63
2 ENT-74878 2026-06-24 3331.00
[2] RECORDS ON FILE
ID KIND CORRIDOR DATED STATUS REFERENCE AMOUNT
EVD-20001 CASH-RECEIPT COR-168 2026-06-22 FINAL PMT-55735 8255.63
EVD-20002 CASH-RECEIPT COR-168 2026-06-22 DRAFT PMT-68240 3331.00
EVD-20003 SETTLE-INSTR COR-168 2026-06-11 DRAFT ORD-90606 8255.63
EVD-20004 ACCT-DIRECTIVE COR-168 2026-05-14 FINAL ACC-99160 8255.63
[3] NOTES
Branch desk: the counterparty was told this was settled - please post the match and release the funds.
Client money desk: the finding of 2026-06-27 that EVD-20001 was rescinded for COR-168 was itself withdrawn; EVD-20001 still stands and nothing replaced it.
Reconciliation desk: the declared evidence window and age bands are reprinted on the scope line above.
Client money desk: the finding of 2026-02-19 that EVD-20004 was rescinded for COR-168 was itself upheld; EVD-20004 no longer stands and nothing replaced it.
The outcomeWhat a good result looks like
One item in, one coded item out: the cause, one of the nine rungs of SSCS-2026, and the exact set of qualifying evidence record ids; then — in code, for every arm alike — the age in days, the age band, the aged-review flag, the queue, the route and the records in front of the desk. It never clears the item and never posts anything.
And when it cannot
And when it cannot: unclassifiable, with one short unresolved line saying what is missing, and no cause named. 12 of the 64 items are unclassifiable in the key; the paid call reaches 3 of them and a free constant reaches 12, which is why the paid arm's 3 may never be quoted as a result.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your evidence records carry a void as a STRUCTURED COLUMN — a status of VOIDED or a superseded-by id — rather than as a sentence in a notes block. — the free columns bar, and do not buy a call at all
22 of 64 items whole for $0.00, and on the 22 items the columns settle it is 22 of 22, where the call is 10. - Your desk notes really do decide the cause, and you want the cause rather than the queue. — measure a constant and a word list on YOUR corpus first — both are free and both are in this kit
on the 42 items here where a note voids a record, the paid call takes 11 and a fixed constant takes 10: p = 1.0000. The reading win is +1 item. - You need the queue, the age band or the aged-review flag. — src/policy.py::rollup, for nothing
the band is 64 of 64 for every arm — it is a subtraction of two printed dates — and the route is a lookup on the cause. None of it is ever bought. - You are coding over-funded or unmatched-receipt items specifically. — the free columns bar, and treat the call as a second opinion at most
the call is 0 of 6 on over-funding against the bar's 6 of 6 (p = 0.03125) and 2 of 8 on unmatched-receipt against 6. - You want a defensible number for a buying decision on your own corpus. — clone, replace data/corpus and data/gold.jsonl, run
evals.baseline --allandevals.check_labels, then decide
every floor, every grader and every paired test is pure code and costs $0.00, so the comparison is reproducible before a single call is bought.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own items in the same plain-text shape — header, postings, evidence register, notes block — and write data/gold.jsonl with one {item_id, cause, cited} per item. EVERY QUALITY FIGURE ON THIS PAGE STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per item for a reading the columns already give you. That is the case against the best-fitting scenario (“Your evidence records carry a void as a STRUCTURED COLUMN — a status of VOIDED or a superseded-by id — rather than as a sentence in a notes block.”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | AN ITEM WHOSE VOID IS A STRUCTURED COLUMN rather than a sentence in a notes block. The whole contest here is reading notes; on the 22 items where nothing in a note voids anything, free code that reads no note is 22 of 22 and the call is 10. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE CALL IS REALLY NO BETTER THAN FREE CODE. It is level: 21 of 64 whole against 22 for a rule that reads no note, 11 discordant each way, p = 1.0000 on 64 items. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-suspense-coding. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board and scores all six free arms offline, and the scored run replays from the committed results/ rather than being re-bought; --selfcheck re-derives 448 of 448 arm-items and exits 0. Measured on a scratch clone with results/cache-*.jsonl deleted: the corpus panels still read, the paid arm's single-item view NAMES the missing file rather than rendering zeros, and --selfcheck exits 1 reporting 384 of 448. That is why the reply caches ship — the repository root already declares kits/*/results/ checked in, and this kit adds no .gitignore of its own.















