The business caseThe problem this solves
A restaurant's tills are counted at close, the cash office bands the money into bags, an armoured carrier collects them and some days later the bank credits the account. Four records have to agree and none of them was written to agree with the others: the drawers, the prep block, the carrier's run sheet and the bank statement. A day that is genuinely short looks exactly like a day whose deposit is still on the road, a bag that was re-banded after a broken seal looks exactly like a second deposit that never arrived, and an adjustment the vault withdrew looks exactly like one that stands. The reconciliation most stores run is a column comparison, and on this corpus the back office's own answer is wrong on 20 of 62 days. Opening one day's file, ticking the prep block against the carrier's run sheet, tracing each bag number through the bank's reference text, reading each cash-office note to see which of them is about this day, and working out whether an uncredited deposit is late or simply inside the carrier's window.
Audience
The cash office desk at a multi-unit restaurant group working a period's deposit exceptions, and the area finance analyst behind it. Whoever reads the output is deciding which days go to the carrier's claims desk, which go to the bank's research desk and which close. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual business day file
The corpus is 62 business day file, 0.16 MB (txt 62). It is generated because it has to be. A real deposit reconciliation is a store's own trading record and a bank statement for a real account, and the exact shapes this kit is about — a bag re-banded after a broken seal, a vault adjustment the carrier withdrew, a sister store's bag mis-referenced to this account — are the ones a group would least want published. Generating it also makes the key DERIVABLE: each day is built as a structure and CDR-2026 is applied to that same structure, so there is no second place the answer lives and a corpus change cannot leave a stale label behind.
The corpus
- The 62 business day filegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every store is an invented code and trading name, every bag number, pickup, bank reference and amount is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a shift is open, mid or close, a drawer is D1, and the only signature on a pickup log is the cash office desk. evals/check_labels.py sweeps all 62 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your business day file. That is the whole change — there is no database to migrate.
==============================================================================
CASH OFFICE DEPOSIT RECONCILIATION -- ONE STORE, ONE BUSINESS DAY
==============================================================================
FILE DEP-0001
STORE 4471 Northgate Crossing
BUSINESS DATE 2026-07-01
PERIOD 2026-07-01 to 2026-07-31
CUT-OFF 2026-07-08
BANK ACCOUNT ****2200 operating, deposits only
CARRIER armoured pickup, contracted settlement window 2 banking days
TOLERANCE 1.00
CURRENCY USD
-- POS CASH DECLARED AT CLOSE, BY DRAWER -------------------------------------
DRAWER SHIFT POS EXPECTED COUNTED OVER/SHORT
D1 open 1,435.00 1,435.15 +0.15
D2 mid 860.00 860.00 0.00
D3 close 970.00 970.00 0.00
TOTAL COUNTED AT CLOSE 3,265.15
LESS FLOATS RETAINED 250.00
DECLARED FOR DEPOSIT 3,015.15
-- CASH OFFICE PREP -- BAGS ON THE PREP BLOCK --------------------------------
BAG SEALED AMOUNT DAY STATUS MEMO
BAG-88013 2026-06-30 22:20 1,800.00 2026-06-30 SEALED prior business day, final close bag
BAG-88000 2026-07-01 21:10 1,507.55 2026-07-01 SEALED drawer skim
BAG-88001 2026-07-01 22:17 1,507.60 2026-07-01 SEALED final close bag
BAG-88011 2026-07-01 19:40 150.00 2026-07-01 VOIDED raised against the wrong business day and cancelled
-- ARMOURED PICKUP LOG -------------------------------------------------------
PU-5000 2026-07-01 06:05 BAG-88013 signed at the cash office desk
PU-5001 2026-07-02 06:05 BAG-88000, BAG-88001 signed at the cash office desk
Abridged — the file continues.
The outcomeWhat a good result looks like
One business day in, one row out: which bags belong to the day, which bank credits settled them with the credit row quoted verbatim, the day's over or short recomputed in code, and one of five CDR-2026 verdicts. A short is a variance with its evidence, routed to the desk that owns it. It never names a person.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed, nothing stopped at the ceiling and no call failed. Where it is wrong it is wrong about a READING: it kept a prep row that a note says was keyed against the wrong business day on 2 of the 5 no-deposit days, and on one day it dropped a bag that was really banked. The station cannot notice either — it re-derives faithfully from whatever reading it is handed and returns a verdict with full confidence.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your bank feed carries per-bag detail and your prep block's business-day column is reliable — the free modal floor, and do not buy a call at all
47 of 62 days for $0.00. Merged credits, split deposits, card-batch decoys, drawer over/shorts and the carrier's settlement window are all decidable from columns and dates, and the floor gets every one of them. - Your cash office records a re-band, a mis-routed bag or a withdrawn adjustment in a NOTE rather than in a column — the paid call, and read the note_decides family's own rate first
The modal floor scores 0 of 15 on those days because nothing in a column says it. The paid call scores 13. - You want the exception record never to name anybody — either arm, and keep src/refusal.py
0 person-shaped names, 0 blame assertions, 0 proposed actions and 0 finality claims on every committed arm AND on all 20 attacked days, including the 5 handed a clause asking in terms for the closing manager to be named.
And where nothing here is good enough:
- Somebody can write a note into the file you send — neither, until you have a control on the reading
A FALSE re-band clause naming two real bags took 5 of 5 attacked days. The kit believes a re-band note, which is exactly what makes it worth buying and exactly what makes it attackable. - Your carrier contract names a bank-holiday calendar — neither, until you add one
src/money.py::banking_days skips Saturday and Sunday and nothing else, so a deposit over a public holiday reads as late when it is not.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own business day files in the same shape and data/stores.json with your own declared figures, tolerances and carrier windows. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per day for arithmetic you already have. That is the case against the best-fitting scenario (“Your bank feed carries per-bag detail and your prep block's business-day column is reliable”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A prep block that is not fixed-width columns. src/deposits.py's row regexes are the shape these files are printed in; a CSV export or a PDF needs its own parser and nothing else changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-deposit-recon. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run. python3 tools/build_corpus.py --check rebuilds all 62 files byte-identically (verified under three PYTHONHASHSEEDs), python3 -m evals.check_labels re-derives the key independently at 0 disagreements, and python3 -m evals.run --run-id b000-deposit-recon-modal --floor modal scores the best free arm. No pip install: the kit is standard library only.






