The business caseThe problem this solves
An audit engagement team sends external confirmations — to a client's bank, its customers and its suppliers — and before a confirmation can go into a workpaper the response has to agree with three other things: the address the team sourced independently, the client's recorded balance at the confirmation date, and the client's own records behind every reconciling item on its schedule. None of those was written to agree with the others. A response can print the sourced address and still have reached the desk through the client; a changed address can be genuine if a call-back verified it before AS AT; a later response can be withdrawn; an outstanding cheque can have cleared before the period end; a receipt can have been reversed. The client's own schedule concludes RECONCILED on 45 of 64 files and is wrong on 40. Opening one confirmation file, checking every logged response's FROM address character for character against the sourced one and its RECEIVED date against AS AT, reading each note for a forwarding, a call-back verification or a withdrawal and its date, matching every reconciling item to the client record it cites by kind, reference, amount and date, signing the supported items, and working out what is left unreconciled — by hand, for every confirmation in the period.
Audience
The confirmation desk and the engagement team on an audit, working a period's external confirmations into the cash, receivables and payables workpapers — and anyone deciding whether a language model should read those files at all. On this corpus the honest answer is no, and the page leads with it. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual confirmation files
The corpus is 64 confirmation files, 0.17 MB (txt 64). It is generated because it has to be. A real confirmation file is an audit client's bank balances, its customers' and suppliers' names and amounts, and desk notes written by named people about named contacts — confidential on every line, and the one thing this kit measures (a sentence about who sent what, and when) is exactly what would be stripped out of anything publishable. So the corpus is built to the job instead: 19 cases, 32 files a sentence decides and 25 carrying a decoy note in the same vocabulary, five date traps, and the client's own schedule wrong on 40 files so the incumbent is measured rather than assumed.
The corpus
- The 64 confirmation filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every engagement, client, confirming party, address, balance and record is invented, every email address sits under the reserved .example domain, and there are no people in the corpus at all — a note speaks for the confirmation desk, the engagement team or one of the client's desks. ECR-2026 is invented too, and no auditing standard is quoted, paraphrased or relied on.
Swap this folder for your own material and the kit is pointed at your confirmation files. That is the whole change — there is no database to migrate.
================================================================================
EXTERNAL CONFIRMATION RECONCILIATION -- ONE REQUEST, ONE AS-AT DATE
================================================================================
FILE ACR-0001
ENGAGEMENT AUD-26-0410 client CL-3100, a regional distributor
CONFIRMATION CNF-41200
TYPE BANK
CONFIRMING PARTY CP-7301 Callowfen Mutual Bank
ACCOUNT BK-10400
CONFIRMATION DATE 2026-06-30
REQUEST SENT 2026-07-01
AS AT 2026-07-21
FILE COMPILED 2026-07-26
CURRENCY USD
-- INDEPENDENTLY SOURCED ADDRESS (the engagement team's, not the client's) -----
SOURCED FROM the bank's published confirmation service page, checked 2026-06-30
CHANNEL email
ADDRESS audit-confirmations@callowfen-mutual-bank.example
-- CLIENT RECORDED BALANCE AT THE CONFIRMATION DATE ----------------------------
LEDGER cash book, bank account BK-10400
RECORDED BALANCE 72,040.99
-- RESPONSES LOGGED BY THE CONFIRMATION DESK -----------------------------------
RESPONSE RECEIVED BALANCE STATED FROM
RSP-1 2026-07-08 65,117.76 audit-confirmations@callowfen-mutual-bank.example
-- CLIENT RECONCILIATION SCHEDULE (as prepared by the client) ------------------
ITEM KIND REFERENCE DATED AMOUNT SUPPORT
RI-1 DEPOSIT IN TRANSIT DEP-80000 2026-06-29 6,923.23 CR-51000
-- CLIENT RECORDS (ledger and bank activity, extracted 2026-07-18) -------------
RECORD KIND REFERENCE DATED AMOUNT
CR-51000 BANK CREDIT DEP-80000 2026-07-01 6,923.23
CR-51001 BANK CREDIT DEP-88000 2026-07-02 2,190.02
Abridged — the file continues.
The outcomeWhat a good result looks like
One confirmation file in, one row out: which responses are authentic at AS AT, which were withdrawn, which reconciling items the client's records support, the unreconciled amount to the cent, every exception and one of five ECR-2026 verdicts — SENDER-MISMATCH, NO-RESPONSE, EXCEPTION, RECONCILED or AGREED — with the response, record, address and date behind each. Never a conclusion, an adjustment, a sign-off or a contact.
And when it cannot
And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 25 files it got wrong are named in the kit README with what it answered: 5 unsupported items kept (an EXCEPTION reported RECONCILED — the costly direction), 15 supported items dropped, 5 printed values compared the wrong way round in its own sentence, 3 withdrawn responses read as not authentic and 1 id list contradicting its own sentence. Confidence does not separate them: median 0.90 on right and wrong files alike.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your confirmation files are mostly decidable from printed columns — addresses, dates, record kinds, amounts — the free modal floor, and do not buy a call at all
On the 32 column-decided files the modal floor gets all 32 for $0.00; the paid call gets 20. - A sentence often decides a reading — forwarded responses, call-backs, withdrawals, reversed or misapplied records — the free vocabulary floor first; the paid call is not ahead of it here
On the 32 files a note decides, vocab gets 20 and the paid call 19. The paid call wins on notes about a RECORD (misapplied support 5 of 5 against the best floor's 3) and loses on notes about a SENDER (5 of 10 against vocab's 7). - You want the client's own reconciliation schedule audited — any arm — every one beats it
The client's printed status is right on 24 of 64 verdicts and returns no reading. It is published as an arm so the comparison is against what is running today rather than against nothing. - The costly error for you is an EXCEPTION reported as clean — vocab or the paid call — they tie at 5, and the modal floor makes it 14 times
The paid call's 5 are all unsupported items on column-decided files that the modal floor gets right; the modal floor's 14 are sentence files it cannot read. A date-and-amount check in code in front of the paid call would remove all 5.
And where nothing here is good enough:
- Your engagement uses a materiality threshold, multi-account bank confirmations or open-invoice receivable confirmations — neither, yet
No file in this corpus does. The unit of work is one balance, one response log and one-item-one-record support, and every percentage here is against that unit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own confirmation files in the same block shape and data/confirmations.json with your own control register, then run python3 -m evals.run --run-id b000-<yours>-vocab --floor vocab — it needs no key and costs nothing, and on this corpus it is ahead of the paid call. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per confirmation for date and amount comparisons code already makes exactly. That is the case against the best-fitting scenario (“Your confirmation files are mostly decidable from printed columns — addresses, dates, record kinds, amounts”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A reconciling item supported by several records, or one record clearing several items — partial payments, one receipt against many invoices, netting. The reduction is one item, one record. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NOT VERIFIED THAT THE PAID CALL BEATS FREE CODE — ON THIS CORPUS IT DOES NOT. 39 of 64 files whole against the vocabulary floor's 40, 16 / 17 discordant, exact p = 1.0; no field-level row against any free floor is significant. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-audit-confirmation. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured rendered the whole board, all four free floors and every committed run; tools/build_corpus.py --check-fillers measured the free floors in well under a second and evals.check_labels returned 0 disagreements. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the Ask the model button and a new scored run.








