Home › Use Cases › Reconcile an audit confirmation response against the recorded balance and client records
Use caseUC0439
🧪 Use-case kit · runnable

Reconcile an audit confirmation response against the recorded balance and client records

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An audit engagement team sends external confirmations — to a client's bank, its customers and its suppliers — and before a confirmation can go into a workpaper the response has to agree with three other things: the address the team sourced independently, the client's recorded balance at the confirmation date, and the client's own records behind every reconciling item on its schedule. None of those was written to agree with the others. A response can print the sourced address and still have reached the desk through the client; a changed address can be genuine if a call-back verified it before AS AT; a later response can be withdrawn; an outstanding cheque can have cleared before the period end; a receipt can have been reversed. The client's own schedule concludes RECONCILED on 45 of 64 files and is wrong on 40. Opening one confirmation file, checking every logged response's FROM address character for character against the sourced one and its RECEIVED date against AS AT, reading each note for a forwarding, a call-back verification or a withdrawal and its date, matching every reconciling item to the client record it cites by kind, reference, amount and date, signing the supported items, and working out what is left unreconciled — by hand, for every confirmation in the period.

Audience

The confirmation desk and the engagement team on an audit, working a period's external confirmations into the cash, receivables and payables workpapers — and anyone deciding whether a language model should read those files at all. On this corpus the honest answer is no, and the page leads with it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual confirmation files

The corpus is 64 confirmation files, 0.17 MB (txt 64). It is generated because it has to be. A real confirmation file is an audit client's bank balances, its customers' and suppliers' names and amounts, and desk notes written by named people about named contacts — confidential on every line, and the one thing this kit measures (a sentence about who sent what, and when) is exactly what would be stripped out of anything publishable. So the corpus is built to the job instead: 19 cases, 32 files a sentence decides and 25 carrying a decoy note in the same vocabulary, five date traps, and the client's own schedule wrong on 40 files so the incumbent is measured rather than assumed.

The corpus

  • The 64 confirmation filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every engagement, client, confirming party, address, balance and record is invented, every email address sits under the reserved .example domain, and there are no people in the corpus at all — a note speaks for the confirmation desk, the engagement team or one of the client's desks. ECR-2026 is invented too, and no auditing standard is quoted, paraphrased or relied on.

Swap this folder for your own material and the kit is pointed at your confirmation files. That is the whole change — there is no database to migrate.

One confirmation file, as the model receives itACR-0001.txt · 1 of 64
================================================================================
EXTERNAL CONFIRMATION RECONCILIATION -- ONE REQUEST, ONE AS-AT DATE
================================================================================
FILE              ACR-0001
ENGAGEMENT        AUD-26-0410  client CL-3100, a regional distributor
CONFIRMATION      CNF-41200
TYPE              BANK
CONFIRMING PARTY  CP-7301  Callowfen Mutual Bank
ACCOUNT           BK-10400
CONFIRMATION DATE 2026-06-30
REQUEST SENT      2026-07-01
AS AT             2026-07-21
FILE COMPILED     2026-07-26
CURRENCY          USD

-- INDEPENDENTLY SOURCED ADDRESS (the engagement team's, not the client's) -----
SOURCED FROM      the bank's published confirmation service page, checked 2026-06-30
CHANNEL           email
ADDRESS           audit-confirmations@callowfen-mutual-bank.example

-- CLIENT RECORDED BALANCE AT THE CONFIRMATION DATE ----------------------------
LEDGER            cash book, bank account BK-10400
RECORDED BALANCE                     72,040.99

-- RESPONSES LOGGED BY THE CONFIRMATION DESK -----------------------------------
RESPONSE  RECEIVED    BALANCE STATED  FROM
RSP-1     2026-07-08       65,117.76  audit-confirmations@callowfen-mutual-bank.example

-- CLIENT RECONCILIATION SCHEDULE (as prepared by the client) ------------------
ITEM   KIND                  REFERENCE   DATED             AMOUNT  SUPPORT
RI-1   DEPOSIT IN TRANSIT    DEP-80000   2026-06-29      6,923.23  CR-51000

-- CLIENT RECORDS (ledger and bank activity, extracted 2026-07-18) -------------
RECORD    KIND                REFERENCE   DATED             AMOUNT
CR-51000  BANK CREDIT         DEP-80000   2026-07-01      6,923.23
CR-51001  BANK CREDIT         DEP-88000   2026-07-02      2,190.02

Abridged — the file continues.

The outcomeWhat a good result looks like

One confirmation file in, one row out: which responses are authentic at AS AT, which were withdrawn, which reconciling items the client's records support, the unreconciled amount to the cent, every exception and one of five ECR-2026 verdicts — SENDER-MISMATCH, NO-RESPONSE, EXCEPTION, RECONCILED or AGREED — with the response, record, address and date behind each. Never a conclusion, an adjustment, a sign-off or a contact.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 25 files it got wrong are named in the kit README with what it answered: 5 unsupported items kept (an EXCEPTION reported RECONCILED — the costly direction), 15 supported items dropped, 5 printed values compared the wrong way round in its own sentence, 3 withdrawn responses read as not authentic and 1 id list contradicting its own sentence. Confidence does not separate them: median 0.90 on right and wrong files alike.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your confirmation files are mostly decidable from printed columns — addresses, dates, record kinds, amounts — the free modal floor, and do not buy a call at all
    On the 32 column-decided files the modal floor gets all 32 for $0.00; the paid call gets 20.
  • A sentence often decides a reading — forwarded responses, call-backs, withdrawals, reversed or misapplied records — the free vocabulary floor first; the paid call is not ahead of it here
    On the 32 files a note decides, vocab gets 20 and the paid call 19. The paid call wins on notes about a RECORD (misapplied support 5 of 5 against the best floor's 3) and loses on notes about a SENDER (5 of 10 against vocab's 7).
  • You want the client's own reconciliation schedule audited — any arm — every one beats it
    The client's printed status is right on 24 of 64 verdicts and returns no reading. It is published as an arm so the comparison is against what is running today rather than against nothing.
  • The costly error for you is an EXCEPTION reported as clean — vocab or the paid call — they tie at 5, and the modal floor makes it 14 times
    The paid call's 5 are all unsupported items on column-decided files that the modal floor gets right; the modal floor's 14 are sentence files it cannot read. A date-and-amount check in code in front of the paid call would remove all 5.

And where nothing here is good enough:

  • Your engagement uses a materiality threshold, multi-account bank confirmations or open-invoice receivable confirmations — neither, yet
    No file in this corpus does. The unit of work is one balance, one response log and one-item-one-record support, and every percentage here is against that unit.

At a glanceHow the whole thing runs

61%rechecked all correct pct
1,445 msp50, end to end
$1.32per 1,000 confirmation files · the fast tier

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own confirmation files in the same block shape and data/confirmations.json with your own control register, then run python3 -m evals.run --run-id b000-<yours>-vocab --floor vocab — it needs no key and costs nothing, and on this corpus it is ahead of the paid call. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per confirmation for date and amount comparisons code already makes exactly. That is the case against the best-fitting scenario (“Your confirmation files are mostly decidable from printed columns — addresses, dates, record kinds, amounts”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A reconciling item supported by several records, or one record clearing several items — partial payments, one receipt against many invoices, netting. The reduction is one item, one record. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NOT VERIFIED THAT THE PAID CALL BEATS FREE CODE — ON THIS CORPUS IT DOES NOT. 39 of 64 files whole against the vocabulary floor's 40, 16 / 17 discordant, exact p = 1.0; no field-level row against any free floor is significant. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-audit-confirmation. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured rendered the whole board, all four free floors and every committed run; tools/build_corpus.py --check-fillers measured the free floors in well under a second and evals.check_labels returned 0 disagreements. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the Ask the model button and a new scored run.

A living map of modern AI — kept current every morning