Home › Use Cases › Work one season's auditor request list item by item back to the trial balance
Use caseUC0459
🧪 Use-case kit · runnable

Work one season's auditor request list item by item back to the trial balance

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An auditor sends a request list at the start of fieldwork — one line per item, each naming an account, a period and what it wants. Somebody in the finance office then works it item by item: find which schedule on file answers this request, check it was prepared for the period asked for, agree its total to the trial balance of the account the REQUEST names, check controller review stands for it at the as-at date, check it has not already gone across this season, and where nothing was produced, write down exactly where it was looked for. The desk keeps a tracker line per item, and by the middle of fieldwork that tracker is wrong on most of them, because the deciding fact is usually a sentence somebody typed in the notes. Opening one request packet, finding the schedule that answers the request among those on file, checking its period, agreeing its total to the trial balance of the account the request names rather than the one the schedule ties to, reading every note for a review completed or withdrawn and checking its date against the as-at date, reading the furnish log for a line already sent and checking ITS date, and writing down the search record where nothing was produced.

Audience

The finance office of a nonprofit or foundation working one season's auditor request list, and the controller's office that owns review. This report's own answer is a qualified one: on the 36 items decidable from printed columns and dates the paid call is two items WORSE than free code, and the whole case for it is the 27 items a sentence decides. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual request-list items

The corpus is 63 request-list items, 0.09 MB (json 4 · jsonl 1 · md 2 · txt 63). It is generated because it has to be. A real auditor request list is an organisation's own working papers: named preparers and reviewers on every schedule, real balances, a real funder's grant conditions, and an audit in progress. None of that can be published, and a redacted version would destroy the very thing being measured — the notes, which is where the deciding fact lives on 27 of the 63 items. Generating it also buys something a real extract could not: the key is DERIVED from the same structure the files are rendered from, so there is no labelling opinion anywhere, and the failure modes are planted by name and counted.

The corpus

  • The 63 request-list itemsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every organisation, funder, account, schedule, request and balance is invented, and there are no people in it at all: a preparer is a desk and a reviewer is the controller's office. No audit standard, assurance framework, regulator or firm is named anywhere in the corpus, the code or the write-up.

Swap this folder for your own material and the kit is pointed at your request-list items. That is the whole change — there is no database to migrate.

One request-list item, as the model receives itAR-0001.txt · 1 of 63
AUDITOR REQUEST LIST ITEM (FY2026 SEASON)
  DOC                 AR-0001
  REQUEST             RQ-2141
  AS AT               2026-08-10
  PERIOD REQUESTED    2026-01-01..2026-06-30
  ACCOUNT NAMED       5110
  SUPPORT NAMED       SCH-1030
  ASKED FOR           The detail behind professional fees for the period requested, agreed to the trial balance.
  TRACKER STATUS      HOLD

TRIAL BALANCE (FY2026, as at 2026-06-30)
  ACCOUNT  DESCRIPTION                                  BALANCE
  4110     GRANT-REVENUE-PUBLIC                    1,364,127.49
  4120     GRANT-REVENUE-PRIVATE                     657,552.93
  4210     SPECIAL-EVENT-REVENUE                   1,185,580.46
  5010     SALARIES-AND-WAGES                      1,445,883.79
  5110     PROFESSIONAL-FEES                         677,249.44

SUPPORT ON FILE
  SUPPORT    PREPARED-BY        AS-AT        PERIOD                    TIES-TO           TOTAL  REVIEW
  SCH-1030   FUND-ACCOUNTING    2026-07-07   2026-01-01..2026-06-30    5110         677,249.44  PENDING
  SCH-1058   REQUEST-DESK       2026-07-07   2026-01-01..2026-06-30    5010       1,445,883.79  REVIEWED

FURNISH LOG (this season)
  LOG        SUPPORT    FURNISHED-ON   RELEASED-BY
  (none)

SEARCH RECORD
  LOCATION   SEARCHED-ON   RESULT
  (none)

NOTES
  - 2026-08-07 (RECORDS-DESK): SCH-1058 was re-indexed on 2026-08-07 and is not the schedule this request asks for.

The outcomeWhat a good result looks like

One request-list item in, one row out: which schedule answers it, whether review stands at the as-at date, whether it has already been furnished this season, where it was searched when nothing was produced, the difference to the cent, the exception set and one ARA-2026 status — with the desk's own tracker line printed beside it so a reader can see where they disagree.

And when it cannot

And what it does when it cannot. On the scored run 62 of 63 replies parsed; the one that did not opened with prose, never opened a JSON object, was closed by the streamed reply stop at 69 output tokens, and is counted WRONG inside the denominator rather than dropped. The 9 items it gets wrong are named in the kit README with what it answered and why, and 8 of the 9 are the same single field — reviewed, read as withdrawn off a note that withdraws nothing.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your request tracker already records review status, furnish status and the schedule id as COLUMNS, and your notes are free text nobody reads twice — the free columns floor at $0.00
    on the 36 items decided by printed columns and dates it scores 33 of 36 against the paid call's 31. You would be paying to be two items worse.
  • Your notes carry decisions in a stable in-house vocabulary and somebody in the office can write the phrase list — the free vocab_neg floor at $0.00
    41 of 63 whole, and it costs a morning of somebody's time rather than a per-item fee. It is this kit's floor of record for exactly that reason.
  • The deciding fact is usually a sentence, the wordings vary, and a wrong REVIEW reading is cheap because a higher rule usually owns the status — the paid call at $0.000786 per item
    54 of 63 whole, +13 over the floor of record at p = 0.0192, and 23 of the 27 note-decided items — the ones a phrase list cannot reach.
  • You hold this kit's own corpus generator and can write a regex against its decoy wordings — the free tuned floor at $0.00
    46 of 63, and the paid call's 54 is NOT significantly better on 63 items (16 against 8, p = 0.152). Published because it is the honest comparison for that reader.
  • You need a defensible record of WHY each item was decided the way it was — the paid call, and keep the run file
    every item carries the four readings, the key's beside them, the station's determination, the confidence and the correctness of all seven fields.

And where nothing here is good enough:

  • You want the arm to do the arithmetic as well as the reading — neither — recompute in code
    the same replies score 51 RAW and 54 RECHECKED, and difference alone goes from 57 to 62 when the station recomputes it in integer cents.

At a glanceHow the whole thing runs

86%rechecked all correct pct
1,562 msp50, end to end
$2.38per 1,000 request-list items · the fast tier

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own request packets in the same six-block shape and data/requests.json with your own request register, then run python3 -m evals.run --run-id b000-<yours>-vocab_neg --floor vocab_neg — no key, no network, no spend — to see what free code already gets on your material before you buy a single call. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: The moment a deciding fact moves into a sentence. On the 27 items a note decides, the same floor gets 4 of 27. That is the case against the best-fitting scenario (“Your request tracker already records review status, furnish status and the schedule id as COLUMNS, and your notes are free text nobody reads twice”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A request list that is not six fixed-width blocks per item — a spreadsheet export, an email thread, a scanned PDF. src/packets.py binds to the literal block headings, so anything else parses to nothing and every reading is empty. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-audit-evidence. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board and every committed run, and scores all six free floors offline. pip install -r requirements.txt installs nothing — the kit is standard library only — and the corpus is rebuilt from the seed in about a second. The only thing a missing key removes is the button that would buy a call, which is disabled and says why.

A living map of modern AI — kept current every morning