Home › Use Cases › Reconcile one school's student activity fund for one month against the bank
Use caseUC0393
🧪 Use-case kit · runnable

Reconcile one school's student activity fund for one month against the bank

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A school office takes money at a window all month — ticket sales, club dues, concession takings — writes a receipt for every dollar, banks it in deposits, and pays club bills out by check. At the month end somebody has to put that ledger back against two numbers the school did not write: the balance carried forward and the bank statement's own closing balance. The work is not the subtraction. It is deciding which POSTED rows really moved money, because a deposit the bank returned unpaid and a check that was voided before it was presented both sit on the ledger reading POSTED, and the only record of either is a sentence typed under the row. Opening one school's monthly pack, adding the deposits up, subtracting the checks, reading the note printed under every posted row to decide whether that row really moved money, chasing each payment's requisition, and comparing the result with the balance the bank actually holds.

Audience

A district finance officer working a monthly pack across a dozen schools, and the activities secretary who assembled it. Who is NOT the audience: anybody deciding whether spending was allowable, or whether a person should answer for a difference. This kit reaches neither. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual fund months

The corpus is 62 fund months, 0.25 MB (txt 62). It is generated because it has to be. A real activity fund ledger is a named school's own record of money taken from named children, and the exact thing this kit measures — a sentence typed under a row saying a deposit came back unpaid — is written by a named person about a named account. There is no public corpus of these and there should not be. Generating it also buys the one thing a scraped corpus cannot: the key is DERIVED from the same structure the file is rendered from, so there is no second place the answer lives and a corpus change cannot leave a stale key behind.

The corpus

  • The 62 fund monthsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every school is an invented trading name, every club, vendor, receipt reference and paying-in slip number is arithmetic on the file index, and there is no personal data at all — evals/check_labels.py sweeps all 62 files for five families of identifier on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your fund months. That is the whole change — there is no database to migrate.

One fund month, as the model receives itAFR-0001.txt · 1 of 62
====================================================================================================
STUDENT ACTIVITY FUND RECONCILIATION                                        AFR-0001
School: SCH-101 - Northgate Ridge High School (invented)
Fund: SAF-101   Month: 2026-01   Procedure: SAF-2026
====================================================================================================

FUND AND MONTH AS THE DISTRICT LEDGER HOLDS IT
  opening bank                     39330.00
  statement bank                   43828.84
  opening undeposited                625.00   receipted and not yet banked
  tolerance pct                        1.00   pct of month activity
  threshold cash                    1500.00   or more
  threshold pct                        3.00   pct or more
  month                          full-month

CLUB ACCOUNTS AS CARRIED FORWARD
  CLUB       NAME                                    OPENING
  CLB-01     Yearbook                                1900.00
  CLB-02     Marching Band                            125.00
  CLB-03     Senior Class                            2345.00
  CLB-04     Drama Club                               515.00
  CLB-05     Student Council                         3440.00

FUND LEDGER AS POSTED BY THE SCHOOL OFFICE
  ROW        DATE               AMOUNT  KIND           STATUS   CLUB     SUPPORT    MEMO
  AFL-0001   2026-01-03        2101.00  receipt        POSTED   CLB-01   RCT-3000   ticket sales at the window
  AFL-0002   2026-01-03        1146.00  disbursement   POSTED   CLB-01   REQ-5000   club supplies, Larkspur Print Co
  AFL-0003   2026-01-05        1478.00  receipt        POSTED   CLB-02   RCT-3001   club dues collected in homeroom

Abridged — the file continues.

The outcomeWhat a good result looks like

One fund month in, one row out: which ledger rows this procedure treats differently from the ledger that printed them, which disbursements the fund cannot show support for, whether an opening correction was in force, and the verdict SAF-2026 derives from those readings — with the undeposited cash and the overdrawn club count beside it.

And when it cannot

And what it does when it cannot. On the scored run 62 of 62 replies parsed, nothing stopped at the ceiling, and there were 0 failures. What it does NOT do is get the arithmetic right when the arithmetic is the reading: see not_good_enough.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your ledger's bad rows announce themselves — every returned deposit says "returned", every void says "void" — the free rules floor, and do not buy a call at all
    The floor already beats the paid arm on the whole-month score (18 against 12) and a keyword list over the notes reaches everything a keyword can reach, for $0.00.
  • Your fund notes routinely record corrections that were QUERIED, REQUESTED or PROPOSED beside the ones that were actually made — the paid call, and read the correction field's own rate first
    The regex applies 11 corrections that were never made; the paid arm applies 0, and reads the field right 62 of 62 against the regex's 51. Every declined wording carries a real effective date inside its own governing window, a real reference and a real amount, so a date test and a shape test both pass and only the sentence says no.
  • Your deposit memos describe what was sold (N bk at P per bk) and the amount column is usually already the money — the free floor, without hesitation
    This is arithmetic, not reading: the memo states both numbers and they either multiply out to the column or they do not. Free code gets the family 3 of 5; the paid arm gets 1 and invents 44 citations across the corpus doing it.
  • You need the refusals held under pressure — no correcting entry, no allowability ruling, no name — either arm, and keep both defences
    0 of 18 attacked months proposed an action, stated one as accomplished fact, ruled on allowability or named a person — including the three handed a clause asking in terms for all three. The schema offers no field that could express any of it and src/refusal.py reads every free-text surface of every arm anyway.

And where nothing here is good enough:

  • Your disbursements' support is often filed somewhere the support column does not name — neither, without reading the two error directions apart
    The paid arm misses 2 unsupported rows against the free rule's 23 — much better. But it also names 15 SUPPORTED rows as unsupported, where the free rule names 0. Which direction costs you more is a question about your desk, not about the model.

At a glanceHow the whole thing runs

19%all five correct pct
1,815 msp50, end to end
$0.00per 1,000 fund months · google/gemini-3-flash

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own fund month files in the same shape and data/schools.json with your own register rows, then run python3 -m evals.check_labels and python3 tools/build_corpus.py --check. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per fund month for a regex you could write in an afternoon. That is the case against the best-fitting scenario (“Your ledger's bad rows announce themselves — every returned deposit says "returned", every void says "void"”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A ledger that is not fixed-width columns. src/ledger.py's row regex is the shape these files print; a CSV export or a bank API feed needs a different parser and nothing above it changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-activity-fund. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors and every committed run. python3 -m evals.run --run-id b000-activity-fund-rules --floor rules makes no call and needs no credential.

A living map of modern AI — kept current every morning