The business caseThe problem this solves
A student stops attending in week six. Some of their aid was earned and some was not; some of what was not earned goes back from the institution and some from the student, and what goes back has to be split across the funds in a prescribed order. THAT PART IS ARITHMETIC, AND THIS KIT DOES NOT PRETEND OTHERWISE -- a hundred lines of Python do it perfectly, free, for all 60 cases in a tenth of a second. The problem is upstream of the arithmetic: WHICH withdrawal date, WHICH funds, WHICH charges, WHICH days. Those inputs are spread across a term calendar, a withdrawal record, an aid ledger, a charges schedule, a document register and a page of free-text case notes, and the records disagree with each other. A dated note can move any input; an undated one moves none; a note dated BEFORE the ledger it contradicts moves none either; and a nine-day campus closure that never reached the term calendar moves nothing at all. Get one input wrong and every figure after it is wrong together. The named trap is the third of those, because it is the one nobody measures: an arm that reads everything and believes all of it is not better than one that reads nothing, it is differently wrong. An aid administrator working a withdrawal by hand: counting the days of the period off the term calendar, deciding which scheduled breaks are long enough to come out and which are not, deciding whether a leave of absence was approved in time to count, finding the date the calculation actually runs to when the withdrawal was unofficial, deciding which ledger rows are in the scheme at all and which have been reversed, totalling the institutional charges, doing the percentage and the two roundings, capping the institution's share at the unearned amount, splitting what is left across the funds in the prescribed order -- and then reading a page of case notes to find out which of those printed figures somebody has already changed, and which perfectly authoritative-looking sentence changes nothing at all.
Audience
Financial aid administrators and student account staff who work a return-of-funds recalculation after a mid-term withdrawal, the auditors who re-perform it afterwards, and the people who build student information and aid packaging tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual withdrawal recalculation cases
The corpus is 60 withdrawal recalculation cases, 0.45 MB (txt 60). The thing being measured has to be PLANTED to be measured. The question is whether an arm reads an input out of prose that contradicts a printed table -- and to count that you have to know which record was supposed to win. A real archive holds the figure that was POSTED, not the figure that was right, and where a recalculation was corrected on appeal the correction lives in a different system from the original. The fields the calculation turns on are financial aid records about identifiable people: a household income band, a dependency status, the reason a student stopped attending. A corpus that redacted them would not carry the question. And the rules had to be invented on top of that, or a good score could mean the arm recalled a real programme's convention rather than reading the one in front of it.
The corpus
- The 60 withdrawal recalculation casesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your withdrawal recalculation cases. That is the whole change — there is no database to migrate.
Return of Aid Recalculation
---------------------------
Case RT-0001
Institution Kelmoor Institute of Technology
Programme BSc Marine Biology
Payment period 2028 Winter Term
Recalculation prepared 2028-02-13
Scheme Rules
------------
The Aurelian Student Aid Scheme (ASAS) is an INVENTED aid programme, written for this kit. It does not exist, the
Aurelian Tertiary Finance Board does not exist, and the five funds named in ASAS-R11 do not exist.
These rules restate no real programme's rules. Every student, institution, date and
amount in this case file is invented too. Nothing here is an aid determination.
The rules of the scheme, in the form a recalculation applies them:
ASAS-R1 The period runs from the first day of instruction to the last day of instruction
named in the Term Calendar of this case, both days counted. Where any other record
in this case states a different first or last day of instruction, the Term Calendar
governs.
ASAS-R2 A scheduled break listed in the Term Calendar is excluded from the period when it
runs for FIVE or more consecutive days, counting both the first and the last day of
the break. A listed break of four days or fewer is counted as part of the period. A
closure, suspension or interruption that is NOT listed in the Term Calendar is never
excluded, however long it lasted and whatever any other record in this case says
about it.
ASAS-R3 A leave of absence is excluded from the period only where the case shows it was
APPROVED, and approved on or before the day the leave began. A leave that wasAbridged — the file continues.
The outcomeWhat a good result looks like
The whole recalculation, not just the answer: the withdrawal date, the days in the period, the days completed, the percentage completed to three decimal places, the aid in the calculation, the aid earned, the aid unearned, the institution's return, the student's share, any post-withdrawal disbursement owed, the split across the funds in the scheme's own order, and the outcome -- each cited to the rules and the records it came from. Beside every one of them, what the strongest free code would have written, and the first row where the two diverge.
And when it cannot
READ THE FREE COLUMN BEFORE THE PAID ONE. The paid arm scored 95.0 pct on the discriminator (57 of 60 cases with all twelve figures right, r001-r2t4-recalc). A KEYWORD SWEEP OVER THE SAME CASES SCORED 95.0 pct AT $0.00 (b002-r2t4-recalc-notesweep) -- a tie, to the case. And the plain calculator, which never opens the case notes at all, took 60.0 pct for nothing. The paid arm is BEHIND FREE CODE on the half free code owns: 95.83 pct against 100.0 pct on the 24 structured cases the printed tables settle outright. What it buys is the other half -- 95.83 pct of the 24 prose cases where the calculator scores 0.0 pct and cannot go at all. AND THE THREE MISSES ARE PUBLISHED UNFIXED WITH THEIR DIAGNOSIS: two are a ONE-CENT rounding disagreement (the arm carried the unrounded percentage into the money step where ASAS-R9 says to apply the rounded one) and evals/tolerance.py shows both disappearing at a $0.05 tolerance; the third is real and it is worth $105.37 in the direction that bills a student -- on RT-0035 the arm subtracted an approved leave of absence from the days COMPLETED and not from the days in the PERIOD, so its own numerator and denominator disagreed about the same six days.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your withdrawal files are structured. The term calendar, the leave record, the aid ledger and the charges schedule are all fields in a system, they are current, and nothing in your process turns on a sentence somebody typed into a notes box. — the free CALCULATOR -- src/calc.py, $0.00, a tenth of a second, no key at all
It scores 60.0 pct on the discriminator, is EXACTLY right on every case the printed tables settle, holds all 12 distractor cases by not reading, and invents nothing. On this half of the job a model is slower, dearer and -- on this run -- marginally worse: 95.83 pct against the calculator's 100.0 pct on the structured cases. - Your withdrawal files carry free text that matters. A disbursement reversed after the ledger was drawn, a leave approved late and recorded in a note, a charge credited back, an unofficial withdrawal whose real last date of activity is in an engagement report rather than in the record field. — the paid arm, with BOTH free columns computed beside it on every case
It takes 95.83 pct of those prose cases where the calculator takes 0.0 pct, and it held the inputs on 100.0 pct of the distractor cases -- the ones where a note LOOKS authoritative and is not. That second column is the one thing here that free code measurably cannot do: a keyword sweep believes any note carrying a date and held only 91.67 pct of them.
And where nothing here is good enough:
- You want a number for a board and you are choosing between arms on price. — neither, yet -- read baseline_note first
A keyword sweep over the dated notes TIED the paid arm on this corpus at 95.0 pct for $0.00. That number is an artefact: its four regular expressions were written against the twenty prose templates this corpus's generator writes, and it inverted them. It is an upper bound on free code HERE and not a claim about prose anywhere.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Rewrite src/rules.py for your own scheme and rebuild the corpus: the rulebook is data, src/calc.py follows it, and the answer key, all three floors and the gate follow from there. THE ONE THING NOT TO COPY IS THE SCHEME. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not use it anywhere a figure can be changed by something outside the fields. It scores 0.0 pct on the 24 cases where a dated note moves an input, and it does not fail loudly there -- it produces a complete, confident, wrong recalculation, because it never opened the note. That is the case against the best-fitting scenario (“Your withdrawal files are structured. The term calendar, the leave record, the aid ledger and the charges schedule are all fields in a system, they are current, and nothing in your process turns on a sentence somebody typed into a notes box.”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A case layout that is not this one. src/case.py is a set of regular expressions written for these headings, these two-space key/value lines and these fixed-width tables; against a real student information system's withdrawal export it parses nothing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | THE ABLATION WAS NOT RUN. python3 -m evals.run --blind removes the Case Notes section and says so in the system prompt; the code path is written, the truncation is anchored on the heading AND its underline, and it raises rather than silently no-opping. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is an OpenAI-compatible endpoint reached over raw HTTP -- src/adapters/__init__.py. The runtime provider is not named on this page; the kit runs on whichever key a forker already holds.; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-26 — r001-r2t4-recalc. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone, python3 -m tools.build_corpus, python3 -m src.app, open the port. No key, no install, no index build: the corpus, the rulebook, the answer key, all three free floors, the tolerance sweep and every recorded run are committed, and the generator is deterministic from one seed so a fresh clone rebuilds the same bytes.



