The business caseThe problem this solves
A subrecipient invoices under a subaward and somebody at the prime institution checks it before it goes forward. The check is against the subaward's own terms: the approved budget by cost category, the period of performance, the negotiated indirect rate and the base it applies to, the unallowable schedule and the prior-approval schedule. Most of that is arithmetic and a date subtraction. What it is NOT is the cost category column, which is a coding decision somebody at the subrecipient made -- and a capital item at $11,900.00 booked as Supplies is still equipment, a subsistence allowance paid direct to a non-employee attendee is still participant support, and a catering bill booked as Travel is still entertainment. A research administrator reading one subrecipient invoice line by line against the subaward agreement and deciding, per line, whether the cost is chargeable and under which article.
Audience
Whoever is deciding whether to put a reading step in front of a subaward invoice review that already has a column reader. The answer this kit gives is NO on the whole board and YES on 22 of 374 lines, and it says which 22. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual subaward invoice
The corpus is 62 subaward invoice, 0.17 MB (txt 62). A subaward invoice is the one document in research administration where a cost's CATEGORY and a cost's NATURE routinely disagree, and where the difference decides whether the prime institution can charge it. That is why the corpus is built as 13 families with exactly one defect each (measured: 0 invoices carry two) and why 22 of the 374 lines are DERIVED, not asserted, to be lines the columns cannot decide -- the rendered invoice is parsed back with src/rules.py and any line where the column answer differs from the key's is flagged. The other 352 are what a regular expression gets for nothing, and publishing that split is the whole point.
The corpus
- The 62 subaward invoicegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your subaward invoice. That is the whole change — there is no database to migrate.
SUBAWARD INVOICE - SUBRECIPIENT BILLING
INVOICE HEADER
Invoice SI-0001
Subaward SUB-2100-01
Prime institution Calder State University
Subrecipient Bexley Health Sciences Institute
Sponsor Federal Office of Materials Research (FOMR), prime award FOMR-4000000
Project Rural clinic staffing and retention
Invoice period 2025-07-01 to 2025-09-30
Submitted 2025-10-11
SUBAWARD TERMS
Period of performance 2025-01-01 to 2026-06-30
Indirect cost rate 22 pct
Indirect cost base MTDC - modified total direct cost, excluding equipment, participant support, and the portion of each subcontract over $25,000.00
Cost share required YES
Unallowable categories Entertainment, Alcohol, Lobbying
Prior approval required Equipment, a single item at or over $5,000.00; Foreign travel; Participant support rebudget
Approved budget by cost category
Personnel $239,000.00
Fringe $71,700.00
Travel $12,000.00
Equipment $26,000.00
Indirect $70,994.00
BILLED TO DATE BY CATEGORY (prior invoices under this subaward, this line excluded)
Personnel $119,500.00
Fringe $30,114.00
Travel $5,040.00
Equipment $6,760.00
Indirect $15,618.68
INVOICE LINES
# Category Service date Current period Cumulative Description
1 Personnel 2025-07-28 $55,194.00 $174,694.00 Salary and effort, Dr. E. Marchetti, 0.30 FTE for the periodAbridged — the file continues.
The outcomeWhat a good result looks like
A verdict and a cited subaward article on every line, with the invoice row copied verbatim for every line that is not payable, and one recommendation a named administrator acts on -- PAY, or HOLD naming the lines.
And when it cannot
An invoice sent forward for payment with an unallowable, out-of-period or over-budget line on it. Measured at 0 of 62 on both scored runs and 0 of 14 under the adversarial arm; the free rules floor does it 11 times.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You want the budget, period, threshold and indirect arithmetic checked — the free pure-code station alone - src/rules.py + src/recheck.py
352 of 352 lines and 44 of 62 whole invoices for $0.00, no key and no network, including 8 of 8 cumulative-over-budget traps. It beats the paid call on this half and it beats it on the whole board. - Your subrecipients code costs to categories that hide what they are — the paid call
17 of the 22 lines whose verdict lives in a description, against 0 for every free arm on this board - paired, discordant 17 to 0, exact two-sided p = 0.00002. That is the only thing on this page that separates, and it separates hard. - You want a guarantee that nothing gets paid, approved or waived — the shape of the answer contract, and read what it cannot express
no field can pay an invoice, release a payment, grant a prior approval, waive a certification, adjust or rebudget a line, or name an approver; src/prompt.py asserts it at import against 37 forbidden names and evals/check_labels.py fails the build on the same list. Four invoices instruct exactly those things and every committed arm answered all four correctly.
And where nothing here is good enough:
- You want a number you can plan against — neither, yet - run it twice first
the two identical runs differ by 3 invoices (43 and 40 of 62) while r001 differs from the free floor by 1. Any single number from one run of this kit is inside its own run-to-run spread.
At a glanceHow the whole thing runs
Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/ with your own rendered invoices and keep the six panel headings, or change LINE_RE in src/rules.py to your layout. DO NOT BRING A REAL SUBRECIPIENT INVOICE TO A KIT. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for arithmetic. It is most of the work in this job and none of the value. That is the case against the best-fitting scenario (“You want the budget, period, threshold and indirect arithmetic checked”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An invoice whose lines are shaped differently. Every column is read off a fixed layout by one regular expression in src/rules.py; a different column order or a wrapped description line returns no parsed lines at all, and the floor then answers nothing rather than answering wrong. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A THIRD run. Two identical passes were fired and they differ by 3 invoices; two points do not give a variance, only a range, and the range is wider than every margin on this page. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-03 — r001-subaward-invoice. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Nothing to fetch and nothing to install - this kit is Python standard library end to end. The corpus, the subaward agreements and the answer key rebuild byte-identically from one seed under two different PYTHONHASHSEEDs, in under a second.



