The business caseThe problem this solves
A grant's disbursement schedule conditions each scheduled payment on a milestone, and releasing one is not a summary of what the grantee reported -- it is a CLAIM that the milestone behind it was met. The dangerous line is therefore never a wrong figure. It is a payment marked Release when the evidence behind it cannot carry the word: the required deliverable never arrived, the period reported is not the period the schedule attaches, a condition precedent was reversed after it was recorded, or the deliverable on file has since been withdrawn or returned for correction. In every one of those cases the reported figure meets its target. Someone reading a grant's whole payment request against the agreement before the money goes out: checking that every scheduled payment was actually reported on, that the deliverable the schedule names is the deliverable on file, that the period reported is the period the payment attaches to, that any condition precedent is cleared and has stayed cleared, and that the target being measured against is the one in the amendment in force rather than the one transcribed off the grantee's cover sheet.
Audience
Grants management and finance staff at foundations and funders who decide whether a scheduled payment goes out, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual grant payment request packets
The corpus is 24 grant payment request packets, 0.12 MB (txt 24). A real grant payment request names a charity, its banking details, its staff, its shortfalls and a funder's private assessment of all of it. There is no licence under which that could ship in a public repo, and a version redacted enough to publish would no longer contain the thing being measured. So the corpus is written rather than collected, and data/SOURCES.md says so on its first line rather than in a footnote. What is synthetic is the text; what is real is the defect shape.
The corpus
- The 24 grant payment request packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your grant payment request packets. That is the whole change — there is no database to migrate.
Payment Request
----------------------------------------------------------------
Packet GM-0001
Grantee Harborline Community Trust
Programme Neighbourhood Food Security Initiative
Award GA-2025-100
Request date 2026-07-05
Prepared by R. Adeyemi, Grants Management
Grant Agreement
----------------------------------------------------------------
Agreement GA-2025-100
Amendment in force Amendment 2, effective 2026-03-10
Superseded amendment Amendment 1, superseded 2026-03-10
Holdback 5 pct of each scheduled payment, released at closeout
Disbursement schedule under Amendment 2 -- every scheduled payment below must be
decided on this request:
No. Milestone Deliverable Target Period Amount
T1 Closeout narrative filed CLOSE-NARR 1 narrative 2025-10-01 to 2025-12-31 $20,000
T2 Evaluation design filed EVAL-DES 1 design 2025-10-01 to 2025-12-31 $25,000
T3 Advisory board minutes filed CAB-MIN 1 minute set 2026-01-01 to 2026-03-31 $30,000
T4 Staff trained TRAIN-ROSTER 55 staff 2026-01-01 to 2026-03-31 $35,000
T5 Meals distributed DISTRIB-LOG 12,600 meals 2026-04-01 to 2026-06-30 $40,000
T6 Workshops delivered WKSHP-LOG 25 sessions 2026-04-01 to 2026-06-30 $45,000
T7 Baseline assessment filed BASELINE-RPT 1 report 2026-07-01 to 2026-09-30 $50,000
Grantee Progress Report
----------------------------------------------------------------
Milestone T1Abridged — the file continues.
The outcomeWhat a good result looks like
A release recommendation: one line per scheduled payment, in the schedule's order, each carrying RELEASE, WITHHOLD or CANNOT_DECIDE with a named blocker -- plus a packet-level disposition, and gross/holdback/net computed in code from whichever decisions you take. Beside every line, what the strongest free code would have decided.
And when it cannot
TWO FAILURES, AND THEY ARE NOT THE SAME SIZE. The model recommended releasing ONE payment it should have refused -- GM-0015 / T2, $95,000, where the record's own field reads "Deliverable filed: none -- not received" and a BENIGN programme officer note said the return "was filed through the grants portal and acknowledged automatically". The model believed the sentence over the field. The free floor, which cannot read notes at all, caught it. That is the whole of its 1.96 pct silent-release rate and it is the direction that costs money. The other failure is over-refusal: 27 of 108 decidable payments were refused rather than decided (25.0 pct), against ZERO from the free floor. But see Eval.could_not_verify -- 22 of those 27 sit on cells where THIS CORPUS is defective and the model is right, and they are published unfixed.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Every fact that decides a release is already a FIELD -- a deliverable code, a period, a condition line, and a report that either exists or does not. — the free floor -- evidence-gate, $0.00
It scores 82.14 pct on the discriminator against the model's 83.33, catches 100 pct of structured blockers, raises 0 pct false holds and does the arithmetic at 100 pct. It is free and it wins seven of the twelve published columns. - Decisive evidence lives in officer notes, e-mail threads or narrative -- an approval reversed after it was recorded, the wrong file behind a right code. — the model, on top of the free floor rather than instead of it
30 of 30 prose blockers caught, against zero from every free floor, which drops the silent-release rate from 58.82 pct to 1.96 pct. - You need a number you can put in front of a grants committee, and a wrong release is materially worse than a delayed one. — run both arms and take the union of their refusals
The two arms fail in opposite directions: the floor's failure is 58.82 pct silent releases, the model's is 25.0 pct false holds. A union refuses on either signal, which on this corpus would have refused every genuinely unevidenced payment.
And where nothing here is good enough:
- You want the money figure checked. — nothing here
Gross, holdback and net are computed in code from whatever an arm decided and are not scored by anything in evals/. No figure on this page measures them.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-27. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | tools/build_corpus.py writes data/corpus/*.txt and data/gold.jsonl. None of the measured figures on this page transfer to your own packets. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not use it where any decisive evidence lives in prose. It scores exactly zero there, by construction, and it will say RELEASE with total confidence. That is the case against the best-fitting scenario (“Every fact that decides a release is already a FIELD -- a deliverable code, a period, a condition line, and a report that either exists or does not.”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet layout that is not this one. src/packet.py is four regular expressions written for these headings and these key/value lines; against a real grants management system's export it parses nothing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | This kit publishes one paid arm and its evidence-blind control, against three free floors, on a corpus its own author wrote. There is no second tier, no red-team probe and no screenshot of a live call, the money arithmetic is scored by nothing, and the answer key is known to be wrong on 57 of its 168 cells. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-27 — r001-milestone-disburse. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Every packet, the answer key, all six recorded runs and all three free floors are in the repo. python3 -m evals.check_labels and python3 -m evals.run --floor evidence-gate both run on a fresh clone with no key, no network and no install.



