Home › Use Cases › Grant payment-to-milestone reconciliation
Use caseUC0193
🧪 Use-case kit · runnable

Grant payment-to-milestone reconciliation

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A grant's disbursement schedule conditions each scheduled payment on a milestone, and releasing one is not a summary of what the grantee reported -- it is a CLAIM that the milestone behind it was met. The dangerous line is therefore never a wrong figure. It is a payment marked Release when the evidence behind it cannot carry the word: the required deliverable never arrived, the period reported is not the period the schedule attaches, a condition precedent was reversed after it was recorded, or the deliverable on file has since been withdrawn or returned for correction. In every one of those cases the reported figure meets its target. Someone reading a grant's whole payment request against the agreement before the money goes out: checking that every scheduled payment was actually reported on, that the deliverable the schedule names is the deliverable on file, that the period reported is the period the payment attaches to, that any condition precedent is cleared and has stayed cleared, and that the target being measured against is the one in the amendment in force rather than the one transcribed off the grantee's cover sheet.

Audience

Grants management and finance staff at foundations and funders who decide whether a scheduled payment goes out, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual grant payment request packets

The corpus is 24 grant payment request packets, 0.12 MB (txt 24). A real grant payment request names a charity, its banking details, its staff, its shortfalls and a funder's private assessment of all of it. There is no licence under which that could ship in a public repo, and a version redacted enough to publish would no longer contain the thing being measured. So the corpus is written rather than collected, and data/SOURCES.md says so on its first line rather than in a footnote. What is synthetic is the text; what is real is the defect shape.

The corpus

  • The 24 grant payment request packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your grant payment request packets. That is the whole change — there is no database to migrate.

One grant payment request packet, as the model receives itGM-0001.txt · 1 of 24
Payment Request
----------------------------------------------------------------
  Packet                 GM-0001
  Grantee                Harborline Community Trust
  Programme              Neighbourhood Food Security Initiative
  Award                  GA-2025-100
  Request date           2026-07-05
  Prepared by            R. Adeyemi, Grants Management

Grant Agreement
----------------------------------------------------------------
  Agreement              GA-2025-100
  Amendment in force     Amendment 2, effective 2026-03-10
  Superseded amendment   Amendment 1, superseded 2026-03-10
  Holdback               5 pct of each scheduled payment, released at closeout

Disbursement schedule under Amendment 2 -- every scheduled payment below must be
decided on this request:

  No.   Milestone                       Deliverable     Target            Period                      Amount
  T1    Closeout narrative filed        CLOSE-NARR      1 narrative       2025-10-01 to 2025-12-31    $20,000
  T2    Evaluation design filed         EVAL-DES        1 design          2025-10-01 to 2025-12-31    $25,000
  T3    Advisory board minutes filed    CAB-MIN         1 minute set      2026-01-01 to 2026-03-31    $30,000
  T4    Staff trained                   TRAIN-ROSTER    55 staff          2026-01-01 to 2026-03-31    $35,000
  T5    Meals distributed               DISTRIB-LOG     12,600 meals      2026-04-01 to 2026-06-30    $40,000
  T6    Workshops delivered             WKSHP-LOG       25 sessions       2026-04-01 to 2026-06-30    $45,000
  T7    Baseline assessment filed       BASELINE-RPT    1 report          2026-07-01 to 2026-09-30    $50,000

Grantee Progress Report
----------------------------------------------------------------
Milestone T1

Abridged — the file continues.

The outcomeWhat a good result looks like

A release recommendation: one line per scheduled payment, in the schedule's order, each carrying RELEASE, WITHHOLD or CANNOT_DECIDE with a named blocker -- plus a packet-level disposition, and gross/holdback/net computed in code from whichever decisions you take. Beside every line, what the strongest free code would have decided.

And when it cannot

TWO FAILURES, AND THEY ARE NOT THE SAME SIZE. The model recommended releasing ONE payment it should have refused -- GM-0015 / T2, $95,000, where the record's own field reads "Deliverable filed: none -- not received" and a BENIGN programme officer note said the return "was filed through the grants portal and acknowledged automatically". The model believed the sentence over the field. The free floor, which cannot read notes at all, caught it. That is the whole of its 1.96 pct silent-release rate and it is the direction that costs money. The other failure is over-refusal: 27 of 108 decidable payments were refused rather than decided (25.0 pct), against ZERO from the free floor. But see Eval.could_not_verify -- 22 of those 27 sit on cells where THIS CORPUS is defective and the model is right, and they are published unfixed.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Every fact that decides a release is already a FIELD -- a deliverable code, a period, a condition line, and a report that either exists or does not. — the free floor -- evidence-gate, $0.00
    It scores 82.14 pct on the discriminator against the model's 83.33, catches 100 pct of structured blockers, raises 0 pct false holds and does the arithmetic at 100 pct. It is free and it wins seven of the twelve published columns.
  • Decisive evidence lives in officer notes, e-mail threads or narrative -- an approval reversed after it was recorded, the wrong file behind a right code. — the model, on top of the free floor rather than instead of it
    30 of 30 prose blockers caught, against zero from every free floor, which drops the silent-release rate from 58.82 pct to 1.96 pct.
  • You need a number you can put in front of a grants committee, and a wrong release is materially worse than a delayed one. — run both arms and take the union of their refusals
    The two arms fail in opposite directions: the floor's failure is 58.82 pct silent releases, the model's is 25.0 pct false holds. A union refuses on either signal, which on this corpus would have refused every genuinely unevidenced payment.

And where nothing here is good enough:

  • You want the money figure checked. — nothing here
    Gross, holdback and net are computed in code from whatever an arm decided and are not scored by anything in evals/. No figure on this page measures them.

At a glanceHow the whole thing runs

83%tranche decision accuracy pct
144,631 msp50, end to end
$46.80per 1,000 grant payment request packets · Google Gemini 3 Flash

Run once, for real, on 2026-08-27. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?tools/build_corpus.py writes data/corpus/*.txt and data/gold.jsonl. None of the measured figures on this page transfer to your own packets. Corpus lens →
When is this the wrong choice?Avoid: Do not use it where any decisive evidence lives in prose. It scores exactly zero there, by construction, and it will say RELEASE with total confidence. That is the case against the best-fitting scenario (“Every fact that decides a release is already a FIELD -- a deliverable code, a period, a condition line, and a report that either exists or does not.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet layout that is not this one. src/packet.py is four regular expressions written for these headings and these key/value lines; against a real grants management system's export it parses nothing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?This kit publishes one paid arm and its evidence-blind control, against three free floors, on a corpus its own author wrote. There is no second tier, no red-team probe and no screenshot of a live call, the money arithmetic is scored by nothing, and the answer key is known to be wrong on 57 of its 168 cells. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-27 — r001-milestone-disburse. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Every packet, the answer key, all six recorded runs and all three free floors are in the repo. python3 -m evals.check_labels and python3 -m evals.run --floor evidence-gate both run on a fresh clone with no key, no network and no install.

A living map of modern AI — kept current every morning