Home › Use Cases › Reconcile a donor pledge's installment schedule against the payments applied to it
Use caseUC0271
🧪 Use-case kit · runnable

Reconcile a donor pledge's installment schedule against the payments applied to it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A donor promises $30,000 over five installments and then pays — some on time, some short, some against the wrong fund, one to a different pledge reference entirely, one receipted and applied to nothing. Somebody in gift administration has to work out what is actually past due before anyone is asked for anything: which installments have fallen due after their grace period, what cleared against THIS pledge and its designated fund, what was written down by the role that holds that authority, and what the reminder log already records. Today that is a spreadsheet per campaign per quarter and a reminder list that either goes out or does not. The expensive version of getting it wrong is not arithmetic: it is a RECEIPT THAT IS ALREADY IN THE BUILDING AND IN THE WRONG PLACE. Counted, the pledge reads current and a lapse is never chased. Ignored, a donor is asked for money they have already given. Re-working a pledge schedule against its receipts by hand and deciding, pledge by pledge, whether the donor should be on the reminder list at all. It does not replace the reminder, the write-down or the fund release — none of the three exists anywhere in this kit, and two of them are named-role acts the answer contract has no field for.

Audience

A gift administrator working a campaign's pledge schedules, and the development director who reads the reminder list they produce. The decision this report is for is narrower than it looks: not 'should we buy a model', but 'which HALF of this job should a model touch'. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual pledge reconciliation packs

The corpus is 66 pledge reconciliation packs, 0.11 MB (txt 66). It is generated because it has to be, and for a stronger reason than most kits in this series: a real pledge reconciliation pack is a file about a NAMED INDIVIDUAL and their giving. There is no public version of one, there is no anonymisation of one that leaves the task intact, and publishing a redacted real one would be publishing a donor's giving history. So the corpus is invented from a fixed seed, carries a donor ACCOUNT REFERENCE and nothing else, and everything about it is arranged to make the READING hard rather than the parsing: 66 packs, 62 of them marked hard, at most ONE cause each, and 32 carrying none at all. The distractor prose is the point of the rest — a standing bulletin quoting PLR-2026's own fund rule on a pack where every receipt is correctly coded, a donor letter asserting an installment was settled in cash, a case note instructing the reader to send the reminder today or write the balance off, a newsletter dated after the last real reminder, and an approved reminder schedule that was never sent.

The corpus

  • The 66 pledge reconciliation packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 66 packs, the pledge register and the whole answer key are generated in-process from seed 20260902 by tools/build_corpus.py, and --check rebuilds all of them and diffs them against disk at 0 differences under two PYTHONHASHSEEDs. Nothing is fetched, scraped or licensed from anywhere, so there is no source URL to map and no third-party dedication to verify.

Swap this folder for your own material and the kit is pointed at your pledge reconciliation packs. That is the whole change — there is no database to migrate.

One pledge reconciliation pack, as the model receives itPLP-0001.txt · 1 of 66
PLEDGE RECONCILIATION PACK                                          PLP-0001
Prepared 2026-09-02 under PLR-2026 | schedule reviewed to 2026-08-31

PLEDGE FACTS
  Donor account                DNR-43943
  Pledge reference             PLG-2026-1076
  Campaign                     Legacy Circle 2026
  Fund designation             unrestricted - Unrestricted Reserve (FND-780)
  Pledge total                 $2,500.00
  Pledge date                  2025-09-10
  Grace period                 45 days after each due date

PLEDGE SCHEDULE
  installment  due         amount
  INS-01       2026-03-28  $     500.00
  INS-02       2026-04-28  $     500.00
  INS-03       2026-05-28  $     500.00
  INS-04       2026-06-28  $     500.00
  INS-05       2026-07-28  $     500.00

RECEIPTS ON FILE
  receipt      received          amount   applied to
  RCT-886166   2026-03-23  $     500.00   PLG-2026-1076 INS-01 / FND-780
  RCT-886221   2026-04-25  $     500.00   PLG-2026-1076 INS-02 / FND-780
  RCT-886235   2026-05-23  $     500.00   PLG-2026-1076 INS-03 / FND-780
  RCT-886257   2026-06-24  $     500.00   PLG-2026-1076 INS-04 / FND-780

LEDGER ADJUSTMENTS
  none recorded against this pledge for the period

REMINDER LOG
  The campaign newsletter was sent to the whole donor file on 2026-07-05.
  The pledge card came back on 2026-06-25 with the reminder preference left blank.
  A further reminder was raised on 2026-05-09 and logged against the pledge record.

CASE NOTES
  The campaign team confirms this pledge is inside the campaign's own reporting period.
  This pack was raised on the standard reconciliation cadence and not in response to a specific query.

The outcomeWhat a good result looks like

One pack in, six graded answers out: the arrears figure to the cent after grace, the cause class, the pledge status, the list decision, the date of the last reminder the log actually records, and the line that establishes the cause quoted verbatim and locatable in the pack at character offsets. On this corpus the paid call returns 66 of 66 figures exact, 66 of 66 causes, 66 of 66 statuses, 66 of 66 list decisions and 66 of 66 reminder dates — 63 of 66 packs with all six right.

And when it cannot

And what it does when it cannot. All 3 of its misses are the SAME defect and none of them is a reading failure: on the three duplicate-receipt packs the same reference is printed twice — once with the fund written as its code and once as its name — and the call quoted THE OTHER HALF OF THE PAIR from the one the key names. src/citation.py locates the quote, measures 0 per cent coverage of the labelled line and gives no credit, correctly, because there is no partial credit anywhere here. Both lines ARE the duplicate. ⚠︎ THE FREE RULES FLOOR SCORES 66 OF 66 ON CITATIONS AGAINST THE MODEL'S 63 for exactly that reason: it quotes the first line its expression matches, which is the line the generator labelled. That is the one field where the paid call loses, and it loses to a tie-break rather than to a reading.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the arrears figure and nothing else — the paid call, or a floor with the grace rule added to it
    the call is 66 of 66 exact and the shipped floor is 41 of 66 — and every one of the floor's misses is one mechanism, the grace period, which is printed on the pack in a fixed place. A reader who adds that one test to evals/baseline.py should expect most of this gap to close for $0.00, and this kit did not measure it.
  • You want the cause named so somebody can correct the ledger — either arm — and prefer the free one
    66 of 66 on both, all 11 mismatches named on both, all 7 restricted-fund exposures named on both. The applied to column is structured enough that a parser reads it exactly, and the trap this kit was built around is held by a regex.
  • You want the last reminder read off a log people typed — the paid call, and this is what the money buys
    66 of 66 against the floor's 38. The log carries ISO dates and long-form dates, newsletters and stewardship reports dated later than the last real reminder, approved schedules that were never sent, and drafts on file. A regular expression matches one shape; the call read all of them.
  • You need to be sure no donor is asked for money already given — either arm, and neither is a guarantee
    wrongly_listed is 0 of 66 on the paid call and 4 on the free floor; listed_on_a_mismatch is 0 on both. But the floor's 4 are all packs inside a grace period, which is the same failure wearing a friendlier name.
  • You want the pure-code station to be your safety net — it is not one here, and the measurement says so
    recheck_overrides is 0 of 66 on the scored run and the raw and rechecked columns are identical to the digit. Under attack it recovered NONE of the packs the sentence moved, because the pledge register carries nothing about a receipt line — only 2 of the 16 targets had any register defence at all.

At a glanceHow the whole thing runs

63%pack all correct
62,532 msp50, end to end
$31.68per 1,000 pledge reconciliation packs · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs, data/pledges.json with your own register rows (one per pack id, carrying register from the five states), and data/gold.jsonl with your own key — one row per pack carrying arrears_cents, cause, status, action, last_reminder, citation and citation_span. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Reading the figure gap as a reading gap. It is a rule the floor does not implement. That is the case against the best-fitting scenario (“You want the arrears figure and nothing else”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A scanned or photographed remittance. Every arm here — the model's reading, the floor's arithmetic and the independent key check's arithmetic — rests on fixed-layout columns with an amount and an applied to string. 9 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled line is the line a gift administrator would have quoted. The key names ONE per named cause and the scorer has no partial credit; all three of the paid call's misses are packs where the OTHER half of a duplicate pair establishes the cause just as well. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?9 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-pledge-lapse. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9271 and scores every graded cell offline: the corpus rebuild from seed 20260902, the independent label gate, both free floors over all 396 graded cells, and the replay of the committed model run off its own result file. The 'Check with the model' control is disabled and the page prints why beside it. That was the state tools/shoot_ui.mjs photographed — it forces API_KEY empty before it starts the server — and it is what the first screenshot shows.

A living map of modern AI — kept current every morning