Home › Use Cases › Decide which expenses actually release a donor's restriction
Use caseUC0174
🧪 Use-case kit · runnable

Decide which expenses actually release a donor's restriction

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A donor gives money for something. The gift agreement says what, and for how long, and sometimes on what condition -- in prose, in a letter. The fund accounting system then records that as a short purpose string and a list of fund codes, and expenses are charged against it. At period end somebody decides, expense by expense, which of those charges actually satisfies the restriction and therefore releases funds from restricted to unrestricted net assets, and what the correct release total for the period is. MOST OF THAT IS A DATE TEST AND A LOOKUP. The part that is not is the part that moves money wrongly: a purpose the ledger records broadly and the donor wrote narrowly; a letter that ends the restricted period earlier than the fund record does; a milestone the Condition column says nothing about; and -- the other direction -- a letter that names the very cost being charged, in the sentence recording that the donor was asked and declined. Somebody with the gift file open beside the general ledger: reading each award letter, deciding whether the expense in front of them is what the donor meant, checking the date against the restriction period the LETTER states rather than the one the fund record carries, noticing that the same evaluation contract has been charged to two gifts, watching the unreleased balance, and totalling the release. The date test and the code lookup in that list are free and this kit ships them; the reading is what it is measuring.

Audience

Nonprofit finance staff and their auditors: whoever prepares the period's release schedule against restricted gifts and whoever tests it afterwards. The reader is not asking 'is this expense real' -- it is in the ledger. They are asking 'does it satisfy what this donor actually wrote', which is a question about a letter, not about a code. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual restriction release review packs, one accounting period each

The corpus is 38 restriction release review packs, one accounting period each, 0.25 MB (txt 38). A real corpus of this shape cannot be published. Restricted gift agreements name donors and carry their instructions; the expense ledger beside them names staff salaries and household-level assistance. There is no redaction that leaves the thing this kit measures intact, because the thing it measures IS the donor's own wording. And the answer key has to be KNOWN rather than inferred: the whole claim is a comparison between what free code settles from the tables and what only the letter settles, and that comparison needs a per-line ground truth nobody argued about afterwards. Each line is built to be a named case and the key is written from the case; evals/check_labels.py then reads the shipped page back through the shipping parser and asserts the planted key still holds.

The corpus

  • The 38 restriction release review packs, one accounting period eachgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your restriction release review packs, one accounting period each. That is the whole change — there is no database to migrate.

One restriction release review packs, one accounting period each, as the model receives itRR-0001.txt · 1 of 38
Restriction Release Review
------------------------------------------------------------------------
Pack: RR-0001
Organization: Harbor Light Trust
Period: 2025-10-01 to 2025-12-31
Period End: 2025-12-31
Prepared: 2026-01-21
Release Policy: RP-2026.1

Release Rules In Force
------------------------------------------------------------------------
RR-1.1  Every line in "Expenses Charged In The Period" is decided once, against the gift it is
        charged to in that table's Charged To column. No line is left out.
RR-1.2  Where more than one rule bears on a line, apply them in this order and stop at the first
        that decides it: RR-7.1, RR-5.1, RR-4.1, RR-3.1, RR-2.2, RR-3.2, RR-2.1, RR-6.1.
RR-2.1  An expense releases restricted funds when its account code is one of the gift's Fund Codes
        and its date falls inside the gift's governing time restriction. An expense whose account
        code is not one of the gift's Fund Codes, and which no award letter in this pack brings
        inside the purpose, does not release.
RR-2.2  An expense dated outside the gift's governing time restriction does not release, however
        well its purpose fits. The governing time restriction is the Restriction Window in
        "Gift Agreements On File" unless an award letter in this pack states a different one.
RR-3.1  Where an award letter states the restricted purpose more NARROWLY than the Ledger Fund
        Purpose recorded against the gift, the letter governs and an expense outside the letter's
        purpose does not release even though its account code is one of the gift's Fund Codes.
RR-3.2  Where an award letter states the restricted purpose more BROADLY than the Ledger Fund

Abridged — the file continues.

The outcomeWhat a good result looks like

A drafted reconciliation: one entry per expense, in the pack's order, each carrying the gift it was decided against, RELEASE / NO_RELEASE / INSUFFICIENT_EVIDENCE with a named ground, the governing rule from the pack's own printed rulebook, the gift agreement and donor letter identifiers the ground needs, the amount released on that line, and the period release total -- with the strongest free floor's answer computed beside every line for nothing.

And when it cannot

⚠︎ THE PAID ARM WINS THIS ONE, AND THE FIRST THING TO SAY IS THAT ONE OF ITS 38 REPLIES CAME BACK EMPTY. r002-restriction-release reached 86.17 pct on the discriminator against the strongest free floor's 78.01 pct -- 8.16 points -- and it got the period release total exactly right on 37 of the 37 periods it answered, against the floor's 11 of 38. ⚠︎ RR-0002 RETURNED NOTHING. finish_reason 'stop', 11,707 output tokens, every one of them provider-side reasoning, and an empty string where the JSON should have been. Not a truncation -- the ceiling was 64,000 tokens and the heaviest reply in the run drew 42,334. Its 7 lines are counted as OMITTED and stay inside the published denominator; the run was not re-fired. ⚠︎ AND THE PAID ARM IS WORSE THAN FREE CODE ON THE RECORD CHANNEL: 85.32 pct against 100.0 pct on the 218 lines the printed tables settle. It buys its headline entirely on the other 64, where every free floor scores 0.0 or 28.12 pct.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Every restriction that decides a line is already a FIELD -- a purpose string, a fund code list, a window, a condition flag -- and your gift files carry no prose anybody has to read. — the free floor -- system-fields, $0.00
    It is at 100.0 pct on the 218 lines the printed tables settle, against the paid arm's 85.32 pct. It gets the duplicate tiebreak, the balance cap and both absences exactly right, invents nothing, and runs the whole corpus in a fraction of a second with no key. On a corpus with no letter channel it is simply the better product.
  • Your award letters narrow, widen or re-date what the ledger recorded, and the release total is signed by somebody. — the paid arm, $0.05482837 per period
    89.06 pct on the letter channel against every free floor's 0.0 or 28.12, and the period release total exact on 37 of the 37 periods it answered against the strongest floor's 11. It over-released nothing at all -- 0 of 136 lines that release nothing -- and it released none of the 16 costs the donor was asked about and declined, where the keyword floor released all 16.
  • You want the cheapest thing that reads a letter at all. — ⚠︎ NOT letter-keyword, even though it wins the free discriminator
    It is 78.01 pct against system-fields' 77.3 -- 0.71 points -- and it buys them by releasing every one of the 16 costs whose letter names them in the sentence recording the donor declined. Its period total is exact on 11 of 38 periods against system-fields' 7, and its median period is out by $43360.39 against system-fields' $25779.79.

At a glanceHow the whole thing runs

86%restriction release accuracy pct
135,691 msp50, end to end
$54.83per 1,000 restriction release packs · Gemini 3 Flash

Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt and data/gold.jsonl, in that order, and read data/SOURCES.md first. ⚠︎ WHAT DOES NOT TRANSFER IS THE BASE RATE, AND IT IS THE FIRST THING PEOPLE CARRY OVER. Corpus lens →
When is this the wrong choice?Avoid: Do not use it where anything is settled in prose: it is at 0.0 pct on all 64 letter-channel lines here and its period total is exact on 7 of 38 periods. That is the case against the best-fitting scenario (“Every restriction that decides a line is already a FIELD -- a purpose string, a fund code list, a window, a condition flag -- and your gift files carry no prose anybody has to read.”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack layout that is not this one. src/pack.py is a set of regular expressions written for these headings, these key/value blocks and these fixed-width tables; against a real fund accounting export it returns EMPTY tables rather than a plausible wrong answer, which is the failure you want. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?The support convention. 10 further misses are lines where every judgement field is right and the arm attached the line identifier where the key wanted the gift -- a requirement the prompt never states. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-26 — r002-restriction-release. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — git clone, then python3 -m src.app and open http://127.0.0.1:9085. No pip install, no index build, no key. The corpus, the answer key and every committed result ship with the kit, so the page renders in full offline and the second button replays what r002-restriction-release actually answered. python3 -m evals.check_labels and all three free floors run with no key at all.

A living map of modern AI — kept current every morning