The business caseThe problem this solves
A conservancy's gift-ops desk drafts an acknowledgment letter for every gift that comes in: the gift amount and the date it was received, the fund it was designated to, the donor it is addressed to, the signer's role, every program line in the fund's printed order and the date the letter falls due. Everything needed is printed on the packet — except that a six-entry gift-ops log beside it can hold a gift for review, close a fund to designations, replace an addressee or re-record a gift's amount and date, and each of those quietly puts a printed value out of use. Reading one packet's gift-ops log against the organisation's own acknowledgment template to decide whether each of T-7's four elements is still usable, and then filling in the letter by hand.
Audience
A gift-ops officer deciding whether a drafted letter can go to review, and the person who owns the organisation's acknowledgment template. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual gift acknowledgment request packets
The corpus is 64 gift acknowledgment request packets, 0.16 MB (json 5 · jsonl 1 · md 2 · txt 64). Because the failure that matters on an acknowledgment is not filling in fields, it is using a value the gift-ops log has quietly put out of use — a gift held for review, a fund closed to designations, an addressee replaced, an amount re-recorded — and then thanking a donor for something that is not what happened. 39 of the 64 packets can only be decided by reading that log, and that is the population the money is bought for.
The corpus
- The 64 gift acknowledgment request packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere. All 64 packets, data/gift_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your gift acknowledgment request packets. That is the whole change — there is no database to migrate.
MARROWFIELD RIVERS CONSERVANCY - GIFT ACKNOWLEDGMENT REQUEST PACKET
SYNTHETIC SAMPLE DATA. GENERATED, NOT REAL. NO REAL ORGANISATION, DONOR OR GIFT.
Packet GAK-0001 Donor record DNR-72342 Prepared 2026-03-24
GIFT ON FILE
Gift reference GFT-63872
Gift amount $1,000
Gift received 2026-03-21
Fund designated FND-2105
PRIOR GIFT ON THIS DONOR RECORD, NOT THIS ONE
Prior gift reference GFT-60075
Fund the prior gift went to FND-2102
ADDRESSEE RECORDS
Acknowledgment addressee on file ADR-81890
Second addressee on record ADR-81381
SIGNER
Acknowledgment signer on record STF-3563
Signer role Gift Operations Officer
Backup signer STF-3336
Backup signer role Stewardship Manager
PROGRAM LINES THIS FUND SUPPORTS (in the order the letter lists them)
1 Habitat restoration
2 Water quality monitoring
3 Land stewardship
4 Schools and youth programme
ACKNOWLEDGMENT TEMPLATES ON FILE
AKT-2025 for letters prepared before 2026-02-16
AKT-2026 for letters prepared on or after 2026-02-16
MAJOR-GIFT THRESHOLD (OPERATOR-SUPPLIED)
Threshold $5,000
This value is set by whoever operates this deployment. It is not an organisation
figure, no external source is cited for it, and without it the tier is not
assessed and is never defaulted.
GIFT-OPS LOGAbridged — the file continues.
The outcomeWhat a good result looks like
Either a drafted letter of seven fields — gift amount, gift received, designation, addressee, signer role, the program lines in order and the acknowledge-by date — or a decline naming exactly ONE element, the first that fails in the template's own precedence order, with no letter fields at all. Plus the template in force, the tier and the route.
And when it cannot
⛔ THIS KIT DOES NOT BEAT ITS FREE FLOOR OF RECORD. Rechecked, the paid call takes 48 of 64 full rows against the domain floor's 43 — +5 packets against the free floor of record, exact McNemar 15/10, p = 0.424356 — NOT SIGNIFICANT. What IS proven, and decisively: against the LOG-BLIND floor the same call takes 48 against 25, 27/4, p = 0.000034: reading the gift-ops log is worth paying for, not reading it better than a careful rules engine. The reading-required slice (39 packets) reads 27 against 18, p = 0.078354 — suggestive only. ⛔ THE HEADLINE IS A COMPOSITE OF 13 CELLS AND ONLY ONE IS CONTESTED. The call beats the best free arm on addressee alone, +9 of 21 packets of room, and is level or behind on the other twelve. Cell by cell, both columns, against the best shipped free arm: template is 64 of 64 free — ZERO packets of room — and tier leaves 2 and the paid call is 4 BEHIND there, so neither may ever be quoted as a result. On the remaining eleven the paid call is ahead on exactly ONE: addressee, 52 against 43, +9 packets of the 21 available. On the other ten it is level or behind, by as much as 6 packets. A +5 row built from one winning cell and twelve level-or-losing ones is a different product claim from '+5 overall', and only the per-cell table shows it. The eligibility rule is in code — evals/scoring.py::card_eligible with CARD_ELIGIBLE_MIN_HEADROOM = 5, joined by evals/baseline.py::free_cell_check — so no surface has to remember it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- An organisation that can write its template's rules and its log wordings out in code — the free rules/regex floor, and no model call
it takes 43 of 64 full rows at $0.00 and no network; the paid call takes 48 and the paired test cannot separate them (p = 0.424356). - An organisation whose gift-ops log is free prose nobody has enumerated — the paid call
against the same code with the log not read the call takes 48 against 25, p = 0.000034. That is what the money buys. - A desk whose main risk is thanking a donor for a gift that is on hold — the paid call, with a person on every decline
it drafts 6 letters that were never owed against the floor of record's 7 and the log-blind floor's 17 — the one error shape where it is ahead of both. - A desk that only needs the addressee right after a confirmed replacement — the paid call
addresseeis the one cell of thirteen with real room where the call is ahead: 52 against 43, +9 of 21 available. - Anyone who needs the tier, the template or the date arithmetic — free code
template is 64 of 64 for every code reader, tier is 62 against the call's 58, and 7 of the 10 packets the call lost that the floor had went onacknowledge_by.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own rendered request packets and data/gift_records.json at the structured half, rewrite src/packet.py for your layout and data/ack-template.md for your template, then run python3 -m evals.baseline to get your own free floors before buying a single call. The boundary is the ANSWER KEY. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a margin this corpus cannot show is real. That is the case against the best-fitting scenario (“An organisation that can write its template's rules and its log wordings out in code”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet not laid out the way src/packet.py parses it. The parser is positional over one rendering; another organisation's gift-ops export is a new parser, not a new prompt. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the +5 packet margin over the free floor of record is real at all. 64 paired packets give 15 against 10 discordant and p = 0.424356; nothing here separates the two arms, and a larger corpus was not bought. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier — the paid call, after the station. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-donation-ack. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Recorded on a clean checkout with no key configured and no network: python3 -m evals.baseline scores the six free readers over all 64 packets and writes data/floors.json in about a second at $0.00; python3 -m evals.run --selftest, python3 -m evals.boughtcheck, python3 -m evals.streamcheck and python3 -m evals.check_labels all pass with no socket; python3 tools/build_corpus.py --check rebuilds all 68 files byte-identical.










