Home › Use Cases › Check a drafted grant award letter and schedule against the board minute
Use caseUC0332
🧪 Use-case kit · runnable

Check a drafted grant award letter and schedule against the board minute

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A board votes an award — a grantee, an amount, a period, a restriction, a reporting date and a number of instalments — and a grants officer then drafts the letter and the payment schedule that go out under it. Between the vote and the signature the paperwork has to be checked against the minute, and the checking is done at the end of a grant round on several dozen packets at once. The two things that cost the most money are the two a reader skims: a letter addressed to a related legal person rather than the one the board approved, and a purpose paragraph that grants broader terms than the board voted. Both look like ordinary drafting. Reading a drafted award letter and a payment schedule against a board minute line by line — the addressee against the approved grantee, the figure against the approved amount, the words against the figures, the period against the approved period, every instalment date against the period window and the instalment sum against the award — and deciding from two paragraphs of prose what restriction the letter actually grants and what date it actually sets.

Audience

A grants officer, a programme director or a foundation's finance lead who reviews drafted award packets before they are signed, and anyone who has to answer an auditor asking how the letter that went out relates to the minute that authorised it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual drafted grant award packets

The corpus is 60 drafted grant award packets, 0.12 MB (json 4 · jsonl 1 · md 2 · txt 60). Because the shape of an award-letter defect is not the arithmetic — it is a paragraph. Every packet had to carry a summary block that a column reader believes and an operative sentence underneath it that governs, and the corpus had to build the disagreement in BOTH directions: 14 checks where the block hides a real defect and 7 where it raises one that is not there. A corpus with only the first direction makes any arm that reads more than the block look infallible.

The corpus

  • The 60 drafted grant award packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 60 packets, data/minutes.json and the whole answer key are generated in-process by the file that sits beside them. Every foundation, grantee, resolution number, EIN, programme, capital item, date, amount and sentence is invented, and there is no personal data anywhere. GRANTQC-2026 is invented too and is not an accounting standard, a charity regulator's guidance, the Uniform Guidance or anybody's grants manual. See data/SOURCES.md.

Swap this folder for your own material and the kit is pointed at your drafted grant award packets. That is the whole change — there is no database to migrate.

One drafted grant award packet, as the model receives itAW-0001.txt · 1 of 60
GRANT AWARD PACKET - LETTER AND SCHEDULE QC

PACKET HEADER
  Packet             AW-0001
  Foundation         The Hallowbrook Trust
  Programme          Community Health
  Prepared by        M. Hallberg, Grants Officer
  Prepared on        2026-12-02

BOARD MINUTE EXTRACT (the board's own record of what it approved - the authority for every check)
  Resolution         RES-2027-100
  Meeting held       2026-11-02
  Grantee approved   Riverbend Community Health Collective
  Grantee EIN        20-1000000
  Award approved     $71,000.00
  Award in words     seventy-one thousand dollars
  Grant period       2027-01-01 to 2027-12-31
  Restriction voted  program-restricted
  Purpose voted      the direct costs of the after-school reading scheme
  Reporting set      2027-04-01
  Instalments voted  2

DRAFT AWARD LETTER (as drafted by the grants officer; not issued, not signed)
  Addressed to       Riverbend Community Health Collective
  Grantee EIN        20-1000000
  Amount in figures  $71,000.00
  Amount in words    seventy-one thousand dollars
  Grant period       2027-01-01 to 2027-12-31
  Restriction stated program-restricted
  First report due   2027-04-01
  Purpose stated     The funds are provided for the after-school reading scheme alone, and any balance unapplied at the end of the period is returnable.
  Reporting stated   A first progress report is due on 2027-04-01, covering activity from the commencement date.

PAYMENT SCHEDULE
   #  Due date     Amount          Note
   1  2027-01-15   $35,500.00      on execution of the award letter
   2  2027-12-05   $35,500.00      mid-period
    SCHEDULE TOTAL $71,000.00

PREPARER NOTES
  This is the second award to this organisation under the current programme.

SIGN-OFF

Abridged — the file continues.

The outcomeWhat a good result looks like

Each of eight numbered checks carries a verdict, the clause of the rulebook it rests on, the money it puts at issue to the cent and with its sign, and one row copied verbatim out of the packet as evidence — then ISSUE or REVISE for the packet, the exact set of checks to revise, and the total. 478 of the 480 checks are answered with the right verdict once the rulebook is re-applied in code to the model's two readings, and 59 of the 60 packets get the right recommendation.

And when it cannot

⚠︎ THE PACKET-LEVEL NUMBER IS 39 OF 60 AND THAT IS NOT A TYPO. packet_all_correct requires all four graded fields right on all eight checks AND both readings AND the recommendation AND the revise set AND the total. The verdict column reads 478 of 480 and the packet column reads 39 of 60, and the whole distance between them is one behaviour: the reply quoted a row on 20 of the 434 checks that pass, where the answer contract says null. Set the citation column aside and the same 60 answers score 58 of 60 packets. Both numbers ship.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your drafting officers code the summary block correctly and your arguments are about names, figures, dates and sums. — the free block-line floor alone — python3 -m evals.run --floor rules
    Every one of those is a string comparison, a date test and two additions. The floor is 459 of 480 check verdicts and 39 of 60 packets for $0.00, no key and no network.
  • Your purpose paragraphs are drafted freely and the summary block is a convenience nobody checks. — the paid call, with the station
    That is the sub-population this kit exists for: 21 checks, and the block-line floor gets 0 of them while the call gets 20.
  • You need the quoted row to be right because the output goes to whoever redrafts the letter. — any of the free floors
    A pure-code arm quotes the row it is talking about by construction — src/rules.py::citation_rows is the same function the key was planted with. The paid arm has to find the row from a prose rule and returned one on 20 checks that pass, where the contract says null.
  • You want a number you can defend to an auditor. — the whole kit, and read data/SOURCES.md first
    The key is derived, an independent gate retypes the rulebook and verifies 459 of the 480 checks without importing the engine, and every arm's result file is committed.

At a glanceHow the whole thing runs

98%packet recommendation correct rechecked pct
3,323 msp50, end to end

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported packets and data/minutes.json at your own minutes. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for a reading you do not need. That is the case against the best-fitting scenario (“Your drafting officers code the summary block correctly and your arguments are about names, figures, dates and sums.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet whose panels are not the six this parser knows. The parser finds panels by their headings and labelled rows by their labels; a differently laid-out packet parses to nothing and every check comes back MISSING. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates a model's variance from a real difference. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-award-letter-qc. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run python3 -m evals.baseline with no key, no network and nothing installed: all four floors score in about a second, and the best of them takes 479 of 480 check verdicts.

A living map of modern AI — kept current every morning