The business caseThe problem this solves
A grant application arrives with its attachments — a budget, a board list, a tax determination letter, audited financials, a project narrative. Before a programme officer reads a word of the proposal, somebody checks it against the foundation's published checklist. Two things make that check harder than a tick list. The right document is often enclosed in the wrong FORM — a compilation where an audit is required, last year's statements, a board list with no affiliations, a determination letter naming the fiscal sponsor rather than the applicant. And several clauses only apply CONDITIONALLY: the audit clause fires above $750,000 of expenses, the fiscal sponsor clause only for a sponsored project. A screen that cannot read the condition chases a volunteer-run applicant for paperwork nobody asked them for, which is exactly how a completeness screen becomes a barrier to the applicants a fund exists to reach. the clause-by-clause read a grants administrator does against the published checklist before an application reaches a programme officer
Audience
The grants administrator who screens the application and the applicant who will be told what to send. The decision is whether a model call per application is worth buying over the checklist you would write instead — and on one of the three questions the honest answer here is no. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual screen packets
The corpus is 40 screen packets, 0.24 MB (json 4 · jsonl 1 · md 2 · txt 40). ⚠︎ EVERY BYTE OF IT IS INVENTED, AND THAT IS STATED BEFORE ANY NUMBER. The Merrivale Trust is a fictional foundation, the Community Resilience Fund a fictional programme, and CRF-7 a fictional checklist written for this kit; its part numbers, clause numbers and dollar thresholds are made up, and the taxpayer identification numbers are EIN-shaped with the prefix 00, which is not issued to anybody. It was generated rather than found because the measurement needs a key at the level of one clause of one application AND needs the corpus shaped against the shortcuts a reader might take: 124 of the 384 attachments cite the clause they answer and a wrong-form one cites it as often as a satisfying one; 86 clauses do not apply to the applicant at all and 35 of those have the document enclosed ANYWAY; 17 attachments name a document and enclose nothing; 53 satisfying attachments are reworded end to end while keeping every required particular; and 159 attachments answer no clause at all.
The corpus
- The 40 screen packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your screen packets. That is the whole change — there is no database to migrate.
GRANT APPLICATION COMPLETENESS SCREEN — APP-0001
Programme: Community Resilience Fund, Round 7 (fictional) · The Merrivale Trust (fictional)
Applicant: Ashcombe Riverside Gardens (fictional) · Taxpayer id 00-1400037 (invented, not a real EIN)
Checklist: Community Resilience Fund, Round 7 — Eligibility and Completeness Checklist (a fictional checklist written for this kit) · Reference CRF7-02217
PART A — WHAT THE APPLICANT DECLARED ON THE COVER SHEET
These declarations are the applicant's own. Several clauses of PART B apply only
when a declaration below meets the condition the clause states.
The applicant is independently incorporated and holds its own exemption determination.
Total expenses in the most recent completed fiscal year were $2,597,093.
The applicant employs no paid staff and is run entirely by volunteers.
The amount requested is $59,000.
No part of the amount requested is for equipment or building works.
No part of the amount requested is passed to another organisation.
The applicant has previously held a grant from this fund, awarded in 2021.
PART B — WHAT THE PUBLISHED CHECKLIST REQUIRES
Each clause below is a completeness requirement. An application answers a clause
only if something actually enclosed responds to THAT clause, in the form the
clause states.
CRF-7-2.4.d Part 2 Application and Signature — Project narrative
Submit a project narrative of not more than eight (8) pages, covering the need addressed, the
activities proposed, the population served and the timeline, in the programme's published
narrative order, with every heading used.
CRF-7-4.1.a Part 4 Financial Documentation — Audited financial statementsAbridged — the file continues.
The outcomeWhat a good result looks like
Every checklist clause gets one of four verdicts and the evidence behind it: the attachment and the words it rests on where the application answers, the applicant's own declaration where the clause does not apply.
And when it cannot
It made 4 errors in 320. Two are a project narrative whose content list drops one item the clause names — the key calls that WRONG_FORM and it answered SATISFIED. One is an in-scope audit clause with nothing enclosed, reported as WRONG_FORM instead of MISSING. And one is the error that matters most: APP-0030 declared $22,000 of equipment against a clause that fires at $25,000, and the run chased it as MISSING. That is 1 of the 86 clauses that do not apply; the tuned free floor got it wrong 9 times and the plain clause checklist 51 times. It also waved 2 real gaps through as satisfied, against the free floor's 56.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- you need a list of documents nobody enclosed at all — the free
bestfloor — 0 calls, $0.00
it scores 231 of 234 on presence against the paid call's 233. 3 clauses to 1, exact two-sided p = 0.625000: on this corpus there is NO MEASURABLE DIFFERENCE, and the code costs nothing and reaches no network - your checklist has conditional clauses and you care about not chasing applicants for paperwork they do not owe — the paid call
it chased 1 of the 86 out-of-scope clauses against the tuned floor's 9 and the plain clause checklist's 51, and it is right on all 35 of the ones where the applicant enclosed the document anyway - you need to know whether the enclosed document is in the FORM the clause asks for — the paid call, and nothing else on this page comes close
171 of 173 against the tuned floor's 120 — 51 discordant clauses to 0, p < 0.000001. The clause checklist scores the CONSTANT: 103 of 173, which is what answering SATISFIED to every enclosed document gets you
And where nothing here is good enough:
- you want to decide anything about the application itself — none of this
there is no endpoint here that declines, scores, ranks or funds. The reply schema has four fields and none of them is a decision, and the four boundary counters were 0 on every one of the 40 calls and 0 under all 12 attacks
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/checklist.json with your own — clause ids, document types, requirement text, the particulars each clause names and the CONDITION sentence that decides whether it applies — and rewrite the clause table in tools/build_corpus.py so the generator can build packets from it. THE MEASURED ACCURACY DOES NOT TRAVEL WITH YOUR CHECKLIST. Corpus lens → |
| When is this the wrong choice? | Avoid: Quoting its four-way verdict — it waves 56 real gaps through as satisfied and finds only 17 of the 70 wrong-form attachments. That is the case against the best-fitting scenario (“you need a list of documents nobody enclosed at all”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a condition that needs TWO declared facts, or a document, or a conversation. Every condition in this checklist is settled by exactly one figure or one fact on the cover sheet, and the whole applicability result rests on that. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | whether any of these rates survives on a REAL checklist. CRF-7 is invented, its wrong forms come from a published list of degraded values, and every one of its six conditions is settled by exactly one declared figure or fact. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-applicant-screen. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the answer key, the committed run and all four free floors come off disk, and the free floors re-score offline at $0.00. That state is what the empty screenshot is shot against: a second server started with API_KEY blanked in its own environment, not a banner drawn to look like one.







