The business caseThe problem this solves
A grantmaking organisation issues a letter of findings to a subrecipient, the subrecipient writes back, and somebody on the monitoring desk has to read the reply against the letter finding by finding. The question is never 'is this a good response' — it is the narrower and much harder one: of the four things a finding page prints that a response could be directed at, WHICH ONE did this paragraph actually reach? Two paragraphs can read almost identically and belong in different columns, because the difference is whether the action named touches the condition the finding states, the cause it names, or the recommendation it makes. The first pass over one letter's worth of corrective-action narrative: which condition each response paragraph reaches, and which evidence element it points at. It replaces nothing downstream — no finding is closed, accepted or rejected, no cost is called allowable or repayable, no amount is computed and no date is produced. Measured from a clean checkout on 2026-09-21 with no key configured: PYTHONHASHSEED=0 python3 tools/build_corpus.py --check rebuilds all 64 files and 214 findings and confirms them byte-identical in under two seconds; python3 -m evals.check_labels re-derives all 856 graded cells and exits 0; python3 -m evals.baseline --write-floors scores all four free arms with no key and no network; python3 -m src.app serves the board on 127.0.0.1:9516. Four minutes is the wall clock for those four commands including reading the printed output. No install, no index build, no key.
Audience
The monitoring officer who works the outstanding list, and the programme officer who has to defend a disposition at the next cycle. The decision in front of them is not 'should we use AI' but 'is a free keyword arm enough' — and on the evidence pointer the honest answer is that free code already does better than the paid call. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual finding-response files
The corpus is 64 finding-response files, 0.32 MB (txt 64). A subrecipient monitoring file is the right shape for this question and an impossible shape to obtain: a real one names the subrecipient, names the people who signed, and carries an auditor's findings about an organisation that did not publish them. So it is generated — and generated so that the answer never sits in a column. ⛔ THE FIRST CORPUS WAS THROWN AWAY AND THE REASON IS THE MOST USEFUL THING IN THIS LENS: every action clause paraphrased its own printed line, word overlap decided everything, and a free arm scored 100.00%. The rebuild is structural rather than cosmetic — every response paragraph now names ALL THREE of its finding's matters (the condition, the cause, the recommendation) in one reference string that is also exactly what the finding page prints, exactly one of them is the object of the verb, and the rest sit in a 'having reviewed …' clause in shuffled order, across three sentence forms that invert which positional cue is right. The same arm scores 52.34% after it. Thirteen case families run from a clean control group to three whole files where one printed signature column overrides every paragraph in them.
The corpus
- The 64 finding-response filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your finding-response files. That is the whole change — there is no database to migrate.
SUBRECIPIENT FINDING-RESPONSE FILE FRF-0001
Prepared 2026-09-21 under SMP-2026 | monitoring cycle 2025-07-01 to 2026-06-30
FILE FACTS
Subrecipient SUB-0001
Award reference AWD-2026-1301
Audit period 2025-07-01 to 2026-06-30
Management decision issued 2026-06-25
Response received 2026-08-08
Response signed by the subrecipient's own auditor
Policy revisions in force SMP-2026.1 to 2026-03-31; SMP-2026.2 from 2026-04-01
FINDINGS AS RECEIVED
[F-2026-01]
Condition The file shows the three months of indirect cost charged at a rate the award's terms do not carry.
Cause The finding attributes this to the rate held in the accounting system, which nobody updated when the award was signed.
Effect Indirect cost on the award is charged on terms the award does not carry.
Recommendation The organisation recommends that the subrecipient apply the rate the award's terms carry.
[F-2026-02]
Condition The file shows the costs charged to two budget lines the approved budget does not carry.
Cause The finding attributes this to the chart of accounts, which was never aligned to the approved budget.
Effect The award's ledger and its approved budget do not describe the same thing.
Recommendation The organisation recommends that the subrecipient charge costs only to budget lines the approved budget carries.
[F-2026-03]
Condition The file shows the three charges posted to the award with no supporting invoice attached.
Cause The finding attributes this to the posting procedure, which did not require a supporting invoice before a charge was approved.Abridged — the file continues.
The outcomeWhat a good result looks like
One row per finding: which condition the response reached, one line of the evidence index quoted verbatim or left deliberately empty, and then — in pure code — the disposition, the outstanding flag, whether the prior-period register carries the finding, the evidence state and which of the seven rules produced it. A monitoring officer confirms a row instead of re-reading a file.
And when it cannot
When it cannot read the file it answers UNCLEAR-NEEDS-REVIEW and the finding goes to REVIEW-REQUIRED — a real answer and the only honest one where a finding states no condition, or where the response was signed by somebody other than the subrecipient's own authorised officer. 18 findings in this corpus are exactly that, and the paid arm gets 9 of them. ⚠︎ ITS WORST FAILURE IS THE OPPOSITE ONE: it clears 20 findings of 136 that the key says were not reached, against the free floor's 14 — and two of those are on a file whose printed signature column overrides every paragraph in it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You want the evidence pointer and nothing else — the free floor — evals/baseline.py mode
domain
86.45% for $0.00 and it invents a pointer on 0 of the 79 findings whose key carries none, where the paid call invents 16. The paid call's 91.12% is bought entirely on the other side of that split and is 8.88 points below the free ceiling. - You want a first pass over a whole cycle's responses and your policy is short — the free floor — evals/baseline.py mode
domain
52.34% of reach classes and 45.33% of findings fully right for $0.00 with no network, and only 14 false clears of 136. On a set of files this regular it answers half the job and it is the cheapest thing on this page. - Your files carry actions directed at the cause or the recommendation rather than the condition, or two conditions in one finding — the paid call, and read the reach class on its own
CAUSE-ONLY 33 of 38, OTHER-CONDITION 16 of 16, PARTIAL-CONDITION 15 of 18 and NOT-ADDRESSED 14 of 14, against a floor that cannot separate them at all. OTHER-CONDITION in particular needs the whole letter in one call — test 4 asks whether the action is directed at a SIBLING finding's condition — and the paid arm takes all 16. - You need the disposition, the outstanding flag or the repeat flag — the pure-code station — src/policy.py, whatever produced the reading
None of them is ever asked of a model. They are the card's seven ordered rules over the reach class, the prior-period register and the file's own printed columns, and the station changed the disposition on 22 of r001's findings relative to what the reach class alone would have produced.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/policy.md and data/policy.json with your own monitoring policy — the same eight tests and seven rules in both, in the same order, which evals/check_labels.py enforces. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for it. And avoid reading the tuned ceiling's 100% as achievable on your own files — it has read the generator. That is the case against the best-fitting scenario (“You want the evidence pointer and nothing else”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A reply that is not seven printed panels. The free floor's parser and the label gate's structural checks are both bound to literal headings; a single narrative letter or a scanned PDF returns UNCLEAR-NEEDS-REVIEW on everything, which fails safe and fails completely. 9 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the evidence pointer measures anything at all outside this corpus. The free tuned arm resolves 214 of 214 because the generator wrote both the response paragraph and the evidence index, so the 100.00% ceiling is an artefact. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 9 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-21 — r001-finding-response. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on 2026-09-21 from a clean checkout with no key configured: tools/build_corpus.py --check rebuilt all 64 files and 214 findings and confirmed them byte-identical under PYTHONHASHSEED 0 and 1; --check-fillers proved that no distractor line is ever the labelled line of the finding it sits on, over 135 labelled lines and 64 policy extracts; evals/check_labels.py re-derived all 856 graded cells and exited 0; all four free arms scored the whole corpus with no network; the --stub pass exercised the prompt assembly, the JSON parse, the normaliser, the station, the citation locator and the scorer end to end; and the board rendered every panel on 127.0.0.1:9516 with the model control disabled and the reason printed beside it. Nothing needed installing. What a clean checkout CANNOT do is call a provider, and the scored run, the adversarial arm and the refusal probe replay from their committed result files instead.







