The business caseThe problem this solves
A declaration of interests is a form somebody fills in about themselves, and the part that matters most is the part where it says nothing. Two different nothings sit on the same page: a NIL RETURN — "- NIL -", "None.", "NONE DECLARED", "this section does not apply to me" — where somebody has put their name to the absence, and a SILENCE, where the heading is printed with nothing under it or the section is not on the form at all. Collapse the second into the first and a hole in a declaration is on the record as a clean nil return; every conflict check run against those rows afterwards passes over it and reports nothing. The forms also arrive in three revisions that share no layout — typed prose, labelled key/value blocks, pipe-delimited tables — and the answer must not change with the layout. The manual pass over one return: reading it in whichever revision it arrived in, deciding which section each item belongs to, normalising the entity name so it can be matched against a client or restricted list later, and recording an absence in a way that says whether anybody signed it. It does NOT replace the conflict check itself — this kit decides nothing about whether anything is a conflict and has no such output.
Audience
Whoever holds a register of interests and has to turn a stack of returns into rows something can be checked against: independence and conflicts teams, a compliance or ethics office, the person who runs the annual declaration cycle. Also whoever has to chase the incomplete ones, because NIL_DECLARED and NOT_STATED are the two answers that decide who gets chased. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual declaration forms
The corpus is 60 declaration forms, 0.05 MB (txt 60). A declaration of interests is a document about a named private individual and their family, so there is no public set and no prospect of one. What this corpus buys instead is a controlled cross: the SAME planted facts printed in three revisions that share no layout, so the by-revision column measures reading rather than the mix of the sample. It also forces the two distinctions the task turns on — 51 of 60 forms carry BOTH a nil return and a silence, so no arm can score well by picking one absence class and applying it everywhere.
The corpus
- The 60 declaration formsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your declaration forms. That is the whole change — there is no database to migrate.
DECLARATION OF INTERESTS | FORM ID-1 (REV. 2025)
REF PR-27979 | DECLARANT Emeric Skarsgaard-Hume | ROLE Associate Director | UNIT Assurance - Group 4
PERIOD 01/01/2025 - 12/31/2025 (MM/DD/YYYY) | DECLARED 06/03/2025 (MM/DD/YYYY)
VALUE BANDS: A $1 - $5,000 | B $5,001 - $25,000 | C $25,001 - $50,000 | D $50,001 - $250,000
[1] SECURITIES AND OTHER FINANCIAL HOLDINGS
ENTITY | NATURE | HELD BY | VALUE | PCT
Ferrisgate Analytics Inc | ordinary shares | self | $22,100.00 | -
[2] DIRECTORSHIPS AND OFFICES HELD
ENTITY | NATURE | HELD BY | VALUE | PCT
NONE DECLARED | | | |
[3] OUTSIDE ENGAGEMENTS, EMPLOYMENT AND APPOINTMENTS
ENTITY | NATURE | HELD BY | VALUE | PCT
Drossmere Publishing Ltd | consultancy engagement | self | $9,000.00 | -
END OF FORM
The outcomeWhat a good result looks like
One row per declared interest, plus exactly one coverage row for every section that declared nothing — never two, never none. Each row carries the section, the outcome, the holder, the entity as printed AND normalised, the interest type, the amount as exact cents or a low/high range or an explicit not_stated, the percentage as an exact decimal string, and how it is held. 60 forms produced 297 rows on this corpus.
And when it cannot
A silence recorded as a signed nil return. The section then looks answered, the check runs over it and reports nothing, and the gap is invisible to everybody downstream — including to the person who would otherwise have chased it. ⚑ MEASURED AT ZERO ON EVERY ARM OF THIS RUN: notstated_read_as_nil is 0 on the free floor, 0 on the model raw and 0 on the model rebuilt. Nothing here separates the arms on the failure the kit exists to catch.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Declarations that arrive in ONE house layout — a form somebody fills in on screen, or a single revision of a paper return — the free rules floor alone
evals/baseline.py already reads 60.0% of forms exactly and 94.3% precision on row identity for $0.00, and eleven added lines take it to 98.3% — the same as the paid arm, class for class. A layout-anchored reader is the right tool for an anchored layout. - Free-typed returns in many house styles, or a register that has accumulated layouts nobody controls — the model, with the pure-code rebuild behind it
The floor's per-revision column is the argument: as shipped it loses 13.1 recall points moving from the 2025 table to the 2019 typed return (93.9% to 80.8%) while the model loses nothing. The floor's weakness is layout-specific by construction; every new layout is another set of patterns somebody has to write. - You need to know WHICH sections nobody answered, so you can chase them — either arm — this is what the kit is for and both arms do it
NOT_STATED is a first-class outcome with its own row, and notstated_read_as_nil is 0 on every arm measured. NIL_DECLARED recall is 100.0% on the model and 89.7% on the floor as shipped (100.0% patched).
And where nothing here is good enough:
- You want the rows to decide something — clear somebody, flag a conflict, approve a return — neither. This kit has no such output.
It produces the rows a person then checks something against. There is no verdict, no flag and no endpoint that adds one.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own forms as .txt into data/corpus/, add one row per form to data/forms.json (id, vintage, case, hard, ref, declarant, role, unit, declared_on) and write data/gold.jsonl with one record per form carrying declared_on, declarant and rows. ⚠︎ A DECLARATION OF INTERESTS NAMES AN INDIVIDUAL, THEIR FAMILY AND WHAT EACH OF THEM OWNS, AND THE WHOLE FORM REACHES YOUR CONFIGURED PROVIDER VERBATIM. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per form for a job a regex and a normalisation function already do. And do not read the floor's 100.0%% columns as portable — those patterns match the column positions and phrasings this generator writes. That is the case against the best-fitting scenario (“Declarations that arrive in ONE house layout — a form somebody fills in on screen, or a single revision of a paper return”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A scanned or photographed form. Everything here is clean text; nothing has been run through OCR. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | The adversarial arm. evals/injection.py ships and HAS NEVER BEEN FIRED — there is no results/eval-x001-holdings-extract.json and this kit publishes no suppression rate, no robustness claim and no injection number in either direction. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-holdings-extract. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — RE-RUN IN THIS CAPTURE PASS, 2026-08-31, in a scratch copy of the kit with no key configured. tools/build_corpus.py rebuilt data/ byte-identically under PYTHONHASHSEED 0, 1, 97531 and 424242 — four seeds, zero differing files, including data/gold.jsonl. python3 -m evals.run --run-id b000-holdings-extract-rules --floor rules reproduced the committed free-floor result exactly (36 of 60 forms row-exact, identity 94.3/89.2). Nothing in this paragraph needed a key or a network. NOT re-run: the paid arm — r001 is replayed from results/cache-r001-holdings-extract.jsonl and no call was made in this pass.



