Home › Use Cases › Tax certification validation
Use caseUC0377
🧪 Use-case kit · runnable

Tax certification validation

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A wealth manager's onboarding desk receives a tax certification form for each account. The form is the client's statement about themselves; the account record is what the firm already holds. Checking one means reading a document that arrives in whatever shape the client sent it — a labelled form, a dotted-leader form, or a prose declaration — and comparing eight fields plus two free-text boxes against the record, then applying a small rulebook about signing capacity, staleness, expiry and benefit eligibility. It is done by eye, one at a time, and the cost of getting it wrong is a client relationship, not a rounding error. the field-by-field comparison an onboarding analyst does by eye between a submitted certification form and the account record it certifies

Audience

The onboarding analyst who works the exception queue, and whoever decides whether a call per form is worth $0.0002. On this corpus the answer is a clear yes for the exception LIST and a clear not-proven for the PASS/QUEUE decision — the free floor already gets 57 of 60 of those. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual submitted certification forms

The corpus is 60 submitted certification forms, 0.07 MB (json 3 · jsonl 1 · md 1 · txt 60). A submitted form is the one document in an onboarding file that the FIRM did not write, so it arrives in whatever shape the client sent it. The corpus makes that the variable: 17 forms in printed layout A (labelled items), 21 in layout B (dotted leaders) and 22 in layout C, where the entity type, the residence and the signing capacity are stated inside a prose declaration in one of four phrasings rather than in a labelled slot; and 28 dates in ISO, 13 as a month name, 8 as dd/mm/yyyy and 11 written out in words. That variety is the whole experiment — it is what separates a reader from a parser.

The corpus

  • The 60 submitted certification formsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your submitted certification forms. That is the whole change — there is no database to migrate.

One submitted certification form, as the model receives itWMC-0001.txt · 1 of 60
NORTHHARBOUR CUSTODY & FUND SERVICES  (fictional)
Form WM-CERT-3 — Client Tax Status Certification
Submitted for validation under standard NH-VAL-2026.
SYNTHETIC DOCUMENT. Every party, jurisdiction, schedule and reference below is invented for evaluation purposes.

Document reference: WMC-0001
Account certified:  ACC-100013
Received by onboarding: 2026-04-07

1.  Name of certifying party .................................. Caldwyn Pension Foundation
2.  Entity type ............................................... PENSION_FUND
3.  Jurisdiction of tax residence ............................. Kerresk
3b. Permanent address (jurisdiction) .......................... Kerresk
4.  Client reference number ................................... TRN-AQC-794699
5.  Treaty benefit claimed .................................... NO
6.  Capacity of signatory ..................................... (left blank)
Item 7   Signed: the twentieth day of March, two thousand and twenty-six
         Stated expiry of this certification: the thirty-first day of December, two thousand and twenty-nine
Item 8   Explanation of address difference:
         (left blank)
Item 9   Beneficial owner declaration:
         (left blank)

Signature: /s/ Mattias Halloway
Printed name of signatory: Mattias Halloway

End of Form WM-CERT-3. This document is synthetic.

The outcomeWhat a good result looks like

Every one of the ten checks carries a verdict, the clause that decides it, and the form value and the account value that disagree — so the query to the client can be written from the report without re-reading the form.

And when it cannot

When a reading is wrong the report is wrong in the most convincing possible way. The station re-applies the rulebook to whatever it was handed, agrees with itself to the letter, quotes the correct clause, and publishes an exception on a form that never had one. On this run that happened 0 times in 720 readings; the board demonstrates it on demand because 0 out of 720 is a measurement about this corpus, not a property of the design.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • you want a worklist — which forms need a human — the free fielded floor
    it gets PASS/QUEUE right on 57 of 60 forms and the paired test cannot separate it from the paid call (p = 0.250000). Paying for that decision buys nothing you can prove.
  • you want to send the client a letter naming the fields that disagree — the paid call
    the exact exception list is 60 of 60 against the floor's 37, p = < 0.000001. A worklist is a flag; a letter has to be right about WHICH clause.
  • your forms arrive in one rigid template from one portal — the free floor, and delete the model
    the strict floor scores 454 of 600 on this corpus and it is the WEAK one; on forms that never vary it would be near-perfect and cost nothing

And where nothing here is good enough:

  • your forms arrive as scans or photographs — neither, yet
    nothing in this kit reads pixels, and every number on this page was measured on typed UTF-8. That job needs a capability stage in front of the model and a different measurement

At a glanceHow the whole thing runs

100%check verdicts pct
1,570 msp50, end to end
$0.57per 1,000 submitted certification forms · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Put real forms in data/corpus as .txt, one per file, and a key line per form in data/gold.jsonl carrying the true readings and judgments; add the matching account records to data/accounts.json. The measured result does NOT travel with the corpus. Corpus lens →
When is this the wrong choice?Avoid: Do not quote the call's 60 of 60 here as if it were a margin — it is not a significant one. That is the case against the best-fitting scenario (“you want a worklist — which forms need a human”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A scanned or photographed form. Every document here is machine-rendered UTF-8 and the kit has no capability stage; a page of pixels reaches the prompt as nothing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE SCORE WOULD HOLD ON A SECOND IDENTICAL RUN. The 60 forms were called once. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-tax-cert-validate. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the 60 account records, Schedule SB-2026, the answer key, the recorded run, the adversarial probe and all three free floors — and scores every free floor offline. evals/check_labels.py and evals/baseline.py both complete with no network. Nothing is pip-installed: the kit is Python standard library end to end.

A living map of modern AI — kept current every morning