The business caseThe problem this solves
A benefits office mails notices — redeterminations, verification requests, appeal rights — and a share of them come back undeliverable. Each returned piece carries a name as printed, an address as printed and a carrier endorsement, and somebody has to decide which case file on the caseload it belongs to before anything can be fixed. Today a clerk does it by eye against a search screen. the case-file lookup a returned-mail clerk does by eye — reading a damaged name and address off an envelope and searching the caseload for the household it belongs to
Audience
The returned-mail desk that works the pile, and the supervisor deciding whether to automate any of it. The decision is not 'is this accurate' — it is 'is this accurate enough that a caseworker can approve its proposals without re-doing them, and safe enough that the ones it gets wrong do not move somebody's address'. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual returned mail pieces
The corpus is 90 returned mail pieces, 0.19 MB (json 3 · jsonl 1 · md 1 · txt 90). ⚠︎ EVERY BYTE OF IT IS INVENTED, AND THAT IS STATED BEFORE ANY NUMBER. This kit's subject is BENEFIT CASE FILES — among the most sensitive records a public agency holds — so there is no real one anywhere in it. There is no Calder County, XA is not the postal abbreviation of anywhere, the five programmes do not exist, and not one name came from any list of real people. It is generated rather than collected for a second reason: the hard part of this job is REFUSING, and a real pile contains the hard shapes by accident at an unknown rate. Here they are planted in counted quantities — 8 buildings where two households share a surname and an initial, 7 envelopes addressed to somebody the caseload has never heard of, 14 cases carrying the same full name as another, 9 pieces where one candidate fits the name and a different one fits the address. 30 of the 90 pieces have no confident match, so refusing on everything is right 33.3% of the time and has read nothing.
The corpus
- The 90 returned mail piecesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your returned mail pieces. That is the whole change — there is no database to migrate.
RETURNED MAIL - UNDELIVERABLE PIECE RM-0001
Calder County Benefits Office - Returned Mail Desk (a fictional office)
Caseload index: CALDER-XA-2026
Notice mailed: 2026-03-28 - Verification request
Carrier endorsement: FORWARD TIME EXPIRED - RETURN TO SENDER
PART A - WHAT IS PRINTED ON THE ENVELOPE
BEATRIZ IYENGAR
360 FAIRHOLM WAY
NEW SHELDON XA 58231
PART B - CANDIDATE CASE FILES PULLED FROM THE CASELOAD
Each candidate below is a case file the desk pulled for this piece. Exactly one of them
may be the household this piece was addressed to - or none of them may be.
CASE EAG-2026-41674 - Energy Assistance Grant
Name on file: Florentina Iyengar
Address on file: 78 Ockham Rd, Apt 3B, Esterhaz XA 58252
On file since: 2026-01-26
Prior addresses: none recorded
Last contact: 2026-11-05
CASE HSA-2026-41610 - Household Support Allowance
Name on file: Bartholomew Mbeki
Address on file: 751 Fairholm St, Unit 6, Quarry Mead XA 58271
On file since: 2025-04-21
Prior addresses: none recorded
Last contact: 2026-05-18
CASE MAP-2026-41987 - Medical Assistance Program
Name on file: Beatriz Iyengar
Address on file: 345 Marlowe Ct, Winnow Park XA 58240
On file since: 2025-04-01
Prior addresses: 360 Fairholm Way, New Sheldon XA 58231 (until 2025-05-10)
Last contact: 2026-09-21
CASE HSA-2026-41083 - Household Support Allowance
Name on file: Nkechi Iyengar
Address on file: 113 Windrose Ter, Bellhaven Flats XA 58264
On file since: 2024-10-04
Prior addresses: none recorded
Last contact: 2026-08-12
CASE CCS-2026-40382 - Child Care SubsidyAbridged — the file continues.
The outcomeWhat a good result looks like
Every returned piece gets one of five answers — four discrepancy reasons naming one case file, or NO_CONFIDENT_MATCH — with the line of the case file the answer rests on, and a list of the candidates that were ruled out.
And when it cannot
On this corpus the paid call named the WRONG household on 5 of 90 pieces (5.56%), all of them at buildings where two case files share a surname and a first initial. The free code floor made ZERO. It also over-refused 6 pieces, which costs a clerk time and costs a household nothing — the two are not the same error and are never averaged here.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- a returned-mail desk with a caseload that has few shared surnames per building — the free record-linkage floor in evals/baseline.py
it scored higher on every question here and made zero wrong matches, for $0.00 and no network - a caseload with many twin buildings — shared surnames, missing unit numbers — the free floor, with its margin threshold raised
the margin test is the control that refuses a tie, and it is four lines of code. Raising it trades over-refusals for safety in the direction a benefits office should trade - you want the REASON as well as the match, in a caseworker's words — the paid call, knowing it is not better at it
it writes a usable note and a ruled-out list that a clerk can check, which the floor does not. It is not more ACCURATE about the reason — 52 against 54, p=0.790527
And where nothing here is good enough:
- you are deciding whether to automate this desk at all — neither, yet — run the free floor over YOUR pile first
it costs nothing and no data leaves the machine, and the mix of hard shapes in your caseload is most of the answer
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/caseload.json with your own case files and data/pieces.json with your own returned pieces, in the same shape, then regenerate the packets and re-derive the key. THE MEASURED ACCURACY DOES NOT TRAVEL WITH YOUR CASELOAD, AND ON THIS KIT THE MIX IS MOST OF THE ANSWER. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a model on the strength of this measurement. That is the case against the best-fitting scenario (“a returned-mail desk with a caseload that has few shared surnames per building”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a caseload with REAL duplicate records — two live case files for one household. Every candidate here is a distinct household by construction, and a desk's hardest pieces are the ones where the duplication is in the caseload rather than in the mail. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | whether a larger model tier would refuse the twin buildings — this kit bought one tier. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-returned-mail. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the answer key, the committed paid run and all four free floors — and scores every free arm offline at $0.00. A key is needed only to make a new call, and the board has no button that does.







