The business caseThe problem this solves
A household applies to more than one benefit program and each program keeps its own enrollment record — food assistance, health coverage, cash assistance, a child care subsidy. By any given date those records disagree: an address updated in one system and not the others, a newborn still PENDING a client index number in one and listed in another, an income that differs by a few dollars or by a few hundred, a member one system no longer lists. Some of those disagreements change a benefit and some are record-keeping, and the thing that decides several of them is a sentence in the case notes — a merge the client index desk recorded, a merge it struck as entered in error, a supervisor review that was completed or only requested. The cross-match interface a caseworker is handed today compares raw strings, reads no note and says MISMATCH on 64 of the 66 files in this corpus. Opening one household file, matching each member across the program systems by client index number, reading every case note to see whether a merge, an assignment or a review was actually recorded, comparing each field under the procedure's own tolerance and normalisation, and deciding which disagreements touch a benefit.
Audience
The eligibility-operations team working cross-program reconciliation exceptions — and the caseworker who has to decide what to do about a household once the disagreements are named. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual household files
The corpus is 66 household files, 0.17 MB (txt 66). It is generated because it has to be. A real cross-program enrollment extract is the most sensitive record a public agency holds — household composition, income, dates of birth and case notes about named people — and a corpus that could be published would have had the one thing this kit measures, the sentence that records or strikes a merge, stripped out first. So the whole thing is invented, declared and generated from one seed, with the household key, the member key and the income tolerance DECLARED on every file because in the real work they are operator-defined questions.
The corpus
- The 66 household filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every household, program case, client index number and address is invented, and THERE ARE NO NAMES IN THIS CORPUS AT ALL — a member row carries a client index name token, and every note is written by a desk. evals/check_labels.py asserts every name field is a token, sweeps for honorifics, cadence language and citations on every run, and reports 0.
Swap this folder for your own material and the kit is pointed at your household files. That is the whole change — there is no database to migrate.
==============================================================================
CROSS-PROGRAM ENROLLMENT RECONCILIATION -- ONE HOUSEHOLD, ONE AS-OF DATE
==============================================================================
FILE XPR-0001
HOUSEHOLD LINK HHL-30400
AS OF 2026-06-10
PROCEDURE XPR-2026
MATCH KEY client index number (CIN); a PENDING row matches only on an identical
name token AND date of birth
INCOME TOLERANCE 25.00 a month
NAMES not carried; each system prints the client index name token
CURRENCY USD
-- LINKED PROGRAM CASES ------------------------------------------------------
SYS PROGRAM CASE STATUS COMPOSITION EFFECTIVE
FA food assistance FA-5510000 OPEN 2026-05-15
HC health coverage HC-7720000 OPEN 2026-02-01
CA cash assistance CA-2200000 OPEN 2026-02-01
-- HOUSEHOLD FIELDS AS EACH SYSTEM HOLDS THEM --------------------------------
ADDRESS
FA 10 Wren Hollow Lane, Unit 1, Carrowmere Township REGION 3
HC 10 WREN HOLLOW LN UNIT 1 CARROWMERE TOWNSHIP REGION 3
CA 10 Wren Hollow Ln., Unit 1, Carrowmere Township REGION 3
HOUSEHOLD SIZE
FA 2
HC 2
CA 2
MONTHLY INCOME
FA 1,500.00
HC 1,500.00
CA 1,500.00
-- MEMBERS AS EACH SYSTEM LISTS THEM -----------------------------------------
ROW CIN NAME TOKEN DATE OF BIRTH ROLE
FA/1 CIN-3000000 NT-AADB 1979-01-01 head
FA/2 CIN-3091873 NT-Q8SU 2008-01-01 child
HC/1 CIN-3000000 NT-AADB 1979-01-01 head
HC/2 CIN-3091873 NT-Q8SU 2008-01-01 child
CA/1 CIN-3000000 NT-AADB 1979-01-01 head
CA/2 CIN-3091873 NT-Q8SU 2008-01-01 child
Abridged — the file continues.
The outcomeWhat a good result looks like
One household file in, one row out: every conflict key, the value each system holds, which conflicts the impact table marks as changing a benefit, which member rows cannot be matched, which conflicts carry a completed supervisor review, and one of five XPR-2026 verdicts. After the station recheck 66 of 66 households come back whole — ⚠︎ against 61 for the best free code, a margin that is NOT statistically significant (exact McNemar p = 0.0625).
And when it cannot
And what it does when it cannot. On the scored run 66 of 66 replies parsed, nothing stopped at the ceiling and no call failed. The failures are in the call's OWN comparison, which the kit discards: the same replies before the recheck are whole on 18 of 66 and the call's own verdict is wrong on 30 — it missed a composition effective date on 10 households, answered a benefit conflict where the key is a record conflict on 9, and answered CONSISTENT on 4 households that have a conflict. A reply that cannot be parsed is counted WRONG and stays in the denominator.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your client index and supervisor desk record merges, assignments and reviews as status columns — the free modal floor, and do not buy a call at all
50 of 66 households whole for $0.00 when no note is read. Closed cases, suffix normalisation, the printed tolerance, age bands, regions and the ladder are all columns. - Your deciding notes are written to a fixed template — the free rules floor
On this corpus a regex copied off seven templates is whole on 61 of 66 and the paid call is not significantly better (p = 0.0625). - Your notes are free-written, and entries are struck, restated and reversed — the paid call for the three readings — and a harder corpus before you believe any number here
The only households the paid call adds are the three struck entries and the two reversed merges, which is exactly this scenario. What it would score on real free text is NOT measured. - You want the cross-match interface's flag audited — either floor — both beat it
The interface compares raw strings, closed cases included, and says MISMATCH on 64 of 66 files. Believing it gets 52 verdicts right; the modal floor gets 53 and the free code 62.
And where nothing here is good enough:
- Your programs share no client index number, or you need a per-field tolerance table — neither, yet
No file in this corpus lacks a CIN, and income is the only field with a tolerance. The unit of work and every percentage on this page assume both.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own household files in the same five-block shape and data/households.json with your own register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per household for arithmetic you already have. That is the case against the best-fitting scenario (“Your client index and supervisor desk record merges, assignments and reviews as status columns”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | Programs that share no client index number. M-1 keys every member on the CIN; without one every row falls to the exact PENDING fallback and comes back UNMATCHED. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-xprog-enrollment. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, re-runs both free floors live and replays every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.








