The business caseThe problem this solves
A state sends a coverage roster every cycle; the plan holds its own enrolment table; the two disagree on thousands of members. The work is not finding the members — a dictionary does that — it is deciding, for each one, whether the COVERAGE INTERVALS agree and what kind of disagreement it is when they do not. Retroactive additions, retroactive terminations, breaks in cover, double coverage, members on one file only and benefit-segment mismatches all look identical in a row-by-row comparison and need completely different correction transactions. An analyst then has to read the case notes, because a caseworker has often already recorded why the extract is wrong. The manual pass an analyst makes over a cycle's roster exceptions before keying corrections.
Audience
An enrolment reconciliation analyst at a health plan, and the enrolment operations lead who decides how many of them the plan needs. Not an eligibility worker: nothing this kit produces is an eligibility determination, and there is no field in its reply schema that could be one. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual roster reconciliation packets
The corpus is 40 roster reconciliation packets, 0.09 MB (json 4 · jsonl 1 · md 2 · txt 40). The corpus is shaped to answer one question honestly: is a language model worth paying for on a job that is mostly interval arithmetic. A corpus made only of date columns would measure whether the model can subtract dates — and it would misrepresent the job, because a real analyst works from the extracts AND the remarks log, where a caseworker has typed the reason the extract is wrong. So 48 of the 200 member cases carry a remark that SUPERSEDES the structured fields, each written in FOUR paraphrases that share the fact and almost none of the vocabulary; 22 more carry a HARD DECOY that uses the same vocabulary and negates it ("an appeal was filed; no decision has been issued"); 13 carry a benign note; and 117 carry none. A floor that reads only the date columns is right about 160 cases and wrong about 40. A floor that treats any remark as an override is wrong about the 22 hard decoys. That gap is the measurement. The corpus was rebuilt once for exactly this reason: the first build had benign decoys only, the tuned floor scored 183 on the class and 200/200 on the pivot date, and a corpus a keyword count can solve cannot answer the question this kit asks.
The corpus
- The 40 roster reconciliation packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your roster reconciliation packets. That is the whole change — there is no database to migrate.
CALDERRA DEPARTMENT OF HEALTH COVERAGE — CRX-9 ROSTER RECONCILIATION PACKET
PACKET RCN-0001 CYCLE 2026-01 REGION Vale County
PLAN OF RECORD: Meridian Community Health Plan PROGRAMME: Calderra Health Assistance Program (CHAP)
An open-ended segment is written THRU 9999-12-31. Both COVERAGE FROM and COVERAGE THRU are
INCLUSIVE. A member may appear on more than one line of either extract.
PART A — STATE ROSTER EXTRACT (CRX-9 segment feed, as transmitted 2026-01-03)
SEG MEMBER ID MEMBER NAME COVERAGE FROM COVERAGE THRU BENEFIT SEGMENT
A01 MBR-ZZ-64994 CORVINO, CASIMIR 2024-12-01 2025-01-31 CHAP-LTS-D7
A02 MBR-ZZ-64994 CORVINO, CASIMIR 2025-06-01 2025-12-31 CHAP-LTS-D7
A03 MBR-ZZ-55753 STROUD, PERPETUA 2025-09-01 2026-01-31 CHAP-CHIP-G8
A04 MBR-ZZ-16230 MARCHETTI, PERPETUA 2025-07-01 2026-06-30 CHAP-STD-A2
A05 MBR-ZZ-83846 SABATINI, CASIMIR 2025-02-01 2025-08-31 CHAP-STD-A2
A06 MBR-ZZ-20364 NAKASHIMA, MIRELA 2025-09-01 2026-06-30 CHAP-CHIP-G8
PART B — PLAN ROSTER EXTRACT (Meridian Community Health Plan enrolment table, as at 2026-01-03)
SEG MEMBER ID MEMBER NAME COVERAGE FROM COVERAGE THRU BENEFIT SEGMENT
B01 MBR-ZZ-64994 CORVINO, CASIMIR 2024-12-01 2025-12-31 CHAP-LTS-D7
B02 MBR-ZZ-55753 STROUD, PERPETUA 2025-09-01 2026-07-31 CHAP-CHIP-G8
B03 MBR-ZZ-16230 MARCHETTI, PERPETUA 2025-07-01 2026-06-30 CHAP-STD-B1
B04 MBR-ZZ-83846 SABATINI, CASIMIR 2025-07-01 2025-08-31 CHAP-STD-A2
B05 MBR-ZZ-20364 NAKASHIMA, MIRELA 2025-09-01 2027-01-31 CHAP-CHIP-G8
Abridged — the file continues.
The outcomeWhat a good result looks like
Every member in the packet carries one of seven findings and, where there is a discrepancy, the single date the correction transaction is keyed from — with the evidence that decided it quoted from the packet.
And when it cannot
The failure that costs something is a discrepancy reported as ALIGNED: nobody looks at that member again, and a coverage record stays wrong until somebody notices downstream. The paid run did that 6 times in 200. This kit's own free code did it 0 times. The second failure is a right finding with the wrong pivot date — the correction is keyed from the wrong day — which happened 8 times on members whose class the run got right.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- you need to reconcile coverage spans between two rosters — the free
intervalfloor — 0 calls, $0.00, no network, no key
it scores 152 of 152 on the cases the date columns answer, which is everything a correctly-transmitted feed contains. The paid call scores 141 on the same cases - your case notes routinely overturn the extracts and you cannot enumerate the wordings — a reader — but measure it on your own notes first, and put the interval algebra in front of it
the paid call scores 28 of 48 on the remark cases against plain interval code's 8, so it does read them. It is the only thing on this page it is better at - your caseworkers write notes from a fixed set of templates or a drop-down — regexes. The
templatefloor scores 200 of 200
if the phrasings are knowable in advance, a regex is exact, free, instant and auditable, and nothing a model does can improve on exact
And where nothing here is good enough:
- you want a number you can quote about your own roster feed — neither, until you re-run the eval on your own labelled cases
CRX-9 is invented, the remarks come from 24 wordings, and the intervals are clean monthly boundaries. Real feeds carry mid-month effective dates, corrected member ids and notes nobody enumerated
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point src/roster.py at your own two extracts. ⚠︎ DO NOT PUT REAL MEMBER DATA THROUGH THE PAID ARM WITHOUT DECIDING THAT FIRST. Corpus lens → |
| When is this the wrong choice? | Avoid: Sending real member data to a provider to do arithmetic. There is no version of this comparison in which that is the cheaper or the safer choice. That is the case against the best-fitting scenario (“you need to reconcile coverage spans between two rosters”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a roster whose end dates are EXCLUSIVE rather than inclusive. Every off-by-one in this problem domain comes from mixing the two conventions, and the prompt states the convention in one line — a feed with the other one silently shifts every RETRO_TERM pivot by a day. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | whether the free floor's margin survives on a REAL remarks log. Its lexicons were written by somebody who had read this corpus, and the corpus generates its superseding remarks from 24 wordings. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is withheld on published pages by standing direction — src/adapters/ is the only place the vendor is named, and .env is the only place the model id is; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier (MEASURED — the published run). Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r002-roster-span. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, python3 evals/run.py --floor best --run-id b000-roster-span-best, and the whole measurement reproduces at $0.00 with no key, no network and no install — the kit is Python standard library end to end. python3 tools/build_corpus.py rewrites data/ to the same bytes under any PYTHONHASHSEED; python3 evals/check_labels.py re-checks the key it just wrote.






