The business caseThe problem this solves
A provider group's credentialing system knows where each clinician practises, what their taxonomy code is, whether they are taking new patients and when their participation started. The payer publishes its own version of all four in a directory a member searches to find care. The two drift — a suite closes and the old row keeps being published, a change is filed and acknowledged and does not appear, a colleague's listing gets keyed to the wrong NPI. Today somebody opens the payer's extract in a spreadsheet next to the roster and eyeballs it, or a monitoring tool string-compares the two and reports every abbreviated address as a defect. Opening one payer extract beside the roster of record, deciding which published row is the clinician's, working out which filed change the payer has already accepted, and comparing four fields by eye.
Audience
The roster desk at a multi-site provider group, working one payer's directory extract at a time, deciding which rows to put on a correction file. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual reconciliation packets
The corpus is 60 reconciliation packets, 0.16 MB (txt 60). It is generated because it has to be. A real payer directory extract beside a real credentialing roster is, by construction, a file of practitioner names, NPIs, addresses and panel status — and this kit's fourth refusal is that a directory defect is never attached to a person. A generated corpus is the only version of this job that can be published at all, and it is also the only way the answer key can be DERIVED rather than typed: each packet is built as a structure and RDR-2026 is applied to the same structure, so a corpus change cannot leave a stale label behind. ⚠︎ Every NPI here begins with 9, which is not an issued prefix, and none carries a valid check digit; evals/check_labels.py asserts it over all 120 distinct identifiers in the corpus.
The corpus
- The 60 reconciliation packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement, including the three defects found by measuring the corpus before a call was bought.
Swap this folder for your own material and the kit is pointed at your reconciliation packets. That is the whole change — there is no database to migrate.
==============================================================================
PAYER ROSTER AND DIRECTORY RECONCILIATION -- ONE CLINICIAN, ONE PAYER, ONE LOCATION
==============================================================================
FILE RST-0001
NPI 9400000017
CLINICIAN REF CLN-4400 (this packet names no person; a clinician is an NPI and a reference)
PAYER PY-31 Cascadia Health Plan
NETWORK commercial PPO
LOCATION LOC-1180 Rowan Field Clinic
SNAPSHOT 2026-07-06 (the payer directory extract this packet reconciles)
UPDATE WINDOW contracted directory update window 2 banking days from acknowledgement
MATCH RULE specialty by taxonomy code; address normalised; flags and dates exact
-- GROUP ROSTER OF RECORD (the credentialing system) -------------------------
FIELD VALUE
panel_status ACTIVE
specialty 207Q00000X Family Medicine
service_address 1180 Rowan Field Road, Suite 300, Bellhaven WA 98004
accepting_new N
effective_from 2026-01-05
terminated_from --
-- PAYER DIRECTORY EXTRACT, LISTINGS AT THIS LOCATION ------------------------
LISTING NPI STATUS SPECIALTY ACCEPT EFFECTIVE ADDRESS
DIR-10001 9400000017 RETIRED 207Q00000X N 2024-12-01 1180 Rowan Field Rd Ste 260, Bellhaven WA 98004
DIR-10002 9400000993 ACTIVE 207Q00000X Y 2026-01-05 1180 Rowan Field Rd Ste 310, Bellhaven WA 98004
DIR-10003 9400000017 ACTIVE 207Q00000X Y 2026-01-05 1180 Rowan Field Road, Suite 300, Bellhaven WA 98004
-- ROSTER SUBMISSIONS FILED WITH THE PAYERS ----------------------------------
SUB FILED PAYER FIELD NEW VALUE ACK STATUSAbridged — the file continues.
The outcomeWhat a good result looks like
One packet in, one row out: which listing is theirs, which submissions are in force, which of the four compared fields disagree, and one verdict from a closed five-value ladder — with the submission row quoted verbatim so the desk can check it without opening the packet.
And when it cannot
And what it does when it cannot. On the scored run 55 of 60 packets came back with all five graded fields right and 5 did not. On two of those five — both not_listed packets — the arm read the listing set correctly and then returned an EMPTY cite set, dropping the in-force submission the key still names. On a third it answered MISSING-FROM-DIRECTORY where the roster shows the NPI terminated and the payer still publishes a live row, which is EXTRA-IN-DIRECTORY. Both free floors get all three.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- your defects are in a COLUMN — a status flag, an NPI, a date — the free modal floor in evals/baseline.py
it gets 39 of 60 packets here for $0.00 and it is the same code path the paid arm's answers go through - your defects are in a SENTENCE — a closure, a withdrawal, a mis-key — the paid call, on the two readings only
14 of 15 on the packets where a note decides the reading, against 0 of 15 for the modal floor - you need the row a member would see checked, not the row the roster holds — this kit unchanged, with your own extract
every chosen listing must agree, not any of them — a row that disagrees is a wrong answer to a member whatever the other row says - you want a defect RATE you can publish to a payer — the free floors first, then the paid call on the families they miss
the two arms disagree on 22 of 60 packets and the disagreement is concentrated in three families; a rate that mixes them is a rate nobody can act on - you need a correction FILED, not a row reported — something else — this kit refuses, in code
the answer contract has no field that could submit, attest, retire or panel, and src/prompt.py fails the build if one is added
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt and data/roster.json with your own packets in the same block layout — header, roster of record, directory extract, submission log, monitoring recap, notes — and everything under src/ and evals/ works unchanged. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per row for a comparison your extract already answers. That is the case against the best-fitting scenario (“your defects are in a COLUMN — a status flag, an NPI, a date”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A listing row whose NPI column is blank or malformed. R-3 turns entirely on that column; with it absent the reading falls back to the address and this kit HAS NO TEST for it. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the two citation floors (coverage 0.60, precision 0.30) are the right two. They were chosen before the first call from the shape of a submission row and have never been swept. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, as answered, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-roster-reconcile. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, replays all four recorded arms, runs all three free floors on any packet and re-scores every committed run — measured, not asserted: the empty frame in docs/shots was shot against a server started with API_KEY blanked.






