The business caseThe problem this solves
A telehealth group's internal roster claims each clinician holds licence X in state Y, of a given class, with an expiry and a status. The state registry record is the authority and it disagrees in ordinary ways. Reconciliation LOOKS like a spreadsheet join and is not one, because it is TWO QUESTIONS THAT COME APART CONSTANTLY: is the roster line wrong, and on which field -- and may this clinician see a patient in that state today? A licence renewed under a new number makes the roster wrong and leaves the right to practise untouched. A married name the register carries and the roster does not makes a name matcher raise a finding every time and changes nothing. A class the roster has promoted means licensed, but not at the class they are being scheduled at. And the expensive one runs the other way: a roster line typed perfectly against somebody else's licence, where name, number, class, status and even the national provider id all agree because the board keyed this clinician's id onto the wrong application, and the only record of it is a sentence in the credentialing correspondence. A credentialing specialist reading a roster line by line against a state registry extract: matching each line to a row, deciding whether a differently-spelled name is the same licensee, checking the number against the renewal chain, the status against any reinstatement, the class against the equivalence table and the board-actions column against the roster -- and then reading the credentialing correspondence to find out which of the resulting differences the file has already answered.
Audience
Credentialing and privileging offices at telehealth groups and multi-state provider networks, the specialists who reconcile a roster against state registry extracts before a credentialing committee sits -- and the scheduling desk that acts on the answer. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual licence roster review packs
The corpus is 48 licence roster review packs, 0.44 MB (txt 48). Because the answer key is the whole product and a real one cannot exist. The measurement is 'which of 288 roster lines really disagrees with the registry, on which field, and may the clinician practise today' -- and no credentialing file comes with that written down. Generating it is the only way to have both the question and a defensible answer, and the generator is gated: evals/check_labels.py re-derives every structured ground FROM THE FINISHED PACKS and asserts fourteen properties, including that pure code reproduces the key exactly on every table-channel line and DISAGREES with it on every case-file line.
The corpus
- The 48 licence roster review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your licence roster review packs. That is the whole change — there is no database to migrate.
Licence Roster Review
----------------------------------------------------------------
Pack LIC-0001
Organisation Harborlight Remote Care
Clinician reference CL-6999
Clinician as rostered Elizabeth Grace Zelenko
National provider id 9367167597
Review date 2026-05-29
Registry extract pulled 2026-05-23
Prepared by Credentialing and Privileging Office
Verification Rules And Index
----------------------------------------------------------------
Roster verification rules
VR-1.1 The state registry record is the authority and a roster line is only a claim about it
VR-1.2 A registry name that differs from the rostered name is matched only where a name change record on file connects the two
VR-2.1 A licence renewed under a new number is matched only where a renewal certificate on file connects the superseded number to the current one
VR-2.2 A registry status other than active bars practice in that state unless a reinstatement notice on file states otherwise
VR-2.3 The licence class the registry prints governs; a rostered class the equivalence table does not carry for it is a mismatch
VR-2.4 A board action printed on the registry record must be carried on the roster line
VR-3.1 A registry record whose national provider id differs is not a record about this clinician
VR-4.1 Every roster line is verified at primary source before each credentialing cycle
Licence class equivalence
RN Registered Nurse
APRN Advanced Practice Registered Nurse
LCSW Licensed Clinical Social Worker
LPC Licensed Professional Counselor
PA-C Physician Assistant
MD Physician and SurgeonAbridged — the file continues.
The outcomeWhat a good result looks like
A drafted reconciliation: one entry per roster line, in the pack's order, each carrying MISMATCH with a named ground of six, the verification rule this pack carries for it, the record that proves it, and WHETHER THE CLINICIAN MAY PRACTISE IN THAT STATE TODAY -- or MATCH, or INSUFFICIENT_EVIDENCE naming what is missing. Plus a whole-file case action. Nothing is suspended, edited, reported or sent.
And when it cannot
⚑ THE PAID ARM WINS THE HEADLINE HERE -- 88.89 pct against the strongest free floor's 75.00 pct over the same 288 roster lines (r001-license-roster against b002-license-roster-registrygate) -- AND THE INTERESTING NUMBERS ARE THE TWO COLUMNS IT STILL LOSES. Free code takes 100.00 pct of the grounds the printed extract proves against the paid arm's 97.50 pct, and it computes the days-to-expiry arithmetic on 100.00 pct of lines against 94.35 pct. Where the pack prints the answer, a paid arm buys nothing and is slightly worse. ⚠︎ AND IT LOSES A THIRD COLUMN THAT MATTERS MORE THAN EITHER: IT WILL NOT SAY IT CANNOT TELL. 12 of 22 unsettleable lines were named as unsettleable, against 22 of 22 for the floor -- and the split is exact: it named every line where the extract carried no row at all, and NOT ONE of the ten where the board had flagged its own record as under revision. It read the printed columns of a record the board itself says is being amended and called them a match. ⚠︎ AND 14 OF ITS 32 MISSES ARE ONE SHAPE WHERE THE FINDING IS ACTUALLY RIGHT: on all 14 wrong-person lines the disposition, the ground and the practice call are correct and the REQUIRED RECORD is not -- it attached the board's own note and the registry row, both printed in the pack, where the key names the primary source verification.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your registry extracts are complete and current, every clinician and every registry row carries the same provider identifier, and your credentialing correspondence rarely contradicts the extract. — the free floor -- registry-gate, $0.00
It scores 75.00 pct on the discriminator, takes 100.00 pct of the grounds the printed extract proves, cites 100.00 pct of their rules, attaches 100.00 pct of their records, names 100.00 pct of the lines it cannot settle, computes the expiry arithmetic on 100.00 pct and invents nothing. It runs all 48 packs in under a second with no key. Sweep NAME_RATIO against your own names first -- it costs nothing. - Your boards e-mail you. Reinstatements, terminated orders, name changes filed with the board and not with you, and keying errors the board has admitted all arrive as prose, and your extract is a nightly cache. — the paid arm -- the fast tier at $0.078569 a clinician
It catches 100.00 pct of the grounds only a sentence settles, including 14 of 14 lines where every printed column agrees and the licence belongs to somebody else. It states a mismatch on 2 of 42 already-answered lines against the floor's 42, and it makes ZERO errors in either dangerous direction on the practice question.
And where nothing here is good enough:
- You want a number to put in front of a committee and you have no answer key of your own. — neither, yet
Every percentage on this page is a comparison against 288 labelled roster lines on a SYNTHETIC corpus. Without your own key you can run this kit and read its output, but you cannot measure it, and the whole claim here is that the second thing is what matters.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace src/roster.py's parse() and keep everything else: the three free floors, the scorer, the prompt and the UI all work off the parsed structure and not off the text. ⚠︎ THE WITHHOLDING SEAM IS A NAMED SECTION AND NOT A REDACTION SYSTEM. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not use it where any decisive fact is a sentence. It catches 0.00 pct of case-file grounds, states a mismatch on 42 of 42 lines the pack has already answered, and clears 14 clinicians who may not practise. That is the case against the best-fitting scenario (“Your registry extracts are complete and current, every clinician and every registry row carries the same provider identifier, and your credentialing correspondence rarely contradicts the extract.”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A pack layout that is not this one. src/roster.py is a set of regular expressions written for these headings, this two-or-more-space registry extract and these key/value blocks; against a real credentialing export it parses nothing and returns empty tables rather than guessing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Repeatability. Every arm was run ONCE. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-26 — r001-license-roster. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, then python3 -m src.app. No install, no key, no network. The corpus, the answer key, all three free floor runs, the threshold sweep and the answer-key gate are committed and reproduce offline; python3 tools/build_corpus.py --seed 4471 rebuilds the corpus byte for byte.



