Home › Use Cases › Reconcile CMS transaction reply files and cluster each reject by root cause
Use caseUC0236
🧪 Use-case kit · runnable

Reconcile CMS transaction reply files and cluster each reject by root cause

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

CMS returns a transaction reply file against the enrollment transactions a plan submitted. Each line carries a reply code, the beneficiary segment CMS holds and the transaction as it went, and somebody has to decide what actually went wrong. THE LOOKUP IS NOT THE JOB: a reply code names its own root cause only some of the time, and where it names two an analyst has to read the transaction against the segment before anyone can file anything. The failure that costs money is never a wrong lookup. It is a reject clustered under the wrong root cause -- which sends the correction down a queue where nobody re-files anything, so the enrolment simply does not exist -- and a corrective transaction drafted with an effective date that re-rejects, which looks filed, clears the worklist, and comes back next month saying the same thing. Neither reports anything downstream. An enrollment analyst reading each reply line against the segment CMS holds -- comparing the package submitted against the packages open at that segment, deciding whether a name that differs by a hyphen is a demographic difference or nothing at all, working out whether the effective date requested clears BOTH entitlement starts and the retro window this file opens, noticing that one reply code is exempt from that window, and then reading the submitting desk's own note on the lines where two explanations are both true and nothing printed decides between them.

Audience

Medicare Advantage and Part D enrollment reconciliation desks working transaction reply files, and the people who build tooling for them. ⚠︎ ITS OWN ANSWER TO THIS AUDIENCE IS A QUALIFIED NO: on this corpus a free floor scores higher than the paid call, and the reason is a fact about the corpus rather than about the job. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reply-file packs

The corpus is 60 reply-file packs, 0.39 MB (txt 60). ⚑ A TRANSACTION REPLY FILE CANNOT BE PUBLISHED, BY ANYBODY, EVER. It is a per-beneficiary record of enrolment activity between a plan and CMS, and the fields being reconciled -- the identifier, the entitlement dates, the demographics as submitted -- ARE the sensitive part. There is no public corpus, no redaction that would make one publishable, and there never will be one to fetch. And the prose is not what is measured: what is measured is a decision and the line behind it, and both need facts that are KNOWN -- which root cause, which queue, which effective date to the day, which lines cannot be settled at all, and exactly which line establishes it, to the character. So the facts are planted, the key is DERIVED from them by src/policy.py, and the packs are written around them.

The corpus

  • The 60 reply-file packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.

Swap this folder for your own material and the kit is pointed at your reply-file packs. That is the whole change — there is no database to migrate.

One reply-file pack, as the model receives itRR-0001.txt · 1 of 60
TRANSACTION REPLY FILE PACK RR-0001 -- daily reply file, afternoon batch

PACK FACTS
  Submitting contract                H9858
  Reply file received                2026-03-05
  CMS processing month               2026-02
  Retro window                       3 months -- the earliest effective date a corrective
                                     transaction may carry is 2025-11-01, unless the reply
                                     code is marked retro-eligible below
  Reconciliation procedure           RR-2026, read as at 2026-08-31

REPLY CODE TABLE IN FORCE -- RC-2026, read as at 2026-08-31

  R-101   accepted   Enrollment accepted as submitted
                     NOT_A_REJECT
  R-104   accepted   Disenrollment accepted as submitted
                     NOT_A_REJECT
  R-107   accepted   Plan benefit package change accepted as submitted
                     NOT_A_REJECT
  R-212   settled    Identifier submitted is not on file
                     IDENTIFIER_INVALID
  R-224   settled    A transaction covering this period is already on file
                     DUPLICATE_TRANSACTION
  R-241   settled    Beneficiary is enrolled with another organisation for the period requested
                     OTHER_PLAN_ENROLLMENT
  R-256   settled    Effective date requested is outside the window this transaction accepts
                     EFFECTIVE_DATE_INVALID; RETRO-ELIGIBLE
  R-231   ambiguous  Beneficiary record does not support the transaction as submitted
                     consistent with DEMOGRAPHIC_MISMATCH or PLAN_CODE_MISMATCH -- the transaction as submitted decides
  R-238   ambiguous  Coverage requested cannot start on the date submitted
                     consistent with EFFECTIVE_DATE_INVALID or ENTITLEMENT_GAP -- the transaction as submitted decides

Abridged — the file continues.

The outcomeWhat a good result looks like

A drafted worklist: one entry per reply line, in the pack's order, each carrying the root-cause cluster, the queue it goes to, the corrective effective date where a correction is drafted, and the line of the pack that establishes it -- plus the three free floors' answers beside every one.

And when it cannot

⚠︎ THE FREE FLOOR WINS THIS ONE, AND IT IS THE FIRST THING TO SAY. b002-reply-reconcile-notes -- a regex over the printed comparisons plus a keyword match over the desk note, costing $0.00 -- scores 240 of 240 reply lines with every graded field right. r001-reply-reconcile scores 239 of 240 for $0.271736. ⚠︎ AND THAT FLOOR'S NUMBER IS A CEILING, NOT A RESULT: tools/build_corpus.py writes every deciding desk note from five sentences per direction and every finding from a small template pool, and the floor's patterns were written by the same author on the same day against that pool. ⚠︎ THE PAID CALL'S ONLY MISS IN 240 JUDGEMENTS IS A CITATION, NOT AN ANSWER. On RR-0019 line 2 it got the cluster, the queue and the date right and quoted the package finding as its evidence, on a line where two comparisons are both true and only the desk note settles which one CMS rejected on: coverage 0 pct, precision 0 pct, no credit. An analyst would file the right correction and be shown the wrong reason for it. ⚠︎ AND THE SECOND NAMED TRAP CAUGHT NOTHING. 51 lines carry a requested effective date below a floor, and date_would_reject is 0 of 51 on every arm including the lookup floor. The trap is planted, it is real, and it is published as unsprung rather than as a win.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • your reply codes each name one root cause and your desk keys them by hand — the reply-code lookup floor (b000), and no model at all
    it takes 110 of 240 lines here for $0.00 and every one of them is a table read. Paying a model for a lookup is paying for a lookup.
  • your reconciliation tool already prints field-by-field comparisons — the printed-comparison floor (b001)
    204 of 240 lines, $0.00, and it takes the plan-code trap 40 of 40 because RR-9 is a written rule a floor author would read.
  • the deciding evidence on your hard lines is FREE TEXT somebody typed — the model call -- but measure the channel first
    36 of these 240 lines are settled only by prose, and the fast tier takes all 36. That is the only band on this corpus where a call is doing something a lookup cannot.
  • you want the drafted effective date to be right — pure code -- src/policy.py::corrective_date()
    RR-5..RR-8 are max() over four printed dates. Every arm here scored 0 of 51 on date_would_rereject, and src/recheck.py re-derives the date from the model's own reading so a fluent wrong date cannot survive.

At a glanceHow the whole thing runs

99.6%line all correct pct
40,720 msp50, end to end
$20.89per 1,000 reply-file packs · google/gemini-3-flash

Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/packs.json with your own reply-file rows -- one object per pack, four lines each, carrying the reply code, the beneficiary segment and the transaction as submitted -- and rewrite tools/build_corpus.py's render() to print them in your file's own layout. ⚠︎ WHAT STOPS BEING TRUE ON YOUR CORPUS IS EVERY NUMBER ON THIS PAGE. Corpus lens →
When is this the wrong choice?Avoid: A paid call on the settled half -- RR-2 discards the reading on those lines anyway. That is the case against the best-fitting scenario (“your reply codes each name one root cause and your desk keys them by hand”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack with more or fewer than four reply lines. The scorer grades a fixed list in a fixed order; a fifth line has no gold row and a missing fourth is counted as unanswered, not absent. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the fast tier's 239 of 240 would hold on a second run. It ran once. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is an OpenAI-compatible endpoint configured in the repo-root .env; only src/adapters/__init__.py knows which; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-01 — r001-reply-reconcile. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9236, rebuilds the corpus to the byte, re-derives the answer key independently, and scores all three free floors offline -- 60 packs, 240 judgements, $0.00. The committed scored run replays from its result file with every quote highlighted at its offsets. Only the "Check with the model" control needs a key, and with none it is disabled and says so.

A living map of modern AI — kept current every morning