The business caseThe problem this solves
A payer returns a remittance: one row per claim line, each with an amount allowed, a contractual adjustment, an amount paid, an amount put to the member and sometimes a denial code. The contract the practice signed says what each of those lines was supposed to allow, and the member's plan says what share of it the member owes. Posting the remittance closes the claim; nobody re-derives the contract's own answer, so a line denied for a reason the contract does not support, a line the contract carves out and the payer allowed anyway, or a rate amendment the payer never loaded all settle silently. Opening one claim's remittance, looking the contracted rate up in the fee schedule, multiplying it out, taking the copay then the deductible then coinsurance in the plan's own order, and reading every remark under every denied line to decide whether the denial holds. Nothing is replaced end to end: the pack produces a row a person reads, and it appeals nothing, rebills nothing and posts nothing.
Audience
A provider's revenue-cycle desk working a remittance worklist, and the person who has to decide which claims are worth a payer conversation this week. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual remittance file
The corpus is 63 remittance file, 0.18 MB (txt 63). It is generated because it has to be. A real remittance is a payer's own adjudication of a named member's claim under a signed contract, and none of that can be published. Generating it also buys the one thing a scraped set cannot: the answer key is DERIVED from the structure the file was rendered from, so there is no second place the answer lives and a corpus change cannot leave a stale key behind.
The corpus
- The 63 remittance filegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement, including the one construction defect that makes a graded field free.
Swap this folder for your own material and the kit is pointed at your remittance file. That is the whole change — there is no database to migrate.
======================================================================================================================
REMITTANCE RECONCILIATION FILE REM-0001
Payer: PAY-2100 - Cascade Ridge Health Plan (invented)
Provider: GRP-4400 - Drumlin Park Virtual Clinic (invented)
Claim: CLM-80000 Contract: CNT-1100 Service: 2026-08-03 Procedure: RRP-2026
======================================================================================================================
CLAIM AND PLAN AS THE CONTRACT HOLDS IT
place of service 10 telehealth, patient at home
member copay 15.00
deductible remaining 400.00
coinsurance pct 0.00 pct of what remains
out-of-pocket remaining 4000.00
tolerance pct 1.00 pct of contract-expected allowed
threshold cost 50.00 or more
threshold pct 3.00 pct or more
remittance original
CONTRACT FEE SCHEDULE AS PRINTED
SCHED CODE NAME MOD UNITS BILLED RATE CODE RATE PER UNIT
FS-0100 SVC-4014 virtual visit, established, standard V1 1 177.00 RC-2200 136.40
FS-0101 SVC-4021 virtual visit, new patient, extended -- 1 308.00 RC-2201 218.80
REMITTANCE AS THE PAYER RETURNED IT
LINE DATE UNITS BILLED ALLOWED TYPE STATUS CODE BASIS
RL-0101 2026-08-03 1 177.00 134.62 service ADJUDICATED --- payer allowance 134.62 per unit x 1Abridged — the file continues.
The outcomeWhat a good result looks like
One claim in, one row out: which remittance lines this procedure treats differently from the remittance that printed them, whether a contracted-rate amendment was in force, and — re-derived in code from those two — the contract-expected allowed amount, the member's share under the plan's own terms, the variance in money and as a percentage, and one verdict.
And when it cannot
And what it does when it cannot. On the scored run 63 of 63 replies parsed and none stopped at the token ceiling, so there is no unanswered-claim behaviour to report from measurement. A reply the parser cannot read is counted WRONG and stays in the denominator; a reply with no supported amount reaches src/policy.py as None and the station returns a null verdict rather than folding it into AS-CONTRACTED.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your payer's remarks say plainly why a denial does not hold — 'authorization on file', 'member eligible', 'duplicate reversed' — in words a keyword list would catch. — the free rules floor, and do not buy a call at all
On the twelve remark-decided claims where every wording carries a keyword the paid arm gets 5; on the eight where none does it also gets 5. The keyword made no difference to it, and a rule is free. - You have claims whose totals net out — a line wrongly denied and a line wrongly allowed of the same size — and nobody ever finds them because the claim total is right. — this kit, for
citesalone
4 of 4 against 0 for the payer's own totals, 0 for the majority reading and 1 for free code. No amount and no verdict can carry that family;citesis the only field that can. - Your allowed columns are sometimes a per-unit rate rather than a line total. — free code
It is arithmetic on the row's own basis text. The free floor gets 2 of 5 and the paid arm gets 0 of 5. - You need a written record that nothing in the pipeline appeals, rebills, writes off or cites a law. — this kit
The answer contract offers no field that could express any of the four — asserted against data/fields.json at import — and src/refusal.py read the sentence on every arm of every run: 0 on all 63 claims and 0 on all 19 attacked ones, including four handed a clause demanding an appeal and naming a statute.
And where nothing here is good enough:
- Your remittances are mostly correct and your desk's cost is chasing false positives. — nothing here yet
On the twelve claims where the payer adjudicated every row correctly, doing nothing scores 12 and this kit scores 4. It manufactures work on exactly the claims you would not have worked.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own remittance files in the same shape and data/contracts.json with your own register, then run python3 -m evals.run --run-id b000-<slug>-rules --floor rules — no key, no spend — to see what free code gets on your files before buying anything. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per claim for a regex you could write in an afternoon. That is the case against the best-fitting scenario (“Your payer's remarks say plainly why a denial does not hold — 'authorization on file', 'member eligible', 'duplicate reversed' — in words a keyword list would catch.”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A remittance that is not fixed-width columns. src/statement.py's row regex is anchored to the printed layout; a CSV or an 835 segment stream parses to nothing and every arm returns nothing rather than something wrong. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval on the headline is claimed. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-09 — r001-expected-remit. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors, the stub and every committed run. tools/build_corpus.py --check and python3 -m evals.check_labels both pass with nothing installed.







