The business caseThe problem this solves
Payer remittances arrive as faxes and scans. A poster keys them in, and a number read onto the wrong claim line moves money silently — the check total still agrees, so nothing downstream objects. Manual keying and a second read of every page
Audience
Revenue-cycle posting teams and the supervisor who signs off exceptions Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual remittance advices
The corpus is 30 remittance advices, 13.42 MB (json 1 · md 1 · png 63). Synthetic, so ground truth is what was PRINTED before it was printed. A real remittance would need every cell hand-transcribed to be scorable, and a hand transcription is an opinion that then gets graded as fact.
The corpus
- The 30 remittance advicesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your remittance advices. That is the whole change — there is no database to migrate.
The corpus is 30 remittance advices, and it is not text — there is no clip to show you here. The pipeline reads these files directly; the folder swap above is still the whole change.
The outcomeWhat a good result looks like
A posting record plus the lines a human must decide, each naming the rule that fired and the two amounts it compared.
And when it cannot
A line posted from a misread page. This kit never posts, so the failure it can still have is an exception queue nobody trusts.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Paper and fax remittances, one payer's layout — the paid batch tier with the model reader
the measured path scores 0.9355 exception F1 - Handwritten annotations on the remittance — measure it before trusting it
handwriting is not in the corpus
And where nothing here is good enough:
- 835 EDI files — none — parse the file
there is no page to read
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus with your own page images and edit VALID_CARC in src/rules.py to your payer's adjustment set. The rules assume allowed = paid + adjustments and that the claim lines sum to the check total. Corpus lens → |
| When is this the wrong choice? | Avoid: Measured on ONE synthetic layout. A payer whose table differs will move these numbers and this kit cannot tell you by how much. That is the case against the best-fitting scenario (“Paper and fax remittances, one payer's layout”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | Measured, not guessed: a claim id the engine misreads takes the whole row with it. The free engine lost 352 of 1102 rows that way and the paid one lost 5. 3 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Only ONE reading VENDOR is priced here. Textract and Rekognition return AccessDenied on this account, Google's credentials are empty and Azure carries a Speech key only, so the two priced tracks are one product at two service levels. 4 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is openai-compatible; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 3 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-09 — r001-remit-posting-model-mistral. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — From a clean checkout we ran: python3 -m venv .venv, pip install -r requirements.txt (Pillow, plus pyobjc for the free macOS arm), then tools/build_corpus.py --check, which re-derived gold byte-identical. Scoring needs no install at all — evals/ and src/ are standard library only.

