Home › Use Cases › Medical bill and records extraction
Use caseUC0505
🧪 Use-case kit · runnable

Medical bill and records extraction

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A demand package arrives on a bodily-injury liability claim: a provider's itemised statement of 11 to 19 charge lines, and the treatment records and correspondence covering the same dates. Somebody has to say, line by line, which charges the records actually evidence, which are the same service billed twice, which were withdrawn and rebilled, and which visit each charge belongs to - and to cite the printed line for every one. Today an adjuster and a nurse reviewer read both documents side by side and build that list by hand. The first pass over a provider bill package: matching each printed charge line to the note that evidences it, deciding whether a repeat is a duplicate or a legitimate second service, and attaching a range-dated or undated charge to a visit.

Audience

Anyone putting an automated reader over a claim file where the expensive mistake is not a wrong total but a charge marked unevidenced that the records plainly evidence. The answer this page gives about its own pack is a flat no: free code does the job better, significantly, on every one of the nine planted line families. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual bill packages

The corpus is 64 bill packages, 0.45 MB (md 64). The smallest corpus that makes the interesting mistake unavoidable. A bill package is the one document pair where the arithmetic is genuinely free - the capture tuple, the statement sum, exact-tuple matching, visit ordering, interval and gap arithmetic are all exact for every arm including a constant - and the judgement genuinely is not: whether two differently worded charges are the same service billed twice, and whether a charge the records never name is unevidenced or just described in other words. So the split between what money buys and what free code does could be NAMED before a call was made and SCORED after. It came out inverted, which is the finding.

The corpus

  • The 64 bill packagesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your bill packages. That is the whole change — there is no database to migrate.

One bill package, as the model receives itpackages/BI-2681-0001.md · 1 of 64
# Medical bill package -- BI-2681-0001

Synthetic package. Every reference, name, date, code and note below was generated by tools/build_corpus.py from a fixed seed. Nothing here is a real claim, a real person, a real provider or a real bill, and the service codes are an invented internal schedule that corresponds to no published code set.

Claim number: BI-2681-0001
Claimant reference: CLMT-0001
Date of loss: 2026-01-27
Adjuster: R. Halvorsen, liability claims
Nurse reviewer: H. Rasmussen, medical review
Demand package: DEM-0001-03, received 2026-06-26
Provider on this statement: imaging centre, PRV-1041
Records window read on this package: 2026-01-27 to 2026-06-26
Gap threshold (operator setting on this claim): 14 days

## What this pack does and does not do

The pack structures and cites what this statement and these records already say and nothing beyond it. It never evaluates medical necessity and never approves, reduces or pays a charge. The window and the gap threshold above are settings on this claim, supplied by the operator; nothing on this package states a rule, threshold or deadline of any other kind.

## Itemised statement -- PRV-1041

| ref | date of service | code | mod | units | amount | description |
|---|---|---|---|---|---|---|
| p2 l01 | 2026-04-16 | SCH-2481 | -- | 1 | 961.20 | advanced imaging, single region |
| p2 l02 | 2026-03-12 | SCH-2410 | -- | 1 | 970.10 | cross-sectional imaging study |
| p2 l03 | 2026-01-28 | SCH-2318 | V2 | 1 | 240.00 | single-region imaging series |
| p2 l04 | 2026-02-12 | SCH-2372 | -- | 1 | 261.60 | standard imaging, one region |
| p2 l05 | 2026-04-16 | SCH-2481 | -- | 1 | 907.80 | advanced imaging, single region |
| p2 l06 | 2026-01-28 | SCH-2301 | -- | 1 | 240.00 | plain imaging study, single region |

Abridged — the file continues.

The outcomeWhat a good result looks like

One {disposition, partner, visit} row per printed charge line, keyed on the line's own p<page> l<line> reference: a disposition from the closed set billable / duplicate / superseded / unevidenced, the line it duplicates or supersedes (or none), and the visit date the records place it on (or unattached). A package is right only when every one of its lines is right on all three.

And when it cannot

It over-flags. Across 878 charge lines it answered unevidenced 327 times where the key says 35, and billable 334 times where the key says 702. Its largest single loss is on the EASIEST lines - 347 of the 622 plain lines, each carrying an exact printed date and exactly one admitted note evidencing it - because it cannot hold the occurrence ledger in printed order across 11 to 19 lines. On the 24 undated charges it is wrong 24 times.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Deciding which printed charges a set of records evidences — free code - the domain arm in evals/floors.py
    It reaches 711 of 878 charge lines and 16 of 64 whole packages for $0.00, against the paid call's 380 and 0, and the loss is significant at p = 1.93433e-57.
  • Flagging the same service billed twice under different wording — the domain arm, with exact-tuple matching beside it
    84 of 109 real duplicate pairs against the paid call's 36, and exact-tuple matching alone gets 28 for free. 29 of the 57 identical printed tuples are NOT duplicates, so the tuple rule needs the word list beside it.
  • Attaching a range-dated or undated charge to the visit it belongs to — free code, and then a person
    The bar attaches 45 of 56 and the paid call 11. On the 24 undated charges the paid call is wrong 24 times out of 24 - it takes the first visit it can see rather than the earliest visit with an unclaimed occurrence.
  • Holding a hard cap - no necessity, no approval, no reduction, no payment — this pack's structural contract
    0 breaches over 24 adversarial calls under four framings, including one claiming a claims manager's direction AND a regulator's bill-review rule and demanding both a necessity finding and a day count, and one ordering the unmatched charges denied and the remainder repriced. The cap is structural - the answer has room for {disposition, partner, visit} and nothing else - rather than a prompt sentence, and src/recheck.py drops anything outside it.

At a glanceHow the whole thing runs

0%package all pct
1,988 msp50, end to end
$58.59per 1,000 bill packages · Claude Fable 5

Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/packages at your own bill packages and rewrite src/record.py::parse to return the same three things: the printed charge lines with their p<page> l<line> references, the dated note entries, and the claim header carrying the loss date, the records window and the gap threshold. Nothing measured here transfers to your packages. Corpus lens →
When is this the wrong choice?Avoid: Paying for this reading on material of this shape. It was measured and it is worse, on every one of the nine planted line families. That is the case against the best-fitting scenario (“Deciding which printed charges a set of records evidences”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A statement whose charge lines do not carry a printed page-and-line reference. The whole answer is keyed on p<page> l<line>; without it there is nothing for a disposition to attach to and no way to cite the line a duplicate points at. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether any of this holds on a real bill package. Every charge line and every candidate pair was labelled by the program that wrote it, so the corpus cannot tell you the rate at which two human reviewers would disagree about whether a note evidences a charge - and that disagreement is the whole job. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is OpenAI-compatible endpoint; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-18 — r001-med-bill-extract - 64 bill packages, 878 printed charge lines, 64 billed model calls, the fast tier, reasoning off. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on this build: a clean checkout with no key configured rebuilds all 66 corpus files byte-identically under two PYTHONHASHSEEDs, re-derives all 878 labels with 0 disagreements, scores all four free arms and the whole scoring path at $0.00, and renders the whole board on port 9605 with the one live control disabled. The three reply caches ship, and that was decided by measurement rather than convention: with them removed the board's "the reply, exactly as it arrived" panel falls from S018's whole 1,260-character reply to the record's 400-character head, and --resume --rescore - which otherwise re-derives every published figure for $0.00 - finds nothing cached and would buy 64 calls.

A living map of modern AI — kept current every morning