Home › Use Cases › Resolve pended and suspended claims by cause class, and keep the release human
Use caseUC0237
🧪 Use-case kit · runnable

Resolve pended and suspended claims by cause class, and keep the release human

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A pended claim is a claim the adjudication system could not finish, and the queue it lands in is where money and deadlines go quiet at the same time. The pend message names the first edit that fired, which is frequently not what is actually wrong; what the payer has SINCE located sits in a different system from the narrative; and the regulatory payment clock has been running since the claim was RECEIVED, not since it pended. A desk works one claim at a time and re-derives all of that per claim. Nothing downstream reports the two failures that matter: a claim that quietly ages past its deadline surfaces later as an interest payment or an audit finding, and a claim released when it should have been held surfaces as money that has to be recovered. Re-deriving four answers per pended claim by hand -- reading the narrative past the pend message, checking a second system for what has been located, and doing the receipt-date arithmetic. It replaces the derivation, never the decision: the resolution is proposed and an examiner applies it.

Audience

The claims examiner or pend-desk lead who has one suspended claim in front of them and a queue to get through -- and the claims leadership deciding, separately and much later, whether any pend category has enough evidence behind it to be released without a person. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual pended claims

The corpus is 60 pended claims, 0.08 MB (txt 60). Because the alternative does not exist. There is no public corpus of pended claims and there never will be one: a pend record is a record of a person's care, a provider's billing and a payer's own adjudication, and the desk notes this kit measures ARE the sensitive part. And even if one existed it would not carry what is being measured as ground truth -- which of eight causes actually pended the claim as against what the system's own message said, what the payer has since located, and exactly which sentence establishes the cause to the character. A human labelling pass over real claims would be one examiner's reading rather than a known fact. So the facts are injected, the key is derived from them by the same engines the model is rechecked with, and the records are written around them.

The corpus

  • The 60 pended claimsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.

Swap this folder for your own material and the kit is pointed at your pended claims. That is the whole change — there is no database to migrate.

One pended claim, as the model receives itPC-0001.txt · 1 of 60
PEND RECORD PC-0001 -- Inpatient claim, held pending development

CLAIM FACTS
  Claim id                      CLM-2026-400037
  Line of business              Medicare Advantage
  Billing provider              contracted
  Claim type                    clean claim
  Received on                   2026-08-10
  Pended on                     2026-08-12
  Register read as at           2026-09-01
  Queue                         general pend queue at Westmoor Community Plan

PEND MESSAGE -- what the adjudication system printed

  EDIT 302 -- precertification required for this place of service and none located; pended.

DESK NOTES

The desk note carries the analyst's initials and the queue it was worked from. The claim arrived in the weekend batch and was worked in the order it was received.

The desk opened the record on the day it suspended and has added to it twice since.

CORRESPONDENCE AND SYSTEM TRAIL

  2026-08-10  claim received and logged into adjudication
  2026-08-12  claim suspended by the adjudication cycle
  2026-08-24  queue reassignment run; the claim stayed on the same desk
  2026-08-31  record opened on the desk worklist

ADJUDICATION REGISTER -- what the payer has actually located

  a matching authorisation was located on the utilisation file, for this code, these units and this provider.

The outcomeWhat a good result looks like

Four answers a person confirms instead of a claim they re-derive: the cause class with the sentence it came from copied verbatim and located at its offsets, the resolution PR-2026 produces from that cause and the adjudication register, where the regulatory clock stands with the rule and the arithmetic beside it, and who applies the result.

And when it cannot

Four directions and they are not comparable. A FALSE RELEASE is a claim let go without a person: it is counted, never averaged, and it is the number the whole kit is built around -- 0 of 60 on both arms, 0 of the 6 bait records. A MISSED ESCALATION is a claim left sitting while its deadline passes and nothing downstream reports it -- 0 of 12 on both arms. RELEASED WHEN IT SHOULD NOT is money leaving on a claim the register contradicts -- 0 on both arms. DENIED WHEN IT SHOULD HAVE MOVED makes a provider appeal a claim they billed properly -- 2 on the paid call's raw answers and 0 after the recheck, which is the clearest thing the pure-code station bought on this run.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A pend queue where the adjudication system's own message usually names the real cause — the free rules floor alone
    The floor and the paid call tie at 44 of 60 records all-correct, and the floor beats the call on the citation 49 to 45. Every register-driven resolution is free: PR-1 6/6, PR-3 12/12 and PR-4 8/8 on the floor, because they come from data/claims.json and no reading is involved. The clock is 60/60 and the release is 60/60 on the floor, at $0.00 and 0.20 seconds for the whole corpus.
  • A pend queue where the printed edit is the first thing that fired rather than what is wrong — the paid call, and keep PR-2026 behind it
    This is the only slice where the reading pays, and it pays a lot: 6 of 6 against 0 of 6 on cause_looks_like_another, and 19 of 21 against 10 of 21 on the resolutions the CAUSE CLASS decides. Keep the recheck: it moved 6 resolutions on this run and every one the right way.
  • You want the clock and the escalation list and nothing else — src/clock.py, with no model and no floor
    It is two dates, three claim facts and an ordered table. Both arms score 60 of 60 and 0 missed escalations because neither is doing anything -- the arithmetic is. It costs nothing and it is the half of this job with a regulatory deadline attached to it.
  • You are deciding whether a pend category could ever be auto-released — this kit as an evidence pipeline, and NOT as an answer
    The false-release count is what the kit publishes and it is 0 of 60 on both arms with 0 of 6 on the bait records. That is an upper bound at n=60, over one cycle, with an adversarial arm that has never been fired. The release gate names four conditions and this run satisfies none of them.

At a glanceHow the whole thing runs

93%cause class accuracy pct
27,147 msp50, end to end
$3.60per 1,000 pended claims · the fast tier, off-peak

Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own pend records as .txt into data/corpus/, add one row per record to data/claims.json (id, claim_id, title, line_of_business as MA or MEDICAID, clean, contracted, received_on, pended_on, register_read_on, and an evidence value from clean / correction / contradicted / awaited / none), and write data/gold.jsonl with the cause class, the resolution, the rule, the clock state, the clock rule, the release, the citation and its span. ⚠︎ THE MEASURED NUMBERS DO NOT TRAVEL, AND ONE CHOICE MOVES THEM MORE THAN ANY OTHER. Corpus lens →
When is this the wrong choice?Avoid: The paid call -- on this slice it buys nothing and costs $0.1676 for 60 claims. That is the case against the best-fitting scenario (“A pend queue where the adjudication system's own message usually names the real cause”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A record whose sections are not the five this kit's generator writes. evals/baseline.py splits on the headings and simply gets an empty section for a missing one -- it does not raise -- but a record with no PEND MESSAGE section hands the floor nothing to read first, which is the whole of how it works. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE CITATION KEY IS THE RIGHT KEY. 13 of the paid call's 15 citation misses quote the ADJUDICATION REGISTER line rather than the pend message. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is an OpenAI-compatible endpoint configured in .env; src/adapters/ is the only file in the kit that knows which; the Prompt lens states what swapping it costs. The published figures come from 2 models on the fast tier and the free rules floor, pure Python. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-01 — r001-pend-resolve. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured rebuilds the corpus, re-derives the answer key and scores the free floor in 0.20 seconds of wall clock, measured on this machine, with no network and nothing to install -- there is no requirements.txt to satisfy. The board renders in full on 127.0.0.1:9237 with the model button disabled and the reason printed beside it, and the committed scored run replays off its result file. What a cold clone CANNOT do is fire the paid arm or the injection arm; both need a key, and the injection arm has never been fired by anybody.

A living map of modern AI — kept current every morning