Home › Use Cases › Draft a patient cost estimate from the plan terms, or name the input that is missing
Use caseUC0475
🧪 Use-case kit · runnable

Draft a patient cost estimate from the plan terms, or name the input that is missing

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A benefits desk is asked what one scheduled service will cost a member. The plan's printed terms are on the packet, the amounts on file for that member and that service are on the packet, and the ladder that turns the two into a number is five integer steps. What is NOT on any column is whether a value that IS printed may still be estimated from: a fee schedule reloaded, an accumulator rebuilt, a benefit category unpublished — each suspending one input, each with an effective date, each cancellable by a later note, and all of it written in prose in the eligibility desk notes. Today a benefits analyst reads those notes by eye, then runs the ladder by hand, and a second person re-reads the ones that produce a number a member will see. Reading one packet's eligibility desk notes against the plan's own estimate terms to decide whether each required input may still be estimated from, then running the five-step ladder by hand — not the decision to send an estimate to a member, and not the person who owns it.

Audience

A benefits analyst drafting one member's cost estimate, and the supervisor who decides whether the number goes out. The answer this report gives them is a no with a measurement attached: the paid call beats the free floor of record by one packet in sixty-four and the paired test cannot separate the two, while free code that reimplements the plan terms beats the paid call outright. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual patient estimate packets

The corpus is 64 patient estimate packets, 0.20 MB (json 5 · jsonl 1 · md 2 · txt 64). Because the shape of a cost-estimate failure is not the arithmetic, it is whether a printed value may still be estimated from, and that lives in prose. EVERY packet carries a suspension that does NOT bite — cancelled by a later note, effective after the service date, or naming the prior-year record whose id is four characters away — so a reader that spots the word and stops is worse than a reader that never looks. That is measured rather than asserted: the domain word list scores 10.94 pct of the full row and the generator-tuned regex 18.75 pct, against the printed columns' 82.81 pct. A corpus of clean ladders would measure a calculator; this one measures a reader.

The corpus

  • The 64 patient estimate packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 64 packets, data/request_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.

Swap this folder for your own material and the kit is pointed at your patient estimate packets. That is the whole change — there is no database to migrate.

One patient estimate packet, as the model receives itEST-0001.txt · 1 of 64
HARROWGATE BENEFIT ADMINISTRATORS - PATIENT COST ESTIMATE PACKET
SYNTHETIC SAMPLE DATA. GENERATED, NOT REAL. NO REAL PLAN, MEMBER OR PROVIDER.

Packet EST-0001   Request REQ-40024   Prepared 2026-03-06

MEMBER AND PLAN
  Member reference                              MBR-75733
  Plan                                          Harrowgate Care Select (HCS-2026)
  Plan terms document                           HCS-TERMS-2026
  Network status of the rendering provider      in-network
  Current plan-year accumulator                 ACC-2026-4543
  Prior plan-year accumulator                   ACC-2025-4699
  Current out-of-pocket accumulator             OPA-2026-5195
  Prior out-of-pocket accumulator               OPA-2025-5558
  Current coinsurance schedule                  CSCH-2026-51
  Prior coinsurance schedule                    CSCH-2025-28
  Current fee schedule                          FS-2026-13
  Prior fee schedule                            FS-2025-33

SCHEDULED SERVICE
  Service code                                  SVC-3528
  Service                                       Virtual nutrition counselling, 45 minutes
  Also on this order, not estimated here        SVC-3514
  Rendering provider                            PRV-3675
  Referring provider                            PRV-3742
  Prior authorisation reference                 AUTH-12670
  Service date                                  2026-03-25
  Estimate requested                            2026-03-04
  Benefit category                              coinsurance-based

AMOUNTS ON FILE
  Provider billed charge                        $544
  Plan allowed amount, in-network               $464
  Plan allowed amount, out-of-network           $424

PLAN TERMS IN FORCE (from HCS-TERMS-2026)

Abridged — the file continues.

The outcomeWhat a good result looks like

Either a drafted estimate carrying seven figures — allowed amount, deductible applied, coinsurance applied, copay applied, out-of-pocket relief, patient pays, plan pays — each naming the term of the plan document that produced it, plus a route and a timing verdict; or exactly ONE named missing input and no figures at all. A packet whose required input is absent or suspended for this service date is declined and never estimated.

And when it cannot

⚠︎ THE HEADLINE IS A NON-RESULT AND IT IS PUBLISHED AS ONE. The measure is the FULL ROW — the decision AND all seven figures AND the timing, on one packet. Rechecked, the paid call takes 54 of 64. The free floor of record, pure code over the printed columns with the desk notes not read at all, takes 53. That is +1 packet and the paired McNemar exact test on the same 64 items is p = 1.000000 — NOT SIGNIFICANT, 9 discordant pairs split 5 and 4. Worse: full, the plan terms reimplemented in free code, takes 62 of 64 and BEATS the paid call at p = 0.007812. And the arm's own end-to-end draft, before the pure-code station, is 39 of 64 — BELOW the free floor. The pure-code station is worth +15 packets; the model is worth +1.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A packet whose required inputs are all printed and none suspended — the free printed-columns floor
    it takes 53 of the 64 full rows at $0.00 and no network, and the paired test cannot separate the paid call from it (p = 1.000000). On a packet with nothing to read, there is nothing to buy.
  • A packet whose answer lives only in the eligibility desk notes — the paid call, with the recheck station — or free code that reimplements the terms
    the paid call takes 5 of these 11 and the printed-columns floor takes 0, which is the only place on this page where the money reaches something free code over the columns cannot. ⚠︎ But full, which reimplements the plan terms in pure code, takes 9 of the same 11 — so the honest pick is whichever you can write.
  • You can afford to write the plan terms out in code once — free code, and do not buy a call at all
    full takes 62 of 64 against the paid call's 54 and beats it at p = 0.007812, for $0.00 and no network. Its two misses are a word-list gap rather than a logic gap.

And where nothing here is good enough:

  • Anything a member will read — neither; a person
    5 packets got a confident seven-figure estimate the plan terms forbid drafting at all, on BOTH columns, and the station cannot see them because it trusts the reading. A wrong number in front of a member is the harm the missing-input rule is written against.

At a glanceHow the whole thing runs

84%full row correct pct
1,656 msp50, end to end
$0.79per 1,000 patient estimate packets · GPT-5.6 Luna

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own rendered estimate packets and data/request_records.json at your own structured half, then rewrite src/packet.py for your layout and data/plan-terms.md + data/plan-terms.json for your benefit design. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call on this population at all. The 53 packets the columns settle are where both arms agree. That is the case against the best-fitting scenario (“A packet whose required inputs are all printed and none suspended”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet whose eligibility desk notes are not in the printed layout src/packet.py parses. The parser is positional over a fixed rendering; another administrator's form is a new parser, not a new prompt. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether any of this holds on a real administrator's packets. Every figure here is against a key derived from a generator, on an invented plan document, for an administrator that does not exist. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-cost-estimate. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on a clean checkout of this repository with no key configured and no network: python3 -m evals.baseline scores all five free floors over all 64 packets and writes the two committed b000 result files, python3 -m evals.check_labels re-derives the whole answer key a second time and reports 0 disagreements, and python3 -m src.app serves the whole board — every panel, the committed run replayed, the probe replayed and all five floors computed live — with the one endpoint that could spend anything disabled and the reason printed beside it. What a clean checkout CANNOT do is buy a call: that needs one credential, and nothing in the free path opens a socket.

A living map of modern AI — kept current every morning