Home › Use Cases › Telehealth benefit and cost-share extraction
Use caseUC0392
🧪 Use-case kit · runnable

Telehealth benefit and cost-share extraction

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A telehealth visit is booked and nobody knows what the member will be asked to pay. The answer is in the member's own benefit document, and that document states it SEPARATELY for a video visit, an audio-only visit and an asynchronous e-visit — sometimes in a benefit grid, sometimes in a paragraph, sometimes twice with two different numbers because a rider restated it, and very often not at all for the modality actually booked. The scheduler reads the first telehealth line they find. When that line was written for a different modality, the member is quoted the wrong cost-share, and the correction arrives after the visit. the read a scheduler does through a member's benefit booklet before a telehealth appointment, on the visits there is time for

Audience

The scheduling or patient-access team filling in a visit record before the appointment, and the operations lead deciding whether a model call per visit is worth buying over the extraction they could write in code. On this corpus the honest answer is no, and the report says so in its first sentence. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual member benefit and cost-share packets

The corpus is 62 member benefit and cost-share packets, 0.16 MB (json 3 · jsonl 1 · md 1 · txt 62). ⚠︎ EVERY BYTE OF IT IS INVENTED, AND THAT IS STATED BEFORE ANY NUMBER. TBP-2026 is a fictional plan benefit document written for this kit — not a real insurer's booklet, not a summary-of-benefits template, and not a reading of any published rule. NO REGULATION, STATUTE OR DISCLOSURE RULE IS NAMED AS GOVERNING ANYWHERE IN THIS KIT: the use-case row it was built from carries BLOCKED-PENDING-ANCHOR, so the rulebook is the member's OWN plan document and nothing else. It was generated rather than found because the measurement needs an answer key at the term level — and, more importantly, because THE CORPUS IS THE EXPERIMENT. 31 packets state the benefit as a structured GRID with the category, modality and term key printed, and 31 state the same benefits as BOOKLET prose, with the same verdict rates from the same generator. The booklet half is shaped against the shortcut: the modality named four different ways (a 'telephone visit' is an audio-only visit), 64 sentences that name it only by back-reference ('a visit of this kind'), a coinsurance stated as what THE PLAN pays, values spelled out, 79 rider statements that restate a term with another number and do not say they supersede, 105 statements about a category the visit is not in, and five decoy numbers per packet.

The corpus

  • The 62 member benefit and cost-share packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your member benefit and cost-share packets. That is the whole change — there is no database to migrate.

One member benefit and cost-share packet, as the model receives itBEN-0001.txt · 1 of 62
MEMBER BENEFIT AND COST-SHARE PACKET — BEN-0001
Plan: Thornbury Benefit Plan (fictional) · benefit document TBP-2026 (fictional)
Member: Alder R. Finnegan · member id 853486283 · group 43484
Scheduled visit SCH-334539 · benefit document format GRID

PART A — THE MEMBER'S BENEFIT DOCUMENT
  Benefit document TBP-2026, plan year 2026, issued to member Alder R. Finnegan (member id 853486283, group 43484).
  The plan-year deductible is USD 2,500.00 for an individual and USD 5,000.00 for a family. The out-of-pocket maximum is USD 10,000.00.

  BENEFIT GRID — telehealth benefits, plan year 2026
  CATEGORY             MODALITY     TERM                 VALUE
  PRIMARY_CARE         AUDIO_ONLY   coinsurance_pct      10% member share
  PRIMARY_CARE         AUDIO_ONLY   originating_site     SITE_REQUIRED
  PRIMARY_CARE         AUDIO_ONLY   vendor_required      ANY_NETWORK_PROVIDER
  PRIMARY_CARE         E_VISIT      deductible_applies   NO
  PRIMARY_CARE         E_VISIT      visit_limit          UNLIMITED
  SPECIALIST           E_VISIT      coinsurance_pct      0% member share

  AMENDMENTS AND RIDERS effective 2026-01-01
  - PRIMARY_CARE E_VISIT deductible_applies is YES

PART B — THE SCHEDULED VISIT (scheduling system record)
  SCHEDULED VISIT        SCH-334539
  VISIT DATE             2026-03-28
  MODALITY               E_VISIT
  BENEFIT CATEGORY       PRIMARY_CARE
  NETWORK STATUS         IN_NETWORK
  ALLOWED AMOUNT         USD 240.00
  DEDUCTIBLE MET         USD 400.00 of USD 2,500.00
  MEMBER SERVICE NOTES: (none recorded)

PART C — TERMS TO EXTRACT FOR THIS VISIT (every term the estimate needs)
  TERM                 WHAT IT IS                                       THE FORM TO RETURN IT IN

Abridged — the file continues.

The outcomeWhat a good result looks like

Seven benefit terms per scheduled visit, each with one of four verdicts and the plan's own words quoted beside it, and then a pure-code expected cost-share — or a refusal that names exactly which term the document does not state.

And when it cannot

It UNDER-claims rather than over-claims, and that direction is the finding. 44 of its 82 errors say the document does not state a term for this visit when the document states it for a DIFFERENT modality of the same category — a distinction the estimator does not act on either way. It quoted the wrong modality's cost-share 0 times in 70 opportunities, where a keyword scan does it 49 times and a constant 70. The errors that cost something are smaller and specific: 6 contradictions resolved to one side, 14 terms the document states reported as contradicted, and one packet where a conservative reading of a single term removed a contradiction the key holds and let the estimator produce a figure the key refuses.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • your benefit terms live in a structured GRID — a table, an export, a benefits configuration record — the keyed lookup in evals/baseline.py, on its own
    it scores 217 of 217 on that half of this corpus against the paid call's 186, for $0.00 and no network. A lookup by category, modality and term is what code is for
  • you only need to know whether a term is USABLE — can the estimator compute with it, or must this go to the benefits desk? — either arm on grid documents; take the free one
    on the grid half the usable question is 212 against 217, p = 0.062500 — this corpus cannot separate them
  • your booklets use several words for the same modality and you can list them — the prose reader in evals/baseline.py, with your vocabulary in data/plan.json
    the single commonest failure of the paid call on this corpus is not recognising that a 'telephone visit' is an audio-only visit. A list of your own wordings is an afternoon's work and it is free
  • your booklets are prose and your term set is large or changes often — measure the call before deciding, on your own labelled packets
    an anchor lexicon needs one line per term and someone to maintain it; a call needs neither. This kit did not measure that trade — its catalogue is seven fixed terms

At a glanceHow the whole thing runs

81%term verdict correct pct
1,863 msp50, end to end
$0.61per 1,000 member benefit and cost-share packets · the fast tier

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/plan.json with your own plan catalogue — your benefit terms and the form each value must come back in, your modalities and the words your documents use for them, your benefit categories, and the phrasings your booklets actually use — and drop your packets into data/corpus/ in the three-part shape tools/build_corpus.py emits. THE MEASURED RESULT DOES NOT TRAVEL WITH YOUR CORPUS, AND ON THIS KIT IT CUTS BOTH WAYS. Corpus lens →
When is this the wrong choice?Avoid: Buying a call per visit for this. On this corpus it is measurably worse. That is the case against the best-fitting scenario (“your benefit terms live in a structured GRID — a table, an export, a benefits configuration record”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a benefit document that scopes a whole SECTION to one modality with a heading rather than qualifying each sentence. Every statement here carries its own qualifier or a back-reference to the sentence before it; a heading that governs a page is a structure this corpus does not have. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?whether the paid call would beat the floor on a booklet that is NOT built from a closed pool of 65 templates. This is the biggest open question on the page and the corpus cannot answer it. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-benefit-costshare. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the answer key, the committed run and all four free floors — and scores every floor offline in pure Python. Shot as kits/UC0392-benefit-costshare/docs/shots/benefit-costshare-empty.png against a server started with API_KEY blanked in its own environment, not against a banner drawn to look like one.

A living map of modern AI — kept current every morning