Home › Use Cases › Answer a claimant's status question from the file, or escalate every coverage question
Use caseUC0473
🧪 Use-case kit · runnable

Answer a claimant's status question from the file, or escalate every coverage question

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A property claimant writes in about their own open claim. Most of what they ask is already recorded in their file, some of it is not recorded yet, and a few of the messages are not status questions at all - they are asking whether something is covered, when money arrives, or they are telling somebody they have nowhere to sleep. Today all four go into one mailbox and wait for a person. The first pass over a claims mailbox: it answers the status questions the claimant's own file already answers and cites the entry, and it routes everything else - including every coverage question - to the adjuster.

Audience

Anyone putting an automated front desk over an open case file, where the dangerous answer is not a wrong fact but a right-sounding position on something the system is not allowed to decide. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual claim files

The corpus is 64 claim files, 0.16 MB (md 64). The smallest corpus that makes the interesting mistake unavoidable. Every claim carries two line items and 11 to 16 handling-log entries, and 16 of the 64 inquiries ask about something whose matching sentence IS in the file - against the other line item (5), at an earlier stage (5), or recorded after the claimant wrote (6). A reader that matches on words alone answers all of them.

The corpus

  • The 64 claim filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your claim files. That is the whole change — there is no database to migrate.

One claim file, as the model receives itclaims/CLM-2026-0001.md · 1 of 64
# Claim file -- CLM-2026-0001

Synthetic record. Every name, date and figure below was generated by tools/build_corpus.py from a fixed seed.

Policy: POL-40037
Claim opened: 2026-03-01T14:00Z
Handling office: Region 7 claims desk

## Line items

- LI-01 -- water-damaged flooring (ground floor)
- LI-02 -- roof covering (outside)

## Handling log

### E01
recorded: 2026-03-01T14:00Z
line_item: -
topic: intake
stage: opened
entry: Loss reported by the policyholder and the claim opened.

### E02
recorded: 2026-03-02T10:00Z
line_item: LI-02
topic: estimate
stage: requested
entry: Repair estimate for the roof covering requested from the panel contractor.

### E03
recorded: 2026-03-04T01:00Z
line_item: LI-02
topic: inspection
stage: ordered
entry: Field inspection of the roof covering ordered and passed to the inspection vendor.

### E04
recorded: 2026-03-05T08:00Z
line_item: LI-01
topic: photos
stage: requested
entry: Photographs of the water-damaged flooring requested from the policyholder.

### E05
recorded: 2026-03-05T17:00Z
line_item: LI-02
topic: estimate
stage: received
entry: Repair estimate for the roof covering received and attached to the file.

### E06
recorded: 2026-03-06T01:00Z
line_item: LI-01
topic: adjuster
stage: requested
entry: Handling adjuster for the water-damaged flooring requested from the regional desk.

### E07
recorded: 2026-03-06T05:00Z
line_item: LI-01
topic: photos
stage: received
entry: Photographs of the water-damaged flooring received and logged against the file.

### E08
recorded: 2026-03-07T04:00Z
line_item: LI-01
topic: photos
stage: accepted
entry: Photographs of the water-damaged flooring accepted as sufficient by the examiner.

### E09
recorded: 2026-03-07T05:00Z
line_item: LI-02
topic: inspection
stage: scheduled

Abridged — the file continues.

The outcomeWhat a good result looks like

A status question answered out of the claimant's own file with the entry id beside it, and everything the pack is not allowed to decide sent to a person.

And when it cannot

It asserts a status the file does not support. Three of 64 inquiries went that way, each one a decoy the corpus was built around: the sentence the claimant asked about exists, against the other line item or at an earlier stage.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • An automated front desk over case files, where some questions are barred outright rather than answered carefully — This shape, as it stands.
    The barred classes are a ROUTE the model must choose, not a topic it must handle delicately - which is why they are gradable in code and why the counters read 0 of 14 and 0 of 10.
  • You need the accuracy number to beat free code — Not this, on this evidence.
    50 of 64 against a word list's 44 of 64 does not survive a paired test at n=64 (p = 0.263). A bigger labelled set would settle it, and that is more calls.
  • Case files too large to send whole — This shape plus a retrieval step, and measure the retrieval separately.
    Everything here assumes the file fits the prompt; the whole 2.7 KB goes in and the decoys are resolvable because nothing was dropped.

At a glanceHow the whole thing runs

78%route exact pct
1,206 msp50, end to end
$19.01per 1,000 claim files · Claude Fable 5

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/claims at your own case files and rewrite src/claimfile.py::parse to return (entry id, recorded at, line item, stage, text) for each log entry. Nothing measured here transfers to your corpus. Corpus lens →
When is this the wrong choice?Avoid: If the barred class is a matter of degree rather than of kind, this shape gives you a precision you have not got. There is no confidence here to tune. That is the case against the best-fitting scenario (“An automated front desk over case files, where some questions are barred outright rather than answered carefully”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A handling log with no recorded time on an entry. The whole contract turns on 'recorded at or before the moment they asked'; an undated entry cannot be admitted and cannot be excluded, and this kit has no third answer for it. 4 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?THE MARGIN OVER FREE CODE. +6 inquiries at n=64, exact McNemar p = 0.263 rechecked and 0.383 raw. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is OpenAI-compatible endpoint; the Prompt lens states what swapping it costs. The published figures come from 3 models on the fast tier and free code, no model and the constant, which decides nothing. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-claimant-status - 64 claimant inquiries, 64 model calls, the fast tier, reasoning disabled. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured builds the corpus in about a second, scores the three free arms at $0.00, and renders the whole board by replaying the committed run record - the one live control is disabled when no key is present.

A living map of modern AI — kept current every morning