Home › Use Cases › Flag which SIU fraud indicators a claim file raises and quote the line behind each
Use caseUC0277
🧪 Use-case kit · runnable

Flag which SIU fraud indicators a claim file raises and quote the line behind each

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A claim file accumulates facts that an SIU indicator library says are worth recording — a document reference used twice, an estimate pricing something the loss description does not mention, a payee on neither the policy nor the estimate, a notice date well outside the file's own norm. Somebody has to walk ten checks over every file, and then do the harder half: notice when the file ITSELF already explains the fact, and notice when what looks like an explanation names a document the file does not carry. Six of the ten indicators are true of large numbers of entirely ordinary claims, so the cost of getting this wrong is not a missed fraud — it is an ordinary claimant recorded against something their own file answers. Walking a ten-check indicator library over a claim file by hand and writing down which checks observed their fact. It does not replace the SIU analyst's disposition, the handler's coverage decision, any referral, or any decision about how the claim is handled — none of those exists anywhere in this pack, structurally.

Audience

An SIU support analyst working a queue of claim files, and the desk handler whose file it is. The decision this output is for is narrower than it looks: not 'is this claim fraudulent' and not 'should this claim be paid', but 'which named, checkable facts are on this file, and how well evidenced is the set'. Those are four different jobs belonging to four different people. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual claim file

The corpus is 64 claim file, 0.11 MB (txt 64). It is generated because it has to be. A real SIU indicator review pack is somebody's claim file: a named policyholder, their address, their loss, their prior claims and an adjuster's notes about them. There is no version of publishing that. And there is no public indicator library to check against either, so the library is invented and says so twice. What is measurable on a generated corpus is exactly what this kit claims: whether an arm applies a STATED library correctly to a STATED file, including the two rules that require reading.

The corpus

  • The 64 claim filegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNothing is fetched, scraped or adapted. One generator, one seed, one pass — see data/SOURCES.md, which also states what the generator costs the measurement and names the one regularity still in the corpus.

Swap this folder for your own material and the kit is pointed at your claim file. That is the whole change — there is no database to migrate.

One claim file, as the model receives itFI-0001.txt · 1 of 64
CLAIM FILE  FI-0001                             indicator review pack prepared 2026-09-02
Library     SIL-2026

POLICY
  Policy reference          POL-145076
  Cover                     household buildings
  Inception                 2025-09-04
  Last coverage change      none recorded
  Policy address            7 Bramsford Way, Marrenby
  Correspondence address    7 Bramsford Way, Marrenby

LOSS
  Loss date                 2026-07-04
  Notice date               2026-07-19
  Loss type                 escape of water
  Loss description          Escape of water from a failed flexible hose under the kitchen sink; damage to the kitchen floor and to the base units below the sink.

DOCUMENTS ON FILE
  DOC-01  2026-07-19  first notice of loss     ref FNL-11964    taken by telephone; description recorded as above
  DOC-02  2026-07-25  repairer estimate        ref EST-81141    kitchen floor and base units
  DOC-03  2026-08-02  photographs              ref PHO-50536    13 images of the damage described

CLAIMS HISTORY
  Prior claims in the last 24 months   1
  CLM-839341  2026-04-28  accidental damage  withdrawn

CROSS-REFERENCE  servicing parties on other open files
  Bracehill Restoration Ltd      appears on 1 other open file   (FI-0047)

PAYMENT INSTRUCTION
  Payee                     the policyholder
  Named on the policy       YES
  Named on the estimate     NO

FILE CHECKLIST
  Independent report expected for this loss type   NO
  Independent report on file                       NO

EXPLANATIONS RECORDED ON THE FILE
  none recorded

ADJUSTER NOTES
  Contact made 2026-07-22; the file is with the desk handler.

The outcomeWhat a good result looks like

One file in, four graded answers out: the raised set of indicator codes, the leading code under the library's reading order, the evidence band the raised set falls into, and one line COPIED VERBATIM out of the file establishing the leading code — with NO line where nothing is raised, because quoting one there is also a wrong answer.

And when it cannot

And what it does when it cannot. On the scored run the call over-flagged 8 of the 64 files and under-flagged none. Every one of the eight is the same failure: a file whose defining fact is present, whose own text records the explanation, and whose explanation names a document the file carries — so the first defeat rule closes it and the answer key raises nothing. The call raised it anyway, at a median confidence of 0.89. It respected 2 of the 10 documented explanations in the corpus. The free rules floor respected 5.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want every checkable fact on the file written down — the paid call
    100.0 pct recall over 640 indicator decisions — it did not miss one. The free floor recalls 92.9 pct.
  • You want as few wrong flags as possible — the free rules floor — evals/baseline.py
    91.2 pct precision against the call's 87.5, 5 over-flagged files against 8, for $0.00. On this corpus paying makes the over-flagging worse.
  • Your files carry assertions from interested parties — the paid call
    8 of 8 resisted, where the named document is not on the file. The floor swallows 4 of the 8 — and on a real queue that is the failure that quietly closes real indicators.
  • You want a whole file right more often than a weekend script gets it — no evidence either way
    55 of 64 against 55 of 64, McNemar exact p = 1.00.

And where nothing here is good enough:

  • Your files carry documented explanations that close a fact — NEITHER, yet
    2 of 10 for the call and 5 of 10 for the floor. This is the unsolved half of the job and the honest answer is that this kit does not have it. A second reader on every file whose raised set is non-empty AND whose explanations panel or notes name a document is where the next improvement is.

At a glanceHow the whole thing runs

86%file all correct pct
38,789 msp50, end to end
$19.69per 1,000 claim file · google/gemini-3-flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own claim files, data/library.json and data/library.md with your own indicator list, and data/gold.jsonl with your own answer key. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: If a wrong flag is expensive to you. This arm buys its recall with the worst over-flag rate of the three: 8 files against the free floor's 5, and every one of them is a claimant recorded against a fact their own file already answers. That is the case against the best-fitting scenario (“You want every checkable fact on the file written down”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A scanned or photographed claim pack. Every arm here — the model's reading, the floor's panel parsing and the independent key check's — rests on a fixed-layout text file. 4 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Run-to-run variance. One paid run of the scored arm; no repeat was bought. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-fraud-indicator. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:8277 and scores every graded cell offline: the corpus rebuild from seed 20260902, the filler check, the independent label gate, both free floors and the replay of the committed paid run. The only thing a key buys is a new call.

A living map of modern AI — kept current every morning