Home › Use Cases › Duplicate claim detection
Use caseUC0362
🧪 Use-case kit · runnable

Duplicate claim detection

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A loss is reported. Before a claim number is issued somebody has to ask whether this loss is already in the book — and the two records almost never agree on paper. The insured rang it in and the broker emailed it the same day; the loss date is typed a day out, or the date the caller rang is entered as the date it happened; one record says animal strike and the other says hit a deer; the road is named by its number on one and by a landmark on the other. Today an intake handler decides by eye, at the pace notices arrive. And the hard half is the other direction: the same vehicle, the same peril and the same week are NOT one loss, and the records that look most alike are exactly the ones where stopping is worst. Reading a new notice against the claims already in the book and deciding whether a second claim file opens. Not the decision itself — the kit stops one step before that, at the pair a handler confirms with both records shown — and not the field comparison either: the free floors ship inside the kit and are computed live on every pair the board draws, because the honest question is what the model buys ON TOP of them.

Audience

The intake handler who takes the notice and the claims manager who owns the duplicate control. The decision is not 'is this matcher accurate' — it is 'what happens the first time it is wrong, and in which direction', because the two directions are not the same size of mistake. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual candidate pair sheets, each one a new notice and a prior claim side by side

The corpus is 64 candidate pair sheets, each one a new notice and a prior claim side by side, 0.15 MB (txt 64). Because a real first notice of loss is somebody's accident, their address, their vehicle, their injury and their money, filed on the day it happened — and the exact shapes this corpus is built to measure are the records a claimant would least want published. There is no public corpus of paired FNOL records and there should not be one. So all of it is invented from one seed, and python3 -m tools.build_corpus --check rebuilds every byte and diffs against disk. What it buys is a labelled negative class: 17 of the 64 pairs are a genuinely DIFFERENT loss sharing the policy, the asset, the peril and often the week. Without those there is no denominator for a false-stop rate, and an arm that stops everything cannot be told from one that reads.

The corpus

  • The 64 candidate pair sheets, each one a new notice and a prior claim side by sidegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your candidate pair sheets, each one a new notice and a prior claim side by side. That is the whole change — there is no database to migrate.

One candidate pair sheets, each one a new notice and a prior claim side by side, as the model receives itPAIR-0001.txt · 1 of 64
CANDIDATE PAIR              PAIR-0001
Blocked on                  policy cpa445100; asset 1fujgldr9klb96620; insured vasquez, loss
                            dates 0 days apart
Book searched               64 claims open, or closed within 90 days

NEW NOTICE

Notice reference            FNL-2026-004000
Policy number               CPA-445100
Prior policy number         -
Named insured               Vasquez Haulage
Also known as               -
Reported by                 Dana Vasquez (named insured)
Report channel              broker email
Reported at                 2026-04-28 14:40
Date of loss                2026-04-27
Time of loss                late afternoon, before dusk
Loss location               Rt 9 northbound, just past the Milbank overpass
Cause of loss               animal strike
Vehicle or property         2019 Freightliner Cascadia tractor, ref 1FUJGLDR9KLB96620
Coverage claimed            comprehensive
Damage described            near side of the cab, mirror housing and lower step
Other claimants             none beyond the insured
Estimate given              -
Notes                       -
Narrative                   Broker advises the insured's tractor struck a deer on the northbound
                            carriageway. Damage to the passenger side and the mirror. Unit is
                            drivable.

PRIOR CLAIM

Claim number                CLM-2026-112420
Policy number               CPA-445100
Prior policy number         -
Named insured               Vasquez Freight LLC
Also known as               Vasquez Haulage
Reported by                 Dana Vasquez (named insured)
Report channel              telephone, out-of-hours service
Reported at                 2026-04-27 09:12
Date of loss                2026-04-27

Abridged — the file continues.

The outcomeWhat a good result looks like

One candidate pair in; three readings, one of five places, and BOTH records' own lines quoted back — so a handler confirms a decision with the evidence in front of them instead of opening two claim files. On this run the pass stopped 43 of 44 pairs that must be stopped and closed NONE of the 17 genuinely separate losses.

And when it cannot

When it cannot, it attaches. Every one of this run's 8 wrong routings went the same way — LINK-SUPPLEMENT, hanging a notice on an open claim instead of escalating it or opening its own — and 2 genuinely separate losses were held at intake that way. It never once closed a real claim, which is the harm it was most at risk of; what it does instead is bury a claim inside another. And on 47 of 64 pairs it could not produce the two lines that show its work.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You need to know whether a duplicate control is safe to switch on — claim-duplicate-stop, read as three numbers
    It is the only grader that separates the two harms. A single accuracy figure is the number the merge every pair floor scores best on.
  • The output goes in front of a person who has to confirm it — claim-duplicate-evidence
    It is the half this kit is named for and the half the paid arm is weakest on: 17 of 64 pairs.
  • You are deciding whether the model is worth the money over free code — claim-duplicate-significance
    It answers the question with a denominator, and it answers it BOTH ways on this run: the routing clears, the whole answer does not.
  • You want the whole row, exactly right, no partial credit — claim-duplicate-pair, pair_all_correct
    It is the strictest thing the kit measures and the honest one to lead with: 15 of 64.

At a glanceHow the whole thing runs

81–83%disposition exact pct
1,618 msp50, end to end
$3.45per 1,000 candidate pair sheets · Gemini 3 Flash

Run once, for real, on 2026-09-10. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point tools/build_corpus.py at your own notices, or write data/corpus/*.txt and data/gold.jsonl directly — one pair sheet per file, one gold row per sheet with the three readings and the two evidence lines. EVERY MEASURED FIGURE HERE STOPS AT YOUR CORPUS BOUNDARY. Corpus lens →
When is this the wrong choice?Avoid: Any single headline. The first number alone is owned outright by an arm that reads nothing and holds everything. That is the case against the best-fitting scenario (“You need to know whether a duplicate control is safe to switch on”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A POLICY RENEWED UNDER A NEW NUMBER THE BOOK DOES NOT RECORD. Planted, and one of the two true links the key cannot make at all. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE MODEL OR THE RULE ENGINE HELD OFF THE INJECTION. All 12 probe calls held and the recorded readings show the model itself answered different_loss, so on this evidence it refused. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-10 — r002-claim-duplicate-reasoning-off. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, replays the committed scored run, computes all three free floors live, re-derives the answer key with a second parser and re-runs the significance test — all of it in pure Python, in under a second of wall time, for $0.00. The single control that would spend is disabled and prints why on the page.

A living map of modern AI — kept current every morning