Home › Use Cases › Extract the facts an incident report states and flag every conflict with the claim file
Use caseUC0249
🧪 Use-case kit · runnable

Extract the facts an incident report states and flag every conflict with the claim file

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A first notice of loss arrives as two accounts of one event. There is the REPORT — the insured's statement as it was taken plus the adjuster's field notes, free text in somebody's own words — and there is the RECORD FACTS header, which is what an intake clerk typed into the claim system while the call was still going. The header is clean, structured, indexed, and read by every downstream system in the carrier, none of which will ever open the narrative. The header is wrong more often than the report is, and it is the one everything reads. Today the only thing that catches that is an investigator opening the file and reading both, and most files are never opened. The read-both-accounts pass over a first notice of loss — the extraction and the discrepancy check, not the resolution, which stays with a named adjuster and whose procedure is a precondition of deploying this at all.

Audience

A claims investigator deciding which of a day's files actually needs reading, and whether the extracted facts in front of them can be trusted well enough to act on. The number they are buying is not per-field accuracy; it is how many of the contradictions in the file would have reached them. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual incident reports

The corpus is 62 incident reports, 0.08 MB (txt 62). A real first notice of loss cannot be published by anybody, ever — the narrative IS the sensitive part and no redaction survives it, because the whole job is reading what a named person said about an event on a named street. So the corpus is invented end to end: no claim, no person, no medical detail, no place, no citation to any regulation. What it exercises is the thing that actually matters and is otherwise unmeasurable — a document that contains TWO accounts of one event which do not agree, one of them clean and structured and read by everything, the other free text and correct. Every property that makes the measurement work is stated as constructed in data/SOURCES.md rather than presented as observed.

The corpus

  • The 62 incident reportsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your incident reports. That is the whole change — there is no database to migrate.

One incident report, as the model receives itIR-0001.txt · 1 of 62
INCIDENT REPORT IR-0001 -- Side impact at a signalled junction

RECORD FACTS  (typed by intake at first notice; this is the claim system's record,
               not the report)
  Claim number                  CLM-2026-04100
  Policy                        PE-8800000        Line: personal-auto
  Reported                      2026-01-02, by the named insured, by telephone
  Loss date                     2026-01-01
  Location type                 intersection
  Injuries reported             Y
  Police report                 filed
  Witnesses recorded            0
  Vehicle                       a 2019 sedan, insured driver

INSURED'S STATEMENT  (verbatim, as taken)

  I did not report it straight away; I finally rang you on Wednesday 2 January 2026.
  It happened on Monday 1 January 2026, at about 6:00 in the morning.
  We were in the middle of the junction of Marston Avenue and Ninth Street when it happened.
  There was a bang and a 2022 crossover had gone into my rear.
  One of the people in the other car was limping and said her knee had gone.
  The police attended and gave me a report number before they left.
  I have photographs on my phone if they are any use.

ADJUSTER FIELD NOTES

  note 1  Spoke to the insured; statement above is as taken and has not been edited.
  note 2  Left a message for the third-party insurer, no response so far.

ATTACHMENTS
  - photographs (2)
  - repair estimate

The outcomeWhat a good result looks like

Four extracted facts per report — the date of loss, the kind of place, whether anybody was hurt, whether police attended — each taken from the report and not from the file, plus a conflict flag wherever the two disagree, with the sentence an adjuster has to read in order to resolve it. The flag leaves OPEN.

And when it cannot

It answers the header instead of the report, and the contradiction never surfaces — an injury written out of the file, a loss date moved, an independent police account nobody goes and gets. Or it floods the investigator with conflicts raised on files that have none, which costs the same read as a real one and is how a flag gets switched off. Or it concludes fault, and a liability position enters the claim file from a machine.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know WHETHER your headers can be trusted, and have never measured it — the free keyword floor first
    It reaches 24 of the 26 contradictions here and all 12 carrier-favourable ones for $0.00. It will over-raise — 31 flags to find 24 — but a noisy first count of how often your file and your narrative disagree is worth more than a clean count you have not taken.
  • An investigator has to work every raised flag, so a false one costs the same read as a real one — the model
    That is the whole margin here. Recall is close — 26 of 26 against the floor's 24 — and precision is not: 100.0 pct against 77.4 pct, which on this file is 5 wasted reads against 0.
  • The narratives are dictated, transcribed, or full of half-corrections — a fact stated one way and then another — the model, and re-measure before believing anything on this page
    The 30 planted distractors are the closest thing this corpus has to that, and they are what the free floor loses on: 87.9 pct field accuracy against 100.0 pct. But they are ONE sentence competing with one sentence, deliberately unambiguous under IX-4, and a real half-corrected narrative is not.

And where nothing here is good enough:

  • A claim file whose RECORD FACTS headers you already trust — the intake is clean and the narratives rarely say anything the header does not — neither arm; run nothing
    The null floor IS trusting the header, and on this corpus it scores 86.3 pct of the field judgements for $0.00. If the contradictions are genuinely rare, the pack's whole output is a queue of flags nobody needs to work.
  • You want the pack to settle the contradiction rather than raise it — nothing here
    IX-2 forbids it and there is no field in this kit in which a resolution could be recorded. That is deliberate and structural: a pack that could record one is a pack that could quietly make one. The conflict-resolution procedure is a precondition of deploying this at all, and it is a person's.

At a glanceHow the whole thing runs

100%field accuracy pct
8,214 msp50, end to end
$5.50per 1,000 incident reports · Gemini 3 Flash

Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt and data/gold.jsonl, and put your own header values in data/records.json — the vocabularies, the conflict codes and their severity order live in data/policy.json. Every measured number stops being true. Corpus lens →
When is this the wrong choice?Avoid: Buying a model to confirm a file you already believe. That is the case against the best-fitting scenario (“A claim file whose RECORD FACTS headers you already trust — the intake is clean and the narratives rarely say anything the header does not”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A report with no RECORD FACTS header at all. IX-5 needs two sides; with nothing to compare against, every fact is extracted and no conflict can be raised. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled contradicting sentence is the sentence a claims investigator would have quoted. The key names one per conflicted report and would score a neighbouring one carrying the same fact at zero. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-01 — r001-incident-extract. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, scores both free floors offline, re-derives the answer key and the entire register independently (evals/check_labels.py, IX-5 retyped inside itself, 0 disagreements), and replays the committed scored run report by report. Nothing was installed: requirements.txt is empty and the kit is Python standard library end to end.

A living map of modern AI — kept current every morning