The business caseThe problem this solves
A loss file lands and somebody has to answer one question before it closes: is there a party other than the insured that a recovery could run against, and what in the file says so? A file tells you WHO DID WHAT, and that is the one thing only a reader can supply. Everything after it -- whether a waiver covers that party, whether they are named on the same schedule as the party claiming, whether the amount clears the threshold, whether the period has run -- is a role lookup and two comparisons. Deciding a recovery is mostly not a reading problem. It is a clause lookup with a reading problem in front of it, and the reading problem is where the file hides the answer: on nine of these sixty files the narrative says the cause was never established and the attached engineer's report names the party who installed the failed part. The reviewer's manual pass over one loss file against the recovery terms -- reading the narrative, the parties block and the attachments for who is responsible, then checking the waiver, the named-insured schedule, the threshold and the limitation period -- before a person confirms the decision. It replaces neither the confirmation nor the recovery action; there is no endpoint here that opens one.
Audience
The recovery or subrogation reviewer who decides whether a file closes or a recovery is opened, and the claims owner who answers for both a recovery that was never opened and an action opened against a party the terms had already given up the right to pursue. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual loss files
The corpus is 60 loss files, 0.07 MB (txt 60). The cases are the ones that make this decision hard, and none of them can be scraped. A contractor who plainly caused the loss and against whom recovery has already been waived -- the file's own terms bar it while every trigger word in the narrative says otherwise. A file whose narrative says the cause was never established and whose attached engineer's report names the party who installed the failed part. Two accounts on one file that contradict each other, each side blaming the other, where there is a party named and there is no answer. A second vehicle that left before details were exchanged -- a third party is responsible and there is nobody to pursue, which is not the same as nobody being responsible. A file where the terms need to know whether an agreement was executed before the loss and the file DOES NOT RECORD IT, where reading absent as 'no' opens a waived recovery. And a decoy: two files that list the supplier who delivered the load first, so an arm that takes the first party who is not the insured names somebody the schedule does not cover. All five clauses are in force on EVERY file, deliberately -- a reader who checks whether a file MENTIONS a waiver learns nothing here, because every file mentions one; the question is always whether the condition is met. A quarter of the set is clean, because a corpus that is all traps measures a different job from the one the reviewer has.
The corpus
- The 60 loss filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your loss files. That is the whole change — there is no database to migrate.
LOSS FILE LF-0001 -- Impact damage -- yard
FILE FACTS
Reference RF-1001
Date of loss 2024-02-03
Amount claimed USD 21,500
Location the north yard shutter, the Ockley depot
Named insureds on schedule Brantley Marine plc
Services agreement none in place
INCIDENT NARRATIVE
A member of the insured's staff reported that the north yard shutter had been struck at some point during the shift. The site was in normal operation and the area was not closed to vehicles.
Setterfield & Co Ltd's driver reversed a box van into the north yard shutter while turning in the yard, and Setterfield & Co Ltd has confirmed in writing that the vehicle and its driver were under its control at the time.
The damage is confined to the north yard shutter. Photographs taken the same day are on the file and the amount claimed is the repair quoted by the insured's own contractor.
PARTIES ON THE FILE
Brantley Marine plc the insured itself
Setterfield & Co Ltd a supplier or visiting party with no services agreement in place
ATTACHMENTS
none
The outcomeWhat a good result looks like
A recovery decision a person confirms instead of a narrative and a set of terms read side by side: the verdict, the party a recovery would run against spelled as the file spells it, the clause that bars it where one does, and ONE SENTENCE COPIED VERBATIM out of the file that establishes the third party's responsibility -- located in the file and highlighted where it was found, never paraphrased.
And when it cannot
Two directions, and they cost different things. A MISSED RECOVERY is a recoverable file called NOT-RECOVERABLE: the money is written off and nothing downstream ever reports it, because there is no queue of recoveries that were never opened. A FUTILE PURSUIT is the reverse -- an action opened against a party a clause in force bars, which is handling time spent for nothing and, where the waiver sits inside an agreement, a breach of that agreement. On this corpus the free rules floor ships 6 missed recoveries of 23 and 3 futile pursuits of 27; the model ships 0 of each. The model's own failure is a third thing neither of those columns catches: on 4 files it named a responsible party where the file establishes none.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Files where the responsible party and the sentence that says so are both in the narrative, on a book with a clean register — the free rules floor alone
On the six plain_third_party and four rear_impact_clear files it gets the verdict and the party right 10 of 10 -- it only loses the citation, quoting an earlier trigger-carrying sentence. And it gets EVERY clause free: 5 of 5 waiver_bars, 4 of 4 waiver_present_not_applicable, 3 of 3 below_threshold, 3 of 3 missing_bar_fact and 3 of 3 narrative_implies_bar_language, all-correct, because the clause work is the terms engine and the terms engine costs nothing. - Files that carry the finding in an attachment while the narrative says the cause was never established — the model, rechecked
9 of 9 against the floor's 0 of 9, and the floor's zero is STRUCTURAL: it reads the narrative only and cannot be tuned to score anything else. On six of those files the floor's answer is a missed recovery and on the other three it is the right verdict for a reason the file does not contain. This row is the whole case for the call.
And where nothing here is good enough:
- Files where two accounts contradict each other and the right answer is to refuse — NEITHER ARM, AND THIS KIT SAYS SO. Refer them by rule.
conflicting_accounts is the ONE case where both arms score 0 of 4 all-correct. The floor answers a confident RECOVERABLE against the party named first. The model calls the verdict right 4 of 4 and cites a real sentence 4 of 4 and then NAMES A PARTY ANYWAY on all four -- its entire party_invented count and 4 of its 5 misses. The kit has no control for it: recheck_overrides is 0 and the terms engine checks clauses, not whether a name belongs.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own loss files as .txt into data/corpus/, add one row per file to data/files.json (id, loss_date, loss_amount, currency, named_insureds, parties with a role from src/verdicts.py::ROLES, and agreement_executed_before_loss as true / false / NULL FOR 'NOT RECORDED' -- the third state is load-bearing, because absent is not 'no'), and write data/gold.jsonl with the verdict, the party and the citation. ⚠︎ A REAL LOSS FILE NAMES THE PARTIES TO A LIVE DISPUTE AND QUOTES STATEMENTS TAKEN FROM PEOPLE, AND THE WHOLE FILE REACHES YOUR CONFIGURED PROVIDER VERBATIM -- the third party's name, the driver's statement, the engineer's report and anything said about anybody. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per file for the clause lookup. Six of the fifteen cases here -- 23 files -- are a tie at 100 pct on both arms, and every one of them is a file whose narrative says who did what. That is the case against the best-fitting scenario (“Files where the responsible party and the sentence that says so are both in the narrative, on a book with a clean register”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A REAL FILE. Sixty files from a generator with fifteen body templates and small phrase pools, with substituted names, places, dates and amounts. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE PURE-CODE STATION ACTUALLY WORKS. recheck_overrides is 0: the terms engine re-applied every clause, threshold and period test to the model's answers and vetoed nothing. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-recovery-flag. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — ⚠︎ PARTIALLY EVIDENCED, AND THE REST IS NOT MEASURED. What is on disk and checkable: the 60 files, the register, RT-2026 in both renderings, the answer key with its citation spans, and TWO COMMITTED KEYLESS RUNS -- results/eval-b000-recovery-flag-rules.json (model 'free-floor:rules', usd null, 29 of 60) and results/eval-t000-recovery-flag-stub.json (model 'stub', usd null, 29 of 60, largest reply 146 tokens) -- neither of which could have been produced with a provider in the loop. The scored run's cache (results/cache-r001-recovery-flag.jsonl) ships too, so the model columns render without a call. NOT MEASURED: no timed fresh-clone run was recorded for this kit -- no HTTP-200-on-every-endpoint sweep of the board and no stopwatch on tools/build_corpus.py or evals/check_labels.py. The claim here is 'the artefacts that make the free half work are committed', not 'a cold clone was timed'.



