Home › Use Cases › Compare an expiring policy against the renewal offered for it, term by term
Use caseUC0398
🧪 Use-case kit · runnable

Compare an expiring policy against the renewal offered for it, term by term

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A commercial policy comes up for renewal and the carrier sends a new schedule of terms. It looks like last year's. Somewhere inside it a sublimit has been cut, a deductible now applies per claim instead of per occurrence, an exclusion has grown a clause, or a valuation basis has moved from replacement cost to actual cash value — and the premium is the same. Nobody notices until somebody reads the expiring schedule against the renewal offer line by line, six coverage parts against six coverage parts, with the expiry date coming. Reading an expiring policy schedule against the renewal offered for it, coverage part against coverage part, and writing down which terms changed and which way. It does not replace deciding what to do about a change, it never says whether the renewal should be taken, and it never prices anything.

Audience

A renewal underwriter or a commercial broker putting the expiring policy beside the renewal offer before it is bound, and the person who has to decide what to do about a change once it is found. This kit produces the FILE they read. It never makes the decision, never prices anything and never advises the insured. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual renewal packs

The corpus is 63 renewal packs, 0.26 MB (txt 63). It is generated because it has to be. A real policy schedule is a carrier's contract document carrying a named insured, a real premium and real filed wording, and publishing an expiring schedule beside its renewal offer would be republishing somebody's contract and their price. More than that: an exclusion clause that LOOKS like real filed wording is the failure mode this whole subject has, so the corpus is built so that the SHAPES are real and every identifier is in a namespace nothing uses. That is also what makes the answer key derivable — the generator gives every cell a position on the term's own scale before any text exists, so the direction is structural rather than a judgement somebody typed.

The corpus

  • The 63 renewal packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 63 renewal packs and the whole answer key is generated. Nothing here is fetched, scraped, licensed or derived from anything that was.

Swap this folder for your own material and the kit is pointed at your renewal packs. That is the whole change — there is no database to migrate.

One renewal pack, as the model receives itRNW-0001.txt · 1 of 63
COMMERCIAL POLICY RENEWAL - TERMS COMPARISON PACK

PACK HEADER
  Policy record       PL-0001
  Line of business    Commercial property and general liability
  Documents           EXP, REN

DOCUMENT EXP
  Document            EXP
  Paper               Expiring policy
  Policy reference    POL-7674-EXP
  Carrier             Carrowmoor Specialty Insurance
  Named insured       Lislevane Marine Services
  Policy period       2025-11-06 to 2026-11-06 (12 months)
  PART  1  Building limit                                USD 12,000,000
  PART  1  Contents limit                                USD 1,000,000
  PART  1  Business interruption sublimit                USD 2,000,000 in the aggregate
  PART  1  Flood sublimit                                USD 500,000
  PART  1  Earth movement sublimit                       USD 1,000,000
  PART  1  Valuation basis                               Replacement cost
  PART  2  Each occurrence limit                         USD 2,000,000 per occurrence
  PART  2  General aggregate limit                       USD 2,000,000 in the aggregate
  PART  2  Products and completed operations aggregate   USD 2,000,000 in the aggregate
  PART  3  Perils basis                                  Named perils only
  PART  3  Named exclusions                              Loss caused by mould; Loss caused by seizure by a public authority
  PART  3  Exclusion carve-back                          This exclusion does not apply to fire or explosion that follows
  PART  4  Property deductible                           USD 25,000 per claim
  PART  4  Wind and hail deductible                      USD 250,000 per occurrence
  PART  4  Liability retention                           USD 25,000 per claim
  PART  4  Deductible basis                              Per occurrence

Abridged — the file continues.

The outcomeWhat a good result looks like

One renewal pack in, one comparison report out: every comparable term with the first rule of RTC-2026 that reaches it, its DIRECTION, the papers that stated a value, both sides quoted verbatim out of their own schedules, the coverage parts referred and one status — plus the pack-level finding this whole kit exists to write, COVER-REDUCED-WITH-NO-PREMIUM-REDUCTION. 50 of 63 reports came back completely correct after the pure-code station.

And when it cannot

⚠︎ AND THE HONEST HEADLINE IS THAT THE PAID CALL DID NOT BEAT THE FREE FLOOR THIS KIT SHIPS. Zero of the 63 calls failed to return a report — every reply parsed, none reached the ceiling, every one of the 132 quotations RTC-2026 owed came back and every one was located in the paper it was attributed to. But 13 of 63 reports are still wrong after the station, and 12 of those 13 wrong reports are the SAME failure: the call answered DIRECTION-NOT-DETERMINABLE on a term whose own scale gives a direction. It saw the change, refused to say which way, and the headline finding never fired. On the cover-cut-with-no-premium-cut finding the free trigram floor is ahead, 19 of 26 against 18. Confidence does not separate right from wrong well enough to route on: the median on a correct report is 0.97 and on a wrong one 0.90, with a floor of 0.80.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your two schedules are already structured and differ only in how an amount or a percentage is printed, or in the order of a list — src/rules.py's R-4 normaliser, with no model at all
    1,107 of the 1,138 comparable terms here are settled by R-4 and R-5 together, at $0.00 and with no network. There is no measured case for spending anything on them.
  • A limit, a sublimit, a deductible, a retention or a premium moved, and both papers state it on the same unit with the same qualifier — src/rules.py's R-5 arithmetic, still with no model
    the direction is subtraction against a scale data/terms.json declares. Every arm on this board gets those terms right and none of them pays for it.
  • An exclusion was re-scoped, a basis was paraphrased, or a deductible's amount held while its qualifier moved — ⚠︎ MEASURE BOTH BEFORE YOU BUY. This is the seam the call is for and on this corpus it does not win it.
    18 of the 31 terms that need a reading, against the free floor's 11 — and on the 16 that are a real reduction in cover the FLOOR is ahead, 8 to 6. The exact paired test on whole reports gives p = 0.19.
  • You need the referral list to be short enough that people keep reading it — the paid call, and this is the one thing it clearly buys
    it wrongly referred 2 packs against the floor's 5, and it inverted no directions at all where the floor inverted 14. A referral list padded with re-wordings and good news is how a real reduction goes past somebody.

At a glanceHow the whole thing runs

79%rechecked report all correct pct
4,075 msp50, end to end
$7.66per 1,000 renewal packs · google/gemini-3-flash

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own renewal packs in the same panel shape — a PACK HEADER, one DOCUMENT <CODE> panel per paper with PART <n> <term> <value> rows, a COMPARABLE TERMS list and PACK NOTES — and re-run both floors, which need no key. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO, AND IT IS THIS KIT'S HEADLINE. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call whose answer a normaliser and a comparison already produced. That is the case against the best-fitting scenario (“Your two schedules are already structured and differ only in how an amount or a percentage is printed, or in the order of a list”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A real policy schedule, which arrives as a PDF. There is no OCR, no PDF reader and no layout model in this kit; the corpus is already columnar text and getting a real schedule into that shape is work this kit does not do and does not measure. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the paid call beats the free floor at all. It does not, on this corpus, at any significance a reader should act on: 50 of 63 reports against 43, exact two-sided McNemar p = 0.19, and the floor AHEAD on the finding the kit exists to produce. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-renewal-terms. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9398 — all 63 renewal packs, both free floors computed live, RTC-2026, the term schedule with every direction scale, and every committed run replayed at $0.00. The one control that would spend is disabled and says why. That state is one of the five published frames.

A living map of modern AI — kept current every morning