Home › Use Cases › Is the replacement packet complete -- and is a replacement even being disclosed
Use caseUC0172
🧪 Use-case kit · runnable

Is the replacement packet complete -- and is a replacement even being disclosed

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

<p>A life application that replaces an existing contract owes a packet, and checking it is two questions rather than one. The first &mdash; are the five forms on the register, signed, with the existing carrier and policy number filled in, and dated in the order the carrier's bulletin requires relative to the application date &mdash; is exactly what code does well. It is presence and arithmetic on dates, and free code settles all 100 of the cells on the disclosed files in this corpus perfectly, for nothing.</p><p><b>The second question is the one that costs money: is a replacement even being disclosed?</b> The application asks it outright at Q7b, and a file can answer <i>No</i> while its own funding source, its surrender request, or a sentence in the producer notes says otherwise. 13 of these 52 files are exactly that. On 7 of them a financial record proves it &mdash; a 1035 exchange, a surrender, a policy loan against a life or annuity contract &mdash; and code catches every one. On the other <b>6 the only evidence is a sentence</b>: <i>&ldquo;once this contract is issued and delivered the applicant will surrender the Kestrel policy&rdquo;</i>. Every register in those files is internally consistent and every form is legitimately absent, because nobody ever knew a packet was owed.</p><p>And it runs the other way too. 4 files carry a surrender request naming another carrier that is <b>not a replacement at all</b> &mdash; a disability income contract being cancelled, which the bulletin's own scope line excludes, or an annuity surrendered fourteen months ago whose proceeds went into a house. A structured check convicts the second pair and sends two clean files to a compliance queue.</p><p>Today that is a new-business examiner reading an application against a forms register, a funding register and two pages of producer prose, and deciding from that whether the answer on the front of the file is true.</p> Nothing is replaced. It produces a READING that an examiner acts on; it never issues, declines, releases or refers a file, and never contacts an applicant, a producer or another carrier. There is no such endpoint in src/app.py and no configuration flag that adds one. What it removes is the first pass: on this corpus pure code settles 220 of the 260 cells for $0.00 in under a second, and the paid arm is only being asked for the 13 files where the application says No and something else might not.

Audience

Whoever releases a life application to underwriting: a new-business examiner, a case manager at an agency, or the compliance desk that gets the file when somebody upstream was not sure. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual life new-business files

The corpus is 52 life new-business files, 0.23 MB (txt 52). A real new-business file carries a named applicant, their policy numbers, their medical requirements and the producer's own commission position. All of it is confidential, which is why there is no public corpus of (file, packet verdict) pairs -- and why publishing a scrubbed real one would be worse than publishing none, because the scrubbing is exactly where the interesting defect hides: the sentence in the producer notes IS the personal data and IS the answer. More to the point, the thing being measured has to be PLANTED to be measured. To count how often an undisclosed replacement is caught you have to know which files were concealing one, and a real archive does not come labelled -- the whole definition of the interesting case is that nobody recorded it.

The corpus

  • The 52 life new-business filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your life new-business files. That is the whole change — there is no database to migrate.

One life new-business file, as the model receives itRC-0001.txt · 1 of 52
Replacement Packet Check
----------------------------------------------------------------
  File: RC-0001
  Carrier: Calderwood Mutual Life Assurance Company
  Product: Calderwood Ten-Year Level Term (Form CML-TRM-2021)
  Applicant: Clemence C. Thurgood
  Application Date: 2026-03-07
  Face Amount Applied For: $210,500
  Modal Premium: $11,000, annual
  Producer: Emmet Cawdrey (Producer No. CM-37979)
  Agency: Northmarch Financial Partners
  File Opened: 2026-03-08

Packet Requirements
----------------------------------------------------------------
  Source: Calderwood New Business Bulletin NB-14, "Replacement Packet", revision 8.
  Where a replacement is being made, the following are owed before this file may be
  released to underwriting. Cite findings to the rule identifiers in this table.

  Rule       Form        Title                                           Requirement
  NB-14.2.1  CML-REPL-3  Producer Certification of Replacement           Signed by the producer. The date must be on or after the producer signature date on the Notice Regarding Replacement and on or before the application date.
  NB-14.2.2  CML-REPL-4  Existing Coverage Identification                The existing carrier name and the existing policy number must both be stated.
  NB-14.2.3  CML-REPL-1  Notice Regarding Replacement                    Signed by the applicant and by the producer. The producer signature date must be on or before the application date.
  NB-14.2.4  CML-REPL-5  Applicant Acknowledgement of Existing Coverage  Signed by the applicant. The date must be on or after the applicant signature date on the Notice Regarding Replacement.

Abridged — the file continues.

The outcomeWhat a good result looks like

One reading per file: the disclosure verdict with the identifiers that decide it, then one finding per required form -- disposition, reason, and the rule identifier this file's own renumbered table prints for it -- plus a file-level action. 99.23 pct of 260 items complete on all of that at once (r001-replacement-check), against the strongest free floor's 84.62 pct.

And when it cannot

⚠︎ IT LOSES TWO COLUMNS TO CODE THAT COSTS NOTHING, AND THEY ARE ON THIS PAGE. The free floor attaches the deciding identifiers on 88.10 pct of cells against the paid arm's 82.14 pct, and it INVENTS NOTHING -- 0 of 260 against the paid arm's 2. It is also level at 8 of 8 on incomplete elements, 8 of 8 on out-of-order dates and 7 of 7 on the structured undisclosed files. ⚠︎ AND THE WIN IS EIGHT FILES WIDE: the entire 14.61-point gap is 40 cells across 8 of the 52 files -- the 6 where the only tell is a sentence and the 2 whose funding record is fourteen months stale. On the other 44 the two arms agree cell for cell, and the model is being paid $0.0224 a file to reproduce them.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your replacement disclosures are trustworthy. When a producer says a replacement is happening it is, and when they say it is not it is not -- because your administration system posts every 1035, surrender and loan against the application before the file is worked, and the funding record is always there. — the free floor -- funding-flag, with FUNDING_TYPES and IN_SCOPE_KINDS re-mapped to your own transaction vocabulary
    It takes 100.00 pct of the presence, completeness and date-order cells, cites 100.00 pct of their rule identifiers correctly, attaches the deciding identifier on 88.10 pct of cells -- ahead of the paid arm's 82.14 pct -- and invents nothing at all, 0 of 260 against the paid arm's 2. Under a second for the whole corpus, with no key.
  • Your producers write notes, and the reason a replacement goes undisclosed is usually that somebody wrote down the plan and nobody read it -- a surrender held until delivery, a bank draft stopped, a loan planned for next year's premium. — a paid tier -- $0.02237071 per file on the projected card
    It is the only arm that scores at all on the column free code cannot reach: 6 of 6 prose-only undisclosed replacements against 0 of 6, and it also cleared both stale surrenders the floor convicts, quoting the fourteen-month gap back. It matched the floor everywhere else -- 8 of 8 blank elements, 8 of 8 out-of-order dates, 89 of 89 rule identifiers -- with 0 false referrals in 39 clean files.
  • You want both and you are willing to run two arms. — the free floor on every file, and a paid tier only where the floor's own verdict is in doubt
    They fail in opposite directions. The floor settles the mechanical half perfectly for $0.00 and cites it better than the paid arm does; the paid arm is the only thing that reads a sentence. On this corpus the paid queue is the 13 files whose application says No against something a register cannot settle -- 25 pct of the volume, about $0.29 rather than $1.16 -- and the UI already computes the disagreement column on every file for nothing.

At a glanceHow the whole thing runs

99%replacement form accuracy pct
46,209 msp50, end to end
$22.37per 1,000 life new-business files · Google Gemini 3 Flash

Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own files into data/corpus/ and your own labels into data/gold.jsonl, one row per file, with a disclosure verdict, the channel that reveals it, and one finding per required form. src/select.py withholds a NAMED SECTION and is not a redaction system. Corpus lens →
When is this the wrong choice?Avoid: Do not use it where the disclosure itself is the question: it catches 0 of the 6 files whose only tell is a sentence, and it refers 2 clean files whose funding record is fourteen months stale. Its 84.62 pct is also not parameter-free -- it depends entirely on FUNDING_TYPES and IN_SCOPE_KINDS matching your system's spelling. That is the case against the best-fitting scenario (“Your replacement disclosures are trustworthy. When a producer says a replacement is happening it is, and when they say it is not it is not -- because your administration system posts every 1035, surrender and loan against the application before the file is worked, and the funding record is always there.”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A file layout that is not this one. src/casefile.py is a set of regular expressions written for these underlined headings and these space-aligned tables. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether 6 of 6 on the prose channel would happen again. That column is the only reason to pay for this kit and its denominator is SIX FILES. 14 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-26 — r001-replacement-check. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — git clone, then python3 -m src.app. No pip install, no key, no network: the corpus, the answer key and every committed result file ship with the kit, all three free floors and the label check are pure Python, and the UI replays the scored run straight off results/eval-r001-replacement-check.json. The read button with no key returns a sentence, not an error.

A living map of modern AI — kept current every morning