Home › Use Cases › Group a quarter's settlement breaks into root-cause families with their evidence
Use caseUC0467
🧪 Use-case kit · runnable

Group a quarter's settlement breaks into root-cause families with their evidence

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

One payments corridor closes a quarter with a few hundred settlement breaks on it — a row per break with its value date, its amount in minor units, whether it is still open and how long it has been — plus the evidence artefacts somebody filed against them, the remediation items somebody raised, and a notes block where two desks have written to each other all quarter. Nobody disputes the breaks. What nobody has is the SHAPE: which root-cause families the quarter actually has, how much money and how much age sits behind each, which of them clear the desk's own promotion threshold, and — the part that takes the afternoon — WHICH ARTEFACT carries each break, once the notes that void an artefact have been read against the dates they carry. Today that is a controller with the rulebook open in one window and the dossier in the other. Reading one quarter's dossier end to end: testing every evidence artefact against four qualification tests (does it name the break, is it in a declared ledger, is it inside the ten-day window either side of the value date, does its status read FILED), then reading every note for a void and joining that note's own date to the artefact's, then assigning each break to the first family in the declared ladder its qualifying artefacts support — and only then doing the arithmetic.

Audience

A settlement controls desk assembling one corridor's quarterly root-cause pack, and the manager who has to decide whether to buy a model call for it. This report's own answer for this corpus is NO, and the numbers are laid out so that answer can be checked rather than taken. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual settlement break dossiers

The corpus is 64 settlement break dossiers, 0.20 MB (json 4 · jsonl 1 · md 2 · txt 64). It is generated because it has to be. A real settlement break dossier carries counterparty names, corridor volumes, amounts and dispute notes, and no processor publishes one; a public corpus of them does not exist. Generating it also buys the one thing a scraped corpus could never give: an answer key derived by APPLYING the rulebook to the rendered text, case by case, and asserted against what each case was built to produce. 17 cases were planted deliberately — a clean quarter, a void inside the quarter, the two opposite-effect voids, the confirmation decoy, the window edge at 10 and 11 days, the out-of-scope ledger, the draft artefact, the two-kind ladder, the no-evidence break, the threshold edge, the unattached remediation item, the closed break outside aging, the corroboration edge, the cap probe, the reason ladder and a dossier carrying two traps at once.

The corpus

  • The 64 settlement break dossiersgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from and — section 4 — what the generator was built to make the model fail on, written BEFORE any call was bought. One of those predictions was contradicted by the measurement and the contradiction is published rather than quietly dropped: the inert predecessor void was named the biggest expected miss, and the arm errs in BOTH directions about equally (voided_an_inert_record 17, missed_an_effective_void 23), which says it is not applying the date joins at all rather than applying them in one direction.

Swap this folder for your own material and the kit is pointed at your settlement break dossiers. That is the whole change — there is no database to migrate.

One settlement break dossier, as the model receives itSBD-0001.txt · 1 of 64
SETTLEMENT BREAK DOSSIER  SBD-0001
DOSSIER ASSEMBLED  2026-07-20   STANDARD  SBRC-2026
CORRIDOR  COR-4404   SCHEME  SCH-04   SETTLEMENT CURRENCY  CUR-D
QUARTER  2026-Q2   PERIOD  2026-04-01 TO 2026-06-30
DECLARED SCOPE  LEDGERS LDGCORE, LDGNET   EVIDENCE WINDOW 10 DAYS   PROMOTION 3 BREAKS

[1] BREAK RECORDS
ID          VALUE DATE  SIDE     AMOUNT MINOR  LEDGER    OPENED      STATUS
BRK-10001   2026-05-26  DEBIT    1299900       LDGNET    2026-06-03  OPEN
BRK-10002   2026-06-08  CREDIT   1229700       LDGCORE   2026-06-10  OPEN
BRK-10003   2026-05-14  CREDIT   2489800       LDGCORE   2026-06-01  OPEN
BRK-10004   2026-04-20  DEBIT    6327900       LDGNET    2026-05-03  OPEN
BRK-10005   2026-04-25  CREDIT   5974700       LDGNET    2026-05-11  OPEN
BRK-10006   2026-04-21  CREDIT   6990800       LDGCORE   2026-04-30  OPEN
BRK-10007   2026-06-01  DEBIT    9265800       LDGNET    2026-06-04  OPEN
BRK-10008   2026-05-29  CREDIT   9323500       LDGCORE   2026-06-18  OPEN
BRK-10009   2026-05-01  CREDIT   3247100       LDGNET    2026-05-03  OPEN
BRK-10010   2026-05-23  CREDIT   6237200       LDGNET    2026-06-03  OPEN
BRK-10011   2026-04-29  DEBIT    6148500       LDGNET    2026-05-04  OPEN
BRK-10012   2026-05-28  DEBIT    1371300       LDGNET    2026-05-29  OPEN

[2] EVIDENCE ARTEFACTS
ID          KIND             SUBJECT     LEDGER    DATED       STATUS
EVD-20001   CUTOFF-LOG       BRK-10011   LDGCORE   2026-04-22  FILED
EVD-20002   CUTOFF-LOG       BRK-10005   LDGCORE   2026-04-25  FILED
EVD-20003   CUTOFF-LOG       BRK-10005   LDGNET    2026-04-28  FILED
EVD-20004   RATE-CARD        BRK-10002   LDGNET    2026-06-12  FILED
EVD-20005   FEE-SCHEDULE     BRK-10004   LDGCORE   2026-04-24  FILED
EVD-20006   DUPLICATE-TRACE  BRK-10001   LDGNET    2026-05-25  FILED

Abridged — the file continues.

The outcomeWhat a good result looks like

One dossier in, one pack out: every break assigned to a root-cause family or carrying exactly one unassigned reason with the artefacts that decided it, then — in code, for every arm alike — family volume, amount in minor units, aging over OPEN breaks only, the corroborated share, a HIGH/MEDIUM/LOW confidence band, PROMOTED against HELD at the declared threshold of three breaks, the remediation join and the unattached items.

And when it cannot

And what it does when it cannot. On the scored run all 64 replies parsed, none stopped at the ceiling and none produced a shape defect — no family outside the vocabulary, no break invented, no break dropped, no citation to an artefact the dossier does not print. What it produces instead is a CONFIDENT WRONG READING: 52 of the 64 dossiers have something wrong in them, at a median self-reported confidence of 0.820 against 0.825 on the twelve it gets completely right. There is no signal in that number and a reader must not route on it.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your corridors' dossiers record a void, a withdrawal or a supersession as a STRUCTURED COLUMN on the artefact rather than as a sentence in a notes block. — the free columns floor, and do not buy a call at all
    22 of 64 dossiers whole for $0.00, and on the 22 where the columns decide it is 22 of 22 — 100% — where the paid arm is 5.
  • Your notes block is where the voids actually live, and a note can take an artefact out of play with a date join nobody has coded. — the call, and expect to lose more than you win
    it is the ONLY arm that scores on those 42 dossiers — 7 of 42 against the floor's 0 — but it pays by wrecking 17 of the 22 the floor gets whole.
  • You want the pack arithmetic, not the reading. — src/policy.py::rollup(), free, with any reading you like
    given the key's reading the station reproduces the key's pack on 64 of 64. None of the volume, amount, aging, corroboration, banding, promotion or remediation join needs a model.
  • You want to know whether the call can be talked into acting. — evals/injection.py, 12 live calls, about $0.006
    the cap held on 12 of 12 under three wordings, including eight trials that asked in plain language for a clearing, a posting, a named fault and a platform fix.

And where nothing here is good enough:

  • You need to know WHICH dossiers a note decides before you route. — nothing here — this is the open problem
    an oracle router over the two arms reaches 29 of 64 (45.3%), which is the best number on this page, and nobody has that oracle: knowing which dossiers a note decides IS the judgement being bought.
  • You are choosing a confidence threshold to route on. — none — do not route on this arm's confidence
    the median is 0.825 where the dossier is completely right and 0.820 where it is not. Five thousandths. There is no signal in it.

At a glanceHow the whole thing runs

19%rechecked all correct pct
2,634 msp50, end to end
$2.01per 1,000 settlement break dossiers · the cheapest published long-context card

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own quarterly dossiers in the same shape — header, DECLARED SCOPE line, break table, evidence artefacts, remediation items, notes — and rebuild the key with python3 -m tools.build_corpus. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per dossier for arithmetic you already have. That is the case against the best-fitting scenario (“Your corridors' dossiers record a void, a withdrawal or a supersession as a STRUCTURED COLUMN on the artefact rather than as a sentence in a notes block.”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A dossier with no DECLARED SCOPE line. The whole reading is 'inside the declared scope', and with nothing declared there is no in-scope ledger list, no evidence window and no promotion threshold — the pack cannot be computed at all, by any arm. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no figure on this page carries a variance. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-settlement-break. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, runs all three free floors live, re-scores every committed arm from its result file and serves all eight screenshots. What it cannot do is buy a call: the 'Ask the model' control disables itself and says why. Measured from this repository with API_KEY unset; nothing else is required beyond Python 3 — requirements.txt pins no package.

A living map of modern AI — kept current every morning