Home › Use Cases › Code why a reconciliation break exists, and name which side owns the fix
Use caseUC0354
🧪 Use-case kit · runnable

Code why a reconciliation break exists, and name which side owns the fix

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A reconciliation runs overnight and produces a queue of breaks. Finding them is the easy part and it is already done; the expensive part is the morning after, when an analyst has to say WHY each one exists under the firm's own cause code list, and two analysts working the same queue reach different answers on the same record. A break coded to a cause that closes itself is not worked by anybody -- and if it was really a failed delivery or an entitlement nobody booked, it does not close, it ages. Nothing. It sits between a reconciliation that already found the break and an ageing report that will find it again in three weeks if nobody coded it right.

Audience

The reconciliation desk at an asset manager or fund administrator, and whoever answers the quarterly question 'show me the cause-code distribution and who confirmed each one'. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reconciliation break records

The corpus is 62 reconciliation break records, 0.12 MB (txt 62). A real break record is a joined extract of a fund accounting book, a custodian position file, a depot feed and an operations queue, and all four carry an account holder and a security a real firm really holds. None of it can ship. THE FIRST DRAFT OF THIS CORPUS WAS DISCARDED before a single call was bought: every panel was a column, so a parser written straight off the card scored 62 of 62 and the paid call would have bought nothing. Seven breaks whose deciding fact is only in a custodian's SENTENCE were added, and seven BENIGN sentences in the same panels on breaks whose cause is something else, so that a sentence being present establishes nothing.

The corpus

  • The 62 reconciliation break recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.

Swap this folder for your own material and the kit is pointed at your reconciliation break records. That is the whole change — there is no database to migrate.

One reconciliation break record, as the model receives itBRK-0001.txt · 1 of 62
BREAK RECORD  BRK-0001
Reconciliation REC-2026-0908-01   As-of date 2026-09-08   Account ACC-4417
Prepared by Marchwood Asset Management, reconciliation desk
==============================================================================
1. BREAK
   Break identifier                            BRK-0001
   Reconciliation type                         position
   Security, internal identifier               SEC-10400  Northgate Utilities Ord
   Security, as the custodian names it         CX-80000  NORTHGATE UTIL ORD
   Position basis, book                        trade date
   Position basis, custodian                   trade date
   Quantity on the book                        6,500
   Quantity at the custodian                   6,380
   Quantity difference, book less custodian    120
   Cash on the book                            0.00 USD
   Cash at the custodian                       0.00 USD
   Cash difference, book less custodian        0.00 USD
   Short side                                  custodian

2. SECURITY MASTER MAPPING
   Mapping row                                 none on file for SEC-10400
   Mapping state                               unmapped
   Mapping effective from                      not applicable

3. TRADE ACTIVITY IN THE WINDOW
   no trades with a trade date on or before the as-of date

4. SETTLEMENT STATUS
   no instruction in the window

5. CORPORATE ACTIONS
   no events with an ex date on or before the as-of date

6. FEES AND ACCRUALS IN THE WINDOW
   no postings in the window

7. OPERATIONS NOTES
   - Desk read on first pass: looks like timing, expect it to clear on the next custodian file.
   - The account's previous quarter-end reconciliation closed with no items outstanding.

Abridged — the file continues.

The outcomeWhat a good result looks like

One coded row per break that an analyst can confirm or overturn in seconds, with the line of the record that establishes it quoted verbatim, the side that owns the fix named, and the pattern-review flag derived in code rather than remembered.

And when it cannot

A break whose cause is wrong is a break worked by the wrong desk, or by nobody. The named harm is a cause that will NOT close by itself coded as one that will: measured at 0 of 42 on the paid arm and 0 of 42 on the free table, so on this corpus neither arm commits it.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your break records are fully structured -- every fact in a column — the free rules table, evals/baseline.py mode rules, and buy nothing
    56 of 62 causes and 56 of 62 on all four fields, for $0.00, no key and no network. It beats the paid call on this corpus and the paired test cannot separate them.
  • Your custodian or counterparty answers queries in prose, not in fields — the paid call
    5 of 7 on the breaks whose deciding fact is only in a sentence, against the free table's 1 of 7. That is the whole case for the money and it is four breaks.
  • Your card's hardest decisions are precedence conflicts — the free rules table
    11 of 11 on the breaks where two tests both have evidence, against the paid call's 6. A table applies an order exactly; a reader argues with it.
  • You want the owning side and the pattern-review flag and nothing else — either arm plus src/recheck.py
    both are lookups on the cause. The rechecked owner is 53 of 62 on the paid arm and 56 on the free table, and the flag 59 against 60 -- the difference is just the cause accuracy carried through.
  • You want to know whether your own cause card is written down properly — evals/check_labels.py, before you buy anything
    it re-derives the whole key from a second reading of the card. If it cannot, the card is not written down yet -- and that is worth knowing for $0.00.

At a glanceHow the whole thing runs

81%rechecked break all correct pct
1,634 msp50, end to end

Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/ and data/gold.jsonl with your own coded history, and data/policy.md and data/policy.json with your own cause list, precedence order and escalation rules. The kit will not tell you whether your cause list is the right one. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call to do what a table already does. And avoid reading the table's 56 as a floor -- its phrase list comes from the card, so it is a CEILING on phrase matching, not a floor under it. That is the case against the best-fitting scenario (“Your break records are fully structured -- every fact in a column”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a break queue whose records are fully structured -- the free rules table wins there, for $0.00. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?run-to-run variance: one scored run, so the 4-against-8 split against the free table is a single observation and the paired p is 0.3877. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, as answered, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-08 — r001-break-cause. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — clone, python3 -m evals.run --run-id b000-break-cause-rules --floor rules, and the board renders with a scored free arm and no key at all.

A living map of modern AI — kept current every morning