The business caseThe problem this solves
An acquirer's merchant-services back office carries every card dispute case through a lifecycle — a retrieval, a chargeback, a representment, pre-arbitration, arbitration — and every stage moves money on the merchant's account: the chargeback debit and its fee, a re-credit when a representment is won, a pre-arbitration debit, an arbitration credit or fee. Before settlement, those postings have to agree with the stage the case has actually reached and the decision on file. The settlement matcher reads COLUMNS — a reference on file, a duplicate check, an amount check. What decides the answer is often a sentence: a chargeback logged against the wrong case, a provisional decision that became final, a decision the network set aside, a journal that reversed a posting the extract still shows. The matcher's own status disagrees with the procedure on 37 of the 64 cases in this corpus. Opening one dispute case, working out which lifecycle events still stand and which network decisions count at the AS AT date, reading each note for an event logged in error, a decision confirmed or set aside, a pre-arbitration withdrawn or a journal reversal, matching every posting to the stage that earns it at the register's amount, and finding the credit the lifecycle owes that nobody posted.
Audience
The merchant-services settlement desk and the disputes desk at an acquirer, working the day's open dispute cases before the settlement run — and whoever owns the merchant's account when an exception is raised. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual dispute cases
The corpus is 64 dispute cases, 0.15 MB (txt 64). It is generated because it has to be. A real dispute book is a network's and an acquirer's confidential record: cardholders and merchants on every case, real reason codes and time limits under a network's licence, and disputes-desk notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 dispute casesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every case, merchant code, issuer code, transaction, amount, fee, lifecycle event, network decision, posting and note is generated from one seed; the network lifecycle rulebook (NLR-2026) and the reconciliation procedure (CLR-2026) are both fictional, and no real network's reason codes, fees or time limits ship. THERE ARE NO PEOPLE IN THIS CORPUS — a merchant is a code and a sector, there is no card number and no cardholder, and a note speaks for the disputes desk, the settlement desk, network liaison or merchant services. evals/check_labels.py sweeps all 64 packs for an honorific or a capitalised two-word name on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your dispute cases. That is the whole change — there is no database to migrate.
================================================================================
CHARGEBACK LIFECYCLE RECONCILIATION -- ONE DISPUTE CASE, ONE AS-AT DATE
================================================================================
FILE CLR-0001
CASE CB-560100
MERCHANT M-7100 an online homewares seller
MERCHANT ACCOUNT MA-81000
ISSUER IS-2200
TRANSACTION TX-9001000 2026-04-06 card-not-present
TRANSACTION AMOUNT 184.50
DISPUTE CATEGORY DC-1 card-absent fraud claimed (NLR-2026)
DISPUTED AMOUNT 184.50
CURRENCY USD
AS AT 2026-05-15
PACK COMPILED 2026-05-19
-- FEE SCHEDULE (merchant agreement, as the case register holds it) ------------
CHARGEBACK FEE 15.00
ARBITRATION FEE 250.00
-- DISPUTE LIFECYCLE (exported 2026-05-13 from the dispute system) -------------
EVENT DATE STAGE ACTION STATUS
EV-40000 2026-04-15 RETRIEVAL RECEIVED RECORDED
EV-40001 2026-04-24 CHARGEBACK RECEIVED RECORDED
EV-40002 2026-05-01 REPRESENTMENT FILED RECORDED
-- NETWORK DECISIONS ON FILE ---------------------------------------------------
DECISION DATE STAGE EVENT FOR NOTICE
ND-7000 2026-05-10 REPRESENTMENT EV-40002 MERCHANT PROVISIONAL
-- MERCHANT ACCOUNT POSTINGS ON THIS CASE --------------------------------------
POSTING DATE TYPE DR/CR AMOUNT REF
PST-300000 2026-04-25 CB-DEBIT DR 184.50 EV-40001
PST-300001 2026-04-25 CB-FEE DR 15.00 EV-40001
PST-300002 2026-05-11 REP-CREDIT CR 184.50 ND-7000
TOTAL DEBITS 199.50Abridged — the file continues.
The outcomeWhat a good result looks like
One case pack in, one row out: which lifecycle events stand at AS AT, which network decisions count, which postings are reversed, every posting's flag, the credit the lifecycle owes, the signed gap in cents, every exception and one of six CLR-2026 verdicts. 59 of 64 cases come back with all six graded fields right, against 43 for the best free floor, 41 for the columns read literally and 1 for the settlement matcher's own status.
And when it cannot
And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 5 cases it got wrong are named in the kit README with what it answered, and every one is a reading: 2 decisions that became final left out — both report a credit owed as RECONCILED, 54.75 and 450.53 — 1 note dated after AS AT followed, 1 journal reversal missed and 1 decoy believed. A reply that cannot be parsed is counted WRONG and stays in the denominator.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your dispute system records an event logged in error, a decision confirmed final and a journal reversal as a STATUS, and your ledger export carries every journal — the free modal floor, and do not buy a call at all
41 of 64 for $0.00. Stages posted twice, debits with no stage, fees never reached, wrong amounts, credits not posted and provisional decisions are all decidable from columns and dates, and on those 37 cases the floor gets 37. - Your notes decide cases, and you can trust that a note with a deciding word always decides — the free date-aware vocabulary floor (vocab_date)
It is 23 of 27 on the note-decided cases, exactly the paid call's figure, and it reads R-7's date for nothing. - Your notes decide cases AND carry decoys — a query that found the chargeback belongs here, a finality chased and not yet given, a journal checked and not posted — the paid call, always behind the station
27 of 28 decoy cases against vocab_date's 0 and modal's 24, and 59 whole against the best floor's 43 (17 / 1, p = 0.000145). - A credit owed to a merchant going unreported is the error you cannot afford — the paid call behind the station, with a date check on every finality note before a case is released as RECONCILED
Both of its costly misses (CLR-0040, CLR-0046) are a decision a note makes final before AS AT, left out. The vocabulary floors get CLR-0040 whole; nothing in this kit combines them yet.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own dispute case packs in the same six-block shape and data/cases.json with your own case register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per case for arithmetic you already have. That is the case against the best-fitting scenario (“Your dispute system records an event logged in error, a decision confirmed final and a journal reversal as a STATUS, and your ledger export carries every journal”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A posting with no REF to an event or a decision — a batch adjustment, a netted settlement line. The whole reduction is the posting-to-stage join; a posting with nothing to join to reads as UNSUPPORTED. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-chargeback-lifecycle. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all five free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.








