The business caseThe problem this solves
A customer writes something about their bill, and somebody has to work out what it is about before anything else can happen. The words are rarely the answer on their own: this corpus's customers name the cause outright on 22 of 60 disputes, describe it without ever using a billing word on 26, name only the symptom on 5, and name the WRONG cause on 7. Nor is the bill the answer on its own -- a promotion that ended leaves no line on the invoice being disputed, and a charge billed twice across a cycle boundary leaves no repeat on either invoice. Today somebody opens the account, reads two bills side by side, decides which of six things this is, and pastes the deciding lines into the ticket. The read-the-dispute-beside-the-account step of billing dispute intake: deciding which of six causes a free-text complaint is about and pulling the specific invoice lines that bear on it. It does not decide the outcome -- the analyst still decides whether a credit is owed.
Audience
Whoever is deciding whether to put a model in front of a billing-dispute queue. The answer this kit gives them is a SPLIT: the paid call beats this kit's best free code on the complete answer (58 of 60 against 51, exact two-sided p = 0.039062) and does NOT beat it on the label alone (58 against 52, p = 0.070312). Anybody who reads only the first of those concludes the money is worth spending; anybody who reads only the second concludes it is not. Both are printed. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual disputes, each with two invoices from its own account put in front of it
The corpus is 60 disputes, each with two invoices from its own account put in front of it, 0.19 MB (json 4 · jsonl 1). Because the alternative is nothing. A billing dispute is a customer's own words about a private account, held beside that account's invoice history -- both halves are personal data about a named person and a commercial relationship, and neither exists in public in a form anybody may publish. There is no public corpus of billing disputes paired with the invoices that decide them and there will not be one. A kit whose corpus cannot ship is a kit nobody can re-run, and a kit nobody can re-run cannot be checked -- so the corpus is synthetic, the generator ships beside it, and the seed is in the dataset version. ⚑ WHAT IT IS HONESTLY GOOD FOR is comparing arms on identical input, which is what it is used for here: the same 60 disputes, the same key, the same scorer, and a paired test between the paid call and the free code.
The corpus
- The 60 disputes, each with two invoices from its own account put in front of itgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your disputes, each with two invoices from its own account put in front of it. That is the whole change — there is no database to migrate.
[
{
"account_id": "ACCT-41000",
"customer": "Ingrid Prendergast",
"segment": "consumer",
"market": "Drumlin Bay",
"plan_code": "MOB-PLS",
"plan_name": "Mobile Plus",
"invoices": [
{
"invoice_id": "INV-41000-202511",
"period": "2025-11",
"issued": "2025-11-02",
"due": "2025-12-24",
"lines": [
{
"line_no": 1,
"kind": "plan",
"code": "MOB-PLS",
"description": "Mobile Plus -- monthly service",
"service_period": "2025-11",
"amount_cents": 5500
},
{
"line_no": 2,
"kind": "fee",
"code": "FEE-REG",
"description": "Regulatory cost recovery",
"service_period": "2025-11",
"amount_cents": 346
},
{
"line_no": 3,
"kind": "promo",
"code": "PRM-BUNDLE",
"description": "Bundle saving -- month 5",
"service_period": "2025-11",
"amount_cents": -2100
}
],
"total_cents": 3746
},
{
"invoice_id": "INV-41000-202512",
"period": "2025-12",
"issued": "2025-12-02",
"due": "2026-01-24",
"lines": [
{
"line_no": 1,
"kind": "plan",
"code": "MOB-PLS",
"description": "Mobile Plus -- monthly service",
"service_period": "2025-12",
"amount_cents": 5500
},
{
"line_no": 2,
"kind": "fee",
"code": "FEE-REG",
"description": "Regulatory cost recovery",
"service_period": "2025-12",
"amount_cents": 346
},
{
"line_no": 3,
"kind": "promo",
"code": "PRM-BUNDLE",
"description": "Bundle saving -- month 6",
"service_period": "2025-12",
"amount_cents": -2100
}
],
"total_cents": 3746
},
{
"invoice_id": "INV-41000-202601",
"period": "2026-01",Abridged — the file continues.
The outcomeWhat a good result looks like
A coded ticket per dispute with its evidence already open: one of six causes, the exact invoice lines that decide it, and the amount at issue to the cent. On run r001-dispute-triage that was right on all three fields for 58 of 60 disputes, and the line references were right on 82 of 86 individual lines.
And when it cannot
It codes the dispute to the wrong cause and the derivation station prices the wrong cause's rule against the wrong lines -- confidently, to the cent. That happened 2 times on 60 disputes. On DSP-0051 the arm answered promotion_expiry while a live -$24.00 credit was still sitting on the disputed bill; the key says rate_applied at $2.59, and the station published $24.00. It flagged its own contradiction and returned the figure anyway, because the figure is what the pipeline would have published and refusing would have hidden the failure.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A queue whose customers mostly name the cause -- a web form with a drop-down, a chat flow that asks 'what is this about?' first — the free keyword floor
It scores 22 of 22 on the plain wordings here, which is every one of them, for $0.00 and no network. If your intake already funnels customers into naming the thing, the reading is not the hard part and there is nothing to buy. - A queue of free text where customers describe the problem in their own words and the bill shows more than one change — the model
This is the only population on this corpus where the margin is real: 25 of 26 against the delta floor's 19, 6 discordant to 0, exact two-sided p = 0.031250. Describing a charge without naming it is what language is, and a diff cannot read it. - A queue that only needs the money -- which charge, how much — the free delta floor, then the model on what it cannot decide
Diffing two structured invoices is exact, free and instant. The floor misstates $192.71 of the $1834.01 at issue across the corpus against the model's $42.29 -- a real difference, and one you can have for $0.00 on the 51 of 60 disputes where the bill decides on its own.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/accounts.json and data/disputes.json with your own rows in the same shape -- an account with a list of invoices, each invoice a list of charge lines carrying kind, code, description, service_period and amount_cents as an INTEGER -- and data/gold.jsonl with one key row per dispute. ⚠︎ WHAT STOPS BEING TRUE the moment you swap the corpus: every accuracy figure on this page, and the comparison between the arms. Corpus lens → |
| When is this the wrong choice? | Avoid: The model -- you would be paying for a margin that does not exist on that population: 22 of 22 against 22 of 22, identical. That is the case against the best-fitting scenario (“A queue whose customers mostly name the cause -- a web form with a drop-down, a chat flow that asks 'what is this about?' first”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A dispute about something that stopped or started more than one billing cycle ago. The window is TWO invoices and that is the hard ceiling -- a promotion that ended three cycles back leaves no trace in anything any arm is shown, model or floor. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether either arm survives a real dispute queue. Every sentence in the corpus came out of four template sets in tools/build_corpus.py; a real customer misspells their account number, attaches a photograph, writes in two languages, or argues about something that happened in March. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-dispute-triage. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Observed on the committed kit this session, with the model key blanked for the process: API_KEY= python3 -m src.app starts the board on port 9376 and every panel renders -- the dispute, both invoices, the answer key, all three free floors, the derivation station and the recorded run's replay -- with the model button genuinely disabled. The whole free half runs with no network at all: the corpus generator in 0.04 s, the key gate in 0.03 s (0 problems over 393 checks), and the delta floor's full 60-dispute scored run in 0.07 s. requirements.txt installs nothing; there is no pip step to fail.





