Home › Use Cases › Assemble the evidence behind a disputed invoice line by line, and name what is missing
Use caseUC0472
🧪 Use-case kit · runnable

Assemble the evidence behind a disputed invoice line by line, and name what is missing

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A customer disputes lines on an invoice, and before anybody can rule on it somebody has to assemble the evidence FILE behind each disputed line: for a PRICE dispute the signed contract and the order; for QTY the entitlement and the usage record; for SERVICE the service record, the ticket and the SLA report; for TAX the tax determination and the exemption certificate; for DUP the two invoices and the payment; for CREDIT the credit note and the credit memo approval. Each kind is either present WITH the record that satisfies it, attached to that line, inside the declared 540-day lookback window and inside the declared systems — or missing for a stated reason. Two approval gates have to be checked, each against its own accepting role. Most of that is a column read: the records are printed on the file with STATUS columns. But on 15 of the 64 files a NOTE somewhere else takes a printed record or an approval out of play, and on 29 more a note looks exactly as if it does and does not. Opening one disputed-invoice evidence file, reading the printed evidence and approval blocks against DEA-2026's per-reason kind sets, testing every candidate record for line attachment, declared system scope and the declared lookback window, reading every note to see whether it takes a record out of play AS AT THE FILE ASSEMBLED DATE or merely mentions it, checking both approval gates against their own accepting roles, and carrying the cited records, systems and approvers out with the answer.

Audience

The collections and AR-operations desk of a software company assembling the evidence file behind a disputed invoice, and the credit and settlement desks whose approvals it cites. The decision it supports is 'is this file complete enough to hand on' — never whether the dispute is valid, never whether a credit should be issued, and never what to say to the customer. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual dispute evidence files

The corpus is 64 dispute evidence files, 0.10 MB (json 3 · jsonl 1 · md 2 · txt 64). It had to be a corpus where the printed columns are NOT enough, and the first build was not: a regex written from the generator's own note phrasing scored 60 of 64. That was fixed STRUCTURALLY rather than by rewording — the distinctive clause of a note that VOIDS a record and of a note that confirms it STANDS is now shared clause for clause, by index, between the two libraries, so the same sentence naming the same record decides on one file and decides nothing on another and only the dates separate them. The tuned regex fell to 36 of 64 and now LOSES to the honest free floor of 49. That is what a corpus looks like when the wording tell is gone.

The corpus

  • The 64 dispute evidence filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your dispute evidence files. That is the whole change — there is no database to migrate.

One dispute evidence file, as the model receives itDSP-0001.txt · 1 of 64
DISPUTED INVOICE EVIDENCE FILE  DSP-0001
FILE ASSEMBLED  2026-04-17   STANDARD  DEA-2026

ACCOUNT  ACC-4421   CONTRACT  CON-8755
INVOICE  INV-26-71430   INVOICE DATE  2026-03-19   CURRENCY  EUR
INVOICE TOTAL  22169.00   DISPUTED TOTAL  20669.00
DISPUTE RAISED  2026-03-27   COLLECTOR  COL-7341
DECLARED SCOPE  SYSTEMS BILLING, CRM   LOOKBACK 540 DAYS FROM 2024-10-03

[1] DISPUTED LINES
LINE  DESCRIPTION                     REASON   QUANTITY  UNIT PRICE  AMOUNT
L11   Unapplied service credit        CREDIT   35        501.40      17549.00
L12   Uptime commitment               SERVICE  15        208.00      3120.00

[2] EVIDENCE RECORDS
ID          LINE  KIND  SYSTEM    DATE        STATUS
EV-40002    L11   CMA   BILLING   2025-09-15  ON-FILE
EV-40001    L11   CRD   BILLING   2024-09-24  ON-FILE
EV-40005    L12   SLA   CRM       2025-07-23  ON-FILE
EV-40003    L12   SVC   BILLING   2025-06-24  ON-FILE
EV-40004    L12   TKT   CRM       2025-10-25  ON-FILE

[3] APPROVALS
ID          GATE                 APPROVER  ROLE                 DATE        STATUS
AP-50000    SETTLEMENT-POSITION  APV-2202  COLLECTIONS-MANAGER  2026-04-13  RECORDED
AP-50001    CREDIT               APV-3355  CONTROLLER           2026-04-16  RECORDED

[4] NOTES
Revenue operations desk: the 2026-04-15 review voided EV-40005 - the extract had been run against a superseded contract version - and it is not carried on this file.
Billing desk: a query was raised on 2026-04-15 about whether record EV-40001 had been pulled from the wrong billing period; it had not, and it does stand as evidence for this line.
Sales desk: can we just credit the disputed lines and close this? The renewal is next week.

The outcomeWhat a good result looks like

One dispute evidence file in, one row out: the file id, every disputed line with the evidence kinds DEA-2026 requires for its reason — each PRESENT with the record ids that satisfy it or MISSING with a reason off that kind's ladder — a band of strong, thin or no-evidence-found on every line, both approval gates, one ASSEMBLED or GAPS-REMAIN verdict about the FILE, and the three lineage lists: the records cited, the systems they came from and the approvers. 61 of 64 files come back with all seven graded fields right once the reading is rechecked, against 49 for the floor of record and 36 for a regex written from the generator's own phrasing.

And when it cannot

And what it does when it cannot. 62 of 64 replies parsed, 0 stopped at the 1,000-token ceiling and 0 were closed by the streaming runaway stop. The 2 that did not parse came back COMPLETE — finish_reason: stop, 517 and 511 output tokens — each missing ONE comma between two elements of the lines array; they are counted WRONG, stay in the denominator, are never re-fired and are deliberately NOT repaired, because inserting the comma would raise the published score on a repair the grader invented. The one file it read wrong (DSP-0016) is named in the kit README with the sentence that decided it.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your billing system records a void, a supersession or a withdrawn approval as a STATUS on the row itself, and your evidence is exported for the declared systems only — the free columns floor, and do not buy a call at all
    49 of 64 files whole for $0.00 — and on the 49 files the columns decide it is 49 of 49 where the paid arm is 47. Line attachment, declared scope, the lookback window and both approval gates are all decidable from printed rows, dates and roles.
  • Record changes arrive as free prose in a notes block — a desk writing that a record was voided, replaced, queried or confirmed to stand — the paid call, and trust it with the READING only
    that is the whole 12-file gap. On the 15 files a note decides, the columns floor scores 0 and the paid arm 14; on the 4 replacement files every free arm scores 0 and the paid arm 3. p = 0.004181 on the whole file.
  • Untrusted text can reach the notes block of a file before it is assembled — provenance on the notes, upstream of this kit
    the adversarial arm moved the bought reading on 1 of 7 trials with an ordinary-sounding revenue-operations sentence, and nothing in this kit fires on it. The cap held 21 of 21; the reading is not defended.

And where nothing here is good enough:

  • You want a single headline percentage for a card or a slide — neither number alone — publish the pair
    RECHECKED 95.3 pct and RAW 75.0 pct are answers to different questions and the gap IS the finding. A per-line band figure is worse than either: a constant reply reaches 83.2 pct of them reading nothing.

At a glanceHow the whole thing runs

95%rechecked all correct pct
2,697 msp50, end to end
$1.59per 1,000 dispute evidence files · the fast tier

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own plain-text evidence files into data/corpus/ in the same four-block layout — header with a DECLARED SCOPE line, disputed lines, evidence records, approvals, notes — add a row per file to data/cases.json, then edit data/policy.json so the dispute reasons, the required kinds, the reason ladder and the two approval gates are YOURS. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per file for arithmetic you already have. That is the case against the best-fitting scenario (“Your billing system records a void, a supersession or a withdrawn approval as a STATUS on the row itself, and your evidence is exported for the declared systems only”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A file with no DECLARED SCOPE line. Every missing-kind reason below record-voided is 'inside the declared systems and inside the declared window'; with neither declared there is nothing to test against and the kit would be inventing a scope. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether DEA-2026 resembles any real collections desk's evidence standard. It is invented for this kit — the six dispute reasons, the thirteen evidence kinds, the six-step missing-kind ladder, the two approval gates and their accepting roles, the three bands and the two verdicts are all ours, and nothing outside this repository agrees with them. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-invoice-dispute. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, runs all four free floors live in the browser and replays every committed arm. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run; without one the button is disabled and says so.

A living map of modern AI — kept current every morning