The business caseThe problem this solves
A customer pays less than the invoice and types a reason on the remittance. Somebody now has to decide whether that short payment is authorised, and it looks like a lookup and is not. The reference the customer cited lives inside their free text in whatever spelling their AP clerk used -- with the hyphen, without it, or as the four digits alone -- and on this corpus 2 of 60 notes cite nothing at all while 5 cite a code that resolves to nothing. The cited commitment may exist and belong to a different customer, or have closed before the invoice shipped, or cover a different product line. And the money has TWO ceilings rather than one: a commitment with $12,915.54 accrued can still justify only $927.89 of a claim, because at its own rate against that invoice's qualifying value that is all it could ever have earned. Get it wrong upward and the money is gone AND the accrual now shows a balance it does not have, so the next claim against that promotion is short and its measured return is wrong for the rest of its life. Working one short payment by hand: reading the remittance note for a reference, finding it in the register, checking the customer, the product line and the invoice date against the period, then computing the draw against two caps in cents. It replaces the building of the answer, not the confirming of it. Nothing here clears a deduction, posts a credit, draws down an accrual, opens a dispute or writes to any system -- there is no such endpoint and no flag that adds one.
Audience
The accounts-receivable or trade-spend analyst who works a deduction queue, and whoever answers for the residual that is or is not put back to the customer. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual deductions, each with the register put in front of it
The corpus is 60 deductions, each with the register put in front of it, 0.08 MB (json 2). The defect mix a deduction queue actually fails on, and no public dataset can provide it. 45 deductions cite a reference the way a pattern expects; the other 15 do not -- 6 write the digits alone, 5 drop the hyphen, 2 cite two references and 2 cite nothing at all. 5 cite a code that resolves to nothing in the register. Then the planted judgement cases: 4 where the cited commitment is held for ANOTHER CUSTOMER, 5 where it covers another PRODUCT LINE, 5 where it closed before the invoice shipped, 6 where the note cites a later-expiring commitment on purpose so that expiry-order drawing takes it off the wrong accrual, 8 where the claim exceeds the accrual, 6 where the RATE CEILING binds below an ample balance, 7 that span two commitments and 4 where the customer holds nothing at all. ⚠︎ AND THE MIX IS CHOSEN, WHICH IS WHY NO RATE HERE IS AN ESTIMATE OF ANYTHING IN THE WORLD. How often a real short payment is unauthorised, how often a customer miscites a promotion, and how much of a real queue is written off untouched are all unknown here and are not claimed.
The corpus
- The 60 deductions, each with the register put in front of itgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your deductions, each with the register put in front of it. That is the whole change — there is no database to migrate.
[
{
"commitment_id": "TPR-2201",
"customer_id": "C-1001",
"customer": "Brightvale Grocers",
"description": "off-invoice scan allowance",
"period_label": "Q1",
"period_start": "2026-01-01",
"period_end": "2026-03-31",
"product_scope": [
"SNK",
"BEV"
],
"rate_bps": 400,
"accrued_balance_cents": 890709
},
{
"commitment_id": "TPD-3301",
"customer_id": "C-1001",
"customer": "Brightvale Grocers",
"description": "co-op advertising allowance",
"period_label": "Q1",
"period_start": "2026-01-01",
"period_end": "2026-03-31",
"product_scope": [
"FRZ"
],
"rate_bps": 500,
"accrued_balance_cents": 640000
},
{
"commitment_id": "TPD-3302",
"customer_id": "C-1001",
"customer": "Brightvale Grocers",
"description": "prior-period display allowance",
"period_label": "Q2",
"period_start": "2026-04-01",
"period_end": "2026-06-30",
"product_scope": [
"SNK"
],
"rate_bps": 450,
"accrued_balance_cents": 585000
},
{
"commitment_id": "TPR-2202",
"customer_id": "C-1002",
"customer": "Northmarch Grocers",
"description": "display and feature allowance",
"period_label": "Q2",
"period_start": "2026-04-01",
"period_end": "2026-06-30",
"product_scope": [
"BEV",
"BEV"
],
"rate_bps": 450,
"accrued_balance_cents": 1063767
},
{
"commitment_id": "TPD-3303",
"customer_id": "C-1002",
"customer": "Northmarch Grocers",
"description": "co-op advertising allowance",
"period_label": "Q2",
"period_start": "2026-04-01",
"period_end": "2026-06-30",
"product_scope": [
"CHL"
],
"rate_bps": 500,
"accrued_balance_cents": 640017
},
{
"commitment_id": "TPD-3304",
"customer_id": "C-1002",
"customer": "Northmarch Grocers",
"description": "prior-period display allowance",Abridged — the file continues.
The outcomeWhat a good result looks like
The dispute list a person works through: for one short payment, the commitments it validly draws against, how much is covered, what is left to dispute -- and the whole draw printed beside it, every row showing what was left before it, the accrued balance, the rate ceiling for this invoice, WHICH OF THE TWO CAPS BOUND and what came out. A covered amount a person cannot check is a covered amount they have to believe.
And when it cannot
An unauthorised deduction paid. It is expensive twice -- the money is gone and the accrual it came off now shows a balance it does not have. On this corpus the null arm that approves every claim in full pays $30,136.36 it should not and approves 13 of the 14 unauthorised claims; the shipped rules floor pays $12,401.47 across 6 deductions; the scored model arm pays $0.00 and 0. ⚠︎ AND THE THIRD NUMBER IS THE ONE TO READ LAST, because AUTHOR COUNTERFACTUAL, measured 2026-08-31 on a keyless scratch copy and NOT a shipped arm of this kit: three lines added to evals/baseline.py step 2, so that a resolved reference is re-checked on customer and product scope -- the same two tests the rulebook already states and the floor's OWN fallback filter already applies -- the same $0.00 arm then pays $0.00 over the key and 0 unauthorised claims. See not_good_enough.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A remittance column where the promotion reference arrives as a structured field, or in one canonical spelling -- an EDI feed, a portal, a customer who always writes the code the same way — the free rules floor, with its step-2 shortcut removed
The reading is what the money is being asked to buy, and there is nothing to read. On this corpus a floor that re-checks customer and product scope after resolving a code gets covered_cents and residual_cents EXACT on 60 of 60 deductions and $0.00 of money wrong, for $0.00 and in 0.08 s with no network. Its remaining 5 misses are all right-money-wrong-accrual. - A free-text remittance column typed by a hundred AP clerks, where the customer's citation is a strong hint about WHICH promotion the money should come off -- and where the promotion's measured return matters as much as the amount — the model, rechecked
This is the one place the paid arm is unambiguously ahead on this corpus. It read 60 of 60 references including the bare four digits and the 2 notes citing nothing, against the floor's 49, and it is 6 of 6 on thecited_priorityfamily where the customer names a later-expiring commitment on purpose -- against 1 of 6 for the floor with or without its shortcut removed. The money is the same either way; the ACCRUAL is not, and a balance drawn off the wrong promotion is wrong for the rest of that promotion's life.
And where nothing here is good enough:
- A deduction queue where short payments net several invoices and claim types into one payment, or where the register expresses scope as customer hierarchies and item groups — neither arm, until something else has split the payment and resolved the hierarchy
This kit starts after somebody has one deduction against one invoice with a flat product code and a flat customer id. Splitting a netted payment and resolving a customer hierarchy are different problems with different failure modes, and a kit that did both would report one number for three jobs.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own short payments into data/deductions.json and your own open commitments into data/commitments.json in the shapes this kit uses -- a deduction carries the customer, the invoice and its date, the product line, the invoice total, the qualifying value, the amount deducted and the remittance note verbatim; a commitment carries the customer it is held for, its period, its product scope, its rate in basis points and its accrued balance. ⚠︎ THE REMITTANCE NOTE AND THE WHOLE REGISTER REACH YOUR CONFIGURED PROVIDER VERBATIM. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per deduction for a job a regex and a filter already do -- and inheriting a 7.6 s p50 to do it. That is the case against the best-fitting scenario (“A remittance column where the promotion reference arrives as a structured field, or in one canonical spelling -- an EDI feed, a portal, a customer who always writes the code the same way”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A REAL REMITTANCE COLUMN. The notes here are templated -- six spellings from one generator, 62 characters on average and never more than 82. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHERE THIS MODEL'S EDGE IS. Every failure bucket is empty and every graded field is 100.0 pct, so the run says what did not happen and cannot say what would. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-deduction-match. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured this session on a copy of the kit with no .env and no API key. python3 tools/build_corpus.py --check rebuilt 60 deductions, 168 commitments and the 180-check key in 0.05 s and reported data/ BYTE-IDENTICAL to a fresh build -- and it did so under PYTHONHASHSEED 0, 12345 and random, which is the two-seed test rather than the claim. python3 -m evals.check_labels re-graded that key with independent arithmetic in 0.21 s: KEY CLEAN, 60 deductions, 180 checks, 168 commitments re-derived. python3 -m evals.run --floor rules scored the free floor in 0.08 s including writing its result file, and reproduced the committed figures exactly (37 of 60 raw, 47 of 60 rechecked, $12,401.47 of money wrong). No call was made by any route.



