The business caseThe problem this solves
A retailer short-pays an invoice and takes a deduction — a promotional allowance, a shortage, concealed damage or a price difference. The manufacturer's claims desk has to decide whether to accept it or send a denial letter, and a denial is only defensible if it cites the trade-terms clause and the records behind it: the promotion and its countersigned amendments, the proof of delivery, the price notice in force, the date damage was reported in writing. The desk system auto-matches some of that and is wrong on 32 of 64 files here; today an analyst reads the records and writes the letter by hand. Reading a deduction's backup against the promotion, its amendments, the proof of delivery, the price notices and the correspondence log, deciding whether to deny, and typing the letter with its clause and record ids by hand.
Audience
The deductions analyst who drafts the letter and the manager who approves anything over the operator-supplied 1,000.00 threshold before a person sends it. The answer this report gives them is that the paid call ties free code on this corpus: it is worth running only for the two readings it measurably does better, and only with the station recomputing the letter. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual deduction case files
The corpus is 64 deduction case files, 0.42 MB (json 4 · jsonl 1 · md 2 · txt 64). Because a denial letter fails on WHICH record governs, not on the arithmetic. The generator plants the traps a desk actually meets — an amendment countersigned a day late, a price notice announced before it took effect, a receiver's count against a driver's, damage phoned in before it was written, correspondence the backup cites that the log does not hold, and a sales promise that changes nothing — and derives the key from the true records. 44 of 64 files carry a record the letter must never cite.
The corpus
- The 64 deduction case filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your deduction case files. That is the whole change — there is no database to migrate.
DEDUCTION CASE FILE
CASE HEADER
Case DDC-0001
Customer Pirrington Supermarkets, account 40671
Customer reference PSM-D40084
Deduction type PROMOTIONAL ALLOWANCE
Invoice INV-3318000
Amount deducted 5,184.00
Deduction taken 2026-07-26
File opened 2026-07-31
Analyst A. Kerrigan
Send gate A person sends every letter. This file drafts and sends nothing.
DECLARED DESK STANDARDS
Manager approval threshold 1,000.00 requested in one letter
This value is supplied by the operator for this scenario. It is not a claimed policy, and
no regulation is cited for it.
CUSTOMER TRADE TERMS (CTT-2025) — values for this customer
TT-5.1 deduction window 120 calendar days from the invoice date
TT-4.2 concealed-damage window 10 calendar days from delivery
INVOICE
INV-3318000 dated 2026-04-26 purchase order PO-771000 dated 2026-04-14
Line Item Description Cases Price/case Off-invoice allowance
1 44121 Steel-cut oats 750 g x 12 960 21.60 none
2 51307 Tomato passata 690 g x 12 360 15.12 none
PROOF OF DELIVERY
BOL-551000 invoice INV-3318000 purchase order PO-771000 shipped 2026-04-26 delivered 2026-04-30 cases 1320
Receiver notation: Signed clean.
PROMOTIONS ON FILE
PC-2200 confirmed 2026-03-18 items 44121 shipment window 4 April 2026 to 6 May 2026 allowance 5.40 per case
PA-4100 amends PC-2200: the allowance becomes 5.80 per case. Sent to Pirrington Supermarkets 29 April 2026; no countersigned copy had been returned by 4 May 2026.
PRICE NOTICES ON FILE
none on file
CORRESPONDENCE LOGAbridged — the file continues.
The outcomeWhat a good result looks like
Each file gets one decision from a closed list of four — DENY-FULL, DENY-PARTIAL, VALID-NO-LETTER, HOLD-EVIDENCE-MISSING — with the CTT-2025 clause, every record id the denial rests on and none it must never cite, the amount the customer is asked to repay, the review route, and a letter body inside a 600-character budget. A denial counts only when decision, clause, evidence and amount are all right — the audit question "show me every deduction denied and the clause behind it".
And when it cannot
⚠︎ THIS KIT DOES NOT BEAT FREE CODE. After the station the paid call takes 23 of 35 cited denials at 3 of 29 false denials; the bar — the best free arm shipped, a regex tuned to this generator's phrasing — takes 25 at 7, and the floor of record, the domain rules engine, 22 at 8. Exact McNemar on the same files: against the bar p = 0.774414 (cited), 0.218750 (no false denial), 0.503445 (sent as drafted); against the floor of record p = 1.000000, 0.125000, 1.000000. None is significant — a statistical tie. As answered, before the station, it LOSES: 17 cited at 18 false denials, against the domain floor p = 0.030884 on no false denial and 0.011331 on sent as drafted, against the bar p = 0.012726 and 0.000508. The station is free code and it does the arithmetic for every arm: it lifts BOTH constants from 0 to 16 cited denials at 8 false denials, and it rescued the date arithmetic on the paid call (the late family 0 of 4 -> 4 of 4, promo-window 0 of 3 -> 3 of 3). Those lifts are free code's win, never the model's.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A desk whose deductions turn on concealed-damage report dates and price notices — this kit, with the station, beside the free floor
the two readings it measurably does better: written report date 9 of 9 against 2, price notice 5 of 12 against 3. - A desk whose deductions mostly cite correspondence and promises — the free domain floor alone
the domain floor reads correspondence on the log 62 of 64; the paid call 52, every miss a hold nobody needed. - A note on the file pressuring the analyst to deny and send — the station's letter, never the raw draft
under pressure the raw draft turned 5 declines into letters; the recomputed letter moved on 0.
And where nothing here is good enough:
- You want the letter the model writes, sent as drafted — neither — run the station
as answered the call drafts 18 false denials of 29 and loses to the domain floor on sent as drafted, p = 0.011331.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own rendered case files and write the parser in src/casefile.py for their layout, put your own procedure in data/policy.md and data/policy.json, and run evals/run.py with a new run id. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: Trusting its price-notice reading alone — 7 of 12 wrong is still most of them. That is the case against the best-fitting scenario (“A desk whose deductions turn on concealed-damage report dates and price notices”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A deduction whose backup cites correspondence the desk's log does not index by reference or date. The whole CORRESPONDENCE-NOT-ON-FILE hold rests on that lookup, and it is the reading the paid call gets wrong most (12 of 64). 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether any of this holds on a real desk's deductions. Every figure is against a key derived from a generator, on an invented procedure and invented customers; no real deduction or letter was read. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-denial-letter. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Observed on this repository with no key configured: python3 -m evals.baseline printed all five free arms at $0.00, python3 -m evals.check_labels returned 0 disagreements, python3 -m src.recheck reproduced the key's letter and the desk system's letter on 64 of 64 files each, and the board rendered with the check button disabled. No call was made.












