The business caseThe problem this solves
A customer disputes a bill, or a billing check flags a charge, and somebody has to write the justification for each requested credit line before an approver will look at it: which record establishes the line, how many days of the cycle it covers, whether the requested amount reconciles, and which approval tier the packet needs. The records are there, but the facts are in their prose — an outage restored "88 hours 20 minutes later", a payment received on time but short — and today an analyst reads them by eye against the drafting standard. Reading every evidence record in a credit justification packet against the drafting standard by eye to write each requested line's justification and decide whether the packet may go to an approver — not the approval, and not the analyst who owns the packet.
Audience
A billing operations analyst drafting credit justification packets, and the approver who decides whether a credit goes ahead. The answer this report gives them is a qualified yes: the paid draft beats free code only through the pure-code station that recomputes everything but two readings, and the kit says so in its first sentence. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual credit justification packets
The corpus is 64 credit justification packets, 0.32 MB (json 3 · jsonl 1 · md 2 · txt 64). Because the shape of a credit justification failure is not the arithmetic, it is WHICH RECORD establishes a line and OVER HOW MANY DAYS, and both are prose. Seven unevidenced lines cite a decoy record; 46 evidence records are decoys whose dates settle nothing — a fault opened or closed inside the cycle, a cancellation requested but effective after it, a payment posted late but received on time. Outage restorations are given as durations, across midnight and across the cycle start; payment dates as days after issue; cancellations as notice periods. 20 evidenced lines sit on packets raised by a billing check that cite nothing at all. A corpus of clean columns would measure a calculator; this one measures a reader.
The corpus
- The 64 credit justification packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 64 packets, data/case_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your credit justification packets. That is the whole change — there is no database to migrate.
CREDIT JUSTIFICATION PACKET
PACKET HEADER
Packet CJP-0001
Operator Quellmarsh Telecom
Drafting standard QMT-CJD-2026
Account QA-32137
Bill cycle 2026-02-01 through 2026-02-28 (28 days)
Candidate source Billing dispute DSP-21434, raised by the customer; the dispute's own anchor is owed upstream
Raised by R. Okafor, Billing Operations
Raised on 2026-03-10
APPROVAL SCENARIO
Second-approver threshold 75.00 dollars
Supplied by the operator running this scenario. It is a scenario value, not a stated policy, and no policy is cited for it.
CREDIT POLICY SOURCES
PS-1 Service interruption credits — formal written policy
PS-2 Late payment fee on a bill paid in full by its due date — formal written policy
PS-3 Charges billed after a cancellation took effect — formal written policy
REQUESTED CREDIT LINES
# Charge Service Billed Requested Policy Cited Description
1 TVBASE SV-46748 33.00 4.71 PS-1 E-1 TV Base package monthly charge
2 LPF BL-532744 8.00 8.00 PS-2 E-5 Late payment fee on bill BL-532744
3 FIB900 SV-26913 72.00 72.00 PS-3 E-4 Fibre 900 broadband monthly charge
Requested total 84.71
EVIDENCE RECORDS
E-1 Fault ticket FT-58388 for SV-46748: service unavailable from 11:30 on 2 February 2026 to 00:05 on 5 February 2026; ticket closed 9 February 2026.
E-2 Cancellation CN-33335 took effect on 10 March 2026 for SV-46748, having been requested on 14 February 2026.
E-3 Fault ticket FT-95629 was opened on 2026-02-01 for SV-46748. The service was lost at 2026-01-31 09:30 and restored at 2026-01-31 22:40.Abridged — the file continues.
The outcomeWhat a good result looks like
Every requested credit line carries a finding from a closed list of six, the evidence record ids that establish it, the affected days and a note of at most 160 characters; the packet carries its lane (not drafted, held, or ready for approval), its tier, the terms it is held for, any policy gaps and its source provenance. A line nobody evidenced, an amount that does not reconcile, a plan-terms reading whose anchor is owed upstream, a line already credited or a tier with no threshold HOLDS the packet.
And when it cannot
⚠︎ AS ANSWERED THE PAID DRAFT DOES NOT BEAT FREE CODE, AND IT IS PUBLISHED AS SUCH. Lines wholly correct — finding, exact evidence ids and affected days together — 69 of 121 as answered against the domain-written floor's 78 (p = 0.242960, behind, not significant). After the pure-code station it is 94 against 78 (p = 0.016589), and the same station takes a constant reply that reads nothing from 0 to 27. What the call buys is nameable: the 44 lines only the records settle, 29 after the station against the floor's 6 (p < 0.000001). The lane — whether a packet reaches an approver — is not separable from free code (p = 0.133801 and 0.263176).
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A packet raised by a billing check that cites no evidence, where the establishing record has to be found among the records — the paid draft, with the station
14 of the 20 such lines after the station against the domain floor's 3 and the columns floor's 0. This is where the money buys something no free arm reaches. - An interruption credit whose outage record gives the restoration as a duration, or crosses the start of the cycle — the free domain floor, or a person with a calendar
the paid draft gets 0 of 6 duration lines and 1 of 6 cycle-start lines after the station; the domain floor gets 2 and 3, for $0.00. - A whole-packet sign-off — every line and every packet field right — the paid draft, with the station, and a person
41 of 64 packets wholly correct after the station against the columns floor's 29 (p = 0.022656).
And where nothing here is good enough:
- Deciding which packets reach an approver — neither alone; a person on every packet the draft sends forward
after the station the paid draft holds 33 of 34 packets that must not go and sends 18 of 30 that should; the columns floor holds 27 and sends 18. On the lane as a whole the corpus cannot separate them (p = 0.133801 and 0.263176).
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own rendered packets and data/case_records.json at their structured half, then rewrite src/packet.py for your layout and data/policy.md + data/policy.json for your drafting standard. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: The draft without the station. As answered it gets 10 of those lines, and on the corpus as a whole it does not beat free code (69 against 78). That is the case against the best-fitting scenario (“A packet raised by a billing check that cites no evidence, where the establishing record has to be found among the records”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet of more than about seven requested lines, extrapolated rather than seen, because no packet here has more than 3. Replies average about 100 output tokens per line (225.5, 331.9 and 421.7 at one, two and three lines) and the longest drew 475 of 1,000 at three. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether any of this holds on a real operator's packets. Every figure here is against a key derived from a generator, on an invented drafting standard, for an operator that does not exist. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-credit-justification. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on 2026-09-16 on a copy of the kit folder with the gitignored call caches removed, no key and no .env: python3 tools/build_corpus.py --check rebuilt the corpus byte-identically, python3 -m evals.check_labels re-derived the key with 0 disagreements, and python3 -m evals.baseline scored every free arm raw and through the station — all three in about one second. python3 -m src.app --port 9583 then served the board and every read route answered 200 (the board, the packet list, one packet, the rules, the prompt, the corpus and the pressure tab), with the draft button disabled for want of a key. What the copy could NOT do is buy a call: that needs one credential, and nothing in the free path opens a socket.







