The business caseThe problem this solves
A telco cancels an order and the unwinding is done by four different systems: the provisioning tasks close in one, the number or circuit is released in another, the equipment comes back through a returns depot, and the recurring charges are stopped in billing. A sweep some weeks later asks one question — was this order actually unwound, and is anything still billing? The cleanup check reads the COLUMNS: a task STATUS, an assignment state, a return log, a stop date. What decides the answer is often a sentence: a circuit released after the inventory export was taken, a stop the rerating job reversed, an effective date corrected after the order was raised, a unit checked in on the sweep date itself. The order system's own cleanup status disagrees with the procedure on 40 of the 64 orders in this corpus. Opening one cancelled order, checking every provisioning task, number, circuit, device and recurring charge against the state the export prints, reading each note to see whether it releases, reverses, reopens, corrects or receives something on this order, prorating every bill line raised after the effective date by actual days, netting the standing adjustments, and routing the residual against the account's declared refund-review threshold.
Audience
The billing-operations and order-management desks running a post-cancellation residual sweep, and the provisioning, inventory and returns desks an exception is routed to. A customer-credit decision belongs to none of them and this kit never makes one. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual cancelled orders
The corpus is 64 cancelled orders, 0.18 MB (txt 64). It is generated because it has to be. A real cancelled-order book is a carrier's own operational record: subscriber numbers, circuit ids, premises, device serials, the amounts billed to named accounts, and desk notes written about named customers. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 64 cancelled ordersgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every account, order, premises, task, number, circuit, device, charge and adjustment is an invented code built from the file index, no carrier, operator, vendor or brand is named, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a customer is an account code and every note speaks for the order desk, the provisioning desk, network inventory, the returns depot, billing operations or customer care. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your cancelled orders. That is the whole change — there is no database to migrate.
================================================================================
CANCELLED ORDER RESIDUAL VERIFICATION -- ONE ORDER, ONE SWEEP DATE
================================================================================
FILE COR-0001
ACCOUNT A-40101 residential
ORDER SO-770103 MOVE
PREMISES P-20104
CANCEL REQUESTED 2026-06-03
CANCEL EFFECTIVE 2026-06-07
SWEEP AS AT 2026-07-11
PACK COMPILED 2026-07-15
RECHECK WINDOW 60 days after the effective date (declared)
REFUND REVIEW second approver above 75.00 (declared)
CURRENCY USD
-- PROVISIONING TASKS (exported 2026-07-10 from the order system) --------------
TASK TYPE SYSTEM STATUS UPDATED
PT-81201 SERVICE DISCONNECT ACTIVATOR CLOSED 2026-06-17
PT-81205 CIRCUIT TEARDOWN DESIGN CLOSED 2026-07-08
PT-81209 EQUIPMENT RETURN LOGISTICS CLOSED 2026-06-16
PT-81213 BILLING STOP BILLING CLOSED 2026-07-08
-- NUMBER AND CIRCUIT INVENTORY (exported 2026-07-10) --------------------------
RESOURCE KIND STATUS UPDATED
CKT-40011 FIBER CIRCUIT RESERVED 2026-05-17
-- CUSTOMER PREMISES EQUIPMENT (exported 2026-07-10 from the returns system) ---
SERIAL DEVICE SHIPPED RETURN STATUS LOGGED
SN-2W20103 OPTICAL TERMINAL 2026-05-19 RETURNED 2026-06-22
-- RECURRING CHARGES (exported 2026-07-10 from billing) ------------------------
CHARGE PRODUCT MONTHLY STATUS STOP DATE
CHG-30102 FIBER BROADBAND 1G 89.00 STOPPED 2026-06-08
-- BILL LINES POSTED ON THE ACCOUNT --------------------------------------------Abridged — the file continues.
The outcomeWhat a good result looks like
One order pack in, one row out: the cancellation date in force, which tasks are still open, which numbers or circuits are still held, which devices are still out, which charges are still live, which adjustments stand, the residual to the cent, every exception, one of six CRV-2026 verdicts and a review tier. 56 of 64 orders come back with all ten graded fields right, against 39 for the best free floor and 0 for the order system's own cleanup status.
And when it cannot
On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 8 orders it got wrong are named in the kit README with what it answered, and every one is a reading: 2 reversed stops it missed, 2 circuits released before the sweep it kept, 2 devices already back it kept out, 1 corrected effective date it did not take — which moved five of that order's ten fields on its own — and 1 note dated two days AFTER the sweep that it followed. Four are false residuals and four are missed residuals. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your order system records a reversed stop, a late release, a corrected effective date and a device check-in as a STATUS with a date, and your billing export lists every line raised since the cancellation — the free vocab_dated floor, and do not buy a call at all
39 of 64 orders for $0.00. The residual arithmetic, the exception set, the verdict ladder and the review tier are all free — the station runs on every arm — and on the 32 orders a column decides the floor gets 32 against the paid call's 31. - Releases, reversals, corrections and check-ins arrive as free text — an order-desk note, a number-administration note, a line pasted from a ticket — the paid call
This is the whole product. On the 32 orders where a sentence decides a reading the paid arm is 25 and the best free floor 15, and on the 40 orders the order system's own status disagrees with the procedure it is 33 against 22. - You want the residual arithmetic and nothing else — no reading at all — src/policy.py on its own, with the columns as the reading
That is the modal floor, 38 of 64 whole for $0.00, and its residual alone is right on 57 of 64.
And where nothing here is good enough:
- Your sweep turns on a boundary date — a release, a check-in or a stop dated ON the sweep date itself — neither, yet
Three of the paid arm's eight misses are exactly this shape (COR-0001, COR-0033, COR-0046) and the date_trap family is 5 for the paid call against 6 for the free floor. Nothing in this kit is good enough at the boundary to sell. - You are tempted to trust the reply's own residual and verdict and skip the station — do not
The same 64 replies score 22 whole on their own arithmetic and 56 through the station. The gap is 34 orders and the station is free.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own cancelled-order packs in the same nine-block shape and data/orders.json with your own register — the declared recheck window and refund-review threshold per order — then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal, which needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per order for arithmetic you already have. That is the case against the best-fitting scenario (“Your order system records a reversed stop, a late release, a corrected effective date and a device check-in as a STATUS with a date, and your billing export lists every line raised since the cancellation”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An order whose blocks are not printed as fixed-width tables — a PDF export, a spreadsheet, an order system that writes its history as one running log. The whole reduction is the id-to-row join and a row with no printable id joins to nothing. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-cancel-residual. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all five free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.










