The business caseThe problem this solves
A customer signs a change to a live order — another circuit, a faster one, a longer term, a different site — and it arrives as a SUPPLEMENT against an order that is already being built. Then another one arrives. By the cut-off there are three records of what this order is: what was sold, what the supplements authorise, and what provisioning actually completed. They disagree, and the disagreements are the expensive kind — a circuit built that nobody authorised, a change the customer signed and nobody worked, a charge billed twice because a completion record was keyed twice. The order desk's own status line is computed from the columns and is wrong on 30 of the 62 orders in this corpus, because the thing that decides it is a sentence in the notes. Opening one order folder, reading every supplement against the order as sold, reading every completion record against the circuit schedule, reading each note to see whether it is about this order and whether it changes anything, and deciding whether an undelivered change is late or still inside its revised due date.
Audience
The order desk and the provisioning fallout desk at a carrier, working a period's change-order exceptions — and the account team who has to call the customer when two supplements in force say different things. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual order folder
The corpus is 62 order folder, 0.14 MB (txt 62). It is generated because it has to be. A real change-order book is a carrier's own contractual record: customer names, circuit ids that map to real buildings, charges that are commercially confidential, and account notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.
The corpus
- The 62 order foldergenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every account is an invented code and trading name, every order number, circuit id, supplement, completion record and charge is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note is from the account team, the build desk or the order desk, and the only signature anywhere is a desk. evals/check_labels.py sweeps all 62 files for a person-shaped name and an honorific on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your order folder. That is the whole change — there is no database to migrate.
==============================================================================
SERVICE ORDER SUPPLEMENT RECONCILIATION -- ONE ORDER, ONE CUT-OFF
==============================================================================
FILE SUP-0001
ACCOUNT 7741 Kestrel Freight
ORDER ORD-40100
ORDER DATE 2026-05-04
CUT-OFF 2026-06-15
PERIOD 2026-05-01 to 2026-06-30
CONTRACT TERM 36 months
BUILD QUEUE metro build queue 4, contracted build window 10 banking days
TOLERANCE 5.00
CURRENCY USD
-- THE ORDER AS SOLD ---------------------------------------------------------
CIRCUIT SITE SPEED TERM MRC DUE
CKT-20400 Hollis Yard 100M 36 1,250.00 2026-06-05
CKT-20401 Rowan Depot 100M 36 1,250.00 2026-06-05
ORDER MRC AS SOLD 2,500.00
SOLD DUE DATE 2026-06-05
-- SUPPLEMENTS RAISED AGAINST THE ORDER --------------------------------------
SUPP RAISED CIRCUIT KIND CHANGE MRC DELTA REVISED DUE STATUS
SUP-3301 2026-05-15 CKT-20400 SPEED speed 100M to 500M +900.00 2026-06-07 ACCEPTED
SUP-3302 2026-05-16 CKT-20401 TERM term 36 to 60 months -150.00 2026-06-08 ACCEPTED
-- PROVISIONING COMPLETION RECORDS -------------------------------------------
REC COMPLETED REFERENCE
PRV-0181 2026-06-08 BUILD METRO BUILD QUEUE 4 ORD-40100 SPEED UPLIFT
circuit CKT-20400 +900.00 500M in service
PRV-0182 2026-06-09 BUILD METRO BUILD QUEUE 4 ORD-40100 TERM CHANGE
circuit CKT-20401 -150.00 60 month term applied
Abridged — the file continues.
The outcomeWhat a good result looks like
One order folder in, one row out: which supplements are in force at the cut-off, which completion records deliver the order with the record quoted verbatim, the authorised and provisioned monthly delta, the variance to the cent, and one of six SOR-2026 verdicts. 56 of 62 orders come back completely right, against 39 for the best free arm and 2 for the status line the order desk prints today.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed, nothing stopped at the ceiling and no call failed. The 6 orders it got wrong are named in the kit README with what it answered: 3 cited a completion record and its own re-key, 2 cited a completion dated after the cut-off, and 1 dropped a supplement from a conflicted order so SR-1 never fired. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your supplement register reliably marks a superseded or withdrawn change, and your provisioning feed de-duplicates its own re-keys — the free modal floor, and do not buy a call at all
39 of 62 orders for $0.00. Conflicts, merged builds, split builds, soak-test records, the cut-off and the revised due date are all decidable from columns and dates, and the floor gets every one of them. - Supersessions and customer withdrawals arrive as free text — an account note, an email pasted into the order, a call log — the paid call
This is the whole product. On the 23 orders where a sentence decides the reading the paid arm is 20 and every free arm is 0. - You need the conflicted orders found reliably, because that is the queue a human works — the free modal floor, and read the paid arm beside it
The modal floor got all 6 conflicts and the paid arm 5 — it dropped a supplement on one order and the conflict disappeared with it. A conflict is decidable from the (circuit, kind) columns once the in-force set is known, so free code is the right tool and the call is the risk. - You want the order desk's own status line audited — either paid or free — both beat it comprehensively
The desk's status line is right on 2 of 62 orders. It disagrees with SOR-2026 on 30, and it is published as an arm precisely so the comparison is against what is running today rather than against nothing.
And where nothing here is good enough:
- Your order book carries two orders for one account at one cut-off, or one order across two cut-offs — neither, yet
No file in this corpus does. The unit of work is one order at one cut-off and every percentage on this page is against that unit.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own order folders in the same shape and data/orders.json with your own register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per order for arithmetic you already have. That is the case against the best-fitting scenario (“Your supplement register reliably marks a superseded or withdrawn change, and your provisioning feed de-duplicates its own re-keys”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A supplement register that does not carry a KIND column. SR-1 groups the in-force supplements by (circuit, kind); without the kind there is nothing to group on and no conflict can be found. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-change-order-recon. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.







