The business caseThe problem this solves
An enterprise customer submits a bulk service-order file against a master agreement — many sites, many circuits, a speed and a term on every row — and somebody has to say, row by row, whether it can be built and whether it is priced right before any of it enters provisioning. Seven things have to hold at once on every row: every required field has to be filled; the product code has to be one the catalogue carries; the same site must not already have been ordered for that code earlier in the same file; the site annexe has to establish what kind of building the address actually is, because the speed ceiling and the price uplift both move with it; the speed has to be one the rate card prices AND one that type of site is offered; the term the row actually commits to has to be on the product's schedule, which is not always the term the column prints; and the monthly charge has to equal the rate card taken through the term factor, the site-type uplift and the contract discount, floored exactly once, with the contract value following it. An ordering desk does that by eye, against a catalogue, an availability matrix, a rate card and a folder of survey notes, on every row of every file. Joining every order row to a product catalogue by code, to an availability matrix by site type, to a rate card by speed, and to three factors and a contract discount — then READING the survey note to decide what kind of building the address is and what term it actually commits to. It does not replace the provisioning desk: nothing here provisions, orders, activates, ceases or amends anything, and nothing waives a rule or re-prices a row.
Audience
The provisioning desk a submitted order file is routed to, and the account team that owns the customer's master agreement and has to send a bad file back. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual submitted bulk order files
The corpus is 60 submitted bulk order files, 0.22 MB (json 4 · jsonl 1 · md 2 · txt 60). Because the job needs a document that can be WRONG in seven different ways at once and a key that knows which of them reached each row first, and because the two facts the measurement turns on — what kind of building an address is, and what term it actually commits to — have to be absent from every column. A real ordering queue cannot supply that: the files are somebody's customers, the survey notes are somebody's property data, and nobody holds a per-row label saying which rule reached the row first. Generating it means the key is DERIVED from the same record the file is rendered from, the family mix is stated rather than sampled, both directions of the money are built on purpose, and the whole thing rebuilds byte-identically. What it costs is that a keyword floor written with the corpus open can solve it completely, and data/SOURCES.md measures exactly that rather than hoping nobody checks.
The corpus
- The 60 submitted bulk order filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 60 order files, data/contracts.json and the whole answer key are generated in-process by the file that sits beside them. Every customer, carrier, account, file number, site id, product code, product name, speed, price, factor, annexe entry and signature is invented, and every site is an opaque SITE-##### id with no street, town, postcode or coordinate anywhere in the kit. BULKORD-2026 is invented too and is not a telecommunications standard, a regulator's rule, an interconnection code or anybody's carrier catalogue. See data/SOURCES.md.
Swap this folder for your own material and the kit is pointed at your submitted bulk order files. That is the whole change — there is no database to migrate.
BULK SERVICE-ORDER FILE - A SUBMITTED ENTERPRISE ORDER FILE AGAINST THE CATALOGUE AND THE CONTRACT
FILE HEADER
File BO-0001
Customer Aldenmere Logistics Group
Carrier Quillhaven Telecom
Account ACC-3100
Currency USD
Order date 2026-01-24
Submitted by M. Hallberg, customer ordering desk
CONTRACT BLOCK
Master agreement MSA-4400
Agreement term 2025-01-01 to 2029-01-01
Volume discount 7.50 pct off every monthly charge
Ordering rule Clause 4: one site may be ordered once per product code in one file
CATALOGUE (the carrier's own product catalogue, in force for this file)
Code Product Terms offered (months)
BB-FTTP Business fibre broadband 12, 24, 36
ETH-DIA Dedicated internet access 12, 24, 36, 60
MPLS-VPN MPLS VPN port 24, 36, 60
SDW-EDGE SD-WAN edge service 12, 24, 36
VOI-SIP SIP trunk 12, 24, 36, 60
SPEED AVAILABILITY BY SITE TYPE (the highest speed each product is offered at, in Mbps, by site type)
Code standalone campus shared-tenant remote
BB-FTTP 1000 500 500 500
ETH-DIA 10000 10000 1000 500
MPLS-VPN 1000 1000 1000 100
SDW-EDGE 1000 1000 500 100
VOI-SIP 100 100 100 50
RATE CARD (Schedule R, the monthly charge before the term factor, the site-type uplift and the contract discount)
Code 50 Mbps 100 Mbps 500 Mbps 1000 Mbps 10000 Mbps
BB-FTTP $70.00 $110.00 $260.00 $430.00 --Abridged — the file continues.
The outcomeWhat a good result looks like
Every row of the file carries a verdict, the thing it rests on, the monthly charge at issue to the cent and signed, and one row copied verbatim as evidence — and the file carries RELEASE or RETURN with the rows named and the gap stated. On the published run the pure-code station reaches the right verdict on 233 of 240 rows and the right recommendation on 59 of 60 files.
And when it cannot
⚠︎ THE FILE-LEVEL NUMBER IS 6 OF 60 AND THAT IS NOT A TYPO. file_all_correct requires every one of the six graded fields on every row, and the arm returned a QUOTED ROW on 123 of the 193 rows that are ACCEPTED, where the answer contract says the citation must be null. Every one of those quotes is locatable in the file and every one scores zero. The raw arm also over-rejects badly — 123 acceptable rows called rejected, against 2 after the station. The station recomputes the verdict and the amount and deliberately does not repair the citation, because rewriting it would erase the evidence of the failure. Read the verdict and recommendation columns for what the product does and the all-correct column for what it costs to trust the reply as written.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Every fact the verdict turns on is in a column — the pure-column floor, evals/baseline.py --floor rules
It is free, deterministic, offline, and it is right on 219 of the 219 rows this corpus settles by column. Nothing a model adds can improve a lookup and a multiplication. - The answer is in a short survey note written in somebody's own words — the paid call, with the rulebook re-applied in code to its two readings
It reads 21 of 21 rows the columns cannot settle, against the column floor's 0, exact two-sided p < 0.000001. - The survey notes are templated, or you can write the patterns yourself — a keyword floor
On THIS corpus a keyword floor written with the key open solves it completely for $0.00. That is a fact about generated prose and it is exactly the case a real ordering queue is not — but if your survey notes come from a form with fixed phrasings, it is your case too. - Your address database already records what kind of building each site is — the pure-column floor, with the site type joined in structurally
The site-type reading is 15 of the 21 rows this kit pays a model for. Supply it as a column and the only reading left is the term, which is 6 rows — and at that point the honest answer is that the model has almost nothing to buy.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own order files into data/corpus/ in the same panel layout and the whole free half runs on them: the parser, the engine, all four floors, the refusal reader and the board. The key is the boundary. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not call a model for a set membership test and one multiplication. That is the case against the best-fitting scenario (“Every fact the verdict turns on is in a column”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An order file whose rows do not match the fixed column layout. src/rules.py's row expression is anchored on it and a row that does not match is silently not a row — the parser reads a file with fewer rows rather than raising. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the margin over the de-memorised keyword floor is real. It is five rows, exact two-sided p = 0.36, and this kit calls that NOT SIGNIFICANT rather than a win. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-11 — r001-bulk-order-check. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board and scores every free floor offline: the corpus, the carrier record, the key, the four floors' committed results and the paid run's committed result all ship, and the board replays them. Nothing is installed — the kit is standard library end to end. The first thing that needs a credential is the one button that calls a model, and it is disabled with the reason printed beside it.





