The business caseThe problem this solves
A utility runs a meter-to-account sweep and the join comes back with unmatched records on both sides: meters on the asset register with no billed service point, and active accounts with no meter. Somebody has to say what each one IS before anybody can act, and the honest answer is rarely on the record itself — it is in the work-order history and the desk notes, where a reference to the point is a serial 22 times in 200 and something else the other 178: the last four digits, the site address written another way, the position on the route, or the month a set was done. Working a quarterly meter-to-account sweep by eye. For every unmatched service point that means deciding whether the register side and the billing side are the same asset written two ways, two genuinely different assets, or a point with no counterpart at all — and then what the unmatched record actually IS: revenue never billed, a legitimately unmetered service, a unit sitting in stores, a removal executed but never closed, or a company-use point. The join keys and every date, band and count are free code; the disposition turns on the work orders and desk notes, which no join reads.
Audience
A meter-data desk running an estate sweep, and the revenue-protection analyst who picks up what it flags. The answer this report gives them is NO on the headline: a domain-written free arm reads more of this corpus correctly than the paid call does, for $0.00, and the paid call breaks ten lines that a rule reading no prose at all already had. What it gives them instead is the shape of the problem and a free floor worth shipping. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual meter-to-account sweep lines
The corpus is 64 meter-to-account sweep lines, 0.17 MB (json 4 · jsonl 1 · md 2 · txt 64). Because the shape of a meter-orphan failure is not the join, it is WHAT the unmatched record turns out to be — and that is prose. 178 of the 200 work orders and desk notes refer to the point by something other than its serial; 80 of them name the unit beside it or the point next door; 53 name this line's own record from outside the sweep window; and the printed register-status column answers the question WRONGLY on 33 of the 64 lines, which is what stops a column-reading rule from being the whole answer.
The corpus
- The 64 meter-to-account sweep linesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 64 sweep lines, data/point_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your meter-to-account sweep lines. That is the whole change — there is no database to migrate.
CALDERLEA ENERGY NETWORKS — QUARTERLY METER-TO-ACCOUNT ESTATE SWEEP
Sweep line SL-01 · MSP-RECON-2026 · unit: one service point
SWEEP LINE
Raised from the billing extract
Record under review SP-70000
OPERATOR'S OWN SWEEP SCHEDULE — this operator's text, not a regulator's
Sweep window 2026-01-01 to 2026-03-31 (operator-supplied)
Ageing window 90 days. It is an operator-supplied value, not a stated policy, and no
policy is cited for it.
Match tolerance references are compared character for character after case,
punctuation and spacing are removed. A near-miss reference is never a match.
Write-off threshold none is supplied for this sweep and none is assumed.
PREMISE UNDER REVIEW
Premise P-40000
Route R-13
Address as billing spells it 5 Ashgrove Terrace
Address as the register spells it 5 Ashgrove Ter
OTHER POINTS ON THIS ROUTE — printed for reference only
premise route address as billing spells it
P-40001 R-13 7 Ashgrove Terrace
ASSET REGISTER EXTRACT
serial status install removal route address as the register spells it
B4-218752 in stores 2019-03-30 2025-10-07 R-13 7 Ashgrove Ter
CIS / BILLING EXTRACT
account service pt meter ref rate status last billed address as billing spells it
A-5000001 SP-70000 K7-132154 RES-01 active 2026-03-01 5 Ashgrove Terrace
WORK ORDERS AND DESK NOTES RAISED ON THIS ROUTE
N-400001 4 November 2025 desk note the point at number 5 on route R-13 is fed from the unmetered schedule and no meter is set at that classAbridged — the file continues.
The outcomeWhat a good result looks like
Every sweep line carries a join from a closed list of three, a disposition from a closed list of seven, the work-order and desk-note ids the disposition rests on, and then — derived in code, identically for every arm — the owning desk, the next step, the ageing band and the age in days. Nothing is created, closed, relinked or billed.
And when it cannot
⛔ AND WHEN IT CANNOT, IT IS PUBLISHED AS SUCH. 27 of 64 lines are wrong: 11 on the join (9 of them a SAME-ASSET pair it failed to join, 2 the wrong side), 15 on the disposition (8 true orphans read as something benign, 3 benign records read as orphans) and 1 on the evidence alone. A disposition off the closed list would become READING-INCOMPLETE and never a guess; that never fired — the arm stayed inside the vocabulary on 64 of 64.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Working a meter-to-account sweep where the register and billing extracts print keys, statuses and rates and you want the unmatched records dispositioned — the free domain floor
42 of 64 sweep lines wholly correct for $0.00, against the paid reading's 37 — exact McNemar 11/16, p = 0.442068 — and it reproduces to the digit on every re-run. - You only care about the lines a rule reading no prose can already settle — the free columns floor
27 of 27 on that slice, against the paid reading's 17. The money buys nothing here and costs ten lines. - You want to know whether reading the work orders and desk notes is worth anything at all on this shape of problem — the paid reading, and read the narrow claim
on the 37 lines only the notes settle it takes 20 againstcolumns' and the constant's 0, p = 0.000002. That is a real finding about the task and it is not a win over the rules: against the floor of record on the same slice it is 20 against 22, p = 0.814529. - Finding the true orphans — revenue that was never billed — the free strict domain floor
12 of the 16 true orphans against the paid reading's 8 and the floor of record's 8. Orphan recall has its OWN bar and it is not the headline's arm.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | replace data/corpus/*.txt and data/gold.jsonl with your own sweep lines and key; the rulebook is data/policy.json and the engine reads it as data, so the join order, the seven dispositions, the owner and next-step tables and the cap all change there and nowhere else. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: Quoting its 65.6% as an accuracy: 27 of the 64 lines are settled by the printed columns alone, so a percentage over all 64 is partly a measurement of the corpus mix. That is the case against the best-fitting scenario (“Working a meter-to-account sweep where the register and billing extracts print keys, statuses and rates and you want the unmatched records dispositioned”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A work order or desk note that refers to the point in a form this corpus does not contain. Eight reference forms are generated — the full serial, the last four digits, the site address, the position on the route, the month a set was done, the service-point id, the billing address and the number-plus-route — and 178 of the 200 references are not a full serial. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | SIX OF THE ELEVEN CELLS THIS KIT MEASURES MAY NEVER CARRY A RESULT ALONE, and evals/scoring.py computes which rather than remembering it. owner, next_step, band and age_days are STATION-DERIVED for every arm alike — a lookup keyed on the disposition and date arithmetic — so no arm can win or lose them and they are published only as composition. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-18 — r002-meter-orphan. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on 2026-09-18 on a copy of the kit folder with the answer caches and any .env removed and no key in the environment: python3 tools/build_corpus.py --check rebuilt all 64 sweep lines, the structured records and the key BYTE-IDENTICALLY; python3 -m evals.check_labels re-derived the whole key from the print with 0 disagreements; all five free arms scored offline at $0.00; and the board served all nine of its routes, byte-identical to the board with the caches present. What could NOT be reproduced without the caches is the pressure probe's clean column: with them it re-scores the paid run's own un-injected replies at $0.00, and without them it falls back to the free tuned arm, where an improvement against the key cannot occur by construction. That is why the caches ship.












