The business caseThe problem this solves
An overnight validation, estimation and editing run raises more meter exceptions than a meter-data desk can read before the day's settlement work starts. Each one is nine panels: the header, the service point, the asset-register extract, the interval arithmetic the run performed, the head-end signals it recorded, the signals it checked and EXCLUDED, the prior closed exceptions on sibling meters, the night desk's own free-text note, and what is attached. Somewhere in those panels is a cause - or there is not one, and saying so is the answer. the meter-data analyst's read of nine panels per exception to decide the cause, the repeat flag and the specialist lane - on this corpus, for the 16 of 64 exceptions whose cause is only in the night desk's sentence, and NOT for the rest.
Audience
A meter-data manager deciding whether to put a model in front of the overnight exception queue. The answer this report gives is a qualified one and the qualification is the point. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual VEE exceptions
The corpus is 64 VEE exceptions, 0.16 MB (json 4 · jsonl 1 · md 2 · txt 64). A real VEE exception cannot be published: it carries a distributor, a customer's service point, a meter identifier and a consumption history. So this one is generated - and generated to DEFEAT a word list rather than flatter one. Every distinctive clause is shared across exceptions that decide differently (183 shared occurrences); 43 exceptions carry an incomplete interval day and 21 a day outside the declared band, and NEITHER is a cause; all 64 print prior exceptions on SIBLING meters, so a repeat flag can only be got by counting the cited meter's own rows; and 6 carry a register row that almost matches.
The corpus
- The 64 VEE exceptionsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your VEE exceptions. That is the whole change — there is no database to migrate.
VEE EXCEPTION VEE-0001
================================================================================================
PANEL 1 — EXCEPTION HEADER
Distributor Marrowdale Energy Networks (synthetic distributor)
Service area Area 2 — northern feeders
Circuit CIR-3300 Halloway spur
Interval day 2026-06-11 (2026-Q2, 1 April to 30 June 2026)
Raised by the overnight VEE run VR-2026-0612
Worked by Revenue protection analyst
Exception opened 2026-06-12 05:40
PANEL 2 — METER AS CITED ON THE EXCEPTION
Meter as written MT-41820-001
Premises use as written domestic, terrace
PANEL 3 — ASSET REGISTER EXTRACT (this distributor's own register)
Meter Service point Form Multiplier Digits Daily band kWh
MT-41820-001 SP-1000000 single-phase whole current 1 5 8 to 42
MT-41820-003 SP-1000007 three-phase whole current 20 6 200 to 800
MT-41820-005 SP-1000014 three-phase transformer fed 40 6 480 to 1480
Circuit default multiplier 20 (for CIR-3300)
PANEL 4 — INTERVAL DAY AS THE RUN SAW IT
Intervals received 96 of 96
Register read at start 91491
Register read at end 91528
Multiplier applied by the run 1
Day total as computed by the run 37 kWh
Prior seven-day average 30 kWh
PANEL 5 — SIGNALS RECORDED BY THE HEAD END
SG-10 the collector recorded no contact with this meter for part of the day
PANEL 6 — SIGNALS CHECKED AND EXCLUDED AT THE RUNAbridged — the file continues.
The outcomeWhat a good result looks like
One row per exception a specialist could act on as written: the meter as the asset register writes it, the cause under card EC-2026 taken in the card's own order, the reason it cannot be placed when it cannot, whether that meter is a repeat offender under the declared threshold of 3, the specialist lane under ER-2026, and the one line of the exception that establishes the cause.
And when it cannot
It leaves the exception UNCLASSIFIED with its reason and sends it to the senior lane. The key rules 16 of 64 exceptions unplaceable - 6 whose identifier is not a register row, 5 with no interval data, 5 where nothing states a cause at all. Leaving one unclassified is the right answer, never a failure to answer.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- the causes and their ORDER are a written card and you have the card — the cards floor - pure code, evals/baseline.py
on this corpus it takes 50 of 64 complete rows for $0.00 and the paid call takes 52, a difference of 2 that McNemar cannot separate (p = 0.8145). Writing the card down as code is most of the job. - the cause is stated only in a night desk analyst's own sentence and no word list reaches it — the paid call
this is the one family where it earns its bill: prose goes 2 of 16 free to 12 of 16 paid, and it is the only family with a positive delta. - you need the exception LEFT unclassified rather than bucketed — pure code, and a person
the cards floor forced 0 exceptions into a cause; the paid call forced 3, and all 3 are the CAUSE-UNSTATED shape where the guard is a judgement rather than a join. - you want one number for a slide — the complete triaged row, with its p-value beside it
it is the only metric here a constant does NOT nearly win: the constant takes 20 of 64. Every per-field figure on this kit is degenerate to some degree. - you are comparing a paid arm against a free one on any kit with a pure-code settlement step — both columns, rechecked against rechecked
the station moves the constant from 9 to 20 and the signal arm from 30 to 36 on this kit. A margin taken from a rechecked paid arm against a raw free arm would be mostly station.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own exceptions in the same nine-panel shape, data/register.json with your own asset register, and data/policy.md with your own cause and routing cards; then rewrite src/taxonomy.py's cause list and src/policy.py's lanes to match. THE MEASURED RESULT DOES NOT TRAVEL WITH THE DATA. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call before the floor exists. Without it every paid figure reads as good. That is the case against the best-fitting scenario (“the causes and their ORDER are a written card and you have the card”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A METER IDENTIFIER THAT IS NOT A ROW OF THE ASSET REGISTER. The match is exact by design, so an identifier typed with a space, a hyphen or a trailing revision suffix separates the exception rather than resolving it - 6 of the 64 are built that way, and that is the intended behaviour, not a tolerance to be widened. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE 2-EXCEPTION MARGIN IS REAL. 52 against 50 is NOT statistically significant and McNemar exact two-sided is p = 0.8145. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, off-peak tariff. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 2 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-14 — r001-vee-triage. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — a clean checkout with NO key configured renders the whole board on 127.0.0.1:9474, replays the committed run out of results/, re-derives all four free arms and re-runs the constant sweep - every number on this page, for $0.00. requirements.txt names no package.







