Home › Use Cases › Estimated meter read true-up and estimation drift
Use caseUC0385
🧪 Use-case kit · runnable

Estimated meter read true-up and estimation drift

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A meter is not read every cycle. When it is not, the utility bills an ESTIMATE, and when an actual read finally arrives the two have to be reconciled: what did the meter really record over the gap, how does that differ from what was billed, in kWh and in money under a tiered seasonal tariff, and is the difference big enough to rebill. That is one question. The second is harder and nobody does it: across this meter's recent history, is the ESTIMATION RULE ITSELF wrong -- persistently high, persistently low, merely seasonal, or explained by something that actually changed at the premise. A single unlucky estimate is noise. A run of them in one direction is a defective rule, and it keeps producing rebills until somebody notices. The line-by-line true-up of an estimated read against the tariff -- registers, rollover, day-weighted allocation, tier ladder, season, threshold -- that a billing analyst does by hand or, far more often, does not do at all; and the cross-cycle drift review that essentially nobody does, because it means reading a meter's whole history alongside its field notes.

Audience

The billing analyst deciding whether an estimated-read true-up is worth rebilling, and the metering supervisor deciding whether a meter's estimation rule should be re-set. Neither decision is 'how much' -- it is 'does this need a human'. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual residential account packs

The corpus is 64 residential account packs, 0.22 MB (txt 64). There is no public corpus of estimated meter reads and there could not be one: a read history is a record of what happened inside somebody's home, cycle by cycle, against a tariff a utility files with a regulator. Every byte here is generated, and generating it is what makes the second question answerable at all -- the drift call needs a KNOWN truth about whether a note records something that actually happened, and no real archive carries that label.

The corpus

  • The 64 residential account packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your residential account packs. That is the whole change — there is no database to migrate.

One residential account pack, as the model receives itCHMU-ER-0001.txt · 1 of 64
CEDAR HOLLOW MUNICIPAL UTILITY
ACCOUNT FILE -- ESTIMATED READ TRUE-UP PACK
File CHMU-ER-0001
============================================================================

SECTION 1 -- ACCOUNT AND SERVICE POINT
  Account                                      6363-18263
  Premise                                      1514 Cotton Mill St, Cedar Hollow
  Rate class                                   Residential, single-phase
  Meter                                        MT-73681
  Register                                     5 digits, rolls at 99999
  Trailing 12-month average cycle consumption  720 kWh

SECTION 2 -- TARIFF IN FORCE (Schedule R-2026)
  Customer charge, per cycle                   $11.95
  Tier 1, first 600 kWh                        10.4300 c/kWh
  Tier 2, next 600 kWh                         12.8700 c/kWh
  Tier 3, above 1200 kWh                       15.9400 c/kWh
  Summer adder, Tier 3 only (Jun-Sep)         +2.1500 c/kWh
  A cycle is a summer cycle when the midpoint of its period falls in June, July, August
  or September. The customer charge is per cycle and no true-up changes it.

SECTION 3 -- ESTIMATION RULE ON FILE
  Rule          EMP-2026 Table B, degree-day weighted against a 12-month base
  Set on        2024-02-28
  Last reviewed 2024-02-28

SECTION 4 -- CLOSED TRUE-UP HISTORY (oldest first)
  id      cycle    period                     estimated   actual   signed   signed
                                                    kWh      kWh      kWh      pct
  TU-01   2025-05  2025-04-03..2025-05-02        580      617      +37     +6.0
  TU-02   2025-05  2025-05-03..2025-05-31        593      624      +31     +5.0
  TU-03   2025-06  2025-06-01..2025-06-28        890      937      +47     +5.0

Abridged — the file continues.

The outcomeWhat a good result looks like

One row per account: the adjustment in kWh and in cents, whether the estimate ran high or low, whether the account is a REBILL CANDIDATE QUEUED FOR A BILLING ANALYST or within tolerance, and separately a drift verdict off a seven-rung ladder with the printed lines that support it quoted verbatim.

And when it cannot

An amount stated as owed or owing. Nothing in this kit may say a customer owes or is owed money -- there is no field for it, it is the second of four refusals stated before anything else in the prompt, and every arm's prose is read afterwards for the vocabulary of a bill. 0 breaches across the scored run and the adversarial probe.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the estimated read trued up and the rebill threshold walked — free code -- evals/baseline.py --floor strict
    it is exactly right on all four fields for all 64 accounts against the paid call's 0, it costs $0.00, it needs no network, and it re-runs in under a second.
  • You want the drift call and you will accept a statistical answer — free code -- the same floor
    53 of 64 against the paid call's 12, p = 5.889e-11, for nothing.
  • You want to know WHY a meter's estimates drifted, and a documented cause named — the paid call, with the arithmetic taken away from it
    on the 11 accounts whose verdict turns on a sentence it beats every free floor 8-0 (p = 7.812e-03) and the floors score 0 of 11 by construction. Run it BEHIND the free floor, on the accounts the floor already calls persistent, and let src/recheck.py keep only its three judgement fields.

At a glanceHow the whole thing runs

0%trueup all correct pct
1,765 msp50, end to end
$0.00per 1,000 residential account packs · openai/gpt-5-6-terra

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs in the same eight-section layout -- the headings in src/pack.py::HEADINGS are exact-match anchors and the three row shapes are regexes -- and rewrite data/gold.jsonl with your own key. The measured result does not travel with them. Corpus lens →
When is this the wrong choice?Avoid: Paying for arithmetic. The call's own adjustment amount is right 0 times in 64. That is the case against the best-fitting scenario (“You want the estimated read trued up and the rebill threshold walked”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A meter read every cycle. There is no estimated history, so the drift window is empty and the second question does not arise. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?REPEATABILITY. The scored arm ran ONCE, and the adversarial probe once. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN), one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-estimated-read. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board -- all 64 packs, the tariff, the derived key and every committed arm -- and runs all four free floors, the stub control, the dry run and evals/check_labels.py on it. Zero calls, zero network, zero dollars. Only ASK THE MODEL and evals/run.py's paid arm need a credential, and both say so rather than failing at the HTTP layer.

A living map of modern AI — kept current every morning