Home › Use Cases › Usage-billing reconciliation
Use caseUC0290
🧪 Use-case kit · runnable

Usage-billing reconciliation

A small, forkable project that does one job end to end. Run twice for real over the same set, and every figure on these pages captured from those runs.

The business caseThe problem this solves

A SaaS vendor meters usage continuously and rates it once a month. Two records come out of that: a metered usage ledger of individual events, and a rated invoice line. NOBODY RECONCILES THEM. When somebody does, the difference runs BOTH ways — usage that was metered and never charged, and usage that was charged and the ledger does not show — and the second direction is the one that costs a customer relationship rather than a quarter. On this corpus 47 of 66 packs carry a difference: 12 under-billed and 35 over-billed, up to $690.00 on a single line. Somebody opening an account's month-end pack beside its invoice: deciding each usage row's billability from its status, its timestamp, its meter and whether its id is an individual event or an aggregated batch; summing the billable quantities in ledger units; converting at the plan's factor and rounding UP; reading the plan's one free-text rating-notes line to decide whether the prior period's unused allowance carries; subtracting the allowance once; pricing it; and only then deciding which way the difference runs and whether the pack can support saying so.

Audience

A revenue operations analyst working one account's month-end pack, and the controller who has to sign any credit that comes out of it. The kit's own answer to them is a qualified NO on the headline: a free regular expression shipped in the same repo beats the paid call on this corpus, and the page says so before it says anything else. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reconciliation pack

The corpus is 66 reconciliation pack, 0.16 MB (json 4 · jsonl 1 · md 2 · txt 66). It is generated because it has to be. A real metering ledger carries the customer's own end-users — an API key per person, a seat per named employee, a source IP per request — and the reconciliation needs none of it. Generating the corpus is what makes that separation visible rather than asserted, and it is what makes the answer key DERIVED: the generator builds each pack as a structure, renders the panels FROM it, and computes the key with the same integer arithmetic the rulebook states, so a corpus change cannot leave a stale key behind.

The corpus

  • The 66 reconciliation packgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md, including the section 'what the generator costs the measurement' where the corpus is attacked by the person who wrote it — five named ways this shape flatters a deterministic reader, and the cases the model was expected to fail, written down before a call was bought.

Swap this folder for your own material and the kit is pointed at your reconciliation pack. That is the whole change — there is no database to migrate.

One reconciliation pack, as the model receives itURP-0001.txt · 1 of 66
USAGE RECONCILIATION PACK                                            URP-0001
Prepared 2026-09-02 under UBR-2026 | billing period 2026-07-01 to 2026-07-31

PACK FACTS
  Account                        Quillon Labs
  Account code                   ACC-45076
  Metered service                seats
  Meter                          MTR-SEATS-STD
  Billing period                 2026-07-01 to 2026-07-31
  Prior period unused allowance  12 seat-month
  Pack reference                 URP-0001

PLAN TERMS
  Plan reference                 PLN-6856
  Effective                      2025-01-01 to 2026-12-31
  Ledger unit                    seat-day
  Rated unit                     seat-month
  Conversion                     30 seat-day to 1 seat-month, rounded UP
  Included units                 31 seat-month per period
  Unit price                     $    15.00 per seat-month
  Rating notes                   Any included allowance not consumed in this period is added to the following period's included units.
  Reconciliation tolerance       none - UBR-2026 sets no tolerance

METERED USAGE EVENTS
  event        timestamp                quantity  meter                status
  EVT-407547   2026-07-05T11:04Z             578  MTR-SEATS-STD        BILLABLE
  EVT-407525   2026-07-15T06:43Z             834  MTR-SEATS-STD        BILLABLE
  EVT-407643   2026-07-16T18:51Z             637  MTR-SEATS-STD        BILLABLE
  EVT-407467   2026-07-18T21:18Z             681  MTR-SEATS-STD        BILLABLE
  EVT-407686   2026-07-20T18:51Z             798  MTR-STORAGE-STD      BILLABLE
  EVT-407630   2026-07-22T01:42Z             606  MTR-SEATS-STD        BILLABLE
  EVT-407597   2026-07-23T20:41Z             183  MTR-SEATS-STD        BILLABLE

Abridged — the file continues.

The outcomeWhat a good result looks like

One pack in, six graded answers out: the reconciled quantity in whole rated units, the difference in SIGNED cents, which of nine causes explains it, which of seven findings UBR-2026's precedence order produces, what happens to the row, and one line copied verbatim from the pack establishing the cause. Every CREDIT-CANDIDATE carries the controller, stamped by the engine at every size.

And when it cannot

⚠︎ AND WHAT IT DOES WHEN IT CANNOT, WHICH ON THIS RUN WAS SIX PACKS: nothing at all. Six of 66 replies were cut off at the 32,000-token ceiling, returned no parseable answer, and were counted WRONG on all six fields rather than dropped. They stay in every denominator on this page.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your metering export is fixed-layout and one invoice line is wrong in at most one way — the free rules floor alone
    it gets 54 of 66 here for $0.00 and the paid call gets 46. On this shape the arithmetic reconstruction is the product.
  • Your plan terms carry real prose — rollover clauses, mid-period changes, carve-outs written by a lawyer — the paid call, on the clause only
    the one place the call beats the floor here is the free-text rating-notes line: 8 packs the floor gets wrong and the call gets right, including one where the floor puts a credit in front of the controller against a customer who owes the money.
  • You want a per-pack answer inside a few seconds — the free floors
    the paid call's p50 is 78.3 s and its p95 is 186.1 s, because 98.7 pct of the output tokens are provider-side reasoning; both floors are instant.

And where nothing here is good enough:

  • You need the short-ledger refusal to be reliable — neither, yet
    the call named 5 of 11 and the floor named 11 of 11 — but the floor only wins because the rule is a PREFIX TEST on an event id, which is a shape this corpus chose. On a real ledger where incompleteness is not announced in an id, neither arm has been measured.

At a glanceHow the whole thing runs

70%pack all correct pct · 2 runs, no ordering
78,341 msp50, end to end
$44.24per 1,000 reconciliation pack · Gemini 3 Flash

Run twice over the same set, for real, the last on 2026-09-02. Every figure on these pages was captured from those runs — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs in the same five-panel layout and data/accounts.json with your own per-account register rows, then run python3 -m evals.run --run-id b000-<yours>-rules --floor rules. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Believing the 81.8 pct transfers to an export this kit has never seen. That is the case against the best-fitting scenario (“Your metering export is fixed-layout and one invoice line is wrong in at most one way”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack whose panels are not the five this generator renders — the section splitter returns empty sections and the arithmetic refuses rather than guessing. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the call would answer the packs it truncated. Six were cut off in r001 and five in r002, four of them the same packs; every one was counted wrong and none was re-fired, per the estate's rule. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-usage-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9290 and scores every graded cell offline: the corpus rebuilds from its seed byte for byte, evals/check_labels.py re-derives the entire answer key independently at 0 failures, both free floors score 54 of 66 and 19 of 66, and the committed paid run replays per pack for $0.00.

A living map of modern AI — kept current every morning