Home › Use Cases › Which lines moved against their own history past a written threshold, and why
Use caseUC0323
🧪 Use-case kit · runnable

Which lines moved against their own history past a written threshold, and why

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A subscriber's bill moving is not a defect. A subscriber's bill moving in a way neither they nor the carrier can account for is what becomes a call, a complaint and sometimes a churn — and the only cheap place to catch it is the pre-send window, before the cycle goes out. The control every carrier runs on itself is a threshold policy somebody wrote down; the work is applying it per line, against that line's own history, and separating the movement nobody expected from the movement the subscriber themselves asked for. The manual pass over a pre-send exception export — reading each moved line against the account's change log, the promotional end dates and the last three cycles, then marking up a review list.

Audience

Billing Operations and revenue assurance in a carrier, working the pre-send review queue on cycle close. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual surveillance cards

The corpus is 64 surveillance cards, 0.15 MB (txt 64). A real pre-send surveillance extract cannot be published and should not be: it is a subscriber's own charge history, cycle by cycle, and the whole point of this kit is that the question is answerable without any of it. So the corpus is built to make that separation visible — a card is a record about charges and dates, and evals/check_labels.py sweeps every one of them for personal-data vocabulary and reports 0 hits. It is also built so the ARITHMETIC is doable by a regular expression and the READING is not, which is the only way a kit can say something honest about whether a model is worth buying here.

The corpus

  • The 64 surveillance cardsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromGENERATED FROM A FIXED SEED — there is no source. All 64 cards are produced by tools/build_corpus.py from a linear congruential generator seeded 20260908, plus the invented market, plan and note lists in that file. Nothing was fetched, scraped, licensed, extracted or anonymised, so there is no source URL to map to, no publisher whose terms could be checked, and no third party with a claim on any byte of it. python3 tools/build_corpus.py --check rebuilds every card, the register, the answer key and the corpus statistics in memory and diffs them against what is committed, exiting non-zero on the first byte that differs — that is the whole provenance claim and it is executable. Run under two different PYTHONHASHSEEDs it is also the determinism claim.

Swap this folder for your own material and the kit is pointed at your surveillance cards. That is the whole change — there is no database to migrate.

One surveillance card, as the model receives itBSK-0001.txt · 1 of 64
SUBSCRIBER BILL-SHOCK SURVEILLANCE CARD                              BSK-0001
Prepared 2026-09-08 under BSS-2026 | cycle 202608, open 2026-08-01 to 2026-08-31

ACCOUNT FACTS
  Account reference              BSK-0001
  Market                         Rothbeck
  Plan on file at cycle open     Connect Plus
  Cycle under surveillance       202608
  Cycle opened                   2026-08-01
  Completed cycles on card       5
  Carried surveillance state     clear
  Tax and regulatory rate        9 pct of the subtotal

CHARGE HISTORY BY COMPONENT
  code    component                                   202603      202604      202605      202606      202607      202608
  RECUR   Recurring plan charge                        31.00       31.00       31.00       31.00       31.00       31.00
  DISC    Promotional discount                          0.00        0.00        0.00        0.00        0.00        0.00
  DEVIC   Equipment instalment                         19.00       19.00       19.00       19.00       19.00       19.00
  DATA    Domestic data overage                         4.00        0.00        0.00        1.00        3.00        0.00
  VOICE   Domestic voice and messaging overage          0.00        0.00        0.00        0.00        0.00        0.00
  ROAM    International roaming usage                   0.00      124.00        0.00        0.00        0.00        0.00
  ONEOF   One-time charges                              0.00        0.00        0.00        0.00        0.00        0.00
  PREM    Third-party and premium content               0.00        0.00        0.00        0.00        0.00        0.00
  TAXES   Taxes and regulatory fees                     4.86       15.66        4.50        4.59        4.77        4.50

Abridged — the file continues.

The outcomeWhat a good result looks like

One row per subscriber-cycle: this line's own baseline under BSS-2026, the movement, the part of it the account's own change log accounts for, the residual the band is assigned on, the component behind that residual, the status and whether the row is queued for pre-send review — with the qualifying change-log entry quoted verbatim so a reviewer can see in one glance what is already explained.

And when it cannot

TWO FAILURES, OPPOSITE IN DIRECTION AND UNEQUAL IN SIZE, AND THIS KIT NEVER AVERAGES THEM. A bill left standing that should have been queued goes out and the subscriber finds it — the paid arm does that once in 64, the free code floor eight times. A bill queued that should not have been costs a reviewer five minutes — the paid arm does that seven times in 64 and the floor three — and enough of those is how a queue stops being read, which produces the first failure. Every one of the paid arm's seven is the same mechanism: it QUOTED the change-log entry correctly and then answered the movement where the standard asks for the residual.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your pre-send export is a fixed-layout table and you have no surveillance today — the free code floor first
    it reaches 82.8 pct of the calls and 100 pct of the baselines for $0.00, in milliseconds, with no key and no network. On a file where you have nothing today, a free arm that is right on five cards in six is already the whole win.
  • What you are afraid of is a subscriber finding the bill before you do — the paid call
    this is the only place the two arms separate in a way that matters. 1 bill left standing against the floor's 8 — and every one of the floor's eight is a change-log sentence its phrase list has no word for.
  • Your change log is free text that agents type, in no fixed form — the paid call
    the three families where the floor scores 0 of 8 are all one thing: a sentence whose meaning is not in its keywords — a reversal that was refused, two dates and neither called effective, a quote by another name.
  • You want the queue to be short as well as complete — the free code floor
    3 wrong queues against the paid call's 7, and 3 of 14 change-log-explained cards against 7 of 14. A short queue is a queue people keep opening.

And where nothing here is good enough:

  • Somebody can write a note on the account before the cycle is read — neither, without a control on who writes notes
    x001-bill-shock: one sentence in a case note claiming the increase was disclosed in advance moved 4 of 24 cards, 3 of them to STANDING. The pure-code station defends none of it, because every queueable card has a clear carried state and the sentence works by moving the residual.

At a glanceHow the whole thing runs

78–84%the CALL — QUEUE / WATCH / STANDING — over the 64 surveillance cards
1,857 msp50, end to end

Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own cards in the same panel order — ACCOUNT FACTS, CHARGE HISTORY BY COMPONENT, PRIOR CYCLE SURVEILLANCE, ACCOUNT CHANGE LOG, CASE NOTES — write data/accounts.json with one carried state per card, and write data/gold.jsonl with the six answers per card. EVERY MEASURED NUMBER ON THIS PAGE STOPS BEING TRUE, and three of them stop being true immediately. Corpus lens →
When is this the wrong choice?Avoid: Buying a call before you have measured what the floor gets on YOUR export. On this corpus the paid call beats it by one card and that is inside the noise. That is the case against the best-fitting scenario (“Your pre-send export is a fixed-layout table and you have no surveillance today”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A card whose charge history is not a fixed-layout table. Both floors parse the panels by regular expression and the free code floor's exact 64 of 64 on the baseline is a fact about that layout. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A SECOND RUN OF THE SAME SET. Every figure here is one run of 64 calls. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 5 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-08 — r001-bill-shock. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, replays the committed run, and scores both free floors and the stub end to end. python3 tools/build_corpus.py --check, python3 tools/build_corpus.py --check-fillers and python3 -m evals.check_labels all pass with nothing installed and no network.

A living map of modern AI — kept current every morning