The business caseThe problem this solves
A customer says the bill is wrong, or an audit asks whether a cycle was billed correctly. Somebody has to recompute it: which days the plan was actually held, what the days-in-period convention says the denominator is, whether the one-time fees on the cycle are the ones the events justify, whether a SIM replacement was under warranty, and whether the same shape has been found on this account before. It is arithmetic against a written convention over records that live in three systems, done on a queue, one cycle at a time. Recomputing a disputed bill cycle by hand against the proration convention and deciding, cycle by cycle, whether the variance is a one-off or the shape of a defect that will recur. It does not replace the correction: it refers, and a person corrects and approves.
Audience
A billing operations reviewer working a cycle, and the supervisor who approves any credit that follows. The decision this report is for is narrow and it is not 'should we buy a model': it is 'does this job need one at all'. This kit's own answer is NO on the corpus it measured, and that is the finding rather than a caveat on it. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual bill cycles
The corpus is 60 bill cycles, 0.08 MB (txt 60). It is generated because it has to be. A real proration dispute is an extract from a live billing system joined to a CRM note trail and a rate card, about a named subscriber, and no operator can hand one over. It is generated the way it is for a second reason: the source row for this use case asks for the days-in-period convention to be 'validated against a written policy rather than taken as the single true convention across all plan types, since some legacy plans may differ' — so the corpus carries BOTH conventions, actual-days as default and a two-plan legacy schedule on a 30-day denominator, because a kit that assumed one convention for everything would be answering an easier question than the real one.
The corpus
- The 60 bill cyclesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states what is invented (all of it), what the generator costs the measurement (six named things), and the sign defect this kit shipped with and fixed.
Swap this folder for your own material and the kit is pointed at your bill cycles. That is the whole change — there is no database to migrate.
PRORATION & ONE-TIME CHARGE VERIFICATION PACKET PRV-0001
Prepared 2026-09-02 under PRC-2026
ACCOUNT FACTS
Account ACC-40137
Bill cycle 2026-06-08 to 2026-07-07
Plan at cycle start Horizon Essential 45
Add-ons at cycle start Roaming day pass bundle
RATE REFERENCE - monthly recurring charge
Horizon Essential 45 $ 45.00
Roaming day pass bundle $ 25.00
ONE-TIME CHARGE SCHEDULE
Activation fee $ 35.00
Late payment fee $ 10.00
Number change fee $ 25.00
SIM replacement fee $ 15.00
SERVICE EVENTS THIS CYCLE
none recorded in this cycle
BILL AS RENDERED
line kind description amount
L1 plan Horizon Essential 45 $ 45.00
L2 addon Roaming day pass bundle $ 25.00
ACCOUNT NOTES
The customer contacted care about coverage at their home address; no billing change was made.
PRIOR CYCLE HISTORY
2026-04-09 to 2026-05-08 billing dispute opened on the plan line and later withdrawn
2026-05-09 to 2026-06-07 complaint reference CMP-1014 recorded against the add-on line
The outcomeWhat a good result looks like
One packet in, eight graded answers out: the recomputed amount per line to the cent, the line verdict, the missing charges, the signed net variance, the warranty exemption, the pattern, the route and the quoted history line. A good result is every cent exact and the route right, because the route is the only field that would ever move money.
And when it cannot
And what it does when it cannot. The paid call referred SIX CORRECT BILLS to the corrective-charge queue — the pack asking a customer to pay more on a bill that was right — on 6 of the 27 clean cycles. All six are one misreading of one rule: on a cycle that opens with an activation it billed the plan at the FULL monthly charge instead of prorating it from the activation date (PR-3), reading the activation event as justifying only the one-time fee. The recompute station catches five of the six. It never sent an undercharge to the credit queue (credit_direction_wrong 0 of 60) and never closed a defective bill (no_action_where_variance 0 of 33): its errors all point one way.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your bill cycles arrive as fixed-layout exports, the way this corpus does — THE FREE RULES FLOOR ALONE — evals/baseline.py plus src/prorate.py
60 of 60 cycles on all eight fields for $0.00, no key and no network, beating the paid call on every column. Buy no inference at all. - You want the money figure and nothing else — src/prorate.py on its own
The engine is integer arithmetic over named records and has no model in it. It is what produces every published figure on both arms; the recompute station is just this engine applied to a reply. - Your reading is genuinely hard — a PDF bill, a CRM note and a rate card in three systems — the paid call WITH the recompute station, and re-measure on your own corpus
This corpus cannot tell you whether the model helps, because its reading problem is a column position. The station is what makes a model answer safe to publish either way: it took the money to 100.0 pct and the route from 90.0 to 98.3 pct on the paid run. - You want the pattern reading across cycles and are happy to recompute in code — the call for
patternandexempt_eventsonly, everything else from the engine
Those are the two fields the station cannot re-derive, because both are readings of prose no record holds. They are also the only two where a model has anything to add here — and on this corpus the floor's phrase lists match them too.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own packets and data/cycles.json with one record per packet id carrying cycle_start, cycle_end, plan_at_start, addons_at_start, the rate table, the events and the billed lines. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: A cycle whose plan or add-on changed mid-period in a way the cycle record does not carry. The engine reads the EVENT FEED; a change that exists only in a care note is invisible to it, and the floor will return a confident, wrong figure with no flag on it — the one failure mode a deterministic arm has and a model might have caught. That is the case against the best-fitting scenario (“Your bill cycles arrive as fixed-layout exports, the way this corpus does”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A scanned or photographed bill. Every arm here — the model's reading, both floors' parsers and the independent key check's parser — rests on fixed-layout panels with an amount in a column. 9 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Run-to-run variance on the paid arm. It was fired ONCE — 60 calls — and no repeat was bought, so nothing here reports a confidence interval on 86.7 or 98.3 pct. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-02 — r001-prorate-verify. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9274 and scores every graded cell offline: the corpus rebuilds from seed 20260902 byte-identically under two PYTHONHASHSEEDs, evals/check_labels.py re-derives all 480 cells and reports 0 failures, both free floors score all 480, and both committed paid runs replay from their result files. Nothing to install — Python standard library end to end. The only thing a key buys is the live 'ask the fast tier' button, and on this corpus it buys nothing else.



