Home › Use Cases › Verify a waste service agreement's base rate against its own annual escalation clause
Use caseUC0320
🧪 Use-case kit · runnable

Verify a waste service agreement's base rate against its own annual escalation clause

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A waste services customer holds a few hundred service agreements, each with a Section 4 that says the base service rate is adjusted once a year. Whether that happened correctly is four separate questions about one line on one invoice: was an adjustment due in this period at all, was it applied on the date the clause fixes, at the percentage the clause yields once its greater-of / lesser-of / cap wording is resolved against a published index, and against the rate the clause says was in force — which is not always the last row of the rate history. Nobody checks. The rate arrives on a monthly invoice beside disposal, tonnage and two percentage fees, it is within a few per cent of last month's, and it is right often enough that looking is expensive and finding nothing is the usual result. Reading one escalation clause, resolving it against a published index table, deciding which row of a rate history the clause calls the rate in force, doing the multiplication and the rounding, and quoting the line that proves it — per account, per year, across a portfolio nobody has time to open.

Audience

The billing analyst or contract administrator who works a portfolio of waste agreements and has to decide which accounts to open. What they get here is a verdict, a variance to the cent, and the one line of the packet it turns on — and, on this corpus, a free floor that outperforms the paid call, which is the first thing they should read. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual rate verification packets

The corpus is 62 rate verification packets, 0.14 MB (txt 62). It is generated because a real one cannot be published. A rate verification packet names a real customer, a real hauler, a real agreement number and a real monthly charge, and the whole point of the artefact is that it shows what a supplier is billing one named account. What CAN be published is the structure — a commencement date, an escalation clause, a published index, a rate history in the operator's own words, and one figure billed — and that structure is what this kit measures a reading of.

The corpus

  • The 62 rate verification packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your rate verification packets. That is the whole change — there is no database to migrate.

One rate verification packet, as the model receives itRSV-0001.txt · 1 of 62
CONTRACT RATE VERIFICATION PACKET
==============================================================================
  Hauler:         Threlkeld Container Services
  Customer:       Delvin Industrial Estate
  Site:           Basement compactor room
  Agreement:      WSA-8648
  Packet:         RSV-0001
  Schedule 1:     Front-load 4 yd container, 1 lift/week
  Commencement:   2023-10-12
  Billed period:  2026-01

CONTRACT EXTRACT - Section 4, base service rate and annual adjustment
  4.1  The base service rate is the recurring monthly charge for the container service described in Schedule 1, exclusive of disposal and tonnage charges, one-time charges, taxes, and any percentage fee however described.
  4.2  The base service rate shall be adjusted with effect from the first day of January in each year, irrespective of the Commencement Date.
  4.3  The adjustment shall be a fixed two and one half percent (2.5%).
  4.4  The adjustment is applied to the base service rate then in force under this agreement, excluding any temporary concession recorded as such in the rate history.
  4.5  An adjusted rate is expressed to the cent, rounding a half cent upward.

BASE SERVICE RATE HISTORY - agreement WSA-8648
    effective     monthly base rate   recorded as
    2023-10-12           1,054.79  initial rate at commencement
    2024-01-01           1,081.16  annual adjustment, contract year 2
    2025-01-01           1,108.19  annual adjustment, contract year 3
    2025-02-01           1,047.24  contract rate amended at the customer's request, with no end date
    2025-05-01             920.63  short-term reduction after the compactor was repaired, ends 2025-07-31

CURRENT PERIOD BILLING - 2026-01
  Base service rate billed this period            1,073.42

Abridged — the file continues.

The outcomeWhat a good result looks like

One packet in, and out: the adjustment date Section 4 fixes, the percentage it yields, the rate in force it applies to, the expected rate, one of six verdicts, the monthly and annualised variance, and one line copied verbatim as evidence.

And when it cannot

And what it does when it cannot. On this corpus the paid call got 18 of 62 packets completely right after the pure-code station re-applied RATECHK-2026 — against the free keyword floor's 43. It answered a percentage on all 5 packets where the right answer is null, called 7 matching rates at variance and 4 variances a match, and quoted a line on 9 packets where the only right answer is silence.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your agreements state the adjustment date and the percentage in structured fields — no model at all
    the arithmetic is R-5 and it is free; there is no reading to buy
  • Your escalation clauses are prose but written from a small set of templates — the clause floor in evals/baseline.py
    on this corpus a regular expression over the clause text takes the adjustment date 62 of 62 and the percentage 53 of 62 for $0.00
  • Your rate histories carry concessions and amendments in free-text notes — measure BOTH before choosing
    this is the field the kit exists for and on this corpus the keyword floor still wins it 47 to 37. Run all four floors on your own corpus first; they cost $0.00
  • You want the reply to reason its way through the clause — reasoning stays OFF
    measured on the same eight packets in the same hour: the provider's default cut off 7 of 8 replies at the ceiling and returned nothing, for thirteen times the bill

At a glanceHow the whole thing runs

29%packet all correct pct
2,446 msp50, end to end
$0.00per 1,000 rate verification packets · Claude Fable 5

Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packets in the same panel layout and data/contracts.json with your own billed period, billed rate, commencement date, rate history and index table keyed by the same ids, then label data/gold.jsonl. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying a provider to do one multiplication. That is the case against the best-fitting scenario (“Your agreements state the adjustment date and the percentage in structured fields”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet whose rate-history rows are not effective rate recorded as in a column layout, or whose Section 4 clauses are not one physical line each numbered 4.1 to 4.5. src/rules.py returns an empty list rather than raising, and src/packet.py::check refuses the run. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether a keyword floor written against the generator's OWN word pools would beat the published floor. It would; it was not built, because a floor tuned to the generator measures the generator. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, reasoning explicitly disabled (r001), one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-08 — r001-rate-escalation. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9320, scores all four free floors live, replays the committed scored run from its result file, and runs the corpus rebuild, the label gate and the stub end to end. Observed on this machine: tools/build_corpus.py --check reported 0 differences under two PYTHONHASHSEEDs, evals/check_labels.py reported 0 problems over 62 packets, and the stub scored 29 of 62 cells without opening a socket.

A living map of modern AI — kept current every morning