Home › Use Cases › Reconcile the rates loaded in a billing system against the agreed rate schedule
Use caseUC0346
🧪 Use-case kit · runnable

Reconcile the rates loaded in a billing system against the agreed rate schedule

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An engagement letter and a set of outside-counsel guidelines agree a rate schedule: a standard rate per timekeeper grade, a discount, an annual increase cap, and a roster of timekeepers the client has approved. Then somebody loads that schedule into a billing system, by hand, once a year, alongside every other client's. What actually sits in the rate table afterwards is a different question from what was agreed, and the two drift for ordinary reasons: a rate letter supersedes a row and nobody closes the old one, a bulk load points at the wrong matter for one cycle, a load runs twice, a row is keyed at the standard rate because the discount is applied at invoice on some matters and at load on others, or a timekeeper is loaded who was never approved. Nobody notices until a bill is queried, and by then the rate has been billing for months. Opening a matter's rate table beside its engagement letter, working down the loaded rows one at a time, reading the memo and the note under each, applying the discount by hand to check the figure, looking each timekeeper up on the approval roster, and multiplying every difference by the budgeted hours to see whether it is worth raising.

Audience

A law firm's billing desk or legal-operations analyst reconciling a matter's rate load, and the client-side legal-operations lead who audits the same table from the other side. Whoever has to answer 'show me every timekeeper billing this matter and the document that approved their rate'. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual rate load review files

The corpus is 62 rate load review files, 0.22 MB (txt 62). It is generated because it has to be. A rate table is a law firm's own commercial record and a client's rates are the most confidential number in an engagement; there is no public corpus of them and there never will be. The exact shapes this kit is about — a row a rate letter superseded and nobody closed, a bulk load that pointed at the wrong matter for one cycle, a row keyed at the standard rate because the discount is applied at invoice — are also the shapes nobody would ever publish. Generating them is what makes the key derivable rather than written, which is what makes every committed run re-scorable by anybody, forever, for nothing.

The corpus

  • The 62 rate load review filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every client is an invented trading name, every timekeeper an id and a grade, and there is no personal data in the corpus at all — evals/check_labels.py sweeps every file for email addresses, telephone numbers, national insurance and social security numbers, payment card numbers and postal addresses on every run and found 0.

Swap this folder for your own material and the kit is pointed at your rate load review files. That is the whole change — there is no database to migrate.

One rate load review file, as the model receives itRLD-0001.txt · 1 of 62
==============================================================================
RATE LOAD REVIEW FILE                                      RLD-0001
Client: CLT-4100 - Halloway Brands (invented)
Matter: MTR-4000   commercial disputes panel   Rate year: 2026-01-01 to 2026-12-31   Procedure: RLR-2026
==============================================================================

MATTER AND RATE YEAR AS THE ENGAGEMENT REGISTER HOLDS IT
  currency unit                       USD
  engagement discount                8.00   pct off standard
  annual increase cap                5.00   pct over prior agreed
  tolerance pct                      0.50   pct of the agreed rate
  threshold exposure             2,500.00   or more
  threshold pct                      3.00   pct or more
  rate year                     full-year

AGREED RATE SCHEDULE AS THE ENGAGEMENT LETTER AND THE OCG PRINT IT
  GRADE      TIMEKEEPER GRADE                     STANDARD       AGREED   PRIOR AGREED
  GR-01      Equity Partner                       1,380.00     1,269.60       1,257.02
  GR-02      Salaried Partner                       780.00       717.60         710.49
  GR-03      Counsel                                725.00       667.00         660.39
  GR-04      Senior Associate                       670.00       616.40         610.29

CLIENT-APPROVED TIMEKEEPER ROSTER
  TIMEKEEPER   GRADE    APPROVED      BUDGET HOURS
  TK-1001      GR-01    2025-11-15             180
  TK-1002      GR-02    2025-11-06             240
  TK-1003      GR-03    2025-11-12             520
  TK-1004      GR-04    2025-11-25             330
  TK-1005      GR-01    2025-11-07              60
  TK-1006      GR-02    2025-11-16             590
  TK-1007      GR-03    2025-11-15             420

Abridged — the file continues.

The outcomeWhat a good result looks like

One matter in, one row out: which loaded rows this procedure treats differently from the table that printed them and why, whether a rate amendment is in force, what the rates in force cost against the agreed schedule over the rate year at the matter's own budgeted hours, and one verdict from an ordered rulebook — with the row quoted verbatim for every line it moved.

And when it cannot

And what it does when it cannot. On the scored run 62 of 62 replies parsed and nothing stopped at the ceiling, so there is no unparsed column to report. What there IS to report is the opposite failure: on 22 of 62 matters the arm cited a row that IS the rate in force and is merely off schedule, the station took it out as instructed, and a term breach came back as a clean matter. clean_where_fault is 1 on the raw column and 16 on the rechecked one.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • auditing a matter whose rate table the billing system already reports as clean — the paid arm
    on the 16 matters where the system's own reconciliation is wrong it is 11 to 0
  • a first pass over a whole panel, ranked by what to look at — the panel floor, then the paid arm on what it flags
    the free arm costs nothing and is right on 34 of 62 matters
  • finding the rows a rate letter superseded and nobody closed — the paid arm
    12 of 14 on that family against 0 of 14 for the panel and 4 for free code
  • deciding whether a rate column is the standard rate or the agreed one — free code
    it is arithmetic and the schedule prints both numbers; the free rules floor gets 2 of 5 for nothing and the paid arm 4
  • deciding whether an increase was countersigned or refused — the paid arm
    every refused increase carries a real grade code, a real figure and a real in-range effective date, so the regex floor applies all 19 of them

At a glanceHow the whole thing runs

61%all five correct pct
2,216 msp50, end to end
$0.00per 1,000 rate load review files · google/gemini-3-flash

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own review files in the same shape and data/matters.json with your own engagement register — the matter's discount, cap, tolerance and thresholds, the agreed schedule per grade, and the approved roster with its budgeted hours. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Trusting the verdict it returns on a matter that was already right — that is the half it loses. That is the case against the best-fitting scenario (“auditing a matter whose rate table the billing system already reports as clean”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A rate table that is not fixed-width columns. src/ratesheet.py's row regex is the shape these files print — id, timekeeper, grade, rate, effective date, kind, status, reference, memo, with the note indented under the row. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-rate-load. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors, all six committed runs and every one of the 62 matters. python3 tools/build_corpus.py --check, python3 -m evals.check_labels and every --floor arm run with no credential and no network.

A living map of modern AI — kept current every morning