Home › Use Cases › Check a time narrative against the billing guideline that governs it, and cite the clause
Use caseUC0469
🧪 Use-case kit · runnable

Check a time narrative against the billing guideline that governs it, and cite the clause

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A law firm's billing coordinator closes a cycle by reading every time entry on every draft pre-bill against the client's outside counsel billing guideline — and the guideline has been amended four times, so the version that governs an entry is the one that was in force on the day the work was done, not the one filed today. The firm's own billing system prints a guideline version on each entry, and that stamp is a coding decision rather than a reading of the amendment notices. Today a coordinator reads the notices, works out which version reaches each work date, and then applies thirteen clauses to six narratives per packet by eye. Reading the amendment notices to work out which guideline version was in force on each work date, then applying thirteen clauses to every narrative on the packet by eye, and writing the clause and version on each flag by hand.

Audience

The billing coordinator who has to send a packet back to a timekeeper or let it assemble, and the billing manager who signs the release. The answer this report gives them is conditional: AS ANSWERED the paid call loses to a free rules engine, and it only wins once the guideline is re-applied in pure code to the readings it made. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual draft pre-bill QC packets

The corpus is 64 draft pre-bill QC packets, 0.62 MB (json 3 · jsonl 1 · md 2 · txt 64). Because the shape of a narrative-QC failure is not the arithmetic, it is WHICH VERSION the arithmetic runs under — and a corpus that does not carry that trap measures a rules engine rather than a reading. The firm's billing system stamps a version on every entry and the stamp is wrong on 118 of 384; 87 of those are entries where believing the stamp reaches a different verdict, and 47 of THOSE are false flags the stamp raises that the guideline does not. A free rules engine handed that column applies the increment, the approved list and the daily cap faultlessly under a version that never governed the day, which is exactly the 55.2 pct / 61 false-flag floor this kit publishes.

The corpus

  • The 64 draft pre-bill QC packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 64 packets, data/matter_records.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.

Swap this folder for your own material and the kit is pointed at your draft pre-bill QC packets. That is the whole change — there is no database to migrate.

One draft pre-bill QC packet, as the model receives itPBL-0001.txt · 1 of 64
PRE-BILL QC PACKET

PACKET HEADER
  Packet                 PBL-0001
  Client                 Crestmere Holdings
  Matter                 CRH-2519  Doramy joint venture wind-down
  Billing guideline      OCG-BILL-2026
  Guideline as filed     v4
  Billing cycle          2025-09
  Packet assembled       2025-09-05
  Prepared by            S. Molyneux
  Release gate           The billing manager releases the invoice. This packet does not.

DECLARED FIRM STANDARDS
  Manager review threshold   2.00 flagged hours in a packet
  QC lead time               5 calendar days between the last work date and assembly
  These are the client's own declared service standards for this engagement.
  No regulation is cited for either.

GUIDELINE VERSION TERMS
  Version  MinIncr  DailyCap  Clerical        Travel  ApprovedLevels
  v1          0.10     12.00  no              no      PARTNER, ASSOCIATE, PARALEGAL
  v2          0.10     10.00  yes             no      PARTNER, ASSOCIATE, PARALEGAL, STAFF ATTORNEY
  v3          0.25     10.00  yes             yes     PARTNER, ASSOCIATE, PARALEGAL, STAFF ATTORNEY
  v4          0.25      8.00  yes             yes     PARTNER, ASSOCIATE, PARALEGAL, STAFF ATTORNEY, CONTRACT ATTORNEY
  Clerical work under v2 and later is billable only in connection with this matter's own
  production or filing during the cycle. Under v1 an entry of 0.20 hours or less need not
  name its subject matter; that allowance was removed in v2.

GUIDELINE AMENDMENT NOTICES
  A-1   Beginning with work performed on 1 January 2024, version v1 governs.
  A-2   Version v2 applies to work performed on or after 1 October 2024. An invoice submitted before 15 December 2024 under the earlier version is not reopened for that reason alone.

Abridged — the file continues.

The outcomeWhat a good result looks like

Every time entry carries one verdict from a closed list of ten, the clause it breaches, the guideline version it was checked against and one short reason — and a flag that does not name both its clause AND its version is not counted as a flag at all. The packet carries a release route, a lead-time verdict and the list of entries to send back. An entry whose governing version the notices do not settle is GUIDELINE-VERSION-NOT-DETERMINABLE and is never guessed; CLEAR-TO-ASSEMBLE is a QC outcome and not a release.

And when it cannot

⚠︎ AS ANSWERED, THE PAID CALL LOSES TO FREE CODE, AND IT IS PUBLISHED AS A LOSS. The measure is cited-flag recall with the false-flag count beside it. Raw, the call takes 47.5 pct at 105 false flags against the domain rules engine's 55.2 pct at 61 — McNemar exact on the same 384 entries, p = 0.000031. It only wins after src/recheck.py re-applies the guideline in pure Python to the five readings it made: 65.0 pct at 39 false flags, p = 0.000355. ⚠︎ AND THE RECHECKED FIGURE NEEDS A RECHECKED FLOOR BESIDE IT: the constant answer, which reads nothing at all, goes from 0.0 pct at 0 false flags to 42.0 pct at 47 once the same station has run. A generator-tuned regex — not a product, the size of the leak — beats every arm at 87.4 pct with 0 false flags.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A pre-bill cycle where the client's guideline has been amended and the billing system's version coding did not follow the amendment notices — this kit, with the station
    the whole margin is the version reading: 283 of 384 against the free engine's 266, and on the 87 entries the notices alone settle the paid arm cites 34 where the free engine cites 0.
  • A guideline with ONE version in force and no amendment history — the free domain rules engine alone
    with nothing to read, every clause in OCG-BILL-2026 is arithmetic over the entry's own columns and the engine does it for $0.00 at 70.0 pct verdict accuracy — the model buys nothing it does not already have.

And where nothing here is good enough:

  • You need the pack to STOP where the guideline stops — neither arm unsupervised
    on the 28 entries where two live notices disagree and the only right answer is to decline, the paid arm declines correctly 9 times and the free engine 0. The station cannot repair a wrong version reading — it applies it.
  • You want a single headline accuracy for a slide — nothing on this page
    a constant reply reaches 62.8 pct verdict accuracy here against a 70.0 pct floor. Any one number from this kit is either a loss, a win, or a figure a reply that reads nothing nearly matches.

At a glanceHow the whole thing runs

65%cited flag recall pct
2,350 msp50, end to end
$1.22per 1,000 draft pre-bill QC packets · GPT-5.6 Luna

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own rendered pre-bill packets and data/matter_records.json at your own structured half, put your own guideline in data/policy.md and data/policy.json, and run python3 -m evals.run --stub first — the whole harness, no socket, $0.00. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Quoting the raw column — as answered it loses to the free engine on the same entries. That is the case against the best-fitting scenario (“A pre-bill cycle where the client's guideline has been amended and the billing system's version coding did not follow the amendment notices”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet with many more than six time entries. Six is the corpus's shape and the whole packet goes into one call; the largest reply was 799 of a 1,000-token ceiling, so eight entries is about the practical bound at this rung and splitting the packet breaks G-7's daily cap and G-10's amendment history, which both read the packet whole. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether any of this holds on a real firm's pre-bills. Every figure here is against a key derived from a generator, on an invented guideline (OCG-BILL-2026), for an invented client. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-narrative-qc. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on a clean checkout of this repository with no key configured and no network: python3 -m evals.baseline prints all three free arms in a few seconds at $0.00, python3 -m evals.check_labels re-derives the whole key at 0 disagreements, and python3 -m src.app serves the board with the model button disabled and the reason printed beside it. Nothing in that path needs pip install.

A living map of modern AI — kept current every morning