Home › Use Cases › Classify each rejected e-billing line by the guideline clause it fails
Use caseUC0462
🧪 Use-case kit · runnable

Classify each rejected e-billing line by the guideline clause it fails

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A client's e-billing platform rejects a line on a law firm's invoice and returns a code. The code is not a reason: the mapping from platform code to guideline clause is built per platform and is incomplete, and on this corpus the code names a DIFFERENT cause from the guidelines on 18 of 64 lines and is unmapped on 6 more. So the billing coordinator opens the matter, the outside-counsel guidelines in force, the agreed rate schedule and the prior billing, and decides by hand which clause the line actually fails — and whether there is still an appeal window to use. the coordinator's read of eight panels per rejected line to decide which guideline clause it fails and whether an appeal window is still open — it does not replace the e-billing platform, the decision to appeal, or the partner who signs one.

Audience

A billing partner or revenue lead deciding whether to put a model in front of a cycle's rejections. This report's answer is NO on the accuracy: a free card walk plus a short keyword list beats the paid call on every graded field, separably. What is worth taking is the framework around it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual rejected e-billing lines

The corpus is 64 rejected e-billing lines, 0.24 MB (json 4 · jsonl 1 · md 1 · txt 64). A real e-billing rejection extract carries client names, matter names, timekeeper names and negotiated rates, so it cannot be shipped and a redacted one cannot be labelled. This one is BUILT as structures and the key is OCG-2026 applied to those same structures, which is the only way to have 64 labelled lines whose key is derivable rather than opinion. It is built to be hard in the ways a real cycle is: 14 lines carry their clause only in a sentence somebody typed, 8 carry evidence for two clauses where the card's order decides, 14 carry a sentence the printed columns REFUTE, 5 have no guidelines attached at all, and the platform's own code names a different cause on 18.

The corpus

  • The 64 rejected e-billing linesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your rejected e-billing lines. That is the whole change — there is no database to migrate.

One rejected e-billing line, as the model receives itEBL-0001.txt · 1 of 64
REJECTED E-BILLING LINE  EBL-0001
====================================================================================================
PANEL 1 — MATTER AND INVOICE HEADER
  Firm                     Marrowgate Legal LLP (synthetic firm)
  Client                   CL-24 (speciality retailer)
  Matter                   MAT-4107  employment claim
  Billing cycle            2026-06 billing cycle
  Billing period ended     2026-05-31
  Invoice                  INV-7100
  Submitted to platform    2026-06-12   (12 days after the period ended)
  Platform                 Meridian Billing Exchange

PANEL 2 — THE REJECTED LINE
  Line                     INV-7100 line 3
  Work date                2026-05-04
  Timekeeper               TK-102  (Partner)
  Task code                TSK-34  settlement analysis
  Activity code            ACT-1
  Hours                    6.5
  Rate billed              795.00
  Line amount              USD 5,167.50
  Narrative                3 task segments recorded in one entry
      1  model settlement range
      2  draft settlement note
      3  call with client
  Amount rejected          USD 5,167.50

PANEL 3 — OUTSIDE-COUNSEL GUIDELINES IN FORCE FOR THIS MATTER
  §2.4     invoices must reach the platform within 30 days of the end of the billing period
  §2.7     work already billed and paid may not be billed again
  §3.1     a time entry covering 3 or more separate tasks and longer than 4.0 hours is block billing
  §3.2     each timekeeper is billed at the rate agreed for them on the matter's rate schedule
  §4.1     task codes open on this matter type (employment claim): TSK-31, TSK-32, TSK-33, TSK-34, TSK-35

Abridged — the file continues.

The outcomeWhat a good result looks like

One row per rejected line a billing coordinator could work as written: the cause under OCG-2026, the guideline clause, the line of the file that establishes it, the appeal deadline, and — derived in pure code, never asked — the review queue, the escalation flag, the platform-code verdict and the cycle rollup by matter and timekeeper.

And when it cannot

It abstains. A line whose outside-counsel guidelines never came through cannot be walked at all, and the answer is NEEDS-REVIEW whatever the billing note says — 5 of 64 lines, and every arm gets all 5 right because the rule is structural. Where the guidelines ARE attached and no clause fits it answers NO-GUIDELINE-BASIS, which is a finding and the one a firm most wants — and the one this arm is least willing to give (4 of 11).

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • your guidelines state the submission window, the rate schedule and the expense caps as NUMBERS, and your export prints them — the free card walk — evals/baseline.py, $0.00
    it is 50 of 50 on the lines settled by printed panels, 8 of 8 on the precedence lines the card's order decides, 14 of 14 on the refutation traps and 11 of 11 on the residual. The paid call is 32, 6, 5 and 4 on the same four. A call there buys nothing and loses you lines the arithmetic already had.
  • the deciding clause lives in a sentence the coordinator typed, and nowhere else — a keyword floor you have actually written — and only then ask whether a call adds anything
    this is the one place the call reads something the plain card walk cannot: 10 of the 14 prose lines against the card walk's 0. It is a real capability. It is also beaten by the keyword floor on the same lines, 13 of 14, which is why the headline here is a loss rather than a trade.
  • the appeal window matters more than the cause, because a missed window ends the appeal — free code, unconditionally
    every free arm here reads OCG §7's two sources and scores 64 of 64. The paid call reached 58 and asserted a date once where the window is unconfirmed — the single unsafe thing the run did.
  • you want what this kit is actually worth on your corpus — the FRAMEWORK — the graded refusal behaviour, the pure-code recheck applied to every arm alike, the derived queue/flag/verdict, the adversarial arm and the measured cost by tariff
    the accuracy result here is a loss, and it is a loss that only exists because a free floor was written and measured first. That is the transferable part: on a corpus whose notes carry more of the decision the same machinery would report the opposite, and it would be believable for the same reason.

At a glanceHow the whole thing runs

66%rechecked cause accuracy pct
1,520 msp50, end to end
$0.82per 1,000 rejected e-billing lines · the fast tier, inside the peak window

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/policy.json with your client's own outside-counsel guidelines — the clauses, the submission windows, the rate-schedule rule, the expense caps and the two appeal-window sources — and src/taxonomy.py's causes and their ORDER with yours. The measured result does not travel. Corpus lens →
When is this the wrong choice?Avoid: Paying per line for a date and an amount comparison. That is the case against the best-fitting scenario (“your guidelines state the submission window, the rate schedule and the expense caps as NUMBERS, and your export prints them”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?an e-billing export whose panels are not the fixed layouts src/panels.py expects — the parser is positional, and a re-laid-out file yields a guidelines panel the recheck reads as NOT ATTACHED, so every arm abstains at once, paid and free alike. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER A PAID CALL IS WORTH MAKING AT ALL ON THIS CORPUS — it is not. It lost to the free floor of record on every graded field and the loss is separable on four of them: cause 42 v 53 (p = 0.0266), clause 38 v 53 (p = 0.0003), deadline 58 v 64 (p = 0.0312), all four 24 v 53 (p < 0.0001). 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, PEAK tariff. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-ebill-rejection. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — The kits repository is private, so this is a measurement of the kit rather than an offer: on a copy of this folder with no key configured, all four free floors, evals/check_labels.py, the committed paid run replayed with every paired test, and the board on port 9462 reproduced at 0 calls, $0.00 and a few seconds of wall clock. Re-running the paid arm needs a provider key.

A living map of modern AI — kept current every morning