Home › Use Cases › Classify each engagement's realization shortfall by cause, citing the evidence line
Use caseUC0449
🧪 Use-case kit · runnable

Classify each engagement's realization shortfall by cause, citing the evidence line

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An engagement closes and the time recorded is worth more than what was collected. The subtraction is already printed — standard value, billed, collected, the total shortfall. The work is saying WHY, because the cause decides who looks at it: a scope overrun goes to pricing, re-performed work to quality, senior hours billed at a junior rate to resourcing, and a shortfall nothing on the file explains goes to the practice leader. Today the write-off reason is whatever the person releasing the invoice keyed, and on this corpus that keyed reason names a different cause from the card on 39 of 64 files, read at face value. the reviewer's read of eight panels per file to decide which cause a shortfall is — it does not replace the billing system, the decision to write anything off, or the practice-leader review.

Audience

A practice finance lead deciding whether to put a model in front of the quarter's write-offs. This report's answer is: the classification is worth doing, and a free parser plus five keyword patterns does it as well as the paid call on this corpus. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual engagement shortfall files

The corpus is 64 engagement shortfall files, 0.19 MB (json 4 · jsonl 1 · md 2 · txt 64). A real write-off register carries client names, fee terms and people's time, so it cannot be shipped and a redacted one cannot be labelled. This one is BUILT as structures and the key is RVC-2026 applied to those same structures, which is the only way to have 64 labelled files whose key is derivable rather than opinion. It is built to be hard in the ways a real register is: 17 files carry their cause only in a sentence someone typed, 8 carry evidence for two causes where the card's order decides, 4 negate the keyword a pattern would match, and the keyed billing-system reason names a different cause from the key on most files.

The corpus

  • The 64 engagement shortfall filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your engagement shortfall files. That is the whole change — there is no database to migrate.

One engagement shortfall file, as the model receives itRLZ-0001.txt · 1 of 64
ENGAGEMENT REALIZATION FILE  RLZ-0001
================================================================================================
PANEL 1 — ENGAGEMENT HEADER
  Firm                   Quillmere Advisory Partners (synthetic firm)
  Practice               Transaction Advisory
  Client                 C-201 (industrial distributor)
  Engagement             E-26-0401  buy-side diligence, 2026
  Quarter                2026-Q2, 1 April to 30 June 2026
  Fee arrangement        HOURLY AT STANDARD RATES
  Engagement letter      EL-26-0401 dated 2026-03-02

PANEL 2 — REALIZATION THIS QUARTER
  Standard value of time recorded                  USD 60,280.00
  Billed                                           USD 60,280.00
  Collected                                        USD 57,530.00
  Billing shortfall (standard minus billed)        USD 0.00
  Collection shortfall (billed minus collected)    USD 2,750.00
  Total realization shortfall                      USD 2,750.00   95.4% realized

PANEL 3 — SCOPE IN THE ENGAGEMENT LETTER
  Task  Description                                        In letter  Change order
  T1    Data room index and gap list                       yes        —
  T2    Quality of earnings bridge                         yes        —
  T3    Debt-like items schedule                           yes        —

PANEL 4 — TIME RECORDED THIS QUARTER, BY TASK AND LEVEL
  NOT ATTACHED — the time detail did not come through with this file. Practice
  finance holds the time records for this engagement.

PANEL 5 — STAFFING PLAN AGAINST ACTUAL HOURS
  Level       Std rate  Planned   Actual   Hours over plan
  Manager       420.00     60.0     60.0   none
  Senior        310.00     82.0     82.0   none
  Associate     230.00     42.0     42.0   none

Abridged — the file continues.

The outcomeWhat a good result looks like

One row per engagement shortfall file that a practice finance reviewer could file as written: the cause under the firm's card RVC-2026, the line of the file that establishes it, the review route and whether it goes on the practice-leader list — plus the clients with a recurring cause.

And when it cannot

It abstains. A file whose time records are not attached cannot be classified from the page, and the walk stops at NEEDS-REVIEW whatever the notes say — 5 of 64 files, and every arm gets all 5 right because the rule is structural. It also answers NO-RECORDED-REASON when the file is complete and nothing on it explains the shortfall, which is a real answer here, not a failure.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • your billing export prints the cause in a column — adjustment reasons on time rows, a cap adjustment, a credit note — the free columns walk — evals/baseline.py, $0.00
    it is 47 of 47 on the files whose cause is printed in a column; the paid call is 42 of 47 on the same files. A call there buys nothing and loses you files the arithmetic already had.
  • the reason lives in the billing notes, and people write negations — the paid call, read against a keyword floor you have actually written
    it is right on all 4 notes that say a thing did NOT happen (the keyword floor 0) and 15 of 17 prose files (the keyword floor 12) — and still not separable from that floor on 64 files, p = 0.8036.
  • both, which is every real register — write the columns floor first, then test a floor-first composition on a FRESH corpus before buying anything
    on this corpus a composition that lets the columns floor decide and uses the paid reply only where the floor falls through reads 61 of 64 on 26 of 64 calls — but its rule was chosen AFTER the paid run's misses were visible. That is post-hoc analysis, not a measured product result: a hypothesis to test, not a number to buy on.

At a glanceHow the whole thing runs

89%rechecked cause accuracy pct
1,354 msp50, end to end
$0.33per 1,000 engagement shortfall files · the fast tier, off-peak

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/policy.md with your own card, data/policy.json's routes and threshold with yours, and src/taxonomy.py's causes and their order with yours. The measured result does not travel. Corpus lens →
When is this the wrong choice?Avoid: Paying per file for an amount match. That is the case against the best-fitting scenario (“your billing export prints the cause in a column — adjustment reasons on time rows, a cap adjustment, a credit note”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a billing export whose panels are not the fixed layouts src/panels.py expects — the parser is positional, and a re-laid-out file yields panels the recheck reads as NOT ATTACHED, so every arm abstains at once. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE PAID CALL BEATS FREE CODE — it did not separate from the best free floor: cause 57 v 55, p = 0.8036; all four fields 48 v 55, p = 0.2100, the floor ahead. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, off-peak. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-realization-cause. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — The kits repository is private, so this is a measurement of the kit rather than an offer: on a copy of this folder with no key configured, all four free floors, evals/check_labels.py, the committed paid run replayed with every paired test, and the board reproduced at 0 calls, $0.00 and 3 seconds of wall clock. Re-running the paid arm needs a provider key.

A living map of modern AI — kept current every morning