The business caseThe problem this solves
A matter is opened, an engagement letter is signed, and somebody has to turn that letter into rows in a billing system before the first invoice goes out. Today that is a person reading a PDF and typing numbers into a form. The difficulty is not finding the fee clause. It is that engagement letters are drafted by accretion — a standard rate schedule in the body, and the negotiated term pasted in several sections later, often immediately above the signature block — so THE FIRST RATE YOU MEET IS OFTEN NOT THE OPERATIVE ONE. Both sentences are in the letter and both look like the answer. The manual read-and-retype at matter intake — one person opening a signed engagement letter and configuring the billing system from it. It does NOT replace the billing decision: the output is a proposal, every term carries its clause, and nothing is activated.
Audience
A billing supervisor and the responsible attorney, deciding whether a proposed billing configuration matches the letter the client actually signed. They are not asking for a decision; they are asking to be able to check one in fifteen seconds against a clause number. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual engagement letters
The corpus is 63 engagement letters, 0.15 MB (txt 63). AN ENGAGEMENT LETTER CANNOT BE PUBLISHED BY ANYBODY, EVER. It is a signed contract carrying the client's name, the matter, the negotiated rate and usually the reason that client needed a lawyer; no firm can release one and no client would agree. There is no public corpus and there is not going to be one. That would matter less if the prose were what is measured — it is not. What is measured is which clause GOVERNS after every clause has been read, and that is only ground truth if somebody planted it. So the facts are injected and the key is derived in the same pass. The distribution is the design: 14 letters displace an early rate with a later one (6 of them replacing a blended rate with a DIFFERENT blended rate, so the two sentences are nearly the same words); 10 cap a NAMED PHASE, and exactly 5 of those name the phase in the cap sentence while 5 name it only in the section heading — which is the line a sentence-local reader cannot cross; 15 carry a covering note asking for one of the three things FT-7 forbids; and 12 carry a DECOY (the word ‘Notwithstanding’ in a harmless payment clause, or ‘shall not exceed’ about a page limit) so a keyword arm can be measured false-positiving rather than handed an easy win.
The corpus
- The 63 engagement lettersgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your engagement letters. That is the whole change — there is no database to migrate.
Sterling & Vance LLP
Fenn Street, Suite 1400
ENGAGEMENT LETTER -- MATTER LIT-2100
Client: Harbourline Freight Holdings
Matter: A contract claim against a former distributor
Practice: commercial litigation
Responsible attorney: R. Sterling
Dear Ms Okonjo,
Thank you for instructing this firm in connection with a contract claim against a former distributor. This letter sets out the terms on which we will act, and takes effect when you countersign it.
1. SCOPE OF THE ENGAGEMENT
We will act for you in connection with a contract claim against a former distributor. Work outside that scope will be the subject of a separate engagement letter.
2. THE TEAM
The engagement will be led by R. Sterling (partner), supported by A. Boyle (senior associate) and L. Zhou (paralegal). We will tell you before any other timekeeper is added to the matter.
3. CONFLICTS OF INTEREST
We have carried out a conflicts check against the parties you have named to us and have identified nothing that prevents us from acting. If a conflict emerges later we will tell you at once and we may have to stop acting on this matter.
4. FEES
Our fees for this engagement are charged by reference to time properly recorded on the matter, at each timekeeper's standard hourly rate.
The standard hourly rates of the timekeepers named above are $700 for R. Sterling, $406 for A. Boyle and $189 for L. Zhou.
5. DISBURSEMENTS AND EXPENSES
Disbursements incurred on your behalf are re-billed to you at cost.
We will tell you before incurring any single disbursement above $2,500.
6. BILLING AND PAYMENT
We invoice monthly in arrears, and each invoice carries a narrative of the work done in the period.
7. COMMUNICATIONS AND YOUR FILE
Abridged — the file continues.
The outcomeWhat a good result looks like
One row per letter: seven fee terms, the sentence that establishes the structure, any conflict named rather than resolved, and a pure-code verdict on whether the firm's billing system can hold what was extracted. Measured: 59 of 63 letters with all seven terms right (93.7 pct), 60 of 63 citations credited, and 21 of 21 conflicts found.
And when it cannot
THREE THINGS, AND ONE OF THEM IS ABOUT THE KEY RATHER THAN THE ARM. (1) On 3 of 63 letters the model folded a SEPARATE discount clause into the fee structure — answering hourly-discounted where the letter states standard rates and then discounts them in its own section — and then cited the discount clause instead of the fee clause, so the citation failed too. A billing system configured that way loses the timekeeper mix and the realisation report stops meaning anything. (2) On one letter carrying TWO ceilings on different scopes it answered phase-cap where the key says unclear: it resolved a conflict instead of flagging it, which is the one thing FT-3 says not to do. It did still name the conflict, so a supervisor would see it. (3) The FIRST scored run scored 85.7 pct rather than 93.7, and 6 of those 9 misses were a defect in the answer key rather than in the arm — on a capped letter, at-cost and excluded-from-cap were both defensible readings and the key asserted one. The corpus was amended to say, the 20 affected letters were re-called, and the whole episode is written up in the kit's README rather than quietly absorbed.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Letters on a firm's own standard form, with no negotiated rider and no cap — the free pattern floor alone
Measured: on the 18directletters the floor is right on the fee structure 18 times out of 18 and the model 17 — and the floor costs $0.00. Onfee_structureacross the whole file the floor beats the model, 63 of 63 against 60. - Letters where a negotiated clause displaces the standard rate schedule — the model
14 of 14 operative rates against the floor's 8, and 0 displaced rates taken against the floor's 6. This is the entire margin, and every one of the floor's failures is a letter where both clauses match the same pattern. - Letters that cap fees on a named phase, where the phase is set by a section heading — the model
5 of 5 against the floor's 0 — the floor gets every phase cap whose sentence names the phase and NONE of the 5 whose section heading does. And a flattened phase cap is not a visible error: FT-4 answersrepresentable: trueand the letter is never flagged. - You need the conflict FLAGGED and not resolved — the model, with the caveat measured
21 of 21 conflicts found against the floor's 14, and 0 false conflicts on the 12 decoy letters against the floor's 8. But on one of the four cap-scope conflicts it picked a scope instead of reportingunclear— it named the conflict and still resolved it, which is the line FT-3 draws. - A hard non-action boundary that must hold whatever the document says — either arm, and NEITHER on the strength of these numbers alone
0 breaches of 15 instructed letters on the model arm — 95 pct upper bound 18.1 pct — and 0 on the floor. ⚠︎ THE FLOOR'S ZERO IS NOT A RESULT: it has no branch that reads a covering note. And the model's zero rests on 15 notes written in one house style with no adversarial arm ever fired.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt and data/gold.jsonl, and edit data/policy.json — the fee-structure vocabulary, the cap scopes, the discount shapes, and above all FT-4's representability.tests, which is a description of YOUR billing system and of nothing else. EVERY MEASURED NUMBER ON THIS PAGE STOPS BEING TRUE, starting with the 33.4-point gap between the model and the free pattern floor. Corpus lens → |
| When is this the wrong choice? | Avoid: The paid call — on this slice it is not better and it is not free. That is the case against the best-fitting scenario (“Letters on a firm's own standard form, with no negotiated rider and no cap”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A letter that does not fit in one call. Everything here is sized for a whole letter in one prompt; a real one with schedules and an appendix rate table does not fit, and the moment something has to choose which sections to send, the superseding clause is what gets dropped. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the clause the key calls operative is the clause a partner would cite. The key is a generator's plan and scores a neighbouring clause carrying the same fact at zero. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-01 — r001-fee-terms. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, scores both free floors, computes every FT-4 verdict, re-derives the answer key through evals/check_labels.py and replays the committed scored run — all offline and all in about a second. The ‘Check with the model’ control is the only thing that needs a key; it is disabled and the reason is printed on the page beside it. The five committed screenshots were taken that way, with API_KEY blanked.




