The business caseThe problem this solves
A professional-services firm signs a statement of work with a list of deliverables on it. Then the engagement starts, and for the next six months the scope moves in e-mail, in meeting notes and in a chat channel — a client asks for something the list does not cover, a deliverable quietly comes out, a team starts work on a Tuesday that nobody has papered. None of it arrives as a change request. It arrives as correspondence, and nobody reads a whole thread against the signed schedule before a monthly scope review, so the changes surface later, in an invoice argument, when the leverage has gone. Reading a six-month correspondence thread against a signed statement of work by hand, the week before a scope review, and finding what moved. Today that is either not done at all or done once a quarter by whoever has the time, and the failure is silent in both cases: work that nobody papered goes on being done.
Audience
The engagement manager preparing a monthly scope review, and the commercial lead who reads what they produce. The decision underneath is whether to raise something with the client this month — which is exactly why this kit stops short of every commercial conclusion. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual engagement correspondence packs
The corpus is 60 engagement correspondence packs, 0.17 MB (txt 60). It is generated because it has to be. Real engagement correspondence IS personal data — named people saying things in their own words about their own work, and often about each other — and the fields that make the task hard are exactly the ones nobody may publish. This corpus is built so the difficulty survives without any of it: what makes a message hard is what it says about SCOPE, not who said it. evals/check_labels.py sweeps every pack for e-mail addresses, telephone numbers, postal addresses and dates of birth and reports 0 hits.
The corpus
- The 60 engagement correspondence packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — every one of the 60 packs, every statement of work, every change order register and the whole answer key are generated in-process from seed 20260903. data/SOURCES.md states what the generator costs the measurement rather than leaving it to be discovered.
Swap this folder for your own material and the kit is pointed at your engagement correspondence packs. That is the whole change — there is no database to migrate.
ENGAGEMENT SCOPE PACK ESP-0001
Engagement Core ledger migration
Client Northmoor Utilities Group
Statement of work SOW-2026-101, signed 2026-04-08
Correspondence window 2026-05-04 to 2026-06-08
Prepared for the engagement manager, before the monthly scope review
STATEMENT OF WORK -- SCOPE BASELINE SOW-2026-101
Deliverables in scope
D1 Two integration test cycles against the client test instance
D2 One handover session for the client finance systems team
D3 A post-migration balance reconciliation for the opening period
Named exclusions
X1 Training of client staff beyond the single handover session
X2 Any work inside third-party systems the client licenses
Change control Article 7 -- no work outside the deliverables listed above is performed
until a change order is signed by both parties.
CHANGE ORDER REGISTER
CO-01 signed 2026-05-05 covers: A third integration test cycle
CO-02 signed 2026-05-08 covers: Extraction of eleven legacy report definitions
CORRESPONDENCE
M01 2026-05-04 email from Ogilvy (client) to Bhatt (firm)
Subject: Housekeeping
The plan on the portal is up to date as of this morning.
The scope review moves to 2026-05-10 and the meeting room on level four is booked.
Copying Bhatt for visibility.
M02 2026-05-09 meeting note scope review, client and firm attending
- The plan on the portal is up to date as of this morning.
- Copying Bhatt for visibility.
- Checking that the single handover session for the finance systems team is part of what we signed; I could not find it in the schedule.
M03 2026-05-15 chat delivery channelAbridged — the file continues.
The outcomeWhat a good result looks like
One pack in, six graded rows out: per message, what it does to the scope baseline, the sentence that evidences it copied verbatim and located at character offsets, the change order on the register that covers it, and whether the work has started — and then, derived in code from those three, the list of messages that are UNPAPERED WORK. On this corpus the paid call found 54 of 54 unpapered messages with 0 false alarms and got the flag set exactly right on 60 of 60 packs, against a phrase-matching floor's 24 of 54 and a null floor's 0.
And when it cannot
And what it does when it cannot. Thirteen of 360 messages came back imperfect and NINE OF THEM ARE THE SAME HALF OF THE JOB: work coming OUT. Four times a removal recorded on a change order was read as no change at all; five times a removal that has not happened yet was marked as started. The other four are a partial quote on a declined request — the right sentence, cut at 39-45 pct coverage against a 60 pct floor. ⚑ NOT ONE OF THE THIRTEEN COSTS A FLAG: SCP-8 excludes REDUCE from unpapered work by construction, and the four quote misses sit on messages whose verdict, reference and started are all right.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You want to know which messages moved the scope at all — the phrase floor is genuinely usable — 87.8 pct on verdicts and 100 pct on the 228 messages that change nothing
SCP-5's four bullets turn into a workable negative list, and on a corpus this regular that is most of the job. - You want the unpapered work, which is what the review is for — the paid call
54 of 54 against the floor's 24 of 54, paired p = 1.9e-09, with 0 false alarms on both. The 30 the floor loses are all the same shape: a message that names a change order which does not cover it.
And where nothing here is good enough:
- You want work coming OUT of scope logged as reliably as work going in — neither arm, yet
REDUCE is 9 of the paid call's 13 imperfect messages and 6 of 6descope_paperedare only 33 pct right. It is the half nobody logs and the half this kit is worst at. - You want a number to put in front of a client — nothing here
This kit never approves a change order, prices one, alleges a breach or concludes that unbilled work is billable, and the answer contract has no field that could. UNPAPERED is a fact to raise internally.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-03. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own packs in the same three-panel shape and write a gold.jsonl beside them. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Believing it transfers. The sentences here are six templates per family; the floor's number is a ceiling on pattern matching, not a floor under it. That is the case against the best-fitting scenario (“You want to know which messages moved the scope at all”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A thread longer than six messages. The input is comfortable; the REPLY is not — six objects plus the provider's reasoning budget already used 92.9 pct of a 32,000-token ceiling, and two of twelve adversarial replies were cut off at it. 8 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the labelled sentence is the sentence a person would have quoted. The key names ONE per changed message and the scorer has no partial credit; four messages on the scored run are right on every other field and score zero on the quote for copying part of it. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-03 — r001-scope-change. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9294 and scores every graded cell offline: the 60 packs, the answer key with its character offsets, both free floors computed live, and the recorded run replayed from its own result file. python3 tools/build_corpus.py --check rebuilds the corpus byte-identically and python3 -m evals.check_labels grades the key, both for $0.00 and with no network.



