Home › Use Cases › Scope changes hiding in engagement correspondence, and work running with no change order
Use caseUC0294
🧪 Use-case kit · runnable

Scope changes hiding in engagement correspondence, and work running with no change order

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A professional-services firm signs a statement of work with a list of deliverables on it. Then the engagement starts, and for the next six months the scope moves in e-mail, in meeting notes and in a chat channel — a client asks for something the list does not cover, a deliverable quietly comes out, a team starts work on a Tuesday that nobody has papered. None of it arrives as a change request. It arrives as correspondence, and nobody reads a whole thread against the signed schedule before a monthly scope review, so the changes surface later, in an invoice argument, when the leverage has gone. Reading a six-month correspondence thread against a signed statement of work by hand, the week before a scope review, and finding what moved. Today that is either not done at all or done once a quarter by whoever has the time, and the failure is silent in both cases: work that nobody papered goes on being done.

Audience

The engagement manager preparing a monthly scope review, and the commercial lead who reads what they produce. The decision underneath is whether to raise something with the client this month — which is exactly why this kit stops short of every commercial conclusion. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual engagement correspondence packs

The corpus is 60 engagement correspondence packs, 0.17 MB (txt 60). It is generated because it has to be. Real engagement correspondence IS personal data — named people saying things in their own words about their own work, and often about each other — and the fields that make the task hard are exactly the ones nobody may publish. This corpus is built so the difficulty survives without any of it: what makes a message hard is what it says about SCOPE, not who said it. evals/check_labels.py sweeps every pack for e-mail addresses, telephone numbers, postal addresses and dates of birth and reports 0 hits.

The corpus

  • The 60 engagement correspondence packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 60 packs, every statement of work, every change order register and the whole answer key are generated in-process from seed 20260903. data/SOURCES.md states what the generator costs the measurement rather than leaving it to be discovered.

Swap this folder for your own material and the kit is pointed at your engagement correspondence packs. That is the whole change — there is no database to migrate.

One engagement correspondence pack, as the model receives itESP-0001.txt · 1 of 60
ENGAGEMENT SCOPE PACK  ESP-0001
  Engagement             Core ledger migration
  Client                 Northmoor Utilities Group
  Statement of work      SOW-2026-101, signed 2026-04-08
  Correspondence window  2026-05-04 to 2026-06-08
  Prepared for           the engagement manager, before the monthly scope review

STATEMENT OF WORK -- SCOPE BASELINE  SOW-2026-101
  Deliverables in scope
    D1  Two integration test cycles against the client test instance
    D2  One handover session for the client finance systems team
    D3  A post-migration balance reconciliation for the opening period
  Named exclusions
    X1  Training of client staff beyond the single handover session
    X2  Any work inside third-party systems the client licenses
  Change control  Article 7 -- no work outside the deliverables listed above is performed
                  until a change order is signed by both parties.

CHANGE ORDER REGISTER
  CO-01  signed 2026-05-05  covers: A third integration test cycle
  CO-02  signed 2026-05-08  covers: Extraction of eleven legacy report definitions

CORRESPONDENCE
  M01  2026-05-04  email  from Ogilvy (client) to Bhatt (firm)
      Subject: Housekeeping
      The plan on the portal is up to date as of this morning.
      The scope review moves to 2026-05-10 and the meeting room on level four is booked.
      Copying Bhatt for visibility.

  M02  2026-05-09  meeting note  scope review, client and firm attending
      - The plan on the portal is up to date as of this morning.
      - Copying Bhatt for visibility.
      - Checking that the single handover session for the finance systems team is part of what we signed; I could not find it in the schedule.

  M03  2026-05-15  chat  delivery channel

Abridged — the file continues.

The outcomeWhat a good result looks like

One pack in, six graded rows out: per message, what it does to the scope baseline, the sentence that evidences it copied verbatim and located at character offsets, the change order on the register that covers it, and whether the work has started — and then, derived in code from those three, the list of messages that are UNPAPERED WORK. On this corpus the paid call found 54 of 54 unpapered messages with 0 false alarms and got the flag set exactly right on 60 of 60 packs, against a phrase-matching floor's 24 of 54 and a null floor's 0.

And when it cannot

And what it does when it cannot. Thirteen of 360 messages came back imperfect and NINE OF THEM ARE THE SAME HALF OF THE JOB: work coming OUT. Four times a removal recorded on a change order was read as no change at all; five times a removal that has not happened yet was marked as started. The other four are a partial quote on a declined request — the right sentence, cut at 39-45 pct coverage against a 60 pct floor. ⚑ NOT ONE OF THE THIRTEEN COSTS A FLAG: SCP-8 excludes REDUCE from unpapered work by construction, and the four quote misses sit on messages whose verdict, reference and started are all right.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know which messages moved the scope at all — the phrase floor is genuinely usable — 87.8 pct on verdicts and 100 pct on the 228 messages that change nothing
    SCP-5's four bullets turn into a workable negative list, and on a corpus this regular that is most of the job.
  • You want the unpapered work, which is what the review is for — the paid call
    54 of 54 against the floor's 24 of 54, paired p = 1.9e-09, with 0 false alarms on both. The 30 the floor loses are all the same shape: a message that names a change order which does not cover it.

And where nothing here is good enough:

  • You want work coming OUT of scope logged as reliably as work going in — neither arm, yet
    REDUCE is 9 of the paid call's 13 imperfect messages and 6 of 6 descope_papered are only 33 pct right. It is the half nobody logs and the half this kit is worst at.
  • You want a number to put in front of a client — nothing here
    This kit never approves a change order, prices one, alleges a breach or concludes that unbilled work is billable, and the answer contract has no field that could. UNPAPERED is a fact to raise internally.

At a glanceHow the whole thing runs

100%unpapered found pct
91,001 msp50, end to end
$42.36per 1,000 engagement scope packs · Gemini 3 Flash

Run once, for real, on 2026-09-03. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs in the same three-panel shape and write a gold.jsonl beside them. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Believing it transfers. The sentences here are six templates per family; the floor's number is a ceiling on pattern matching, not a floor under it. That is the case against the best-fitting scenario (“You want to know which messages moved the scope at all”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A thread longer than six messages. The input is comfortable; the REPLY is not — six objects plus the provider's reasoning budget already used 92.9 pct of a 32,000-token ceiling, and two of twelve adversarial replies were cut off at it. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled sentence is the sentence a person would have quoted. The key names ONE per changed message and the scorer has no partial credit; four messages on the scored run are right on every other field and score zero on the quote for copying part of it. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-scope-change. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9294 and scores every graded cell offline: the 60 packs, the answer key with its character offsets, both free floors computed live, and the recorded run replayed from its own result file. python3 tools/build_corpus.py --check rebuilds the corpus byte-identically and python3 -m evals.check_labels grades the key, both for $0.00 and with no network.

A living map of modern AI — kept current every morning