The business caseThe problem this solves
A content licence is agreed as a two-page DEAL MEMO — territory, window, exclusivity, holdbacks, media, fee, payment schedule, MFN — and drafted into a LONG-FORM AGREEMENT weeks later by people who were not all in the room. Somebody in licensing operations then compares the two, term by term, and puts every material difference on business affairs' deviation register for the cycle. It is not a diff. The long-form restates every term in drafting language, so most of the text disagrees and almost none of it matters — and the difference that costs the most is not in the text at all. A territory, a holdback or an MFN the memo granted and the long-form never addresses reads, to any comparison of the clauses both documents contain, exactly like agreement. It is found when the licensee exercises the right and finds it is not there, or when their auditor asks to see what was signed against what is being enforced. The analyst's own pass over one memo and one long-form for one term — locating the clause on each side, deciding whether the difference is substance or drafting, and writing the register row. It replaces the LOCATING and the READING and none of the deciding: LC-2026 is applied in code either way, the disposition is fixed, and the locating half turns out not to need a model at all.
Audience
The licensing operations analyst working a comparison cycle, and the business affairs lawyer who disposes of every entry the cycle raises. Behind both of them, the licensee's auditor, whose question is the one the row this kit answers records: show me that what we signed matches what you are enforcing — and if it does not, show me who approved the difference. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual comparison packs
The corpus is 60 comparison packs, 0.23 MB (txt 60). Content licensing paperwork is confidential by construction: a deal memo and the long-form drafted from it are commercially sensitive to two parties at once, and there is no public corpus of memo/long-form pairs. There will not be one. So the choice was between measuring on documents nobody else can see — which makes every number here unverifiable — and modelling the SHAPE of the work on invented text anybody can re-run. What is modelled from reality is the two-document structure, the set of terms a comparison turns on, the fact that a long-form RESTATES rather than copies, version amendment, the file-note tail, and the four ways the comparison goes wrong. What is invented is every value. The cost of that trade is stated in data/SOURCES.md and in breaks_on rather than in a footnote.
The corpus
- The 60 comparison packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your comparison packs. That is the whole change — there is no database to migrate.
DEAL COMPARISON PACK CI-0001
Deal: D-0001 - "Harbour Lights" (8 x 47' drama)
Licensee: Meridian Streaming Ltd
Comparison cycle: 2026-Q1
Deal memo: DM-D-0001, dated 14 October 2025
Long-form compared: LF-D-0001 v1, dated 11 January 2026
TERM UNDER COMPARISON: TERRITORY
This pack is compared against long-form version v1 and no other. Report the comparison for the
term named above only.
SECTION A - DEAL MEMO DM-D-0001 (dated 14 October 2025), in full
--------------------------------------------------------------------------
M-1 TERRITORY. The Territory is the United States, Canada and the United Kingdom.
M-2 LICENCE TERM AND WINDOW. The Window runs for 24 months from 1 January 2026.
M-3 EXCLUSIVITY. The Licence is exclusive in the Territory for the first 12 months of the Window.
M-4 HOLDBACKS. Licensor holdback on all other linear and streaming exploitation in the Territory for 6 months from first transmission.
M-5 MEDIA AND RIGHTS GRANTED. Rights granted are subscription video on demand and catch-up streaming.
M-6 LICENCE FEE. Licence fee is USD 360,000 for the Window.
M-7 PAYMENT SCHEDULE. Fee payable in three equal instalments: on signature, on delivery, and on first transmission.
M-8 MOST-FAVOURED-NATION. Most-favoured-nation applies against comparable licensees in the Territory for the Window.
M-9 SUBLICENSING. No sublicensing without Licensor's prior written consent.
M-10 DELIVERY AND TECHNICAL. Delivery in ProRes 422 HQ with textless elements and English subtitles.
M-11 AUDIT RIGHTS. Licensor may audit the Licensee's records once in any twelve-month period on thirty days' notice.
M-12 MARKETING AND CREDIT. Licensor credit in the end board of each episode and in the Licensee's synopsis metadata.
Abridged — the file continues.
The outcomeWhat a good result looks like
One row per material term that a lawyer can dispose of without re-reading either document: what the long-form did to the term, the memo clause and the long-form clause it did it in, the operative sentence from EACH document copied verbatim, and the LONG-FORM VERSION the comparison was made against — so a later re-comparison against a further-amended long-form cannot be confused with this one. Where the long-form is silent, the row says SILENT rather than saying nothing.
And when it cannot
Three directions and they are not comparable. A SILENT TERM READ AS AGREEMENT is a right the memo granted that nobody looks at again until somebody tries to exercise it — the failure this kit exists for. NEITHER ARM PRODUCED ONE: silence_read_as_agreement is 0 on both, across the 12 items whose long-form addresses the term nowhere. A RESTATEMENT REPORTED AS DRIFT is a business affairs review that did not need to happen — the paid call raised 0 of 6, the free floor 2 of 6. AN ENTRY DISPOSED OF RATHER THAN REFERRED is the cap breaking: 0 of 60 on the raw replies, including 0 of the 8 packs whose own file notes ask for the entry to be approved and the licence released, and 0 of 12 under a dedicated adversarial arm.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You need to know whether a term the memo granted is in the long-form at all — the free rules floor
12 of 12 on this corpus, $0.00, no network, and a heading scan is the right instrument for a structural question - You need to tell a restatement from a real narrowing — the paid call
3 of 6 against the floor's 0 of 6; LC-2 is a judgement about substance and a token comparison is not making one - You need the register a licensing cycle owes, complete — the paid call
31 of 32 entries at 94 pct precision against the floor's 26 at 87 — five entries, and every one is a class the floor cannot express - You need to be sure nothing approves a deviation or releases a licence — src/policy.py's LC-11, on either arm
the disposition is forced in pure code on every arm, and the raw breach is counted before the force so the guarantee is measured rather than asserted: 0 of 60, 0 of the 8 packs whose notes ask, 0 of 12 under a dedicated attack
At a glanceHow the whole thing runs
Run twice over the same set, for real, the last on 2026-09-01. Every figure on these pages was captured from those runs — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own packs as .txt into data/corpus/ in the same three-section shape (SECTION A the memo, numbered M-1..M-n with an upper-case heading and one sentence; SECTION B the long-form, numbered X.Y with a heading and one sentence; SECTION C the file notes), add one row per pack to data/records.json naming the deal, the cycle, the term under comparison and the long-form version, and run the two free arms — python3 -m evals.run --run-id b000-<yours>-rules --floor rules scores immediately with no key. ⚠︎ WHAT STOPS BEING TRUE. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for it — the call ties, it does not win. That is the case against the best-fitting scenario (“You need to know whether a term the memo granted is in the long-form at all”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A LONG-FORM WITH SCHEDULES AND EXHIBITS. Every clause here is one sentence in a consistent drafting voice. 8 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE LABELLED LINE IS THE LINE A LAWYER WOULD HAVE QUOTED. The key names one operative sentence per side and scores everything else zero. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-01 — r002-dealmemo-compare. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — ⚠︎ PARTIALLY EVIDENCED, AND THE REST IS NOT MEASURED. What was observed on this machine: with API_KEY unset, python3 tools/build_corpus.py rewrites all 60 packs and the key, python3 -m evals.check_labels grades that key at 0 problems, python3 -m evals.run --floor rules scores all 60 items offline at wall_seconds 0.0, and the board on 127.0.0.1:9252 renders every panel and replays the committed paid run — the six screenshots in this lens were taken that way, with the key blanked. What was NOT measured: a clone into an empty directory on a machine that has never held this repository, and any Python other than the 3.x on this box.





