Home › Use Cases › Term by term, is anyone comparable on better terms and what does the clause require
Use caseUC0231
🧪 Use-case kit · runnable

Term by term, is anyone comparable on better terms and what does the clause require

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A most-favoured-nation clause promises that nobody inside a group the clause itself defines is on better terms, on the terms it lists, subject to carve-outs it writes down. Checking it looks like a spreadsheet job: put the counterparty's other agreements in columns, put the protected one beside them, flag anything smaller. Today that is what gets built, and on this corpus 59.7 pct of everything it raises is not a breach -- 56.9 pct once the same free, deterministic recheck this kit ships is run over it, which is still more than half -- because almost none of the work is comparing numbers. It is deciding WHICH numbers the clause is even about. Building the MFN comparison by hand or in a spreadsheet -- reading which of the counterparty's agreements the clause actually reaches, which side letters change an effective value, which carve-out covers what, and then computing the remedy the clause requires.

Audience

Whoever signs the letter that goes to the counterparty's lawyer -- the commercial or legal reviewer who has to stand behind both the finding and the figure. The decision this board informs is narrow and it is not 'is the model good': it is whether a paid reading step is worth buying on top of a column comparison that costs nothing and already runs. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual review packs

The corpus is 60 review packs, 0.38 MB (txt 60). The three judgements an MFN review actually turns on, planted deliberately, because no public dataset can carry them: what a counterparty pays its OTHER distributors is a negotiating position and is never published. 19 rows put the better number OUTSIDE the group the clause defines (8 wrong territory, 5 wrong tier, 3 another service, 3 outside the look-back); 14 put it inside the group and behind a written carve-out; 8 turn on the clause's own netting rule, 4 where it neutralises the differential entirely and 4 where it survives partially; and 7 riders move an effective value so that the breach is invisible on the printed face. Those are the 33-plus-7 rows a two-box checker cannot express in either direction, and they are the whole reason the kit has four verdicts rather than two.

The corpus

  • The 60 review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your review packs. That is the whole change — there is no database to migrate.

One review pack, as the model receives itMFN-0001.txt · 1 of 60
MOST-FAVOURED-NATION REVIEW
  File reference        MFN-0001
  Review date           31 August 2026
  Protected party       Northgate Cable Partners
  Counterparty          Meridian Media Group
  Agreement under review  AG-0100, executed 1 January 2024
  Covered terms         rate, committed term, exclusivity, payment timing

THE CLAUSE  (AG-0100, clause MFN-1, as executed)
----------------------------------------------------------------------------------------------------------------------
  MFN-1.1  Comparison group
     Agreements of the Counterparty for the Service "Meridian Sports Network", with distributors in Tier 1,
     in the territory United Kingdom, executed within 18 months of the review date. An agreement
     failing any one of those attributes is not a Comparison Agreement.

  MFN-1.2  Covered terms, and what each one means
     Rate             the recurring per-subscriber fee, net of all RECURRING credits,
                      rebates and support payments. A one-off or signing payment is NOT
                      netted, whatever its size.
     Committed term   the months during which the distributor may not terminate for
                      convenience. A unilateral termination right SHORTENS it; an option
                      to renew does not change it.
     Exclusivity      the number of exclusivity holdbacks the agreement imposes.
     Payment timing   the days allowed to settle in the ordinary course. An extended
                      settlement rider LENGTHENS it; a suspension of the clock during a
                      formal dispute does not, being unavailable in the ordinary course.

  MFN-1.3  Carve-outs
     CO-1   A rate granted within the first 6 months following that distributor's launch

Abridged — the file continues.

The outcomeWhat a good result looks like

One verdict per covered term that a person can confirm as written: 240 term rows over 60 review packs, each CARVED_OUT, NOT_COMPARABLE, BREACH or COMPLIANT, with the offending agreement named and the new rate, the arrears months and the credit computed to the cent in pure Python. Measured: 238 of 240 verdicts right (99.2 pct) and 58 of 60 packs right on all four terms (96.7 pct), against the same free code put through the same free recheck at 200 and 26. THE CALL BUYS 38 TERM ROWS AND 32 PACKS. (Until 2026-08-31 the free column read 196 and 23, because the floor arm never ran the recheck every model arm ran; nothing about the paid arm changed.)

And when it cannot

⚠︎ IT FILED THE WRONG ACTION ON THE MOST VALUABLE PACK IN THE CORPUS, AND THE FREE FLOOR FILED THE RIGHT ONE. On MFN-0010 the model called the rate row NOT_COMPARABLE and returned file_action NO_ADJUSTMENT_DUE. The key says BREACH against AG-0490 with a credit of $3,175,200.00 -- the largest single credit in the corpus -- and the rules floor returns that row exactly right in every field: the verdict, the agreement, the new rate 1.7500, 12 arrears months and the credit to the cent. That one row is 100 pct of the model's entire money shortfall ($3,175,200.00 unclaimed of $30,609,000.00 at stake). A wrong verdict here is not a number on a scoreboard; it is a claim not raised, and the clause's arrears cap means the months that pass are not recoverable later.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A counterparty whose paper is uniform and whose clause has no carve-outs and no netting step -- a flat parity undertaking over one tier in one territory — the free rules floor alone
    On the 21 CLEAN packs of this corpus the floor already scores 19 of 21 and the paid call scores 20. One pack. Where nothing sits outside the group, nothing is carved out and nothing is netted, a column comparison is the answer and it costs $0.00 in 0.097 s for all sixty.
  • A clause that defines its comparison group by attributes and writes down carve-outs -- which is what an MFN clause normally does — the paid reading step, with the pure-code recheck behind it
    This is where all the money is. On the 39 HARD packs free code through the same recheck scores 7 and the paid arm 38 (the raw floor, this kit's free column until 2026-08-31, scores 4). ALL 33 of free code's rechecked false breaches are rows a two-box checker cannot express at all -- 19 NOT_COMPARABLE and 14 CARVED_OUT -- and no amount of threshold tuning creates a verdict the vocabulary does not have. Raw it raised 37, the extra 4 being the clause's netting rule, which is the one thing the recheck hands back.
  • Any clause whose remedy is a figure that goes in a letter — the pure-code station, always, whatever reads the pack
    Both arms hand their reading to the same engine and the arithmetic is then exact -- and free code is handed the same engine, so this is not a line the paid call buys. Rechecked, free code still gets the money wrong on 4 of the 25 breaches it finds, by $16,214,300.00, and all four name an agreement outside the comparison group entirely: a reading error the arithmetic cannot reach. Raw, before the recheck, it was 8 rows and $24,729,000.00, the other 4 being the netting rule the recheck hands back. A claim that is right with an indefensible number is the version that gets settled at a discount.

And where nothing here is good enough:

  • A pack assembled from paper the counterparty controls, going to a provider verbatim — nothing on this page -- the measurement does not exist yet
    The adversarial arm is written, paired and unfired. There is no suppression rate here in either direction, and the recheck's trust in carve_outs means the pure-code station is inside the blast radius rather than outside it.

At a glanceHow the whole thing runs

99%verdict accuracy pct
54,100 msp50, end to end
$23.82per 1,000 review packs · Google Gemini 3 Flash

Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own review packs into data/corpus/ in the printed layout this kit uses -- the clause as executed, the protected agreement, a COMPARISON AGREEMENTS table, a TERMS AS PRINTED ON THE FACE table and a RIDERS AND SIDE LETTERS block -- and mirror their structured form in data/agreements.json, which is what the engine and the free floor read. ⚠︎ THE PACK REACHES YOUR CONFIGURED PROVIDER VERBATIM, AND IT IS ASSEMBLED FROM THE COUNTERPARTY'S OWN PAPER. Corpus lens →
When is this the wrong choice?Avoid: Paying per pack for a judgement the clause does not require anybody to make -- and inheriting a 54.1 s p50 to do it. That is the case against the best-fitting scenario (“A counterparty whose paper is uniform and whose clause has no carve-outs and no netting step -- a flat parity undertaking over one tier in one territory”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A REAL PACK, which is a PDF. Everything here is fixed-width columns in one layout -- a COMPARISON AGREEMENTS table, a TERMS AS PRINTED ON THE FACE table and a RIDERS AND SIDE LETTERS block, in that order. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?ROBUSTNESS TO A HOSTILE PACK. evals/injection.py is written, paired and has NEVER been fired -- no x001 result file exists and every figure it would produce is unmeasured. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-31 — r001-mfn-check. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on a copy of the kit with no .env and no API key: tools/build_corpus.py rebuilt all 60 packs, the answer key and the stats in 0.072 s and the result was byte-identical to the committed corpus; evals/check_labels.py re-graded the key with an independent parser and independent arithmetic in 0.050 s and reported no disagreement on any of the 240 rows; the free rules floor scored all 60 packs in 0.097 s and its result file matched the committed b000 record in every field except its own timestamp; the wiring stub ran in 0.090 s. The board and the corpus render with no key, and /api/review returns 200 saying so rather than failing. Nothing was bought and nothing was reached over a network.

A living map of modern AI — kept current every morning