Home › Use Cases › Compare a revised procedure against the revision it supersedes, item by item
Use caseUC0403
🧪 Use-case kit · runnable

Compare a revised procedure against the revision it supersedes, item by item

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A procedure is revised, and the revision has to agree with three things at once: the revision it supersedes, its own revision history, and the register of forms and procedures it cites. All three disagree quietly. A step loses a verification while gaining words. A criterion's floor becomes a ceiling with the number untouched. A form is cited at a revision that was superseded last month. A cross-reference points at a step the revision no longer contains. And the revision history — the document's own account of what changed — does not mention the section the change sits in. nothing. It is a reading step in front of a review that happens anyway.

Audience

the named document owner, and whoever reviews a revision before it goes to a change board Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual change packs

The corpus is 63 change packs, 0.30 MB (txt 63). A revision comparison needs two things a public corpus cannot give at once: two versions of the SAME document with a known relationship between them, and a document-control register saying which of the things they cite is current. Real ones are confidential and real regulatory text cannot be reproduced. So the MECHANISM is carried into an invented quality system and the relationship is DECIDED before any text exists — every cell carries a scale position and every history line carries the sections it accounts for, so the key is computed and there is nowhere in the generator to type one.

The corpus

  • The 63 change packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 63 change packs and the whole answer key is generated. Nothing here is fetched, scraped, licensed or derived from anything that was.

Swap this folder for your own material and the kit is pointed at your change packs. That is the whole change — there is no database to migrate.

One change pack, as the model receives itSCP-0001.txt · 1 of 63
STANDARD OPERATING PROCEDURE - REVISION CONSISTENCY PACK

PACK HEADER
  Procedure           SOP-4001
  Title               Preparation and use of cleaning solution for product-contact surfaces
  Quality system      Merrowfield Bioscience quality system
  Change control      CC-2001
  Documents           SUP, REV

DOCUMENT SUP
  Document            SUP
  Paper               Superseded revision
  Procedure           SOP-4001
  Revision            Rev 05
  Document owner      Quality Operations
  Step list           4.1, 4.2, 4.3
  SECTION  1  Purpose statement                  Describes the preparation of the cleaning solution and its use on product-contact surfaces.
  SECTION  2  Performing role                    Production operator
  SECTION  2  Second-person check                Performed by one operator
  SECTION  2  Record timing                      Recorded at each step as it is performed
  SECTION  3  Cleaning agent                     Alkaline detergent AD-2
  SECTION  3  Minimum agent concentration        at least 1.5 percent by volume
  SECTION  3  Equipment identifier               Vessel V-311
  SECTION  4  Step 4.1 instruction               Rinse the vessel with purified water.
  SECTION  4  Step 4.2 instruction               Record the conductivity reading on Form F-214 Rev 02 and have it verified by a second person.
  SECTION  4  Step 4.3 instruction               Return the vessel to service once it has been left to stand.
  SECTION  5  Hold time                          up to 30 minutes after the final rinse
  SECTION  5  Rinse count                        at least 2 rinses of the vessel
  SECTION  5  Conductivity limit                 not more than 5 uS/cm at the final rinse
  SECTION  5  Residue limit                      not more than 0.001 percent

Abridged — the file continues.

The outcomeWhat a good result looks like

One consistency report per change pack: every comparable item with the first rule of SCP-2026 that reaches it, the direction, which revisions stated anything, both sides quoted verbatim out of their own panels, the sections each history line accounts for, the document-level findings, the sections referred and one status.

And when it cannot

A requirement that really went missing reported as a re-wording, or a re-wording reported as a change. The first sends a revision forward with something quietly dropped; the second fills a referral list with sections nobody needed to open, which is how a referral list learns to be ignored — and that is what produces the first.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your two revisions are already structured and differ only in how a number is printed with its unit, or in the order of a ;-separated list — src/rules.py's R-4 normaliser, with no model at all
    1,107 of the 1,130 comparable items here are settled by R-4 and R-5 together, at $0.00 and with no network.
  • A hold time, a rinse count, a limit, a concentration or a retention period moved, and both revisions state it on the same unit with the same qualifier — src/rules.py's R-5 arithmetic, still with no model
    the direction is subtraction against a scale data/items.json declares, and the free floor gets every one of them.
  • You need to know whether a cited form revision is current, a cross-referenced procedure is retired, or a step cross-reference still resolves — src/pack.py's citation derivation and R-18's three lookups, with no model
    they are register and step-list comparisons. No arm on this board is credited for them and none should be.
  • A revision-history line names no section and you need to know what it accounts for — the paid call
    27 of 28 against the keyword floor's 14, and the paired exact test holds at p = 0.000244. This is the one place on this page where the money is supported by evidence.
  • A step instruction was re-worded and you need to know whether it asks more or less — a human, and read the two lines
    13 of 23 for the paid call against the free floor's 6, and the paired test does NOT support the gap (p = 0.065). On the sub-family where the instruction got LONGER and asked LESS, every arm on this board scores 0 of 5.

At a glanceHow the whole thing runs

82%rechecked report all correct pct
4,657 msp50, end to end
$8.54per 1,000 change packs · google/gemini-3-flash

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own change packs in the same panel shape and everything else runs unchanged — but read data/SOURCES.md first. The answer key is DERIVED from the generator's own structure. Corpus lens →
When is this the wrong choice?Avoid: Paying for a call whose answer a normaliser and a comparison already produced. That is the case against the best-fitting scenario (“Your two revisions are already structured and differ only in how a number is printed with its unit, or in the order of a ;-separated list”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a real controlled procedure, which arrives as a PDF — no OCR, no PDF reader, no layout model. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?THE KEY ON THE 23 ITEMS AND 28 HISTORY LINES THAT NEED A READING. There the key is the generator's own structural scale position and section set, decided before any text existed. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, reached only through src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-sop-consistency. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone, python3 -m evals.run --run-id b000-sop-consistency-rules --floor rules, and a scored free arm comes back in under a second with no key, no install and no network.

A living map of modern AI — kept current every morning