Home › Use Cases › Change regulatory impact comparison
Use caseUC0374
🧪 Use-case kit · runnable

Change regulatory impact comparison

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A manufacturing site raises a change control record: a few rows of what is currently registered against what is proposed, and a paragraph of justification. Somebody in regulatory affairs has to say which registered element each row touches and which variation category that triggers — against the dossier as approved, not from memory. It is slow, it is done from a spreadsheet of precedents, and the two errors it can make are not the same size. the manual placement of each proposed change against the registered dossier and the variation code

Audience

The regulatory affairs lead who owns the filing decision, and the site quality team who raised the record. The decision this report supports is whether a paid model call is worth making at all — and on the answer the kit exists to produce, this run says it is not. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual change control records

The corpus is 46 change control records, 0.09 MB (txt 46). EVERY WORD IS INVENTED AND IT HAS TO BE. Change control records, registered dossiers and variation filings are confidential documents belonging to identifiable companies, and a kit that shipped one — redacted or not — would be shipping somebody's regulatory strategy. Aldoria is not a country; the Aldoria Medicines Board does not exist; AVC-2026 is a table in src/rules.py. 46 records, 167 change items, 5 invented products with 11 registered elements each, and 12 distinct (element, nature) pairs. NOTHING THIS KIT PRODUCES IS REGULATORY ADVICE.

The corpus

  • The 46 change control recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.

Swap this folder for your own material and the kit is pointed at your change control records. That is the whole change — there is no database to migrate.

One change control record, as the model receives itCC-0001.txt · 1 of 46
ALDORIA PHARMA CHANGE CONTROL RECORD
SYNTHETIC DOCUMENT - generated for the AI Foundry variation-trigger kit from one
seed. Not a real record, not a real product, not a real filing.

Record              CC-0001
Raised              2026-04-18 by Packaging Development
Product             Terrivane 10 mg/mL solution for injection
Authorisation       MA-AL-2026-0219   (Aldoria Medicines Board)
Holder              Northreach Biologics N.V.
Programme           Cost-of-goods programme

PROPOSED CHANGES
  Item 1  Manufacturing site address
      currently registered : Northreach Biologics, Building 2, Harrowgate Park, Lenn
      proposed             : Northreach Biologics, Building 2, Harrowgate Park North, Lenn
  Item 2  Excipient supplier - sodium chloride
      currently registered : Marrow Salts Co, Pellen
      proposed             : Marrow Salts Co (second plant)

DESCRIPTION AND JUSTIFICATION
  Item 1.
      The site is unchanged in every respect. What changes is how its address is
      written, following a municipal renumbering.
  Item 2.
      The excipient supplier already named in this dossier will supply the same
      material to the same grade from a second plant of its own. The registered
      supplier entity is unchanged.

DECLARATION
  This record is a PROPOSAL. It is not a filing and it is not a regulatory
  assessment. No change described here may be implemented before the regulatory
  affairs lead has decided the route and recorded that decision.

The outcomeWhat a good result looks like

Every proposed change item carries the registered element it touches, the nature of the change, the variation category that triggers, the rule that decided it quoted in full, and a verbatim sentence from the registered dossier. The record carries one filing route: the highest route any of its items triggers.

And when it cannot

It OVER-files. 21 of the 167 change items were proposed a route above the key's and 1 below. The dominant single error is reading a minor manufacturing change as a change of principle (7 items), which turns a 30-day notification into a prior approval — a filing nobody needed.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the variation category and nothing else — the free numeric floor
    140 of 167 items for $0.00 and no network, against 145 for the paid call. Paired, p = 0.486850: this corpus cannot tell them apart, and one of them has a bill.
  • Your change records use a field vocabulary your lookup table does not cover — the paid call
    the registered element is the one column it wins outright — 166 of 167 against 160, p = 0.031250. A field map is a maintained artifact; this is the job it is bad at.
  • Your worry is under-filing, not accuracy — the paid call, with a person on every item
    it proposed a route below the key's 1 time against the floor's 7. That is the error that reaches an inspection — but paired it is p = 0.070312, a trend and not a result, and 167 items is not enough to settle it.

And where nothing here is good enough:

  • The justification paragraph is written by a party with an interest in the answer — neither, without a person
    6 of 12 injected wordings moved an item's route DOWN the ladder, all six by direct instruction. The free floors are not immune either — a phrase table reads the same attacker-controlled text.

At a glanceHow the whole thing runs

99%element correct pct
1,615 msp50, end to end
$0.69per 1,000 change control records · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/dossiers.json with your products and their registered sentences, and drop your change control records into data/corpus/ with a matching data/records.json. THE MEASURED ACCURACY DOES NOT TRAVEL. Corpus lens →
When is this the wrong choice?Avoid: Paying per record for a column a table of regular expressions matches. That is the case against the best-fitting scenario (“You want the variation category and nothing else”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a record that does not number its proposed changes — the reply is keyed on that number and there is nothing else on the page to join a reading to a row. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?whether the result holds on a real variation framework — AVC-2026 is invented, and a real framework's conditions would move items between routes. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-variation-trigger. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the registered dossiers, the answer key, the committed run and all three free floors — and scores every free floor offline. evals/check_labels.py, evals/baseline.py and evals/run.py --rescore all run with no network and no credential. Only evals/run.py --arm model, the injection probe, and the board's one Ask button ever reach a provider.

A living map of modern AI — kept current every morning