Home › Use Cases › Draft one engineering change's cutover checklist, or name the step its records fail
Use caseUC0513
🧪 Use-case kit · runnable

Draft one engineering change's cutover checklist, or name the step its records fail

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An engineering change order is raised on a manufactured item, and before the new revision can cut in somebody has to draft the cutover checklist: which of the plant's eleven controlled steps this change actually requires, which record on file supplies each one, and which step has nothing behind it. The steps required turn on TWO readings nobody has made yet — the change CLASS (does it alter a characteristic the item's own drawing lists as controlled?) and the EFFECTIVITY TYPE (when does it cut in?) — and on which revision of the plant's standard governs, which turns on the date the change was raised. The change system's own panel ticks a class box and proposes a checklist; on this corpus that box is wrong on 35 of 64 packets and the panel's checklist is the wrong SHAPE on 40. Today an engineer reads the packet and drafts the checklist by hand. Reading one engineering change order packet — the narrative, the drawing's controlled characteristics, the ticked class box, the effectivity clause and twelve plant records with their remarks — against the revision of ECP-2026 in force on the day it was raised, and typing the cutover checklist with a record id or a gap beside every step.

Audience

The change analyst who drafts the cutover checklist and the engineer who approves it before anything cuts in. What this report tells them is narrow and it was pre-registered before a cent was spent: the paid call is worth buying for the CLASS reading and therefore the SHAPE of the checklist, and for nothing else on this corpus. Every other cell free code already does as well or better. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual engineering change order packets

The corpus is 64 engineering change order packets, 0.23 MB (json 4 · jsonl 1 · md 2 · txt 64). Because a cutover checklist fails on WHICH STEPS ARE REQUIRED, not on the arithmetic. The generator plants the traps a change desk actually meets — a ticked class box that is wrong on 35 of 64 packets, a characteristic named by a synonym rather than the words the drawing prints, 10 Class C packets that quote an old and a new value for something the drawing does not control, 18 packets raised under one revision and reviewed under the next, a run-out with a backstop date, an immediate cut that also quotes a serial, and 35 records whose own remark takes the record back — and derives the key from the true records. reading_required is 50 of 64: on those the plant's own panel reaches a different checklist than the standard does over the true records.

The corpus

  • The 64 engineering change order packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your engineering change order packets. That is the whole change — there is no database to migrate.

One engineering change order packet, as the model receives itECO-0001.txt · 1 of 64
ENGINEERING CHANGE ORDER PACKET

CHANGE HEADER
  Change                ECO-0001
  Plant                 Harrowgate Instrument Works, Tressilick plant
  Raised                2026-03-08
  Reviewed              2026-04-16
  Target cutover        2026-05-10
  Originator            P. Vantrell
  Affected programme    SABRETHORN
  Class box as ticked   CLASS-B
  Approval gate         Engineering approves. This packet drafts a checklist and does
                        nothing else.

DECLARED PLANT STANDARDS
  Change-control standard      ECP-2026
  Rev 4 in force from          2026-04-01
  The governing revision is the one in force on the RAISE date above.
  These values are supplied by the operator for this scenario. They are not a claimed
  policy, and no regulation is cited for any of them.

AFFECTED ITEM
  Item                  40-1100   Spindle housing
  Revision              A to B
  Drawing               HD-401100
  Work instruction      WI-401100-01
  Supplier              Vorlaine Components
  CONTROLLED CHARACTERISTICS, from drawing HD-401100: bore diameter, mass, thread pitch

REASON FOR CHANGE
  Raised by P. Vantrell following a process capability study. On item 40-1100, the bore moves
  from 12.40 mm to 12.55 mm. Drawing HD-401100 is reissued at rev B to record it.
  Parts built to the new revision do not fit the previous assembly.

EFFECTIVITY STATEMENT
  Effective on exhaustion of current stock of item 40-1100 at rev A.

DISPOSITION NARRATIVE
  On hand at the time of raising: 426 at rev A. Work in process: 52 units at operation 10.

RECORDS ON FILE
  CNO-30007   customer notification   programme SABRETHORN   sent 2026-04-30
      remark: "a quotation pack was issued from this release; manufacture is authorised"

Abridged — the file continues.

The outcomeWhat a good result looks like

One packet in, one call out, one cutover checklist. The decision comes from a closed list of four — DRAFT-COMPLETE, DRAFT-WITH-GAPS, HOLD-CLASS-UNREADABLE, HOLD-EFFECTIVITY-UNREADABLE — with the change class, the effectivity type, the ordered step list ECP-2026 requires, a state per step (SUPPLIED, ABSENT or NOT-COVERING) with the record id that supplies it, the approval route, and a note of at most 700 characters. THE SHAPE OF THE CHECKLIST IS THE ANSWER: a checklist counts only when every step the standard requires is present and none is extra.

And when it cannot

⚠︎ THE HEADLINE IS A PAIR AND IT IS SPLIT; IT IS NEVER AVERAGED. Against the bar — tuned, the best free arm shipped, at $0.00 — the paid call takes the shape of 44 of 56 checklists against 42, and raises 24 extra steps against 31; THE BAR WINS THE OTHER DIRECTION, 2 steps not reported against 6. And it loses outright on false gaps: 11 step rows the records genuinely cover called unevidenced, against the bar's 0. Both constants also score 0 false gaps, by reading nothing at all, so that cell can never reach a card on its own. It loses on every other cell phase 1 called already free: effectivity 55 of 56 against 56 of 56, decision 52 against 60, route 52 against 53, step rows 283 against 302, gap recall 33 of 37 against 36, holds 7 of 8 against 8. This kit ships NO paired significance test — the margins are counts and are quoted as counts.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A change desk whose packets carry a ticked class box that is often wrong — this kit, with the station, beside the free bar
    the class reading is the one cell it measurably buys: 47 of 60 against the bar's 44, and the shape follows it at 44 of 56 against 42.
  • A desk whose effectivity clauses are written in the standard's own words — the free bar alone
    the bar reads effectivity 56 of 56 for $0.00 and the paid call 55. There is no headroom to buy.
  • You care most about never sending an engineer after a record that is on file — the free bar
    false gaps 11 for the paid call against 0 for the bar, over 267 step rows the records cover.

And where nothing here is good enough:

  • You want the checklist the model itself drafted, sent as drafted — neither — run the station
    the contract returns no checklist at all. The arm returns three readings; every step, state, decision and route on every arm is free code's.

At a glanceHow the whole thing runs

79%checklists shape exact pct
2,101 msp50, end to end
$0.79per 1,000 engineering change order packets · GPT-5.6 Luna

Run once, for real, on 2026-09-21. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own rendered change packets and write the reader in src/packet.py for their layout, put your own standard in data/policy.md and data/policy.json, and run evals/run.py with a new run id. The boundary is the ANSWER KEY, not the packets. Corpus lens →
When is this the wrong choice?Avoid: Trusting its covering-records map — exactly right on only 24 of 64 packets; the station's step-table union is what rescues the shape. That is the case against the best-fitting scenario (“A change desk whose packets carry a ticked class box that is often wrong”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet whose drawing does not print its controlled characteristics as a list. The whole CLASS-A test is that lookup, and the class reading is where both the free bar and the paid call lose most of what they lose (bar 44 of 60, paid 47 of 60). 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether any of this holds on a real plant's change orders. Every figure is against a key derived from a generator, on an invented standard, an invented plant and invented records; no real change order or cutover checklist was read. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-21 — r001-eco-cutover. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Observed on this repository with no key configured: python3 -m evals.baseline printed all six free arms at $0.00, python3 -m evals.check_labels returned 0 disagreements over 64 packets and 432 re-derivations, python3 -m src.app --selftest reported 0 disagreements over 2,304 cells, and the board rendered with the draft button disabled. No call was made.

A living map of modern AI — kept current every morning