Home › Use Cases › Code a shop-floor downtime event onto the plant's own reason card
Use caseUC0344
🧪 Use-case kit · runnable

Code a shop-floor downtime event onto the plant's own reason card

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

Every morning somebody works yesterday's downtime queue: one MES event at a time, deciding what actually happened and where the lost minutes belong. The record in front of them carries a note an operator typed between two other jobs, and on a large share of events that note says the line was waiting on material - it sits on 28 of these 62 events, across 7 different reason codes. Sometimes it is right and the upstream asset is recorded stopped. Sometimes the infeed buffer is printed at forty per cent and the machine's own fault alarm is two lines further down. The two look identical in the note, and coding the second as a starve moves the minutes off this line onto the asset that feeds it, leaves the fault unworked, and the same stop happens again next week. The first pass over a downtime queue. It replaces nothing after that: no OEE figure is touched, no work order is opened or closed, no root cause is stated and no corrective action is proposed.

Audience

A production or reliability lead who owns a line's downtime coding, and the shift supervisor who has to defend a code at the morning meeting. Not a planner, not a maintenance planner, and not anybody who needs an OEE number changed - this kit produces a coded row and nothing downstream of it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual downtime event records

The corpus is 62 downtime event records, 0.09 MB (txt 62). A real downtime record is a joined extract of an MES event table, a controller's alarm historian and a shift log, and all three carry names and a work-order trail that identifies a plant. None of it can ship. So the corpus is built from structures and the key is the card applied to those same structures - and that is what makes the key derivable rather than written, the free floors real arms rather than strawmen, and every committed run re-scorable at zero cost.

The corpus

  • The 62 downtime event recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your downtime event records. That is the whole change — there is no database to migrate.

One downtime event record, as the model receives itDTE-0001.txt · 1 of 62
DOWNTIME EVENT RECORD                                   DTE-0001
Prepared 2026-09-09 under DRC-2026 | production period 2026-07-01 to 2026-09-30

EVENT FACTS
  Event kind                              stop
  Line                                    L-02 Fill and Seal
  Asset                                   AS-2220 rinser
  Asset register reference                AR-2220
  Shift                                   A
  Event began                             2026-09-07 16:15
  Event ended                             2026-09-07 16:33
  Duration                                18 minutes
  Production order open at the event      yes - production order WO-41822

OPERATOR NOTE
  stopped, see shift log

MACHINE LOG
  16:15:00  STATE    RUN -> STOP
  16:15:04  ALARM    J-130 short stop, product jam at the discharge guide
  16:33:30  STATE    STOP -> RUN

LINE CONTEXT
  Upstream asset state at the event       running
  Downstream asset state at the event     running
  Infeed buffer at the event              28 pct
  Outfeed buffer at the event             37 pct

QUALITY AND RATE
  Scrap recorded during the event         0 units (the card's threshold is 25)
  Cycle rate during the event             at standard

EVENT AS LOGGED BY THE MES
  MES event code                          E-22
  MES suggested reason                    Asset stopped, cause not determined by the controller.

SHIFT NOTES
  Shift-board reminder from the coding procedure: a stop is only coded as a material starve where the record shows the infeed empty or the upstream asset stopped -- what somebody typed at the machine is not the test.
  The event was read against the MES extract on file.
  Worked in the ordinary daily coding queue and no exception was raised.

The outcomeWhat a good result looks like

A coded row somebody can work: one of eleven reason codes, the line of the record that establishes it, the account the minutes go on, whether the loss fell on the asset that sets the line's rate, which of the six big losses it reduces to, and which queue it goes to next.

And when it cannot

The row is wrong in one of two directions and each costs something different. A stop that was the machine's own, coded as a starve or a block, sends the minutes to a neighbouring asset and leaves a recurring fault nobody is looking at. An unplanned stop coded as planned time leaves the loss columns entirely, where nobody sees it again. Both are counted on their own denominators here and neither is ever folded into an accuracy figure.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the charge and the constraint flag and nothing else — the free rules table plus the engine - evals/baseline.py mode rules
    53 of 62 on the charge and 60 of 62 on the flag for $0.00, because both are lookups on a code and the table's code is right 52 times out of 62.
  • You want the reason code itself, which is the thing a coder acts on — the paid call
    59 of 62 against the table's 52, and paired over the same events it is right where the table is wrong on 8 and wrong where the table is right on 1 - McNemar's exact two-sided p = 0.0391.
  • You want all four fields on the row and you are counting cost — the free rules table, and measure before you buy anything
    it takes 51 of 62 four-field rows against the paid arm's 48, for $0.00, and the paired test cannot separate them (p = 0.6291).
  • Your plant's MES extracts are complete on every event — the free rules table
    every one of the table's ten misses is an event whose log is thin or missing. On a corpus with none, the table and the call would be much closer.
  • You want to know whether reasoning should be on — OFF, and then measure it yourself
    this kit ran 72 calls with the provider's disable shape sent and no reasoning tokens reported, at $0.000520 per event and a largest reply of 169 tokens.
  • You need an OEE figure produced, adjusted or defended — something else entirely
    the answer contract has no field that could express one and src/prompt.py fails at import if a future edit adds one; the card's section 2 refuses it in the model's first paragraph.

At a glanceHow the whole thing runs

77%rechecked record all correct pct
1,572 msp50, end to end
$1.85per 1,000 downtime event records · openai/gpt-5-6-luna

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own events into data/corpus/*.txt in the same seven panels, put one row per event in data/assets.json (the id, the register reference, the line, the asset, its position and whether it is the line's constraint) and label them in data/gold.jsonl. What stops being true the moment you do: every number this kit publishes was measured on a corpus this kit generated, against a card this kit invented. Corpus lens →
When is this the wrong choice?Avoid: Paying for it. And avoid reading the table's flag score as competence - the MODAL arm, which reads nothing at all, gets 56 of 62 on the same field, and the asset register alone gets 56. That is the case against the best-fitting scenario (“You want the charge and the constraint flag and nothing else”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A plant with ninety reason codes in a three-level tree. Eleven codes fit in a prompt; a real reason tree does not, and at that size the card becomes a retrieval problem rather than a prompt part. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the citation gap is a defect in the arm or in the key. Seven records establish the same fact twice, section 4 of the card permits either line, and the key labels one. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, as answered, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-downtime-code. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clone with no key renders the whole board, scores both free floors live, replays the committed scored run for $0.00 and re-derives the answer key from scratch. python3 tools/build_corpus.py --check rebuilds all 62 records byte-identically; python3 evals/check_labels.py re-derives every gold row with the rulebook retyped. Neither needs a network.

A living map of modern AI — kept current every morning