Home › Use Cases › Reconcile a work order's confirmed production against what reached the ledger
Use caseUC0314
🧪 Use-case kit · runnable

Reconcile a work order's confirmed production against what reached the ledger

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

⚠︎ This kit has not been graded

0 completed runs, none scored.

Everything below is coverage, latency and cost. No figure on any of these pages says whether a brief is any good.

The business caseThe problem this solves

A plant confirms what it produced and a ledger records what was posted, and the two do not agree. The difference is three numbers per work order — good, scrap and rework — and the numbers are the least useful part of the answer: a unit that was never converted, a posting raised from the neighbouring order's confirmation, a confirmation that reached the books twice and one that never arrived all produce a gap, and three of them produce the SAME gap. Worse, an order can reconcile to the unit on all three quantities and still be wrong, because a posting sits on an operation its own confirmation does not name. Nothing in the totals can say so. Opening one period's pack, adding up every confirmation for an order while applying whatever a shift note says about the printed figures, converting each posting into the order's base unit, subtracting, and then deciding which of six causes produced the difference — by hand, for every order on the pack.

Audience

A plant controller or cost accountant closing a period — the person who reads the row and decides whether a posting needs investigating. It is not the person who posts the correction, and nothing here posts anything. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reconciliation packs (one plant, one line, one period, five work orders)

The corpus is 60 reconciliation packs (one plant, one line, one period, five work orders), 0.42 MB (txt 60). It is generated because the answer has to be DERIVABLE rather than typed — each pack is built as a structure of true production events, the two extracts are rendered from that structure, and the key is OPR-2026 applied to the same structure, so a change to the corpus cannot leave a stale key behind. ⚑ AND IT WAS ATTACKED BEFORE IT WAS PUBLISHED, WHICH THE FIRST BUILD FAILED. Reading shapes and posting faults were drawn from one mix, so on a reading pack every order tied and the null arm — answer TIES, read nothing — scored 27 of 60, 45 pct. The reading shapes were rewarding the arm that did not read. A pack now carries a posting fault and, independently, a reading shape, on DIFFERENT orders; the null arm now scores 12 of 60 and the column arm 42 of 60, both measured by running them.

The corpus

  • The 60 reconciliation packs (one plant, one line, one period, five work orders)generated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 60 packs, the order register and the whole answer key are generated. Nothing is fetched, scraped, licensed or derived from anything that was, and there is no personal data of any kind: a confirmation as modelled here is three integers, an operation code and a date.

Swap this folder for your own material and the kit is pointed at your reconciliation packs (one plant, one line, one period, five work orders). That is the whole change — there is no database to migrate.

One reconciliation packs (one plant, one line, one period, five work orders), as the model receives itMER-0001.txt · 1 of 60
==============================================================================================
OUTPUT POSTING RECONCILIATION PACK                            MER-0001
Site: Plant 03 - Fern Hollow Works (invented)
Line: Assembly line 2          Period: 2026-P02 (2026-02-02 to 2026-02-27)
Procedure: OPR-2026
==============================================================================================

WORK ORDERS IN THIS PACK
  ORDER       ITEM       BASE  ALTERNATE UNIT        PLANNED
  WO-41000    ITEM-3100  KG    DRUM = 6 KG               279
  WO-41001    ITEM-3103  KG    DRUM = 6 KG               772
  WO-41002    ITEM-3106  EA    CASE = 4 EA               329
  WO-41003    ITEM-3109  M     ROLL = 10 M               718
  WO-41004    ITEM-3112  M     ROLL = 10 M               746

FLOOR SYSTEM CONFIRMATIONS
  What the line's execution system recorded as each operation closed. Quantities are
  in the order's own base unit.
  CONF        ORDER      OP    CLOSED         GOOD  SCRAP  REWORK
  CNF-51001   WO-41000   0010  2026-02-12       84     12       6
      Material issued to the order (not a production quantity) 217 units
  CNF-51002   WO-41000   0020  2026-02-12      132      0       6
      Setup and teardown time (not a production quantity)      2.3 h
      Machine hours (not a production quantity)                3.2 h
      Material issued to the order (not a production quantity) 836 units
  CNF-51003   WO-41000   0030  2026-02-24       48      6       6
      Machine hours (not a production quantity)                7.6 h
      Unplanned downtime (not a production quantity)           30 min
      Material issued to the order (not a production quantity) 750 units
  CNF-51004   WO-41001   0010  2026-02-09      245      9       7

Abridged — the file continues.

The outcomeWhat a good result looks like

One pack in, five reviewer's rows out: for each work order the good, scrap and rework gap in the order's own base unit, signed, one of seven verdicts, and each document behind that verdict with its row copied verbatim from the pack. 15 of 60 packs came back with every field right on every one of their five orders on the scored run; 226 of 300 individual orders were completely right.

And when it cannot

And what it does when it cannot. 60 of 60 calls returned a usable reply, nothing was truncated at the ceiling, and every miss stays in every denominator. The harms are published apart and never averaged: 14 orders the key says are faulty came back TIES — a wrong posting settled as a clean order, which nobody looks at again — and 37 orders that tie were reported as findings, which costs a reviewer a look. The dominant single failure is the unit conversion: of the 7 orders whose true verdict is UOM-ERROR the arm got 1, and it invented UOM-ERROR on 27 orders that tie.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your packs are clean, columnar exports and shift notes never correct a figure — the free column floor alone — evals/baseline.py, --floor rules
    It finds EVERY structural fault on this corpus — 0 missed of 48 — and settles 42 of 60 packs completely for $0.00. A posting-to-confirmation join, a unit conversion and an ordered rule table are things code does not get wrong.
  • Your confirmations carry correction notes, or runs continue onto unnumbered lines — the column floor to ROUTE, and a call only on the packs it flags
    The floor can SEE that a note or a continuation line exists even though it cannot apply it, so the routing rule is free. On this corpus that is 18 packs of 60 — and it is the only set where a call is worth anything.
  • You need the order that ties on every quantity and is still wrong — the column floor — it caught 13 of 13
    WRONG-OPERATION leaves all three gaps at zero, so no total finds it. The operation is a COLUMN, which is exactly what a parser is good at; the paid arm caught 6 of 13.

And where nothing here is good enough:

  • The packs come from a party with a reason to want a variance quiet — neither arm unsupervised — and read the security block first
    An instruction planted in the pack notes deleted a true finding on 5 of the 7 packs it was tried on. The station recovers only the two verdicts it derives from the ledger itself.

At a glanceHow the whole thing runs

18–25%pack all correct
2,545 msp50, end to end
$3.75per 1,000 reconciliation packs · google/gemini-3-flash

Run once, for real, on 2026-09-06. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt and data/orders.json with your own packs and register in the same shapes; src/packet.py's fixed-width layout is four named regexes at the top of that file. You will have no answer key, so nothing can be scored and every accuracy figure on this page stops applying to you — none of them transfers. Corpus lens →
When is this the wrong choice?Avoid: Paying for any of it. It is most of the job and none of the value. That is the case against the best-fitting scenario (“Your packs are clean, columnar exports and shift notes never correct a figure”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack too large for one call. Chunking splits a continuation line from its confirmation and silently loses the units — G-3 has no defence against it. 4 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the cited row is the row a reviewer would have quoted. The key names the documents OPR-2026's cite rule names; a reviewer might also accept the confirmation row beside a posting fault, and this key scores that zero. 4 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-06 — r002-mes-erp-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone, run python3 tools/build_corpus.py, open the board. No key, no install, no index build, no network: requirements.txt names no package, every recorded run ships in results/, and the two free floors, the stub, the key re-derivation and the whole grader run offline.

A living map of modern AI — kept current every morning