Home › Use Cases › Blend and rework genealogy reconciliation
Use caseUC0435
🧪 Use-case kit · runnable

Blend and rework genealogy reconciliation

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A specialty-chemicals blending site fills an output lot from several input lots — prime material and reworked material — and before QA can release it four records have to agree about what went in: the recorded genealogy, the batch ticket's charges at the vessel, the inventory ledger's movements against the batch order, and the input lot register with each lot's hold status and origins back to supplier receipts. The desk's MES check reads those records' COLUMNS — a STATUS printed on an export, a lot type, a posted quantity. What decides the answer is often a sentence: an issue the stores desk reversed, a pick keyed against the wrong batch order, a quarantine that spanned the charge date, a pending issue that posted before AS AT, a lot the lab re-graded. The desk's own status line disagrees with the procedure on 27 of the 64 genealogy packs in this corpus. Opening one genealogy pack, netting each input lot's standing issues and returns at the AS AT date, reading each note to see whether it reverses, re-keys, posts or holds something, checking every hold against the date the lot was charged, walking each lot's origins back to receipts, and testing the mass balance and the rework share against the blend instruction.

Audience

The QA genealogy desk and the stores desk at a blending site, reconciling a day's blended and reworked output lots before they reach the release authority — and the lab and blend floor that own the queries an exception raises. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual genealogy packs

The corpus is 64 genealogy packs, 0.23 MB (txt 64). It is generated because it has to be. A real blend genealogy is a site's formulation: component identities that are trade secrets, supplier receipts, customer orders, and QA and stores notes written by named people about named people. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, masked and generated from one seed with the key DERIVED by the same rulebook the kit applies.

The corpus

  • The 64 genealogy packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every site, product, lot, receipt and batch order is an invented code, every component is a masked code (the formulation is a trade secret and the corpus never had one), and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note speaks for the QA desk, the stores desk, the blend floor or the lab. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your genealogy packs. That is the whole change — there is no database to migrate.

One genealogy pack, as the model receives itBGP-0001.txt · 1 of 64
================================================================================
BLEND AND REWORK GENEALOGY RECONCILIATION -- ONE OUTPUT LOT, ONE AS-AT DATE
================================================================================
FILE              BGP-0001
SITE              SITE-07 blending hall
PRODUCT           PRD-4410  an industrial degreaser blend
OUTPUT LOT        LOT-600006
BATCH ORDER       BO-71000
BATCH TICKET      BT-81000
LEDGER EXPORTED   2026-06-03
AS AT             2026-06-05
PACK COMPILED     2026-06-08
COMPOSITION       masked -- component codes only
QUANTITY UNIT     kg

-- BLEND INSTRUCTION LIMITS AND THE FILL RECORD --------------------------------
OUTPUT QUANTITY FILLED               5,736.2
MASS BALANCE TOLERANCE                0.50 %
REWORK LIMIT                         10.00 %
LINK TOLERANCE                           0.5

-- RECORDED GENEALOGY (the genealogy record for the output lot) ----------------
ROW   INPUT LOT   COMPONENT   RECORDED KG
G1    LOT-600012  CMP-231         3,000.0
G2    LOT-600014  CMP-362           600.1
G3    LOT-600023  CMP-886           143.0
G4    LOT-600039  CMP-100         2,000.0
RECORDED TOTAL                    5,743.1

-- BATCH TICKET CHARGES (as ticketed at the vessel) ----------------------------
ROW   CHARGE DATE  INPUT LOT   COMPONENT    CHARGED KG
C1    2026-06-01   LOT-600039  CMP-100         2,000.0
C2    2026-06-01   LOT-600012  CMP-231         2,999.8
C3    2026-06-01   LOT-600014  CMP-362           600.1
C4    2026-06-01   LOT-600023  CMP-886           143.0

-- INVENTORY MOVEMENTS (ledger extract for these lots, exported 2026-06-03) ----
MOVEMENT   POSTED ON   INPUT LOT   TYPE    ORDER             KG  STATUS
MV-500014  2026-05-29  LOT-600039  ISSUE   BO-69000       200.0  POSTED

Abridged — the file continues.

The outcomeWhat a good result looks like

One genealogy pack in, one row out: which movements stand at AS AT, which charged lots were held at charge, which inputs are rework, every link flag, every chain's state, the consumed total to a tenth of a kilogram, the mass balance and rework tests, every exception and one of six BGR-2026 verdicts. 55 of 64 packs come back with all six graded fields right, against 43 for the best free floor and 0 for the QA desk's own status line.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 9 packs it got wrong are named in the kit README with what it answered, and every one is a reading: 3 holds that span the charge date left out — RECONCILED on a lot that took in a held input — 2 lots printed ON HOLD after their charge counted as held, 3 postings on or before AS AT left out, and 1 movement dropped after a same-day cancelled reversal. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your ledger records a reversal, a mis-keyed issue and a late posting as a STATUS, and your quality system stamps each hold with its start and end dates — the free floors, and do not buy a call at all
    42 of 64 packs for $0.00 from the columns alone. Links, chains, limits, returns and printed holds are all decidable from columns and dates, and on the 37 packs the columns decide the modal floor is whole on all 37 — the paid call on 35.
  • Reversals, mis-keyed picks and re-grades arrive as free text — a stores desk note, a lab note pasted into the pack — the paid call
    This is the whole product. On the 27 packs where a sentence decides a reading the paid arm is 20, the modal floor 5 and the best vocabulary floor 13; reversed issues 6 of 6 against 1, mis-keyed issues 5 of 5 against 3.
  • You want the QA desk's own MES genealogy check audited — either paid or free — both beat it comprehensively
    The desk's status line disagrees with BGR-2026 on 27 of 64 packs and gets 0 whole: it returns no reading at all. It is published as an arm so the comparison is against what is running today rather than against nothing.

And where nothing here is good enough:

  • You want every blend that took in a held lot kept OUT of release above all — neither alone — read the hold dates in code beside the paid call
    The paid arm let 3 held lots through as RECONCILED (BGP-0001, BGP-0005, BGP-0046 — holds spanning the charge date) and counted 2 lots that were not held (BGP-0026, BGP-0063). A date comparison between a hold note and the CHARGE DATE, which is not built, recovers all five.
  • Pending issues routinely post after the ledger export and before the reconciliation date — neither alone — re-export the ledger at AS AT
    Both vocabulary floors beat the paid call here, 3 of 5 against 1: it admitted none of the three postings dated on or before AS AT. A fresh extract makes the reading a column again.
  • Your blends charge a lot in several additions, across shifts or into more than one vessel — neither, yet
    No file in this corpus does. The unit of work is one charge row per lot per ticket joined by LOT id, and every percentage on this page is against that unit.

At a glanceHow the whole thing runs

86%rechecked all correct pct
1,648 msp50, end to end
$1.48per 1,000 genealogy packs · the fast tier

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own genealogy packs in the same nine-block shape and data/lots.json with your own lot register, then run python3 -m evals.run --run-id b000-<yours>-vocab_neg --floor vocab_neg — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per genealogy pack for arithmetic you already have. That is the case against the best-fitting scenario (“Your ledger records a reversal, a mis-keyed issue and a late posting as a STATUS, and your quality system stamps each hold with its start and end dates”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A blend that charges one lot in several additions, across shifts or into more than one vessel. The whole reduction is one charge row per lot per ticket joined by LOT id; the join does not survive it. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-blend-genealogy. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — The board served with no key configured — a second server with API_KEY blanked in its own environment — rendered every committed run and ran all four free floors on every pack with no network call, and the scored run re-scored from its cache for $0.00. requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning