Home › Use Cases › Check every required entry in a batch record is present and dated
Use caseUC0270
🧪 Use-case kit · runnable

Check every required entry in a batch record is present and dated

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A batch manufacturing record pack arrives for review with ten steps and up to 38 required entries. Somebody has to confirm every one of them is there and dated before the pack goes on — and the entries that matter most are the ones that are absent, which is the hardest thing on a page to see. The pack that costs the most is not the incomplete one; it is the one where a blank entry sits on the same step as a remark recording an event nobody logged a deviation for, because returning it for completion closes the gap and buries the deviation. Reading 38 required entries off a pack by eye to confirm each is present and dated. It does not replace the QA review, the deviation, or any release decision — the answer contract has no field that could express one.

Audience

A batch record reviewer in quality ops working a release queue, and the QA lead who reads what they produce. The decision this report is for is narrower than it looks: not 'should we buy a model' but 'is this job a parser's job', and on this corpus the answer is yes. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual batch record review packs

The corpus is 65 batch record review packs, 0.17 MB (txt 65). It is generated because it has to be. A real batch record is a regulated manufacturing document containing a customer's formulation, and no site would licence one for a public kit. Generating it also buys the thing that makes the measurement possible: the key is derived from the same structure the text is rendered from, so every count, every finding and every citation span is computed rather than typed, and evals/check_labels.py re-derives all of it from a rulebook retyped from scratch.

The corpus

  • The 65 batch record review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came from⚠︎ NOWHERE — THIS CORPUS WAS NOT FETCHED, SCRAPED, LICENSED OR ADAPTED FROM ANY REAL DOCUMENT, so there is no source URL to map and no publisher whose copyright page could be checked. All 65 packs, the answer key and the statistics are written by tools/build_corpus.py from the fixed seed 20260902, and python3 tools/build_corpus.py --check rebuilds every one of them in memory and diffs the bytes — 67 files, 0 mismatches, under two different PYTHONHASHSEEDs. data/SOURCES.md carries the seed, the licence pointer (MIT, in LICENSE-PUBLIC), the case-family map, and the generator's own attack on itself under "what the generator costs the measurement".

Swap this folder for your own material and the kit is pointed at your batch record review packs. That is the whole change — there is no database to migrate.

One batch record review pack, as the model receives itBRP-0001.txt · 1 of 65
BATCH MANUFACTURING RECORD -- COMPLETENESS REVIEW PACK
Record id: BRP-0001
Product: Resin blend RB-17 (product code PRD-017)
Batch/Lot ID: BR-26-0412
Work order: WO-75002    Master record: MBR-T10 rev 7    Batch size: 2,000.0 kg
Record format: paper-transcribed (typed from the executed pages by the records clerk; transcriber notes in square brackets)
Pages in record: 4    Pages in this pack: 4
Manufactured: 07-Apr-2026 to 08-Apr-2026

--- Page 1 of 4 ---
Section A  Batch header: see above.
Section B  Process steps (performed-by = operator initials or e-signature and date; verified-by = second person, where the step requires it)
Step 1  Line clearance  --  DP 07-Apr-2026 performed  /  EV 07-Apr-2026 verified  /  comment: label roll changed, same lot of labels
Step 2  Dispense raw materials (target 660.0 kg)  --  DP 07-Apr-2026 performed  /  EV 07-Apr-2026 verified  /  balance id: BAL-08  /  net weight: 659.8 kg  /  comment: supervisor present for the step
Step 3  Charge mixer  --  DP 07-Apr-2026 performed  /  EV 07-Apr-2026 verified  /  mixer id: MX-3  /  comment: shift change at 14:00, no impact on step
--- Page 2 of 4 ---
Step 4  Heat to set point (80-90 C)  --  DP 07-Apr-2026 performed  /  temperature recorded: 88 C  /  comment: shift change at 14:00, no impact on step
Step 5  Hold and agitate (45 min)  --  DP 07-Apr-2026 performed  /  comment: supervisor present for the step
Step 6  Add component B  --  DP 07-Apr-2026 performed  /  EV 07-Apr-2026 verified  /  comment: none
--- Page 3 of 4 ---
Step 7  In-process sample  --  AK 08-Apr-2026 performed  /  sample id: IPS-0412-1  /  comment: step completed within the planned time
Step 8  Discharge and filter  --  AK 08-Apr-2026 performed  /  comment: label roll changed, same lot of labels

Abridged — the file continues.

The outcomeWhat a good result looks like

One pack in, five graded answers out: the count of absent entries, the step the first one is on, the finding under BRC-2026, the disposition, and the line that shows it, quoted verbatim and locatable in the pack.

And when it cannot

And what it does when it cannot. On 3 of 65 packs the call named the wrong finding — every one of them a missing-date pack in the HYBRID format, where the page carries a signed-but-undated entry and the audit trail carries a timestamp for it. It read an omission as a contradiction, answered record-conflict, and moved the pack from RETURN-FOR-COMPLETION to HOLD-PENDING-DEVIATION. It found the gap every time: the count, the first step and the quoted line are the key's on all three. The direction is one-way — 3 packs held that were merely clerical, and 0 forwarded that should have been held.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want a completeness review over packs like these — the free rules floor — evals/baseline.py --floor rules
    100.0% of packs fully correct for $0.00 on this corpus, with no wins for the paid arm anywhere against it.
  • Your packs have a layout the parser does not know — write the fifth handler first, then re-measure both arms
    The floor's score is a parser meeting four known layouts. Neither arm was measured on a fifth and neither number transfers.
  • Your remarks are written by people, in phrasings no list anticipated — the paid arm, and measure it — this is the one place it could earn its cost
    G5 is the only judgement in the job. The ablation shows the model ahead by 5 packs when the phrase list is removed, at p=0.23 — a direction, not a proven size.

And where nothing here is good enough:

  • You need a release decision — neither — this kit cannot produce one
    The cap is batch-release and it is enforced in the answer contract: none of the six fields can express a release, an approval, a rejection or a certification.

At a glanceHow the whole thing runs

95%record all correct pct
26,867 msp50, end to end
$18.70per 1,000 batch record review packs · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs, data/policy.json with your own completeness rulebook and data/fields.json with your own required-entry set, then write data/gold.jsonl with one key row per pack. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying for arithmetic and lookup. Presence is what a parser is best at, and on this corpus the paid arm has zero wins against the free one. That is the case against the best-fitting scenario (“You want a completeness review over packs like these”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A scanned or photographed page. Every arm here — the model's reading, both floors and the independent key check — rests on the pack being text. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled line is the line a reviewer would have quoted. The key names ONE per pack and the scorer has no partial credit for a different true line. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-batch-record. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9270 and scores every graded cell offline: the corpus rebuild from seed 20260902, the filler check, the independent label gate, all three free floors over the whole corpus, and a replay of both paid runs from their committed caches. Nothing in that list reaches a network.

A living map of modern AI — kept current every morning