The business caseThe problem this solves
A restaurant's recipe book says how much of an ingredient each menu item takes and the point of sale says how many were sold, so theoretical usage is arithmetic. What it has to be reconciled against is a count sheet and a movement ledger that a store keyed by hand — and on 24 of these 62 items the back office's own actual-usage figure is built on product that never moved: a transfer raised and cancelled, a delivery billed and short-shipped, one truck keyed twice, a quantity column entered in cases. Somebody has to open the file, read the note printed under every row, decide which movements are real, notice whether the recipe was re-specced, and say where the item really stands before anybody investigates it. Opening one item's variance file, multiplying the recipe book out against the sales mix by hand, adding the count sheet up, reading each note under a movement row to decide whether the product really moved, checking the store notes for a re-spec, valuing the difference at the item's unit cost and deciding whether it breaks either investigate threshold.
Audience
A restaurant group's food cost desk working a weekly variance list, and the operations analyst behind it. Whoever decides that this item gets escalated, gets a note, or gets nothing at all this period. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual variance file
The corpus is 62 variance file, 0.13 MB (txt 62). It is generated because it has to be. A real count sheet is a franchisee's own trading record, and the exact shapes this kit measures — a transfer raised against the wrong site, a delivery billed and never received, a district manager asking for a shift to be named — are the rows a restaurant group would least want published. Generating it also makes the key DERIVED rather than written: each file is built as a structure, the panels are rendered from it, and TVA-2026 is applied to the same structure by src/policy.py. There is no second place the answer lives.
The corpus
- The 62 variance filegenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every store is an invented trading name, every item code, spec code, order number and movement reference is arithmetic on the file index, and there is no personal data at all — evals/check_labels.py sweeps all 62 files for five families of identifier on every run and reports 0.
Swap this folder for your own material and the kit is pointed at your variance file. That is the whole change — there is no database to migrate.
==============================================================================
FOOD COST VARIANCE FILE FCV-0001
Store: STR-3100 - Northgate Crossing (invented)
Item: CHK-BRST boneless chicken breast Period: 2026-07-06 to 2026-07-12 Procedure: TVA-2026
==============================================================================
ITEM AND PERIOD AS THE MASTER HOLDS IT
stock unit lb
unit cost 3.28 per lb
opening count 132.000 lb
closing count 76.467 lb
tolerance pct 1.00 pct of theoretical
threshold cost 75.00 or more
threshold pct 3.00 pct or more
period full-week
RECIPE BOOK AS PRINTED
SPEC MENU ITEM PER PORTION UNITS SOLD
RS-4000 Crispy Chicken Sandwich 0.250 2220
RS-4013 Grilled Chicken Wrap 0.180 2510
MOVEMENT LEDGER AS POSTED BY THE BACK OFFICE
MOVEMENT DATE QTY TYPE STATUS REF MEMO
MOV-0100 2026-07-06 320.000 purchase POSTED OD-4001 8 cs at 40.000 lb per cs, order OD-4001
MOV-0101 2026-07-07 320.000 purchase POSTED OD-4002 8 cs at 40.000 lb per cs, order OD-4002
MOV-0102 2026-07-12 320.000 purchase POSTED OD-4000 8 cs at 40.000 lb per cs, order OD-4000
VARIANCE AS THE BACK OFFICE REPORTS IT
theoretical usage 1006.800 lb
actual usage 1015.533 lb
variance 8.733 lb
variance cost 28.64
status IN-LINE
Abridged — the file continues.
The outcomeWhat a good result looks like
One item in, one row out: which movements this procedure treats differently from the ledger that printed them, whether a recipe spec change was in force, and one verdict from a closed set of five — COUNT-BROKEN, ESCALATE, OVERUSE, UNDERUSE or IN-LINE. Actual usage, theoretical usage, the variance and its valuation are re-derived in pure code from those readings and the item master, so they follow TVA-2026 whatever the reply said.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed and nothing stopped at the ceiling, so there is no unparsed row to report. What it gets WRONG is published by name: 5 items came back carrying product that was cancelled, short-shipped or keyed twice (counted_movement_that_never_moved), and the pure-code station cannot see it — the row is real, the date parses, the quantity is a whole number of thousandths, and the item master knows nothing about any movement.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your movement ledger's bad rows announce themselves — every cancelled transfer says "cancelled", every duplicate says "duplicate" — the free rules floor, and do not buy a call at all
Across the 24 items where the back office's own figure is wrong the two arms TIE at 14 each. A keyword list over the notes reaches everything a keyword can reach, for $0.00. - Your store notes routinely describe a credit, a short-ship or a cancellation OF SOMETHING ELSE — an earlier delivery, another item on the same invoice — the paid call, and read the keep-note family's own rate first
The floor gets that family 1 of 14 because a keyword rule cannot tell which row a sentence is about. The paid call gets 13. - Your recipe specs change mid-quarter and the change is recorded in prose beside the ones that were costed and refused — the paid call
8 of 11 refused re-specs for the paid arm against 0 for the regex. Every refused wording carries a real effective date inside the period, a real spec code from that file's own book and a real quantity, so a date test and a shape test both pass and only the sentence says no. - You need the counted usage figure and nothing else — either arm — and derive the figure in code from whichever reading you have
The arm's OWN subtraction is right on 25 of 62 and the free floor's on 41. Put either arm's cited set through thirty lines of integer arithmetic and the paid one gives 52. The arithmetic is not what a model is for.
And where nothing here is good enough:
- Somebody downstream acts on an ESCALATE without a second reader — neither, until that changes
Both arms return a clean verdict on items TVA-2026 flags — 5 of 62 on the paid run. The kit escalates an ITEM; it does not decide what happens to it, and it is not built to be the last reader.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own variance files in the same shape and data/stores.json with your own item master, then rebuild the key by labelling them. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per item for a regex you could write in an afternoon. That is the case against the best-fitting scenario (“Your movement ledger's bad rows announce themselves — every cancelled transfer says "cancelled", every duplicate says "duplicate"”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A ledger that is not fixed-width columns. src/countsheet.py's row regex is the shape these files print; a CSV or a back-office API export needs a different parser and nothing above it changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-08 — r001-food-cost-variance. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors, all six committed runs and every screenshot. python3 -m evals.check_labels and python3 tools/build_corpus.py --check both run on a machine with nothing installed.






