Home › Use Cases › Stand sales vs inventory reconciliation
Use caseUC0451
🧪 Use-case kit · runnable

Stand sales vs inventory reconciliation

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A concession stand at a venue closes an event with four records that have to agree: the count sheet (opening, closing, recounts), the transfer log as the stand kept it and as the commissary or the next stand kept it, the waste log, and POS unit sales — which only become inventory through the recipes that say a hot dog is one frank and one bun, with each recipe's version and effective date. The stand's own inventory-system variance report reads those records' COLUMNS. What decides the answer is often a sentence: a recount abandoned half way, a close count nobody took, a transfer cancelled at the dock or sent back unopened, a new build withdrawn before the event. The report disagrees with the procedure on 42 of the 64 packs in this corpus. Opening one stand's event pack, deciding which count rows stand after every recount and note, pairing the stand's transfer log against the other side's and reading which transfers really moved, putting the right recipe version in force for the event date, exploding POS sales into inventory, and testing every item's variance against its class tolerance.

Audience

The concessions finance desk and inventory control at a venue operator, working an event's stands before the numbers go into commission and settlement — and the commissary and POS admin desks an exception is routed to. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual stand-event packs

The corpus is 64 stand-event packs, 0.27 MB (txt 64). It is generated because it has to be. A real venue's stand records are an operator's own commercial data: named cashiers on every shift, supplier-branded products, negotiated recipes, and notes written by named people about named people — and a stand variance is exactly the kind of number that gets pinned on a cashier, which is this kit's first refusal. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides a reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.

The corpus

  • The 64 stand-event packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every venue is a code and a description, every stand a concourse code, every item a generic description under an invented SKU, and every count row, transfer, recipe, waste row and POS sale is invented. THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note speaks for inventory control, the stand lead's desk, the commissary, POS admin, the concessions desk or venue finance. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your stand-event packs. That is the whole change — there is no database to migrate.

One stand-event pack, as the model receives itSDR-0001.txt · 1 of 64
================================================================================
CONCESSION STAND RECONCILIATION -- ONE STAND, ONE EVENT
================================================================================
FILE                SDR-0001
VENUE               VN-2  an outdoor ground, 41,000 seats
EVENT               EV-2600  cup tie
EVENT DATE          2026-08-19
STAND               ST-C13  concourse C, stand 13
COUNT SHEET OPENED  15:51
COUNT SHEET CLOSED  23:45
PROCEDURE           SCR-2026
PACK COMPILED       2026-08-21

-- ITEMS STOCKED AT THIS STAND (inventory master) ------------------------------
SKU       DESCRIPTION             CLASS      TOLERANCE
SKU-104   SODA CUP 32OZ           PACKAGING          6
SKU-203   SOFT PRETZEL            PORTION            4
SKU-204   CHEESE CUP 3OZ          PORTION            4
SKU-205   NACHO TRAY              PACKAGING          6
SKU-301   POPCORN BAG             BAGGED             3
SKU-303   CANDY BAR               BAGGED             3

-- RECIPES (how each POS item depletes inventory) ------------------------------
ROW      POS    POS NAME              SKU      QTY  EFFECTIVE   VERSION
RX-0101  P-113  FOUNTAIN SODA         SKU-104    1  2026-03-01        1
RX-0102  P-121  PRETZEL               SKU-203    1  2026-03-01        1
RX-0103  P-122  PRETZEL WITH CHEESE   SKU-203    1  2026-03-01        1
RX-0104  P-122  PRETZEL WITH CHEESE   SKU-204    1  2026-03-01        1
RX-0105  P-123  NACHOS                SKU-205    1  2026-03-01        1
RX-0106  P-123  NACHOS                SKU-204    1  2026-03-01        1
RX-0107  P-130  POPCORN               SKU-301    1  2026-03-01        1
RX-0108  P-132  CANDY                 SKU-303    1  2026-03-01        1

Abridged — the file continues.

The outcomeWhat a good result looks like

One stand-event pack in, one row out: which count rows stand, which transfers moved stock, which recipe rows were in force, every item's variance against its class tolerance, every exception and one of five SCR-2026 verdicts. 53 of 64 packs come back with all six graded fields right, against 47 for the best free floor and 0 for the stand's own variance report — and, the part that is significant, 1 of 21 real shortfalls explained away against the best free floor's 10.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 11 packs it got wrong are named in the kit README with what it answered; only 3 move a derived field — one shortfall explained away (SDR-0029) and two false alarms (SDR-0003, SDR-0022) — and the other 8 are a reading graded wrong on a pack whose items, exceptions and verdict the station still gets right. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your stand system records a voided recount, a count not taken, a cancelled or reversed transfer and a withdrawn recipe as a STATUS, and both sides log every transfer — the free modal floor, and do not buy a call at all
    On the 37 packs the printed columns decide, the modal floor gets 37 and the paid call 31. Recounts that stand, missing counts, one-sided transfers, unmapped POS items, recipe effective dates and variance at tolerance are all decidable from columns and dates.
  • You want whole packs right and a stand closing wrongly as RECONCILED is no worse than a false alarm — the free tie rule
    47 of 64 for $0.00 against the paid call's 53, and the difference is not significant (p = 0.286). It never raises a false alarm.
  • A shortfall closed as reconciled is the expensive mistake — commission or settlement runs on it — the paid call
    This is the whole product. It explains away 1 of 21 real shortfalls against the tie rule's 10 and calls 0 of 34 exception stands RECONCILED against 11, both significant.
  • Notes about recounts, transfers and recipe builds arrive as free text from the stand lead, the commissary or POS admin — the paid call
    On the 27 packs where a note decides a reading the paid arm is 22, the tie rule 13 and the modal floor 3; on the 10 that hide a shortfall it is 9, against 7 for the best free floor and 0 for the tie rule.
  • You want the stand's own variance report audited — either paid or free — both beat it comprehensively
    The report disagrees with SCR-2026 on 42 of 64 packs and gets 0 whole: it returns no reading, and its POS map holds one item per button, so it flags 56 of 64 OVER-TOLERANCE. It is published as an arm so the comparison is against what is running today rather than against nothing.

And where nothing here is good enough:

  • You want a season view of which stands show a real pattern — neither, yet
    One pack is one stand at one event. No roll-up across events is computed or scored, and a review threshold is an operator's choice SCR-2026 does not make.

At a glanceHow the whole thing runs

83%rechecked all correct pct
1,931 msp50, end to end
$1.36per 1,000 stand-event packs · the fast tier

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own stand-event packs in the same block shape and data/stands.json with your own stand register, then run python3 -m evals.run --run-id b000-<yours>-tie --floor tie — it needs no key and makes no call. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per stand for arithmetic you already have. That is the case against the best-fitting scenario (“Your stand system records a voided recount, a count not taken, a cancelled or reversed transfer and a withdrawn recipe as a STATUS, and both sides log every transfer”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A POS item that depletes by weight or by pour — a draught line, a cheese pump, a bulk popcorn kettle. The whole reduction is recipe explosion in whole units; expected usage that is not an integer breaks every class tolerance. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?THAT IT BEATS THE BEST FREE RULE ON WHOLE PACKS. 53 of 64 against the tie rule's 47 is 14 packs only the paid call gets whole and 8 only the rule gets whole, McNemar exact p = 0.286 — not significant. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-stand-depletion. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — What we reproduced from a clean checkout with no key configured: the board started on a scratch port with API_KEY blank, every route rendered with 0 provider-name hits, the five free floors ran through the station on 13 packs with 0 mismatches against their committed results, and Ask the model answered 400 no key. requirements.txt lists nothing — the kit is standard library only. What we could not reproduce without a key: a new scored run.

A living map of modern AI — kept current every morning