Home › Use Cases › Box office vs ticketing sales reconciliation
Use caseUC0442
🧪 Use-case kit · runnable

Box office vs ticketing sales reconciliation

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

After an event the promoter's settlement desk has two counts that should agree price level by price level: the box office desk's certified count (sold, comp, kill, open) and the ticketing system's sales by channel. Between them sits the desk's own log of reconciling items — walk-up batches keyed late, kill orders, pass lists, sponsor allocations — and the desk's own tie-out calling the difference explained. What decides whether an item really explains a gap is usually a SENTENCE and a CLOCK: a batch uploaded before the export, a kill released at 20:31, an allocation bought at face, a memo that never defines walk-up. The desk's own status disagrees with the procedure on 24 of the 64 packs in this corpus, every one of them EXPLAINED over a true gap. Opening one event's tie-out pack, checking every logged walk-up, kill, pass list and allocation against the deal memo's definitions, reading each note to see whether the item was uploaded, released, voided or bought at face, comparing that clock with the ticketing export time, and moving each standing item's tickets level by level to see what difference is left.

Audience

The settlement accountant and the box office desk at a live-events promoter or venue, tying out an event's counts before the settlement is prepared — and the ticketing and production desks an unexplained gap is routed back to. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual event packs

The corpus is 64 event packs, 0.20 MB (txt 64). It is generated because it has to be. A real promoter's tie-out book is commercial record: a named act's deal memo, a venue's manifest, sponsor terms and box office notes written by named people about named people — and a box office gap is exactly the finding that gets pinned on whoever worked the window. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the note and its clock — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.

The corpus

  • The 64 event packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every event, venue, promoter, deal memo, item and note is invented and synthetic, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — an event is a code and a plain description, and every note speaks for a desk. evals/check_labels.py sweeps all 64 files for a person-shaped name on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your event packs. That is the whole change — there is no database to migrate.

One event pack, as the model receives itBOX-0001.txt · 1 of 64
BOX OFFICE TIE-OUT PACK  (BOT-2026)
EVENT             EV-4300  a touring family ice show, evening performance
VENUE             VN-2200  a 7,400-seat venue
PROMOTER          PR-5100  an independent regional promoter
SHOW DATE         2026-08-01
DEAL MEMO         DM-4300
TICKETING EXPORT  2026-08-01 20:40
PACK COMPILED     2026-08-02
CURRENCY          USD

-- DEAL MEMO DEFINITIONS (DM-4300, as signed) ----------------------------------------------
D-1 COMP      Tickets issued at no charge, whether from the promoter's pass list or in the ticketing system under channel COMP, are comps.
D-2 KILL      A kill is a seat taken off sale by a production kill order; it is a kill while the order stands.
D-3 WALK-UP   Window sales on the show date are walk-up and reach the ticketing system when their batch is uploaded.
D-4 CHANNEL   Channel SPONSOR is a non-revenue allocation and counts as a comp. RADIO and ARTIST tickets are bought by the station and the act at face and are sales.

-- PRICE LEVELS (the venue manifest for this event) ----------------------------------------
LEVEL  SECTION                  FACE  MANIFEST
P1     floor                   89.00     1,200
P2     lower bowl              65.00     3,400
P3     upper bowl              42.00     2,800

-- BOX OFFICE COUNT (certified by the box office desk as at the export) --------------------
LEVEL     SOLD    COMP    KILL    OPEN  MANIFEST
P1         866      74       0     260     1,200
P2       2,837      12       0     551     3,400
P3       1,976      58       0     766     2,800
TOTAL    5,679     144       0   1,577     7,400

-- TICKETING SYSTEM EXPORT BY CHANNEL (exported 2026-08-01 20:40) --------------------------
LEVEL   ONLINE   PHONE  OUTLET  WINDOW   GROUP SPONSOR   RADIO  ARTIST    COMP    OPEN  MANIFEST

Abridged — the file continues.

The outcomeWhat a good result looks like

It calls 1 of 32 true-gap levels explained where the best free floor calls 13 (p 0.0005), but gets only 32 of 64 events whole where that same free floor gets 47 (p 0.0167): a trustworthy EXPLAINED, bought with over-flagged gaps. It over-drops items that stand — 30 of 185 fine price levels reported as true gaps and 8 of the 10 date traps misread.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 32 events it got wrong are named in the kit README by case: every re-keyed walk-up batch it dropped (0 of 6 right), every item whose LOG STATUS changed after the export (0 of 4), 8 of the 10 date traps, and 8 events with standing items where it returned none at all. It waved exactly one true gap through — BOX-0009, a kill released before the export that it kept. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the most events tied out whole with no human reading, and your desk re-reads every EXPLAINED anyway — the free memo floor, and do not buy a call at all
    47 of 64 events whole for $0.00 against the paid call's 32 (p 0.0167). It never drops an item the column says stands.
  • An EXPLAINED must be trustworthy before a settlement relies on it, and a person will re-read every true gap — the paid call
    1 of 32 true-gap levels called explained against memo's 13, desk's 24 and modal's 19 (p 0.0005 against memo). The cost is 30 of 185 fine levels flagged for a second look.
  • You want a clean EXPLAINED and cannot pay per call — the vocabulary rule (vocab), with the same caveat
    2 of 32 true-gap levels called explained — not significantly different from the paid call's 1 (p 1.0) — at $0.00. It pays with 21 false TRUE-GAP levels and 39 events whole.
  • You want the box office desk's own tie-out audited — either paid or free — both beat it on the costly direction
    The desk's status disagrees with BOT-2026 on 24 of 64 packs, every one EXPLAINED over a true gap; accepting its log calls 24 of 32 true-gap levels explained.

And where nothing here is good enough:

  • Walk-up batches are routinely re-keyed after a failed upload, or items change after the export — neither the paid call nor a vocabulary rule — export the item log with its timestamps and use the column
    The paid arm is 0 of 6 on re-keyed batches and 0 of 4 on items whose status changed after the export; the column floors get all ten.
  • Your settlements carry a tolerance, split items across levels, or reconcile money rather than tickets — neither, yet
    No pack in this corpus does. The unit of work is whole tickets on one level per item, and every percentage here is against that unit.

At a glanceHow the whole thing runs

50%rechecked all correct pct
1,604 msp50, end to end
$1.14per 1,000 event packs · the fast tier

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own event tie-out packs in the same printed shape, with your own deal memo definitions on each, then run python3 -m evals.run --run-id b000-<yours>-memo --floor memo — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per event for a reading that gets fewer events whole than free code. That is the case against the best-fitting scenario (“You want the most events tied out whole with no human reading, and your desk re-reads every EXPLAINED anyway”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An item that spans more than one price level, or a transfer between levels. R-7 moves an item's tickets on its OWN level only; a split item has nowhere to go. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?ONE SCORED RUN, BOTH HALVES: false EXPLAINED 1 of 32 levels vs the memo floor's 13 (p 0.0005), but whole events 32 of 64 vs its 47 (p 0.0167); run-to-run spread unmeasured. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-boxoffice-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured rendered the whole board, all five free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning