Home › Use Cases › Reconcile an event-day ancillary stream against counted entries and authorisations
Use caseUC0503
🧪 Use-case kit · runnable

Reconcile an event-day ancillary stream against counted entries and authorisations

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An ancillary stream — a surface lot, a garage, a coat check, an equipment rental desk, a premium-access upgrade — is closed for one event date and somebody has to say whether the revenue is complete before the settlement pack goes forward. The arithmetic is easy and free code owns it: counted entries times the governing rate class equals expected revenue, posted minus expected is the variance to the cent. The hard part is the exception rows — entries that were counted and produced no revenue. Each is either authorised by something on file or it is leakage, and the authority lives in a validation memo, a season-parking manifest line, a permit list or a sentence in the lot log, never in a column. 102 of these 223 rows have an EMPTY cleared by column and 54 more carry one that is wrong. Checking one event-day ancillary stream close against the counted entries and the authorisation file, by eye. For every counted entry that posted no revenue that means deciding which rate-card version governs the event date, and whether a validation memo, a season manifest line, a staff or contractor permit, or an operations manager's event-day override actually covers THIS event date and THIS rate class — or whether the entry is leakage. Not the posting, not the billing, and not the revenue manager who signs the variance off.

Audience

A venue revenue accountant closing an event day, and the revenue manager who decides whether the variance goes forward. The answer this report gives them is NO on the headline: the paid reading is behind a free pass-eligibility join on every slice but one, and behind a free tag reader on that one too. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual event-day ancillary stream close packs

The corpus is 64 event-day ancillary stream close packs, 0.26 MB (json 5 · jsonl 1 · md 2 · txt 64). Because the shape of an ancillary-revenue failure is not the arithmetic — free code gets that perfect — it is whether a counted entry that produced no revenue is AUTHORISED, and the authority is prose. 605 authorisations across 64 packs, written in seven sentence forms over four document kinds; 167 windows pinned to a rate-card version boundary, 82 stated as a count of home dates, 66 as a list of event dates, 66 with no stated end and 30 pointing at another document; 78 scopes that mean 'whatever the card makes pass-eligible', 120 stated as an exclusion, 65 meaning 'whatever the governing version prices' and 64 pointing at another document; 323 long-form dates against 282 ISO. 22 rows where the tag, the scope and the window all match and the card still says no. 88 of the 223 rows cannot be settled from any column at all. ⚠︎ The ceiling arm reads 223 of 223 and 100% on every metric, which is the check item 85 asks for: no field is silently dropping rows, so where an arm fails it is failing to read rather than failing to parse.

The corpus

  • The 64 event-day ancillary stream close packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 64 close packs, data/close_records.json, the rulebook and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.

Swap this folder for your own material and the kit is pointed at your event-day ancillary stream close packs. That is the whole change — there is no database to migrate.

One event-day ancillary stream close pack, as the model receives itANC-0001.txt · 1 of 64
EVENT-DAY ANCILLARY REVENUE CHECK — STREAM CLOSE PACK
Marloway Venue Services, ancillary revenue desk. Prepared for review.
This pack names an ancillary-revenue variance. It does not release, post, bill, credit or
adjust a figure.

CLOSE HEADER
  Close id            ANC-0001
  Stream              STR-462 — surface lot (ancillary, non-ticketed)
  Event               EVT-6269, home fixture 8 of the season
  Event date          2025-11-15
  Rate card           CARD-7000 (operator agreement — this operator's own schedule)
  Close prepared by   the gate supervisor on duty
  Recheck run         ARC-0001

SEASON HOME DATES
  2025-08-09, 2025-08-23, 2025-09-02, 2025-09-19, 2025-10-01, 2025-10-22, 2025-11-04,
    2025-11-15, 2025-12-03, 2025-12-17, 2025-12-27, 2026-01-16, 2026-01-28, 2026-02-13,
    2026-02-24, 2026-03-15, 2026-03-28, 2026-04-12, 2026-04-22, 2026-05-09

RATE CARD ON FILE
  This operator's own rate schedule under the venue operator agreement. These are contract
  terms. No governing authority is cited for any rate, class, pass rule or window below.
    v1  effective 2025-07-15 through 2025-09-18
        priced classes   standard 14.50 | reserved 32.00 | accessible 0.00
        pass-eligible    standard
    v2  effective 2025-09-19 through 2026-02-23
        priced classes   standard 15.50 | oversize 28.50 | reserved 31.00 | accessible 0.00
        pass-eligible    oversize, accessible
    v3  effective 2026-02-24 onwards
        priced classes   standard 15.50 | reserved 34.50 | accessible 0.00 | surcharge 8.00
        pass-eligible    surcharge

COUNTED ENTRIES BY CLASS
  standard       321
  oversize       137
  reserved       215
  accessible     284
  surcharge        1

POSTED REVENUE LINES
  RL-1  standard       321 entries       4,975.50

Abridged — the file continues.

The outcomeWhat a good result looks like

Every exception row carries a disposition from a closed list of four in cascade order — UNRATED-CLASS, DUPLICATE-TAG, AUTHORISED, LEAKAGE — and, where it is AUTHORISED, the id of the ONE document that clears it; every close carries the governing rate-card version and its hold list. A close whose class the governing version does not price is held for rate-card-gap; a close with any unauthorised counted entry is held for leakage.

And when it cannot

⛔ THE PAID READING LOSES TO FREE CODE AND IT IS PUBLISHED AS SUCH. Exception rows wholly correct — the disposition AND the clearing document together — 143 of 223 after the pure-code station against the free eligible_log arm's 167: exact McNemar 25 only-paid against 49 only-floor, p = 0.007084, SIGNIFICANT and running AGAINST the call. Headroom above the bar is 56 ROWS, never a percentage. As answered, before the station, it is 141 — the station is worth 2 rows here and 16 to a fixed reply, so the gap is a difference in READING and not in station. It beats a single fixed reply decisively (143 against 108, p = 0.000082) and the tag-only reader (143 against 92, p = 0.000006), so it is a real reader; it is just worse than the rules.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Closing an event-day ancillary stream where the authorisation file is on the pack and the gate supervisor's cleared by column is what you have — the free eligible_log arm
    167 of 223 exception rows wholly correct for $0.00, against the paid reading's 143 after the station (p = 0.007084 against the call). It also falsely clears ZERO rows where the paid arm falsely clears 21 and names the wrong document on 8 more.
  • The 54 rows whose clearing document expresses its window or scope by pointing at another record — the free tagonly arm, or a person with the pack
    the paid reading gets 23 of the 54 against tagonly's 38 (5/20, p = 0.004077, AGAINST) and columns' 30 (8/15, p = 0.210). The 23-against-0 figure that looks like a win is measured against the floor of record, eligible, domain_strict and the constant — and all four score zero there BY CONSTRUCTION, because none of them ever follows a pointer.
  • The two traps the station decides — an unpriced class and a duplicate tag — any arm at all, including a fixed reply
    16 of 16 for every arm in both columns. R-3 and R-4 are applied by src/ancillary.check_close from the rate-card table and the printed exception table; no arm can win or lose them by reading better.

At a glanceHow the whole thing runs

64%row correct pct
2,001 msp50, end to end
$0.85per 1,000 event-day stream closes · GPT-5.6 Luna

Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?replace data/corpus/*.txt and data/gold.jsonl with your own close packs and key; the rulebook is data/policy.md and the engine reads it as data. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: The paid call on this job. It is behind the floor on the headline, on the 88 reading-required rows and on the 169-row complement of its own predicted slice. That is the case against the best-fitting scenario (“Closing an event-day ancillary stream where the authorisation file is on the pack and the gate supervisor's cleared by column is what you have”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An authorisation whose window or scope is expressed by POINTING at something else — at another document on the pack, or at a rate-card version boundary. 54 of the 223 rows are cleared that way (36 window-relational, 33 scope-relational, 15 both) and they are the rows this kit was built to test. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the declining-phrase test catches a forbidden act phrased in words it does not list. It cannot read intent, the list is in src/refusal.py in full, and a miss says nothing about a refusal — which is why the board names the instrument beside the row rather than leaving a ✗ to be read as a failure to decline. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-18 — r001-ancillary-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on 2026-09-18 on a copy of the kit folder with any .env removed and no key in the environment: python3 tools/build_corpus.py --check rebuilt all 64 packs, the structured records and the key BYTE-IDENTICALLY in 0.09 s; python3 -m evals.check_labels re-derived the key from its own retyped cascade and authorisation reader with 0 disagreements over 223 rows, 605 authorisations and 64 closes in 0.09 s; and python3 -m evals.arms re-expressed all eight free arms in the paid reading contract and reproduced data/floors.json exactly on all six metrics plus the override counts and the card-eligibility verdict in 0.23 s. python3 -m src.app --port 9603 then served the board and every read route answered 200 — the board, the close list, one close, the rules, the prompt, the corpus tab and the pressure tab — with the Read this close now control genuinely disabled for want of a key; a deliberate POST to the one route that can spend returned skipped: true and nothing was billed. ⚑ THE REPLY CACHES SHIP, and the reason is measured rather than assumed: with results/*.jsonl moved away the board is UNCHANGED in every panel, but evals/run.py --resume --rescore resumes 0 of 64 and refuses — so the caches are what make every published number re-derivable for $0.00 on a clone, which on a kit whose headline is a loss is the evidence that matters most.

A living map of modern AI — kept current every morning