The business caseThe problem this solves
Across a gaming day one patron transacts at several windows -- the cage, a table buy-in, a slot ticket purchase, a ticket redemption, the poker room -- sometimes on a player card, sometimes under a name a clerk wrote down as they heard it, sometimes with no identification at all. A currency-reporting programme has to decide which transactions aggregate to the same patron within the gaming day, total cash in and cash out SEPARATELY, and say whether that patron crossed the day's reporting threshold -- and if it did, which transactions belong to it. A cage system that groups on the player-card number alone gets the carded transactions right and is confidently wrong about everyone who also transacted without a card: split across three unlinked identities, each leg reads as under the threshold, and that is the exact failure a currency-reporting programme exists to prevent. A cage system's automatic card-based aggregation, which this kit's own free floor (card-only grouping, arithmetic, a threshold compare) reproduces for nothing -- what it adds is reading the floor notes and resolving the name variants a cage system does not.
Audience
A compliance officer deciding whether a currency report is required for a gaming day, and which transactions belong to it -- and the person who has to justify that call if it is ever reviewed. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual gaming-day transaction-log packs
The corpus is 42 gaming-day transaction-log packs, 0.21 MB (txt 42). Synthetic and generated rather than sourced, because a real answer key here would mean publishing statements about named patrons' cash transactions in a public MIT repository, permanently, with no way to correct one that was wrong. Building the key first and printing the pack from it is what lets evals/check_labels.py re-derive every structured fact from the rendered document and gate the key before any arm is scored against it -- 20 properties, 0 violations, red-proofed against 7 seeded defects.
The corpus
- The 42 gaming-day transaction-log packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your gaming-day transaction-log packs. That is the whole change — there is no database to migrate.
Gaming Day
----------------------------------------------------------------
Gaming day GD-0001
Property Ambermoor Casino
Licence LIC-8041
Gaming day opens 2026-06-23 06:00
Gaming day closes 2026-06-24 05:59
Currency unit CU
Prepared by Cage shift supervisor, end of gaming day
Reporting Rulebook
----------------------------------------------------------------
RB-1 A reportable aggregate is the set of transactions made by ONE patron
within ONE gaming day at this property.
RB-2 Cash in and cash out are totalled SEPARATELY. They are never netted
against one another and a patron may cross on either or on both.
RB-3 The reporting threshold for this gaming day is 8,000.00 CU. An aggregate is
reportable when a total is ABOVE the threshold. A total that lands
exactly on the threshold is not above it and is not reportable.
RB-4 A transaction timestamped outside the gaming day window belongs to the
adjoining gaming day. It is not aggregated here.
RB-5 Identification. A player-card number is conclusive. A name written at
the window identifies the patron where it resolves to one patron and
nothing in the record contradicts it. A transaction presented with
neither cannot be attributed from this pack.
RB-6 Where any transaction inside the window cannot be attributed, the
gaming day is HELD: the aggregates that CAN be closed are still stated,
and the day's determination waits on the unattributed transaction.
Patron Register
----------------------------------------------------------------
card name on the account enrolled
PC-426572 Liorelm Haldtellar 2024-11-17Abridged — the file continues.
The outcomeWhat a good result looks like
A drafted aggregation: every patron who transacted inside the gaming-day window, their cash-in and cash-out totals kept apart, whether they crossed the threshold, and which transactions the finding rests on -- with the net shown but never compared to anything, because netting the two directions is the trap this kit is built to catch.
And when it cannot
A day the kit gets wrong is UNDER-attribution, never over-merging: on the 742-of-764 scored run, over_aggregation_pct is 0.0 (0 of 729 attributable cells -- no two strangers were ever merged) and under_aggregation_pct is 1.68 pct (12 of 714). Every one of the twelve is a transaction the model left UNATTRIBUTED that the answer key says belonged to a specific patron -- a floor note it read too literally, never a patron it invented or a stranger it folded in. The failure direction that would put a real name on a report that should not have one (false_report_pct, contaminated_report_pct) is 0.0 pct on this run.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your cage already captures the player card on nearly every transaction, and the rare uncarded transaction is either genuinely unidentifiable or settled by an exact name match at the window. — the free floor -- day-gate, $0.00
It scores 81.54 pct on the discriminator (742-of-764's floor equivalent), catches 86.13 pct of channel-structured transactions, and reaches 73.20 pct of reportable aggregates -- all in under a second, with no key. Sweep src/checks.NAME_MATCH against your own name conventions first (evals/tolerance.py) -- it costs nothing and moves the floor's own score by 23.29 points across six modes. - Your floor notes carry real identifying information -- a supervisor's recollection, a photograph match, a note that separates two patrons a name matcher would otherwise merge -- and getting the prose-only cases right is worth the per-day cost. — the paid arm -- the fast tier at $0.041664 a gaming day
It reaches 62.86 pct on the prose channel against the floor's 0.0 pct, and its over_aggregation_pct is 0.0 across all 729 attributable transactions -- it never merges a stranger into an aggregate, which is the harm that can name an innocent patron on a report. Its false_report_pct and contaminated_report_pct are both 0.0 on this run.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point tools/build_corpus.py's own generation parameters at your own case mix and threshold values to reshape the corpus without touching src/, or replace data/corpus/*.txt and data/gold.jsonl with your own gaming-day packs and labels in the same shape (see check_labels.py's twenty properties for what 'the same shape' means precisely) and point src/day.py's regular expressions at your own layout. Every figure on this kit's pages is a property of THIS synthetic corpus and its invented rulebook -- the 97.12 pct headline says nothing about accuracy against a real cage's transaction log, a real jurisdiction's reporting threshold, or a real property's floor-note conventions, none of which this run ever saw. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not use it where identity depends on a sentence rather than a card or a clean name match. It cannot read the Floor Notes at all -- 0.0 pct on the prose channel, by construction -- and it is the reason its day-level accuracy (64.29 pct) trails the paid arm's by 16.66 points even though its per-transaction score is closer. That is the case against the best-fitting scenario (“Your cage already captures the player card on nearly every transaction, and the rare uncarded transaction is either genuinely unidentifiable or settled by an exact name match at the window.”). 2 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A real cage's transaction-log layout. src/day.py reads fixed-offset regular expressions against THIS generator's own layout -- point it at a real system's export and it returns empty tables rather than guessing, which is the safe failure and also means it settles nothing until a forker writes their own parser. 4 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Repeatability. Every arm -- the paid run and all three free floors -- was fired ONCE. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-26 — r001-ctr-aggregate. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no API_KEY renders every gaming day, the free floor and the answer key for $0.00 -- only the model column stays empty until a key is set.



