Home › Use Cases › Regenerate every metric feeding the annual risk assessment from its source records
Use caseUC0456
🧪 Use-case kit · runnable

Regenerate every metric feeding the annual risk assessment from its source records

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A gaming operator's annual risk assessment carries an appendix of volume and channel metrics, and every one of them has to be regenerated from the source records before the cycle closes. The pack a data governance desk gets prints the source contribution rows as exported — business area, channel, period bucket, record count, value, source system, extract reference and extract as-of date — and reading the COLUMNS gets a figure. What decides whether that figure can go in the appendix is often a sentence: an extract a note recalls with no replacement, a blank extract column a note supplies, a printed as-of that looks stale until a note gives the true cut date, records recoded out of the channel after the export, last cycle's figure restated for this metric. Read the columns only and 27 of the 64 packs in this corpus come out wrong. Opening one metric refresh pack, deciding which printed source rows belong to the metric at all, reading every note for a recalled, supplied or late-cut extract and for records recoded out of the channel, checking each counted row's extract as-of against the cut-off and the period end, finding the comparison baseline through any restatement, and applying the five-rung ladder in order before the figure goes into the appendix.

Audience

The data governance and compliance desks that assemble the metric appendix before an officer authors the annual assessment — and the officer, who needs to know which metrics were regenerated with lineage that stands and which were excluded rather than silently included. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual metric refresh packs

The corpus is 64 metric refresh packs, 0.15 MB (txt 64). There is no public corpus of annual risk-assessment metric refresh packs, and there could not be: the material is an operator's own source-record exports and its own desks' notes. So it is generated, and the generator is the measurement's own honest cost. What choosing it exercises is the one thing that matters here — the split between what a COLUMN says and what a SENTENCE says about it. 15 of the 64 packs are decidable from the printed columns alone, 27 turn on a note, and all 64 carry at least one DECOY note written in the same vocabulary, with the same row ids and the same dates, that decides nothing. Sixteen case families, three packs at exactly 1500 basis points and three at 1499, so the threshold is measured at its edge rather than in the middle.

The corpus

  • The 64 metric refresh packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every metric, business area, channel, source system, extract id, figure and note is invented, and RMR-2026 is an invented in-house procedure of a fictional operator.

Swap this folder for your own material and the kit is pointed at your metric refresh packs. That is the whole change — there is no database to migrate.

One metric refresh pack, as the model receives itAMR-0001.txt · 1 of 64
ANNUAL RISK ASSESSMENT - METRIC REFRESH PACK

PACK                AMR-0001
METRIC              MET-3101
METRIC NAME         Transaction count, table games buy-in and payout transactions, floor and host channels
BUSINESS AREA       TABLES
CHANNELS            FLOOR HOST
MEASURE             RECORDS
CYCLE               2026
PERIOD              2025-07-01 to 2026-06-30
PERIOD END          2026-06-30
BUCKETS IN PERIOD   2025-Q3 2025-Q4 2026-Q1 2026-Q2
EXTRACT CUT-OFF     2026-07-15
STATED BASIS        BASIS-C4
PROCEDURE           RMR-2026 as of 2026-07-01
PREPARED            2026-07-22

DEFINITION AS STATED IN THE METRIC INVENTORY
  Count of table games buy-in and payout transactions recorded against the TABLES business area
  in the floor or host channel, for each quarter bucket inside the cycle period. One
  contribution row per business area, channel and bucket. A row outside the business area,
  outside the channels in scope, or in a bucket outside the period is not part of this metric.

PRIOR CYCLE
  CYCLE         2025
  PUBLISHED     84,201
  STATED BASIS  BASIS-C4

SOURCE CONTRIBUTIONS
  ROW   AREA        CHANNEL  BUCKET     RECORDS         VALUE US$  SYSTEM     EXTRACT   AS OF
  C-01  TABLES      FLOOR    2025-Q3     12,400      3,534,000.00  SYS-TABLE  EX-40001  2026-07-18
  C-02  TABLES      FLOOR    2025-Q4     31,200     12,854,400.00  SYS-TABLE  EX-40002  2026-07-05
  C-03  TABLES      FLOOR    2026-Q1     24,600      3,911,400.00  SYS-TABLE  EX-40003  2026-07-08
  C-04  TABLES      FLOOR    2026-Q2     15,200      9,636,800.00  SYS-TABLE  EX-40004  2026-07-10
  C-05  TABLES      HOST     2025-Q3     16,800      3,822,000.00  SYS-TABLE  EX-40005  2026-07-13

NOTES
  Data desk, 2026-07-16: a later cut of EX-40002 exists, dated 2026-07-21; it is not the one

Abridged — the file continues.

The outcomeWhat a good result looks like

One metric refresh pack in, one appendix row out: which source rows count, which counted rows carry no standing lineage, which extracts do not cover the period, the comparison baseline, whether the basis changed, the regenerated figure, the movement in basis points, every exception and one of five RMR-2026 verdicts. 57 of 64 packs come back with all seven graded fields right, against 37 for the free floor of record.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the 1,000-token ceiling, 0 were closed by the reply stop and no call failed. The 7 packs it got wrong are named in the kit README with what it answered and why: 2 late-cut rescues it missed, 2 suppressed movements, 2 metrics it excluded that stand, 1 restatement, 1 supplied lineage, 1 recoding it counted and 1 decoy it moved on. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your metric inventory records a withdrawn extract, a supplied extract, a recoding and a restatement as STRUCTURED FIELDS on the export — the free modal floor, and do not buy a call at all
    37 of 64 packs for $0.00, and on the 15 packs the columns decide it gets 15 against the paid call's 14.
  • Those changes arrive as free text — a desk note, an email pasted into the pack, a comment column — the paid call
    This is the whole product. On the 27 packs where a sentence decides a reading the paid arm is 21 and the modal floor 0; on the withdrawn-extract family 5 of 5 against 0, and on basis changes 4 of 4 against 1.
  • You need the appendix to be defensible about what was EXCLUDED rather than accurate about what moved — either arm, with the station
    raised_a_suppressed_metric is 0 on every committed arm, free ones included, because M-1 beats M-4 in src/policy.py and no reply can reach the ladder. The station is what makes that true, and the station is free.

And where nothing here is good enough:

  • Somebody can put a sentence into the pack — neither arm, unchanged
    The cap holds — 0 scores set, 0 risks characterised, 0 assessments authored, 0 people named across all 22 attacked packs — and an asserted threshold moves the station on none of 6. But a FALSE recalled-extract note took 5 of 5, and 7 of the 22 packs right on the clean run were lost.

At a glanceHow the whole thing runs

89%whole pack pct
1,630 msp50, end to end
$2.30per 1,000 metric refresh packs · the fast tier

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs in the same block shape — header, definition, prior cycle, the printed contribution rows with EXTRACT and AS OF columns, the notes, the materiality block — and write one line per pack into data/gold.jsonl with the five readings. ⚠︎ EVERY PERCENTAGE ON THIS PAGE STOPS APPLYING THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per metric for a reading your export already gives you. That is the case against the best-fitting scenario (“Your metric inventory records a withdrawn extract, a supplied extract, a recoding and a restatement as STRUCTURED FIELDS on the export”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A PACK CARRYING MORE THAN ONE PRIOR CYCLE. R-6 and R-7 each read exactly one carried figure and one carried basis. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-assessment-metrics. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run, and re-scores every committed run at $0.00. pip install -r requirements.txt installs nothing — the kit is standard library only. tools/build_corpus.py --check rebuilds all 159,566 bytes and python3 -m evals.check_labels re-derives the key independently, together in about two and a half seconds. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning