The business caseThe problem this solves
A programme grants a tier, then issues benefits against it for a year. Nobody re-derives the tier. The qualification record is a table of activity with a TYPE column somebody else chose, and the exclusions that actually matter — a stay reversed after it posted, a booking settled from the member's own balance, a reservation through an unlisted agency, a promotional uplift, a service-desk gesture — are coded as ordinary activity and live in the description. One unread sentence moves the qualifying totals, which moves the supported tier, which moves every benefit sitting at or above it. The individually invisible one is the allowance: a benefit line whose COUNT is comfortably inside the tier's cap while the CUMULATIVE figure printed beside it is not. Every issuance looks fine, so nobody queries it. A sampled manual re-derivation. The work per account is reading eight-ish activity descriptions, totalling two columns, comparing against two thresholds, applying roll-over and soft landing, then checking four to seven benefit lines against a schedule and a cumulative cap. It is entirely doable and entirely tedious, which is why it is sampled rather than done. The kit does not replace the reading — it replaces the ARITHMETIC, and it measures how much of the reading is worth buying.
Audience
Programme governance and membership operations analysts auditing granted tiers and the benefits issued under them Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual member account records
The corpus is 62 member account records, 0.27 MB (txt 62). It is generated because a real one cannot be published and because the ANSWER HAS TO BE DERIVABLE. Each account is built as a record of typed activity lines carrying their TRUE NATURE, typed benefit lines, and the programme's own rules; the text is RENDERED from that record and the key is then src/policy.py run over the SAME record. There is nowhere to type a verdict, a tier or a determination, so a corpus edit cannot leave a stale key behind. ⚠︎ AND IT WAS ATTACKED BEFORE IT WAS TRUSTED, because two kits on this estate had to rebuild a corpus after the fact. tools/attack_corpus.py ships with the kit and runs three attacks fitted ON THE ANSWER KEY, which the shipped floors may never see. Both leaks it found are recorded in data/SOURCES.md with the measurement that closed them.
The corpus
- The 62 member account recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 62 accounts, the programme register and the whole answer key are generated. Nothing here is fetched, scraped, licensed or derived from anything that was, and there is no personal data in it anywhere: evals/check_labels.py sweeps every record for five real-identifier shapes (email, telephone, card-length digit runs, national-id patterns and any mention of a date of birth) and reports 0.
Swap this folder for your own material and the kit is pointed at your member account records. That is the whole change — there is no database to migrate.
MEMBER ACCOUNT AUDIT - TIER STATUS AND BENEFIT ENTITLEMENT
ACCOUNT HEADER
Account MA-0001
Programme Meridian Circle
Operator Harrowfield Holdings
Member reference MBR-40000
Home market Coastal
Qualification period 2025-01-01 to 2025-12-31
Review date 2026-02-10
PROGRAMME RULES
Tier thresholds (BOTH must be reached)
SILVER 20 units 8 qualifying transactions
GOLD 45 units 18 qualifying transactions
PLATINUM 80 units 32 qualifying transactions
Roll-over units above the tier attained carry forward, capped at 15 units;
roll-over units are units and are never transactions
Soft landing a member who held a tier in the prior period, does not earn it
this period and reaches 80 pct of BOTH of that tier's thresholds
retains it for one further period, once only
Excluded activity activity reversed after posting, activity settled from the member's
own accrued balance, activity booked through an intermediary the
programme does not list, extra units posted by a promotion, and
units added by the service desk as a gesture - WHATEVER THE TYPE
COLUMN SAYS
Separately coded Redemption, Promotion
Benefit schedule
UPG Confirmed upgrade at booking min SILVER allowance SILVER 2, GOLD 6, PLATINUM unlimited
AMN Welcome amenity on arrival min SILVER allowance SILVER 3, GOLD 9, PLATINUM unlimited
WVR Change fee waived on one booking min SILVER allowance SILVER 1, GOLD 3, PLATINUM 6Abridged — the file continues.
The outcomeWhat a good result looks like
One worked audit row per account: the tier the record supports, how it compares with the tier granted, the override state, a verdict and a cited rule on every benefit issued, the benefit lines named for review, and one quoted line from the record behind each. 50 of 62 accounts come back completely right from FREE code with no model and no network; the 12 whose tier lives in a description are the only ones a call can add anything to.
And when it cannot
The two errors are not the same size and this kit never averages them. An account the record sends to REVIEW and the arm called CLEAR is a defect nobody will now look at — 7 of 62 on the paid arm, 12 on the free column floor, 19 on the null floor that simply believes the account. The reverse, an account wrongly sent to review, costs an analyst ten minutes — 8 on the paid arm and 0 on both non-attacking floors. There is no F-score anywhere in this kit.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your activity descriptions add nothing the type column does not already say, and the exclusions your programme cares about are the ones it codes separately — the free columns floor, alone
on this corpus that arm answers 50 of 62 accounts completely correctly at $0.00 with no network and no key, and it BEATS the paid arm's 26. There is no measured case here for spending anything on those accounts. - You want the allowance nobody queries — a benefit whose per-line count sits inside the tier's cap while the cumulative printed beside it does not — the free columns floor, alone
11 of 11 such lines caught, every time, because it is two printed numbers and a schedule lookup. The paid arm caught 6. - The exclusions that matter to you are written into descriptions and coded as ordinary activity, and you can afford to check the output — the paid arm, routed only at the accounts the columns cannot settle
the call reaches 8 of the 12 accounts whose tier lives in a description, which the columns floor reaches 0 of. The columns floor can identify which accounts those are before anything is spent, so the routing signal is free. - You want to use the model's own confidence to decide which rows a human reads — the structural signal instead — whether the answer lives in a column or in a description
the returned confidence does not separate: 0.95 at the median on the accounts it got completely right and 0.92 on the ones it did not.
And where nothing here is good enough:
- You want a machine to act on the tier, the benefit, the override or the unit balance — nothing in this kit
the answer contract offers no field that could, asserted at import against the forbidden-name list, and the free sentence and every returned key are phrase- and name-tested on every arm.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-06. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own records into data/corpus/ as text and point data/accounts.json at them, or rewrite tools/build_corpus.py's family plan. ⚠︎ THE MOMENT YOU BRING YOUR OWN RECORDS YOU HAVE NO ANSWER KEY, AND EVERY NUMBER THIS KIT PUBLISHES IS AGAINST ONE. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a reading that a column already proves. That is the case against the best-fitting scenario (“Your activity descriptions add nothing the type column does not already say, and the exclusions your programme cares about are the ones it codes separately”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A REAL PROGRAMME'S TYPE VOCABULARY. Six types are coded here; a real operator has dozens, and which of them the programme codes separately is the whole free half of T-5.1. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | RUN-TO-RUN VARIANCE. One paid run was bought. 4 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-06 — r001-status-benefit-audit. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, python3 tools/build_corpus.py, then python3 -m evals.run --run-id <id> --floor columns. No key, no network, no install, and it reproduces the headline finding — 50 of 62 accounts completely right — in a few seconds.




