The business caseThe problem this solves
A risk appetite statement is not a set of thresholds. It is a set of thresholds plus a handful of sentences about TIME -- a second consecutive Amber is reported Red, a second consecutive Red goes to the board, an indicator coming back from Red is held at Amber for one period. Two indicators reading exactly the same number against exactly the same thresholds get different verdicts, because one of them was Amber last month and the other was Green. Nothing in the period you are holding tells you which one it is, so somebody goes back through the prior packs by hand, every indicator, every cycle. Somebody reading one indicator's period report, then going back through the previous reports to work out whether this is the second consecutive Amber, whether the board has already been told, and whether last quarter's Red is still on a recovery hold.
Audience
Second-line risk functions and the people who build their reporting: anyone who owns an appetite statement with consecutive-period escalation rules in it, and anyone deciding whether a language model has any business near one. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual period reports
The corpus is 150 period reports, 0.30 MB (jsonl 1 · txt 150). A real KRI pack names a real institution's loss rate, its liquidity coverage, its complaint volume and the person who owns each number, which is exactly the material that never leaves a bank. There is no public corpus of (appetite statement, indicator history, adjudicated escalation) for the same reason there is no public corpus of production alert histories. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/appetite.replay's output over the planted patterns, so a fiddly consecutive-period rule cannot carry its author's misreading into the score.
The corpus
- The 150 period reportsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your period reports. That is the whole change — there is no database to migrate.
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes
no real institution, indicator or person. Unit KRI-0001, period P1.
Indicator
----------------------------------------------------------------
Name : Operational loss rate
Function : Operational Risk
Definition : basis points of quarterly operating revenue
Unit : bps
Direction : higher_is_worse (higher readings are worse)
Appetite Thresholds
----------------------------------------------------------------
Amber at : 30.5 bps
Red at : 50.4 bps
Escalation rules stated in the appetite statement:
1. An indicator reported Amber in two CONSECUTIVE periods is reported Red in the second.
2. An indicator reported Red in two CONSECUTIVE periods requires board notification.
3. An indicator whose reading returns to Green in the period immediately after a Red is held at
Amber for that one period (recovery hold) and may return to Green thereafter.
4. The consecutive count resets on any Green reading. Two Ambers separated by a Green are not
consecutive.
Period Reading
----------------------------------------------------------------
Period : P1
Reading : 22.54 bps
Driver Breakdown
----------------------------------------------------------------
Contribution to this period's reading, by driver:
Processing errors 44.3 % of the reading
External fraud 28.1 % of the reading
Legal settlements 17.0 % of the reading
Physical damage 10.6 % of the reading
Internal Contacts
----------------------------------------------------------------
Indicator owner : T. NakamuraAbridged — the file continues.
The outcomeWhat a good result looks like
Per period: the band to REPORT after the escalation rules have been applied, the name of the rule that was applied, and the driver contributing most to the reading -- plus a four-field carried state the next period is judged against, written by code from the arithmetic and never from the model's reply.
And when it cannot
It names the wrong rule, and on this run it does so in exactly one place. All 10 miss cells are third periods, 9 of them on AMBER-AMBER-AMBER: four report AMBER / NONE where the key says RED / ESCALATED_FROM_AMBER, and one reports BOARD_NOTIFICATION. The tenth is a single FALSE ALARM -- a BOARD_NOTIFICATION on a quiet RED-RED-AMBER third period, 1 of 106. Under-escalating and over-escalating cost differently and are never averaged: the first leaves a sustained breach unreported, the second takes something to a board that should not go there.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- An appetite statement whose rules are all about the CURRENT reading -- a threshold breach, a limit, a traffic light — the free floor, at $0.00
The floor gets the band right on every period where no history rule fires and names the top driver on 150 of 150 -- exactly what the model scores. With the state removed the model returned literally the same 450 answers. - A FIRST escalation, or a recovery hold after a Red — the fast tier WITH the carried state
100 pct, everywhere. 7 of 7 GREEN-AMBER-AMBER third periods, 16 of 16 recovery holds, 11 of 11 counter-reset third periods -- the trap the corpus over-represents on purpose. Both the stateless arm and the free floor score 0 on all of them. - Deciding whether the board should actually be notified — a person, reading the reported band and the rule the kit named
Both directions of the rule-2 error are here: one period claims a board notification the key does not want, and one names it where the key wants ESCALATED_FROM_AMBER. Two rows is not a rate, and board notification is the highest-consequence output this kit has.
And where nothing here is good enough:
- A THIRD consecutive breach on an indicator that has already escalated — neither, as built -- this is the named failure
0 of 5. Every AMBER-AMBER-AMBER third period in the corpus is wrong, four of them citing rule 1's 'reported Amber' wording against a period 2 that was reported RED. It is the only cell of the 24 that is not 100 pct correct, and it is uniformly wrong rather than noisy, so a bigger sample would not average it away. - An indicator whose history is longer than three periods — nothing here yet -- measure it first, and expect it to be worse
Every rule in this corpus resolves inside three periods. The one shape the run cannot do at all is a breach sustained past its first escalation, and that is exactly the shape a longer history produces more of.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-22. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace src/appetite.py FIRST. Every figure on this page is a property of a three-period corpus with unambiguous bands and a clear top driver, measured once. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a model to do arithmetic a regex already does perfectly. That is the case against the best-fitting scenario (“An appetite statement whose rules are all about the CURRENT reading -- a threshold breach, a limit, a traffic light”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | Three periods per indicator. Every rule in this appetite statement resolves within three periods, so nothing here measures a rule that needs four -- a rolling twelve-month breach count, an annual limit, a trend over a year. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE ANSWER KEY READS RULES 1 AND 2 THE WAY A READER WOULD -- AND ONLY HALF OF THAT IS PRICED. Both rules are written about the band an indicator was REPORTED; src/appetite.py resolves both against the RAW reading as well. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-22 — r001-kri-breach. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes every pre-flight assertion (python3 -m evals.check_labels), scores the free floor (python3 -m evals.run --run-id b000-kri-breach-rawband --baseline) and re-scores the recorded run under the other reading of rule 2 (python3 -m evals.ambiguity). Four commands, plain python3, no install step -- requirements.txt names nothing because the kit imports nothing outside the standard library. What it cannot do without a key is re-run the scored eval or its stateless control.