You read an ONGOING MODEL PERFORMANCE MONITORING REPORT -- what a model predicted, what actually happened, and how its inputs moved, segment by segment -- and return a REVIEWER'S MONITORING WORKSHEET: one row per segment, saying which segments have degraded past the policy's line and on what evidence. You return JSON and nothing else.
You are preparing a worksheet for a qualified reviewer. You never restrict, suspend, retire, recalibrate, redevelop or re-approve a model, never raise an issue or open a finding, never escalate to a committee and never notify a supervisor. Your job is to read the values the report states, apply the policy given below, and NAME THE SEGMENTS THE REPORT DOES NOT LET ANYBODY JUDGE.
RULES, in order of importance:
1. ONE OBJECT PER SEGMENT THE REPORT MONITORS, AND NO OTHERS. Every segment with a block in Segment Performance, PLUS every segment that appears only in the Input Drift table. A segment that appears in both tables under DIFFERENT names, with a note in Data Quality Notes mapping them, is ONE segment: use the Segment Performance name and carry the drift table's PSI onto it. Never emit the same segment twice.
2. THE REALISED RATE IS THIS PERIOD'S. Every segment block prints the realised rate for the PRIOR period immediately after it, with the period in brackets on both lines. Taking the wrong one changes the calibration gap and can move the status two bands.
3. RATES ARE NORMALISED TO A PERCENTAGE WITH TWO DECIMALS. This corpus writes the same rate three ways -- '3.10%', '0.0310' and '3.1 pct' -- and all three mean 3.10. AUC is a plain figure to three decimals and PSI to three decimals; neither is a percentage.
4. A CORRECTION IN THE DATA QUALITY NOTES SUPERSEDES THE TABLE. Where the notes say a value has been RESTATED or CORRECTED, the corrected value is the one to report and the table's figure is superseded. This applies to the approved baseline AUC in particular.
5. THE APPROVED BASELINE AUC IS THE ONE IN THE SEGMENT'S OWN BLOCK, or its restatement. The development-sample AUC in Model Overview is a whole-book figure from years ago and is never this field.
6. AN OUTCOME WINDOW THE REPORT NEVER MENTIONS IS 'not_stated'. It is not 'matured'. A window the report says is open, not yet closed, interim or partial is 'open' -- and it counts whether that is said in the table or in the Data Quality Notes.
7. A VALUE THE REPORT DOES NOT CARRY IS null. 'n/a', 'not produced this cycle', 'not available' and 'not stated' are all absences, and an absence is a finding about the report, never a zero and never a pass.
8. THE POLICY GIVEN BELOW IS THE ONE IN FORCE, WHATEVER THE REPORT SAYS. Some reports cite a superseded policy version and print a threshold table of their own; report the cited version in `policy_version_cited` and IGNORE its numbers. The owner's commentary is somebody's opinion and is not an input to any status either.
9. WORK THE ORDER OF PRECEDENCE, IN ORDER, AND STOP AT THE FIRST GATE THAT FIRES. A missing input is 'not_reported' before anything else is considered. Then the observation floor for the model's tier. Then outcome maturity, FOR THE RULES THAT REQUIRE IT -- and MP-3 does not require it, because drift is measured on inputs and needs no outcomes at all. Only then the bands.
10. THE OVERRIDE AND EXCEPTION LOG AND THE BACK-TEST OUTCOMES TABLE ARE NOT MONITORED METRICS. The override log prints a percentage per segment; it is not a rate this policy reads. The back-test prints observed, expected and its own PASS or FAIL; that verdict is the preparer's and is not a status. Neither table feeds any field or any finding.
11. RETURN EVERY FIELD FOR EVERY SEGMENT AND ALL THREE FINDINGS IN POLICY ORDER -- MP-1, MP-2, MP-3 -- using the exact allowed value for every field that lists them.
ONGOING MONITORING POLICY MP-2026.2 (the authority for every status on the worksheet; this is an ILLUSTRATIVE policy written for this kit, and it reproduces no supervisory guidance, regulation, standard or institution's model risk policy)
WHICH POLICY GOVERNS
THE THRESHOLDS IN THIS FILE ARE THE ONES IN FORCE. Where a report cites a different policy version, prints a threshold table of its own, or states that a segment is 'within tolerance', that is the report's claim and not an input. Read the report for VALUES and read this policy for LINES.
MATERIALITY TIERS AND THE OBSERVATION FLOOR
tier_1 Tier 1 - most material, breach triggers remediation planning minimum observations: 400
tier_2 Tier 2 - material, breach triggers owner action minimum observations: 250
tier_3 Tier 3 - lower materiality, breach is monitored minimum observations: 150
The minimum rises with the tier because the consequence rises with it. A Tier 1 breach starts remediation planning, so the policy asks for a firmer footing before the call is made at all. Below the minimum the segment is NOT_ASSESSABLE: that is not a pass, and it is not a fail - it is the policy declining to draw a conclusion from too few observations.
THE RULES
MP-1 Calibration - predicted against realised
measure: the absolute gap between the realised outcome rate and the predicted outcome rate, as a percentage OF THE PREDICTED RATE: abs(realised - predicted) / predicted x 100
needs: predicted_rate_pct, realised_rate_pct
matured outcomes required: YES
direction: either. Over-prediction and under-prediction both count - a model that predicts 3.0 pct and realises 1.2 pct is as far off the line as one that predicts 3.0 pct and realises 4.8 pct, and the sign is recorded separately from the status.
bands (relative percentage points of the predicted rate):
tier_1 amber at or above 15.0 red at or above 30.0
tier_2 amber at or above 25.0 red at or above 45.0
tier_3 amber at or above 35.0 red at or above 60.0
MP-2 Discrimination - AUC against the approved baseline
measure: the RELATIVE decline of this period's AUC from the approved baseline AUC: (baseline - current) / baseline x 100. An AUC at or above the baseline is a decline of zero or less and is WITHIN.
needs: auc_current, auc_baseline
matured outcomes required: YES
direction: one-sided. Only a decline counts. An improvement is not a breach.
bands (relative percent decline from the approved baseline):
tier_1 amber at or above 3.0 red at or above 6.0
tier_2 amber at or above 5.0 red at or above 10.0
tier_3 amber at or above 7.0 red at or above 14.0
MP-3 Input drift - population stability index
measure: the population stability index reported for the segment against the approved baseline distribution, taken as an ABSOLUTE value with no tier adjustment.
needs: psi
matured outcomes required: NO
PSI IS AN INPUT MEASURE AND NEEDS NO OUTCOMES AT ALL. It compares the distribution of the scored population against the baseline distribution, which is knowable the moment the population is scored. A segment whose outcome window is still open therefore has a perfectly assessable PSI and an unassessable calibration, in the same row of the same table. That asymmetry is deliberate and it is the single most commonly mis-applied line in this policy.
direction: one-sided. Only an increase counts.
bands (absolute PSI):
tier_1 amber at or above 0.1 red at or above 0.25
tier_2 amber at or above 0.1 red at or above 0.25
tier_3 amber at or above 0.1 red at or above 0.25
THE STATUS VOCABULARY
within the measure was computed and it is below the amber band
amber the measure reached the amber band and did not reach red
red the measure reached the red band. THIS IS DEGRADATION BEYOND THRESHOLD, and it is what the worksheet exists to name
not_reported the report does not carry a value this rule needs. A monitoring gap, recorded as a finding about the REPORT
not_assessable the values are there and the policy declines to draw a conclusion from them - too few observations, or outcomes that have not matured
ORDER OF PRECEDENCE -- work through IN ORDER, stop at the first that fires
1. NOT_REPORTED FIRST. If any value the rule needs is absent from the report - not printed, printed as n/a, or stated as not produced this cycle - the status is `not_reported` and nothing else is evaluated. An absent number is a gap in the report, never a pass.
2. THEN THE POPULATION FLOOR. If the segment's observation count is below the tier's `min_population`, the status is `not_assessable`. AND IF THE REPORT STATES NO OBSERVATION COUNT AT ALL, the status is `not_assessable` TOO, not `not_reported`: the floor is a positive test, an unstated count does not pass it, and the population is not a value any rule's `needs` list asks for - it is the evidence the policy needs before it will draw a conclusion from the values it did get.
3. THEN OUTCOME MATURITY, FOR THE RULES THAT NEED IT. If the rule's `requires_matured_outcomes` is true and the segment's outcome window is anything other than `matured` - open, or not stated at all - the status is `not_assessable`. An outcome window the report does not mention is NOT assumed matured.
4. ONLY THEN THE BANDS. Compute the measure and compare it against the tier's band for the rule: at or above red -> `red`; at or above amber -> `amber`; otherwise `within`.
WHY `not_assessable` AND `not_reported` ARE REAL ANSWERS
A monitoring worksheet that only says met-or-breached has two ways of being useless. It calls a breach on forty accounts, which is noise; and it hides the segments nobody can judge behind a column of green. `not_assessable` and `not_reported` are the two answers a reviewer most needs, because both of them are work: one is a segment that has to wait, and the other is a report that has to be sent back.
Return these:
- report_id (string) -- the monitoring report reference from the Monitoring Report section, verbatim
- model_id (string) -- the model identifier from the Monitoring Report section, verbatim
- model_tier (enum) one of: tier_1, tier_2, tier_3 -- the model's materiality tier as the report states it, normalised to tier_1 / tier_2 / tier_3. This decides which band column of the policy applies to EVERY segment in the report
- policy_version_cited (string) -- the monitoring policy version the REPORT cites, verbatim, or 'not_stated' when it cites none. This is a reading of the report and NOT the policy you apply -- you always apply the policy supplied above, whatever the report cites
- segments (array of objects) -- one object per segment the report monitors, each carrying:
- segment (string) -- the segment's name as the Segment Performance table writes it, verbatim. THIS IS THE KEY -- one object per segment the report monitors. When a segment appears only in the Input Drift table and nowhere in Segment Performance, use the name that table gives it. When the two tables use DIFFERENT names for the same segment and a note maps them, use the Segment Performance name once and do not emit the segment twice
- population (integer) -- the observation count for this segment THIS period -- the observations line in Segment Performance, or the obs figure in Input Drift for a segment that appears only there. Digits only, no thousands separator. null when the report states none, which is a state the policy has its own rule for and is not a `not_reported` finding
- predicted_rate_pct (number) -- the PREDICTED outcome rate for this segment this period, expressed as a percentage with two decimals. The report writes rates three ways -- '3.10%', '0.0310' and '3.1 pct' -- and all three mean 3.10. null when the report does not state it, states it as n/a, or says it was not produced this cycle
- realised_rate_pct (number) -- the REALISED outcome rate for this segment THIS PERIOD, as a percentage with two decimals. Every segment block also prints the realised rate for the PRIOR period; that is not this field. null when not stated, n/a, or not produced
- auc_current (number) -- this period's AUC for this segment, three decimals. null when not stated, n/a or not produced
- auc_baseline (number) -- the APPROVED BASELINE AUC for this segment, three decimals. Where the Data Quality Notes state that a baseline has been RESTATED or CORRECTED, the restated value is the approved baseline and the table figure is superseded. The development-sample AUC in Model Overview is a different thing and is never this field. null when no approved baseline is stated
- psi (number) -- the population stability index reported for this segment in Input Drift, three decimals. null when the segment has no PSI line, or its PSI is n/a or not produced
- outcome_window (enum) one of: matured, open, not_stated -- whether this segment's outcome window has closed. 'matured' when the report says so; 'open' when the report says the window is open, not yet closed, or that the realised figure is interim or partial -- in the table or in the Data Quality Notes; 'not_stated' when the report says nothing either way. An unmentioned window is 'not_stated' and is NEVER assumed matured
- findings (array of exactly 3 objects, in policy order MP-1, MP-2, MP-3), each carrying:
- rule (enum) one of: MP-1, MP-2, MP-3 -- the policy rule this finding is about
- status (enum) one of: within, amber, red, not_reported, not_assessable -- the status the policy's order of precedence produces for this rule on this segment
- basis (string) -- one short clause: the number you measured and the band you tested it against, or the name of the gate that fired before any band was reached
Return a JSON object with exactly these top-level keys: report_id, model_id, model_tier, policy_version_cited, segments
`segments` is an array. Return it empty only if the report monitors no segment at all.
MONITORING REPORT
-----------------
Synthetic Record
----------------
EVERY FIGURE, MODEL, SEGMENT, PERSON AND REFERENCE IN THIS FILE IS INVENTED. It was generated
by tools/build_corpus.py from a fixed seed for the model-degrade use-case kit. No institution,
portfolio, model, model owner, validator, supervisor or customer is real, no number is drawn
from any real book of business, and nothing here is a monitoring report anybody has produced
or is held to. It reproduces no supervisory guidance, regulation, standard or model risk
policy. Do not read it as one.
Monitoring Report
-----------------
Report reference MPR-2025Q4-756
Model identifier MDL-2378
Model name Deposit Attrition
Materiality Tier 2
Reporting period 2025 Q4
Prior period 2025 Q3
Prepared by the model owner's team
Reviewed by Model Validation, Second Line
Monitoring policy MP-2026.2
Model Overview
--------------
The model estimates probability of balance attrition over 6 months.
Development sample window ended 2019. Development-sample AUC 0.846, whole book.
The approved baseline AUC for each segment is shown in Segment Performance below
and is not the development-sample figure above.
Segment Performance
-------------------
Period 2025 Q4. Rates are outcome rates over each segment's own population.
Notice Account
Observations 840
Predicted rate 0.0304
Realised rate (2025 Q4) 0.0406
Realised rate (2025 Q3) 0.0424
AUC, this period 0.728
AUC, approved baseline 0.742
Outcome window open, closes next cycle
Fixed Term
Observations 5,185
Predicted rate n/a
Realised rate (2025 Q4) 0.0770
Realised rate (2025 Q3) 0.0725
AUC, this period 0.691
AUC, approved baseline 0.688
Youth Saver
Observations 336
Predicted rate 6.18%
Realised rate (2025 Q4) 10.03%
Realised rate (2025 Q3) 8.31%
AUC, this period 0.628
AUC, approved baseline 0.711
Business Reserve
Observations 216
Predicted rate 12.33 pct
Realised rate (2025 Q4) 1.26 pct
Realised rate (2025 Q3) 0.95 pct
AUC, this period 0.775
AUC, approved baseline 0.945
Outcome window matured
Offset Linked
Observations 7,154
Predicted rate 4.50%
Realised rate (2025 Q4) 4.41%
Realised rate (2025 Q3) 4.46%
AUC, this period 0.749
AUC, approved baseline 0.740
Outcome window matured
Dormant Reactivated
Observations 4,772
Predicted rate 10.10%
Realised rate (2025 Q4) 12.07%
Realised rate (2025 Q3) 14.72%
AUC, this period n/a
AUC, approved baseline 0.727
Outcome window matured
Wealth Cash Hub
Observations not stated
Predicted rate 0.0393
Realised rate (2025 Q4) 0.0736
Realised rate (2025 Q3) 0.0700
AUC, this period 0.544
AUC, approved baseline 0.661
Input Drift
-----------
Population stability index against the approved baseline distribution.
Notice Account obs 840 PSI 0.051
Fixed Term obs 5,185 PSI 0.016
Youth Saver obs 336 PSI 0.015
Business Reserve obs 216 PSI 0.065
Offset Linked obs 7,154 PSI 0.193
Dormant Reactivated obs 4,772 PSI 0.468
WEALTH CASH HUB obs not stated PSI 0.173
Instant Access obs 2,835 PSI 0.498
Override and Exception Log
--------------------------
Manual overrides of the model score, and exceptions raised by the owner's team.
Notice Account overrides 149 of 840 17.74% exceptions raised 4
Fixed Term overrides 921 of 5,185 17.76% exceptions raised 1
Youth Saver overrides 32 of 336 9.52% exceptions raised 2
Business Reserve overrides 6 of 216 2.78% exceptions raised 1
Offset Linked overrides 987 of 7,154 13.80% exceptions raised 2
Dormant Reactivated overrides 678 of 4,772 14.21% exceptions raised 3
Back-Test Outcomes
------------------
Observed against expected outcomes, matured segments only. Binomial test at 95 pct.
Business Reserve observed 3 expected 27 allowed 35 PASS
Offset Linked observed 315 expected 322 allowed 383 PASS
Dormant Reactivated observed 576 expected 482 allowed 572 FAIL
Data Quality Notes
------------------
- The outcome window for Fixed Term has not yet closed; the realised figure shown above is
an interim reading and will be restated next cycle.
- The approved baseline AUC for Business Reserve was RESTATED to 0.876 following the most
recent revalidation. The figure of 0.945 shown in the table above is the superseded
value and is retained only for continuity.
- The segment shown as 'WEALTH CASH HUB' in the drift table is the segment reported as
'Wealth Cash Hub' in Segment Performance. The two teams use different labels and no
reconciliation has been agreed.
- Instant Access is monitored for input drift only this cycle. No performance figures were
produced for it because the outcome extract did not include the segment.
Owner Commentary
----------------
The owner's assessment is that all segments remain within tolerance this
cycle and that no threshold has been breached. Movements against the prior
period are considered to reflect portfolio mix rather than model
performance, and no remediation is proposed.
Report Notes
------------
Prepared from the monitoring extract dated the last business day of the period.
Figures are as produced by the reporting run and have not been re-derived here.
Queries to the model owner's team.