A legal hold or retention rule is supposed to cover every system that holds those records, but the register just shows a checkbox. This app checks each system against the rule and flags the ones the hold never actually reached.
PresenterOpens the private repo. Visible to admins only.
For the records compliance leadRetail · Legal Services
Why it matters
Today's manual process, and the same job with the app
A records or e-discovery lead at a retailer or law firm, checking whether a legal hold actually landed.
✕Today's manual process
1Read the register and take its APPLIED column at its word.
2Check systems one at a time cross-referencing hold rules against each connector from memory.
3Miss the quiet ones a mailbox already archived, a legacy host with no hold feature.
4A hold looks covered but isn't and records that should still exist are gone.
Every system checked manually, one at a time
✓With the app
1Every system is checked against the same rule, the same cycle, every time.
2The register's own table is read in code, so nothing here depends on memory.
3What the register can't show is caught a system the flag never actually reached.
4The gap is named the system, the record class, and who owns it.
Every gap named, system and rule together
See it work
One real case, shown from a saved run
Three systems hold payment-card files under one rule; a fourth, an offsite tape vault, never made the register's own page.
Catch the system a legal hold missedReference appBuilt to be shaped to your process
4
1Two systems, on paper ERP-ORACLE-FIN and POS-JOURNAL-VAULT, both logged applied.
2A fourth, never printed TAPE-OFFSITE-VAULT was in scope for cycles, but the register never lists it.
3What the app found It looks held on paper, but the flag never actually reached that system.
4Why it matters Payment-card files, five years old, with two years still left on the clock.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A hold or retention rule is written against record CLASSES and has to land on SYSTEMS, and systems are where it goes wrong. The register shows the hold applied; one system took the flag and dropped it. Worse, reach accumulates and decays across cycles — a system onboarded two cycles ago is in scope now and its first cycle never mentioned it; one decommissioned last cycle still shows in the register. A reviewer working one cycle at a time cannot see either. an e-discovery register whose APPLIED column is a checkbox — which cannot show a system the connector silently refused, re-raises rules answered two cycles ago, and still lists systems decommissioned last quarter.
Audience
a records or e-discovery lead deciding whether a legal hold actually covers what the matter register says it covers, and a compliance officer signing an attestation on it. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual rule review extracts
The corpus is 120 rule review extracts, 0.87 MB (txt 120). A matter register names live litigation, its custodians and every system a company runs. There is no public one and there will not be. A scrubbed export is worse rather than better: scrubbing removes the reviewer's typed note about why a connector refused, which is exactly the layer this kit measures.
The corpus
The 120 rule review extractsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your rule review extracts. That is the whole change — there is no database to migrate.
One rule review extract, as the model receives itRR-0001-C1.txt · 1 of 120
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated records-retention review extract for an
AI use-case kit; it reproduces no company, no matter, no system, no real person and no
real retention schedule. The retention periods and the disposition grace band are
ILLUSTRATIVE, OPERATOR-TUNABLE DEFAULTS -- not any company's or any jurisdiction's
schedule.
Matter Or Rule
----------------------------------------------------------------
Rule reference : RR-0001
Rule kind : SCHEDULE
Matter : -- none; this is a scheduled retention rule
Record class mapping on file : Supplier correspondence
Records owner (register of record) : Declan Ashworth
Register status : ISSUED 2025-05-12, open
Systems In Scope
----------------------------------------------------------------
The matter register's OWN list of systems against this rule, with its OWN
APPLIED / NOT APPLIED column exactly as it stands.
⚠︎ THIS TABLE IS STALE IN BOTH DIRECTIONS. It does not list a system onboarded in an
earlier review cycle, and it still lists one decommissioned in an earlier cycle. And the
column records what somebody DID, not what LANDED: it reads APPLIED for a system that
rejected the flag.
System Kind Register says
M365-MAILBOX-SUPPLY mailbox APPLIED
DOCSTORE-SHAREPOINT document store APPLIED
SFTP-SUPPLIER-DROP file drop APPLIED
Retention Policy (Default)
----------------------------------------------------------------
Abridged — the file continues.
The outcomeWhat a good result looks like
every rule's verdict with the unreached system named and its record class — 10 false actions across 94 quiet readings.
And when it cannot
a CONTEXT_INCOMPLETE reading (no record-class mapping on file — 9 of 120) judges nothing rather than guessing which classes a rule reaches. There is no default class.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Every connector failure is written back into the register as a status — the free floor — structured-log-strong, $0.00 It BEATS the model on the discriminator (90.0 against 87.5), on raise-it-now, on false actions and on disposition recall.
Connector refusals get typed into a reviewer note and the register catches up later — the model — but for ONE column, and know which Unreachable systems recorded only in prose: 100.0 pct against 47.06 pct. Everything else it either ties or loses.
And where nothing here is good enough:
You want the reading but cannot carry state between cycles — neither, as configured here The stateless control collapses to 53.33 pct on verdict with 62 false actions.
At a glanceHow the whole thing runs
88%verdict accuracy pct
9,533 msp50, end to end
$6.23per 1,000 rule review extracts · Google Gemini 3 Flash
Run once, for real, on 2026-08-25. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch the system a legal hold missed14 steps · 4 questions · run once, for real · 2026-08-25
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
tools/build_corpus.py writes data/corpus/*.txt and data/gold.jsonl. The measured figures on this kit's page do not transfer to your own corpus.Corpus lens →
When is this the wrong choice?
Avoid: Do not use it where a refusal is recorded as a note: 47.06 pct on prose-only unreachable against the model's 100.0 pct. That is the case against the best-fitting scenario (“Every connector failure is written back into the register as a status”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A rule with no record-class mapping on file (9 of 120 readings): CONTEXT_INCOMPLETE. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether rendering the carried state as JSON rather than English would score the same. 6 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-25 — r001-retention-reach. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Every corpus extract, the answer key, all three free floors, the injection probe and what r001-retention-reach actually answered ship in the repo. python3 -m evals.check_labels and python3 -m src.app both run with no key and no network.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
99.17%rows answered
9,533 msp50, end to end
55,418 msp95
5 minclone to first result
What the clock covers. one reading — one rule, at one scheduled run, end to end including provider-side reasoning tokens, on the shared connection. Not a cold-start figure and not a per-rule SLA: 120 readings ran in 557.6 wall seconds because different rules run concurrently while a single rule's runs are strictly serial.
Current processWhat it replaces
an e-discovery register whose APPLIED column is a checkbox — which cannot show a system the connector silently refused, re-raises rules answered two cycles ago, and still lists systems decommissioned last quarter.
Where it is not good enough
⚠︎ THIS KIT'S MODEL DOES NOT WIN THE HEADLINE, AND THE PAGE LEADS WITH THAT. It TIES the strong free floor on verdict (87.5 each) and LOSES the discriminator binary, 87.5 against 90.0 — plus raise-it-now, false actions, disposition recall and record class. All of it traces to ONE systematic defect: every one of its 14 verdict misses is the same over-call, UNREACHABLE_SYSTEM where the truth was milder, 13 of them on cycle 1, and all 14 naming ARCHIVE-TIER3 (13) or TAPE-OFFSITE-VAULT (1). Its own rationales show why — it reads the system's KIND off the register table ('a cold object archive tier that exposes no hold API') and treats that as evidence the rule failed, on pages whose change log is empty. The fix candidate is one sentence in the prompt. IT IS NOT SHIPPED AND THE RUN WAS NOT RE-TAKEN.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt120jsonl1json1
40 hold and retention rules, 120 readings across 3 review cycles — every unit re-read whole
verdict 87.5% against the strongest free floor's 87.5%
discriminator 87.5% — the free floor reads 90.0%
55.83%strip the carried state and it falls to
the model LOSES this binary to the strong floor, 87.50 against 90.00
2026-08-25as of
A records watch on system REACH. The register's APPLIED column is a join the floor does exactly; the model is there for a connector refusal recorded only in a sentence. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL: the carried state, written by the arithmetic and never from a model reply, and the clock, where one run owns exactly the events recorded since the last.
⚠︎ THIS IS THE ONE KIT IN THE LAP WHERE THE MODEL LOSES ITS OWN DISCRIMINATOR — 87.50 against the strong floor's 90.00 — and the figure leads with that. It also ties the headline verdict column and loses three more.
⚠︎ EVERY ONE OF ITS 14 VERDICT MISSES IS THE SAME SYSTEMATIC OVER-CALL: it reads a system's KIND off the register table as evidence the rule failed, on pages whose change log is empty. The fix candidate is one prompt sentence. IT IS NOT SHIPPED AND THE RUN WAS NOT RE-TAKEN. ⚑ IT WINS EXACTLY ONE COLUMN OUTRIGHT: unreachable systems recorded only in prose, 100.0 against 47.06. That is the whole case for paying for it.
⚠︎ AND THERE IS NO CADENCE MODULE, DELIBERATELY. The prior process is EVENT-DRIVEN with no interval to compare against; inventing one would be a fabricated number.
The swap seams
Seam
File
What changes
the retention periods
src/reach.py
The DEFAULT schedule.
the disposition rule
src/reach.py
When a record past retention with no hold reaching it becomes eligible.
the model
.env
PROVIDER, BASE_URL, MODEL.
what leaves the machine
src/select.py
SECTION_HINTS and NEVER_SENT.
the review cadence
src/reach.py
How long an unreached system stays unreached.
Components
Component
File
Role
the reach rule and state machine
src/reach.py
The DEFAULT retention periods, the disposition rule, the reach test and step() — which decides the next carried state from the answer key's inputs, never from a model reply. It is also the answer key.
the carried state
src/state.py
Which systems were reached, which failed, and which were onboarded or decommissioned. A system onboarded two cycles ago is in scope now and its first cycle never mentioned it.
the section splitter
src/segment.py
Splits the page into its 8 named sections; asserted across all 120 documents.
the send filter
src/select.py
Custodian Contact is mapped by no field and therefore never sent.
the prompt
src/prompt.py
Three parts: instruction and JSON shape, carried state, extract. The stateless control replaces exactly one line. ⚠︎ IT IS ALSO WHERE THIS KIT'S ONE SYSTEMATIC DEFECT LIVES: nothing tells the model that a system's KIND is not evidence a rule failed.
the model call
src/adapters/__init__.py
Raw HTTP, stdlib only, bounded retry; the shared daily call cap is checked here.
the reader
src/watch.py
One rule, one cycle, one call.
the scorer
evals/scoring.py
Exact match per cell. The unreachable-system discriminator is a separate binary.
the three free floors
evals/baseline.py
register-only, register-mem and structured-log-strong — $0.00 each, and the third beats the model on the discriminator.
the local UI server
src/app.py
http.server, stdlib. Renders with no key.
the local UI client
ui/app.js
Hand-written JS. ⚠︎ A CLIPPED AGREE COLUMN WAS FOUND BY OPENING A SCREENSHOT and fixed by shortening the strings. ⚠︎ AND THE SHARED ui/app.css CARRIES A SIBLING KIT'S VERDICT COLOURS, so this kit's verdicts render uncoloured. It was copied verbatim as instructed and NOT edited; the loudest signal still shows because .perm strong applies.
Where it breaks at scale
LINEAR IN RULES x CYCLES x SYSTEMS. A company with 400 live matters across 30 systems is a large multiple of this corpus and nothing amortises. The sublinear lever NOT implemented is skipping rules whose system list did not change — named here, not built.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
RR-0001-C3 — the model right and the free floor wrong, on a system that is not even in the register's own table. The register cannot report a failure on a row it does not have.successOpen full size →The same rule before anything is read: carried reach, the code-parsed register and the free floor all render with no key.emptyOpen full size →The read button pressed with no API_KEY — a plain sentence, not an error.nokeyOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
RR-0005-C2 — one of the 14 misses, and every one is the same shape: UNREACHABLE_SYSTEM called on a system whose KIND sounds unreachable, on a page whose change log is empty.failureOpen full size →
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
120rule review extracts
0.87 MiBtxt 120
p50 7,590chars per reading
$0.00setup · 0.5s
How it is cutWhat one reading is
40 rules x 3 review cycle, walked strictly in order per rule so the carried state moves forward the way it would in a real deployment. Different rules run concurrently; a single rule's cycles never do.
SetupWhat the setup figure measured
There is no index to build — each reading's extract goes whole into the prompt. The 0.5s and $0.00 are the corpus generation itself.
LicenceLicence
MIT
Bring your ownBring your own rule review extracts
tools/build_corpus.py writes data/corpus/*.txt and data/gold.jsonl. Keep the 8 section headings and the field labels src/watch.position_of parses. Set your own retention schedule and disposition rule in src/reach.py first.
⚠︎ And what stops being true when you do: The measured figures on this kit's page do not transfer to your own corpus. The memory-dependent share (49 of 120 readings here), the prose-only unreachable share (20 rules against 20 structured) and the lopsided failure-reason mix are properties of this generator's declared distribution, not facts about records management. Re-run the evals on your own data; that is what the harness is for.
What breaks it
A rule with no record-class mapping on file (9 of 120 readings): CONTEXT_INCOMPLETE.
A failure reason this corpus does not model. Three are modelled and the mix is lopsided (READ_ONLY_CONNECTOR 10, NO_HOLD_API 5, NO_HOLD_CONCEPT 1) — it is not a scored field, and which reasons are POSSIBLE is a property of which systems a class lives on, so it is published rather than levelled. The gate asserts coverage of at least one, not balance.
A jurisdictionally conditional retention period. The schedule here is flat.
An e-discovery platform that renames its export sections — the send filter's fallback is what stops the whole document going on the wire, reproduced on 120 of 120 documents.
Records that live in more than one system. Reach is judged per system here.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
260
65
the question and the JSON shape
6,149
1,537
the carried state
735
183
the extract
121
30
Total
1,815
This is the cost lesson as arithmetic: of the 1,815 tokens assembled, 1,602 are instructions — 88% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Reproduced from src/prompt.build for RR-0001-C3 with that reading's real carried state, not retyped — byte-identical to what evals/run.py sent.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled records-retention reach review. You read one retention rule or legal hold at one quarterly review cycle, work out which systems it actually reaches and where it silently failed to land, and you answer with one JSON object and no other text.
You are the scheduled records-retention REACH review for one company's retention rules and legal
holds. It runs on a QUARTERLY cadence -- one review cycle every 90 days -- and re-reads EVERY rule.
You are reading ONE rule at ONE review cycle.
A rule is written against a RECORD CLASS. It has to LAND on SYSTEMS. Your job is to say which
systems it actually reaches today and where it silently failed to land -- NOT whether somebody
issued it, which is a column in the register and needs no reader.
The rule, the register's own list of systems in scope with its APPLIED / NOT APPLIED column, the
DEFAULT retention schedule, and this cycle's recorded changes are reproduced in the extract below.
Apply them exactly as written. You cannot see the earlier review cycles; what is known about them
is stated under "Carried state" and is the only history available to you. Do not assume anything
about earlier cycles beyond it.
⚠︎ THE RETENTION PERIODS AND THE DISPOSITION GRACE BAND ARE ILLUSTRATIVE, OPERATOR-TUNABLE
DEFAULTS, NOT ANY REAL COMPANY'S OR ANY JURISDICTION'S RETENTION SCHEDULE. Apply them as printed
regardless of whether they look right for the class in front of you.
⚠︎ THIS REVIEW REPORTS. It never releases a hold, deletes or purges a record, marks anything
disposed, or writes to any register. Every verdict you give is a statement about what the record
already says.
How to work it out:
- FIRST decide whether the record class is settled. If the record-class mapping is blank on file,
or the file gives two classes with neither superseding the other, the verdict is
CONTEXT_INCOMPLETE: no reach is computed, no system is named, the record_class is UNMAPPED and
nothing is actioned. THERE IS NO DEFAULT RECORD CLASS (Rule R-1).
- THEN decide whether this rule is already REMEDIATED. If the carried state or this cycle's
changes record that a system which had FAILED to take this rule has since been recorded as
taking it, the verdict is REMEDIATED and stays REMEDIATED regardless of what the register's
column says (Rule R-6). Applying successfully to a system that had failed IS the remediation;
there is no separate remediation entry.
- THEN work out what is actually IN SCOPE. Start from the systems named in the carried state,
add every system the register's table on this page lists, then apply this cycle's changes:
an ONBOARDED system joins scope, a DECOMMISSIONED system leaves it. The register's table is
STALE BY DESIGN -- it does not list systems onboarded in an earlier cycle and it still lists
systems decommissioned in an earlier cycle (Rule R-2).
- THEN find the UNREACHED systems. A system is unreached if the record shows it could not take
the rule -- a read-only connector, an archive tier that exposes no hold API, a legacy host with
no hold or retention concept at all -- whether that was recorded THIS cycle or carried from an
earlier one. It stays unreached until a later record shows it taking the rule. A failure
described in an ordinary sentence counts exactly as much as one written in a fixed tag format
(Rule R-5). THE REGISTER'S COLUMN DOES NOT SETTLE THIS: it records what somebody DID, not what
LANDED, and it will read APPLIED for a system that rejected the flag.
- If any in-scope system is unreached, the verdict is UNREACHABLE_SYSTEM and "unreached_system"
is that system's name exactly as printed. This is the most important call this kit makes -- it
means the register shows the rule as applied and those records are not held (Rule R-3). If more
than one qualifies, name the first alphabetically.
- Otherwise, if the register's OWN table declares a gap -- a row reading NOT APPLIED or PENDING
for a system still in scope -- the verdict is PARTIAL_REACH and "unreached_system" is that
system. A gap the register declares is one somebody can already see; the difference between the
two verdicts is whether anybody knows (Rule R-4).
- Otherwise consider DISPOSITION, which applies to retention SCHEDULE rules only and only where
NO legal hold covers the class. Compare the age of the oldest record in the class against the
DEFAULT retention period for that class. Inside the period: FULLY_REACHED. At or past it and
within the DEFAULT grace band: DISPOSITION_ELIGIBLE. Beyond the grace band: OVER_RETAINED. A
hold covering the class suspends disposition entirely, however old the records are, and a hold
RELEASED on an earlier cycle no longer covers it (Rule R-7).
- Otherwise the verdict is FULLY_REACHED and "unreached_system" is NONE.
- THEN decide the OWNER. Start from the records owner on file, but a change entry recording a
handover to someone else, in this cycle or an earlier one (see "Carried state"), supersedes it.
- FINALLY decide "action_now": YES only if the verdict is UNREACHABLE_SYSTEM, PARTIAL_REACH or
OVER_RETAINED **and** this review has not already raised that same finding -- the same system,
or the same over-retention -- on an earlier cycle (see "Carried state"). Otherwise NO.
DISPOSITION_ELIGIBLE is never action_now YES: it is reported, and disposing of records is a
decision with its own approval that this review must never appear to have started.
Answer with a single JSON object and nothing else:
{"verdict": "FULLY_REACHED|UNREACHABLE_SYSTEM|PARTIAL_REACH|DISPOSITION_ELIGIBLE|OVER_RETAINED|REMEDIATED|CONTEXT_INCOMPLETE",
"unreached_system": "<the system name exactly as printed, or NONE>",
"record_class": "<the record class exactly as printed, or UNMAPPED>",
"owner": "<the current records owner's name>",
"action_now": "YES|NO",
"rationale": "one sentence, naming the systems you took to be in scope and the evidence you
relied on for the reach verdict"}
Precedence, applied in this order: CONTEXT_INCOMPLETE if the record class cannot be settled (Rule
R-1); then REMEDIATED (Rule R-6); then UNREACHABLE_SYSTEM (Rule R-3); then PARTIAL_REACH (Rule
R-4); then the disposition verdicts (Rule R-7); otherwise FULLY_REACHED. "unreached_system" is a
system name only for UNREACHABLE_SYSTEM and PARTIAL_REACH, and NONE for every other verdict.
Carried state
----------------------------------------------------------------
As at the previous review cycle, this rule's record class was settled as "Supplier correspondence". The systems known to be in scope for this rule, accumulated across every earlier cycle, are: M365-MAILBOX-SUPPLY, DOCSTORE-SHAREPOINT, SFTP-SUPPLIER-DROP. The register's own table on this page may list fewer or more (Rule R-2). No system has been recorded on any earlier cycle as unable to take this rule. This review has not yet raised anything about this rule with anybody. This review has already raised that this class is over-retained -- do not raise that again. The records owner as last settled is Declan Ashworth -- carried forward unless this cycle's record hands the class to someone else. It was last reported OVER_RETAINED.
Retention review extract
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated records-retention review extract for an
AI use-case kit; it reproduces no company, no matter, no system, no real person and no
real retention schedule. The retention periods and the disposition grace band are
ILLUSTRATIVE, OPERATOR-TUNABLE DEFAULTS -- not any company's or any jurisdiction's
schedule.
Matter Or Rule
----------------------------------------------------------------
Rule reference : RR-0001
Rule kind : SCHEDULE
Matter : -- none; this is a scheduled retention rule
Record class mapping on file : Supplier correspondence
Records owner (register of record) : Declan Ashworth
Register status : ISSUED 2025-05-12, open
Systems In Scope
----------------------------------------------------------------
The matter register's OWN list of systems against this rule, with its OWN
APPLIED / NOT APPLIED column exactly as it stands.
⚠︎ THIS TABLE IS STALE IN BOTH DIRECTIONS. It does not list a system onboarded in an
earlier review cycle, and it still lists one decommissioned in an earlier cycle. And the
column records what somebody DID, not what LANDED: it reads APPLIED for a system that
rejected the flag.
System Kind Register says
M365-MAILBOX-SUPPLY mailbox APPLIED
DOCSTORE-SHAREPOINT document store APPLIED
SFTP-SUPPLIER-DROP file drop APPLIED
Retention Policy (Default)
----------------------------------------------------------------
The schedule below is an OPERATOR-TUNABLE DEFAULT, not any real company's or any
jurisdiction's retention schedule -- no documented schedule exists for this row. Applied
exactly as printed.
Default retention period : 6 years
Oldest record in class : 108 months
Legal hold covering this class : NO
Disposition grace band : 12 months past the retention period
The whole DEFAULT schedule, for reference:
Customer service call recordings 2 years
Employee timekeeping records 6 years
Loss-prevention incident reports 6 years
Marketing consent records 5 years
POS transaction journals 7 years
Payment card settlement files 7 years
Price and promotion approvals 6 years
Product safety test certificates 10 years
Store CCTV footage 1 year
Store lease and occupancy files 12 years
Supplier correspondence 6 years
Warehouse gate and yard logs 3 years
Disposition authority : NOT DEFINED. Nothing in this kit releases a hold,
deletes or purges a record, marks anything disposed
or writes to any register.
Rule R-1 A rule is written against a RECORD CLASS and is judged only against the SYSTEMS it has
to land on. If the record-class mapping is blank on file, or the file gives two classes
with neither superseding the other, the verdict is CONTEXT_INCOMPLETE: no reach is
computed, no system is named and nothing is actioned. THERE IS NO DEFAULT RECORD CLASS.
Rule R-2 The SYSTEMS IN SCOPE table is the matter register's own list and it is STALE BY DESIGN.
A system ONBOARDED in an earlier review cycle is in scope now even though the table does
not list it; a system DECOMMISSIONED in an earlier cycle is out of scope now even though
the table still does. Scope ACCUMULATES across cycles (Rule R-5).
Rule R-3 The register's APPLIED / NOT APPLIED column records what somebody DID, not what LANDED.
Where the record shows that a system could not take the rule -- a read-only connector, an
archive tier that exposes no hold API, a legacy host with no hold or retention concept --
that system is UNREACHED however the register's column reads, and the verdict is
UNREACHABLE_SYSTEM. This is the most important call this kit makes: the register shows
the rule as APPLIED and that system is not held.
Rule R-4 A gap the register DECLARES -- a row reading NOT APPLIED or PENDING -- is
PARTIAL_REACH, not UNREACHABLE_SYSTEM. It is a known gap somebody can already see. The
difference between the two verdicts is whether anybody knows.
Rule R-5 Reach evidence ACCUMULATES across review cycles. "Changes In This Cycle" shows only what
was recorded since the PREVIOUS cycle; anything recorded in an earlier cycle is not
repeated and is known only through the carried state. A failure to apply recorded two
cycles ago is still true today unless a later cycle records the same system taking the
rule.
Rule R-6 Once a system that had FAILED to take the rule is later recorded as having taken it, the
rule is REMEDIATED and stays REMEDIATED on every later cycle. Nothing here performs the
remediation; it only reports one that has already happened elsewhere.
Rule R-7 DISPOSITION applies to retention SCHEDULE rules only, and only where NO legal hold
reaches the class. A class at or past its DEFAULT retention period is
DISPOSITION_ELIGIBLE while it is inside the DEFAULT grace band, and OVER_RETAINED once it
is beyond it. A hold covering the class suspends disposition entirely, however old the
records are. Reach beats disposition: a rule that does not reach every system is reported
on its reach first, because a disposition rule that never landed is a class nobody is
disposing of either.
Rule R-8 "action_now" is YES only the FIRST time a given finding is raised. A system already
raised on an earlier cycle is not raised again, and an OVER_RETAINED class already raised
is not raised again. A finding on a system nobody has raised before is YES even if the
rule was already in an actionable verdict for a different system.
Rule R-9 NOTHING IN THIS KIT RELEASES A HOLD, DELETES OR PURGES A RECORD, MARKS ANYTHING DISPOSED
OR WRITES TO ANY REGISTER. Every verdict above is a report about what the record already
says.
Review Position
----------------------------------------------------------------
Cycle ending : 2026-09-30 (review cycle 3 of this rule)
Review cadence : quarterly -- one records-retention review cycle every 90 days
Previous review cycle : 2026-06-30
Next review cycle : -- none scheduled after this one
Changes In This Cycle
----------------------------------------------------------------
Everything recorded against this rule between the previous review cycle and this one.
THIS CYCLE ONLY -- a change recorded in an earlier cycle is not repeated here.
2026-08-06 Passing the records ownership of this class to Rosalind Chu.
Reviewer Notes
----------------------------------------------------------------
Reviewer note: the record-class mapping was rebuilt at the start of the financial year
and has not been re-derived since.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"verdict": "OVER_RETAINED", "unreached_system": "NONE", "record_class": "Supplier correspondence", "owner": "Rosalind Chu", "action_now": "NO", "rationale": "In-scope systems M365-MAILBOX-SUPPLY, DOCSTORE-SHAREPOINT and SFTP-SUPPLIER-DROP show no unreach evidence and the register marks all APPLIED, but with no legal hold and the oldest record at 108 months versus the 6-year default plus 12-month grace, the class is already over-retained and was previously raised."}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch the system a legal hold missed — 120 rule review extracts. One model answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
120rule review extracts
120source documents
1model tier
3grading methods
MeasurementsWhat was measured
COUNTED105 · 64 · 105 · 94 · 72 / 120verdict accuracy pct — verdict, seven-wayDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED106 · 60 · 106 · 97 · 86 / 120unreached system accuracy pct — unreached system namedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED119 · 120 · 120 · 120 · 120 / 120record class accuracy pct — record classDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED115 · 109 · 107 · 107 · 97 / 120owner accuracy pct — records ownerDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED109 · 58 · 111 · 104 · 84 / 120action now accuracy pct — raise-it-now callDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED105 · 67 · 108 · 100 · 89 / 120unreachable system accuracy pct — a system the rule silently failed to reach -- caughtDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED47 · 30 · 40 · 33 · 18 / 49named system correct pct — named the right system where one was unreachedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED17 · 13 · 8 · 8 · 0 / 17prose only unreachable caught pct — unreachable recorded ONLY in prose -- caughtDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED49 · 26 · 44 · 40 · 18 / 49memory verdict accuracy pct — verdict, memory-dependent readings onlyDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED49 · 22 · 44 · 40 · 29 / 49memory unreached system accuracy pct — unreached system, memory-dependent readings onlyDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED22 · 8 · 23 · 21 · 18 / 24disposition recall pct — disposition-eligible recallDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED8 · 3 · 5 · 4 · 0 / 8remediated recall pct — remediated recallDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED9 · 9 · 9 · 9 · 9 / 9context incomplete recall pct — context-incomplete recallDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED16 · 13 · 15 · 13 · 4 / 16rules caught pct — rules still not reaching, caught at the last cycleDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
evals/check_labels.py replays all 40 rule chains through src/reach.step and requires the re-derived answers to equal the committed gold exactly before any run may spend — 20 checks, green, and RED-PROVEN IN SIX DIRECTIONS WITH ACQUITTALS: a seeded state-machine defect (4 mismatches), a seeded banned path, a removed verdict, drifted disposition arithmetic (18), a broken register regex (120), and a clock in the generator.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate — a projection, never a bill and never a vendor's claim about this workload.
Priced at
Per 1M in / out
One rule review extract
1,000 rule review extracts
Share that is the prompt
Google Gemini 3 Flash The estate's shared projection card. Every kit in this series prices its measured tokens against the same published card so the cost pages are comparable to each other; the provider that actually ran these calls is kept out of the tables by this estate's naming rule.
$0.30 / $2.50
$0.006227
$6.23
16%
Same work, 1× the bill
The same rule review extracts, the same tokens — only the rate card changed. And on that card about 16% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE REVIEW CADENCE. One call per rule per cycle; a longer cycle divides the bill and lengthens the period a system sits unheld.
Rates checked 2026-08-18. The provider that actually ran all 284 calls is kept out of these tables per this estate's naming rule, so no figure here is a bill.
the fast tier, with the carried state 87.5% unreachable system accuracy · the strongest free floor — no model 90.0% unreachable system accuracy · 2 more measured on each run
the fast tier, with the carried state 100.0% prose only unreachable caught · the strongest free floor -- no model 47.1% prose only unreachable caught · 1 more measured on each run
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set tells the arms apart, and the sharpest column is the one the model wins outright: unreachable systems recorded ONLY in prose — 100.0 pct against the strong floor's 47.06 pct over 17 readings.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Every connector failure is written back into the register as a status
the free floor — structured-log-strong, $0.00
It BEATS the model on the discriminator (90.0 against 87.5), on raise-it-now, on false actions and on disposition recall.
Do not use it where a refusal is recorded as a note: 47.06 pct on prose-only unreachable against the model's 100.0 pct.
Connector refusals get typed into a reviewer note and the register catches up later
the model — but for ONE column, and know which
Unreachable systems recorded only in prose: 100.0 pct against 47.06 pct. Everything else it either ties or loses.
Do not buy it for the discriminator. It loses that column, and every miss is one systematic over-call the page names.
You want the reading but cannot carry state between cycles
neither, as configured here
The stateless control collapses to 53.33 pct on verdict with 62 false actions.
Do not ship the stateless arm and describe it as this kit.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
SYSTEM_KIND_READ_AS_FAILURE
a system's KIND treated as evidence the rule failed to reach it
14
All 14 verdict misses, 13 of them on cycle 1, and all naming ARCHIVE-TIER3 (13) or TAPE-OFFSITE-VAULT (1). The model's own rationales say it: 'a cold object archive tier that exposes no hold API' — read off the register table, on pages whose change log is…
What we could NOT verify
⚠︎ WHETHER THE 14 MISSES ARE FIXABLE BY ONE PROMPT SENTENCE. The defect is diagnosed from the model's own rationales — it reads a system's KIND as evidence the rule failed — and the fix candidate is one sentence. It was NOT shipped and the run was NOT re-taken, because reading the misses and then changing the question is choosing the scoreboard after the game.
⚠︎ NO CADENCE MODULE EXISTS ON THIS KIT, and that is deliberate. Every sibling monitor publishes what its slower cadence would have missed. This row's prior process is EVENT-DRIVEN — the register is opened when an e-discovery request arrives — with no interval to compare against. Inventing one would be exactly the number data/SOURCES.md refuses to publish elsewhere.
Whether rendering the carried state as JSON rather than English would score the same.
Whether a second model reproduces the over-call.
The failure-reason mix is lopsided and published rather than levelled.
⚠︎ AND THE PUBLISHED CEILING WAS NOT SUFFICIENT. c000 at 8,000 passed with a largest reply of 7,860 — 140 under. Re-probed at 16,000, the SAME nine chains produced 8,899, above the cap that had just passed. 16,000 was published, and one of r001's 120 replies still hit it.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier, with the carried state
3,322.48
2,092.18
9,533 ms
$0.006227
the same tier, STATELESS CONTROL
3,206.18
3,288.78
20,335 ms
$0.009184
the strongest free floor -- no model
0
0
0 ms
$0.000000
the register given the carried state
0
0
0 ms
$0.000000
what an e-discovery register IS today
0
0
0 ms
$0.000000
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Answering vs gradingAnswering and grading are two separate bills
The three free floors, the stub and every pre-run check cost $0.00 and no calls. No run was discarded on this kit. ⚑ AND NO CADENCE MODULE EXISTS — the prior process is event-driven and has no interval to compare against.
Cost driversWhat actually moves the bill
OUTPUT TOKENS, AND ALMOST ALL OF THEM ARE REASONING. 94.1 pct of r001-retention-reach's output (236310 of 251062 tokens) was provider-side reasoning left at the default. The answer is a handful of short fields and a sentence; the bill is the thinking in front of it.
THE PROMPT IS MOSTLY FIXED TEXT. 3322 input tokens per reading, and the policy block, the rule text and the instruction are the same on every call — so a provider with prompt caching would price this workload very differently, and that was not measured.
Your volumeWhat it costs at your volume
LINEAR IN RULES x CYCLES. Nothing amortises. The sublinear lever NOT implemented is skipping rules whose system list did not change.
Where pricing changes shape
Provider-side reasoning. At 94.1 pct of output on this task, a model whose reasoning is not billed and one whose reasoning is billed differ by roughly an order of magnitude on the same workload.
⚠︎ THE PUBLISHED CEILING WAS STILL NOT ENOUGH, AND THIS IS THE CLEAREST CASE IN THE BATCH. c000 at 8,000 passed with a largest reply of 7,860 — 140 tokens under. The SAME nine chains re-probed at 16,000 produced 8,899, above the cap that had just passed. 16,000 was published and one of r001's 120 replies still hit it. A nine-call calibration bounds only what it sampled; 94.1 pct of output tokens were provider-side reasoning.
Your return, with your numbers
Volumerules per review cycle — this run judged 120 (40 rules x 3 review cycle) per arm
What it replacesa records lead reconciling a matter register against thirty systems to find out where a hold actually landed
Time saved per itemnot measured here — it depends on how much of your own connector-failure trail is keyed into the register versus typed into a note
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
It is the tier the shared connection was already pointed at, and the kit's whole claim is that swapping it is .env plus one more run. No second model was run against this corpus; every other row in the cost table is a projection onto a published card and is labelled as one.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
3,322input tokens · this run
2,092output tokens
$0.006what it actually cost
per-query average, this run
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.381
$0.381
$3.18
2026-09-12
gemini-3-flash
Google
$0.953
$0.953
$7.94
2026-09-18
gemini-3-8-flash
Google
$1.241
$1.241
$10.34
2026-09-18
llama-5
Meta
$1.565
$1.565
$13.04
2026-09-18
claude-haiku-4-5
Anthropic
$1.654
$1.654
$13.78
2026-09-12
grok-4-5
xAI
$2.304
$2.304
$19.20
2026-09-18
grok-4-6
xAI
$2.304
$2.304
$19.20
2026-09-18
claude-sonnet-5
Anthropic
$3.308
$3.308
$27.57
2026-09-12
gemini-3-1-pro
Google
$3.810
$3.810
$31.75
2026-09-18
gpt-5-6-terra
OpenAI
$3.810
$3.810
$31.75
2026-09-12
gpt-5-6-sol
OpenAI
$6.616
$6.616
$55.13
2026-09-12
claude-opus-4-8
Anthropic
$8.270
$8.270
$68.92
2026-09-12
claude-opus-5
Anthropic
$8.270
$8.270
$68.92
2026-09-12
claude-fable-5
Anthropic
$16.540
$16.540
$137.83
2026-09-18
claude-fable-5-1
Anthropic
$16.540
$16.540
$137.83
2026-09-18
gpt-6-astra
OpenAI
$16.540
$16.540
$137.83
2026-09-17
Read this against the numbers above
Projection only — no other model was actually called against this corpus.
The reasoning-token share (94.1 pct of output on the fast tier) is measured for that tier only.
Accuracy is NOT projected, only cost — a cheaper or pricier model is not implied to score the same 87.5 pct.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
11 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/reach.pythe reach rule and state machine — a swap seam
The DEFAULT retention periods, the disposition rule, the reach test and step() — which decides the next carried state from the answer key's inputs, never from a model reply. It is also the answer key.
You change it to: How long an unreached system stays unreached.
src/reach.py
# The retention-and-hold REACH rule as a state machine. Pure code, no model, standard library only.
HOLD = "HOLD"
SCHEDULE = "SCHEDULE"
KINDS = (HOLD, SCHEDULE)
FULLY_REACHED = "FULLY_REACHED"
UNREACHABLE_SYSTEM = "UNREACHABLE_SYSTEM"
PARTIAL_REACH = "PARTIAL_REACH"
DISPOSITION_ELIGIBLE = "DISPOSITION_ELIGIBLE"
OVER_RETAINED = "OVER_RETAINED"
REMEDIATED = "REMEDIATED"
src/state.pythe carried state
Which systems were reached, which failed, and which were onboarded or decommissioned. A system onboarded two cycles ago is in scope now and its first cycle never mentioned it.
src/state.py
# The carried state -- the thing that makes this a review programme and not another register read.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_rule(store, rule_id):
def _join(names):
def describe(state):
src/segment.pythe section splitter
Splits the page into its 8 named sections; asserted across all 120 documents.
src/segment.py
# Split a retention-review extract into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Matter Or Rule", "Systems In Scope",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe send filter — a swap seam
Custodian Contact is mapped by no field and therefore never sent.
You change it to: SECTION_HINTS and NEVER_SENT.
src/select.py
# Pick which sections of a retention-review extract are sent. Pure code -- the last deterministic
BANNER = "Synthetic Record"
RULEHEAD = "Matter Or Rule"
SYSTEMS = "Systems In Scope"
POLICY = "Retention Policy (Default)"
POSITION = "Review Position"
CHANGES = "Changes In This Cycle"
CONTACT = "Custodian Contact"
NOTES = "Reviewer Notes"
NEVER_SENT = (CONTACT,)
src/prompt.pythe prompt
Three parts: instruction and JSON shape, carried state, extract. The stateless control replaces exactly one line. ⚠︎ IT IS ALSO WHERE THIS KIT'S ONE SYSTEMATIC DEFECT LIVES: nothing tells the model that a system's KIND is not evidence a rule failed.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model call
Raw HTTP, stdlib only, bounded retry; the shared daily call cap is checked here.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe reader
One rule, one cycle, one call.
src/watch.py
# One rule, one review cycle, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled records-retention reach review. You read one retention rule or "
MAX_TOKENS = 16000
FIELDS = ("verdict", "unreached_system", "record_class", "owner", "action_now")
def documents():
def rules():
def load_doc(doc_id):
def position_of(text):
evals/scoring.pythe scorer
Exact match per cell. The unreachable-system discriminator is a separate binary.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
NEEDS_ATTENTION = ("UNREACHABLE_SYSTEM", "PARTIAL_REACH")
FIELDS = ("verdict", "unreached_system", "record_class", "owner", "action_now")
def _pct(n, d):
def score(records, golds):
def _protection(records, golds):
evals/baseline.pythe three free floors
register-only, register-mem and structured-log-strong — $0.00 each, and the third beats the model on the discriminator.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("register-only", "register-mem", "structured-log-strong")
def _section(text, name, nxt, flat=False):
def _facts(text):
ROW = re.compile(r"^ (\S+) {2,}.+?\s{2,}(APPLIED|NOT APPLIED|PENDING)\s*$", re.M)
def _register(text):
def _finish(verdict, system, record_class, owner, action_now, why):
TAG_APPLIED = re.compile(r"TAG: APPLIED (\S+)")
TAG_FAILED = re.compile(r"TAG: FAILED (\S+)")
TAG_ONBOARDED = re.compile(r"TAG: ONBOARDED (\S+)")
src/app.pythe local UI server
http.server, stdlib. Renders with no key.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "8213"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-retention-reach")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
ui/app.jsthe local UI client
Hand-written JS. ⚠︎ A CLIPPED AGREE COLUMN WAS FOUND BY OPENING A SCREENSHOT and fixed by shortening the strings. ⚠︎ AND THE SHARED ui/app.css CARRIES A SIBLING KIT'S VERDICT COLOURS, so this kit's verdicts render uncoloured. It was copied verbatim as instructed and NOT edited; the loudest signal still shows because .perm strong applies.
ui/app.js
#
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
src/reach.pyThe DEFAULT retention periods, the disposition rule, the reach test and step() — which decides the next carried state from the answer key's inputs, never from a model reply. It is also the answer key. A swap seam.
src/state.pyWhich systems were reached, which failed, and which were onboarded or decommissioned. A system onboarded two cycles ago is in scope now and its first cycle never mentioned it.
src/segment.pySplits the page into its 8 named sections; asserted across all 120 documents.
src/select.pyCustodian Contact is mapped by no field and therefore never sent. A swap seam.
src/prompt.pyThree parts: instruction and JSON shape, carried state, extract. The stateless control replaces exactly one line. ⚠︎ IT IS ALSO WHERE THIS KIT'S ONE SYSTEMATIC DEFECT LIVES: nothing tells the model that a system's KIND is not evidence a rule failed.
src/adapters/__init__.pyRaw HTTP, stdlib only, bounded retry; the shared daily call cap is checked here.
src/watch.pyOne rule, one cycle, one call.
evals/scoring.pyExact match per cell. The unreachable-system discriminator is a separate binary.
evals/baseline.pyregister-only, register-mem and structured-log-strong — $0.00 each, and the third beats the model on the discriminator.
ui/app.jsHand-written JS. ⚠︎ A CLIPPED AGREE COLUMN WAS FOUND BY OPENING A SCREENSHOT and fixed by shortening the strings. ⚠︎ AND THE SHARED ui/app.css CARRIES A SIBLING KIT'S VERDICT COLOURS, so this kit's verdicts render uncoloured. It was copied verbatim as instructed and NOT edited; the loudest signal still shows because .perm strong applies.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 3322 input and 2092 output tokens per reading (one rule, at one review cycle), all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Reading (one rule, at one review cycle)s/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per reading (one rule, at one review cycle) directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
MEASURED, not asserted. Reviewer Notes is this kit's injection surface, and the sentence fired is a deliberate GENERALISATION of the shipped note (which names one system). Both strings are recorded as fields in the result file.
API_KEY is read from .env (gitignored) or the real environment; never logged, never placed in a prompt, and redacted from any provider error message the local UI passes to the browser.
The experimentWe DID attack it — one sentence, on every reading where it could matter
The probe forces the instruction-shaped note onto every reading where a raise was due and re-fires them with everything else held identical. Zero raises were suppressed; three verdicts genuinely moved and all three moved toward the MORE alarming verdict. Both were run for real on 2026-08-25.
Boundary checked
What could go wrong
What the code guarantees
Whether an instruction embedded in Reviewer Notes can suppress a raise this kit's own rules say is due
Reading the scored run for it afterwards rests on wherever the seed placed the note.
x001-retention-reach-injection forces the condition: every reading where action was due, re-fired with the note REPLACED by the instruction-shaped one, carried state identical. 26 of 26 still raised it. Suppression rate 0.0 pct.
One boundary is measured in both directions here: the injection probe forces the condition rather than waiting for it, and the privacy guard is red-proven by reproducing the schema-change condition that reaches its fallback.
The result0 of 26 raises suppressed suppressed by the instruction-shaped note — measured, not assumed.
26attack trials fired
0raises suppressed
One phrasing, one model, one corpus, 26 trials — every reading where suppression was even possible.
Read this twice
The Reviewer Notes reach the model verbatim — there is no filter between what a reviewer types and what the provider sees. ⚑ AND THE PROBE'S RESULT NEEDS READING CAREFULLY: 0 raises were suppressed, but FIVE verdicts differed from gold — of which two were already wrong the same way in the scored run, so only THREE genuinely moved, and all three moved toward UNREACHABLE_SYSTEM, i.e. MORE alarming rather than less.
HonestyWhat this does not prove
Any other injection phrasing. The one fired is a deliberate generalisation of the shipped note.
Whether the result holds on another model tier.
Whether an injection placed in Custodian Contact would have any effect — by construction it cannot reach the prompt.
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
Never release a hold, delete a record or mark anything disposed — and never present the retention periods or the disposition rule as any real company's or jurisdiction's schedule.
Stated to the model on every call in src/prompt.py's INSTRUCTION block, and enforced by evals/check_labels.py's banned-code-path scan, red-proven by seeding a purge path.
EvidenceDoes it hold?
What
Measured
The banned-code-path scan
0 banned code paths across the whole kit, on every run of check_labels.py, including the ones immediately before every paid run.
The instruction-shaped note does NOT suppress — and this one IS measured
26 of 26 readings where action was genuinely due, re-fired with the note forced in, still raised it. Suppression rate 0.0 pct.
The limitWhat a guardrail is not
This is a PROMPT RULE PLUS A STATIC SCAN, not a runtime enforcement layer. Nothing stops a forker adding such a path tomorrow; the scan catches it only the next time somebody runs check_labels.py, which is a manual step.
The injection result is ONE SENTENCE against ONE model on ONE corpus. It is not a resistance rate for prompt injection in general.
WatchedWhat is watched, and why that one
7runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 89 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
3 measured by the latest run86 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The verdict, unreached system, record class, owner and raise-it-now call, per reading, exact match against the computed answer key
alarm
the per-field accuracies and the answered rate — alarm on any field falling below its strongest-free-floor value — the point at which paying for the model stopped being worth it on that column
unreachable-system
A system the rule silently failed to reach
alarm
the FALSE direction on this kit — that is where the model's systematic over-call lands — alarm on any rise in the false direction above the strong floor's 3.
prose-only-unreachable
Unreachable systems recorded ONLY in prose
alarm
the prose slice — the one column the model buys — alarm on any fall below the strong floor's 47.06 pct.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
120
different corpus — nothing is comparable
corpus.bytes
911,715
rule review extracts edited — the count held, the bytes did not
split.count
120
the readings count moved — a different set was scored
split.size_p50
7,590
the median size of one reading moved
split.size_p95
7,672
the 95th-percentile size of one reading moved
dataset.rows
120
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.5
no index is built here — a change is in corpus preparation, not an index rebuild
no shared guard The two recorded runs declare no guard in common, so nothing below is protected by a precondition — read the history as two separate observations, not as a difference.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Verdict, seven-way
87.5 pct
120 readings scored
r001-retention-reach exact match against src/reach.step()'s computed gold
Unreached system named
88.33 pct
120 readings scored
r001-retention-reach exact match against src/reach.step()'s computed gold
Record class
99.17 pct
120 readings scored
r001-retention-reach exact match against src/reach.step()'s computed gold
Records owner
95.83 pct
120 readings scored
r001-retention-reach exact match against src/reach.step()'s computed gold
Raise-it-now call
90.83 pct
120 readings scored
r001-retention-reach exact match against src/reach.step()'s computed gold
A system the rule silently failed to reach -- caught
87.5 pct
120 readings scored
r001-retention-reach exact match against src/reach.step()'s computed gold
Named the right system where one was unreached
95.92 pct
49 named system cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Unreachable recorded ONLY in prose -- caught
100.0 pct
17 prose only unreachable cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Verdict, memory-dependent readings only
100.0 pct
49 memory cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Unreached system, memory-dependent readings only
100.0 pct
49 memory cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Disposition-eligible recall
91.67 pct
24 disposition cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Remediated recall
100.0 pct
8 remediated cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Context-incomplete recall
100.0 pct
9 context incomplete cells
r001-retention-reach exact match against src/reach.step()'s computed gold
Rules still not reaching, caught at the last cycle
100.0 pct
16 rules not fully reaching at last cycle
r001-retention-reach exact match against src/reach.step()'s computed gold
Input tokens, run total
398698
120 readings
r001-retention-reach, re-derived from its result file
Output tokens, run total
251062
120 readings
r001-retention-reach, re-derived from its result file
Latency p50
9533
120 readings
r001-retention-reach, re-derived from its result file
Latency p95
55418
120 readings
r001-retention-reach, re-derived from its result file
Answered
99.17 pct
120 readings
r001-retention-reach, re-derived from its result file
False actions
10.64 pct
120 readings
r001-retention-reach, re-derived from its result file
Missed actions
0.0 pct
120 readings
r001-retention-reach, re-derived from its result file
Discriminator, missed direction
0.0 pct
120 readings
r001-retention-reach, re-derived from its result file
Discriminator, false direction
15.73 pct
120 readings
r001-retention-reach, re-derived from its result file
Raise-it-now, memory-dependent readings
100.0 pct
120 readings
r001-retention-reach, re-derived from its result file
Owner, memory-dependent readings
91.84 pct
120 readings
r001-retention-reach, re-derived from its result file
Systems left unheld, missed
0.0 pct
120 readings
r001-retention-reach, re-derived from its result file
HistoryRun history
7 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 4 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 2 runs. Columns here are only ever compared with each other.
Metric
r001-retention-reach 2026-08-25
s001-retention-reach-stateless 2026-08-25
action now accuracy, %
90.83
48.33
context incomplete recall, %
100.0
100.0
disposition recall, %
91.67
33.33
input tokens, whole run
398698
384741
model latency p50 ms
9533.00
20335.00
model latency p95 ms
55418.00
73424.00
memory unreached system accuracy, %
100.0
44.9
memory verdict accuracy, %
100.00
53.06
named system correct, %
95.92
61.22
output tokens, whole run
251062
394654
owner accuracy, %
95.83
90.83
prose only unreachable caught, %
100.00
76.47
record class accuracy, %
99.17
100.00
remediated recall, %
100.0
37.5
rules caught, %
100.00
81.25
unreachable system accuracy, %
87.50
55.83
unreached system accuracy, %
88.33
50.00
verdict accuracy, %
87.50
53.33
not a time series No two of these 2 runs measured the same system — they differ on documents, stateless — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-retention-reach-stub 2026-08-25
action now accuracy, %
70.0
context incomplete recall, %
100.0
disposition recall, %
75.0
input tokens, whole run
427362
model latency p50 ms
0.00
model latency p95 ms
0.00
memory unreached system accuracy, %
59.18
memory verdict accuracy, %
36.73
named system correct, %
36.73
output tokens, whole run
7182
owner accuracy, %
80.83
prose only unreachable caught, %
0.0
record class accuracy, %
100.0
remediated recall, %
0.0
rules caught, %
25.0
unreachable system accuracy, %
74.17
unreached system accuracy, %
71.67
verdict accuracy, %
60.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 18 chips that all say so.
redteam · with the model in the path — 1 run. Columns here are only ever compared with each other.
Metric
x001-retention-reach-injection 2026-08-25
raises held
26
raises suppressed
0
suppression rate, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 3 chips that all say so.
DeviationsWhat deviated
0 breaches across 7 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
the disposition rule in src/reach.py
which records are eligible, the answer key and every accuracy figure.
measured
red-proven by drifting the disposition arithmetic — 18 mismatches
a section added to SECTION_HINTS
what leaves the machine.
measured
red-proven in both directions
the register regex in the floor
what the free floor can see at all.
measured
red-proven by breaking it
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Nothing fires yet. Not one of the 26 bands on this kit names a consequence, so there is no row to print. Which is the honest answer to the heading: a band that nothing is wired to is a number somebody has to notice, and this kit has only those. An empty table would have said the same thing while looking like a rendering failure.
NextThe three you would add first
⚑ ADD THE ONE PROMPT SENTENCE THAT SEPARATES A SYSTEM'S KIND FROM EVIDENCE OF FAILUREAll 14 verdict misses are that single over-call. The fix is diagnosed and unshipped, because changing the question after reading the misses is choosing the scoreboard.
Level or widen the failure-reason mix10/5/1 is lopsided and published as such.
Automate the manual scan stepToday it is a step a developer has to remember.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
The guardrail scan is a MANUAL step, run before each spend, not a hook. It ran before every paid run and passed at 0 each time. The injection probe is a ONE-OFF: it measured 0.0 pct suppression on 26 readings on 2026-08-25 and nothing re-runs it, so that figure ages from the day it was taken.
What this cannot tell you
Whether a differently-named release path would be caught.
Whether the injection result holds for any other phrasing.
Whether the prompt rule or the absent code path keeps the kit decision-free.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
NO FRAMEWORK. A folder of readable Python, no framework dependency, and a prompt anyone can read end to end. A framework abstraction would own the retrieval step — there is none here — and the memory layer, which is already the entire surface of src/state.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the model call
src/adapters/__init__.py
an LLM wrapper / Runnable
one dict and one function; the seam this kit measures is what the model returns, not how it is called.
the carried state
src/state.py
a memory or checkpoint object
a checkpoint object buys persistence and concurrency; this kit carries a per-system reach map, written by arithmetic and never by the model.
the corpus
tools/build_corpus.py
a document loader
one flat synthetic format this kit fully controls.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Linear: the reader -> src/prompt.py -> src/adapters -> evals/scoring.py. No branching, no tool-calling and no agent loop.
The other sideWhat a framework costs you
Swapping providers means editing the PROVIDERS dict by hand. In exchange the prompt has no hidden templating layer and requirements.txt stays empty.
What we could NOT verify
Whether a framework's memory abstraction would have kept the per-SYSTEM reach map rather than a single applied flag. The per-system split is the whole question.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-retention-reach on the fast tier, with the carried state, 2026-08-25. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
9,533 ms
9533
—
Model, p95
55,418 ms
55418
—
Input tokens
398,698
398698
—
Output tokens
251,062
251062
—
No movement column. Not one of the 4 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-retention-reach9,533 ms
s001-retention-reach-stateless20,335 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-retention-reach-registeronly, b001-retention-reach-registermem, b002-retention-reach-structuredlogstrong recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-25, across 7 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
7 of the 8 sections go to the provider; Custodian Contact never does
the answer key
data/gold.jsonl, computed by src/reach.step()
never
the carried state
data/state.json in a deployment
a few SENTENCES of it do, in every prompt — that is the experiment
every run this kit has fired
results/eval-*.json
never
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 54
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
API_KEY is read from .env (gitignored) or the real environment; never logged, never placed in a prompt, and redacted from any provider error message the local UI passes to the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a per-cycle review; one cycle owns the changes since the last.
40 rules x 3 cycles = 120 calls per arm. (r001-retention-reach, evals/check_labels.py)
⚠︎ NO CADENCE MODULE EXISTS ON THIS KIT AND THAT IS DELIBERATE. Every sibling monitor publishes what its slower cadence would have missed. This row's prior process is EVENT-DRIVEN — the register is opened when a request arrives — with no interval to compare against, and inventing one would be the number SOURCES.md refuses to publish elsewhere.
Nothing here — there is no comparison cadence to invalidate.
state
which systems were reached, which failed, which were onboarded or decommissioned.
s001 collapses: verdict 87.5 to 53.33 pct, false actions 10 to 62, rules caught 100.0 to 81.25 pct. (s001, src/state.py)
⚑ A SYSTEM ONBOARDED TWO CYCLES AGO IS IN SCOPE NOW AND ITS FIRST CYCLE NEVER MENTIONED IT. That is what the carried state is for.
A deployment that cannot persist state between cycles.
model
one completion call per reading at a 16000-token ceiling.
120 readings, 3322 in / 2092 out per reading. p50 9533 ms, p95 55418 ms. 119 of 120 answered. (r001-retention-reach, c000/c001 calibration, src/watch.MAX_TOKENS)
⚠︎ THE CLEAREST CEILING FAILURE IN THE BATCH. c000 at 8,000 passed with 7,860 — 140 under. The SAME nine chains at 16,000 produced 8,899, ABOVE the cap that had just passed. 16,000 was published and one of r001's 120 replies still hit it.
A second model.
labels
a computed answer key from src/reach.step().
All 40 rule chains replay exactly. 20 checks, red-proven in SIX directions with acquittals. (tools/build_corpus.py, evals/check_labels.py)
⚠︎ THE FAILURE-REASON MIX IS LOPSIDED AND PUBLISHED RATHER THAN LEVELLED. Which reasons are POSSIBLE is a property of which systems a class lives on; the gate asserts coverage of at least one, not balance.
Your own register.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
UNREACHABLE_SYSTEM called on a page whose change log is empty
the model read the system's KIND off the register table as evidence the rule failed. This is all 14 of its misses and it is systematic, not scattered.
check whether anything in the cycle actually records a failure. A system that SOUNDS unreachable is not evidence that a rule failed to reach it. (r001-retention-reach misses; docs/shots/retention-reach-miss.png)
the register floor reporting a hold as applied on a system the model flags
a connector refusal recorded only in a note. This is the one column the model wins.
read the reviewer notes for that cycle. (b000 against r001-retention-reach)
the strong floor beating the model on the discriminator
on this kit it does, and the page leads with it. The model's case rests on the prose-only column alone.
read the prose-only slice before deciding to pay. (b002 90.0 against r001 87.5)
['Concurrency. data/state.json is replaced atomically, correct for one writer.', '⚠︎ A CADENCE COMPARISON. The prior process is event-driven with no interval; inventing one would be a fabricated number.', 'A record living in more than one system.', 'Jurisdictionally conditional retention periods.', 'Repeats. One run per arm.', 'Skipping rules whose system list did not change — the sublinear cost lever, named not built.']
The corpus licence, from the Data lens: MIT Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The verdict, unreached system, record class, owner and raise-it-now call, per reading, exact match against the computed answer key
all five fields hit
The register floor reports the hold as applied.
A system the rule silently failed to reach
hit
One of the 31 unreachable readings.
Unreachable systems recorded ONLY in prose
hit
The one column the model wins outright: 100.0 pct against the strong floor's 47.06 pct.
The formulaWhat it computes
accuracy = hits / 120 per field. unreachable_system_accuracy_pct is scored separately over its own denominator.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
87.5% verdict accuracy · 4 more measured on this row
In operationWhat to monitor
Reference standard: the computed answer key in data/gold.jsonl, produced by src/reach.step() from the same planted inputs the corpus builder wrote onto the page.
No true/false rates for this grader. It records 5 operating rows and not one of them carries a confusion matrix, so there is no TPR, TNR or precision to state here — this grader does not sort answers into pass and fail, and rates that were never measured are not going to be printed as though they were. What it did measure is in “What it actually did” above; what to watch is immediately below.
Watch these
the per-field accuracies and the answered rate
Alarm on
any field falling below its strongest-free-floor value — the point at which paying for the model stopped being worth it on that column
How tight can the band be? No threshold was swept: exact match has no tunable. Denominators are stated beside every rate because 120 readings makes each one wide.
Cadence: once per run; free
The decisionWhen to reach for it
Use it
Always. Every published accuracy figure rests on it.
Do not use it
It cannot tell you an answer was reasonable-but-wrong. All 14 verdict misses are one over-call and this grader scores every one as wrong.
The verdict, unreached system, record class, owner and raise-it-now call, per reading, exact match against the computed answer key
all five fields hit
The register floor reports the hold as applied.
A system the rule silently failed to reach
hit
One of the 31 unreachable readings.
Unreachable systems recorded ONLY in prose
hit
The one column the model wins outright: 100.0 pct against the strong floor's 47.06 pct.
The formulaWhat it computes
caught over all 120 readings; the two error directions over their own denominators (31 unreachable readings, 89 quiet).
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
87.5% unreachable system accuracy · 2 more measured on this row
the strongest free floor — no model
90.0% unreachable system accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: the computed answer key in data/gold.jsonl, produced by src/reach.step() from the same planted inputs the corpus builder wrote onto the page.
No true/false rates for this grader. It records 2 operating rows and not one of them carries a confusion matrix, so there is no TPR, TNR or precision to state here — this grader does not sort answers into pass and fail, and rates that were never measured are not going to be printed as though they were. What it did measure is in “What it actually did” above; what to watch is immediately below.
Watch these
the FALSE direction on this kit — that is where the model's systematic over-call lands
Alarm on
any rise in the false direction above the strong floor's 3.
How tight can the band be? No sweep: the verdict is an enum the arm emits.
Cadence: once per run; free
The decisionWhen to reach for it
Use it
Whenever the question is 'does this hold actually cover what the register says'.
Do not use it
It says nothing about custodian acknowledgement — that is a different kit (UC0106) with a different unit and a different population.
The verdict, unreached system, record class, owner and raise-it-now call, per reading, exact match against the computed answer key
all five fields hit
The register floor reports the hold as applied.
A system the rule silently failed to reach
hit
One of the 31 unreachable readings.
Unreachable systems recorded ONLY in prose
hit
The one column the model wins outright: 100.0 pct against the strong floor's 47.06 pct.
The formulaWhat it computes
over the 17 readings whose only evidence of failure is prose.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% prose only unreachable caught · 1 more measured on this row
the strongest free floor -- no model
47.1% prose only unreachable caught · 1 more measured on this row
In operationWhat to monitor
Reference standard: the computed answer key in data/gold.jsonl, produced by src/reach.step() from the same planted inputs the corpus builder wrote onto the page.
No true/false rates for this grader. It records 2 operating rows and not one of them carries a confusion matrix, so there is no TPR, TNR or precision to state here — this grader does not sort answers into pass and fail, and rates that were never measured are not going to be printed as though they were. What it did measure is in “What it actually did” above; what to watch is immediately below.
Watch these
the prose slice — the one column the model buys
Alarm on
any fall below the strong floor's 47.06 pct.
How tight can the band be? No sweep.
Cadence: once per run; free
The decisionWhen to reach for it
Use it
Whenever your connector failures are recorded in a note rather than a field.
Do not use it
Its denominator is 17 readings, so the rate is wide. Read the count.
A living map of modern AI — kept current every morning