A lawsuit's custodian list goes stale the moment nobody re-checks it: new people should be told to preserve evidence, and people already told can go quiet. This app re-reads the list on every scheduled check and flags who needs a notice or a nudge.
PresenterOpens the private repo. Visible to admins only.
For the legal operations teamBanking · Legal Services
Why it matters
Today's manual process, and the same job with the app
A legal operations or e-discovery analyst at a company facing a lawsuit or investigation.
✕Today's manual process
1Re-read the staff list manually every few weeks, against the case's departments.
2Try to remember who was already told and who already answered.
3Miss a quiet acknowledgement because nobody is tracking the clock on it.
4A missed name means evidence nobody was ever told to keep.
Every case re-checked from memory
✓With the app
1The staff list is re-read on every scheduled check, against the case's own rules.
2New names are matched and flagged for a notice the moment they qualify.
3Overdue acknowledgements are called out so nothing quiet gets forgotten.
4A closed case stops being asked about so nobody chases a matter that is already over.
Every case checked the same way
See it work
One real case, read by the app, step by step
MTR-0005-R1: a data-breach case where three engineers newly gained system access.
Catch a company's litigation hold going quietReference appBuilt to be shaped to your process
5
1The matter under review MTR-0005-R1: the case has grown, a notice is due today.
2Who moves onto the hold two names now need a notice, based on their access.
3The third new custodian EMP-0043 matches on the same access record.
4Who stays clear the rest of the roster keeps its normal status.
5Before any notice goes out a legal-hold administrator reviews this list first.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A litigation hold is not a document, it is a duty that runs for the life of the matter. New custodians surface as a case develops -- a document review names someone, a reorganisation moves a relevant function to a new department -- and a hold notice issued once at intake does not see any of that. Today a paralegal or compliance analyst re-reads a headcount export by eye every few weeks, tries to remember who was already told, and has no systematic way to notice that an acknowledgement has gone quiet. A custodian list typed once into a spreadsheet at issuance from a department name and a headcount export, and not revisited until someone remembers to.
Audience
A legal operations or e-discovery analyst deciding whether today's scheduled review needs to issue a new notice or escalate a quiet one, and outside counsel who will ask, if this is ever tested, how the custodian list was derived and when. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual hold matter files
The corpus is 72 hold matter files, 0.39 MB (txt 72). A real litigation-hold roster is built from a company's live HR and access-control records, cross-referenced against a real, often privileged, description of a real dispute. There is no public one and there will not be. A scrubbed real one would be worse: scrubbing removes exactly the departments-and-dates reasoning this kit exists to measure. Generating it lets the three hard cases (REORG, NAMED, ALIAS) be planted, named and counted rather than hoped for.
The corpus
The 72 hold matter filesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your hold matter files. That is the whole change — there is no database to migrate.
One hold matter file, as the model receives itMTR-0001-R1.txt · 1 of 72
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated legal-hold monitoring extract for an AI use-case kit; it reproduces no real matter, no real company, no real employee and no real litigation-hold policy.
Matter
----------------------------------------------------------------
Matter reference : MTR-0001
Matter type : AML/BSA Investigation
Opened : 2025-10-01
Relevant period : 2024-08-06 to 2025-08-03
Matter status : OPEN
Case Scope
----------------------------------------------------------------
Relevant departments (on file) : Compliance, Risk Management
Relevant systems (on file) : Case Management System, Core Banking System
Case narrative:
A suspicious-activity referral was escalated internally and is now the subject of a formal
AML/BSA investigation covering a set of retail accounts. Compliance and Risk Management
reviewed the underlying transaction pattern before the referral was filed.
Candidate Roster
----------------------------------------------------------------
Department and system-access fields below are AS OF THIS RUN'S SNAPSHOT (2025-10-06) and are NOT synchronised in real time between scheduled runs (Rule L-8).
EMP-0001 Priya Lindqvist Compliance
Compliance Analyst
Systems: Case Management System (since 2023-11-09), Core Banking System (since 2023-11-09), Email (since 2023-11-09)
EMP-0002 Marcus Duarte Compliance
Compliance Analyst
Systems: Case Management System (since 2021-10-26), Core Banking System (since 2021-10-26), Email (since 2021-10-26)
Abridged — the file continues.
The outcomeWhat a good result looks like
Every candidate who newly matches the matter's relevant departments, systems and time period gets flagged for a notice on the run that first sees them match, exactly once; every custodian whose acknowledgement is overdue gets escalated exactly once; and a matter that has been released stops being asked about.
And when it cannot
Three ways, and this row is HIGH risk because the first one is not recoverable after the fact. A MISSED custodian is data that is never preserved because nobody was ever told to preserve it -- if it is later deleted in the ordinary course, that is a spoliation exposure with no undo. An OVER-INCLUDED custodian is a false positive somebody reviews and clears -- real cost, recoverable. A MISSED escalation is a live compliance gap sitting unreviewed. The scored run made zero of the first: 0 of 22 matters that needed a new notice were missed, and 0 of 204 gold-custodian judgements were over-included. It also made zero missed escalations. Its real failure is 39 candidates re-flagged as newly overdue when they had already acknowledged -- the cheap direction of error -- concentrated entirely in the 90-day close-out review.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
deciding whether a matter needs a new hold notice issued today — the model -- with the carried state 100 pct discriminator accuracy, zero missed across 22 exercisable matters, and the cost of a miss on this row is a spoliation exposure with no undo.
a fast, free, zero-cost first pass to shrink the candidate pool before any model call — the roster-rule-mem floor 0 pct over-inclusion and 100 pct escalation recall for $0.00 -- it is a legitimate structural pre-filter, not a strawman.
deciding whether an already-noticed custodian's acknowledgement is genuinely overdue, at a long-gap review — the roster-rule-mem floor, or a rewritten carried-state prompt that states who HAS acknowledged the floor never over-escalates (0 false escalations); the model over-escalates 39 candidate-instances at the 90-day review specifically because the current prompt never states the acknowledged set in words.
At a glanceHow the whole thing runs
82%stage accuracy pct
14,846 msp50, end to end
$4.56per 1,000 hold matter files · OpenAI GPT-5.6 Luna
Run once, for real, on 2026-08-24. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch a company's litigation hold going quiet14 steps · 4 questions · run once, for real · 2026-08-24
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py's DEPARTMENTS/SYSTEMS/MATTER_TYPES at your own vocabulary and re-run it, or write your own data/corpus/*.txt keeping the eight section headings in src/segment.py::SECTIONS (the parser asserts all eight before a run may spend) and a matching data/gold.jsonl derived from src/legalhold.step(), never hand-typed. The measured figures on this kit's page are properties of THIS corpus and do not transfer.Corpus lens →
When is this the wrong choice?
Avoid: The model with no carried state (the stateless control): the same discriminator collapses to 43.06 pct with 82 pct false-alarm rate, because it cannot tell a persisting custodian from a new one without memory. That is the case against the best-fitting scenario (“deciding whether a matter needs a new hold notice issued today”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A recent departmental reorganisation the access-control system has not yet absorbed: the on-file roster shows the OLD department and nothing on the extract says it changed. Measured as unrecoverable on this corpus's 3 planted REORG matters (9 of 720 units) -- 0 of 9 caught cold, at issuance, by either arm. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether rewriting src/state.py::describe() to explicitly list who HAS acknowledged (not only who is pending or escalated) fixes the 39 run-3 false escalations -- named as the concrete next experiment, not built or re-scored in this run. 4 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-24 — r001-legal-hold. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no API_KEY renders every matter and reading via the checked-in corpus, gold answer key and the three free floors -- all pure Python, $0.00. The model column stays blank until a key is set; the UI's replay button shows what the scored run r001 actually answered, from the committed result file, labelled as a replay.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
14,846 msp50, end to end
94,072 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt (the fixed instruction, the carried-state sentences, and the matter extract), and the reply parsed to a matter stage, ten per-candidate hold statuses and an escalate list. p95 (94.1s) is 6.3 times p50 (14.8s) -- the tail is driven by how much carried-state and case-scope reasoning a reading needs, not by page length: every extract is within a few hundred bytes of every other.
Current processWhat it replaces
A custodian list typed once into a spreadsheet at issuance from a department name and a headcount export, and not revisited until someone remembers to.
Where it is not good enough
⚑ THE DISCRIMINATOR THAT MATTERS MOST ON A HIGH-RISK ROW IS PERFECT, AND THE PAGE LEADS WITH THAT. "Does at least one candidate need a new hold notice today" is 100.0 pct (22 of 22) with zero missed and zero false alarms. Per-candidate hold-status accuracy is 99.31 pct (715 of 720 units) and over-inclusion is zero across all 516 true-negative units. ⚠︎ THE 5 MISSES THAT REMAIN ARE THE ROW'S OWN DOCUMENTED OPEN ITEM, MEASURED RATHER THAN HYPOTHESISED. All 5 are candidates whose true custodian status (later confirmed by outside counsel) turns on a department the on-file roster does not show, because a reorganisation had not yet been absorbed into the access-control system (Rule L-8). Cold, at issuance, with no carried state to help: the model caught 0 of 3 on the matters where nothing on the page hints at it (REORG), and 1 of 3 where the case narrative names the person by role (NAMED). Once ANY run gets one of these six right, the carried state states it as settled and every later run inherits it correctly -- see Eval.repeat-shaped finding under could_not_verify. THIS IS NOT A MODEL DEFECT: no reader, human or model, derives a fact an extract does not contain, and a real deployment needs a human to check the access snapshot's freshness before treating a proposed custodian list as final (Data.breaks_on). ⚠︎ AND ONE REAL DEFECT WAS FOUND AND PUBLISHED RATHER THAN QUIETLY FIXED. At the 90-day close-out review, the model re-flags an already-acknowledged custodian as newly overdue in 13 of 19 eligible matters -- 39 candidate-level false escalations, all among CORRECTLY identified true custodians (zero come from over-inclusion). The carried-state sentence names who is pending and who is already escalated, but never states in words who has already acknowledged, and the model does not reliably infer "absent from both lists means fine" once 90 days have passed. This is the cheaper of the two escalation errors -- an extra review, never a missed one (escalation recall stayed 100 pct) -- and the fix (state who HAS acknowledged, not only who has not) was named and not built or re-scored here; see Eval.could_not_verify. ⚑ A HUMAN MUST VERIFY THE PROPOSED CUSTODIAN LIST BEFORE ANYONE ACTS ON IT. This kit produces a watchlist, never a hold notice -- see Architecture for the non-configurable guardrail.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt72jsonl1json1
24 matters, 72 readings across 3 scheduled runs — the full 10-person candidate roster re-read whole, every run
dept-only 41.67% (no memory) · dept-only-mem 68.06% · roster-rule-mem 79.17%
the strongest floor TIES escalation recall and over-inclusion at 100%/0%, $0.00
Recorded failureroster-rule-mem loses only to narrative: it reports all 5 ALIAS matters (15 readings) scope-INCOMPLETE rather than resolving them from the case narrative, missing 45 custodian units the model got right from the same sentence
72 readings x (6-way stage + 10 hold statuses + escalate set), free, no judge
missed custodian, over-inclusion and false escalation reported apart
Recorded failure39 of 720 units are false escalations, ALL at the 90-day close-out review (13 of 19 eligible matters) — the carried state never states who has ALREADY acknowledged, only who is pending or already escalated
100.0%does a notice need issuing today — the discriminator, zero missed
81.94% six-way stage vs the free floor's 79.17%
99.31% per-candidate accuracy; the 18 REORG/NAMED units score 72.22%
39 false escalations, all at the 90-day review — named, not patched
2026-08-24as of
It produces a watchlist for a legal-hold administrator to act on — which candidates newly need a notice, whose acknowledgement is overdue, whether the matter is still open — and never issues, escalates or releases a hold; there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is a small per-custodian record written by src/legalhold.step() from the arithmetic, never from the model's reply — the stateless control makes that concrete: stage accuracy falls from 81.94 to 43.06 pct and escalation recall drops from 100 pct to exactly zero. The clock is the second: three scheduled runs, and 48 of 72 readings are memory-dependent. ⚑ THIS ROW IS RISK=HIGH AND THE DISCRIMINATOR IS PUBLISHED FIRST BECAUSE OF IT. Does a candidate need a new hold notice today is 100.0 pct, 0 of 22 missed — the one failure direction this row's own risk rating cannot absorb, since a missed custodian is a spoliation exposure with no undo.
⚠︎ AND THE 5 MISSES THAT REMAIN ARE THE ROW'S OWN DOCUMENTED LIMITATION: 9 of 720 units encode a fact — a department the access-control system had not yet absorbed after a reorganisation — that the extract itself does not contain, by design. No reader, human or model, derives a fact the record does not have; a real deployment needs a human to check the access snapshot's freshness before treating a proposed custodian list as final. ⚑ ONE REAL DEFECT WAS FOUND AND PUBLISHED, NOT QUIETLY FIXED: at the 90-day close-out review the model re-flags an already-acknowledged custodian as newly overdue in 13 of 19 eligible matters, because the carried-state sentence never states who HAS acknowledged. This is the cheaper of the two escalation errors — an extra review, not a missed one — and it was not patched before this page was published.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER/BASE_URL/MODEL in .env -- any OpenAI-compatible or Anthropic endpoint
the corpus
tools/build_corpus.py
point it at your own matter types, departments, systems and roster; keep the eight section headings in src/segment.py::SECTIONS
the escalation window
src/legalhold.py
ESCALATION_DAYS, and Rule L-1's department/system mapping, both invented for this kit
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 24 matters x 3 scheduled runs = 72 extracts, 10 candidates each, from a fixed seed (SEED = 20260824). Plants three hard cases -- REORG, NAMED, ALIAS -- at named indices for reproducible coverage. Gold is computed by src/legalhold.step() from the matter model, never typed by hand.
the custodian arithmetic
src/legalhold.py
Rule L-1 through L-8 as pure code: department + system-access + date-overlap against the relevant period decides who matches; persistence, escalation, release and the six-way stage precedence are computed deterministically. No model, no judgement. ⚠︎ The eight rules, the matter types and the escalation window are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. A per-custodian record (notice-sent run, acknowledged, escalated) plus a released flag, written from the arithmetic and never from the model's reply, rendered into English for the prompt. ⚠︎ Removing it drops stage accuracy from 81.94 pct to 43.06 pct and escalation recall from 100 pct to exactly zero -- see Eval.repeat.
the section splitter
src/segment.py
Splits an extract into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 72 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Candidate Personal Details -- home address, personal mobile, national ID -- is mapped by no field and is subtracted unconditionally, so the one section whose subject is personally identifying data never leaves the machine even when a source-system rename makes every hint match nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state sentences, and the selected sections in document order. The stateless build replaces the sentences with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text.
the watch
src/watch.py
One matter, one scheduled run, one call. Parses run dates and structured fields off the page with a regex and parses NOTHING about which departments are relevant or who is a custodian -- deciding that is the entire task. Holds MAX_TOKENS = 20000, set from measurement (c000/c001 calibration).
the local UI
src/app.py
One matter, one scheduled run, its carried state and the free floor's verdict, on 127.0.0.1:8204. Renders with no key. Shows the carried sentences verbatim, the free floor's per-candidate answer beside the model's, and the withheld-section list as evidence.
the three free floors
evals/baseline.py
dept-only (a spreadsheet's department column, no memory, $0.00), dept-only-mem (the same, given the carried state), roster-rule-mem (full Rule L-1, structured fields only -- the column the model has to beat). All three read ONLY the structured roster fields, never the case narrative.
the scorer
evals/scoring.py
Pure code, exact match per cell. Scores the six-way stage, the SCOPING discriminator, per-candidate hold status (as both readings and finer-grained units), and the escalation set, with missed/false counted apart rather than averaged.
Where it breaks at scale
The roster is re-read WHOLE on every scheduled run rather than diffed against the last one, so cost and latency scale with (matters x candidates), not with what actually changed. A portfolio in the thousands of matters needs the candidate pool pre-filtered by a cheap structural pass before the model ever sees it -- which is exactly what roster-rule-mem already computes for free, and a real deployment would likely run it as a pre-filter rather than treat it purely as a rival column.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
With no API_KEY configured, the page still renders in full off the checked-in corpus and the free floor -- only the model column stays blank until a key is set.successOpen full size →MTR-0005-R1 (an ALIAS matter): the case-scope field is 'not classified -- see narrative', and the model correctly resolves EMP-0041/42/43 as new custodians by reading the narrative's alias phrasing. The free floor, reading only the structured field, reports the matter INCOMPLETE and misses all three.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
MTR-0002-R3 (the 90-day close-out review): the model re-flags EMP-0012 and EMP-0013 as newly overdue for escalation, even though both had already acknowledged their notices between run 1 and run 2. The free floor, applying Rule L-6 exactly, correctly reports no escalation needed. This is the one systematic defect this run found -- see Eval.taxonomy's OVERESCALATE_R3.failureOpen full size →
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
72hold matter files
0.39 MiBtxt 72
p50 5,606chars per legal-hold monitoring extract
$0.00setup · 0s
How it is cutWhat one legal-hold monitoring extract is
one matter x one scheduled run = one document
SetupWhat the setup figure measured
There is no index -- the full candidate roster is re-read whole on every scheduled run; nothing is embedded or chunked.
LicenceLicence
MIT
Bring your ownBring your own hold matter files
Point tools/build_corpus.py's DEPARTMENTS/SYSTEMS/MATTER_TYPES at your own vocabulary and re-run it, or write your own data/corpus/*.txt keeping the eight section headings in src/segment.py::SECTIONS (the parser asserts all eight before a run may spend) and a matching data/gold.jsonl derived from src/legalhold.step(), never hand-typed.
⚠︎ And what stops being true when you do: The measured figures on this kit's page are properties of THIS corpus and do not transfer. The clearest one to state out loud: a candidate's true department/system access on your own roster is exactly as current as your organisation's HR-to-access synchronisation, and this kit has no way to detect a lag it was not told about -- where that lag exists on your data, both the rule and the model will be wrong in the same direction this corpus's REORG matters are, and by design, not by defect.
What breaks it
A recent departmental reorganisation the access-control system has not yet absorbed: the on-file roster shows the OLD department and nothing on the extract says it changed. Measured as unrecoverable on this corpus's 3 planted REORG matters (9 of 720 units) -- 0 of 9 caught cold, at issuance, by either arm.
A matter whose relevant-department field is populated with a synonym or informal team name your own case-management system uses but this kit's prompt has not been told about -- the model reads prose reasonably well but was not measured against your organisation's specific vocabulary.
A candidate named only by an external contractor ID or a system account that does not appear on the HR roster at all -- this kit's Rule L-1 is keyed on employee department and system-access records, not on log-in identifiers.
More than ten live candidates per matter reliably fitting in one call's context at the published token ceiling -- untested past this corpus's fixed roster size.
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
Fixed instruction: rules L-1 to L-8, the required JSON shape, the stage precedence
4,305
1,076
Carried state -- who is already a custodian, pending, escalated or released, from the previous run
475
119
The legal-hold monitoring extract itself -- matter, case scope, candidate roster, this window's events, watch position, administration notes
4,484
1,121
Total
2,316
This is the cost lesson as arithmetic: of the 2,316 tokens assembled, 1,121 are extracts — 48% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Replayed from src/prompt.py::build() against MTR-0001-R1's committed carried-state (empty, this is run 1) and the corpus file on disk -- byte-identical to what evals/run.py actually sent, which check_labels.py's one-line-diff assertion is what makes trustworthy.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are the scheduled legal-hold custodian watch for one matter. It runs on issuance, a 30-day
follow-up and a 90-day close-out review, and re-reads the FULL candidate roster on every scheduled
run. You are reading ONE matter at ONE scheduled run.
The matter file, the case scope, the candidate roster, this window's notice/acknowledgement events
and the administration rules are reproduced in the extract below. Apply them exactly as written.
You cannot see the earlier scheduled runs; what is known about them is stated under "Carried state"
and is the only history available to you. Do not assume anything about earlier runs beyond it.
How to work it out:
- FIRST settle the matter's relevant departments and systems. Where "Case Scope" states them
directly, use those. Where it says the departments are not classified in the case-management
system, read the case narrative and identify the department(s) it describes, even where the
narrative uses an informal name for a team or a role rather than the roster's exact title or
department label. THERE IS NO DEFAULT: never assume a department the file does not support
(Rule L-4).
- THEN for EVERY candidate on the roster, decide CUSTODIAN or NOT_CUSTODIAN under Rule L-1: their
department AND system access must overlap the matter's relevant departments/systems, AND their
access to a relevant system must have been active at some point during the matter's relevant
period (their access start date on or before the period end, and their access end date -- if
any -- on or after the period start). A candidate already carried as a custodian from an earlier
run STAYS a custodian regardless of what today's roster shows for them (Rule L-2), unless the
matter has been released.
- THEN check whether the case narrative or roster notes name a specific person by role rather than
by employee ID -- for example a title described informally rather than by its exact roster
label -- and match that reference to the correct candidate by role, not by exact string.
- THEN apply this window's notice/acknowledgement events and the carried state: a candidate
already on hold who has now acknowledged is no longer pending; a candidate already on hold who
has NOT acknowledged and whose notice was sent 14 or more calendar days before this run's
snapshot date is OVERDUE and should be escalated, UNLESS the carried state says they were
already escalated on an earlier run (Rule L-6).
- THEN check for a Hold Release Memo recorded in this window. If one is recorded, every current
custodian is released this run (Rule L-5) and every candidate should be reported NOT_CUSTODIAN.
- If the case scope cannot be settled at all -- two scope memos naming different departments with
neither superseding the other, or no department and no narrative that resolves one -- report
matter_stage INCOMPLETE, every candidate NOT_CUSTODIAN, and escalate empty. Never guess a
department (Rule L-4).
Answer with a single JSON object and nothing else:
{"matter_stage": "SCOPING|MONITORING|ESCALATION_NEEDED|RELEASING|CLOSED|INCOMPLETE",
"candidates": [{"id": "<candidate id, exactly as listed>", "hold_status": "CUSTODIAN|NOT_CUSTODIAN"}, ...one entry per candidate on the roster, in the order listed...],
"escalate": ["<candidate id>", ...zero or more, ALWAYS a subset of candidates already on hold before this run...],
"rationale": "one sentence naming which candidates (if any) are newly a custodian this run and why, or why the matter is INCOMPLETE, RELEASING or CLOSED"}
Precedence for "matter_stage", applied in this order:
INCOMPLETE beats everything; then CLOSED if the matter was already released on an earlier run
(the carried state says so); then RELEASING if a Hold Release Memo is recorded in THIS window;
then SCOPING if at least one candidate becomes a custodian THIS run who was not one on the carried
state; then ESCALATION_NEEDED if at least one existing custodian is newly overdue; otherwise
MONITORING.
"escalate" lists ONLY candidates who were ALREADY custodians before this run and are newly overdue
this run -- never a brand-new custodian (they have not been sent a notice yet, so there is nothing
to be overdue on) and never a candidate already escalated on an earlier run (Rule L-6).
Carried state
----------------------------------------------------------------
No earlier scheduled run has identified any custodian on this matter. This is either the first appearance of this matter on the watch, or no candidate has yet matched Rule L-1. Nothing is known to be already noticed, acknowledged or overdue. Access and department fields on the roster are refreshed as of each scheduled run's snapshot date and are NOT synchronised in real time -- see Rule L-8.
Legal-hold monitoring extract
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
Every field below is invented. This is a generated legal-hold monitoring extract for an AI use-case kit; it reproduces no real matter, no real company, no real employee and no real litigation-hold policy.
Matter
----------------------------------------------------------------
Matter reference : MTR-0001
Matter type : AML/BSA Investigation
Opened : 2025-10-01
Relevant period : 2024-08-06 to 2025-08-03
Matter status : OPEN
Case Scope
----------------------------------------------------------------
Relevant departments (on file) : Compliance, Risk Management
Relevant systems (on file) : Case Management System, Core Banking System
Case narrative:
A suspicious-activity referral was escalated internally and is now the subject of a formal
AML/BSA investigation covering a set of retail accounts. Compliance and Risk Management
reviewed the underlying transaction pattern before the referral was filed.
Candidate Roster
----------------------------------------------------------------
Department and system-access fields below are AS OF THIS RUN'S SNAPSHOT (2025-10-06) and are NOT synchronised in real time between scheduled runs (Rule L-8).
EMP-0001 Priya Lindqvist Compliance
Compliance Analyst
Systems: Case Management System (since 2023-11-09), Core Banking System (since 2023-11-09), Email (since 2023-11-09)
EMP-0002 Marcus Duarte Compliance
Compliance Analyst
Systems: Case Management System (since 2021-10-26), Core Banking System (since 2021-10-26), Email (since 2021-10-26)
EMP-0003 Elena Whitfield Risk Management
Risk Analyst
Systems: Case Management System (since 2023-07-04), Core Banking System (since 2023-07-04), Email (since 2023-07-04)
EMP-0004 David Ibarra Compliance
Compliance Analyst
Systems: Case Management System (since 2025-10-14), Core Banking System (since 2025-10-14), Email (since 2025-10-14)
EMP-0005 Aisha Tanaka Risk Management
Risk Coordinator
Systems: Shared Drive (since 2022-03-21), Email (since 2022-03-21)
EMP-0006 Thomas Rousseau Marketing and Communications
Marketing Specialist
Systems: Email (since 2024-02-29), Shared Drive (since 2024-02-29)
EMP-0007 Yuki Petrova Consumer Lending Servicing
Consumer Specialist
Systems: Email (since 2021-11-26), Loan Origination System (since 2021-11-26), Core Banking System (since 2021-11-26)
EMP-0008 Carlos Ohara Marketing and Communications
Marketing Specialist
Systems: Email (since 2023-03-02), Shared Drive (since 2023-03-02)
EMP-0009 Fatima Njoroge Retail Banking Operations
Retail Specialist
Systems: Email (since 2022-07-23), Core Banking System (since 2022-07-23), Shared Drive (since 2022-07-23)
EMP-0010 Owen Kessler Marketing and Communications
Marketing Specialist
Systems: Email (since 2021-10-15), Shared Drive (since 2021-10-15)
Notice and Acknowledgement Events
----------------------------------------------------------------
Everything recorded against this matter between the previous scheduled run and this one. THIS WINDOW ONLY -- an event recorded in an earlier window is not repeated here.
No events recorded in this window.
Watch Position
----------------------------------------------------------------
Snapshot taken : 2025-10-06 (scheduled run 1 of this matter)
Watch cadence : issuance, then a 30-day follow-up, then a 90-day close-out review
Previous scheduled run : -- none, this is the first
Next scheduled run : 2025-11-05
Administration Notes
----------------------------------------------------------------
Roster compiled from the current HR and access-management extracts as of the snapshot date
above. This report is a watchlist only; it does not itself issue, escalate or release a
litigation hold.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"matter_stage": "SCOPING", "candidates": [{"id": "EMP-0001", "hold_status": "CUSTODIAN"}, {"id": "EMP-0002", "hold_status": "CUSTODIAN"}, {"id": "EMP-0003", "hold_status": "CUSTODIAN"}, {"id": "EMP-0004", "hold_status": "NOT_CUSTODIAN"}, {"id": "EMP-0005", "hold_status": "NOT_CUSTODIAN"}, {"id": "EMP-0006", "hold_status": "NOT_CUSTODIAN"}, {"id": "EMP-0007", "hold_status": "NOT_CUSTODIAN"}, {"id": "EMP-0008", "hold_status": "NOT_CUSTODIAN"}, {"id": "EMP-0009", "hold_status": "NOT_CUSTODIAN"}, {"id": "EMP-0010", "hold_status": "NOT_CUSTODIAN"}], "escalate": [], "rationale": "EMP-0001, EMP-0002 and EMP-0003 are newly custodians because they are in Compliance or Risk Management and had active access to a relevant system during the relevant period, while all other candidates lack relevant department or system access within that period."}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch a company's litigation hold going quiet — 72 hold matter files. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the matter model -- the relevant period, the relevant departments/systems per run, the candidate roster and this window's notice/acknowledgement events -- via src/legalhold.step(), rather than hand-authored. evals/check_labels.py re-derives all 72 rows from a fresh build_matter() + legalhold.step() and refuses to let a run spend if the committed key disagrees.
72hold matter files
72source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED59 · 31 / 72stage accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED72 · 31 / 72scoping discriminator accuracy pct — readings -- does a hold notice need issuing todayDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED715 / 720unit accuracy pct — per-candidate hold-status unitsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED702 / 702unit accuracy ordinary pct — units excluding the REORG/NAMED special candidatesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED13 · 2 / 18unit accuracy special pct — REORG/NAMED units -- see could_not_verifyDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED5 / 204missed custodian pct — gold-custodian units -- the spoliation-risk directionDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 516over inclusion pct — gold-not-custodian unitsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED4 · 0 / 4escalation recall pct — readings where an acknowledgement was genuinely overdueDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED35 / 48memory stage accuracy pct — memory-dependent readings (run 2 and run 3)Decided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction: evals/check_labels.py replays all 24 matter chains through a freshly-built matter model and src/legalhold.step(), and refuses to let a run spend if any of the 72 rows disagrees with the committed key. It also red-proves the privacy guard in both directions and asserts the stateful/stateless prompts differ on exactly one line.
3,435.86output tokens · the fast tier, with the carried state · 14,846 ms p50
3,140.39output tokens · the same tier, memory removed (THE CONTROL) · 15,050 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 1.0× as long, and lands one row apart on 72. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One hold matter file
1,000 hold matter files
Share that is the prompt
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.004560
$4.56
10%
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.011400
$11.40
10%
Anthropic Claude Haiku 4.5 a cheap frontier-family tier
$1.00 / $5.00
$0.019364
$19.36
11%
Anthropic Claude Sonnet 5 a mid frontier tier
$3.00 / $15.00
$0.058093
$58.09
11%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.193645
$193.64
11%
Same work, 42× the bill
The same hold matter files, the same tokens — only the rate card changed. And across all 5 cards between 10% and 11% of what you pay is the prompt this pipeline sends, not the answer it writes.
THE CANDIDATE-POOL SIZE. roster-rule-mem's structural pre-filter (department + system + date-overlap) is free and already narrows a large HR export down to a short candidate list before any model call -- running it as a pre-filter rather than a rival is the one change that moves this kit's bill by a multiple without touching accuracy on the ordinary units (100.0 pct either way).
Rates checked 2026-08-24. The provider that actually ran all 153 calls (72 scored + 72 stateless control + 9 calibration) is kept out of these tables per this estate's naming rule, so nothing here is what was paid.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors. The only money on this kit is the readings themselves.
The gradersOne way to grade, and why it is the only one
roster-rule-mem ties the model on escalation recall and over-inclusion (both 0/100 pct) and is 6.9 points behind on the discriminator (93.06 vs 100.0). Its stage accuracy (79.17 pct) is dragged down by a single named cause, not a general weakness: it reads only the structured department field, so a matter whose scope is typed in narrative rather than a dropdown (the 5 ALIAS matters, 15 readings) looks IDENTICAL to a genuinely disputed one -- it reports all 15 INCOMPLETE, missing 45 custodian units the model resolved correctly from the same sentence. The naive floor without the full rule (dept-only, no system/date check) scores 77.5 pct on units with 99 over-inclusions -- the cost of skipping the date-overlap check Rule L-1 requires.
the fast tier, with the carried state 81.9% stage accuracy · the same tier, memory removed (THE CONTROL) 43.1% stage accuracy · the strongest free floor, no model 79.2% stage accuracy · the same rule, WITHOUT the system/date-overlap check 68.1% stage accuracy · the same rule, WITHOUT memory 41.7% stage accuracy · 3 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set separates the arms cleanly: five arms scored through one scorer land at 99.31 / 97.36 / 91.25 / 78.19 / 77.5 pct unit accuracy on the same 720 units -- a 21.8-point spread between the model and the naive floor -- and they disagree on WHICH units as well as how many. The stateless control is the sharpest separation on the discriminator specifically: identical prompt but for one block, and the SCOPING call falls from 100.0 to 43.06 pct.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
deciding whether a matter needs a new hold notice issued today
the model -- with the carried state
100 pct discriminator accuracy, zero missed across 22 exercisable matters, and the cost of a miss on this row is a spoliation exposure with no undo.
the model with no carried state (the stateless control): the same discriminator collapses to 43.06 pct with 82 pct false-alarm rate, because it cannot tell a persisting custodian from a new one without memory.
a fast, free, zero-cost first pass to shrink the candidate pool before any model call
the roster-rule-mem floor
0 pct over-inclusion and 100 pct escalation recall for $0.00 -- it is a legitimate structural pre-filter, not a strawman.
trusting it as the FINAL answer on a matter whose scope is written in narrative rather than a dropdown field: it reports all such matters INCOMPLETE rather than resolving them, which is safe but escalates work a five-minute read would have closed.
deciding whether an already-noticed custodian's acknowledgement is genuinely overdue, at a long-gap review
the roster-rule-mem floor, or a rewritten carried-state prompt that states who HAS acknowledged
the floor never over-escalates (0 false escalations); the model over-escalates 39 candidate-instances at the 90-day review specifically because the current prompt never states the acknowledged set in words.
publishing the model's escalate list unreviewed at a close-out run -- see Business.not_good_enough.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
REORG_BLIND
Unrecoverable department mismatch (this row's documented limitation)
5
MTR-0003-R1: EMP-0026's on-file department is Marketing and Communications; gold says CUSTODIAN because outside counsel later confirmed their actual duties. Nothing on the extract explains the discrepancy. Model answered NOT_CUSTODIAN.
OVERESCALATE_R3
Re-flags an already-acknowledged custodian as newly overdue at the 90-day close-out review
39
MTR-0002-R3: EMP-0012 and EMP-0013 acknowledged their notices between run 1 and run 2 (recorded in that window's events). The carried state for run 3 lists them only under 'already a custodian,' never under an explicit 'already acknowledged' heading. The…
NAMED_COLD_MISS
Cold miss on a narrative role-reference, corrected once carried forward
2
MTR-0004-R1: the narrative names 'the branch supervisor who wrote the final performance review' (EMP-0036); model answered NOT_CUSTODIAN at issuance, then CUSTODIAN at every later run once the carried state stated it as already settled.
What we could NOT verify
Whether rewriting src/state.py::describe() to explicitly list who HAS acknowledged (not only who is pending or escalated) fixes the 39 run-3 false escalations -- named as the concrete next experiment, not built or re-scored in this run.
Whether a real HR/access-control synchronisation lag looks like this corpus's REORG scenario (a clean, silent department mismatch with zero other signal) or is usually detectable some other way on a real system (an audit trail, a provisioning ticket). The corpus states it as unrecoverable from the extract; a real environment may carry a weaker or stronger version of that fact.
Performance past 10 live candidates per matter in one call, or past 24 concurrently-open matters -- untested at this corpus's fixed size.
Whether disabling provider-side reasoning (never sent in this run; settings.thinking is null throughout) would hold accuracy at lower cost -- unmeasured, same open question sibling kits in this series carry.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
OpenAI GPT-5.6 Luna
Google Gemini 3 Flash
Anthropic Claude Haiku 4.5
Anthropic Claude Sonnet 5
Anthropic Claude Fable 5
the fast tier, with the carried state
2,185.18
3,435.86
14,846 ms
$0.004560
$0.011400
$0.019364
$0.058093
$0.193645
the same tier, memory removed (THE CONTROL)
2,121.38
3,140.39
15,050 ms
$0.004193
$0.010482
$0.017823
$0.053470
$0.178233
The token counts are measured on a real run and belong to this pipeline — top_k, chunk size, prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-24. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
evals/scoring.py and all three free floors make no model call -- the only money spent on this kit is the 153 readings themselves.
Cost driversWhat actually moves the bill
OUTPUT TOKENS. The reply carries a matter stage, ten per-candidate hold-status judgements and a written rationale for each reading -- output averages 3,435.86 tokens against 2,185.18 input, and output is priced 6x input on the projection card.
TEN CANDIDATES PER READING, FIXED. The unit of work is (matter x scheduled run x candidate), so the roster size is the single biggest lever on both tokens and what the reply has to justify.
THE CADENCE, WHICH MULTIPLIES EVERYTHING. One call per live matter per scheduled run. A caseload of 500 open matters on this three-run cadence is 1,500 calls per matter lifecycle; a monthly cadence instead of issuance/30-day/90-day would multiply it further for as long as each matter stays open.
THE RULE TEXT AND CASE NARRATIVE ON EVERY PAGE, deliberately: the rule the model applies is read off the page so a forker can replace it without touching the prompt.
Your volumeWhat it costs at your volume
LINEAR in matters x scheduled runs, and nothing here amortises -- there is no index and each reading is independent of every other matter. Ten times the caseload is ten times the calls at the same per-reading cost.
Where pricing changes shape
None measured on this run -- see Data.breaks_on for the untested ceiling above 10 candidates per matter, which is a quality question rather than a pricing one.
Your return, with your numbers
Volume1 matter, 3 scheduled runs, 10 candidates -- this kit's own measured setup
What it replacesa manual re-read of a headcount export at issuance and, inconsistently, at later checkpoints
Time saved per itemnot measured here -- see pages/roi.html to compute your own from your caseload and reviewer cost
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
157,333input tokens · this run
247,382output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 72 readings, one completion call each, one tier.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.328
$0.328
$4.56
2026-09-12
gemini-3-flash
Google
$0.821
$0.821
$11.40
2026-09-18
gemini-3-8-flash
Google
$1.046
$1.046
$14.52
2026-09-18
llama-5
Meta
$1.248
$1.248
$17.33
2026-09-18
claude-haiku-4-5
Anthropic
$1.394
$1.394
$19.36
2026-09-12
grok-4-5
xAI
$1.799
$1.799
$24.99
2026-09-18
grok-4-6
xAI
$1.799
$1.799
$24.99
2026-09-18
claude-sonnet-5
Anthropic
$2.788
$2.788
$38.73
2026-09-12
gemini-3-1-pro
Google
$3.283
$3.283
$45.60
2026-09-18
gpt-5-6-terra
OpenAI
$3.283
$3.283
$45.60
2026-09-12
gpt-5-6-sol
OpenAI
$5.577
$5.577
$77.46
2026-09-12
claude-opus-4-8
Anthropic
$6.971
$6.971
$96.82
2026-09-12
claude-opus-5
Anthropic
$6.971
$6.971
$96.82
2026-09-12
claude-fable-5
Anthropic
$13.942
$13.942
$193.64
2026-09-18
claude-fable-5-1
Anthropic
$13.942
$13.942
$193.64
2026-09-18
gpt-6-astra
OpenAI
$13.942
$13.942
$193.64
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of.
ACCURACY IS NOT PROJECTED. Every figure here is a price for the same token counts; nothing says another model would answer the same way, and on this task the strongest free floor already ties the measured model on escalation recall and over-inclusion for $0.00.
OUTPUT TOKENS DOMINATE AND REASONING DOMINATES OUTPUT. 694 of the example row's 974 output tokens (71.3 pct) are provider-side reasoning; output is priced 3x to 6x input on every card above, so this is the whole spread between rows.
THE REORG/NAMED LIMITATION IS A PROPERTY OF THE INPUT, NOT THE MODEL, AND TRAVELS TO EVERY ROW HERE. No model priced on this page was tested against this corpus's 9 unrecoverable REORG units; nothing says a more expensive model would do better on a fact the extract does not contain.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
11 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Three of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 24 matters x 3 scheduled runs = 72 extracts, 10 candidates each, from a fixed seed (SEED = 20260824). Plants three hard cases -- REORG, NAMED, ALIAS -- at named indices for reproducible coverage. Gold is computed by src/legalhold.step() from the matter model, never typed by hand.
You change it to: point it at your own matter types, departments, systems and roster; keep the eight section headings in src/segment.py::SECTIONS
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260824
MATTERS = 24
RULE = "-" * 64
CANDIDATES_PER_MATTER = 10
DEPARTMENTS = ("Compliance", "Risk Management", "Retail Banking Operations", "Commercial Lending",
src/legalhold.pythe custodian arithmetic — a swap seam
Rule L-1 through L-8 as pure code: department + system-access + date-overlap against the relevant period decides who matches; persistence, escalation, release and the six-way stage precedence are computed deterministically. No model, no judgement. ⚠︎ The eight rules, the matter types and the escalation window are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
You change it to: ESCALATION_DAYS, and Rule L-1's department/system mapping, both invented for this kit
src/legalhold.py
# The legal-hold custodian rule, as arithmetic wherever it can be arithmetic. Pure code, no
SCOPING = "SCOPING"
MONITORING = "MONITORING"
ESCALATION_NEEDED = "ESCALATION_NEEDED"
RELEASING = "RELEASING"
CLOSED = "CLOSED"
INCOMPLETE = "INCOMPLETE"
STAGES = (SCOPING, MONITORING, ESCALATION_NEEDED, RELEASING, CLOSED, INCOMPLETE)
CUSTODIAN = "CUSTODIAN"
NOT_CUSTODIAN = "NOT_CUSTODIAN"
src/state.pythe carried state
SEAM 2 -- the thing that makes this a monitor. A per-custodian record (notice-sent run, acknowledged, escalated) plus a released flag, written from the arithmetic and never from the model's reply, rendered into English for the prompt. ⚠︎ Removing it drops stage accuracy from 81.94 pct to 43.06 pct and escalation recall from 100 pct to exactly zero -- see Eval.repeat.
src/state.py
# The carried state -- the thing that makes this a monitor and not a one-time scoping memo.
CADENCE_NOTE = ("Access and department fields on the roster are refreshed as of each scheduled "
def describe(state, candidate_ids):
src/segment.pythe section splitter
Splits an extract into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 72 documents before a run may spend.
src/segment.py
# Split a legal-hold monitoring extract into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Matter", "Case Scope", "Candidate Roster",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector
Decides which sections reach the model. Candidate Personal Details -- home address, personal mobile, national ID -- is mapped by no field and is subtracted unconditionally, so the one section whose subject is personally identifying data never leaves the machine even when a source-system rename makes every hint match nothing.
src/select.py
# Pick which sections of an extract are sent. Pure code -- the last deterministic step before
BANNER = "Synthetic Record"
MATTER = "Matter"
SCOPE = "Case Scope"
ROSTER = "Candidate Roster"
EVENTS = "Notice and Acknowledgement Events"
POSITION = "Watch Position"
PERSONAL = "Candidate Personal Details"
NOTES = "Administration Notes"
NEVER_SENT = (PERSONAL,)
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state sentences, and the selected sections in document order. The stateless build replaces the sentences with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string
INSTRUCTION = """\
def _instruction():
def build(text, carried, candidate_ids, stateless=False):
def strip_instruction_for_diff(prompt_text):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text.
You change it to: PROVIDER/BASE_URL/MODEL in .env -- any OpenAI-compatible or Anthropic endpoint
src/adapters/__init__.py
# SEAM 1 -- the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe watch
One matter, one scheduled run, one call. Parses run dates and structured fields off the page with a regex and parses NOTHING about which departments are relevant or who is a custodian -- deciding that is the entire task. Holds MAX_TOKENS = 20000, set from measurement (c000/c001 calibration).
src/watch.py
# One matter, one scheduled run, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled legal-hold custodian watch. You read one matter's roster against "
MAX_TOKENS = 20000
FIELDS = ("matter_stage", "candidates", "escalate")
def documents():
def matters():
def load_doc(doc_id):
def position_of(text):
src/app.pythe local UI
One matter, one scheduled run, its carried state and the free floor's verdict, on 127.0.0.1:8204. Renders with no key. Shows the carried sentences verbatim, the free floor's per-candidate answer beside the model's, and the withheld-section list as evidence.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "8204"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-legal-hold")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def _describe(carried, candidate_ids):
evals/baseline.pythe three free floors
dept-only (a spreadsheet's department column, no memory, $0.00), dept-only-mem (the same, given the carried state), roster-rule-mem (full Rule L-1, structured fields only -- the column the model has to beat). All three read ONLY the structured roster fields, never the case narrative.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("dept-only", "dept-only-mem", "roster-rule-mem")
def _section(text, name, nxt):
def parse_depts(text):
def parse_systems(text):
def parse_period(text):
def parse_roster(text):
def parse_events(text):
def parse_position(text):
def review(text, carried=None, mode="roster-rule-mem"):
evals/scoring.pythe scorer
Pure code, exact match per cell. Scores the six-way stage, the SCOPING discriminator, per-candidate hold status (as both readings and finer-grained units), and the escalation set, with missed/false counted apart rather than averaged.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("matter_stage",)
LIVE_ATTENTION = ("SCOPING", "ESCALATION_NEEDED")
def _pct(n, d):
def score(records, golds, special_ids=None):
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 24 matters x 3 scheduled runs = 72 extracts, 10 candidates each, from a fixed seed (SEED = 20260824). Plants three hard cases -- REORG, NAMED, ALIAS -- at named indices for reproducible coverage. Gold is computed by src/legalhold.step() from the matter model, never typed by hand. A swap seam.
src/legalhold.pyRule L-1 through L-8 as pure code: department + system-access + date-overlap against the relevant period decides who matches; persistence, escalation, release and the six-way stage precedence are computed deterministically. No model, no judgement. ⚠︎ The eight rules, the matter types and the escalation window are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. A per-custodian record (notice-sent run, acknowledged, escalated) plus a released flag, written from the arithmetic and never from the model's reply, rendered into English for the prompt. ⚠︎ Removing it drops stage accuracy from 81.94 pct to 43.06 pct and escalation recall from 100 pct to exactly zero -- see Eval.repeat.
src/segment.pySplits an extract into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 72 documents before a run may spend.
src/select.pyDecides which sections reach the model. Candidate Personal Details -- home address, personal mobile, national ID -- is mapped by no field and is subtracted unconditionally, so the one section whose subject is personally identifying data never leaves the machine even when a source-system rename makes every hint match nothing.
src/prompt.pyThree parts: the fixed instruction, the carried-state sentences, and the selected sections in document order. The stateless build replaces the sentences with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text. A swap seam.
src/watch.pyOne matter, one scheduled run, one call. Parses run dates and structured fields off the page with a regex and parses NOTHING about which departments are relevant or who is a custodian -- deciding that is the entire task. Holds MAX_TOKENS = 20000, set from measurement (c000/c001 calibration).
evals/baseline.pydept-only (a spreadsheet's department column, no memory, $0.00), dept-only-mem (the same, given the carried state), roster-rule-mem (full Rule L-1, structured fields only -- the column the model has to beat). All three read ONLY the structured roster fields, never the case narrative.
evals/scoring.pyPure code, exact match per cell. Scores the six-way stage, the SCOPING discriminator, per-candidate hold status (as both readings and finer-grained units), and the escalation set, with missed/false counted apart rather than averaged.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 2185 input and 3435 output tokens per query at top-k 10, of which about 150 tokens are fixed (system prompt, chat framing, the question itself) and the rest is retrieved context. Change the retrieval depth and the input scales with it. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model and top-k follow the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Retrieval is free here and stops being free at scale.This kit scores every chunk in Python; a corpus a hundred times larger needs a vector store, which is infrastructure this arithmetic does not price.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py, which invents all of it from a fixed seed. In a real deployment, the field a party outside the legal-ops team could most plausibly influence is the Case narrative -- drafted by an analyst, sometimes pasted in from an intake email or outside counsel's summary.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged, never written to results/, and never rendered by the UI. src/app.py replaces it and the base URL with placeholders before any provider error reaches the browser.
The experimentWe did NOT attack it -- and this is the surface we left open
An indirect prompt injection needs a field somebody outside the process can write into, and this kit's closest candidate is the Case narrative -- sent, deliberately, because hiding it would hide the surface from the run supposed to measure it. Unlike a sibling kit in this series, no instruction-shaped sentence was planted in this corpus's narratives, so whether the model would follow one was neither exercised nor measured here. No attack was fired at this kit and none is claimed. The one boundary below was checked in code on 2026-08-24, and it is red-proven in both directions.
Boundary checked
What could go wrong
What the code guarantees
Does a candidate's home address, personal mobile or national ID ever leave the machine?
Every extract carries a Candidate Personal Details section -- ten employees' home addresses, personal mobile numbers and the last four digits of a national ID. A kit that sends "the document" sends all of it, to a third party, on every reading, forever.
src/select.py maps no answered field to that section and _fallback() subtracts it unconditionally, so it cannot be reached even when a source-system rename makes every hint match nothing. evals/check_labels.py measures BOTH directions before a run may spend: with the guard, 0 of 72 documents leak it; simulating the naive or list(secs) fallback every sibling kit in this series once shipped, 72 of 72 leak.
One boundary is measured in both directions -- the privacy guard. There is no second gate to report, because this kit has no field analogous to critical-date's Tenant Contact playing double duty as an attack vector; Candidate Personal Details is withheld outright and never analysed as an injection surface, since nothing routes it into the prompt at all.
The result0 attack trials, one boundary checked -- the privacy boundary red-proven by removing the guard and watching all 72 documents leak.
1field a real deployment's outside party could plausibly influence (sent, not hidden)
0attack trials fired
72documents that leak the personal-details section without the guard
0redaction holes found -- unlike critical-date, no provider-refused screenshot was captured on this run
The Case narrative is the field an outside party would most plausibly influence in a real deployment, and this kit sends it. Nothing measured what a planted instruction in it would do to a reading, because none was planted.
Read this twice
The Case narrative reaches the model verbatim, and no sentence in this corpus's shipped narratives was written to test whether an instruction embedded there would be followed. This kit produces a watchlist and cannot act on any instruction it reads -- see Guardrails -- but a suppressed or altered row is still the kind of loss this vertical cares about, and it was not measured.
HonestyWhat this does not prove
Whether a real deployment's Case narrative -- prose a legal-ops analyst writes freely, sometimes pasted from an intake email -- would carry an instruction the model follows. Not applicable to this synthetic corpus, and not measured.
Whether the guardrail holds against a code path named something evals/check_labels.py does not know. It asserts the absence of names it knows.
Whether the withheld section stays withheld under a corpus that carries an eighth section under a different name -- the selector's guard is a subtraction and should, but nothing tests a shape this corpus cannot produce.
Whether the departments, systems and escalation window -- all invented -- would need different handling if a forker pointed this at a real caseload. Almost certainly yes, and nothing here helps with it.
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No hold notice issued, escalated or released, non-configurable. This kit produces a matter stage, a per-candidate hold status and an escalate list. It never sends a notice, records an escalation email or files a release memo, and there is no setting that makes it.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writer in the kit is evals/run.py (results/*.json). There is no outbound path of any kind except the one completion call in src/adapters.
EvidenceDoes it hold?
What
Measured
Nothing in this kit issues, escalates or releases a hold
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for nine such names and passes at zero, with the checker itself exempted because it is the only file that has to spell them.
No default department is ever assumed
Both INCOMPLETE matters (6 of 72 readings) report every candidate NOT_CUSTODIAN rather than a guessed department, on the model, on all three floors, and in gold.
Candidate Personal Details never reaches the provider
0 of 72 with the guard, 72 of 72 without it, both measured before any run spent.
The model's answer never becomes the next run's memory
src/state.py is written only from gold's recorded outcome in evals/run.py, never from the model's own reply. A reading that errored still advances the carried state, so one transport failure cannot turn into three scored ones.
A run measured under a non-published token ceiling cannot be mistaken for a scored one
evals/run.py refuses --max-tokens unless the run id begins with c. Both ceiling runs here are c000 and c001 and neither is quoted as a score.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The arithmetic is correct whatever the model says, which means a wrong reading is a wrong watchlist row and not a blocked action. The guardrail stops the kit acting; it does not stop it being wrong -- see the 39 false escalations and 5 missed customers this run recorded, none of which the guardrail has any opinion about.
IT IS NOT A SCHEDULER. This is the half of a monitor the kit does not ship. evals/run.py is INVOKED, not woken, and nothing here detects a missed run.
IT IS NOT LEGAL ADVICE AND THE RULES ARE INVENTED. Whether a candidate is truly a custodian, and what a preservation duty actually requires, depends on the matter, the jurisdiction and outside counsel's instructions, none of which this kit models.
WatchedWhat is watched, and why that one
8runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 25 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
22 measured by the latest run3 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
The matter stage, all ten per-candidate hold statuses and the escalate list, per reading, exact match against the computed answer key
alarm
stage_accuracy_pct; unit_accuracy_pct; scoping_discriminator_accuracy_pct; escalation_recall_pct; missed_custodian; false_escalation — alarm on false_escalation or missed_custodian above zero on any arm about to be trusted unreviewed. The scored run has 39 of the former and 5 of the latter, both named in Eval.taxonomy.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
72
different corpus — nothing is comparable
corpus.bytes
404,819
hold matter files edited — the count held, the bytes did not
split.count
72
the legal-hold monitoring extracts count moved — a different set was scored
split.size_p50
5,606
the median size of one legal-hold monitoring extract moved
split.size_p95
5,906
the 95th-percentile size of one legal-hold monitoring extract moved
dataset.rows
72
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (answered 72, cadence issuance, then a 30-day follow-up, then a 90-day close-out review, documents 72, escalation_cells 4, gold_custodian_units 204, gold_not_custodian_units 516, matters 24, memory_cells 48, readings_scored 72, scoping_cells 22, stateless False, units_scored 720, units_special 18) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
does a notice need issuing today (the discriminator)
not yet known
72 readings
A band is the spread between repeats and this kit fired one run per arm. What IS known is the spread between ARMS on the same 72 readings: 43.06 (stateless) to 100.0 (stateful) pct.
per-candidate hold-status accuracy
not yet known
720 units
One run per arm: model 99.31 pct, strongest floor 91.25 pct, naive floor 77.5 pct.
escalation recall
not yet known
4 readings
100 pct on one run, over a small denominator -- one cell moves this figure 25 points.
false escalation
not yet known
720 units
39 on one run, all at the 90-day close-out review, all among correctly-identified true custodians.
Answered
100.00 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
False scoping (count)
0 on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
False scoping
0.00 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Input tokens, whole run
157,333 on r001-legal-hold
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Latency, median ms
14,846 ms on r001-legal-hold
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Latency, 95th percentile ms
94,072 ms on r001-legal-hold
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Memory-dependent stage accuracy
72.92 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Missed custodian (count)
5 on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Missed custodian
2.45 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Missed escalation
0 on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Missed scoping (count)
0 on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Missed scoping
0.00 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Output tokens, whole run
247,382 on r001-legal-hold
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Over inclusion (count)
0 on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Over inclusion
0.00 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Stage accuracy
81.94 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Unit accuracy ordinary
100.00 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
Unit accuracy special
72.22 pct on r001-legal-hold
72 readings
One run per arm, so no repeat spread exists yet; the figure is r001-legal-hold's own, re-derived from its result file by build/measured/runlog.py.
HistoryRun history
8 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-legal-hold-deptonly 2026-08-24
b001-legal-hold-deptonlymem 2026-08-24
b002-legal-hold-rosterrulemem 2026-08-24
answered, %
100.0
100.0
100.0
escalation recall, %
0.0
100.0
100.0
false escalation
0
32
0
false scoping
31
0
0
false scoping, %
62.0
0.0
0.0
input tokens, whole run
0
0
0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
1.00
memory stage accuracy, %
27.08
66.67
79.17
missed custodian
63
63
63
missed custodian, %
30.88
30.88
30.88
missed escalation
4
0
0
missed scoping
5
5
5
missed scoping, %
22.73
22.73
22.73
output tokens, whole run
0
0
0
over inclusion
99
94
0
over inclusion, %
19.19
18.22
0.00
scoping discriminator accuracy, %
50.00
93.06
93.06
stage accuracy, %
41.67
68.06
79.17
unit accuracy ordinary, %
79.49
80.20
93.59
unit accuracy, %
77.50
78.19
91.25
unit accuracy special, %
0.0
0.0
0.0
not a time series No two of these 3 runs measured the same system — they differ on floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 4 runs. Columns here are only ever compared with each other.
Metric
c000-legal-hold-calibration 2026-08-24
c001-legal-hold-ceiling 2026-08-24
r001-legal-hold 2026-08-24
s001-legal-hold-stateless 2026-08-24
answered, %
83.33
100.00
100.00
100.00
escalation recall, %
100.0
—
100.0
0.0
false escalation
2
3
39
0
false scoping
0
0
0
41
false scoping, %
0.0
0.0
0.0
82.0
input tokens, whole run
12994
6530
157333
152739
model latency p50 ms
20282.00
20465.00
14846.00
15050.00
model latency p95 ms
55842.00
48902.00
94072.00
87860.00
memory stage accuracy, %
66.67
50.00
72.92
14.58
missed custodian
0
0
5
16
missed custodian, %
0.00
0.00
2.45
7.84
missed escalation
0
0
0
4
missed scoping
0
0
0
0
missed scoping, %
0.0
0.0
0.0
0.0
output tokens, whole run
24742
11270
247382
226108
over inclusion
0
0
0
3
over inclusion, %
0.00
0.00
0.00
0.58
scoping discriminator accuracy, %
83.33
100.00
100.00
43.06
stage accuracy, %
66.67
66.67
81.94
43.06
unit accuracy ordinary, %
100.00
100.00
100.00
99.57
unit accuracy, %
100.00
100.00
99.31
97.36
unit accuracy special, %
—
—
72.22
11.11
not a time series No two of these 4 runs measured the same system — they differ on answered, documents, escalation_cells, gold_custodian_units, gold_not_custodian_units, matters, max_tokens, memory_cells, readings_scored, scoping_cells, stateless, units_scored, units_special — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-legal-hold-stub 2026-08-24
answered, %
100.0
escalation recall, %
0.0
false escalation
0
false scoping
31
false scoping, %
62.0
input tokens, whole run
167072
model latency p50 ms
0.00
model latency p95 ms
0.00
memory stage accuracy, %
27.08
missed custodian
63
missed custodian, %
30.88
missed escalation
4
missed scoping
5
missed scoping, %
22.73
output tokens, whole run
11328
over inclusion
3
over inclusion, %
0.58
scoping discriminator accuracy, %
50.0
stage accuracy, %
41.67
unit accuracy ordinary, %
93.16
unit accuracy, %
90.83
unit accuracy special, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 22 chips that all say so.
DeviationsWhat deviated
0 breaches across 8 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the carried state
stage 43.06 -> 81.94 pct, the discriminator 43.06 -> 100.0, escalation recall 0 -> 100, REORG/NAMED units 11.11 -> 72.22
measured
s001-legal-hold-stateless against r001-legal-hold
whether the floor checks system access and date-overlap, not just department
one calibration reading cut off at 8,000 with zero output; the same reading finished cleanly at 24,000 (6,348 tokens)
measured
c000-legal-hold-calibration against c001-legal-hold-ceiling
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
does a notice need issuing today (the discriminator)
nothing yet. This is the column with the widest arm spread on the whole board.
per-candidate hold-status accuracy
nothing yet.
escalation recall
nothing yet. 0 missed on 4 -- a small base rate this corpus does not stress heavily.
false escalation
nothing yet, and this is the one number on the board this run's own could_not_verify names a concrete fix for.
Answered
nothing yet — a second scored run is what would give this column a spread to fire on.
False scoping (count)
nothing yet — a second scored run is what would give this column a spread to fire on.
False scoping
nothing yet — a second scored run is what would give this column a spread to fire on.
Memory-dependent stage accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Missed custodian (count)
nothing yet — a second scored run is what would give this column a spread to fire on.
Missed custodian
nothing yet — a second scored run is what would give this column a spread to fire on.
Missed escalation
nothing yet — a second scored run is what would give this column a spread to fire on.
Missed scoping (count)
nothing yet — a second scored run is what would give this column a spread to fire on.
Missed scoping
nothing yet — a second scored run is what would give this column a spread to fire on.
Over inclusion (count)
nothing yet — a second scored run is what would give this column a spread to fire on.
Over inclusion
nothing yet — a second scored run is what would give this column a spread to fire on.
Stage accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Unit accuracy ordinary
nothing yet — a second scored run is what would give this column a spread to fire on.
Unit accuracy special
nothing yet — a second scored run is what would give this column a spread to fire on.
NextThe three you would add first
A scheduler, and something that notices when a scheduled run did not happenThe kit's whole argument is that re-reading the roster on a schedule catches what a one-time memo misses; a schedule nobody enforces is the same one-time memo with extra steps.
A real HR/access-control freshness checkThe REORG scenario (9 of 720 units, unrecoverable by design) is exactly the failure mode a freshness check on the access-sync timestamp would flag as 'do not trust this department field yet' rather than silently answering wrong.
A second reader on any matter the file cannot settle6 of 72 readings here are INCOMPLETE and every arm gets 100 pct of them right by declining to guess -- which means the machine reliably hands a human a queue, and nothing downstream of this kit works that queue.
An explicit 'already acknowledged' sentence in the carried-state promptThis run's one real defect -- 39 false escalations at the 90-day close-out review -- traces to the carried state naming who is pending and who is escalated but never who has acknowledged. Named in Eval.could_not_verify as the concrete next experiment, not built here.
Provenance on each custodian match, carried alongside itThe carried state says a candidate IS a custodian; it does not say which department/system match earned that call. A legal-ops analyst who has to defend the list to outside counsel needs the second half.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free, seconds) on any change to tools/build_corpus.py, src/legalhold.py, src/segment.py, src/select.py or src/prompt.py -- it is the gate that would catch a roster-parsing regression before a run spends. Re-run evals/run.py --floor roster-rule-mem (free) on any change to Rule L-1 or the escalation window to see the structural-only ceiling move before spending on the model.
What this cannot tell you
One run per arm. Whether the 0 missed-custodian / 0 over-inclusion / 100 pct escalation-recall record is stable across repeats is not measured.
Whether the guardrail holds against a code path named something the checker does not know.
Whether a planted instruction in the Case narrative could suppress or alter a row. The surface is sent and named; no attack was fired.
Whether rewriting the carried-state prompt to name who HAS acknowledged fixes the 39 false escalations -- named as the concrete next experiment, not built here.
Whether a matter added to the caseload mid-horizon, or closed and reopened, behaves sensibly. All 24 matters here run their full three scheduled reviews.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, datetime, argparse, urllib, http.server and concurrent.futures, all standard library.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior runs)
those carry a TRANSCRIPT; this carries a per-custodian RECORD -- notice-sent run, acknowledged, escalated -- written by arithmetic. You would gain persistence, concurrency and a query interface over history, and lose the thing that matters most on a HIGH-risk row: a legal-ops analyst can be shown 'EMP-0012 acknowledged on 2026-04-12' and check it against the notice log, and a checkpointed graph state cannot be checked against anything that plainly.
the model
src/adapters/__init__.py
LiteLLM, LangChain chat models, any provider-abstraction layer
you would gain dozens of providers, retries, streaming and callbacks. 130 lines of urllib is the whole abstraction here and it returns the token counts lens 05 publishes and lens 07 prices; on this kit 71.3 pct of the output is reasoning tokens (694 of 974 on the example row), which is exactly the field such layers most often drop or normalise away.
the custodian rule
evals/baseline.py
a rules engine (Drools, a business-rules DSL), or a trained classifier over HR records
you would gain something that generalises past this corpus's ten fixed candidate slots. You would lose the point of the floor, which is that it is readable in one sitting and a forker can see exactly where it breaks -- on this corpus, a matter whose department field is blank and only resolvable from prose.
the schedule
(not shipped)
cron, Airflow, Temporal, any durable scheduler
you would gain the half of a monitor this kit does not have -- runs that actually happen. It is the first thing to add on a real deployment, and this row's own guardrails.add_first entry says so; it is left out here because the product is a folder of readable Python and a scheduler would be the largest thing in the folder.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each matter is a chain of three scheduled runs and the only edge between them is a per-custodian record passed forward -- there is no branching, no loop and no second agent.
The other sideWhat a framework costs you
No scheduler. This is a watch that carries state and is still INVOKED rather than woken.
No persistence layer. The carried state lives for the length of one evals/run.py invocation; a real deployment needs to write it somewhere between scheduled runs, which this kit does not ship (see environment.state).
No provenance on the carried custodian match. It says WHO is a custodian, not which department/system pairing earned that call, and a legal-ops analyst defending the list to outside counsel needs both.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
The scheduling and persistence seams are the ones that matter most on a real deployment and they are the ones with no code at all, so nothing about them has been tried.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-legal-hold on the same tier, memory removed (THE CONTROL), 2026-08-24. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
14,846 ms
14,846 ms on r001-legal-hold
—
Model, p95
94,072 ms
94,072 ms on r001-legal-hold
—
Input tokens
157,333
157,333 on r001-legal-hold
—
Output tokens
247,382
247,382 on r001-legal-hold
—
No movement column. Not one of the 6 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
c000-legal-hold-calibration20,282 ms
c001-legal-hold-ceiling20,465 ms
r001-legal-hold14,846 ms
s001-legal-hold-stateless15,050 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-legal-hold-deptonly, b001-legal-hold-deptonlymem, b002-legal-hold-rosterrulemem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-24, across 8 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
legal-hold monitoring extracts
data/corpus/MTR-<n>-R<k>.txt -- 72 files, generated once from a fixed seed by tools/build_corpus.py
seven of the eight sections go to the provider in the prompt; Candidate Personal Details -- home address, personal mobile, national ID -- never does, by src/select.NEVER_SENT, and evals/check_labels.py measures both directions before a run may spend
the answer key
data/gold.jsonl -- one row per reading, computed, not typed
never. It is read by evals/scoring.py and src/app.py on your machine and no part of it is ever put in a prompt -- a key in the prompt would be the answer in the question
the carried state
held in memory for the length of one evals/run.py invocation; a real deployment would persist one small record per matter between scheduled runs, which this kit does not ship
a few SENTENCES of it do, in every prompt -- that is the experiment. They name who is already a custodian, who is pending, who is already escalated and whether the matter is released, and nothing else; there is no transcript and no earlier extract in them
every run this kit has fired
results/eval-*.json
never. They are written locally and committed to the public kits repo on purpose, so a reader with no key can replay what the scored run answered
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 35
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged, never written to results/, and never rendered by the UI. src/app.py replaces it and the base URL with placeholders before any provider error reaches the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
issuance, a 30-day follow-up, a 90-day close-out review -- three scheduled runs per matter. One run owns exactly what changed since the last one, and the six-way stage precedence (Rule L-7) is defined against it, not against a fixed margin.
24 matters x 3 scheduled runs = 72 calls per arm, 0 gaps; evals/check_labels.py asserts all 24 sequences are complete before a run may spend. Wall clock 261.9s at 12 workers for the scored run. (r001-legal-hold, evals/check_labels.py, src/legalhold.RUN_LABELS)
What a missed run costs is not measured here the way critical-date prices a missed month -- this corpus does not vary the cadence. What IS measured is what removing the run-to-run MEMORY costs regardless of cadence: escalation recall falls from 100 pct to exactly zero, because whether a notice is overdue cannot be known without knowing when it was sent.
every figure on this page is per THIS three-run cadence. A caseload run monthly instead of issuance/30-day/90-day would change which runs are memory-dependent and how many false-escalation opportunities exist at a long-gap review -- untested.
state
a small per-custodian record (src/state.py), written by src/legalhold.step() and never from a model reply: which candidates are already custodians, their notice-sent run, whether they have acknowledged, and whether they have already been escalated -- plus one flag for whether the matter is released. Rendered into a few English sentences for the prompt.
48 of the 72 readings are memory-dependent (run 2 and run 3). Removing the state moves stage accuracy 81.94 -> 43.06 pct, the SCOPING discriminator 100.0 -> 43.06 pct, and escalation recall 100 -> 0 pct exactly -- see results/eval-s001-legal-hold-stateless.json against results/eval-r001-legal-hold.json. (r001-legal-hold against s001-legal-hold-stateless)
Not measured as a ceiling on this corpus -- state size here is fixed (one record per custodian, at most 10 per matter) and never grows unboundedly the way a conversation transcript would.
the carried-state prose currently names who is pending and who is escalated but never who has acknowledged -- this is this run's one real defect (39 false escalations at the 90-day review). Rewriting the sentence to name the acknowledged set is the concrete next experiment; see Eval.could_not_verify.
model
one provider, one key, one model per run (SEAM 1, src/adapters/__init__.py). The published run used the fast tier at MAX_TOKENS=20000.
8,000 tokens truncated one calibration reading to zero output (c000-legal-hold-calibration); 24,000 finished the same reading cleanly at 6,348 (c001-legal-hold-ceiling). 20,000 is the published ceiling, with headroom for run-to-run variance. (c000-legal-hold-calibration, c001-legal-hold-ceiling, r001-legal-hold)
Not tested past 20,000 output tokens -- no reading in the scored run came within 4,000 tokens of the ceiling (largest was 16,263).
a model that reasons more verbosely than the fast tier, or one with a smaller context window, could need a different ceiling or hit the context limit on the assembled prompt (2,185 input tokens average) sooner -- untested.
labels
data/gold.jsonl, 72 rows, derived mechanically from src/legalhold.step() by tools/build_corpus.py -- never hand-typed.
evals/check_labels.py re-derives every row from a fresh build_matter() + legalhold.step() and refuses to let a run spend if the committed key disagrees -- 0 mismatches on the shipped corpus. (evals/check_labels.py)
Not applicable -- the label set is generated, not collected, so there is no volume ceiling to measure.
a hand-authored label set on a real caseload would need its own verification step; this kit's guarantee (the key replays from the same code that produced it) does not transfer to labels somebody typed.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
a candidate reported NOT_CUSTODIAN on one run and CUSTODIAN on the next, with the roster looking unchanged
either a genuinely new match (a late system-access grant, a scope expansion) or the carried state correcting an earlier miss once a later run's information settles it -- both are visible in this corpus. See MTR-0003 (REORG): missed at issuance, correctly CUSTODIAN from run 2 onward once the true_override fact is reflected in the carried record.
check Eval.taxonomy's REORG_BLIND and NAMED_COLD_MISS entries and the matter's own carried-state history in results/eval-r001-legal-hold.json's per_reading list. (results/eval-r001-legal-hold.json, MTR-0003-R1 against MTR-0003-R2)
matter_stage ESCALATION_NEEDED with an escalate list naming a candidate who acknowledged in an earlier window
the run-3 defect: the carried-state prose states who is pending and who is already escalated but never who has acknowledged, and the model does not reliably infer 'absent from both lists means fine' after a 90-day gap.
check Eval.taxonomy's OVERESCALATE_R3 entry and the free floor's answer on the same reading -- roster-rule-mem never makes this error (0 false escalations). (results/eval-r001-legal-hold.json, MTR-0002-R3)
matter_stage INCOMPLETE on a matter whose case narrative plainly describes a department
the reader (or the floor) is treating a blank structured field the same as a genuinely disputed one, rather than resolving it from prose. This is roster-rule-mem's own signature failure on all 5 ALIAS matters -- it is not a signal that the matter is actually unresolvable.
read the Case narrative yourself before trusting an INCOMPLETE call from a structural-only reader; the model resolves all 5 ALIAS matters correctly from the same sentence. (results/eval-b002-legal-hold-rosterrulemem.json against results/eval-r001-legal-hold.json, MTR-0005-R1)
['Concurrency past EVAL_WORKERS=12 / 24 matter chains -- this run never saturated the retry path.', 'A missed scheduled run. Nothing here detects one or marks the readings it would have produced as late.', 'A population that changes between runs. All 24 matters here run their full three scheduled reviews; none is added or closed mid-horizon outside the planted CLOSE2/CLOSE3 scenarios.', 'Whether disabling provider-side reasoning holds the accuracy. 694 of 974 output tokens on the example row (71.3 pct) were reasoning, and thinking was never sent.', 'Repeats. One run per arm, so no band on any figure -- see guardrails.bands.', 'Whether English or a structured format is the better rendering of the carried state. A design choice, not a measurement.', 'GPU or local-compute sizing -- not applicable; this kit does no local inference or embedding.', "Provider-side data retention on the configured endpoint -- not checked against any provider's policy."]
The corpus licence, from the Data lens: MIT Your corpus’s licence is yours to verify, and the labels you write are about your documents.
The matter stage, all ten per-candidate hold statuses and the escalate list, per reading, exact match against the computed answer key
Catch a company's litigation hold going quiet
PresenterOpens the private repo. Visible to admins only.
In one lineThe matter stage, all ten per-candidate hold statuses and the escalate list, per reading, exact match against the computed answer key
For each of the 72 readings, did the reply's matter_stage equal the key; for each of the 720 (reading, candidate) pairs, did hold_status equal the key; and for each reading, did the escalate set equal the key's. Stage and hold_status are words from a closed list; escalate is a set of candidate ids compared for exact equality.
$0.00per 1,000 hold matter files
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function scores the model, the stateless control and all three free floors.
The inputOne real row, seen by every grader
reading
MTR-0002-R3
candidate
EMP-0012
gold: hold status
CUSTODIAN
gold: should escalate
no
model: hold status
CUSTODIAN
model: escalated (WRONG)
yes
floor: hold status
CUSTODIAN
floor: escalated
no
note
EMP-0012 acknowledged their notice between run 1 and run 2. The model re-flags them as newly overdue at the run-3 close-out review anyway; the floor, applying Rule L-6 exactly, does not.
Grader
Verdict
Why
The matter stage, all ten per-candidate hold statuses and the escalate list, per reading, exact match against the computed answer key
hold status hit; escalate MISS
EMP-0012's hold status is CUSTODIAN and both arms get it right. The escalate call is where the model is wrong: EMP-0012 acknowledged their notice between run 1 and run 2, and the model re-flags them as newly overdue at the run-3 close-out review anyway. This is the published, unpatched defect -- the carried state never states who HAS acknowledged, only who is pending or already escalated, so it recurs in 13 of 19 eligible matters. The floor, applying Rule L-6 exactly, does not make it.
The formulaWhat it computes
unit_accuracy_pct = unit hits / 720. stage_accuracy_pct = stage hits / 72. A reading whose reply did not parse counts as a MISS on every field, never as an exclusion.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
81.9% stage accuracy · 3 more measured on this row
the same tier, memory removed (THE CONTROL)
43.1% stage accuracy · 3 more measured on this row
the strongest free floor, no model
79.2% stage accuracy · 3 more measured on this row
the same rule, WITHOUT the system/date-overlap check
68.1% stage accuracy · 3 more measured on this row
the same rule, WITHOUT memory
41.7% stage accuracy · 3 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by tools/build_corpus.py over the planted matter model at generation time and re-derived from src/legalhold.step() by evals/check_labels.py before any run may spend. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is known and stated as a limitation: the 9 REORG units encode a fact (later confirmed by outside counsel) that the extract itself does not support, by design.
Watch these
stage_accuracy_pct
unit_accuracy_pct
scoping_discriminator_accuracy_pct
escalation_recall_pct
missed_custodian
false_escalation
Alarm on
false_escalation or missed_custodian above zero on any arm about to be trusted unreviewed. The scored run has 39 of the former and 5 of the latter, both named in Eval.taxonomy.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. Denominators are printed beside every rate because several are small: 22 readings that must raise the discriminator, 4 with a genuine overdue acknowledgement, and 18 REORG/NAMED units, so a handful of cells move those figures by several points.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/legalhold.py, src/segment.py, src/select.py or src/prompt.py. Re-run the scored eval AND its stateless control together (paid, 144 calls) on any change to src/prompt.py or src/state.py. All three free floors are free and should be re-run on any change at all.
The decisionWhen to reach for it
Use it
The truth is known and every answered field is a word from a closed list or a set of ids from a fixed roster.
Do not use it
The truth is not known -- the normal state of a real caseload, where whether a person is truly a custodian is decided by outside counsel reviewing the matter. That is why this corpus is generated rather than captured.
A living map of modern AI — kept current every morning