Check an agency's expiring records for holds before destruction
A records series reaches the end of its retention, but a legal hold, an audit or an open records request may still freeze it. This app reads each review, names the hold that freezes the series, and flags one already queued for destruction.
PresenterOpens the private repo. Visible to admins only.
For the records officerCross-domain · Government & Public Sector
Why it matters
Today's manual process, and the same job with the app
A records officer at a government agency, working through series whose retention period has ended.
✕Today's manual process
1Open each review and read every line of the hold registry.
2Judge each hold's scope against the series' category, project and closed date.
3Follow every continued hold back to the one it continues, then check the dates.
4One miss destroys records under a live hold, or nothing is ever destroyed.
Every hold read and judged manually
✓With the app
1Each review is read, and twelve fields are filled in, each showing where it came from.
2Each hold's scope is judged against this series, and the hold that freezes it is named.
3Continued holds are followed back, so a released hold with a live successor still counts.
4Frozen series in the queue are flagged to pull out. A records officer approves every release.
The officer reviews proposals and flags only
See it work
One real case, read by the app, step by step
Series RS-2018-8284, Bell Harbor financial files: the officer's note says nothing is outstanding, but hold PR-2025-041 still covers it.
Check an agency's expiring records for holds before destructionReference appBuilt to be shaped to your process
5
1The record series Bell Harbor financial files, read straight from the review.
2Retention has ended December 2025, so the series is up for destruction review.
3The hold that binds PR-2025-041, an active hold that covers this series.
4What does not count the officer's note says nothing is outstanding. The hold says otherwise.
5Pull it out frozen, yet already queued for destruction. A records officer must act.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Check an agency's expiring records for holds before destruction
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A record series reaches the end of its retention period. Before anything can be destroyed, somebody has to establish that nothing still freezes it — an active legal hold, an audit hold, an open public-records request, or a longer-retention series that captures the same records under a different schedule item. The retention lookup is arithmetic. The hold check is not: a hold's scope is written in prose ('all correspondence relating to the Riverside project, 2019 onward') and has to be judged against this series' own metadata, and the two commonest errors go in opposite directions — freezing everything with a hold anywhere near it, so nothing is ever destroyed and the schedule stops meaning anything, or releasing a series whose hold was released and immediately continued by another. Someone opening each disposition review, reading every line of the hold registry, judging each hold's prose scope against that series' own category, project and closed date, following any 'continues the scope of' reference back to the hold it continues, and only then checking two expiry dates against the review date before deciding whether the series may go to a records officer for destruction approval.
Audience
Records officers, information-governance and archives teams who work a disposition queue, and the counsel who signs the destruction certificate. The decision they are making is whether to put a proposal in front of a person — never whether to destroy something. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual disposition reviews
The corpus is 55 disposition reviews, 0.04 MB (txt 55). Plain text, one format, invented rather than fetched — a real disposition review names a real agency's series, a real custodian and a real matter, and a records officer's own remark about a specific file, so it is not publishable under any licence. What is publishable is the STRUCTURE: jurisdictions publish general records schedules as (item code, series description, retention period), and that shape is modelled here with entirely invented contents. One format keeps the reader's attention on the judgement the kit is about — prose scope against structured metadata — rather than on parsing.
The corpus
The 55 disposition reviewsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your disposition reviews. That is the whole change — there is no database to migrate.
One disposition review, as the model receives itRDS-0001.txt · 1 of 55
Record Series
-------------
RS-2006-3779
Custodian Office
----------------
Fairview Municipal Archives
Series Title
------------
Personnel action files - Riverside
Record Category
---------------
personnel
Related Project
---------------
Riverside
Record Closed
-------------
2006-03
Retention Schedule
------------------
GS-22-06, 7 years
Retention Expires
-----------------
2013-03
Overlapping Series
------------------
none on file
Hold Registry
-------------
PR-2016-370 | active | Scope: all personnel records relating to the Riverside project, 2009 onward. | Successor: none
Disposition Queue
-----------------
queued
Officer Notes
-------------
Ordinary cycle item, no correspondence pending on it.
The outcomeWhat a good result looks like
A twelve-field extracted record per review, plus two derived answers: WHICH hold freezes this series (or none), and whether it may be proposed for disposition. Then one pure-code routing decision taken from two of those fields: a series that is frozen and already sitting in the destruction queue is flagged to be pulled back out. On the fast tier that was 660 of 660 cells exact, 55 of 55 binding holds identified and 55 of 55 eligibility verdicts right, with 16 of 16 frozen-and-queued series flagged and no false alarms.
And when it cannot
When the reply is missing either value the routing rule needs, src/extract.py::compute() returns None and the app prints 'not computed — one of the two values the rule needs was missing' rather than a 'no'. An unknown is never a pass here: this flag exists to stop a destruction, so silence must not read as clearance. And when a call does not come back at all, the review is recorded as a failure with its error text and is never scored as a correct answer — which is exactly what happened once on the deliberating tier, and what the grader got wrong before it was fixed. See not_good_enough.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Screening a disposition batch before a records officer opens it, to find series that must not be in it — either tier — they tie on every published grader over the reviews both answered 100.00 pct eligibility accuracy on both, with 1.00 recall on the frozen class, against the over-cautious floor's 41.82 pct and the tone floor's 60.00 pct. The floors are not strawmen: freezing everything with a hold nearby and trusting the officer's note are the two shortcuts a real desk actually takes.
Deciding which tier to buy for a monthly run — the fast tier It costs 6.9 pct less per review and answers in 3873 ms at the median against 7369 ms, and no published figure separates the two. The deliberating tier also lost the only review either tier lost.
Pointing this at a real hold registry — the kit's shape — one call per review, the registry sent whole, the rule in pure code downstream The parts that must not be probabilistic are not: the three-condition derivation, the review flag and the scoring are all code, and the model is asked only for the prose judgement code cannot do.
At a glanceHow the whole thing runs
100%extraction accuracy
3,873 msp50, end to end
$2.29per 1,000 disposition reviews · Google Gemini 3 Flash
Run once, for real, on 2026-08-22. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Check an agency's expiring records for holds before destruction14 steps · 4 questions · run once, for real · 2026-08-22
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace data/corpus/*.txt with your own disposition reviews, write your own data/fields.json, and supply a gold row per review. Every number on this page stops being true the moment you do.Corpus lens →
When is this the wrong choice?
Avoid: The over-cautious floor for this specifically. It looks safe and is not: it releases 10 genuinely frozen series, because a series can be frozen with a perfectly clean registry — by a longer-retention overlapping series, or simply by its own retention not having elapsed. That is the case against the best-fitting scenario (“Screening a disposition batch before a records officer opens it, to find series that must not be in it”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A review whose sections are not underlined headings. src/segment.py falls back to ONE whole-document segment rather than inventing a carve-up — the extraction still runs and every span collapses to 'document', so the source column stops being evidence. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether either tier's perfect result survives a hold notice written by a lawyer rather than generated from four structured pieces. Every scope in this corpus is one templated sentence; that is the shape of the judgement and not its full difficulty. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 2 models on the fast tier and the deliberating tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-22 — r001-hold-conflict. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Checked on a fresh checkout with API_KEY left blank: python -m src.app starts, the 55-review picker populates, any review renders in full, and Extract returns a plain sentence saying nothing was called. python -m evals.check_labels passes, and both free floors (--baseline holds, --baseline notes) score a complete run with no key at all. What a cold clone cannot reproduce is either paid run — those need a key, and the committed result files are what stands in for them.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
3,873 msp50, end to end
10,887 msp95
2 minclone to first result
What the clock covers. model call only, one per records-disposition review
Current processWhat it replaces
Someone opening each disposition review, reading every line of the hold registry, judging each hold's prose scope against that series' own category, project and closed date, following any 'continues the scope of' reference back to the hold it continues, and only then checking two expiry dates against the review date before deciding whether the series may go to a records officer for destruction approval.
Where it is not good enough
Four things, and only one of them is a score.
1. NEITHER TIER PRODUCED A SINGLE WRONG ANSWER ON THIS CORPUS. 660 of 660 cells, 55 of 55 binding-hold identifications and 55 of 55 eligibility verdicts on the fast tier; the same on all 54 reviews the deliberating tier answered. That is a statement about the corpus and not a licence. A 55-review set that cannot separate two tiers has stopped measuring model quality; what it still does is convict both free floors, at 41.82 pct and 60.00 pct, which is the finding that survives.
2. THE ONE REAL FAILURE IN 110 PAID CALLS WAS NOT A MODEL FAILURE. Review RDS-0018 died on the deliberating tier with <urlopen error [Errno 60] Operation timed out> — a client-side socket timeout at the adapter's 120-second default. That adapter retries only the HTTP statuses in its TRANSIENT set, and a socket timeout raises urllib's URLError, never HTTPError, so it is never retried and never will be. One review of 55 was lost and r002 is published as 54 of 55. It is a one-line fix and it is deliberately NOT applied here, because the published numbers have to come from the code that produced them.
3. A PERFECT SCORE HID A GRADER DEFECT until somebody disbelieved it. The binding-hold grader compared the model's named hold against gold's — and gold names NO hold on RDS-0018, so the missing reply's null matched, and the grader published 55 of 55 over 54 answers. No gate caught it; reading the two numbers side by side did. It is fixed in evals/judge.py (an unanswered review is now its own verdict and never a correct call), and both paid runs were re-scored from their own recorded cells with python -m evals.run --rescore rather than re-purchased.
4. THE SCOPE PROSE IS A TEMPLATE. Every hold scope here is one sentence generated from a category list, a project and a date span. A real hold notice is drafted by counsel, runs to paragraphs, and describes its reach in language that was negotiated rather than templated. The SHAPE of the judgement this kit measures is right; its full difficulty is not, and no number on this page should be read as though it were.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The kit PROPOSES and a records officer RELEASES — nothing here destroys anything, and the guardrail runs the other way: it routes a series that is frozen and ALREADY in the destruction queue, 16 of 16 fired with no false alarms on both tiers. The hard part is prose: a hold's scope is judged on three tests that must all pass, and 17 of the 55 reviews carry a hold that passes two of them and covers nothing. Neither tier got one wrong. The failure worth naming is the free over-cautious floor (evals/baseline.py, any hold on file means freeze it): 22 series frozen that nothing holds, and 10 genuinely frozen series RELEASED, because 12 of the 28 frozen series have a completely clean registry. No red-team run exists for this kit; this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
SECTION_HINTS
src/select.py
map fields to your own review's headings; anything unmatched falls back to the whole document
eligibility() and AS_OF
src/extract.py
the three-condition rule and the review date. Change either and gold, the prompt and the scorer all move together, because all three read the same function and the same constant — which is the point, and also why every published verdict figure is invalidated by the edit
fields.json
data/fields.json
which twelve fields are asked for, their types and their allowed values; the hint on binding_hold_id is where the coverage test is stated to the model in words
PROVIDERS
src/adapters/__init__.py
add a provider in one function and one dict entry; it must return the same shape including token counts, or its run cannot be published
the floors
evals/baseline.py
holds and notes are two shortcuts a real records office actually takes; add your own and the same judge scores it against the same gold
Components
Component
File
Role
segment
src/segment.py
cut the disposition review into addressable sections, pure code
select
src/select.py
map each field to the sections that could state it; the custodian office is mapped by nothing and is never sent
prompt
src/prompt.py
assemble one prompt per review — system rules, field schema, mapped sections
adapters
src/adapters/__init__.py
one chat completion over raw HTTP; provider chosen by .env, no vendor SDK
extract
src/extract.py
parse the reply, locate every value back to its own section, then run the pure-code eligibility derivation and the business-condition flag
budget
src/budget.py
count and cap live calls across every kit sharing one key, before the call
judge
evals/judge.py
four graders, all pure code, plus a per-class table and a no-gold consistency diagnostic
baseline
evals/baseline.py
two free floors that fail in opposite directions — the over-cautious clerk and the tone reader
check_labels
evals/check_labels.py
refuse to let a run spend until gold agrees with its own derivation and every hard case is present in quantity
Where it breaks at scale
One call per review, no concurrency and nothing shared between reviews: 55 reviews took 248.3 s on the fast tier and 422.5 s over 54 answered on the deliberating tier, which is one review per round trip and nothing else. A monthly disposition run at ten thousand series is 10,000 sequential calls and about 11 hours at the fast tier's 3873 ms median — the fix is concurrency, not a different design, because the calls are independent.
AND THE THING THAT ACTUALLY BROKE, AT 55. The adapter's retry set is HTTP statuses (TRANSIENT = 408/429/5xx). A CLIENT-SIDE socket timeout raises urllib's URLError, which carries no status and is never retried — so one review in 55 was simply lost on the deliberating tier. At 55 that is a footnote you can see; at 10,000 it is roughly 180 silently missing series in a queue whose whole purpose is that nothing goes missing. Any real deployment needs the timeout in the retry set and a completeness check that reconciles reviews-in against verdicts-out before the batch runs.
The registry itself does not scale in the model's favour either: every hold line is sent in full, so an agency with forty live holds on one series sends forty scope sentences in one prompt and asks one call to check all of them. Nothing here fans out.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Before anything is asked of the model. Twelve named fields with their own types and allowed values, plus a second panel for the routing decision taken afterwards in pure code — which hold binds, whether the series may be proposed for disposition, whether it is already in the destruction queue, and whether somebody has to pull it out today. The banner above the controls says what the page will not do: there is no destroy control and no write path.successOpen full size →RDS-0004, the sharpest review in the corpus: a hold whose scope covers this series exactly and which reads released, plus a second, ACTIVE registry line whose entire scope is 'continues the scope of AH-2023-234'. The officer's note reads 'Routine end-of-retention item. Nothing outstanding on this series.' — the wrong register. The model followed the reference, named PR-2025-041 as the binding hold, answered disposition_eligible=no, and because the series is already queued the pure-code flag fired: pull it out.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
With no API_KEY configured, Extract returns a plain sentence saying nothing was called rather than an error — the page still renders and every field reads 'not extracted yet' instead of a blank cell. The routing row reads 'not computed — one of the two values the rule needs was missing', which is the behaviour that matters on a kit like this: an unknown must never render as a clearance.failureOpen full size →
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
55disposition reviews
0.04 MiBtxt 55
660sections · p50 48 chars
$0.00setup · 0.0022s
How it is cutWhat one section is
cut on underlined section headings; a review with none falls back to one whole-document segment
SetupWhat the setup figure measured
There is no index. Preparation is segmentation only — 55 reviews cut into 660 sections in about two milliseconds, with no model and no network. Nothing is retrieved; the segment exists so an extracted value can name the section it was read from.
LicenceLicence
MIT — this repository's own licence. Every series id, custodian office, project name, schedule code and hold id is invented, no person is named anywhere, and no real agency, published general records schedule, disposition authority or litigation matter is named or reproduced.
Bring your ownBring your own disposition reviews
Replace data/corpus/*.txt with your own disposition reviews, write your own data/fields.json, and supply a gold row per review. SECTION_HINTS in src/select.py maps fields to your headings; anything unmatched falls back to the whole document, so a wrong map costs tokens rather than correctness. Then rewrite eligibility() and AS_OF in src/extract.py to your own disposition rule and your own review date — those two are the kit, and everything else is plumbing around them. Run python -m evals.check_labels before you spend anything: it refuses a run whose gold disagrees with its own values.
⚠︎ And what stops being true when you do: Every number on this page stops being true the moment you do. The eligibility rule here is three conditions this kit invented; yours will have a schedule authority, event-triggered retentions and a hold procedure negotiated with counsel behind it. And the hold scopes here are one templated sentence each — a model that reads them perfectly has not been shown to read a hold notice drafted by a lawyer.
What breaks it
A review whose sections are not underlined headings. src/segment.py falls back to ONE whole-document segment rather than inventing a carve-up — the extraction still runs and every span collapses to 'document', so the source column stops being evidence.
A hold registry that is not one line per hold. The scope prose is judged by the model, not parsed, so free-form registry text still works — but the free holds floor, evals/baseline.py's regex and check_labels' successor assertions all key on the four-column line format and go quiet rather than loud.
A scope that references a hold NOT in the registry. effective_scope() resolves exactly one level and returns None when the referenced id is absent, so the continuing line binds nothing. That is deliberate — a chain nobody wrote is not a rule — but a real registry with a two-step supersession would be silently under-frozen.
A retention period expressed in anything but whole years from the cutoff. Every date here is 'YYYY-MM' and check_labels asserts the month carries through from closed to expiry; a schedule with a fiscal-year cutoff, an event-triggered retention ('5 years after final payment') or a permanent item has no expiry month to compare and this rule cannot read it.
More than one hold binding the same series. binding_hold_id is a single id and the rule takes the FIRST active covering line in registry order. Two live holds on one series is ordinary in a real programme; here the second one is invisible, which matters the moment the first is released.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
3,258
839
field schema
2,455
599
review sections
781
267
Total
1,705
This is the cost lesson as arithmetic: of the 1,705 tokens assembled, 839 are instructions — 49% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Measured token-for-token via evals/prompt_tokens.py's nested-prefix method on RDS-0004 — three calls at max_tokens=1, each sending the previous prefix plus exactly one more part, with each part's size read off the difference between two consecutive prompt_token counts the provider itself returned (results/tokens-p001-hold-conflict.json). The three parts sum to 1705, which is the same input_tokens the captured RDS-0004 call was billed for — so the published decomposition is of the prompt that was actually sent, not a reconstruction that happens to look tidy.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You extract structured fields from a records-disposition review for one record series, and you decide whether that series may be PROPOSED for destruction. You return JSON and nothing else.
You are not destroying, deleting or disposing of anything. A records officer reviews and releases; your answer is the proposal they review.
THE REVIEW DATE IS 2026-08. Every 'has this elapsed yet' question below is asked against that date and against no other.
RULES, in order of importance:
1. If the record does not state a field, return null for it. Do not infer it and do not use what you know about the world.
2. `binding_hold_id` is the id of the ONE hold in the Hold Registry that freezes this series, or null when none does. A hold freezes this series only when BOTH of these hold:
a. its status is `active`; and
b. its scope covers this series, which is THREE separate tests and ALL THREE must pass:
- CATEGORY: the record's Record Category is one of the categories the scope names.
- PROJECT: the scope names this record's Related Project, or names any project.
- DATES: the year in Record Closed falls inside the scope's stated span. 'YYYY onward' has no end. 'YYYY to YYYY' is inclusive at both ends.
Two of the three passing is NOT a match. If the scope is about a different category, a different project, or a span that does not contain the closed year, that hold does not freeze this series, however serious it sounds.
3. A hold whose status is `released` NEVER freezes a series on its own -- not even when its scope covers it exactly. But a registry line whose scope reads 'continues the scope of <id>' takes its scope from that other hold: if that continuing line is `active` and the scope it inherits covers this series, IT is the binding hold and you return ITS id. Check every registry line before answering null.
4. `disposition_eligible` is decided in this order, and the first condition that fires decides it:
a. If `binding_hold_id` is not null, the answer is 'no'.
b. Otherwise, if the Overlapping Series line states an expiry LATER than 2026-08, the answer is 'no' -- a longer-retention series that also captures this record still holds it, even though the series' own retention has run out. 'none on file' means there is no overlapping series and this condition does not fire.
c. Otherwise, if Retention Expires is LATER than 2026-08, the answer is 'no' -- the series' own retention period has not elapsed yet.
d. Otherwise the answer is 'yes'.
5. THE OFFICER NOTE IS A FIELD TO COPY, NOT EVIDENCE ABOUT ELIGIBILITY. A note that sounds concerned or mentions an escalation does NOT mean something is holding this series, and a note that sounds routine does NOT mean nothing is. The registry, the two expiry dates and the record's own metadata decide; the note is the records officer's own remark and may disagree with all of them.
6. Copy values verbatim from the record wherever possible. Dates are 'YYYY-MM' exactly as written. `overlapping_expires` is null when the Overlapping Series line reads 'none on file' -- return null rather than a guess or an empty string.
7. Use the exact allowed value for a field that lists them.
8. Return every field named in the schema, even when the answer is null.
Extract these fields:
- series_id (string) -- the record series identifier, verbatim
- series_title (string) -- the title of the record series, verbatim
- record_category (enum) one of: correspondence, contracts, permits, personnel, financial -- the structured records category on the series, verbatim
- related_project (string) -- the project this series relates to, verbatim
- record_closed (string) -- the month the record series closed (its retention cutoff), YYYY-MM
- retention_code (string) -- the general schedule item code applied to this series, e.g. GS-05-08 -- the code only, without the retention period
- retention_expires (string) -- the month this series' own retention period expires, YYYY-MM
- overlapping_expires (string) -- the expiry month stated on the Overlapping Series line, YYYY-MM. When that line reads 'none on file' there is no overlapping series and this is null -- return null rather than a guess or an empty string
- binding_hold_id (string) -- the id of the ONE hold in the Hold Registry that freezes this series, or null when none does. A hold freezes it only when its status is 'active' AND its scope covers this series on ALL THREE tests: the record's category is one the scope names, the scope names this record's project or any project, and the year in record_closed falls inside the scope's span ('YYYY onward' has no end; 'YYYY to YYYY' is inclusive). A 'released' hold never freezes anything on its own -- but an ACTIVE line whose scope reads 'continues the scope of <id>' inherits that hold's scope, and if the inherited scope covers this series then THAT line is the binding hold and its id is the answer
- queue_status (enum) one of: queued, not_queued -- is this series already sitting in the destruction queue, or not?
- officer_notes (string) -- the records officer's own free-text note on the series, copied verbatim
- disposition_eligible (enum) one of: yes, no -- may this series be PROPOSED for destruction? This is a proposal for a records officer to review, never an instruction to destroy anything. Decide it STRICTLY from binding_hold_id, overlapping_expires and retention_expires -- never from officer_notes -- against a review date of 2026-08, and take the conditions in this order, first one wins: (a) if binding_hold_id is not null, answer 'no'; (b) otherwise, if overlapping_expires is later than 2026-08, answer 'no'; (c) otherwise, if retention_expires is later than 2026-08, answer 'no'; (d) otherwise answer 'yes'
Return a JSON object with exactly these keys: series_id, series_title, record_category, related_project, record_closed, retention_code, retention_expires, overlapping_expires, binding_hold_id, queue_status, officer_notes, disposition_eligible
Use null for any field the record does not state.
RECORDS DISPOSITION REVIEW
--------------------------
Record Series
-------------
RS-2018-8284
Series Title
------------
Financial transaction files - Bell Harbor
Record Category
---------------
financial
Related Project
---------------
Bell Harbor
Record Closed
-------------
2018-12
Retention Schedule
------------------
GS-30-01, 7 years
Retention Expires
-----------------
2025-12
Overlapping Series
------------------
none on file
Hold Registry
-------------
PR-2025-041 | active | Scope: continues the scope of AH-2023-234. | Continues: AH-2023-234
AH-2023-234 | released | Scope: all financial records relating to the Bell Harbor project, 2015 onward. | Superseded by: PR-2025-041
Disposition Queue
-----------------
queued
Officer Notes
-------------
Routine end-of-retention item. Nothing outstanding on this series.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Check an agency's expiring records for holds before destruction — 55 disposition reviews. Two tiers of one model family answered, and every answer was then graded Four different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
The rulerHow it was graded
Every grader here is pure code. Gold is not a second opinion about the officer's note — it is the kit's own three-condition derivation, re-run inside the grader over the values each review itself states, so the truth cannot drift from the rule the prompt states and the app applies.
55disposition reviews
55source documents
2model tiers
110graded answers
4grading methods
MeasurementsWhat was measured
COUNTED660 · 648 / 660extraction accuracy — cellsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED55 · 54 / 55binding hold identified — reviewsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED55 · 54 / 55eligibility accuracy — reviewsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED28 · 28 / 28frozen-series recall — frozen seriesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer is pure code, exact match with light normalisation against a mechanically-derived gold, so there is no judge to validate — but that is not the same as nothing to check, and this kit paid for the difference. evals/check_labels.py validates the LABELS before any run may spend: it re-runs the eligibility derivation over every gold row, asserts every class produced the verdict it is named for, asserts the successor case is findable only by following the reference and the mirror case carries no successor at all, asserts each hard case appears at least five times, and asserts the review date is the same constant in three files. It does NOT validate the graders themselves — and the binding-hold grader was wrong, counting an unanswered review as a correct 'nothing binds' until a perfect score was disbelieved. See Business.not_good_enough.
485.69output tokens · the fast tier · 3,873 ms p50
542.48output tokens · the deliberating tier · 7,369 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 1.9× as long, and lands one row apart on 55. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate — a projection, not a bill. Nobody paid any of these numbers.
Priced at
Per 1M in / out
One disposition review
1,000 disposition reviews
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows are comparable across use cases.
$0.50 / $3.00
$0.002293
$2.29
36%
Same work, 1× the bill
The same disposition reviews, the same tokens — only the rate card changed. And on that card about 36% of what you pay is the prompt this pipeline sends, not the answer it writes.
which tier is called — and the two tiers tie on every published grader, so the deliberating tier's 7.4 pct premium buys measurably nothing here. The second lever is the field schema: it is 599 tokens on every call, and the binding_hold_id hint alone is most of it, because the coverage test is stated to the model in words.
Rates checked 2026-08-18. The provider that actually ran r001 and r002 publishes no rate card this repo commits, so nothing here is what the run cost. What is published is the token count, which is the part that transfers.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Free means free: evals/judge.py makes no model call, and neither do the two baseline floors. The only thing a re-run costs is the extraction itself, which is priced in the Cost lens. Re-scoring an EXISTING run costs nothing at all — python -m evals.run --rescore <file> rebuilds the reply set from the result file's own recorded cells.
The gradersFour ways to grade
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Per-field exact match, light normalisation Does the model's value for each of the twelve fields match gold, after trimming whitespace and trailing punctuation and lowercasing? A cell is a hit, a miss (nothing returned where gold has a value) or wrong (something else returned).
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 100.0% · the free over-cautious floor 91.1%
Which hold freezes this series Did the run name the same hold gold names — INCLUDING naming none when none binds? This is the prose scope judgement everything else rests on: three tests (category, project, date span) that must ALL pass, plus a reference to follow when an active line's whole scope reads 'continues the scope of <id>'.
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 98.2% · the free over-cautious floor 50.9% · the free tone floor 70.9%
disposition_eligible against gold's own derivation Of every series that really is frozen, how many did the run refuse to release — and how many genuinely eligible series did it freeze? FROZEN IS THE POSITIVE CLASS: a series under a live hold that gets called eligible is the one that reaches a destruction batch, and it is the failure a records programme actually pays for.
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 98.2% · the free over-cautious floor 41.8% · the free tone floor 60.0%
needs_review — frozen AND already in the destruction queue Is this the series somebody has to pull out of the destruction queue today? It fires when disposition_eligible is no AND queue_status is queued. Note the direction: it never proposes a destruction, it asks for one to be stopped.
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 98.2% · the free over-cautious floor 67.3% · the free tone floor 78.2%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Between the models and both free floors, decisively; between the two tiers, not at all. The over-cautious floor scores 41.82 pct on eligibility and the tone floor 60.00 pct where both tiers score 100.00 pct on every review they answered, so the set convicts the two shortcuts a real records desk actually takes. It cannot rank the tiers: every published grader is identical, and the only thing that separated them in this run was a network timeout. A corpus that cannot separate two models has stopped measuring model quality, and saying so is the honest reading of a perfect score.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Screening a disposition batch before a records officer opens it, to find series that must not be in it
either tier — they tie on every published grader over the reviews both answered
100.00 pct eligibility accuracy on both, with 1.00 recall on the frozen class, against the over-cautious floor's 41.82 pct and the tone floor's 60.00 pct. The floors are not strawmen: freezing everything with a hold nearby and trusting the officer's note are the two shortcuts a real desk actually takes.
the over-cautious floor for this specifically. It looks safe and is not: it releases 10 genuinely frozen series, because a series can be frozen with a perfectly clean registry — by a longer-retention overlapping series, or simply by its own retention not having elapsed.
Deciding which tier to buy for a monthly run
the fast tier
It costs 6.9 pct less per review and answers in 3873 ms at the median against 7369 ms, and no published figure separates the two. The deliberating tier also lost the only review either tier lost.
reading the tie as evidence the tiers are equivalent. It is evidence this 55-review corpus cannot separate them, which is a fact about the corpus.
Pointing this at a real hold registry
the kit's shape — one call per review, the registry sent whole, the rule in pure code downstream
The parts that must not be probabilistic are not: the three-condition derivation, the review flag and the scoring are all code, and the model is asked only for the prose judgement code cannot do.
treating the measured accuracy as transferable. Every scope here is one templated sentence; a hold notice drafted by counsel is a different reading task, and nothing on this page has measured it.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
no-model-side-failure
Neither tier produced a wrong answer of any kind on this corpus
0
Across 109 answered replies and 1308 scored cells there was no field miss, no wrong binding hold, no wrong eligibility verdict and no false review-flag alarm. The hardest cases went the same way: all 6 released-with-a-live-successor reviews, all 17 reviews…
call-lost-to-a-client-side-timeout
A review that never came back, and was never going to be retried
1
RDS-0018 on the deliberating tier: <urlopen error [Errno 60] Operation timed out>. src/adapters/__init__.py retries the HTTP statuses in TRANSIENT (408/429/5xx); a client-side socket timeout raises urllib's URLError, which carries no status, so it falls…
grader-counted-a-non-answer-as-an-answer
The binding-hold grader scored an unanswered review as a correct 'nothing binds'
1
evals/judge.py's binding-hold grader compared norm(got) against norm(want). On RDS-0018 gold's binding_hold_id is null and the missing reply's value was null, so the comparison succeeded and the grader published 55 of 55 over 54 answers. The confusion…
What we could NOT verify
Whether either tier's perfect result survives a hold notice written by a lawyer rather than generated from four structured pieces. Every scope in this corpus is one templated sentence; that is the shape of the judgement and not its full difficulty.
Whether the result holds on a second seed. One corpus, one seed (20260822), two tiers. Regenerating with a different seed and re-running is two more paid runs and was not done.
Whether a review with TWO binding holds is handled sensibly. binding_hold_id is a single id and the rule takes the first active covering line in registry order; the corpus never plants a second, so nothing here measures it.
Whether the models would resist an instruction hidden in officer_notes. The corpus plants confusable REGISTER — a routine-sounding note on a frozen series — which is a different test from adversarial text, and no attack run was fired.
What the deliberating tier would have answered on RDS-0018. The call timed out, the review was never retried, and it is published as unanswered rather than filled in from the other tier's reply.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
1,671.18
485.69
3,873 ms
$0.002293
the deliberating tier
1,671.24
542.48
7,369 ms
$0.002463
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
Both free floors (evals/baseline.py) and all four graders (evals/judge.py) are pure code and cost $0.00 — no key, no model, no network. Re-scoring an existing run costs nothing either: python -m evals.run --rescore <file> rebuilds the reply set from the result file's own recorded cells, which is how the binding-hold grader was fixed after two paid runs without buying a third. The only cost of evaluation is the extraction itself: $0.1261 for the fast tier's 55 reviews and $0.1330 for the deliberating tier's 54, on the projected card.
Cost driversWhat actually moves the bill
The fixed instruction block, which is most of every call. The system prompt (839 tokens) and the field schema (599 tokens) are byte-identical on every review and together are 84 pct of the input; the review's own sections are 267 tokens. The bill therefore tracks HOW MANY reviews you have, not how long any of them is.
The reasoning pass you do not see. The captured example billed 576 output tokens for a JSON body of about 290, with reasoning_tokens: 425 — roughly three quarters of the output price, and output is 6x the input rate on the published card. This is the single biggest lever on the bill and it is not visible in the reply.
One call per review, always. There is no batching and no shared context between reviews, so cost is exactly linear in the size of the disposition queue.
The size of the hold registry. Every hold line is sent whole; a series with forty live holds sends forty scope sentences. On this corpus that is at most two lines and invisible — on a real registry it would not be.
Your volumeWhat it costs at your volume
Linear in reviews: each call is independent, carries the same fixed prompt and shares nothing with its neighbours, so ten times the queue is ten times the bill and ten times the wall clock. Nothing amortises — there is no index to build and nothing is cached between calls. What does NOT scale linearly is the reliability: the one review lost at 55 to an unretried socket timeout becomes roughly 180 lost at 10,000, in a queue whose whole purpose is that nothing goes missing.
Where pricing changes shape
Context repricing. Every published card here prices one band and this kit's calls are about 1705 tokens, nowhere near any threshold — but a real registry with dozens of live holds pushes input up without warning, and several vendors reprice the WHOLE request past a context threshold rather than the excess.
The reasoning ceiling. MAX_TOKENS is 6,000 and the observed maximum across 115 live calls was 1849, so nothing has come close — but output is billed at 6x input on the published card, and a model family that reasons harder on a harder registry moves the bill through the output side with no change to the prompt at all.
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier is what a monthly disposition batch would actually be run on, and this kit's fitment recommends it: it ties the deliberating tier on every published grader, costs 6.9 pct less per review and answers roughly twice as fast at the median.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
91,915input tokens · this run
26,713output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 55 records-disposition reviews, one completion call each, on the fast tier. The deliberating tier's own run is recorded separately in Cost.cost_by_model and is not projected here.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.050
$0.050
$0.92
2026-09-12
gemini-3-flash
Google
$0.126
$0.126
$2.29
2026-09-18
gemini-3-8-flash
Google
$0.169
$0.169
$3.07
2026-09-18
claude-haiku-4-5
Anthropic
$0.225
$0.225
$4.10
2026-09-12
llama-5
Meta
$0.228
$0.228
$4.15
2026-09-18
grok-4-5
xAI
$0.344
$0.344
$6.26
2026-09-18
grok-4-6
xAI
$0.344
$0.344
$6.26
2026-09-18
claude-sonnet-5
Anthropic
$0.451
$0.451
$8.20
2026-09-12
gemini-3-1-pro
Google
$0.504
$0.504
$9.17
2026-09-18
gpt-5-6-terra
OpenAI
$0.504
$0.504
$9.17
2026-09-12
gpt-5-6-sol
OpenAI
$0.902
$0.902
$16.40
2026-09-12
claude-opus-4-8
Anthropic
$1.127
$1.127
$20.50
2026-09-12
claude-opus-5
Anthropic
$1.127
$1.127
$20.50
2026-09-12
claude-fable-5
Anthropic
$2.255
$2.255
$41.00
2026-09-18
claude-fable-5-1
Anthropic
$2.255
$2.255
$41.00
2026-09-18
gpt-6-astra
OpenAI
$2.255
$2.255
$41.00
2026-09-17
Read this against the numbers above
Every row below prices the FAST TIER's own 55-call run (r001-hold-conflict) -- the deliberating tier's own token counts are on Cost.cost_by_model[1] and are not separately projected here.
Neither tier's run left anything to disable -- src/adapters/__init__.py's thinking parameter is only sent when a caller passes one, and this kit's own harness never does (see LLM.settings) -- so there is no reasoning-on/reasoning-off discrepancy to caveat here. What there IS: the captured example billed 425 reasoning tokens of its 576 output tokens, so roughly three quarters of the output price on these rows buys thinking nobody reads.
The deliberating tier lost one review of 55 to a client-side socket timeout, so its per-review figures are measured over 54 answered reviews and the fast tier's over 55. The two are comparable per review and NOT comparable as run totals.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Nine modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/segment.pysegment
cut the disposition review into addressable sections, pure code
src/segment.py
# Cut a document into addressable sections. Pure code — no model, no network.
def sections(text):
def locate(text, value):
def span_label(secs, start):
src/select.pyselect — a swap seam
map each field to the sections that could state it; the custodian office is mapped by nothing and is never sent
You change it to: map fields to your own review's headings; anything unmatched falls back to the whole document
src/select.py
# Pick which sections plausibly carry each field. Pure code -- the last deterministic step
SECTION_HINTS = {
def for_field(secs, field):
def plan(secs, fields):
src/prompt.pyprompt
assemble one prompt per review — system rules, field schema, mapped sections
src/prompt.py
# Assemble the extraction prompt. One prompt per records-disposition review, all twelve fields
AS_OF = "2026-08"
SYSTEM = (
def field_schema(fields):
def build(doc_text, secs, fields, selector):
def parse(raw, fields):
src/adapters/__init__.pyadapters — a swap seam
one chat completion over raw HTTP; provider chosen by .env, no vendor SDK
You change it to: add a provider in one function and one dict entry; it must return the same shape including token counts, or its run cannot be published
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/extract.pyextract — a swap seam
parse the reply, locate every value back to its own section, then run the pure-code eligibility derivation and the business-condition flag
You change it to: the three-condition rule and the review date. Change either and gold, the prompt and the scorer all move together, because all three read the same function and the same constant — which is the point, and also why every published verdict figure is invalidated by the edit
src/extract.py
# Extract one records-disposition review's fields: segment, select, prompt, one model call, then
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
FIELDS = os.path.join(HERE, "data", "fields.json")
CORPUS = os.path.join(HERE, "data", "corpus")
MAX_TOKENS = 6000
AS_OF = "2026-08"
def load_fields():
def load_doc(case_id):
def documents():
def _ym(value):
src/budget.pybudget
count and cap live calls across every kit sharing one key, before the call
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
evals/judge.pyjudge
four graders, all pure code, plus a per-class table and a no-gold consistency diagnostic
evals/judge.py
# Score a disposition run. PURE CODE -- gold is exact and the answer is one value per cell, so
def norm(v):
def equal(field, got, want):
def score(fields, records, golds):
def _matrix(rows, positive):
def score_flags(records, flags, golds):
evals/baseline.pybaseline — a swap seam
two free floors that fail in opposite directions — the over-cautious clerk and the tone reader
You change it to:holds and notes are two shortcuts a real records office actually takes; add your own and the same judge scores it against the same gold
evals/baseline.py
# Two free, rules-and-regex extractors. No model, no key, no spend -- scored by the same judge as
WORRIED_KEYWORDS = ("escalat", "second look", "not confident", "manual audit", "disputed",
def _section(text, name):
def _first_line(s):
def _code(s):
def _overlap_expires(s):
def _hold_ids(s):
def extract_one(text, fields, floor):
def extract(text, fields, floor="holds"):
evals/check_labels.pycheck_labels
refuse to let a run spend until gold agrees with its own derivation and every hard case is present in quantity
evals/check_labels.py
# Check the gold set before anything is scored against it. Run this before spending money.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
NULLABLE = {"overlapping_expires", "binding_hold_id"}
FROZEN_CLASSES = {"hold_active", "hold_successor", "overlap_longer", "retention_open"}
ELIGIBLE_CLASSES = {"scope_category_miss", "scope_project_miss", "scope_date_miss",
def bad(msg):
def main():
Start hereThe shortest path into it
src/segment.pycut the disposition review into addressable sections, pure code
src/select.pymap each field to the sections that could state it; the custodian office is mapped by nothing and is never sent A swap seam.
src/prompt.pyassemble one prompt per review — system rules, field schema, mapped sections
src/adapters/__init__.pyone chat completion over raw HTTP; provider chosen by .env, no vendor SDK A swap seam.
src/extract.pyparse the reply, locate every value back to its own section, then run the pure-code eligibility derivation and the business-condition flag A swap seam.
src/budget.pycount and cap live calls across every kit sharing one key, before the call
evals/judge.pyfour graders, all pure code, plus a per-class table and a no-gold consistency diagnostic
evals/baseline.pytwo free floors that fail in opposite directions — the over-cautious clerk and the tone reader A swap seam.
evals/check_labels.pyrefuse to let a run spend until gold agrees with its own derivation and every hard case is present in quantity
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1671 input and 485 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's disposition reviews are entirely synthetic (tools/build_corpus.py, seed 20260822): no real agency, custodian, matter or person exists in the corpus, no person is named anywhere in it, and nothing was fetched. The only outbound traffic the kit makes is one chat-completion request per review to the configured provider, carrying the mapped sections of one review — the custodian office is mapped by no field and never leaves the machine. Nothing is written outside the kit directory, there is no database, no auth and no multi-tenancy, and the local UI binds 127.0.0.1 only.
Read from the shared .env or the real environment only, never committed, never logged and never rendered. .env is gitignored from the first commit. On an adapter error the app redacts the key and base URL out of the message before returning it, so a misconfigured endpoint cannot echo a credential into a browser. Result files record the model id and the provider adapter name only.
The experimentWe did not attack it — and three of five boundaries hold
The three boundaries that hold were confirmed by reading the code, not by a run: no code path destroys anything or writes anywhere; the routing rule reads enums and date strings, so the officer's note cannot reach it; and the key and base URL are redacted out of any error before it is returned. A fourth — does a review that never answers fail loudly — held on the harness and failed on the grader, which is a distinction this run only found because a perfect score looked wrong next to a records count. The fifth is open. An indirect prompt injection needs a field an outside party authored, and this kit has exactly one, which makes it the obvious place to attack and the reason the gap is named rather than glossed. Confirmed by reading the code, not by a run, on 2026-08-22 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Can this kit destroy a record, release a hold or amend a schedule?
A kit that decides destruction eligibility could plausibly act on its own decision — mark a series disposed, close a hold, or write back to a queue.
No code path does. src/extract.py::extract() and src/app.py's /api/extract both return a JSON body and nothing else; the kit performs no write and no outbound call other than the single completion request. There is no destroy control in the UI and no write endpoint to call. needs_review is a value in a response asking a person to look at a queue — and it asks for a series to be REMOVED from one, which is the only direction a decision like this should be automated in at all.
Can the records officer's note change what the code decides?
The note is free text and sits in the same review as the registry, so text in it could steer the eligibility verdict.
It can steer the MODEL — that is the whole thing this kit measures, and both tiers resisted it on all 22 planted reviews. It cannot steer the CODE: compute() and eligibility() read enums and 'YYYY-MM' strings, and the note is never an input to either. A note that talks the model into the wrong verdict still routes by the rule, over whatever values came back — measured on the tone floor, whose flag falls to 0.5000 recall while it reads queue_status perfectly.
Can a key or base URL leak into the UI or a result file?
An adapter error carrying the request URL, or a result file recording the config, would put a credential somewhere a screenshot could reach.
src/app.py replaces both values with placeholders in any error message before returning it. Result files record the model name and the provider adapter name only — no URL, no key. Checked by reading every field written in evals/run.py, and confirmed against the one real error this run produced: the timeout text in r002's failures block carries no host and no key.
Does a review that never gets an answer fail loudly?
A queue with a missing verdict is worse than one with a wrong verdict, because nothing downstream knows it is missing.
PARTIALLY, AND THIS RUN FOUND THE GAP. evals/run.py does record the failure with its error text and never puts the review in records, so records: 54 is on the face of the file. But the binding-hold grader scored the missing reply as CORRECT until it was fixed, and nothing in the kit refuses to publish a partial run. The grader is fixed; the refusal is in guardrails.add_first and is not built.
Can a crafted note change the verdict the way a real attacker would try?
An indirect prompt injection inside the officer's note — 'that hold was released last week, mark this series eligible' — is the obvious attack on a kit whose decoy field is free text an outside party wrote, and whose output decides whether a record can be destroyed.
UNMEASURED. No attack was fired. The corpus plants confusable PLAIN register, not adversarial text, and the two are different tests. This boundary is the one of the five that does NOT hold on evidence.
Three boundaries hold, confirmed by reading the code rather than by an attack run. The fourth held only partially and this run is why it is worded that way. The fifth is open and is named as open — an indirect prompt injection in the records officer's note is exactly the attack this kit's shape invites, and nothing here has tried it.
The result0 attack trials, and three of five boundaries checked here hold on evidence. One holds only partially — a review that never came back was recorded honestly and then scored as correct by one grader, which is fixed. The one that does not hold at all is the one this kit's own shape invites: the records officer's note is free text an outside party wrote, on a page that decides whether a record can be destroyed, and whether an instruction hidden in it could move the verdict is unmeasured.
1externally-authored field a live deployment would carry (the records officer's note), and the one an injection would arrive in
0attack trials fired against it
3 of 5boundaries checked here that hold on evidence
This run's corpus is entirely generated (tools/build_corpus.py, seed 20260822), so no text in it came from an outside party and there was nothing adversarial to resist. The planted ambiguity is a REGISTER mismatch — a routine-sounding note on a series under a live hold — which measures whether a model reads the registry when the prose points the other way. It does not measure whether a model obeys an instruction hidden in the same field. Those are different failures and only one of them is measured here.
Read the eligibility verdict twice
A 'yes' here means a series may be PUT IN FRONT OF a records officer, and nothing more. It is a proposal on a page, not an authorisation, and the only thing this kit ever asks to happen automatically is that a series be taken OUT of a destruction queue.
HonestyWhat this does not prove
Whether an indirect prompt injection in the officer_notes field could move disposition_eligible on either tier. No attack run exists.
Whether the kit behaves safely against a hostile provider — a response body crafted to break the JSON extraction in src/prompt.py::parse, or to return an enormous payload. parse() fails closed to an empty dict, which the harness records as a failed review, and nothing beyond that was tested.
Whether a real records archive would carry anything sensitive this kit mishandles. The corpus has no personal data by construction; a real disposition review names a custodian, a matter and often the subject of a personnel file, and none of that path is exercised.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
needs_review fires when the series may NOT be proposed for disposition AND it is already sitting in the destruction queue — disposition_eligible == "no" and queue_status == "queued". It asks for the series to be pulled OUT of the queue; it never proposes putting one in.
src/extract.py::compute(), called by extract() on every review, by the local app on every click, and by evals/judge.py over both the models' output and both free floors' — the same function in all four places.
EvidenceDoes it hold?
What
Measured
The flag fires on exactly the reviews where both conditions hold, and on no others.
16 of 16 on both tiers, 39 left alone, 0 false alarms — 1.00 recall and 1.00 precision against the same rule run over gold's own values.
An unknown is never a pass. If either value is missing the rule returns None and the app prints 'not computed', never 'no'.
Exercised on every no-key render (see the extract-nokey shot) and on the one review that never came back: RDS-0018 is unanswered in the flag matrix, not a true negative.
The officer's note cannot reach the rule.
compute() reads two enums and nothing else; officer_notes is never an input. A note that talks the MODEL into the wrong verdict still routes by the rule, over whatever came back — which is what happens to the tone floor, whose flag falls to 0.5000 recall while it reads queue_status perfectly every time.
The limitWhat a guardrail is not
It is NOT a check on whether the extracted values are right. If the model names the wrong binding hold and then derives eligibility from it consistently, the series is routed (or not routed) on a wrong reading and this rule cannot tell. The consistency diagnostic in evals/judge.py is the closest thing to that check, it is reported separately, and it is blind to exactly the same case — which is why the binding-hold grader exists and why it needs gold.
It is NOT a real records programme's escalation policy. Frozen-and-queued is this kit's own simplification, invented for this corpus. No published general records schedule, disposition authority or hold procedure was consulted, and none is reproduced. A real records office weighs how close the destruction batch is, who issued the hold, whether counsel has been notified, and whether the series has already been certified for destruction.
It is NOT a disposition, and it is not the opposite of one either. Nothing in this kit destroys, deletes, releases a hold, amends a schedule or notifies anybody. The flag is a value in a JSON response asking a person to look at a queue.
It is NOT a completeness check. It says nothing about reviews that never produced a reply, and the run record is the only place that does. One review of 110 paid calls vanished and this rule was silent about it, correctly — silence is not its job, and it is why the run harness records failures with their error text.
WatchedWhat is watched, and why that one
4runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 20 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
7 measured by the latest run13 need the model half
Metric
Owner
Role
Why this one
field-exact-match-with-normalisation
Per-field exact match, light normalisation
alarm
binding_hold_id and overlapping_expires — the two nullable fields, where a returned null is often correct and a default-to-miss scorer would punish it; retention_code, where the schedule line states a code AND a period and only the code is asked for; officer_notes, which must come back verbatim including its double hyphen — alarm on Any drop in extraction_accuracy below 100 pct on either tier — both runs were exact on every cell they answered (660 and 648), so the first miss is a signal and not noise.
binding-hold-identification
Which hold freezes this series
alarm
missed_a_binding_hold — the expensive direction. A hold that binds and is not named releases a frozen series; the 6 hold_successor reviews, where the covering hold reads released and an active line continues it by reference; the 5 released_no_successor reviews, the mirror case that stops 'released means look for a successor' from being a working shortcut; the 17 reviews whose registry carries a hold that looks relevant and covers nothing — 6 wrong category, 5 wrong project, 6 outside the date span — alarm on Any missed_a_binding_hold or named_the_wrong_hold at all. Both tiers were at 0 across every review they answered.
eligibility-confusion-matrix
disposition_eligible against gold's own derivation
alarm
false_negative — a frozen series called eligible. This is the expensive direction and the one the over-cautious floor fails 10 times and the tone floor 13 times; the 6 overlap_longer reviews, where NOTHING in the registry is wrong and the series is frozen anyway; the 17 reviews carrying a hold that covers nothing — where the honest answer is 'release it' and the safe-looking answer is not — alarm on Any false negative at all. Both tiers were at 0 across every review they answered, so the first one is a signal and not noise.
review-flag-confusion-matrix
needs_review — frozen AND already in the destruction queue
alarm
false_negative — a frozen, queued series NOT flagged. That is a record destroyed under a live hold; the flag inherits whatever disposition_eligible says, so every verdict error upstream can become a flag error here — which is what happens to both floors — alarm on Any false negative. Both tiers fired on all 16 and raised no false alarms.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
55
different corpus — nothing is comparable
corpus.bytes
40,925
disposition reviews edited — the count held, the bytes did not
split.count
660
the sections count moved — a different set was scored
split.size_p50
48
the median size of one section moved
split.size_p95
140
the 95th-percentile size of one section moved
dataset.rows
55
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0022
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (refusal_cells 0) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Extraction accuracy
exact match at 100 pct on both tiers — 660 of 660 cells on the fast tier and 648 of 648 on the deliberating tier's answered reviews
660 cells on the fast tier, 648 on the deliberating tier (54 answered reviews x 12 fields)
two independent tiers (r001-hold-conflict, r002-hold-conflict), both exact -- see Eval.taxonomy for why a zero here is a fact about the corpus.
Binding hold identified
55 of 55 on the fast tier; 54 of 55 with 1 unanswered on the deliberating tier
55 disposition reviews per tier — every gold review, not every reply
evals/judge.py::score_flags. The denominator is deliberately the gold set: an unanswered review is its own verdict and never a correct 'nothing binds', which this grader got wrong until a perfect score was disbelieved.
Binding hold identified, free floors
50.91 pct for the over-cautious floor (24 invented holds), 70.91 pct for the tone floor (16 binding holds missed, because it never opens the registry)
55 disposition reviews
evals/baseline.py, no key and no model (b000-holds, b001-notes).
Eligibility verdict
100 pct on the fast tier and on every review the deliberating tier answered — 28 of 28 frozen series held, 27 of 27 eligible ones released, 1.00 recall and 1.00 precision on both
55 disposition reviews per tier (54 answered on the deliberating tier)
evals/judge.py::score_flags against gold's own derivation, r001 and r002. FROZEN is the positive class — releasing a series under a live hold is the failure that is not recoverable.
Eligibility verdict, free floors
41.82 pct for the over-cautious floor (22 frozen wrongly, and 10 frozen series released anyway) and 60.00 pct for the tone floor (13 frozen series released)
55 disposition reviews
evals/baseline.py, no key and no model. The two floors fail in OPPOSITE directions, which is the finding: over-freezing is the visible failure of a records programme and under-freezing is the expensive one.
Review flag
16 of 16 fired, 0 false alarms on both tiers — 1.00 recall and 1.00 precision
55 disposition reviews per tier
the same two-value rule run over gold's values; a business condition, so it needs labels and says so.
Review flag, free floors
67.27 pct accuracy for the over-cautious floor (0.5625 recall, 0.4500 precision) and 78.18 pct for the tone floor (0.5000 recall, 0.6667 precision)
55 disposition reviews
both floors read queue_status correctly every time by regex; the flag still fails, because it inherits a wrong eligibility verdict.
Span rate
100 pct — 421 of 421 returned values on the nine spannable fields located back to their own section of the review on the fast tier, 414 of 414 on the deliberating tier
421 spannable values on the fast tier, 414 on the deliberating tier
src/extract.py::_locate, which searches the sections src/select.py maps each field to BEFORE falling back to the whole document. The three enum fields are not spannable and are excluded rather than counted as misses.
Hallucinations
exact match at 0 on both tiers — no value was returned that the review does not state
660 cells on the fast tier, 648 on the deliberating tier
evals/judge.py counts a cell as wrong when a value is returned that gold does not carry; there were none on either tier, and every spannable value located back to its own section.
Latency
3873 ms p50 / 10887 ms p95 on the fast tier; 7369 ms p50 / 10645 ms p95 on the deliberating tier — about 90 pct slower at the median for no measured gain
one model call per review, 55 and 54 timed calls
measured end to end around the completion call in evals/run.py. The p95 gap is much narrower than the p50 gap because the fast tier has a long tail of its own — its slowest review took 15123 ms.
Tokens
1671.18 input and 485.69 output tokens per review on the fast tier, 1671.24 and 542.48 on the deliberating tier — input is identical because the prompt does not change
55 reviews on the fast tier, 54 on the deliberating tier
provider-reported usage on every call. Most of the output is not the reply: the captured example billed 576 output tokens for a ~290-token JSON body, with 425 of them reasoning tokens.
Reviews answered
55 of 55 on the fast tier; 54 of 55 on the deliberating tier — one call lost to a client-side socket timeout the adapter does not retry
55 disposition reviews per tier
evals/run.py records every failure with its error text and never scores a missing reply as an answer. THE GRADER DID, until it was fixed — see Eval.taxonomy's grader-counted-a-non-answer-as-an-answer.
HistoryRun history
4 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
b000-holds 2026-08-22
b001-notes 2026-08-22
r001-hold-conflict 2026-08-22
r002-hold-conflict 2026-08-22
extraction accuracy
0.9106
0.9424
1.0000
1.0000
invented values
0
0
0
0
values with a span
0.000
0.000
1.000
1.000
input tokens, whole run
0
0
91915
90247
model latency p50 ms
0.00
0.00
3873.00
7369.00
model latency p95 ms
0.00
0.00
10887.00
10645.00
output tokens, whole run
0
0
26713
29294
not a time series No two of these 4 runs measured the same system — they differ on documents, extraction_cells, failures, provider, thinking, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 4 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
which tier is called (fast vs deliberating)
output tokens (+12 pct) and median latency (+90 pct) move together; cost per review moves +7.4 pct. NOTHING ELSE MOVES — extraction, binding-hold identification, eligibility verdicts, the review flag and the consistency diagnostic are identical on every review both tiers answered.
measured
r001-hold-conflict vs r002-hold-conflict: 660/660 vs 648/648 cells, 55/55 vs 54/54 eligibility verdicts, 16/16 vs 16/16 review flags, 3873 vs 7369 ms p50.
freezing any series with a hold on file instead of reading the hold's scope
eligibility accuracy falls from 100 pct to 41.82 pct, and the review flag from 1.00/1.00 recall and precision to 0.5625/0.4500 — even though queue_status, the other field it reads, is extracted perfectly every time. And it is not merely over-cautious: 10 genuinely frozen series are RELEASED, because a clean registry does not mean nothing freezes the series.
measured
b000-holds against r001/r002 on the same 55 reviews, scored by the same judge.
reading eligibility off the records officer's note instead of the registry
eligibility accuracy falls to 60.00 pct and the review flag to 0.5000 recall. 13 frozen series released and 9 eligible ones frozen — exactly the reviews where the note is written in the register that contradicts the facts.
measured
b001-notes against r001/r002 on the same 55 reviews, scored by the same judge.
changing eligibility() or AS_OF in src/extract.py
gold, the prompt and the scorer, all three, in the same edit — because all three read the same function and the same constant. Every published verdict figure is invalidated and both paid runs would have to be fired again.
reasoning
evals/check_labels.py asserts the review date is identical in src/extract.py, src/prompt.py and tools/build_corpus.py, and refuses to let a run spend otherwise.
fixing the adapter's retry set to cover connection-level errors
reviews answered, and nothing else on this page. It would have made r002 55 of 55 instead of 54, and would not have changed a single verdict — every review the deliberating tier answered, it answered correctly.
reasoning
the single failure in r002 is <urlopen error [Errno 60] Operation timed out>; TRANSIENT in src/adapters/__init__.py is a set of HTTP status codes and urllib's URLError carries none.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
Binding hold identified
on any missed, invented or wrongly-named hold — 0 of those on either tier
Binding hold identified, free floors
on every review whose registry entry does not actually cover it, and on every successor line the first floor reads past
Eligibility verdict, free floors
for the first floor, on every review with any hold on file; for the second, on every review whose note is written in the register that contradicts the registry
Review flag
on a series that is frozen and already sitting in the destruction queue — a request to pull it OUT, never to put one in
Review flag, free floors
wherever the floor's own verdict happens to say frozen and the series is queued
Reviews answered
on any review that produced no parseable reply, for any reason
NextThe three you would add first
A completeness reconciliation: assert that every review sent produced a scored verdict before any batch acts on the output.the one real failure in this run was a review that never came back, and NOTHING in the kit refused because of it — a grader even scored the missing reply as correct. In a destruction queue, a series with no verdict must block the batch, not fall out of it.
Retry connection-level errors, not just HTTP statuses, in src/adapters/__init__.py.TRANSIENT is a set of status codes; a urllib socket timeout raises URLError and carries none, so it falls straight through a retry loop written for a busy provider. One line, and it is the difference between 54 and 55.
Re-derive the binding hold by pure code from the registry lines and compare it against the model's answer.the flag's blind spot is a confidently wrong hold identification, and on THIS corpus the registry is machine-written in a fixed four-column format, so a regex could re-run the coverage test exactly. It would be a real second opinion here and it would transfer to no real registry, which is worth knowing before building it.
A second binding hold. Make binding_hold_id a list.two live holds on one series is ordinary in a real programme and impossible to express here, which means the second hold is invisible the moment the first is released — the exact moment somebody looks at the record again.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-check the flag's evidence on any change to compute() or eligibility() in src/extract.py, on any change to AS_OF, on any regeneration of the corpus, and on any provider or model change. Re-scoring an existing run is free — python -m evals.run --rescore <file> — so a grader change never needs a paid re-run.
What this cannot tell you
Whether 'frozen and queued' is the condition a real records office would want flagged. It is this kit's own two-value simplification and no records programme was consulted.
Whether the flag holds when a review carries two binding holds, or a hold released part-way through the retention period. Neither shape exists in this corpus.
Whether the 1.00 precision survives a corpus where the queued share is not 30 of 55. The flag is a conjunction, so its base rate moves with the queue.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library. requirements.txt is empty on purpose: the corpus is generated in process, the model is reached over urllib, and the UI is http.server plus one file of vanilla JavaScript. There was nothing left to add.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
55 disposition reviews, generated from a fixed seed, never fetched. The class composition is EXACT and then shuffled rather than drawn per review, so the numbers on the page are the numbers the design asked for.
segmentation
src/segment.py
none -- one regex over underlined headings
There is no chunker and no tokeniser here. A review with no headings falls back to one honest whole-document segment rather than a spurious carve-up.
the model call
src/adapters/__init__.py
none -- raw HTTP over urllib
Two providers, one function each. No vendor SDK, so a forker installs nothing to run this on whichever key they already hold. Adding a provider is one function and one dict entry, and it must return token counts or its run cannot be published.
the rule
src/extract.py
none -- three conditions and a boolean pair
Everything that must not be probabilistic is plain Python: the eligibility derivation, the review flag and the span location. The model is asked only for the prose judgement.
scoring
evals/judge.py
none -- exact match and two confusion matrices
No judge model, so there is no judge to validate. What there IS to validate is the labels, and evals/check_labels.py does that before any run may spend.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per review — segment, select, prompt, one call, parse, derive, flag — and no branch, no loop and no agent. The only fan-out is over reviews, and they are independent.
The other sideWhat a framework costs you
No framework means no framework upgrade, no breaking release and no transitive dependency; it also means no retry policy, no rate limiter and no observability except what src/budget.py and the run record write. This run paid for that: a client-side socket timeout that any HTTP client library would retry by default lost a review.
Hand-rolled JSON extraction (src/prompt.py::parse) fails closed to an empty dict, which the harness records as a failed review. That is the right direction and it is not the same as a schema-enforcing client.
The UI is http.server on 127.0.0.1. It is a demonstration, not a deployment, and it says so.
What we could NOT verify
Whether a framework would have caught the timeout. Almost certainly yes — retrying connection-level errors is table stakes in every HTTP client — but nothing here was measured against one.
Whether the same prompt behaves identically through a vendor SDK rather than raw HTTP. The wire payload is the documented shape, and no SDK path was run.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-hold-conflict on the fast tier, 2026-08-22. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
3,873 ms
3873 ms p50 / 10887 ms p95 on the fast tier; 7369 ms p50 / 10645 ms p95 on the deliberating tier — about 90 pct slower at the median for no measured gain
—
Model, p95
10,887 ms
3873 ms p50 / 10887 ms p95 on the fast tier; 7369 ms p50 / 10645 ms p95 on the deliberating tier — about 90 pct slower at the median for no measured gain
—
Input tokens
91,915
1671.18 input and 485.69 output tokens per review on the fast tier, 1671.24 and 542.48 on the deliberating tier — input is identical because the prompt does not change
—
Output tokens
26,713
1671.18 input and 485.69 output tokens per review on the fast tier, 1671.24 and 542.48 on the deliberating tier — input is identical because the prompt does not change
—
No movement column. Not one of the 3 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-hold-conflict3,873 ms
r002-hold-conflict7,369 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
2 runs not plotted. b000-holds, b001-notes recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — each document goes whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-22, across 4 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
records-disposition reviews
data/corpus/*.txt — 55 files, generated once from a fixed seed, 40,925 bytes in total
read whole by src/segment.py and src/select.py; never modified, never uploaded, and only the mapped sections leave the machine — the custodian office is mapped by no field and is never sent
gold labels
data/gold.jsonl — 55 rows, written by the same generator
never. Gold is read only by evals/judge.py and evals/check_labels.py, both pure code running in process; it is never put in a prompt
run records
results/eval-*.json — two paid runs, two free floors, one ceiling calibration
committed to the repo. They record the model id and the provider ADAPTER name only — no base URL and no key
the provider key
the shared <repo>/.env, or the real environment
as one authorization header per call and nowhere else. It is gitignored, never logged, and redacted out of any adapter error before src/app.py returns it
the call ledger
<repo>/.calls-ledger.jsonl, beside the shared .env
never. One line per call, written BEFORE the call, so a crash over-counts rather than under-counts; it is what caps every kit sharing one key
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never committed, never logged and never rendered. .env is gitignored from the first commit. On an adapter error the app redacts the key and base URL out of the message before returning it, so a misconfigured endpoint cannot echo a credential into a browser. Result files record the model id and the provider adapter name only.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
model
one call per DISPOSITION REVIEW, carrying the mapped sections plus the fixed system prompt and field schema. Twelve fields come back, two of which are judgements rather than reads.
3873 ms p50 / 10887 ms p95 on the fast tier, 7369 ms p50 / 10645 ms p95 on the deliberating tier; 1671.18 input and 485.69 output tokens per review on the fast tier, 1671.24 and 542.48 on the deliberating one. (the fast-tier and deliberating-tier runs, 2026-08-22 -- see results/eval-r001-hold-conflict.json and eval-r002-hold-conflict.json)
one call per review, no concurrency and nothing shared between calls, so throughput is one review per round trip. And the retry set is HTTP statuses only: a client-side socket timeout is never retried, which lost one review of 55.
point src/adapters/__init__.py at a different provider or model and every number on this page is a different number.
corpus refresh
nothing, on purpose. The corpus is generated once from a fixed seed and read straight off disk — segmented into 12 addressable sections per review in about two milliseconds, with no model, no network and no index to invalidate.
660 sections across 55 reviews; 48 characters at the median, 140 at p95; 0.0022 s and $0.00 for the whole corpus. Regenerating it is one command and is byte-identical every time. (measured in process over the committed corpus; there is no run record because there is nothing to record — no call is made. tools/build_corpus.py, seed 20260822.)
there is nothing to refresh incrementally, which is the ceiling. A changed review is simply re-read and re-sent whole; nothing is cached between calls, so a corpus that changes daily costs a full run daily. And a review whose sections are not underlined headings collapses to ONE whole-document segment — extraction still works, every span reads 'document' and stops being evidence.
change the heading convention in tools/build_corpus.py and both the section map in src/select.py and every published span figure move; regenerate the corpus and every score on this page is void.
labels
55 gold rows in data/gold.jsonl, written by the same generator that wrote the reviews. Gold's verdict is not a typed opinion — it is the kit's own three-condition derivation re-run over the values each review states, and evals/judge.py re-derives it again at scoring time so the truth cannot drift from the rule.
55 of 55 eligibility verdicts on the fast tier and 54 of 54 on the deliberating tier's answered reviews; the flag fired 16 of 16 with 0 false alarms on both. Every class is asserted to produce the verdict it is named for before a run may spend. (evals/check_labels.py before either paid run; evals/judge.py::score_flags over both run records.)
scoring stops where gold stops. The derivation collapses the hold search to its RESULT, so no ungraded diagnostic can see a reply that names the wrong hold and then reasons about it correctly — that is the step the corpus exists to test, and only gold can grade it. One seed, 55 reviews, and no second corpus.
edit eligibility() or AS_OF in src/extract.py and gold, the prompt and the scorer all move in one edit — every verdict figure here is void and both paid runs would have to be fired again.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
binding_hold_id naming a hold whose own registry line reads released
a bug, and a specific one — a released hold never binds on its own. What SHOULD appear on those reviews is the id of the ACTIVE line whose scope reads 'continues the scope of '. Both tiers got all 6 of these right.
read the whole registry, not the first matching line: the binding hold is the active one, and on a successor review its scope is only a reference. (the 6 hold_successor reviews, both runs, all 6 answered with the successor's id.)
disposition_eligible answered 'no' on a review whose registry is empty
usually correct, and it is the case the over-cautious floor is blind to. A series is frozen with a clean registry when a longer-retention overlapping series has not expired, or when its own retention has not elapsed — 12 of the 28 frozen series here are frozen with no hold on them at all.
check the two expiry dates against the review date before concluding the registry decided anything. (the 6 overlap_longer and 6 retention_open reviews; both tiers answered all 12 correctly, the over-cautious floor got 2 of 12.)
disposition_eligible answered 'yes' on a review carrying an active hold that quotes this exact project
often correct, and it is the answer that feels wrong. 17 reviews carry a hold that matches on two of the three coverage tests and fails the third — wrong category, wrong project, or a date span that excludes the closed year.
check all three tests explicitly before treating a hold as covering. Two of three is the commonest wrong answer on this corpus. (the 6 scope_category_miss, 5 scope_project_miss and 6 scope_date_miss reviews; both tiers answered all 17 correctly, the over-cautious floor got 0.)
a review in the run record with an error and no fields
the call did not come back. On this run it was `` — a client-side socket timeout, which the adapter's retry loop does not cover because it keys on HTTP status.
do not read the run's percentages until you have read records next to the corpus size. A grader can score a missing reply as correct, and one here did. (RDS-0018, r002-hold-conflict — 1 of 110 paid calls.)
Whether a 55-review, single-seed run's clean result generalises to a real disposition queue. Both tiers answered every review they received correctly, which means this corpus has stopped separating them and is only still separating the model from the two free floors. Also unmeasured: a second seed, a registry with more than two holds on one series, a hold notice written by counsel rather than generated from four structured pieces, and whether an instruction hidden in the officer's note could move the verdict.
The corpus licence, from the Data lens: MIT — this repository's own licence. Every series id, custodian office, project name, schedule code and hold id is invented, no person is named anywhere, and no real agency, published general records schedule, disposition authority or litigation matter is named or reproduced. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one linePer-field exact match, light normalisation
Does the model's value for each of the twelve fields match gold, after trimming whitespace and trailing punctuation and lowercasing? A cell is a hit, a miss (nothing returned where gold has a value) or wrong (something else returned).
$0.00per 1,000 disposition reviews
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score, in-process, no key and no model. The same function evals/run.py calls after every run, and the same one --rescore re-runs over a committed result file's own cells.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The disposition review
RDS-0004
The field this row is about
binding_hold_id
The records category
financial
The project it relates to
Bell Harbor
When the series closed
2018-12
When its own retention expires
2025-12
The hold registry, verbatim
PR-2025-041 | active | Scope: continues the scope of AH-2023-234. | Continues: AH-2023-234
AH-2023-234 | released | Scope: all financial records relating to the Bell Harbor project, 2015 onward. | Superseded by: PR-2025-041
In the destruction queue?
queued
What the records officer wrote
Routine end-of-retention item. Nothing outstanding on this series.
Which hold the model named
PR-2025-041
Which hold actually binds
PR-2025-041
What the model answered
no
What the derivation says
no
Pulled out of the queue
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
All twelve fields matched gold, including officer_notes copied back verbatim with its double hyphen and overlapping_expires returned as null against a line reading 'none on file'.
Which hold freezes this series
correct
The registry carries AH-2023-234, whose scope covers this series exactly — 'all financial records relating to the Bell Harbor project, 2015 onward', and the series is financial, Bell Harbor, closed 2018-12 — and which reads released. It also carries PR-2025-041, ACTIVE, whose entire scope is 'continues the scope of AH-2023-234'. The model followed the reference and named PR-2025-041. Reading the released line and stopping answers null, and would have released a series under a live hold.
disposition_eligible against gold's own derivation
correct
A binding hold is named, so the first condition fires and the answer is 'no' — even though the series' own retention expired 2025-12, there is no overlapping series, and the officer's note reads 'Routine end-of-retention item. Nothing outstanding on this series.' Three separate reasons to say yes, and one that outranks all of them.
needs_review — frozen AND already in the destruction queue
correct
Not eligible AND queue_status is queued, so the pure-code rule fires: this series is one approval away from being destroyed under a live hold and has to come out of the queue today.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 100.0%
the free over-cautious floor
scored 91.1%
In operationWhat to monitor
Reference standard: tools/build_corpus.py's gold, read back off the same series metadata, schedule item, overlapping series and hold registry the generator wrote into the review text. Every gold value is asserted to be stated verbatim in the document it labels before the corpus is written.
These rates are UNKNOWN, on purpose
This grader IS the reference standard for field values — it cannot be scored against itself. What can go wrong is the normalisation: a model that returns 'GS-30-01, 7 years' where gold says 'GS-30-01' scores wrong, which is the intended reading (the hint asks for the code only) and would be the wrong reading on a corpus where the period belonged in the field.
Watch these
binding_hold_id and overlapping_expires — the two nullable fields, where a returned null is often correct and a default-to-miss scorer would punish it
retention_code, where the schedule line states a code AND a period and only the code is asked for
officer_notes, which must come back verbatim including its double hyphen
Alarm on
Any drop in extraction_accuracy below 100 pct on either tier — both runs were exact on every cell they answered (660 and 648), so the first miss is a signal and not noise.
How tight can the band be? There is no tolerance band. It is exact match after trimming whitespace and trailing punctuation and lowercasing; there are no numeric fields in this kit, so there is no rounding to forgive.
Cadence: Re-score on any change to src/prompt.py, data/fields.json or tools/build_corpus.py — the first two change what is asked for and the third changes what the answer is.
The decisionWhen to reach for it
Use it
Gold is read back off the same values the review states, never from a separate target — true of every kit corpus generated rather than collected.
Do not use it
The true field values are not known in advance — the normal state of a real disposition queue, and the reason this kit ships a corpus rather than pointing at one.
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineWhich hold freezes this series
Did the run name the same hold gold names — INCLUDING naming none when none binds? This is the prose scope judgement everything else rests on: three tests (category, project, date span) that must ALL pass, plus a reference to follow when an active line's whole scope reads 'continues the scope of <id>'.
$0.00per 1,000 disposition reviews
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_flags, in-process, no key and no model.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The disposition review
RDS-0004
The field this row is about
binding_hold_id
The records category
financial
The project it relates to
Bell Harbor
When the series closed
2018-12
When its own retention expires
2025-12
The hold registry, verbatim
PR-2025-041 | active | Scope: continues the scope of AH-2023-234. | Continues: AH-2023-234
AH-2023-234 | released | Scope: all financial records relating to the Bell Harbor project, 2015 onward. | Superseded by: PR-2025-041
In the destruction queue?
queued
What the records officer wrote
Routine end-of-retention item. Nothing outstanding on this series.
Which hold the model named
PR-2025-041
Which hold actually binds
PR-2025-041
What the model answered
no
What the derivation says
no
Pulled out of the queue
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
All twelve fields matched gold, including officer_notes copied back verbatim with its double hyphen and overlapping_expires returned as null against a line reading 'none on file'.
Which hold freezes this series
correct
The registry carries AH-2023-234, whose scope covers this series exactly — 'all financial records relating to the Bell Harbor project, 2015 onward', and the series is financial, Bell Harbor, closed 2018-12 — and which reads released. It also carries PR-2025-041, ACTIVE, whose entire scope is 'continues the scope of AH-2023-234'. The model followed the reference and named PR-2025-041. Reading the released line and stopping answers null, and would have released a series under a live hold.
disposition_eligible against gold's own derivation
correct
A binding hold is named, so the first condition fires and the answer is 'no' — even though the series' own retention expired 2025-12, there is no overlapping series, and the officer's note reads 'Routine end-of-retention item. Nothing outstanding on this series.' Three separate reasons to say yes, and one that outranks all of them.
needs_review — frozen AND already in the destruction queue
correct
Not eligible AND queue_status is queued, so the pure-code rule fires: this series is one approval away from being destroyed under a live hold and has to come out of the queue today.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 98.2%
the free over-cautious floor
scored 50.9%
the free tone floor
scored 70.9%
In operationWhat to monitor
Reference standard: tools/build_corpus.py::binding_hold_id(), the same registry-order search with the same one-level reference resolution the prompt states in words. Gold's binding hold is asserted to be an ACTIVE registry line before the corpus is written.
These rates are UNKNOWN, on purpose
Whether the judgement transfers to a hold notice written by a lawyer. Every scope here is one templated sentence built from a category list, a project and a date span; a real notice runs to paragraphs and its reach was negotiated rather than generated.
Watch these
missed_a_binding_hold — the expensive direction. A hold that binds and is not named releases a frozen series
the 6 hold_successor reviews, where the covering hold reads released and an active line continues it by reference
the 5 released_no_successor reviews, the mirror case that stops 'released means look for a successor' from being a working shortcut
the 17 reviews whose registry carries a hold that looks relevant and covers nothing — 6 wrong category, 5 wrong project, 6 outside the date span
Alarm on
Any missed_a_binding_hold or named_the_wrong_hold at all. Both tiers were at 0 across every review they answered.
How tight can the band be? No threshold — it is an id or a null. A review that produced no reply is unanswered and is never counted as a correct null, which is the defect this grader shipped with and had fixed.
Cadence: Re-run on any change to the scope wording in tools/build_corpus.py, to the binding_hold_id hint in data/fields.json, or to rules 2 and 3 in src/prompt.py.
The decisionWhen to reach for it
Use it
The registry is present in the review and its scopes are readable prose. This is the grader that says whether the model did the reasoning or got lucky on the verdict.
Do not use it
Holds live in a separate case-management system the review does not quote. Then there is nothing in the prompt to reason over and this kit is the wrong shape for the job.
disposition_eligible against gold's own derivation
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one linedisposition_eligible against gold's own derivation
Of every series that really is frozen, how many did the run refuse to release — and how many genuinely eligible series did it freeze? FROZEN IS THE POSITIVE CLASS: a series under a live hold that gets called eligible is the one that reaches a destruction batch, and it is the failure a records programme actually pays for.
$0.00per 1,000 disposition reviews
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_flags, in-process, no key and no model.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The disposition review
RDS-0004
The field this row is about
binding_hold_id
The records category
financial
The project it relates to
Bell Harbor
When the series closed
2018-12
When its own retention expires
2025-12
The hold registry, verbatim
PR-2025-041 | active | Scope: continues the scope of AH-2023-234. | Continues: AH-2023-234
AH-2023-234 | released | Scope: all financial records relating to the Bell Harbor project, 2015 onward. | Superseded by: PR-2025-041
In the destruction queue?
queued
What the records officer wrote
Routine end-of-retention item. Nothing outstanding on this series.
Which hold the model named
PR-2025-041
Which hold actually binds
PR-2025-041
What the model answered
no
What the derivation says
no
Pulled out of the queue
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
All twelve fields matched gold, including officer_notes copied back verbatim with its double hyphen and overlapping_expires returned as null against a line reading 'none on file'.
Which hold freezes this series
correct
The registry carries AH-2023-234, whose scope covers this series exactly — 'all financial records relating to the Bell Harbor project, 2015 onward', and the series is financial, Bell Harbor, closed 2018-12 — and which reads released. It also carries PR-2025-041, ACTIVE, whose entire scope is 'continues the scope of AH-2023-234'. The model followed the reference and named PR-2025-041. Reading the released line and stopping answers null, and would have released a series under a live hold.
disposition_eligible against gold's own derivation
correct
A binding hold is named, so the first condition fires and the answer is 'no' — even though the series' own retention expired 2025-12, there is no overlapping series, and the officer's note reads 'Routine end-of-retention item. Nothing outstanding on this series.' Three separate reasons to say yes, and one that outranks all of them.
needs_review — frozen AND already in the destruction queue
correct
Not eligible AND queue_status is queued, so the pure-code rule fires: this series is one approval away from being destroyed under a live hold and has to come out of the queue today.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 98.2%
the free over-cautious floor
scored 41.8%
the free tone floor
scored 60.0%
In operationWhat to monitor
Reference standard: Gold's verdict is re-derived inside the grader by the same three-condition rule the kit publishes — a binding hold first, then a longer-retention overlapping series that has not expired, then the series' own retention, all against the review date 2026-08 — so the truth this matrix grades against can never be a separately-typed label that drifted from the rule.
These rates are UNKNOWN, on purpose
Whether the verdict is right on a review shaped unlike these: a hold scoped by custodian rather than by series, a partial release covering some custodians and not others, an event-triggered retention with no expiry month, or a series reclassified mid-retention.
Watch these
false_negative — a frozen series called eligible. This is the expensive direction and the one the over-cautious floor fails 10 times and the tone floor 13 times
the 6 overlap_longer reviews, where NOTHING in the registry is wrong and the series is frozen anyway
the 17 reviews carrying a hold that covers nothing — where the honest answer is 'release it' and the safe-looking answer is not
Alarm on
Any false negative at all. Both tiers were at 0 across every review they answered, so the first one is a signal and not noise.
How tight can the band be? No threshold — the verdict is one of two allowed values, and a reply that returns neither is counted as unanswered rather than folded into the correct-negative cell.
Cadence: Re-run whenever eligibility() or AS_OF in src/extract.py changes, whenever the corpus is regenerated, and on any provider or model change.
The decisionWhen to reach for it
Use it
The true eligibility is derivable from the review's own values — which is exactly when this kit is worth running at all.
Do not use it
The review does not carry the registry, both expiry dates and the series' own metadata together. The derivation returns None rather than guessing, and the row is not scored as a pass.
needs_review — frozen AND already in the destruction queue
Check an agency's expiring records for holds before destruction
PresenterOpens the private repo. Visible to admins only.
In one lineneeds_review — frozen AND already in the destruction queue
Is this the series somebody has to pull out of the destruction queue today? It fires when disposition_eligible is no AND queue_status is queued. Note the direction: it never proposes a destruction, it asks for one to be stopped.
$0.00per 1,000 disposition reviews
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_flags, in-process, no key and no model. The same function src/extract.py runs on every reply and src/app.py runs on every click.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The disposition review
RDS-0004
The field this row is about
binding_hold_id
The records category
financial
The project it relates to
Bell Harbor
When the series closed
2018-12
When its own retention expires
2025-12
The hold registry, verbatim
PR-2025-041 | active | Scope: continues the scope of AH-2023-234. | Continues: AH-2023-234
AH-2023-234 | released | Scope: all financial records relating to the Bell Harbor project, 2015 onward. | Superseded by: PR-2025-041
In the destruction queue?
queued
What the records officer wrote
Routine end-of-retention item. Nothing outstanding on this series.
Which hold the model named
PR-2025-041
Which hold actually binds
PR-2025-041
What the model answered
no
What the derivation says
no
Pulled out of the queue
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
All twelve fields matched gold, including officer_notes copied back verbatim with its double hyphen and overlapping_expires returned as null against a line reading 'none on file'.
Which hold freezes this series
correct
The registry carries AH-2023-234, whose scope covers this series exactly — 'all financial records relating to the Bell Harbor project, 2015 onward', and the series is financial, Bell Harbor, closed 2018-12 — and which reads released. It also carries PR-2025-041, ACTIVE, whose entire scope is 'continues the scope of AH-2023-234'. The model followed the reference and named PR-2025-041. Reading the released line and stopping answers null, and would have released a series under a live hold.
disposition_eligible against gold's own derivation
correct
A binding hold is named, so the first condition fires and the answer is 'no' — even though the series' own retention expired 2025-12, there is no overlapping series, and the officer's note reads 'Routine end-of-retention item. Nothing outstanding on this series.' Three separate reasons to say yes, and one that outranks all of them.
needs_review — frozen AND already in the destruction queue
correct
Not eligible AND queue_status is queued, so the pure-code rule fires: this series is one approval away from being destroyed under a live hold and has to come out of the queue today.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 98.2%
the free over-cautious floor
scored 67.3%
the free tone floor
scored 78.2%
In operationWhat to monitor
Reference standard: The same two-value rule (src/extract.py::compute) run over GOLD's own values. It is a business condition, so unlike a self-consistency check it genuinely needs labels — and saying so is half of what makes the number believable.
These rates are UNKNOWN, on purpose
Whether 'frozen and queued' is the right condition for a real records office. It is this kit's own simplification; a real one weighs how close the batch is, who issued the hold, and whether the series has already been certified.
Watch these
false_negative — a frozen, queued series NOT flagged. That is a record destroyed under a live hold
the flag inherits whatever disposition_eligible says, so every verdict error upstream can become a flag error here — which is what happens to both floors
Alarm on
Any false negative. Both tiers fired on all 16 and raised no false alarms.
How tight can the band be? No threshold. Two values, both enums; if either is missing the rule returns None and the row is unanswered, never a quiet 'no'.
Cadence: Re-check on any change to compute() in src/extract.py, and whenever the corpus's queued share changes.
The decisionWhen to reach for it
Use it
Both values come back. The flag is pure code over the model's own output, so it is as reliable as the field it reads and no more.
Do not use it
Your queue state is not in the review. Then this is a join against another system, not a value in a reply.
A living map of modern AI — kept current every morning