Every trading partner over its limit needs someone who can say who signed off, and that answer is often buried in old committee notes. This app reads the review file and names who decided, when, and whether today's review needs a new alert.
PresenterOpens the private repo. Visible to admins only.
For a commodity credit officerAgriculture & Food
Why it matters
Today's manual process, and the same job with the app
A commodity trading desk's credit officer, reviewing counterparties on a fixed schedule.
✕Today's manual process
1Read every counterparty's file looking for anyone whose exposure now tops its credit limit.
2Trace the committee notes back through past reviews to find who actually signed off.
3Remember what was already flagged so the same episode is not raised as new twice.
4Miss a stale approval and a partner stays cleared long after its file needed another look.
Every counterparty re-checked manually
✓With the app
1Every file is read and any partner over its limit or growing fast is flagged.
2The decision is followed including into the committee notes, or marked not stated if it never resolves.
3Earlier alerts are remembered so this review is never treated as the first sighting.
4One alert per episode raised only on the review that first sees it active.
Only active files reach a person
See it work
One real case: what the app found, step by step
CPTY-0043's second review is over its limit, and the file cannot yet say who approved it.
Track who approved a risky trading partnerReference appBuilt to be shaped to your process
4
1The file's numbers credit limit, current exposure, and the review this covers.
2The verdict still escalated, and the officer who signed off is not stated.
3Why, and what's next the committee note never resolves, and this alert was already raised.
4What stays on this machine one file never leaves it, whatever the review decides.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
"Is this counterparty over its limit" is a SELECT and nobody needs a language model for it. The question a commodity credit desk actually has to answer is the next one: has somebody with the authority to do so already reviewed this file, and if so who, recorded when, on what decision. That fact lives in prose -- a credit-decision-log entry, sometimes naming an officer directly, sometimes only a committee minute number -- and it is written once, in one review's window, and never restated on a later review's page. On this corpus, 3 of the 55 governed readings reference a committee minute that is never actually resolved in that window's notes, where the only honest answer is that nobody can currently say who decided. Today a credit analyst works down an exposure report by eye, tracing each minute reference back through the relationship notes and remembering whether anybody already raised it. A credit analyst re-reading every active counterparty's decision log by eye each review: tracing a committee-minute reference through the relationship notes to find who actually signed off, telling a decoy note about an unrelated counterparty apart from a real one, and remembering whether anybody already raised this exact episode.
Audience
A commodity-desk credit officer deciding what to escalate this review, and the credit manager who has to say whether a given decision is still the one governing a file. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual counterparty review files
The corpus is 150 counterparty review files, 0.31 MB (json 1 · jsonl 1 · txt 150). Fifty counterparties and eight archetypes -- deliberately, so both arms of Rule X-1 (breach and fast growth) fire, at least one committee reference resolves and at least one does not, at least one decoy note about an unrelated counterparty appears, and at least one decision is reversed by a later one -- were chosen on ANSWER-CLASS BALANCE alone, before any reader existed. No model had been called and no floor had been written when the seed was picked.
The corpus
The 150 counterparty review filesgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your counterparty review files. That is the whole change — there is no database to migrate.
One counterparty review file, as the model receives itCPTY-0001-R1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This is a synthetic counterparty credit file generated for the cpty-watch kit (adhana-ai/adhana-foundry-kits, MIT). No real counterparty, position or officer is described here.
Counterparty Profile
----------------------------------------------------------------
Counterparty ID : CPTY-0001
Counterparty name : Elkhorn Capital Partners
Counterparty type : Input Retailer
Region : Texas Panhandle
Onboarded : 2022-08-01
Review sequence : (review 1 of 3 scheduled reviews)
Review date : 2026-01-05
Credit Limit File
----------------------------------------------------------------
Credit limit on file : 2,990,000.00 USD
Limit set by : Sofia Bergstrom, VP, Credit Risk
Limit set on : 2025-12-07
Review cadence : monthly, first business day
Position Ledger
----------------------------------------------------------------
DEFERRED-PAYMENT corn 235,000 bu @ 4.00/bu Exposure: 851,199.94 USD
FWD-SALE corn 170,000 bu @ 5.84/bu Exposure: 628,616.59 USD
BASIS-CONTRACT corn 68,000 bu @ 3.50/bu Exposure: 758,450.15 USD
Total current exposure : 2,238,266.68 USD
Exposure History
----------------------------------------------------------------
Current exposure (this review) : 2,238,266.68 USD
Same review, one year ago : 1,889,542.43 USD
Year-over-year change : +18.5 pct
Credit Decision Log
----------------------------------------------------------------
No new entries this review.
Relationship Notes
Abridged — the file continues.
The outcomeWhat a good result looks like
Every counterparty whose file is active (breached or growing 35 pct or more year over year) is on the watchlist with its governing decision named -- type, officer and date -- and exactly one new alert per episode, raised on the review that first sees it active, not re-raised on every review after.
And when it cannot
Two directions, and they cost different things. Reporting a file CLEAR or APPROVED_CONTINUE when the answer key says it is still active means the desk stops watching exposure that is genuinely open. Raising WATCH or worse on a file the key says is fine costs an analyst's afternoon re-reading a file that did not need it. The scored run made NEITHER: 0 of 52 active files wrongly cleared, 0 of 98 quiet files falsely alarmed. The strongest free floor also made zero of both -- the two readers disagree nowhere on WHETHER a file needs watching, only on WHO is named as having decided it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Confirming a counterparty is over its limit or growing fast — the naive floor 100.0 pct accurate, $0.00, and the model adds nothing measurable on this column.
Naming who currently governs a straightforward, directly-recorded decision — keyword-mem matches the model exactly whenever the entry names an officer directly.
Resolving a committee-minute reference, or knowing when it does not resolve — the model 100.0 pct officer accuracy and 100.0 pct NOT-STATED recall against the floor's 90.0 pct and 0.0 pct.
At a glanceHow the whole thing runs
100%watch status accuracy pct
3,174 msp50, end to end
$2.00per 1,000 counterparty review files · Google Gemini 3 Flash
Run once, for real, on 2026-08-24. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Track who approved a risky trading partner14 steps · 4 questions · run once, for real · 2026-08-24
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py at your own counterparty book, keeping the eight section headings src/segment.py expects (Synthetic Record, Counterparty Profile, Credit Limit File, Position Ledger, Exposure History, Credit Decision Log, Relationship Notes, Counterparty Contact Record). A forker's own corpus will not carry this kit's specific committee-reference mechanic unless it is built in deliberately.Corpus lens →
When is this the wrong choice?
Avoid: Paying for a model call to answer a comparison a spreadsheet already gets right. That is the case against the best-fitting scenario (“Confirming a counterparty is over its limit or growing fast”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
A committee-minute reference whose resolving text lives outside the Relationship Notes window entirely -- e.g. in a separate minutes archive this kit never reads. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether a real credit committee's minutes are typically resolvable in the same system of record the decision log lives in was not measured -- the whole officer-accuracy finding rests on this corpus's own, invented committee-reference mechanic. 4 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-24 — r001-cpty-watch. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — python3 -m src.app renders immediately with no key: the corpus, the answer key, all three floors and the r001/s001 results ship in the repo. It cannot make a live call -- the read button says so in words rather than erroring -- and the second button replays what run r001 actually answered, straight off the committed result file.
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
3,174 msp50, end to end
6,437 msp95
1 minclone to first result
What the clock covers. END-TO-END per reading: one HTTP request carrying the assembled prompt, parsed to a watch status, a governing decision (type, officer, date) and an alert call. Segmenting the file, withholding the Counterparty Contact Record section, parsing exposure and limit off the page, and advancing the carried state all happen outside this measurement and cost no network at all.
Current processWhat it replaces
A credit analyst re-reading every active counterparty's decision log by eye each review: tracing a committee-minute reference through the relationship notes to find who actually signed off, telling a decoy note about an unrelated counterparty apart from a real one, and remembering whether anybody already raised this exact episode.
Where it is not good enough
⚑ THE MONETARY STORY THIS KIT'S SIBLINGS TELL DID NOT MATERIALISE HERE, AND THE PAGE SAYS SO RATHER THAN MANUFACTURING ONE. Breach and fast-growth detection is pure arithmetic that every reader tested -- including the naive floor with no memory and no reading of the decision log at all -- gets right 100 pct of the time on this corpus; nothing was ever wrongly cleared, on any of the four readers. The entire measured gap between the model and the strongest free floor is 15 reader-officer misattributions in 150 readings (90.0 pct v. 100.0 pct), concentrated exactly on committee-referenced decisions the floor cannot resolve by keyword. That is real -- naming who decided is the actual business question on this kit's row -- but it is an ACCOUNTABILITY finding, not an exposure-dollars finding, and a reader expecting a dollars-at-risk headline the way this kit's monitor siblings publish one will not find it here. ⚠︎ ONE MODEL, ONE RUN, NO REPEAT. The 100 pct figures on watch status, decision type, decision date and alert timing are exact on this corpus's 150 readings; nothing here establishes they would hold on a second draw of counterparties, and a corpus built for ANSWER-CLASS BALANCE across archetypes is not a corpus built to be REPRESENTATIVE of any one real book's mix. ⚠︎ THE COMMITTEE-REFERENCE MECHANIC ITSELF IS INVENTED FOR THIS KIT. Whether a real credit committee's minutes are typically resolvable in the same system of record the decision log lives in, or scattered across email and a separate minutes archive this kit's model would never see, was not measured -- and it is exactly the assumption that decides whether this kit's headline finding survives contact with a real credit desk.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
→Sources
txt150jsonl1json1
50 counterparties, 150 readings across 3 scheduled reviews — every counterparty in the book re-read whole, on a clock
It produces a review watchlist for a credit officer to action — which counterparty's file needs watching, what governs it, who is on record as having decided — and never sets a credit limit or records a trade-on decision as taken; there is no setting that makes it. ⚑ TWO STATIONS MAKE THIS A MONITOR AND NEITHER IS THE MODEL. The carried state is five scalars written by src/state.py from the arithmetic, never from the model's reply, rendered into English the model cannot revise — the stateless control makes that concrete: overall accuracy falls from 100.0 to 88.0 pct, alert accuracy from 100.0 to 74.0, and the 44 memory-dependent readings fall to 61.36 pct. The clock is the second: one review owns the window since the last one, and this kit ships no equivalent of a sibling's cadence.py to price what a missed review would cost — an honest gap, not a hidden one. ⚑ THE HEADLINE IS NARROW ON PURPOSE. On the four columns that are mostly arithmetic, the strongest free floor TIES the model exactly at 100.0 pct: nothing on this corpus was ever wrongly cleared or falsely alarmed by any of the four readers tested. The entire measured difference is naming who decided — decision-officer accuracy 100.0 pct model against 90.0 pct floor, concentrated on committee-referenced decisions the floor cannot resolve by keyword, and NOT-STATED recall 100.0 pct against 0.0 pct on the 3 readings that genuinely cannot be resolved at all.
⚠︎ THAT IS AN ACCOUNTABILITY FINDING, NOT A DOLLARS-AT-RISK ONE, AND THE PAGE SAYS SO RATHER THAN MANUFACTURING A FIGURE TO MATCH ITS SIBLINGS' SHAPE.
⚠︎ AND THE MECHANIC ITSELF IS INVENTED: whether a real credit committee's minutes are typically resolvable in the same system of record the decision log lives in was not measured, and it is exactly the assumption this kit's whole headline rests on.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
the carried state
src/state.py
What is carried between reviews, and how it is worded to the model. Five scalars today, rendered as English rather than JSON -- a design choice, not a measurement, and it is in could_not_verify rather than described as a finding.
the credit policy
src/exposure.py
Rules X-1 through X-7, and GROWTH_FAST_PCT (35.0). The rule text is reproduced in the prompt precisely so it can be read, disbelieved and replaced with the credit appetite your own desk actually runs.
the decision reader (the floor)
evals/baseline.py
The keyword table and the officer fallback. Widen it and the floor gets stronger, which is the direction this kit wants: a floor built to lose measures nothing.
the corpus
tools/build_corpus.py
The counterparties, the archetypes, the officer and place-name pools, and the seed. Keep the eight section headings or src/segment.py's assertion refuses to start.
Components
Component
File
Role
the corpus generator
tools/build_corpus.py
Generates 50 counterparties x 3 scheduled reviews = 150 review files from a fixed seed (SEED = 20260824). Each counterparty follows one of eight archetypes -- clear, trips into an episode with no decision yet, trips and gets kept/restricted/held/escalated, already decided before review 1, or a decision reversal -- chosen by weighted random draw. The gold labels are src/exposure.step's output over the planted decision events, never typed by hand.
the credit-policy arithmetic
src/exposure.py
The rule as pure code: exposure vs. limit, year-over-year growth vs. 35 pct, and the six-way watch status against whichever decision currently governs. No model, no judgement. ⚠︎ The seven rules and the decision vocabulary are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
the carried state
src/state.py
SEAM 2 -- the thing that makes this a monitor. Five scalars per counterparty (alert raised, governing decision type/officer/date, watch status last reported), written from the arithmetic and never from the model's reply, rendered into two or three English sentences for the prompt. ⚠︎ alert_raised and decision_type are two separate facts and merging them would be a real bug: one boolean would either re-alert every review or stop watching an episode the moment it named a decision.
the section splitter
src/segment.py
Splits a review file into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 150 documents before a run may spend.
the section selector
src/select.py
Decides which sections reach the model. Counterparty Contact Record -- the finance contact's name, direct line and email -- is mapped by no field and is subtracted unconditionally, so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
the prompt
src/prompt.py
Three parts: the fixed instruction, the carried-state sentences, and the selected sections in document order. The stateless build replaces the sentences with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
the model adapter
src/adapters/__init__.py
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text.
the watch
src/watch.py
One counterparty, one scheduled review, one call. Parses exposure and limit off the page with a regex and parses NOTHING about the decision log -- deciding which decision governs and who made it is the entire task, and a kit that regexed the decision log would be measuring its own regex. Holds MAX_TOKENS = 8000, the published ceiling.
the local UI
src/app.py
One counterparty, one scheduled review, its carried state and its verdict, on 127.0.0.1:8201. Renders with no key. Shows the carried sentences verbatim, the free floor's answer beside the model's, and the withheld-section list underneath as evidence.
the three free floors
evals/baseline.py
naive -- the rule a credit-risk spreadsheet already runs, no memory, no reading of the decision log; naive-mem -- the same rule given the true carried state, which isolates what memory alone buys; keyword-mem -- the strongest free version: carried state plus a keyword scan of the window's Credit Decision Log for a directly-named officer, falling back to the credit-limit-setting officer's name when the entry instead points at a committee minute. 0 calls, $0.00, all scored through the identical scorer.
the scorer
evals/scoring.py
Exact match per cell against the computed gold, split ways an average would hide: five fields, wrongly-cleared and false-alarm counted apart with the exposure at risk beside them, the memory-dependent subset, and NOT-STATED recall on its own.
the pre-flight
evals/check_labels.py
Sixteen things that must be true before a run may spend: the eight sections parse, every counterparty's review sequence is complete and gap-free, no document restates an earlier window's decision text, the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path sets a credit limit or records a trade-on decision, the answer key replays from src/exposure.step, every parsed input matches the answer key, and the corpus actually exercises memory-dependence, the NOT-STATED case, a decoy note and a decision reversal.
Where it breaks at scale
The unit cost is a call per counterparty per review, so a book of 4,000 counterparties on a monthly cadence is 4,000 calls a month regardless of how many actually changed -- linear, and this kit ships no batching. Three things break before the call count does. First, nothing here detects a missed review or catches up a skipped one -- src/state.py is what a deployment would carry forward, and a scheduler that does not run is invisible to it. Second, a real deployment's carried-state file is one file replaced atomically -- correct for one writer, not a concurrency model. Third, and specific to this kit: the whole memory story rests on decision-log entries being WINDOWED -- shown once, on the review that logged them, and never restated. That is a property this corpus was authored to have. A real credit system that periodically re-prints a standing decision in a monthly summary (an entirely normal, helpful thing for a system to do for a human reader) would hand every review the same information a stateless build gets, and this kit's whole demonstrated memory value would evaporate with it -- untested here, because the corpus was built not to do that.
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
CPTY-0043-R2, the whole argument on one screen, and it was chosen by reading the answer key for the row where the free floor and the truth disagree -- not by looking for a flattering frame. The Credit Decision Log reads "ESCALATE TO COMMITTEE per Q2 Credit Committee, minute #38 -- see Relationship Notes", and that window's Relationship Notes says only that minute #38 is pending transcription. The keyword floor cannot read Relationship Notes at all: it falls back to the only other named officer on the page -- Aisha Cho, who set the credit limit -- and reports her as having escalated this counterparty. She did not. The model correctly reports the deciding officer as NOT STATED, matching every other cell exactly. ⚠︎ THE MODEL COLUMN HERE IS REPLAYED FROM THE COMMITTED RESULT FILE, NOT A LIVE CALL, and the column header says so in as many words.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
Before anything is asked. The two things this page has to get right are already on it: the carried state, verbatim as it goes into the prompt -- here "an earlier review already raised an alert ... it was last reported WATCH" -- and the free floor's answer, computed for nothing, already visible before any model is called. The model column is dashes throughout: a page that pre-fills a plausible-looking answer before the button is pressed would be indistinguishable from one that already spent.failureOpen full size →The same page with no API_KEY configured. It does not error and it does not go blank: the review file, the carried state, the parsed exposure and limit, the withheld-section list and the entire free floor are computed locally, so everything except the model column still renders, and a plain sentence explains that nothing was called.failureOpen full size →
How it is cutWhat one reviews (3 scheduled reviews each) is
No split, no chunking. The unit is a COUNTERPARTY -- three consecutive scheduled reviews processed strictly in order, because February's prompt contains carried state settled by January's reading. Each review file goes whole into one call, minus the withheld Counterparty Contact Record section.
SetupWhat the setup figure measured
No index is built. The corpus is generated once by tools/build_corpus.py and each review file goes whole into one call; build_seconds/build_cost_usd are 0.0 because there is no separate preparation step, not because one was skipped.
LicenceLicence
MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every counterparty, position, officer, committee minute and dollar figure is invented here. Verified against the repository's own LICENSE file on 2026-08-24.
Bring your ownBring your own counterparty review files
Point tools/build_corpus.py at your own counterparty book, keeping the eight section headings src/segment.py expects (Synthetic Record, Counterparty Profile, Credit Limit File, Position Ledger, Exposure History, Credit Decision Log, Relationship Notes, Counterparty Contact Record). Replace src/exposure.py's Rules X-1 through X-7 -- the growth threshold, the six-way watch status, the decision vocabulary -- with your own desk's credit appetite and governance process before publishing anything measured against real counterparties.
⚠︎ And what stops being true when you do: A forker's own corpus will not carry this kit's specific committee-reference mechanic unless it is built in deliberately. Without it, the officer-accuracy finding this kit's whole headline rests on -- the model beating the floor specifically on committee-referenced decisions -- has nothing to measure.
What breaks it
A committee-minute reference whose resolving text lives outside the Relationship Notes window entirely -- e.g. in a separate minutes archive this kit never reads. The corpus only ever resolves within the same window or not at all.
A credit limit that changes mid-relationship. This kit models no limit amendment; the limit on file is fixed across a counterparty's whole three-review history.
More than one governing decision landing in the SAME review's window -- the corpus and the arithmetic both assume at most one new decision per window.
A decision-log date that does not parse as YYYY-MM-DD, or a section-heading rename that defeats src/segment.py's parser (see select.py's fallback guard, which check_labels.py proves is load-bearing).
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
216
54
the question
3,268
817
Carried state
233
58
Credit policy note
42
11
Counterparty review file
2,052
513
Total
1,453
This is the cost lesson as arithmetic: of the 1,453 tokens assembled, 882 are instructions — 61% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Reassembled from src/prompt.build for CPTY-0043-R2 with that reading's real carried state, not retyped -- the harness does not log the assembled prompt string into the result file, only its length in characters (see run.py's prompt_parts). tokens.input above is a character-based estimate (chars / 4) for THIS exact prompt, 1453; the run's own provider-measured average across all 150 readings was 1411.67, within 3 pct, which is the basis for trusting the estimate rather than a separately logged per-reading figure.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a scheduled counterparty exposure review in a commodity merchandiser's credit risk team. You apply a written credit policy to one counterparty at one review. You answer with one JSON object and no other text.
You are the scheduled counterparty exposure review for a commodity merchandiser's credit desk. It
wakes monthly, re-reads EVERY counterparty's credit file, and reports each one. You are reading
ONE counterparty at ONE scheduled review.
The credit policy is reproduced below. Apply it exactly as written. Whether an alert has already
been raised and whether a decision already governs this counterparty are things you cannot see
directly -- what is known about them is stated under "Carried state" and is the only history
available to you. Do not assume anything about earlier reviews beyond it. A decision recorded in
an earlier review is NOT restated in this review's Credit Decision Log unless it changed.
How to decide:
- The file is ACTIVE this review if current exposure is OVER the credit limit on file, or if
exposure has grown 35 pct or more against the same review a year ago (Rule X-1). If the file is
not active, "watch_status" is CLEAR, "governing_decision_type" is NONE, and the episode is
closed even if a decision or alert was standing before (Rule X-7).
- If the file is ACTIVE: with no governing decision on record (from the carried state or from a
NEW entry in THIS review's Credit Decision Log window), "watch_status" is WATCH.
- If a decision IS on record, "watch_status" follows it: KEEP TRADING -> APPROVED_CONTINUE,
RESTRICT -> RESTRICTED, HOLD NEW TRADES -> HOLD_NEW, ESCALATE TO COMMITTEE -> ESCALATED
(Rule X-3).
- The GOVERNING decision is the most recently recorded one, never the first one on file (Rule
X-4). If THIS review's Credit Decision Log window logs a new decision, it replaces whatever was
carried, and you must report the NEW one -- do not keep echoing an older carried decision once
a newer one has been logged.
- Some log entries name a committee or a minute number rather than an officer directly. Where
that is so, find who the minute names as having decided it in THIS SAME review's Relationship
Notes -- never from a different counterparty's note, and never by assuming a name from the
Counterparty Profile or elsewhere. If the minute is not actually resolved in this window's
Relationship Notes, report "governing_decision_officer" as "NOT STATED" rather than guessing
(Rule X-5). Relationship Notes occasionally carry a line about a different, unrelated matter --
do not treat an unrelated sentence as a decision or as naming who made one.
- "governing_decision_date" is the date the governing decision was recorded, in the file's own
YYYY-MM-DD form, or "NONE" if no decision governs.
- "alert_now" is YES only on the review that FIRST finds the file active with no alert already
standing for the current episode (Rule X-6); otherwise NO, even while the file stays active
across several reviews.
Answer with a single JSON object and nothing else:
{"watch_status": "CLEAR|WATCH|APPROVED_CONTINUE|RESTRICTED|HOLD_NEW|ESCALATED",
"governing_decision_type": "NONE|KEEP_TRADING|RESTRICT|HOLD_NEW|ESCALATE_COMMITTEE",
"governing_decision_officer": "<name and title, or NONE, or NOT STATED>",
"governing_decision_date": "<YYYY-MM-DD, or NONE>",
"alert_now": "YES|NO",
"rationale": "one sentence, naming the rule you applied and what in the file drove it"}
Carried state
----------------------------------------------------------------
An earlier review already raised an alert for the episode this counterparty is currently in; do not treat this review as the first sighting of it. No governing decision is on record from an earlier review. It was last reported WATCH.
Credit policy
----------------------------------------------------------------
See Rule X-1 through X-7 as applied above.
Counterparty review file
----------------------------------------------------------------
Synthetic Record
----------------------------------------------------------------
This is a synthetic counterparty credit file generated for the cpty-watch kit (adhana-ai/adhana-foundry-kits, MIT). No real counterparty, position or officer is described here.
Counterparty Profile
----------------------------------------------------------------
Counterparty ID : CPTY-0043
Counterparty name : Salt Fork Port Grain Inc.
Counterparty type : Ethanol Plant
Region : Kansas Wheat Belt
Onboarded : 2025-09-11
Review sequence : (review 2 of 3 scheduled reviews)
Review date : 2026-02-02
Credit Limit File
----------------------------------------------------------------
Credit limit on file : 2,100,000.00 USD
Limit set by : Aisha Cho, VP, Credit Risk
Limit set on : 2025-09-11
Review cadence : monthly, first business day
Position Ledger
----------------------------------------------------------------
DEFERRED-PAYMENT soybeans 102,000 bu @ 4.06/bu Exposure: 1,001,658.10 USD
FWD-SALE corn 225,000 bu @ 6.74/bu Exposure: 695,097.62 USD
DEFERRED-PAYMENT corn 295,000 bu @ 4.89/bu Exposure: 931,722.25 USD
Total current exposure : 2,628,477.97 USD
Exposure History
----------------------------------------------------------------
Current exposure (this review) : 2,628,477.97 USD
Same review, one year ago : NOT AVAILABLE (onboarded within the last year)
Year-over-year change : NOT COMPUTABLE (no comparable prior-year review)
Credit Decision Log
----------------------------------------------------------------
2026-02-02 ESCALATE TO COMMITTEE per Q2 Credit Committee, minute #38 -- see Relationship Notes.
Relationship Notes
----------------------------------------------------------------
Minute #38 is pending transcription; full committee notes were not available at review time.
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{"watch_status": "ESCALATED", "governing_decision_type": "ESCALATE_COMMITTEE", "governing_decision_officer": "NOT STATED", "governing_decision_date": "2026-02-02", "alert_now": "NO", "rationale": "The file is ACTIVE because current exposure exceeds the credit limit, and the newly logged ESCALATE TO COMMITTEE decision governs, but the committee minute is pending transcription so the deciding officer is NOT STATED under Rule X-5; alert_now is NO because an earlier review already raised the alert for this episode under Rule X-6."}
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Track who approved a risky trading partner — 150 counterparty review files. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
Every cell is compared for exact equality against a gold set computed by tools/build_corpus.py from the planted decision-window events -- never from the rendered page. No model grades anything in this kit and no judge prompt exists. 150 of 150 readings parsed on the scored run (r001); 149 of 150 parsed on the stateless control (s001), with one reading (CPTY-0010-R1) cut off at the published 8,000-token ceiling with zero visible output, apparently consumed entirely by provider-side reasoning.
150counterparty review files
150source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED150 · 132 / 150watch status accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 / 150decision type accuracy pct — readings, with the carried stateDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 / 150decision officer accuracy pct — readings -- exact string match on who the file names as having decidedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 / 150decision date accuracy pct — readings, exact date match on the governing decisionDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED150 · 111 / 150alert accuracy pct — readings -- whether THIS review must raise a new alertDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED3 / 3not stated recall pct — readings whose committee-minute reference never actually resolvesDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 52wrongly cleared pct — readings whose file the answer key says needed watchingDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 98false alarm pct — readings whose file the answer key says was fineDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED44 · 27 / 44memory watch status accuracy pct — memory-dependent readingsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED44 / 44memory officer accuracy pct — memory-dependent readingsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The METHOD is validated by construction and the construction is ASSERTED: evals/check_labels.py re-derives every one of the 150 gold rows from the per-review decision-window events through src/exposure.step and requires them to match the committed key exactly (0 mismatches), and separately requires every value src/watch.position_of parses off the rendered page to equal the value the generator used (0 fields drifted). It also asserts no document restates an earlier window's decision text and that the Counterparty Contact Record guard would fail on 150 of 150 documents without the code that makes it hold on 0.
431.52output tokens · the fast tier, with the carried state · 3,174 ms p50
518.12output tokens · the same tier, STATELESS CONTROL · 2,765 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 0.9× as long, and lands one row apart on 150. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One counterparty review file
1,000 counterparty review files
Share that is the prompt
Google Gemini 3 Flash the shared card every kit in this series projects onto, so the rows compare across use cases rather than across price lists
$0.50 / $3.00
$0.002000
$2.00
35%
OpenAI GPT-5.6 Luna the cheapest row on the published set, for the low end of the range
$0.20 / $1.20
$0.000800
$0.80
35%
Anthropic Claude Fable 5 a frontier tier, for the other end of the range
$10.00 / $50.00
$0.035693
$35.69
40%
Same work, 45× the bill
The same counterparty review files, the same tokens — only the rate card changed. And across all 3 cards between 35% and 40% of what you pay is the prompt this pipeline sends, not the answer it writes.
Rates checked 2026-08-18. The provider that actually ran all 300 calls is kept out of these tables per this estate's naming rule, so nothing here is what was paid. The real spend for this kit's whole build is on the shared call ledger, not on this page.
Grading unitWhat the grading figure prices
freegrading cost, as measured
The GRADER costs nothing: evals/scoring.py makes no model call, and neither do the three free floors it is compared against. What cost money is the two model arms -- 150 readings each, r001 and s001.
The gradersOne way to grade, and why it is the only one
the fast tier, with the carried state 100.0% watch status accuracy · the same tier, memory removed (THE CONTROL) 88.0% watch status accuracy · the strongest free floor, no model 100.0% watch status accuracy · the same rule, GIVEN the carried state but not the window text 82.0% watch status accuracy · the rule a credit-risk spreadsheet already runs 70.7% watch status accuracy · 4 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
This labelled set DOES separate the readers, unlike some sibling kits' sets: the four arms (naive/naive-mem/keyword-mem/model) score 70.67/82.0/90.0-100.0/100.0 on decision officer accuracy respectively, a clean ordering with real gaps between each step.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Confirming a counterparty is over its limit or growing fast
the naive floor
100.0 pct accurate, $0.00, and the model adds nothing measurable on this column.
paying for a model call to answer a comparison a spreadsheet already gets right
Naming who currently governs a straightforward, directly-recorded decision
keyword-mem
matches the model exactly whenever the entry names an officer directly.
assuming the floor also works when the entry points at a committee minute instead
Resolving a committee-minute reference, or knowing when it does not resolve
the model
100.0 pct officer accuracy and 100.0 pct NOT-STATED recall against the floor's 90.0 pct and 0.0 pct.
trusting either reader's officer field without checking whether this corpus's committee mechanic resembles your own desk's real minute-taking
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
officer-misattribution
Floor names the limit-setting officer instead of the true decision-maker
9
CPTY-0043-R2: keyword-mem reports 'Aisha Cho, VP, Credit Risk' (who set the credit limit) as having escalated the file to committee; the true answer is NOT STATED, because the referenced minute #38 never resolves in that window's Relationship Notes.
stateless-realert
Without carried state, the model re-raises an alert on an episode already flagged
6
s001 (stateless control) false-alarms on 6 of 98 quiet readings that r001 (stateful) reports correctly, because the alert-raised fact is only ever stated once, in the carried-state paragraph the stateless build removes.
ceiling-cutoff
A reply consumed entirely by provider-side reasoning, cut off before any visible JSON
1
CPTY-0010-R1 in the stateless control: 8,000 of 8,000 output tokens spent, finish_reason=length, zero characters of visible answer text.
What we could NOT verify
Whether a real credit committee's minutes are typically resolvable in the same system of record the decision log lives in was not measured -- the whole officer-accuracy finding rests on this corpus's own, invented committee-reference mechanic.
Per-reading input token counts were not separately logged by this kit's harness, only run-level totals; the LLM lens's tokens.input for the published example is a character-based estimate, not a value read back from the provider for that specific call.
Whether CPTY-0010-R1's stateless-run ceiling cutoff would have resolved under a higher token ceiling was not tested -- no calibration run was fired to avoid spending an unbudgeted extra call on a control run's edge case.
Only one model tier was run. Nothing here measures whether a stronger or weaker tier holds the same 100 pct officer accuracy, or whether it degrades gracefully instead.
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
OpenAI GPT-5.6 Luna
Anthropic Claude Fable 5
the fast tier, with the carried state
1,411.67
431.52
3,174 ms
$0.002000
$0.000800
$0.035693
the same tier, STATELESS CONTROL
1,391.44
518.12
2,765 ms
$0.002250
$0.000900
$0.039820
The token counts are measured on a real run and belong to this pipeline — top_k, chunk size, prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
evals/scoring.py and evals/baseline.py make no model call; grading this kit's whole eval costs nothing beyond the two model arms already priced above.
Cost driversWhat actually moves the bill
Provider-side reasoning tokens -- 72.4 pct of r001's output budget and 77.1 pct of s001's, and the provider bills for them whether or not the visible answer uses any of it.
The size of the counterparty book, linearly -- one call per counterparty per review, no batching.
The review cadence -- halving the interval doubles the bill for the same book.
Your volumeWhat it costs at your volume
Linear. Ten times the counterparties is ten times the calls at the same $0.0020/reading; nothing here batches, caches or shares context across counterparties, so there is no volume discount built into the pipeline itself.
Where pricing changes shape
Your return, with your numbers
Volumeyour own counterparty book size and review cadence
What it replacesa credit analyst re-reading every active counterparty's decision log by eye each review, tracing committee-minute references and remembering whether an alert already stands
Time saved per itemnot measured here -- this kit publishes the per-reading cost and accuracy; how many analyst-minutes a correctly-named governing decision saves is a fact about your own desk's process, not something this run can measure
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The same fast tier every kit in this series measures, so the officer-accuracy finding is comparable across the estate rather than confounded by a different model choice.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
211,751input tokens · this run
64,728output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every stateful number on these pages: 150 readings, one completion call each, one tier.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.120
$0.120
$0.80
2026-09-12
gemini-3-flash
Google
$0.300
$0.300
$2.00
2026-09-18
gemini-3-8-flash
Google
$0.402
$0.402
$2.68
2026-09-18
claude-haiku-4-5
Anthropic
$0.535
$0.535
$3.57
2026-09-12
llama-5
Meta
$0.540
$0.540
$3.60
2026-09-18
grok-4-5
xAI
$0.812
$0.812
$5.41
2026-09-18
grok-4-6
xAI
$0.812
$0.812
$5.41
2026-09-18
claude-sonnet-5
Anthropic
$1.071
$1.071
$7.14
2026-09-12
gemini-3-1-pro
Google
$1.200
$1.200
$8.00
2026-09-18
gpt-5-6-terra
OpenAI
$1.200
$1.200
$8.00
2026-09-12
gpt-5-6-sol
OpenAI
$2.142
$2.142
$14.28
2026-09-12
claude-opus-4-8
Anthropic
$2.677
$2.677
$17.85
2026-09-12
claude-opus-5
Anthropic
$2.677
$2.677
$17.85
2026-09-12
claude-fable-5
Anthropic
$5.354
$5.354
$35.69
2026-09-18
claude-fable-5-1
Anthropic
$5.354
$5.354
$35.69
2026-09-18
gpt-6-astra
OpenAI
$5.354
$5.354
$35.69
2026-09-17
Read this against the numbers above
NOT MEASURED. Not one of these rows was run. They are this run's measured token counts multiplied by a published rate, and the rate is as of each row's own rates_as_of.
THE OUTPUT SIDE IS 72-77 PCT REASONING AND THAT IS WHAT MAKES THIS TABLE MOVE. A model that does not emit reasoning tokens, or does not bill them, would land nowhere near its row here for the same answers. Output is priced 3x to 6x input on every card, so this is the whole spread.
ACCURACY IS NOT PROJECTED. Every figure here is a price for the same token counts; nothing says another model would answer the same way, and on this task the strongest free floor already ties the measured model on four of five fields.
ONE READING'S CEILING FAILURE IS NOT IN THESE TOTALS. It occurred on the stateless control (s001), not on r001, whose tokens are what this table projects.
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
12 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pythe corpus generator — a swap seam
Generates 50 counterparties x 3 scheduled reviews = 150 review files from a fixed seed (SEED = 20260824). Each counterparty follows one of eight archetypes -- clear, trips into an episode with no decision yet, trips and gets kept/restricted/held/escalated, already decided before review 1, or a decision reversal -- chosen by weighted random draw. The gold labels are src/exposure.step's output over the planted decision events, never typed by hand.
You change it to: The counterparties, the archetypes, the officer and place-name pools, and the seed. Keep the eight section headings or src/segment.py's assertion refuses to start.
tools/build_corpus.py
# Generate the corpus and the answer key. Deterministic: one seed, one corpus, forever.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
CPTY_JSON = os.path.join(HERE, "data", "counterparties.json")
STATS = os.path.join(HERE, "data", "corpus-stats.json")
SEED = 20260824
N_CPTY = 50
REVIEW_DATES = ["2026-01-05", "2026-02-02", "2026-03-02"]
RULE = "-" * 64
src/exposure.pythe credit-policy arithmetic — a swap seam
The rule as pure code: exposure vs. limit, year-over-year growth vs. 35 pct, and the six-way watch status against whichever decision currently governs. No model, no judgement. ⚠︎ The seven rules and the decision vocabulary are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus.
You change it to: Rules X-1 through X-7, and GROWTH_FAST_PCT (35.0). The rule text is reproduced in the prompt precisely so it can be read, disbelieved and replaced with the credit appetite your own desk actually runs.
src/exposure.py
# The breach-and-decision rule as arithmetic. Pure code, no model, no library.
CLEAR = "CLEAR"
WATCH = "WATCH"
APPROVED_CONTINUE = "APPROVED_CONTINUE"
RESTRICTED = "RESTRICTED"
HOLD_NEW = "HOLD_NEW"
ESCALATED = "ESCALATED"
WATCH_STATUSES = (CLEAR, WATCH, APPROVED_CONTINUE, RESTRICTED, HOLD_NEW, ESCALATED)
NONE_DECISION = "NONE"
KEEP_TRADING = "KEEP_TRADING"
src/state.pythe carried state — a swap seam
SEAM 2 -- the thing that makes this a monitor. Five scalars per counterparty (alert raised, governing decision type/officer/date, watch status last reported), written from the arithmetic and never from the model's reply, rendered into two or three English sentences for the prompt. ⚠︎ alert_raised and decision_type are two separate facts and merging them would be a real bug: one boolean would either re-alert every review or stop watching an episode the moment it named a decision.
You change it to: What is carried between reviews, and how it is worded to the model. Five scalars today, rendered as English rather than JSON -- a design choice, not a measurement, and it is in could_not_verify rather than described as a finding.
src/state.py
# The carried state -- the thing that makes this a monitor and not a one-shot classifier.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DEFAULT_PATH = os.path.join(HERE, "data", "state.json")
def load(path=DEFAULT_PATH):
def save(store, path=DEFAULT_PATH):
def for_cpty(store, cpty_id):
def describe(state):
src/segment.pythe section splitter
Splits a review file into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 150 documents before a run may spend.
src/segment.py
# Split a counterparty review document into its named sections. Pure code, no model.
SECTIONS = ("Synthetic Record", "Counterparty Profile", "Credit Limit File", "Position Ledger",
RULE = "-" * 64
def split(text):
def render(secs):
src/select.pythe section selector
Decides which sections reach the model. Counterparty Contact Record -- the finance contact's name, direct line and email -- is mapped by no field and is subtracted unconditionally, so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/select.py
# Pick which sections of a review are sent. Pure code -- the last deterministic step before the
BANNER = "Synthetic Record"
PROFILE = "Counterparty Profile"
LIMIT = "Credit Limit File"
POSITIONS = "Position Ledger"
EXPOSURE = "Exposure History"
DECISIONS = "Credit Decision Log"
NOTES = "Relationship Notes"
CONTACT = "Counterparty Contact Record"
NEVER_SENT = (CONTACT,)
src/prompt.pythe prompt
Three parts: the fixed instruction, the carried-state sentences, and the selected sections in document order. The stateless build replaces the sentences with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/prompt.py
# Assemble the one prompt this kit sends. No template engine, no framework -- string
INSTRUCTION = """\
def build(text, carried, stateless=False):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text.
You change it to: PROVIDER, BASE_URL, API_KEY and MODEL in .env. Adding a provider is one function and one entry in PROVIDERS; it must return token counts, because lens 05 publishes them and lens 07 prices them.
src/adapters/__init__.py
# SEAM 1 -- the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens, thinking=None):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
THINKING_OFF = {"type": "disabled"}
def complete(cfg, system, user, max_tokens=1024, thinking=None):
src/watch.pythe watch
One counterparty, one scheduled review, one call. Parses exposure and limit off the page with a regex and parses NOTHING about the decision log -- deciding which decision governs and who made it is the entire task, and a kit that regexed the decision log would be measuring its own regex. Holds MAX_TOKENS = 8000, the published ceiling.
src/watch.py
# One counterparty, one scheduled review, one model call. The runtime half of the kit.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
SYSTEM = ("You are a scheduled counterparty exposure review in a commodity merchandiser's credit "
MAX_TOKENS = 8000
FIELDS = ("watch_status", "governing_decision_type", "governing_decision_officer",
def documents():
def counterparties():
def load_doc(doc_id):
def position_of(text):
src/app.pythe local UI
One counterparty, one scheduled review, its carried state and its verdict, on 127.0.0.1:8201. Renders with no key. Shows the carried sentences verbatim, the free floor's answer beside the model's, and the withheld-section list underneath as evidence.
src/app.py
# The minimal local UI. Standard library only -- python3 -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
PORT = int(os.environ.get("PORT", "8201"))
RECORDED_RUN = os.environ.get("RECORDED_RUN", "r001-cpty-watch")
GOLD_ROWS = {r["doc_id"]: r for r in
def carried_for(doc_id):
class H(BaseHTTPRequestHandler):
def main():
evals/baseline.pythe three free floors — a swap seam
naive -- the rule a credit-risk spreadsheet already runs, no memory, no reading of the decision log; naive-mem -- the same rule given the true carried state, which isolates what memory alone buys; keyword-mem -- the strongest free version: carried state plus a keyword scan of the window's Credit Decision Log for a directly-named officer, falling back to the credit-limit-setting officer's name when the entry instead points at a committee minute. 0 calls, $0.00, all scored through the identical scorer.
You change it to: The keyword table and the officer fallback. Widen it and the floor gets stronger, which is the direction this kit wants: a floor built to lose measures nothing.
evals/baseline.py
# THE FREE FLOORS. Three of them, none a strawman, all of them $0.00.
MODES = ("naive", "naive-mem", "keyword-mem")
DECISION_WORDS = [
DIRECT_RE = re.compile(r"recorded by ([A-Za-z][A-Za-z .]+), ([A-Za-z][A-Za-z, ]+)\.")
DATE_RE = re.compile(r"^(\d{4}-\d{2}-\d{2})\s")
LIMIT_OFFICER_RE = re.compile(r"^Limit set by\s+: ([A-Za-z][A-Za-z .]+), ([A-Za-z][A-Za-z, ]+)$",
def _decision_window_text(text):
def _guess_decision(text):
def review(text, carried=None, mode="keyword-mem"):
evals/scoring.pythe scorer
Exact match per cell against the computed gold, split ways an average would hide: five fields, wrongly-cleared and false-alarm counted apart with the exposure at risk beside them, the memory-dependent subset, and NOT-STATED recall on its own.
evals/scoring.py
# Score a run against the answer key. Pure code, exact match per cell. No model grades anything.
FIELDS = ("watch_status", "governing_decision_type", "governing_decision_officer",
ACTIVE = ("WATCH", "RESTRICTED", "HOLD_NEW", "ESCALATED")
def _pct(n, d):
def score(records, golds):
evals/check_labels.pythe pre-flight
Sixteen things that must be true before a run may spend: the eight sections parse, every counterparty's review sequence is complete and gap-free, no document restates an earlier window's decision text, the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path sets a credit limit or records a trade-on decision, the answer key replays from src/exposure.step, every parsed input matches the answer key, and the corpus actually exercises memory-dependence, the NOT-STATED case, a decoy note and a decision reversal.
evals/check_labels.py
# Everything that must be true BEFORE a run is allowed to spend. Free, and it runs in seconds.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GOLD = os.path.join(HERE, "data", "gold.jsonl")
CPTY_JSON = os.path.join(HERE, "data", "counterparties.json")
FAILED = []
def check(name, ok, detail=""):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 50 counterparties x 3 scheduled reviews = 150 review files from a fixed seed (SEED = 20260824). Each counterparty follows one of eight archetypes -- clear, trips into an episode with no decision yet, trips and gets kept/restricted/held/escalated, already decided before review 1, or a decision reversal -- chosen by weighted random draw. The gold labels are src/exposure.step's output over the planted decision events, never typed by hand. A swap seam.
src/exposure.pyThe rule as pure code: exposure vs. limit, year-over-year growth vs. 35 pct, and the six-way watch status against whichever decision currently governs. No model, no judgement. ⚠︎ The seven rules and the decision vocabulary are INVENTED and the kit ships that as a stated limitation -- see Data.why_this_corpus. A swap seam.
src/state.pySEAM 2 -- the thing that makes this a monitor. Five scalars per counterparty (alert raised, governing decision type/officer/date, watch status last reported), written from the arithmetic and never from the model's reply, rendered into two or three English sentences for the prompt. ⚠︎ alert_raised and decision_type are two separate facts and merging them would be a real bug: one boolean would either re-alert every review or stop watching an episode the moment it named a decision. A swap seam.
src/segment.pySplits a review file into its eight named sections. Pure code; evals/check_labels.py asserts all eight are present in all 150 documents before a run may spend.
src/select.pyDecides which sections reach the model. Counterparty Contact Record -- the finance contact's name, direct line and email -- is mapped by no field and is subtracted unconditionally, so the one section whose subject is a PERSON never leaves the machine even when a source-system rename makes every hint match nothing.
src/prompt.pyThree parts: the fixed instruction, the carried-state sentences, and the selected sections in document order. The stateless build replaces the sentences with a line saying no history is available -- byte-identical everywhere else, which is what makes the control a control, and evals/check_labels.py asserts exactly that.
src/adapters/__init__.pySEAM 1 -- the model. Raw HTTP for every provider, standard library only. Returns input/output tokens, reasoning tokens and finish_reason beside the text. A swap seam.
src/watch.pyOne counterparty, one scheduled review, one call. Parses exposure and limit off the page with a regex and parses NOTHING about the decision log -- deciding which decision governs and who made it is the entire task, and a kit that regexed the decision log would be measuring its own regex. Holds MAX_TOKENS = 8000, the published ceiling.
evals/baseline.pynaive -- the rule a credit-risk spreadsheet already runs, no memory, no reading of the decision log; naive-mem -- the same rule given the true carried state, which isolates what memory alone buys; keyword-mem -- the strongest free version: carried state plus a keyword scan of the window's Credit Decision Log for a directly-named officer, falling back to the credit-limit-setting officer's name when the entry instead points at a committee minute. 0 calls, $0.00, all scored through the identical scorer. A swap seam.
evals/scoring.pyExact match per cell against the computed gold, split ways an average would hide: five fields, wrongly-cleared and false-alarm counted apart with the exposure at risk beside them, the memory-dependent subset, and NOT-STATED recall on its own.
evals/check_labels.pySixteen things that must be true before a run may spend: the eight sections parse, every counterparty's review sequence is complete and gap-free, no document restates an earlier window's decision text, the privacy guard holds AND would fail without it, the two prompt builders differ on exactly one line, no code path sets a credit limit or records a trade-on decision, the answer key replays from src/exposure.step, every parsed input matches the answer key, and the corpus actually exercises memory-dependence, the NOT-STATED case, a decoy note and a decision reversal.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1411 input and 431 output tokens per query at top-k 1, of which about 150 tokens are fixed (system prompt, chat framing, the question itself) and the rest is retrieved context. Change the retrieval depth and the input scales with it. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model and top-k follow the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Retrieval is free here and stops being free at scale.This kit scores every chunk in Python; a corpus a hundred times larger needs a vector store, which is infrastructure this arithmetic does not price.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
No untrusted external input surface in the shipped version. Every field that reaches the prompt comes from tools/build_corpus.py, which invents all of it from a fixed seed. In a real deployment exactly one field is writable by somebody outside the credit desk -- the Relationship Notes, which a relationship manager types freely and into which a counterparty's own email could be pasted.
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged, never written to results/, and never rendered by the UI. src/app.py replaces it and the base URL with placeholders before any provider error reaches the browser -- untested against a provider that echoes a masked form of the key (see could_not_verify).
The experimentWe did NOT attack it -- and this is the surface we left open
An indirect prompt injection needs a field somebody outside the process can write into, and this kit's closest candidate is the Counterparty Contact Record -- a finance contact's name, direct line and email, carried verbatim on every review. No adversarial trial was fired against it. Both were run for real on 2026-08-24.
Boundary checked
What could go wrong
What the code guarantees
Does the counterparty's finance contact -- name, direct line, email -- ever leave the machine?
Every review file carries a Counterparty Contact Record. A kit that sends 'the document' sends all three, to a third party, on every reading, forever.
src/select.py maps no answered field to that section and _fallback() subtracts it unconditionally, so it cannot be reached even when a source-system rename makes every hint match nothing. evals/check_labels.py measures both directions: with the guard 0 of 150 leak, and with the naive or list(secs) fallback under a renamed schema 150 of 150 do.
Can anything here set a credit limit or record a trade-on decision as taken?
A counterparty-monitoring system that can compute a watch status is one commit away from being a system that acts on it -- setting a limit or approving continued trading is a commercial act, not a report.
There is no such code path, no such endpoint and no configuration flag that adds one. The only writers in the kit are evals/run.py (results/*.json) and src/state.py (data/state.json). evals/check_labels.py greps every .py and .js file for the names of such paths and passes at zero.
Can a provider error print the key into the browser?
Provider errors are passed through verbatim so a reader can see what was actually said, and provider errors routinely quote the request.
src/app.py replaces the api_key and base_url strings with [API_KEY] and [BASE_URL] before the message is serialised. ⚠︎ EXACT-MATCH ONLY, untested against a masked echo -- see could_not_verify.
The privacy guard is the one boundary measured in both directions -- with it, 0 of 150 reviews leak the contact section; simulating its absence, all 150 would.
The result0 attack trials, three boundaries checked -- the privacy boundary red-proven by removing the guard and watching all 150 documents leak, and the credit-decision boundary confirmed absent by a code-path grep.
1field an outside party could influence (sent, not hidden)
0attack trials fired
150documents that leak the contact section without the guard
The Relationship Notes ARE the field an outside party would influence in a real deployment, and this kit sends them -- including, deliberately, a decoy note about an unrelated counterparty on 5 of 150 documents. Nothing measured what a followed instruction embedded in a note would do to a reading.
Read this twice
The Relationship Notes reach the model verbatim — there is no filter between what a relationship manager types and the prompt. This kit produces a watchlist and sets no limit and stops no trade, so the worst a followed instruction can do here is misname who decided, which is exactly the accountability failure this kit measures.
HonestyWhat this does not prove
Whether a real deployment's Relationship Notes -- prose a relationship manager writes freely, often after a phone call with the counterparty -- would carry an instruction the model follows. Not applicable to this synthetic corpus, and not measured.
Whether the key redaction in src/app.py holds against a provider that echoes a masked form of the key rather than the verbatim string, matching a defect found by screenshot on a sibling kit. Not exercised here.
Whether the guardrail holds against a code path named something evals/check_labels.py does not know. It asserts the absence of names it knows.
Whether the withheld section stays withheld under a corpus that carries a NINTH section. The selector's guard is a subtraction and should, but nothing tests a shape this corpus cannot produce.
Whether the counterparty names, positions and officer names -- all invented -- would need different handling if a forker pointed this at a real book. Almost certainly yes, and nothing here helps with it.
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
No credit decision taken, non-configurable. This kit produces a watch status, a governing decision name and an alert call. It never sets a credit limit, never records a trade-on decision as taken, and never contacts a counterparty, and there is no setting that makes it.
Everywhere and nowhere -- it is a property of what is ABSENT. The only writers in the kit are evals/run.py (results/*.json) and src/state.py (data/state.json). There is no outbound path of any kind except the one completion call in src/adapters.
EvidenceDoes it hold?
What
Measured
Nothing in this kit sets a limit, records a decision as taken, or contacts a counterparty
0 code paths. evals/check_labels.py greps every .py and .js file in the kit for such names and passes at zero, with the checker itself exempted because it is the only file that has to spell them.
No officer name is ever guessed where the file does not resolve one
100.0 pct NOT-STATED recall over 3 readings on r001; 0.0 pct on the strongest free floor, which is exactly the defect -- it names the limit-setting officer whether or not a decision actually resolved to them.
The counterparty finance contact section never reaches the provider
0 of 150 with the guard, 150 of 150 without it, both measured before any run spent.
The model's answer never becomes the next review's memory
src/state.py is written only by src/exposure.step. evals/run.py advances the carried state from the answer key's own inputs even on a reading that errored, so one transport failure cannot turn into three scored ones.
A run measured under a non-published token ceiling cannot be mistaken for a scored one
evals/run.py refuses --max-tokens unless the run id begins with c. No calibration run was fired for this kit, so no run claims a ceiling the page does not name.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON WHETHER THE ANSWER IS RIGHT. The arithmetic is correct whatever the model says, which means a wrong reading is a wrong watchlist row and not a blocked action. The guardrail stops the kit acting; it does not stop it being wrong.
IT IS NOT A SCHEDULER. This is the half of a monitor the kit does not ship. evals/run.py is INVOKED, not woken, and nothing here detects a missed review.
IT IS NOT CREDIT POLICY AND THE RULES ARE INVENTED. Whether a given exposure figure should trigger escalation, restriction or a hold is a risk-appetite decision this kit does not make and should not be trusted to make.
WatchedWhat is watched, and why that one
6runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 25 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
18 measured by the latest run7 need the model half
Metric
Owner
Role
Why this one
cell-exact-match
Watch status, decision type, decision officer, decision date and the alert call, per reading, exact match against the computed answer key
alarm
watch_status_accuracy_pct; decision_type_accuracy_pct; decision_officer_accuracy_pct; decision_date_accuracy_pct; alert_accuracy_pct; unparsed_replies — alarm on unparsed_replies above 0. A reading that returns nothing is scored as a miss in all five fields, so a reliability failure arrives disguised as a quality failure. The scored run had 0 of 150; the stateless control had 1 of 150 at the identical ceiling, so this is a property of removing the carried-state paragraph, not of the ceiling alone.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
150
different corpus — nothing is comparable
corpus.bytes
322,147
counterparty review files edited — the count held, the bytes did not
split.count
50
the reviews count moved — a different set was scored
split.size_p50
3
the median size of one review moved
split.size_p95
3
the 95th-percentile size of one review moved
dataset.rows
150
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (active_cells 52, answered 150, cadence monthly, first business day, clear_cells 98, counterparties 50, documents 150, exposure_at_risk_usd_gold 114044274.15, failures 0, false_alarms 17, memory_cells 44, not_stated_cells 3, readings_scored 150, stateless False, wrongly_cleared 0) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
watch status accuracy
not yet known
150 readings
A band is the spread between repeats and this kit fired one run per arm. What IS known is the spread between ARMS on the same 150 readings: 70.67 to 100.0 pct.
decision officer accuracy
not yet known
150 readings
One run per arm.
NOT-STATED recall
not yet known
3 readings
0 or 100 pct on a denominator of 3 -- one reading moves this figure by a third.
alert accuracy
not yet known
150 readings
One run per arm.
the memory-dependent subset
not yet known
44 memory-dependent readings
One run per arm. This is the subset the whole monitor claim rests on.
latency
not yet known
150 answered readings
One run per arm, on a shared provider account that multiple kits were hitting at once, so the tail here is partly queueing.
Answered
100.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Decision date accuracy
100.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Decision type accuracy
100.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Exposure missed share
0.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Exposure missed USD
$0.00 on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
False alarm
0.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Input tokens, whole run
211,751 on r001-cpty-watch
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Memory-dependent decision type accuracy
100.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Output tokens, whole run
64,728 on r001-cpty-watch
the whole run
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
Wrongly cleared
0.00 pct on r001-cpty-watch
150 readings
One run per arm, so no repeat spread exists yet; the figure is r001-cpty-watch's own, re-derived from its result file by build/measured/runlog.py.
HistoryRun history
6 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked. They are split into 3 blocks because a free floor is a baseline rather than a peer, and a run of another kind measures another thing; a column in one block is never differenced against a column in another.
triage · no model in the path — a baseline, not a peer column — 3 runs. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
b000-cpty-watch-naive 2026-08-24
b001-cpty-watch-naivemem 2026-08-24
b002-cpty-watch-keywordmem 2026-08-24
alert accuracy, %
74.67
100.00
100.00
answered, %
100.0
100.0
100.0
decision date accuracy, %
70.67
82.00
100.00
decision officer accuracy, %
70.67
82.00
90.00
decision type accuracy, %
70.67
82.00
100.00
exposure missed share, %
0.0
0.0
0.0
exposure missed usd
0.0
0.0
0.0
false alarm, %
17.35
11.22
0.00
input tokens, whole run
0
0
0
model latency p50 ms
0.00
0.00
0.00
model latency p95 ms
0.00
0.00
0.00
memory decision type accuracy, %
15.91
54.55
100.00
memory officer accuracy, %
15.91
54.55
77.27
memory watch status accuracy, %
15.91
54.55
100.00
not stated recall, %
0.0
0.0
0.0
output tokens, whole run
0
0
0
watch status accuracy, %
70.67
82.00
100.00
wrongly cleared, %
0.0
0.0
0.0
not a time series No two of these 3 runs measured the same system — they differ on false_alarms, floor, the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
triage · with the model in the path — 2 runs. Columns here are only ever compared with each other.
Metric
r001-cpty-watch 2026-08-24
s001-cpty-watch-stateless 2026-08-24
alert accuracy, %
100.0
74.0
answered, %
100.00
99.33
decision date accuracy, %
100.0
88.0
decision officer accuracy, %
100.0
88.0
decision type accuracy, %
100.0
88.0
exposure missed share, %
0.0
0.0
exposure missed usd
0.0
0.0
false alarm, %
0.00
6.12
input tokens, whole run
211751
207325
model latency p50 ms
3174.00
2765.00
model latency p95 ms
6437.00
7895.00
memory decision type accuracy, %
100.00
61.36
memory officer accuracy, %
100.00
61.36
memory watch status accuracy, %
100.00
61.36
not stated recall, %
100.0
100.0
output tokens, whole run
64728
77200
watch status accuracy, %
100.0
88.0
wrongly cleared, %
0.0
0.0
not a time series No two of these 2 runs measured the same system — they differ on answered, documents, failures, false_alarms, stateless — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
stub · no model in the path — a baseline, not a peer column — 1 run. A floor is what the kit scores with no model in the path; it is the line the model runs have to beat, and it is never differenced against them.
Metric
t000-cpty-watch-stub 2026-08-24
alert accuracy, %
74.67
answered, %
100.0
decision date accuracy, %
70.67
decision officer accuracy, %
70.67
decision type accuracy, %
70.67
exposure missed share, %
0.0
exposure missed usd
0.0
false alarm, %
17.35
input tokens, whole run
212867
model latency p50 ms
0.00
model latency p95 ms
0.00
memory decision type accuracy, %
15.91
memory officer accuracy, %
15.91
memory watch status accuracy, %
15.91
not stated recall, %
0.0
output tokens, whole run
6750
watch status accuracy, %
70.67
wrongly cleared, %
0.0
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 18 chips that all say so.
DeviationsWhat deviated
0 breaches across 6 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
whether the prompt carries the carried state
watch status 88.0 -> 100.0 pct, alert accuracy 74.0 -> 100.0, memory-dependent watch status 61.36 -> 100.0
measured
r001-cpty-watch against s001-cpty-watch-stateless
whether the floor may read this window's Relationship Notes for a committee-referenced officer
decision officer accuracy 82.0 -> 90.0 pct (naive-mem to keyword-mem), never reaching the model's 100.0
measured
b001 against b002, both free
whether the floor is given the true carried state at all
watch status accuracy 70.67 -> 82.0 pct (naive to naive-mem)
measured
b000 against b001, both free
the published token ceiling
coverage 100.0 pct at 8,000 (r001); 99.33 pct at the same ceiling once memory is removed (s001), one reading cut off with zero visible output
measured
r001-cpty-watch against s001-cpty-watch-stateless
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
watch status accuracy
nothing yet. The strongest free floor ties the model exactly at 100.0 pct.
decision officer accuracy
nothing yet. This is the column the model WINS -- 100.0 pct against the strongest free floor's 90.0.
NOT-STATED recall
nothing yet. 100.0 pct on the model, 0.0 pct on every floor.
alert accuracy
nothing yet. 100.0 pct stateful, 74.0 pct stateless on the same readings.
the memory-dependent subset
nothing yet. 100.0 pct stateful against 61.36 pct stateless on the same 44 cells.
latency
nothing yet. p95 is roughly 2x p50.
Answered
nothing yet — a second scored run is what would give this column a spread to fire on.
Decision date accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Decision type accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Exposure missed share
nothing yet — a second scored run is what would give this column a spread to fire on.
Exposure missed USD
nothing yet — a second scored run is what would give this column a spread to fire on.
False alarm
nothing yet — a second scored run is what would give this column a spread to fire on.
Memory-dependent decision type accuracy
nothing yet — a second scored run is what would give this column a spread to fire on.
Wrongly cleared
nothing yet — a second scored run is what would give this column a spread to fire on.
NextThe three you would add first
A scheduler, and something that notices when a review did not happenNothing here detects a missed review, and unlike this kit's monitor siblings, no cadence.py experiment prices what one would cost on this book.
Provenance on the governing decision, carried alongside itThe carried state says a decision was recorded and by whom; it does not carry a document or minute reference a credit officer could pull up to verify it. That is four more bytes.
A second reader on any file whose committee reference does not resolve3 of 150 readings here are genuinely unresolvable and the model gets all 3 right by saying so -- which means it is reliably handing a human a short, real queue, and nothing downstream of this kit works that queue.
A cadence check against your own book's review intervalThis kit ships no equivalent of evals/cadence.py. Before trusting a monthly cadence on your own book, measure what a skipped review would cost against your own limits and growth rates.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-run evals/check_labels.py (free, seconds) on any change to tools/build_corpus.py, src/exposure.py, src/segment.py, src/select.py or src/prompt.py -- it is the gate that would catch a floor regression or a corpus drift. There is no cadence.py to re-run here; adding one is named in guardrails.add_first.
What this cannot tell you
One run per arm. Whether the 0 wrongly-cleared / 0 false-alarm record is stable across repeats is not measured.
Whether the guardrail holds against a code path named something the checker does not know.
Whether a followed instruction in the Relationship Notes could suppress or alter a reading. The surface is sent and named; no attack was fired.
Whether the officer-accuracy finding transfers to a real credit committee's minute-taking practice. It is a property of this generator's own invented mechanic.
Whether a counterparty added to the book mid-horizon, or removed from it, behaves sensibly. All 50 are live at all three reviews here.
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies at all -- json, os, re, random, http.server and concurrent.futures, all standard library. There is not even a templating library, which on a kit whose whole subject is assembling one prompt from named sections is the conspicuous omission: the entire need is string concatenation you can read in src/prompt.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the carried state
src/state.py
agent memory layers and conversation stores (LangGraph checkpointers, LangChain memory, a vector store of prior runs)
those carry a TRANSCRIPT; this carries five scalars written by arithmetic. You would gain persistence, concurrency and a query interface over history, and lose the thing that matters most on this task: a credit officer can be shown 'a RESTRICT decision was recorded on 2026-02-02, by Kwame Okonkwo' and check it against the file, and a checkpointed graph state cannot be checked against anything that simply.
the model
src/adapters/__init__.py
LiteLLM, LangChain chat models, any provider-abstraction layer
you would gain dozens of providers, retries, streaming and callbacks. Under 150 lines of urllib is the whole abstraction here and it returns the token counts lens 05 publishes and lens 07 prices; on this kit 72-77 pct of the output is reasoning tokens, which is exactly the field such layers most often normalise away or drop.
the decision reader
evals/baseline.py
a rules engine, or a trained extractor over credit-file text
you would gain something that generalises past this kit's own committee-minute phrasing, which the floor demonstrably does not. You would lose the point of the floor, which is that it is readable in one sitting and a forker can see exactly where it breaks -- on this corpus, any decision recorded by committee reference rather than a direct name.
the schedule
(not shipped)
cron, Airflow, Temporal, any durable scheduler
you would gain the half of a monitor this kit does not have -- reviews that actually happen -- and lose nothing this kit values. It is the first thing to add, named in guardrails.add_first. It is left out because the product here is a folder of readable Python and a scheduler would be the largest thing in the folder.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph, and the one place it looks like one is a straight line. Each counterparty is a chain of three scheduled reviews and the only edge between them is five scalars. Counterparties are independent of each other and run 50-wide.
The other sideWhat a framework costs you
No scheduler. This is a watch that carries state and is still INVOKED rather than woken, and nothing here measures what a missed review would cost -- unlike its monitor siblings, this kit ships no cadence.py.
No persistence layer. A real deployment's data/state.json is one file replaced atomically: correct for one writer, not a concurrency model.
No provenance on the carried decision. It says WHAT was decided and by WHOM, not by which document or minute, and a credit officer defending the call needs the second half.
What we could NOT verify
No port to any framework was built, so every row above is reasoning about the seams rather than a measured alternative implementation.
The scheduling seam is the one that matters most here and it is the one with no code at all, so nothing about it has been tried.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-cpty-watch on the fast tier, with the carried state, 2026-08-24. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
3,174 ms
not yet known
nothing yet. p95 is roughly 2x p50.
Model, p95
6,437 ms
not yet known
nothing yet. p95 is roughly 2x p50.
Input tokens
211,751
211,751 on r001-cpty-watch
—
Output tokens
64,728
64,728 on r001-cpty-watch
—
No movement column. Not one of the 4 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-cpty-watch3,174 ms
s001-cpty-watch-stateless2,765 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
3 runs not plotted. b000-cpty-watch-naive, b001-cpty-watch-naivemem, b002-cpty-watch-keywordmem recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — the population is re-read whole on each scheduled run, and the prior run's carried state is what makes a reading a CHANGE.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
9 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-24, across 6 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
counterparty review files
data/corpus/CPTY-<n>-R<k>.txt -- 150 files, 322147 bytes, generated once from a fixed seed by tools/build_corpus.py
seven of the eight sections go to the provider in the prompt; Counterparty Contact Record -- the finance contact's name, direct line and email -- never does, by src/select.NEVER_SENT, and evals/check_labels.py measures both directions before a run may spend
the answer key
data/gold.jsonl -- one row per reading, computed, not typed
never. It is read by evals/scoring.py and src/app.py on your machine and no part of it is ever put in a prompt -- a key in the prompt would be the answer in the question
the carried state
data/state.json in a deployment; scoped to the run and never written to disk inside evals/run.py
two or three SENTENCES of it do, in every prompt -- that is the experiment. They carry a decision type, an officer, a date and two booleans and nothing else; there is no transcript and no earlier document in them
every run this kit has fired
results/eval-*.json
never. They are written locally and committed to the public kits repo on purpose, so a reader with no key can replay what the scored run answered
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 38
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
The provider key lives only in .env (gitignored) or the shared repo-root .env, read by src/config.py and never logged, never written to results/, and never rendered by the UI. src/app.py replaces it and the base URL with placeholders before any provider error reaches the browser -- untested against a provider that echoes a masked form of the key (see could_not_verify).
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
cadence
a MONTHLY watch on the first business day -- 2026-01-05, 2026-02-02, 2026-03-02. One review owns exactly the window since the last one; a decision or an alert not logged in THIS window is invisible to this review and must be carried.
50 counterparties x 3 scheduled reviews = 150 calls per arm, 0 gaps and 0 out-of-order readings; evals/check_labels.py asserts all 50 sequences are complete before a run may spend. A book of 4,000 live counterparties on this cadence is 48,000 calls a year, about $96 on the shared projection card. (r001-cpty-watch, evals/check_labels.py, tools/build_corpus.py REVIEW_DATES)
⚠︎ NOTHING HERE MEASURES WHAT A MISSED REVIEW COSTS. Unlike this kit's monitor siblings, evals/ ships no cadence.py -- no experiment re-derives the answer key on a reduced schedule. That gap is recorded in not_measured, not argued around.
every figure on this page is per a MONTHLY watch on a population that stays fixed at 50 counterparties across all three reviews. A book that gains or loses counterparties between reviews, or a review interval that changes, was not exercised.
state
five scalars per counterparty, written by src/exposure.step and never from a model reply: whether an alert has been raised for the current episode, the governing decision type/officer/date, and the watch status last reported. Rendered into two or three English sentences for the prompt.
44 of the 150 readings are memory-dependent, measured rather than declared -- s001 removes exactly this paragraph and nothing else. Removing it moves watch-status accuracy 100.0 -> 88.0 pct, alert accuracy 100.0 -> 74.0 pct, and memory-dependent watch-status accuracy 100.0 -> 61.36 pct. (r001-cpty-watch against s001-cpty-watch-stateless; data/gold.jsonl)
⚠︎ IT DOES NOT GROW, AND THAT IS THE POINT AT SCALE. Five fields cost the same on review 30 as on review 2. What it does NOT carry is provenance for the governing decision beyond a name and a date -- not which committee minute, not a document a credit officer could pull up. Named in guardrails.add_first rather than shipped.
a real deployment's carried-state file is one file replaced atomically -- correct for one writer, not a concurrency model. The state is also scoped to the RUN inside evals/run.py, never to disk, because an eval that carried state between runs could not be re-run or compared with its own control.
model
one provider, one key, one completion per reading, MAX_TOKENS = 8000, thinking never sent. Swapping it is PROVIDER / BASE_URL / API_KEY / MODEL in .env and the same run again.
150 readings, 211751 input tokens and 64728 output, of which 46851 (72.4 pct) was provider-side reasoning left at the default. p50 3174ms, p95 6437ms. On the stateless control, one reading (CPTY-0010-R1) hit the 8,000 ceiling and returned zero visible output; not re-fired at a higher ceiling. (r001-cpty-watch, results/eval-s001-cpty-watch-stateless.json)
⚠︎ THE CEILING WAS NOT SEPARATELY CALIBRATED. No c000 probe was fired before spending; MAX_TOKENS=8000 was set from reasoning about task shape, not measurement. It held for the scored run (largest reply 2,066 of 8,000) but not for the stateless control, where reasoning alone consumed the full budget on one reading.
every cost figure here is a projection of THIS model's token counts onto a published card, and 72-77 pct of the output is reasoning. A model that does not emit reasoning tokens, or does not bill them, lands nowhere near its row for the same answers. Nothing here measures whether another model would hold the same 100 pct officer accuracy.
labels
data/gold.jsonl -- one row per reading carrying the exposure, limit and growth inputs, the watch status, the governing decision and an alert call. Computed by tools/build_corpus.py from each counterparty's plan, never hand-authored, and re-derived from src/exposure.step by evals/check_labels.py before any run may spend.
150 rows, 0 mismatches on replay. The corpus carries 81 CLEAR, 25 WATCH, 17 APPROVED_CONTINUE, 7 RESTRICTED, 6 HOLD_NEW and 14 ESCALATED readings, so every watch status is exercised; evals/check_labels.py asserts memory-dependence, the NOT-STATED case, a decoy note and a decision reversal are all present. (data/gold.jsonl, data/corpus-stats.json, evals/check_labels.py)
⚠︎ THE COMMITTEE-REFERENCE MECHANIC IS A PROPERTY OF THIS GENERATOR, NOT OF CREDIT COMMITTEES IN GENERAL. Every case where a minute resolves or fails to resolve is a decision this kit's author made when writing tools/build_corpus.py, not a measurement of how real minutes are kept.
the labels are a property of a generator built for ANSWER-CLASS BALANCE, not for resemblance to any one real book's mix. The 90.0 pct officer accuracy the strongest free floor scores is a fact about nine invented committee references, not a general property of keyword floors.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
the same officer name reported on a counterparty every review, long after the committee minute it came from
the reader is treating the credit-limit-setting officer as a permanent stand-in for the decision-maker -- the keyword floor's specific, repeatable failure. It never resolves a committee reference and never says so
check decision_officer_accuracy_pct split against not_stated_recall_pct. The strongest free floor scores 90.0 pct and 0.0 pct respectively; the model scores 100.0 pct on both (results/eval-b002-cpty-watch-keywordmem.json against results/eval-r001-cpty-watch.json, CPTY-0043-R2)
an alert raised on a counterparty every month it stays active, rather than once
the carried state is not reaching the reading, or was rebuilt without it -- alert_raised is the only thing distinguishing the review that first sees an episode active from every review after it
compare alert_accuracy_pct against the stateless control. r001 scores 100.0 pct; s001 (memory removed) drops to 74.0 pct (results/eval-r001-cpty-watch.json against results/eval-s001-cpty-watch-stateless.json)
a reply with zero visible answer text and a finish_reason of length
provider-side reasoning consumed the entire token ceiling before any JSON was written -- a coverage gap, not a judgement error
check output_tokens against max_tokens on the failing reading. CPTY-0010-R1 in the stateless control spent exactly 8,000 of 8,000 (results/eval-s001-cpty-watch-stateless.json, failures[0])
a governing decision reported for a counterparty whose file is CLEAR this review
Rule X-7's reset was not applied -- an episode that closes must clear its standing decision along with its alert, not carry it forward into a quiet file
check the gold set for watch_status=CLEAR rows with a non-NONE governing decision type; there are none, because evals/check_labels.py's replay check would have failed the corpus build itself (data/gold.jsonl, src/exposure.step Rule X-7)
['Concurrency. data/state.json is replaced atomically in a deployment, which is correct for one writer.', "A missed review. Nothing here detects one, back-fills it, or marks the readings it produced as late -- and unlike this kit's monitor siblings, no cadence.py experiment measures what a missed review would cost.", 'A population that changes between reviews. All 50 counterparties are live at all three.', 'Whether disabling provider-side reasoning holds the accuracy. 72.4-77.1 pct of output was reasoning and thinking was never sent.', 'Repeats. One run per arm, so no band on any figure.', 'Whether English or JSON is the better rendering of the carried state. A design choice, not a measurement.', "Whether a real credit committee's minutes are typically resolvable in the same window the decision log lives in. This kit's whole officer-accuracy finding rests on its own invented mechanic.", "Per-reading input token counts. Only run-level totals were logged; the LLM lens's worked example uses a character-based estimate."]
The corpus licence, from the Data lens: MIT -- this repository's own licence. Nothing is fetched and no third-party grant is involved; every counterparty, position, officer, committee minute and dollar figure is invented here. Verified against the repository's own LICENSE file on 2026-08-24. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Watch status, decision type, decision officer, decision date and the alert call, per reading, exact match against the computed answer key
Track who approved a risky trading partner
PresenterOpens the private repo. Visible to admins only.
In one lineWatch status, decision type, decision officer, decision date and the alert call, per reading, exact match against the computed answer key
For each of the 150 readings and each of the five answered fields, did the reply equal the computed answer key? Watch status, decision type and alert are words from a closed list; decision officer and decision date are exact-string and exact-date comparisons.
$0.00per 1,000 counterparty review files
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function scores the model run, the stateless control and all three free floors, so every arm is comparable by construction.
The inputOne real row, seen by every grader
Document
CPTY-0043-R2
The file
Ethanol plant, exposure 2,628,477.97 USD against a 2,100,000.00 USD limit, growth undetermined (onboarded within the last year)
Credit Decision Log
2026-02-02 ESCALATE TO COMMITTEE per Q2 Credit Committee, minute #38 -- see Relationship Notes.
Relationship Notes
Minute #38 is pending transcription; full committee notes were not available at review time.
Watch status, decision type, decision officer, decision date and the alert call, per reading, exact match against the computed answer key
all five hit for the model; the free floor hits four and misses the decision officer
CPTY-0043-R2 is one of the 3 readings whose referenced committee minute never resolves. The key says NOT STATED and the model says NOT STATED. The strongest free floor answers 'Aisha Cho, VP, Credit Risk' -- the credit-limit-setting officer -- which is the guess this kit exists to catch: 100.0 pct against 0.0 pct NOT-STATED recall, on a field that names who is accountable for a decision.
The formulaWhat it computes
accuracy = hits / 150 per field. A reading whose reply did not parse counts as a MISS in every field, never as an exclusion -- which is why coverage is published beside every accuracy.
The analysisWhat it actually did
Model
Result
the fast tier, with the carried state
100.0% watch status accuracy · 4 more measured on this row
the same tier, memory removed (THE CONTROL)
88.0% watch status accuracy · 4 more measured on this row
the strongest free floor, no model
100.0% watch status accuracy · 4 more measured on this row
the same rule, GIVEN the carried state but not the window text
82.0% watch status accuracy · 4 more measured on this row
the rule a credit-risk spreadsheet already runs
70.7% watch status accuracy · 4 more measured on this row
In operationWhat to monitor
Reference standard: data/gold.jsonl, computed by tools/build_corpus.py over the planted decision-window events at generation time and re-derived from src/exposure.step by evals/check_labels.py before any run may spend. This grader IS the reference, so its own TPR and TNR are not measured -- scoring it against itself would be circular.
These rates are UNKNOWN, on purpose
This grader's own error rate is not measured and cannot be: it IS the reference. What can be wrong is the KEY, and one specific thing about it is a decision this kit's author made rather than a measurement: whether an unresolved committee minute should score as NOT STATED (the choice made here) or as a miss. Both the model and the strongest floor are scored against the same choice.
Watch these
watch_status_accuracy_pct
decision_type_accuracy_pct
decision_officer_accuracy_pct
decision_date_accuracy_pct
alert_accuracy_pct
unparsed_replies
Alarm on
unparsed_replies above 0. A reading that returns nothing is scored as a miss in all five fields, so a reliability failure arrives disguised as a quality failure. The scored run had 0 of 150; the stateless control had 1 of 150 at the identical ceiling, so this is a property of removing the carried-state paragraph, not of the ceiling alone.
How tight can the band be? No tuned threshold anywhere on the model path -- the reply decides. The denominators are printed beside every rate because two of them are small: 3 readings whose committee reference never resolves, and 44 memory-dependent readings, so one row moves those figures by a third and by ~2.3 points respectively.
Cadence: Re-run evals/check_labels.py (free) on any change to tools/build_corpus.py, src/exposure.py, src/segment.py, src/select.py or src/prompt.py. Re-run the scored eval AND its stateless control together (paid, 300 calls) on any change to src/prompt.py or src/state.py -- the headline is a difference, so one arm re-run alone is not comparable with the other's old figure. All three free floors are free and should be re-run on any change at all.
The decisionWhen to reach for it
Use it
The truth is known and three of the five answered fields are words from a closed list.
Do not use it
The truth is not known -- the normal state of a real credit book, where whether a given exposure figure should trigger escalation, restriction or a hold is a risk-appetite decision a committee makes, not a fact a generator can compute. That is why this corpus is generated rather than captured.
A living map of modern AI — kept current every morning