Home › Use Cases › Suspense Account Monitoring and Clearing-Entry Proposal
Use caseUC0076
🧪 Use-case kit · runnable

Suspense Account Monitoring and Clearing-Entry Proposal

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A suspense-clearing policy is not a set of thresholds. It is a handful of sentences about WHAT HAS ALREADY HAPPENED TO THIS ITEM -- a proposal already with the approver is not re-proposed, a proposal the approver returned is not re-sent on the same evidence, a proposal whose evidence has since collapsed is pulled back, and evidence that is partial for the second cycle running goes to the reconciliation controller. Two items with identical balances, identical evidence registers and identical ages get different verdicts, because one of them has a clearing entry sitting on somebody's desk and the other does not. Nothing in the review pack you are holding tells you which one it is, so somebody goes back through the previous worksheets and the approval queue by hand, every item, every cycle. Somebody opening a suspense item's review pack, then going back through the last two cycles' worksheets and the approval queue to work out whether a clearing entry is already sitting with an approver, whether the approver sent the last one back, and whether the evidence was partial last cycle as well.

Audience

Wealth and custody operations teams who work a suspense or unallocated-cash ledger on a cycle, the reconciliation controllers those items escalate to, and anyone deciding whether a language model has any business near an account that money moves out of. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual review packs

The corpus is 150 review packs, 0.61 MB (jsonl 2 · txt 150). A real suspense-ageing pack names a client, a counterparty, an account number, an unmatched cash amount and -- on every page -- the operations analyst who owns the item and the person who signs its clearing entry off. That is exactly the material that never leaves a custodian. There is no public corpus of (suspense item, evidence register, adjudicated clearing action) for the same reason there is no public corpus of production reconciliation breaks. Generating it also bought the one thing a captured corpus cannot give: the answer key is src/policy.replay's output over the planted review patterns, so a rule set with this much precedence in it cannot carry its author's misreading into the score.

The corpus

  • The 150 review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your review packs. That is the whole change — there is no database to migrate.

One review pack, as the model receives itSUS-0001-R1.txt · 1 of 150
Synthetic Record
----------------------------------------------------------------
This file is synthetic. It was generated for an open evaluation kit and describes
no real institution, account, client or person. Item SUS-0001, review R1.

Suspense Item
----------------------------------------------------------------
Item reference     : SUS-0001
Account class      : Settlement suspense
Suspense account   : 7831-SUS-SET
Currency           : AUD
Balance            : 373,980.04 AUD (debit balance)
Source system      : Custody position keeper
Opened on          : 2026-04-19

Clearing Policy
----------------------------------------------------------------
Attention age for this account class : 45 days

Evidence strength is read from the Evidence Register below:
  FULL     a line whose match strength is 95.0 % or more AND whose counterparty reference and
           amount are BOTH shown as matched.
  PARTIAL  no line meets the full-match test, but at least one line reaches 60.0 %.
  NONE     no line reaches 60.0 %.
The clearing route is the category of the register's strongest line by match strength. Where the
evidence is NONE the route is UNRESOLVED.

Clearing actions, tested IN THIS ORDER. The first rule that applies decides the action.
  1. WITHDRAWAL. If a clearing entry proposed for this item is still with the approver and the
     evidence at this review is no longer FULL, withdraw the proposal.     -> WITHDRAW_PROPOSAL
  2. WITH THE APPROVER. If a clearing entry proposed for this item is still with the approver,
     take no further action on it at this review.                          -> AWAIT_APPROVAL
  3. RETURNED. If the approver returned the last clearing entry proposed for this item and the

Abridged — the file continues.

The outcomeWhat a good result looks like

Per review: the action the policy requires, the clearing route the evidence supports, and the register line that is strongest -- plus a four-field carried state the next review is judged against, written by code from the parsed evidence and never from the model's reply, and a proposal record whose approver slot is empty and whose posted flag is false.

And when it cannot

The published run has no wrong answers to describe: 150 of 150 reviews, 450 of 450 cells. What a wrong answer would COST is not symmetric and the kit never averages the two directions. Under-acting leaves an aged item unescalated and a collapsed proposal standing in front of an approver; over-acting sends a controller an item that is quietly progressing, or re-proposes a clearing entry the approver has already declined. The one direction measured at zero on every arm is over-action on the 37 quiet reviews -- 0, including on the free floor, which cannot raise one.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A clearing policy whose rules are all about the item in front of you -- an evidence threshold, an age limit, a match test — the free floor, at $0.00
    The floor is 100 pct on all four of this policy's memory-free rules (104 of 104 reviews), names the route on 150 of 150 and the strongest evidence line on 150 of 150 -- exactly what the model scores. With the state removed the model returns the same answers on 448 of 450 cells and is wrong on one the floor gets right.
  • Any rule about what happened at an EARLIER review -- a proposal already out, a proposal returned, a second consecutive partial — the fast tier WITH the carried state
    46 of 46, and even across all four such rules. Both the stateless arm and the free floor score 0 on every one of them, and they do it confidently rather than by abstaining.
  • Deciding whether a clearing entry should actually be POSTED — a person -- there is no other option, and the kit has no path to one
    Posting is outside the cap by construction: src/proposal.py has no post(), no journal handle and no approval setter, evals/check_labels.py fails the build if one appears under src/, and approved_by is written as None with nothing anywhere able to write a value into it. The pack's whole output is a proposal record and an evidence record.

And where nothing here is good enough:

  • An item that is BOTH past its attention age and carrying a live proposal — nothing here yet -- measure it first
    Rules 1, 2 and 3 sit above rule 5 in the written order, and NO ITEM IN THIS CORPUS EXERCISES THAT PRECEDENCE. It is the ordinary shape of a genuinely aged item and this run says nothing about it.
  • A policy a careful reader could read two ways — nothing here -- expect it to be worse, and price the ambiguity before you ship
    This policy was written not to be arguable, which is the single biggest reason the headline is 100 pct. A sibling monitor kit on an arguable rule lost six of its answers and had to publish two figures for the same run.
  • An item history longer than three reviews — nothing here yet -- measure it first
    Every rule in this policy resolves inside three cycles. A fourth consecutive partial and a twice-returned proposal are both entirely ordinary and neither is in this corpus.

At a glanceHow the whole thing runs

100%action accuracy pct
4,307 msp50, end to end
$2.35per 1,000 review packs · Google Gemini 3 Flash

Run once, for real, on 2026-08-23. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace src/policy.py FIRST. Every figure on this page is a property of a three-review corpus with an unambiguous policy, an unambiguous evidence test and a stated attention age, measured once. Corpus lens →
When is this the wrong choice?Avoid: Paying for a model to do table-sorting and four comparisons a regex already does perfectly. That is the case against the best-fitting scenario (“A clearing policy whose rules are all about the item in front of you -- an evidence threshold, an age limit, a match test”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?Three reviews per item. Every rule in this policy resolves within three cycles, so nothing here measures a rule that needs four -- a fourth consecutive partial, a proposal returned twice, an item that ages past two thresholds. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER 100 PCT REPEATS. One scored run, one control, one perturbation arm, each fired once. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, with the carried state, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-23 — r001-suspense-age. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A fresh clone with NO key configured reproduces the corpus byte-identically (python3 tools/build_corpus.py), passes all 21 pre-flight assertions including both red-proofs (python3 -m evals.check_labels), scores the free floor (python3 -m evals.run --run-id b000-suspense-age-packonly --baseline) and proves the wiring (python3 -m evals.run --run-id t000-suspense-age-stub --stub). Four commands, plain python3, no install step -- requirements.txt names nothing because the kit imports nothing outside the standard library. What it cannot do without a key is re-run the scored eval, its stateless control or the state-perturbation arm.

A living map of modern AI — kept current every morning