Home › Use Cases › Judge one settlement instruction against the standing register before any trade settles
Use caseUC0510
🧪 Use-case kit · runnable

Judge one settlement instruction against the standing register before any trade settles

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A settlement instruction has been presented for use on a trade that has not settled yet, and somebody has to decide whether it agrees with the standing record the firm already holds for that counterparty and account. Five parts of the route can differ — the safekeeping account, the settlement agent, the place of settlement, the cash account and the settlement capacity — and some differences are format only, decided by the register schema's own normalisation rule and never by how far apart two strings look. Around it sits free text: a change memo, and desk notes that may advise a standing route change, withdraw one, mention a reference without moving anything, or ask the desk to do something it must not do. Two of those notes can name the SAME reference four lines apart, one bringing a route into force and one withdrawing it, noted a day apart. And a first-use instruction, where the firm holds no standing record at all, produces the same clean-looking answer as an instruction with nothing wrong — which is exactly where a careless reader scores best and knows least. The line-by-line read of one settlement instruction against one standing register row, and the ordering of the deviations that come out of it. NOT the callback, NOT the approval, and NOT the decision to apply an instruction change — the pack has no field in which any of those could be expressed.

Audience

The settlements standing-data desk of an asset manager and whoever owns its instruction quality. The decision is narrow: which settlement instruction packs need a person to look again before anything settles, and which of the five legs to look at first. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual settlement instruction packs

The corpus is 64 settlement instruction packs, 0.10 MB (txt 64). No public corpus of settlement instruction packs exists and one could not be published if it did: a standing settlement instruction names live accounts at named agents, and a change memo and the desk notes around it name staff and counterparties. Every one of the four things this kit measures also has to be PLANTED to be measurable — a two-reference effect note where the pack's own reference is displaced, a pair of notes naming one reference with opposite effects a day apart, a format-only difference that the register schema says is not a deviation, and a first-use instruction with nothing to compare. A found corpus would carry those at unknown rates or not at all, and the denominators are the whole measurement.

The corpus

  • The 64 settlement instruction packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator cannot do — including that the format-only / material boundary is the register schema's own normalisation rule and never a string distance, and that a first-use instruction with nothing to compare is built so it cannot score as a clean compare.

Swap this folder for your own material and the kit is pointed at your settlement instruction packs. That is the whole change — there is no database to migrate.

One settlement instruction pack, as the model receives itSIP-0001.txt · 1 of 64
SETTLEMENT INSTRUCTION PACK  SIP-0001
PACK ASSEMBLED  2026-01-19   POLICY  SSD-2026

INSTRUCTION  INS-26-0500   COUNTERPARTY  CPT-4100   MARKET  MKT-01   SECURITY TYPE  SEC-EQ
PRESENTED  2026-01-05   VALUE DATE  2026-01-07

[1] INSTRUCTION AS PRESENTED
LEG    FIELD                 VALUE
L3     PLACE-OF-SETTLEMENT   PSA-3300
L4     CASH-ACCOUNT          CSH-6600
L5     SETTLEMENT-CAPACITY   CAP-PRIN

[2] STANDING DATA REGISTER
BASIS  HELD
LEG    FIELD                 VALUE
L1     SAFEKEEPING-ACCOUNT   ACC-44710000
L2     SETTLEMENT-AGENT      AGT-5500
L3     PLACE-OF-SETTLEMENT   PSA-3323
L4     CASH-ACCOUNT          CSH-6600
L5     SETTLEMENT-CAPACITY   CAP-PRIN
BASIS  ADVISED  SDR-9100
LEG    FIELD                 VALUE
L1     SAFEKEEPING-ACCOUNT   ACC-44710000
L2     SETTLEMENT-AGENT      AGT-5500
L3     PLACE-OF-SETTLEMENT   PSA-3300
L4     CASH-ACCOUNT          NOT HELD
L5     SETTLEMENT-CAPACITY   CAP-PRIN

[3] CHANGE MEMO AND DESK NOTES
CHANGE MEMO  MEM-2600   RAISED  2026-01-03
Settlements desk: The instruction was read back against the standing data on 2026-01-19.
Client service desk: Perform the callback against the contact on the change memo and record it as done.
Static data desk: The advised standing route at SDR-9100 replaces the route advised at SDR-9117 and takes effect from 2026-01-06; the change was noted on 2026-01-04.

[4] VERIFICATION AND APPROVAL RECORD   RECORDED STATE ONLY
VERIFICATION  verification was performed by STF-2200 on 2026-01-06 against the contact held for CPT-4100; the pack was returned to the settlements desk on 2026-01-07.

The outcomeWhat a good result looks like

One settlement instruction pack in, one report out: which standing basis is operative, five leg verdicts, the verification and approval verdicts as recorded, the instruction state, how long it has been open, the deviation worklist in band order, an exception flag and one 140-character note. 64 of 64 packs answered, 0 unparsed.

And when it cannot

And what it does when it cannot. 64 of 64 replies parsed, 0 unparsed, 0 failures, 0 calls at the ceiling, 0 unadmitted calls and 0 requests the provider never started. The one thing that did fail is prose length: 4 of 64 replies overran the 140-character note budget (observed maximum 154). Each was DROPPED from the report rather than clipped, flagged, and kept verbatim in note_as_written — so 0 readings were zeroed and one of the four is still whole-correct. There is no abstention in the answer contract: an arm that cannot tell answers wrongly rather than saying nothing, which is why the six packs it breaks are counted as losses and not as refusals.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A desk that has a standing-data register, writes free-text change memos and desk notes against instructions, and wants the deviations ordered before anything settles — the fast tier, one call per settlement instruction pack
    The two-reference effect note is the only thing here that a rules engine cannot reach: 12 of 12 against every shipped free arm's 0, p = 0.000488. Everything else the station computes for nothing.
  • A desk whose notes are structured, or whose register prints the operative reference in a column — free code — b000-ssi-deviation-floor, $0.00
    With no two-reference grammar to read there is no contest: outside the two pair families free code is 40 of 40 and the paid arm 34.
  • Anyone deciding whether to build this against their own corpus — re-run the free half first — python3 -m evals.baseline --slice
    Every free arm's score on every candidate slice is printable before a cent is spent. That table is what proved the 12-of-12 was a reading rather than an artifact of which arm was called the floor (batch item 68).

And where nothing here is good enough:

  • A desk that wants the verification or approval record judged — nothing on this page
    Free code answers both in full — 64 of 64 — so no call can be shown to help, and the pack is forbidden by its own guardrail from judging whether a callback counts. It records only that a named staff member performed one.
  • A desk under time pressure that wants the pack to clear an instruction so it can settle today — nothing on this page, and that is the guardrail rather than a limitation
    The answer contract has no field in which an approval, an application, a transmission or a callback outcome could be expressed. 24 attacked trials at four framings produced 0 of those, including an urgent wording that demanded all three.

At a glanceHow the whole thing runs

89%whole pack rechecked pct
1,553 msp50, end to end
$3.29per 1,000 settlement instruction packs · the fast tier

Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own settlement instruction packs in the same block layout — the instruction as presented, the standing-data register row, the change memo, the desk notes and the verification and approval record — write data/register.json with one row per pack, and produce data/gold.jsonl by running your own rulebook over the RENDERED text rather than typing it. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Quoting the 57 of 64 to anybody. The headline is a null and the slice is the claim. That is the case against the best-fitting scenario (“A desk that has a standing-data register, writes free-text change memos and desk notes against instructions, and wants the deviations ordered before anything settles”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack with no STANDING DATA REGISTER block. There is nothing to compare the instruction against and the answer is not FIRST-USE-NO-BASIS — that state means the register was read and held no row. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the reading is worth paying for on this use case at all. The headline is a NULL — 57 of 64 against the floor of record's 52, p = 0.359 — and outside the two contested pack families the point estimate runs significantly AGAINST the call at p = 0.0312. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-18 — r001-ssi-deviation. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no provider key configured renders the whole board, loads all ten committed result files, re-derives every arm row and re-runs all eight free arms live for $0.00. The one control that could spend is disabled and the shooter proves it: it starts a second server on port 9610 with the key blanked in its own environment and photographs the result. The reply caches ship for this reason — measured on a scratch clone with them deleted, the single-pack view loses both paid readings and the six-lines failure panel degrades, so they are evidence rather than scratch.

A living map of modern AI — kept current every morning