The business caseThe problem this solves
A trade exception lands on a wealth manager's trade support desk with ten panels behind it: the block booking and its allocation rows, what the counterparty sent back and the settlement instructions held for it by effective date, position and cash by account with movements dated around the settlement date, the matching system's status codes, the events on the trade, prior closed exceptions with the same counterparty, the desk's own cutoff settings, what is attached, and the client service notes people typed. Somebody has to say which root cause it is under the manager's own card, where it sits against the desk's cutoff, and whose queue it goes to - or say that the card cannot place it. Nothing, on this corpus. The analyst's read of ten panels per exception is done better by code written from the same three cards: 49 of 64 exact rows for $0.00 against the paid call's 34.
Audience
A head of trade support deciding whether a model belongs in front of the exception queue. The answer this report gives is no, on this corpus, and it is measured rather than argued. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual trade exceptions
The corpus is 64 trade exceptions, 0.21 MB (json 3 · jsonl 1 · md 1 · txt 64). A real trade exception cannot be published: it carries a manager, a counterparty, accounts and positions. So this one is generated - and generated to defeat a reader that trusts the loudest sentence. Every one of the 12 distinctive note sentences decides the cause on one exception and decides nothing on another; 29 exceptions carry an instruction note timed BEFORE the block was booked; 24 carry an ex-day note outside the settlement window; 7 carry a net amount that differs only inside the tolerance; 6 carry a reference one keystroke from a real allocation; and 4 ask outright for the approval or booking this pack never performs.
The corpus
- The 64 trade exceptionsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your trade exceptions. That is the whole change — there is no database to migrate.
TRADE EXCEPTION TXE-0001
================================================================================================
PANEL 1 — EXCEPTION HEADER
Manager Quillmere Wealth Partners (synthetic manager)
Exception raised at 2026-06-19 06:46
Raised because pending item carried from the previous check
Trade reference as written T-260617-0579-A3
Counterparty CP-2041
Security SID-20938 (listed equity)
Market MKT-BR
Worked by Client service associate
PANEL 2 — BLOCK BOOKING EXTRACT (this manager's own blotter)
Block booked at 2026-06-17 13:25
Allocation ref Account Side Quantity Price Net amount Trade date Settle date Counterparty instruction
T-260617-0579-A1 AC-29730 SELL 12,900 72.44 934,476.00 2026-06-17 2026-06-19 D2/32137-04
T-260617-0579-A2 AC-41164 SELL 24,900 72.44 1,803,756.00 2026-06-17 2026-06-19 D2/32137-04
T-260617-0579-A3 AC-12014 SELL 7,100 72.44 514,324.00 2026-06-17 2026-06-19 D2/32137-04
PANEL 3 — COUNTERPARTY SIDE
Confirmation received at (none received)
Instructions held for CP-2041 on this desk's reference data:
effective from 2025-08-21 D5/73269-01
effective from 2026-04-19 D2/32137-04
PANEL 4 — POSITION AND CASH BY ACCOUNT (settled figures and scheduled movements)
AC-29730 settled position 1,200
AC-29730 cash balance 101,207.00
AC-41164 settled position 6,500
AC-41164 due 2026-06-18 deliver 2,200
AC-41164 cash balance 161,328.00
AC-41164 due 2026-06-22 cash in 75,474.00Abridged — the file continues.
The outcomeWhat a good result looks like
One row per exception a supervisor could act on as written: the trade reference exactly as an allocation row writes it, the root cause under TX-2026 taken in the card's own order, the reason it cannot be placed when it cannot, the position against this desk's own cutoff settings under TS-2026, the queue under TR-2026, and the one line of the exception that establishes the cause.
And when it cannot
It leaves the exception UNCLASSIFIED with its reason and sends it to the operations senior review queue. The key rules 16 of 64 exceptions unplaceable - 6 whose written reference is not an allocation row, 5 with nothing from the counterparty and nothing attached, 5 where nothing that passes its gate states a cause. Leaving one unclassified is the right answer, never a failure to answer.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- the causes, their ORDER, the note gates and the desk's cutoff settings are written cards and you have the cards — the domain-written floor - pure code, evals/baseline.py
on this corpus it takes 49 of 64 exact rows for $0.00 and the paid call takes 34 - the call is behind by 15 and McNemar says the loss is real (p = 0.0015). Writing the cards down as code is the job. - you need the exception LEFT unclassified rather than bucketed — pure code, and a supervisor
the domain floor forced 0 exceptions into a cause; the paid call forced 8 in its own row and 2 after the station (TXE-0020, TXE-0038) - and those are readings the station cannot undo. - you want one number for a slide — the exact row of five, raw and rechecked, with the floor and the p-value beside it
it is the metric a constant does NOT nearly win: 21 of 64 against the floor's 49. Quote the run-to-run variation with it: a second draw of the same prompt moved 11 exact rows. - you compare a paid arm with a free one on any kit with a pure-code settlement step — both columns - raw against raw, rechecked against rechecked
the station moves the constant from 0 to 21 and the paid call from 11 to 34 on this kit.
And where nothing here is good enough:
- the cause is stated only in a client service note and no word list reaches it — neither, yet - measure it on your own notes first
this is the one family where the call is ahead of the floor of record: prose 4 of 16 paid against 3 free. But a constant reply takes 6 of 16 prose rows through the station - more than the call - so the figure cannot be quoted, and on every other family the call gives back more than it gains (plus 1 on prose, minus 16 elsewhere).
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own exceptions in the same ten-panel shape and data/gold.jsonl with the row your desk would assign; then rewrite src/taxonomy.py's causes, reasons, status codes and note gates and src/policy.py's cutoff settings and queues to match your own cards. THE MEASURED RESULT DOES NOT TRAVEL WITH THE DATA. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call before the floor exists. Without it every paid figure reads as good. That is the case against the best-fitting scenario (“the causes, their ORDER, the note gates and the desk's cutoff settings are written cards and you have the cards”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A WRITTEN REFERENCE THAT IS NOT AN ALLOCATION ROW. The match is exact by design, so a reference one keystroke from a real allocation separates the exception rather than resolving it - 6 of the 64 are built that way, and that is the intended behaviour, not a tolerance to be widened. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether free code would still win on a real desk's exceptions. On this corpus it did: the domain floor took 49 of 64 exact rows to the call's 34, p = 0.0015. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, off-peak tariff - run of record. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 2 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r002-trade-exception. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — with NO key configured the board served every panel on 127.0.0.1:9579 from the committed run of record - driven on 2026-09-16 across all five floors and all 64 exception views, with the live control refusing - and the five free arms and the constant sweep were scored by pure code with no provider reached, for $0.00. requirements.txt names no package.








