Home › Use Cases › Classify why an account transfer came back rejected, and what cures it
Use caseUC0322
🧪 Use-case kit · runnable

Classify why an account transfer came back rejected, and what cures it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An account transfer between two brokerage firms was submitted and came back REJECTED. Every morning somebody works the return queue: one rejection record at a time, deciding what actually went wrong and what happens next. The record is a stack of panels - the registration as instructed against the registration the delivering firm holds, an authorization panel with a signature count and an age, the instructed positions, and a return message - and it always arrives with a sentence of the delivering firm's own boilerplate on top. On a large share of returns that sentence says the registration does not agree, whatever the defect turned out to be, and coding the return to the sentence sends a corrected title field back out against a defect that is still there. The first pass over a morning's return queue: which of eight reasons each rejection is, the line that establishes it, and which desk it goes to. It replaces nothing after that - no instruction is corrected, resent, cancelled or amended by anything here.

Audience

The operations analyst who works the return queue, and the supervisor who has to explain six months later why an account sat for three cycles. The decision is a coded row: the reason, the line that establishes it, whose desk the next step is on, and whether it can be corrected and resent today. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual transfer rejection records

The corpus is 62 transfer rejection records, 0.11 MB (txt 62). A transfer return queue is the right shape for this question and the wrong shape to obtain: a real rejection record is a joined extract of another firm's account master, a clearing venue's return message and an operations queue, and all three carry an account holder's name, an account number and a tax identification. So it is built from STRUCTURES - a registration on each side, an authorization panel, a positions list, a return message and a notes panel - and the key is TRC-2026 applied to those same structures by src/policy.decide(). Nothing is typed. There is no personal data anywhere and the label gate sweeps nine patterns to say so: 0 hits.

The corpus

  • The 62 transfer rejection recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your transfer rejection records. That is the whole change — there is no database to migrate.

One transfer rejection record, as the model receives itXFR-0001.txt · 1 of 62
TRANSFER REJECTION RECORD                                 XFR-0001
Prepared 2026-09-08 under TRC-2026 | transfer cycle 2026-07-01 to 2026-09-30

RETURN FACTS
  Record type                             full transfer
  Instruction submitted                   2026-08-09
  Returned by the delivering firm         2026-08-11
  Delivering firm code                    DF-6193
  Client file reference                   CF-2001
  Account located at the delivering firm  yes - open

REGISTRATION AS INSTRUCTED
  Registration type as instructed         CONSERVATORSHIP
  Parties as instructed                   1
  Tax identification as instructed        ending 6748

REGISTRATION ON THE DELIVERING FIRM'S RECORD
  Registration type on record             CONSERVATORSHIP
  Parties on record                       1
  Tax identification on record            ending 6748

AUTHORIZATION
  Signatures received                     1 of 1 required
  Authorization dated                     2026-06-15
  Authorization age at submission         55 days (the card allows 90)
  Supporting document on the instruction  a court appointment of the conservator dated 2026-07-23

POSITIONS INSTRUCTED
  1  listed equity                        position ref P-1018
  2  corporate bond                       position ref P-1025
  3  exchange-traded fund                 position ref P-1032

REJECTION AS RETURNED
  Reject code                             R-29
  Delivering firm narrative               The account is not available for delivery at this time.

OPERATIONS NOTES
  Desk reminder from the transfer procedures memo: a registration mismatch is corrected on the instruction and resent the same day where a current authorization is on file.
  An earlier transfer on this account is still open at the delivering firm.

Abridged — the file continues.

The outcomeWhat a good result looks like

One return in, four graded answers out: which of eight reasons, one line of the record quoted verbatim as the evidence for it, the cure route, and a same-day-resubmit flag. The reason also selects the CURE - a sentence from a table of eight, never a sentence a model wrote.

And when it cannot

And what it does when it cannot. UNCLEAR-NEEDS-REVIEW is a real answer and it means the return record does not carry the lines the card is applied to; the return goes back to whoever assembled the file rather than being coded on a guess. 4 of the 62 returns are that case and the paid call answered all 4 correctly, coding none of them anyway. ⚠︎ It also fails in a way the page names: on ALL 18 returns whose client file holds a current authorization it answered the same-day flag false, because the register is deliberately not in the prompt.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the cure route and the same-day flag and nothing else — the free rules table plus the engine - evals/baseline.py mode rules
    52 of 62 on the route and 61 of 62 on the flag for $0.00, because both are lookups on a reason and the table's reason is right 52 times out of 62.
  • You want the reason itself, which is the thing an analyst acts on — the paid call, rechecked
    61 of 62 against the table's 52, and paired over the same returns it is right where the table is wrong on 10 and wrong where the table is right on 1 - McNemar's exact two-sided p = 0.0117.
  • Your returns all arrive as a structured clearing message — the free rules table, and measure before you buy anything
    every one of the table's 10 misses is a return where the deciding fact is in prose. On a corpus with none, the table and the call would be much closer.
  • You want to know whether reasoning should be on — OFF, and then measure it yourself
    this kit ran 75 calls with the provider's disable shape sent and no reasoning tokens reported, at $0.000513 per return and a largest reply of 309 tokens.
  • You want the citation to be the audit trail — the paid call
    56 of 58 credited where the key names a line, mean coverage 0.97, and 0 unlocatable quotes in 58 citations returned - nothing was invented.

At a glanceHow the whole thing runs

97%record all correct pct
1,656 msp50, end to end
$1.64per 1,000 transfer rejection records · openai/gpt-5-6-luna

Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/policy.md and data/policy.json with your own coding card - the same rules in both, stated in the same order. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying for it. And avoid reading the table's flag score as competence - the MODAL arm, which reads nothing at all, gets 54 of 62 on the same field. That is the case against the best-fitting scenario (“You want the cure route and the same-day flag and nothing else”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A return that is not seven fixed panels. The free floor's readers, the label gate's structural checks and the citation locator are all bound to literal headings; a fixed-width clearing message or a PDF needs a different front end. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Reasoning ON against OFF on this corpus. Every call this kit made sent the provider's disable shape; the comparison is a second scored run and it was not bought. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, as answered, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-08 — r001-transfer-reject. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on 2026-09-08 from a clean checkout with no key configured: tools/build_corpus.py rebuilt all 62 returns byte-identically under two PYTHONHASHSEED values, evals/check_labels.py re-derived the key with 0 disagreements, both free floors scored the whole corpus, and src/app.py served every panel of the board. Nothing was installed, nothing was fetched and nothing was spent.

A living map of modern AI — kept current every morning