The business caseThe problem this solves
A trade finance presentation pack lands on a checker's desk with twelve panels behind it: the presentation and its document set, the invoice lines with goods codes and prices, the transport documents and the route, the parties and the notify chain, the bank's own restricted-list extracts for parties and vessels, the dual-use goods list, the benchmark price bands, the licences and their validity dates, the payment directions, the prior presentations on the same corridor, this desk's own escalation window settings, and the timed checker notes people typed. Somebody has to say which red flag it carries under the bank's own card, where the pack sits against the desk's escalation window, and whose review queue it goes to - or say that the card places no flag on it. Nothing, on this corpus. The checker's read of twelve panels per pack is done better by code written from the same three cards: 51 of 64 exact rows for $0.00 against the paid call's 36.
Audience
A head of trade finance compliance deciding whether a model belongs in front of the screening queue. The answer this report gives is no, on this corpus, and it is measured rather than argued. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual trade finance presentation packs
The corpus is 64 trade finance presentation packs, 0.30 MB (json 3 · jsonl 1 · md 1 · txt 64). A real trade finance presentation pack cannot be published: it carries a bank, an applicant, a beneficiary, a carrier and a price. So this one is generated - and generated to defeat a reader that trusts the loudest sentence. Every one of the 12 distinctive note sentences decides the red flag on one pack and decides nothing on another; 64 packs carry a dual-use list entry for a goods code this presentation does not carry; 26 carry a list entry one character from a party on the pack and 12 one character from the carrying vessel; 14 carry an authorisation day on or after the shipment date; 8 carry a price below the band that the declared volume terms legitimately reach; 6 carry a reference one keystroke from a register row; and 4 ask outright for the clearance, block, release or report this pack never performs.
The corpus
- The 64 trade finance presentation packsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your trade finance presentation packs. That is the whole change — there is no database to migrate.
TRADE FINANCE PRESENTATION PACK TFP-0001
========================================================================================================
PANEL 1 — PRESENTATION HEADER
Bank Calderhaven Trade Bank (synthetic bank)
Pack raised at 2026-06-18 04:30
Raised because not examination-complete at the scheduled check
Presentation reference as written TP-260612-0699-D1
Applicant PTY-95065
Beneficiary PTY-46486
Corridor CORR-BR
Product documentary credit
Examined by Documentary examination analyst
PANEL 2 — PRESENTATION REGISTER EXTRACT (this bank's own register)
Register entry made at 2026-06-12 09:52
Presentation ref Notify Invoice value Invoice date Payment due
TP-260612-0699-D1 PTY-85597 5,365.00 2026-06-12 2026-06-18
TP-260612-0699-D2 PTY-94472 1,260,987.90 2026-06-12 2026-06-18
TP-260612-0699-D3 PTY-47921 420,661.70 2026-06-12 2026-06-18
PANEL 3 — COMMERCIAL INVOICE LINES (one line per drawing)
Presentation ref Goods Description Control Quantity Unit Unit price Line total
TP-260612-0699-D1 GD-1920 polymer resin NONE 1,450 kg 3.70 5,365.00
TP-260612-0699-D2 GD-3911 frequency converters NONE 7,170 unit 175.87 1,260,987.90
TP-260612-0699-D3 GD-6163 ceramic substrates NONE 5,690 unit 73.93 420,661.70
PANEL 4 — TRANSPORT DOCUMENT (the carrier's side)Abridged — the file continues.
The outcomeWhat a good result looks like
One row per pack a reviewer could act on as written: the presentation reference exactly as the register extract writes it, the red flag under TF-2026 taken in the card's own order, the reason when no flag fits, the position against this desk's own escalation window under TW-2026, the queue under TQ-2026, and the one line of the pack that establishes the flag.
And when it cannot
It leaves the pack UNCLASSIFIED with its reason and sends it to the TBML senior review queue. The key rules 16 of 64 packs unplaceable - 6 whose written reference is not a register row, 5 with nothing from the carrier and nothing attached, 5 where nothing that passes its gate states a red flag. Leaving one unclassified is the right answer, never a failure to answer, and NO-FLAG-ON-THE-CARD-FITS is not a clearance.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- the red flags, their ORDER, the note gates and the desk's escalation windows are written cards and you have the cards — the domain-written floor - pure code, evals/baseline.py
on this corpus it takes 51 of 64 exact rows for $0.00 and the paid call takes 36 - the call is behind by 15 and McNemar says the loss is real (p = 0.0026). Writing the cards down as code is the job. - you need the pack LEFT unclassified rather than bucketed — pure code, and a named reviewer
every written free floor forced 0 packs into a red flag; the paid call forced 4 in its own row and 3 after the station - and those are readings the station cannot undo. - you want one number for a slide — the exact row of five, raw and rechecked, WITH the per-cell free-arm scores and the 13-pack headroom count beside it
the row is a composite and two of its five cells are 64/64 for every free arm including the constant. Quoted alone it implies twenty points of contest that do not exist: the room above the floor is 13 packs. - you compare a paid arm with a free one on any kit with a pure-code settlement step — both columns - raw against raw, rechecked against rechecked
the station moves the constant from 1 to 19 and the paid call from 22 to 36 on this kit, on 11 and 1 overrides respectively. - you want to know whether the money bought the READING — the 13-pack headroom slice, scored against EVERY free arm
the 13 packs the floor of record misses, scored against EVERY free arm: the floor of record 0/13, panels-only 0/13, THE PAID CALL 4/13, gates-off 9/13, the tuned regex 13/13 (p = 0.0039 against the paid call), the constant 3/13. Against the floor alone it reads 4-to-0; against every free arm the money is worth NINE PACKS LESS than the best free arm.
And where nothing here is good enough:
- the red flag is stated only in a timed checker note and no word list reaches it — neither, yet - measure it on your own notes first
prose is the one family where the call is ahead of the floor of record: 6 of 16 paid against 4 free. But that is 16 packs of a 64-pack corpus, and on the 13 packs the floor actually misses the call takes 4 where a free gates-off regex takes 9 and the tuned regex takes all 13 (p = 0.0039 AGAINST the call).
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own presentation packs in the same twelve-panel shape and data/gold.jsonl with the row your desk would assign; then rewrite src/taxonomy.py's red flags, reasons, system codes and note gates and src/policy.py's escalation windows and queues to match your own cards. THE MEASURED RESULT DOES NOT TRAVEL WITH THE DATA. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call before the floor exists. Without it every paid figure reads as good. That is the case against the best-fitting scenario (“the red flags, their ORDER, the note gates and the desk's escalation windows are written cards and you have the cards”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A WRITTEN REFERENCE THAT IS NOT A REGISTER ROW. The match is exact by design, so a reference one keystroke from a real register row separates the pack rather than resolving it - 6 of the 64 are built that way, and that is the intended behaviour, not a tolerance to be widened. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | HOW MUCH OF THE HEADLINE WAS EVER CONTESTED. The exact row of five is a COMPOSITE and two of its cells - presentation_ref and escalation_position - are 64/64 for EVERY free arm including the constant, while a constant EMPTY reason takes 48 of 64. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, off-peak tariff - run of record, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-tbml-screen. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — with NO key configured the board served every panel on 127.0.0.1:9591 from the committed run of record - driven on 2026-09-17 across all six arms and every pack view, with the live control disabled and saying why - and the five free arms and the constant sweep were scored by pure code with no provider reached, for $0.00. requirements.txt names no package: an AST walk over src/, evals/ and tools/ found 18 modules imported and none outside the standard library.












