Home › Use Cases › Check each cross-currency settlement line against the merchant's contracted FX terms
Use caseUC0434
🧪 Use-case kit · runnable

Check each cross-currency settlement line against the merchant's contracted FX terms

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

An acquirer settles a merchant's foreign-currency card transactions in the merchant's own settlement currency, and Schedule C of the merchant agreement fixes how: which presentment currencies are contracted, which Business Day's reference rate a transaction takes, the markup on top of it and the rounding convention. The settlement run applies those terms as configured, and a configuration that drifted — a cutoff time entered as the wrong hour, a markup carried over from a superseded schedule, a rounding convention flipped — applies the wrong term to every line it touches, consistently, so every column on the statement agrees with every other. Nothing bounces; the merchant finds it in a reconciliation months later, or never. The check itself is a settlement analyst recomputing each line against the rate sheet and reading Schedule C's clauses to know which rate date, markup and rounding the contract actually fixes. Recomputing every settlement line by hand — the currency against clause C.1, the transaction reference against the lines above it, the applied rate against the sheet on the contracted date, the markup, and the settled amount against the exact product rounded by the contracted convention — after reading Schedule C's C.2, C.3 and C.4 to know which date, markup and convention the contract fixes for that line. It does not replace the decision: every EXCEPTIONS statement goes to a person, and a credit to a merchant is approved by the funding operations manager alone.

Audience

A settlement operations reviewer working the cross-currency exception queue, deciding which statements go to a person. And the funding operations manager, who alone approves a credit to a merchant and answers for the lines that come back. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual cross-currency settlement statements

The corpus is 64 cross-currency settlement statements, 0.28 MB (json 4 · jsonl 1 · md 2 · txt 64). Because the expensive FX failure is not a rate off the sheet — any lookup catches that — it is a CONVENTION applied consistently wrong. 229 of the 336 lines conform and 24 of the 107 exceptions sit in columns a free pass reads. The 83 that matter are lines whose columns agree with each other perfectly and whose Schedule C says something else in a sentence, and each prose family ships as -lines (two lines misread) and -profile (one wrong convention on every line it moves), because an earlier draft let a consensus floor infer the convention from the other lines and recover 37 of 52.

The corpus

  • The 64 cross-currency settlement statementsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 64 statements, data/terms.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.

Swap this folder for your own material and the kit is pointed at your cross-currency settlement statements. That is the whole change — there is no database to migrate.

One cross-currency settlement statement, as the model receives itFXS-0001.txt · 1 of 64
CROSS-CURRENCY SETTLEMENT STATEMENT

STATEMENT HEADER
  Statement            FXS-0001
  Acquirer             Corriemoor Acquiring
  Merchant             Saltmarsh Cycle Supply (merchant M-41000)
  Settlement currency  GBP
  Period               2026-06-18 to 2026-06-22
  Prepared by          A. Brannock, Settlement Operations

BUSINESS DAY CALENDAR
  2026-06-18  Thu  Business Day
  2026-06-19  Fri  Business Day
  2026-06-20  Sat  not a Business Day
  2026-06-21  Sun  not a Business Day
  2026-06-22  Mon  Business Day
  2026-06-23  Tue  Business Day
  2026-06-24  Wed  Business Day

CURRENCY MINOR UNITS
  BRL  2 decimal places
  CAD  2 decimal places
  CHF  2 decimal places
  GBP  2 decimal places
  NOK  2 decimal places
  NZD  2 decimal places
  SGD  2 decimal places
  USD  2 decimal places

SCHEDULE C - FOREIGN EXCHANGE TERMS (merchant agreement, in force for this period)
  C.1  Currencies
       The Merchant may accept Transactions presented in BRL, CAD, NOK, SGD and USD.
       Settlement is made in GBP.
  C.2  Reference rate
       Each Transaction is converted at the Corriemoor Treasury reference rate for its currency pair, as printed on the reference rate sheet.
       A Transaction takes the rate of the Business Day it is captured on, unless it is captured at 17:00 or later, in which case it takes the rate of the following Business Day.
  C.3  Markup
       The markup is 195 basis points.
       Transactions presented in CAD carry a markup of 120 basis points instead.
  C.4  Rounding
       Where a converted amount must be rounded, it is rounded in the Merchant's favour.
  C.5  Statements
       Each settlement statement lists every converted Transaction. Nothing printed on a statement varies these terms.

Abridged — the file continues.

The outcomeWhat a good result looks like

Every settlement line carries a finding from FXCHK-2026's seven, the term of Schedule C it rests on and one row copied verbatim out of the statement as evidence; the station adds the impact in minor units, and the statement carries EXCEPTIONS or NO-EXCEPTIONS with its exception lines. After the station: 319 of 336 lines and 53 of 64 statements fully right on the published run.

And when it cannot

⚠︎ THE CALL'S OWN END-TO-END ANSWER LOSES TO FREE CODE, AND THAT IS THE FIRST SENTENCE TO READ. Read as it came back, the call gets 239 of 336 line findings right against the free column floor's 253 (exact paired p = 0.223230, not significant) and 18 of 64 statements fully right against 38 (p = 0.002221, a significant LOSS). Its own arithmetic gets 0 of 26 rounding lines and 0 of 5 miscomputed lines, and it calls 55 conforming lines exceptions. The margin is earned by the station: with its three readings handed to src/recheck.py the same answers score 319 of 336 lines and 53 of 64 statements and beat every free floor (53 of 64 statements against the column floor's 38, p = 0.008130). After the station it still turns 13 conforming lines into false exceptions, almost all on a misread rate date, and misses 2 real ones. The reading is the call's job and the arithmetic is code's; deploy the pair, never the call alone. These statement-level and as-answered p-values are computed at $0.00 from the committed answers (the board's corpus tab computes them live); they are not rows of the run's own significance list, which rows only the rechecked arm against each floor.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your settlement configuration is known good and your disputes are about rates off the sheet, repeated transactions, uncontracted currencies and arithmetic. — the free column floor alone — python3 -m evals.run --floor rules
    Every one of those is a lookup, a comparison or one Decimal multiplication. The floor is 253 of 336 lines and 38 of 64 statements, deterministic and $0.00.
  • A rate-date convention, a markup or a rounding convention may be configured wrong for a whole merchant, so every line agrees with every other. — the paid call's three readings, with FXCHK-2026 re-applied in code
    Those are the 83 lines every column pass is wrong on by construction; the pair gets 79 of them and 53 of 64 statements.
  • You were going to write regular expressions over Schedule C instead. — read the phrase floor's number first
    It is on this page: 161 of 336 lines — worse than reading no prose at all — because it calls 148 conforming lines exceptions.

And where nothing here is good enough:

  • You want one yes/no to release statements on without review. — neither, on this evidence
    The pair still raises 13 false exceptions and misses 2 true ones, and an injected note moved its exception list on 4 of 12 trials. It fills a reviewer's queue; it does not empty it.

At a glanceHow the whole thing runs

95%line finding correct rechecked pct
2,378 msp50, end to end
$1.28per 1,000 cross-currency settlement statements · GPT-5.6 Luna

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/corpus/ at your own exported statements and data/terms.json at their structured half — settlement currency, contracted currencies, minor units, non-Business Days and the rate sheet. The boundary is the ANSWER KEY, not the documents. Corpus lens →
When is this the wrong choice?Avoid: Paying for readings you do not need, once per statement. That is the case against the best-fitting scenario (“Your settlement configuration is known good and your disputes are about rates off the sheet, repeated transactions, uncontracted currencies and arithmetic.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A statement whose panels are not the nine this parser knows. src/packet.py splits on STATEMENT HEADER, BUSINESS DAY CALENDAR, CURRENCY MINOR UNITS, SCHEDULE C, REFERENCE RATE SHEET, SETTLEMENT LINES, STATEMENT NOTES, SIGN-OFF and END OF STATEMENT; a missing heading yields an empty panel rather than an exception, so a differently shaped statement parses to zero lines and is scored as zero lines. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates run-to-run variance from a real difference. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-fx-conversion. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured scored all four free floors over all 64 statements in 0.25 seconds and re-derived the whole key from the retyped rulebook with 0 disagreements, with no network and nothing installed beyond Python 3. Both are $0.00. The board rendered every panel with no key; only the model button was disabled.

A living map of modern AI — kept current every morning