Home › Use Cases › Match a network's cardholder alert to the transaction row and say what the merchant owes
Use caseUC0210
🧪 Use-case kit · runnable

Match a network's cardholder alert to the transaction row and say what the merchant owes

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A prevention alert lands -- a network email, a portal export, a row off a CSV feed, a note typed off a phone call -- and somebody on a dispute desk reconciles it before a refund is issued or a chargeback is accepted: which row on the merchant's own ledger it is, whether a refund already covers it, whether the same transaction already carries an alert somebody worked, and whether the network's response window is still open. Almost none of that is a reading problem. It is a join and a subtraction, in front of prose that writes the date as 'the 12th of July' and relays the merchant descriptor truncated to a field width. The desk's manual reconciliation of each prevention alert against the merchant's ledger -- the blocking, the amount tolerance, the refund coverage sum and the response-window clock -- before a person confirms the record.

Audience

The dispute desk that confirms the record, and the finance owner who answers for the chargeback ratio when an alert is worked wrongly or not at all. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual prevention alerts

The corpus is 60 prevention alerts, 0.03 MB (txt 60). A defect mix a dispute desk would recognise and no public dataset can provide: descriptors the network truncated, abbreviated or stripped of punctuation over two rows blocking cannot separate; a window that closes on the next calendar date; a refund posted between the alert arriving and the desk working it; a decimal slip that is a disagreement rather than an absence; a currency the merchant does not settle in; a partial dispute stated in a sentence rather than a field; a transaction that already carries an alert somebody worked. Two thirds of the alerts are planted because the clean case teaches nothing about the arithmetic -- and a third are clean because a corpus that is all traps measures a different job from the one the desk has.

The corpus

  • The 60 prevention alertsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your prevention alerts. That is the whole change — there is no database to migrate.

One prevention alert, as the model receives itAL-0001.txt · 1 of 60
From: alerts@sentinel-network.example
To: disputes@bellrock-supply.example
Subject: [SENTINEL] Prevention alert AL-0001 - response required

PREVENTION ALERT
Alert ID:             AL-0001
Issued (UTC):         2026-06-24T06:53:00Z
Card:                 **** **** **** 9730
Cardholder:           HANA THORNE
Transaction amount:   USD 1184.00
Disputed amount:      USD 1184.00
Transaction date:     2026-06-21
Merchant descriptor:  BELLROCK SUPPLY #0512 SPOKANE WA
Reason given:         Cardholder does not recognise the charge

Resolve or decline within the network response window. An alert left unworked when the
window closes converts to a chargeback and cannot be recalled.

The outcomeWhat a good result looks like

A worked alert record a person confirms instead of an alert a person reconciles: the ledger row, the disposition, the refund to the cent, whether the desk beat the window, a verify flag when the issuer's name is a different surname, and one sentence naming the rule.

And when it cannot

A false close. An alert that should have moved money, closed as already refunded, not ours or a duplicate, is a chargeback the merchant pays with a fee on top and a count that moves toward the scheme's monitoring thresholds. On this corpus the RAW model ships 0 of those in 33 opportunities and the FREE FLOOR ships 3 -- $1,915.00 of refunds it sends to a queue rather than pays -- which is the reverse of the direction a page selling a model would expect.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A feed that is already structured -- a CSV export or a portal API with named fields, on a merchant with one store — the free rules floor alone
    It scored 95.0 pct for $0.00 on this corpus and read every field of every alert correctly. Its only losses are alerts where two ledger rows cannot be separated without reading a mangled descriptor -- a problem a single-store merchant does not have.
  • Many stores or terminals sharing a descriptor prefix, with a network that truncates or abbreviates it — the model, rechecked
    3 ROWS of 60 -- 60 of 60 rechecked against the free floor's 57 of 60. All three are the descriptor tie-break: 7 of 7 two-candidate matches against the floor's 4 of 7, worth $1,915.00 of refunds on 60 alerts that the floor sends to a queue. The recheck keeps the reading and re-derives everything else, so the model's own arithmetic mistakes never reach the record -- and it is run on BOTH arms, which is what makes the 3 honest.
  • Free-text alerts with no field structure at all -- transcribed calls, faxes, screenshots typed out by hand — the model, rechecked
    The floor's call-log reading holds only because these logs were generated to a template; a real transcript has no labels to regex and the reading is the whole job.

At a glanceHow the whole thing runs

88%record all correct pct
14,676 msp50, end to end
$8.98per 1,000 prevention alerts · Google Gemini 3 Flash

Run once, for real, on 2026-08-30. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own alerts as .txt files into data/corpus/ and add one desk-record row per alert in data/alerts.json -- the network, the issue timestamp and the moment the desk worked it. ⚠︎ A REAL PREVENTION ALERT IS CARDHOLDER DATA AND IT REACHES YOUR CONFIGURED PROVIDER VERBATIM. Corpus lens →
When is this the wrong choice?Avoid: Paying per alert for a job a regex and a join already do. That is the case against the best-fitting scenario (“A feed that is already structured -- a CSV export or a portal API with named fields, on a merchant with one store”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A REAL FEED. The corpus is four consistent formats from a generator with small phrase pools, and BOTH arms read every field of every alert correctly -- card, both amounts, currency, date and reason all 60 of 60. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE MODEL WOULD EVER SET THE VERIFY FLAG. It missed 4 of 4, which is the whole denominator this corpus offers. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-30 — r001-prevention-alert. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on a copy of the kit with no .env and no API key, on this machine: python3 -m tools.build_corpus rebuilt the 60 alerts, the 247-row ledger and the key; python3 -m evals.check_labels re-graded that key with independent arithmetic (KEY CLEAN, 60 alerts, 15 cases); python3 -m evals.run --floor rules scored the free floor at 57 of 60; and python3 -m src.app served the board (HTTP 200 on every endpoint) with the model button disabled and saying why. The corpus, the ledger, the key and both recorded model runs ship in the repo, so the whole product renders before anyone decides to spend.

A living map of modern AI — kept current every morning