Home › Use Cases › Propose a code from the onboarding file's evidence, and say when it may not stand alone
Use caseUC0289
🧪 Use-case kit · runnable

Propose a code from the onboarding file's evidence, and say when it may not stand alone

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A merchant applies to be boarded. The file carries what the merchant says it does, in its own words; what a capture of its website shows; the channel mix; and sometimes a code a previous processor had it under. An intake analyst reads all of that and proposes a category code, because the code drives which programmes the acquirer runs over the account. Most files are obvious. The ones that are not are the ones where the description and the site do not agree, and those are also the ones that look most ordinary. Reading a merchant onboarding file panel by panel and deciding, file by file, whether the evidence in it is good enough to categorise from.

Audience

An onboarding intake analyst working a queue of applications, and the underwriter who reads what they hand over. The decision is whether this file can be categorised from its own evidence or whether somebody has to resolve something first — and the answer this kit gives to that is a flag, not a code. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual merchant onboarding files

The corpus is 62 merchant onboarding files, 0.07 MB (txt 62). It is generated because it has to be. A real onboarding file is a live business's application and a live website capture, and the third input — the acquirer's own category policy — does not exist outside one acquirer. None of the three can be published. So the choice was a generated corpus that ships and can be inspected, or a real one that cannot be shown, and a kit whose evidence cannot be inspected is a kit nobody can check.

The corpus

  • The 62 merchant onboarding filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 62 files, the ACC-2026 catalogue, the policy and the whole answer key are generated or written in this repository. data/SOURCES.md says so at length and then attacks the corpus on its own terms.

Swap this folder for your own material and the kit is pointed at your merchant onboarding files. That is the whole change — there is no database to migrate.

One merchant onboarding file, as the model receives itMCA-0001.txt · 1 of 62
MERCHANT ONBOARDING FILE

Application id            MCA-0001
Intake week               2026-W36, programme ACQ-ONB-04
Catalogue in force        ACC-2026 as of 2026-09-03

APPLICATION FACTS

Legal name                Copperfen Group Ltd
Trading name              Quillon Shop
Registered town           Fallowbridge
Trading since             2007
Card channel mix          96 pct card present, 4 pct card not present
Average ticket            $25
Projected monthly volume  $65,700
Prior category on file    none on file

STATED BUSINESS DESCRIPTION

We run a sit-down restaurant with a fixed menu and table service, open six days a week. One
counter, one terminal, no orders taken remotely.

WEBSITE EVIDENCE

Capture result            page retrieved, a self-built site
Page title                Quillon Shop — Hello
Captured line             Menu · Reservations · Private dining · Our kitchen

ACQUIRER FILE NOTES

None recorded at intake.

CATEGORISATION RECORD

Proposed by               not yet proposed
Assigned by               an underwriter, after this file is read

The outcomeWhat a good result looks like

One file in, six graded answers out: how the evidence stands, the line of business, the channel, the ACC-2026 code those give, the disposition, and one line copied verbatim out of the file that establishes the activity.

And when it cannot

And what it does when it cannot. On the scored run one file of 62 came back flagged that the key says stands on its own — MCA-0059, where a previous processor's code was read as a second line of business. That is an underwriter's few minutes. Under the adversarial arm, two files of 17 went the other way and were proposed unflagged, and the pure-code station took back neither.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know whether anything went forward with a code and no flag — proposed_without_the_flag, ALWAYS beside flagged_where_supported
    it is the only metric that distinguishes a file an underwriter will look at again from one nobody will, which is the only distinction the job turns on
  • You want to know whether the reading is worth paying for — activity_correct_pct and evidence_basis_correct_pct
    those two are the only fields a term list cannot get from the fixed layout, and the activity is where the paired margin is significant (10/1, p = 0.012)
  • You want the number a buyer would argue with — record_all_correct_pct
    all six fields at once is the strictest thing this kit measures
  • You want to know what the pure-code station buys — recheck_overrides
    it counts how often an arm read the file correctly and then wrote the wrong code or disposition beside it

At a glanceHow the whole thing runs

98%disposition exact pct
29,326 msp50, end to end
$19.33per 1,000 merchant onboarding files · google/gemini-3-flash

Run once, for real, on 2026-09-03. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own onboarding files, data/catalogue.json with your own activities, bands and codes, and data/policy.json/.md with your own rules; then write data/gold.jsonl for your files. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Quoting it alone. The free flag-all floor scores 0 on it for $0.00 by returning conflict on every file, so on its own it is a number a regex owns. That is the case against the best-fitting scenario (“You want to know whether anything went forward with a code and no flag”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?The vocabulary is templated: one stated sentence and one site line per activity, drawn from a fixed table. That is why the rules floor reaches 87.1 pct on the disposition. 10 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?A SECOND RUN OF THE SAME SET. Every figure here is one run of 62 calls. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 2 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-mcc-propose. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9289 and scores every free floor over all 62 files: python3 -m evals.run --run-id b000-mcc-propose-rules --floor rules makes no network call at all. tools/build_corpus.py --check rebuilds every byte and diffs against disk, and it is asserted to agree under two different PYTHONHASHSEEDs.

A living map of modern AI — kept current every morning