The business caseThe problem this solves
A cheque photographed at remote deposit has to be read and then JUDGED: do the two amounts agree, is it signed, is it endorsed, is the payee somebody on this account, is it stale or post-dated, and has this exact item already been deposited. A clerk opens the image, reads it, and checks each of those against the account by hand. the clerk's manual read of the item and its seven checks
Audience
A deposit-operations lead deciding whether to buy document OCR, and the officer who signs the bill. The answer this kit gives them is specific and slightly uncomfortable: the reading is close to solved and the cheapest measured path is already at the ceiling, so the money question is not which engine — it is whether the model call is buying anything the free floor does not. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual cheques captured at remote deposit
The corpus is 40 cheques captured at remote deposit, 12.76 MB (json 42 · png 80). Synthetic for two reasons, and the second is the important one. Real cheque images carry account numbers belonging to real people and could not be published at all; and they would need every field HAND-TRANSCRIBED to be scorable, which makes the answer key somebody's reading rather than what was printed. Here the key is recorded BEFORE the ink goes down, so it is exact by construction and every grader is arithmetic. It also makes the result CONSERVATIVE in one direction and optimistic in another, and the kit says which: the fonts are cleaner than handwriting, and the phone pipeline is harsher than a bank's own capture app, which deskews before an engine ever sees the image.
The corpus
- The 40 cheques captured at remote depositgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your cheques captured at remote deposit. That is the whole change — there is no database to migrate.
{
"account_holder_names": [
"Greta Lampeter",
"Lampeter Print Works"
],
"deposit_date": "2026-09-05",
"prior_items": []
}
The outcomeWhat a good result looks like
a field record plus the exception list a clerk decides, each exception carrying the rule that fired and the two values it compared
And when it cannot
It flags an item that is fine, or reads a field wrong and flags on that. On the winning track 33 of 33 exceptions are caught and 0 are raised in error across 40 cheques — a clerk opens the item either way, which is why the kit's own recommendation is a draft a clerk signs and never an automatic post.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- you need the fields and nothing else — the free floor, and spend the money on a checksum instead
it scores 0.975 against the model's 0.981 on the same OCR text, for $0.00 - you need the exception list a clerk acts on — the cheapest OCR track plus the model call
exception recall 1.000 against the floor's 0.848 — the model catches 5 exception(s) the floor never raises, and it is the missed one that reaches a customer
And where nothing here is good enough:
- you want an item posted without a clerk — neither, yet
the routing number — what a duplicate check keys on — is read exactly right on 95% of cheques and no checksum guards it, and nothing here verifies a signature against a specimen
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop <id>-front.png, <id>-back.png and <id>-context.json in data/corpus/ and a gold.jsonl beside them with one line per cheque: the twelve printed fields under truth, the exception set, and the account context. The measured numbers do NOT carry onto real cheques. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per cheque for field accuracy that regex already has. That is the case against the best-fitting scenario (“you need the fields and nothing else”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a cheque with two registered names that differ by one token — the payee rule is set membership, and a payee that is a subset of one registered name and not another is accepted, which on a joint account is right and on a business account may not be. 4 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | AWS Textract Detect Text — the cheapest primary-verified row in this kit's stage. tools/run_textract.py is written, wired and correct; the call returns AccessDeniedException: "User: arn:aws:iam::[account]:user/[the kit's IAM user] is not authorized to perform: textract:DetectDocumentText because no identity-based policy allows the textract:DetectDocumentText action". 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is openai-compatible; the Prompt lens states what swapping it costs. The published figures come from 2 models on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-08 — r001-deposit-exceptions-mistral. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured re-scores every committed run offline — the graders are pure code, so re-scoring costs nothing and returns the same answer. python3 evals/score.py needs no install at all. Regenerating the IMAGES needs Pillow (pip install -r requirements.txt) and reproduces them byte-for-byte from seed 20260908; tools/build_corpus.py --check asserts the answer key still reproduces. Only the OCR tracks and the model stage need credentials.

