Home › Use Cases › The invoices one incoming payment was meant to pay
Use caseUC0370
🧪 Use-case kit · runnable

The invoices one incoming payment was meant to pay

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

Cash arrives in the bank before anybody knows what it is for. A payer sends one amount and types something on the remittance advice -- sometimes an invoice number, sometimes a purchase order, sometimes “the March 2026 statement”, often nothing usable at all -- and the amount is frequently short of the invoice, because they took a settlement discount, or drew on a credit note, or simply underpaid. Today a person opens the remittance beside the open-receivables ledger and decides which invoices the money is against and how much each one gets. What they cannot decide goes to unapplied cash, where it ages: the invoice keeps appearing on the collections list, the customer gets chased for money they have already paid, and the balance sheet says the receivable is still open. The read-the-remittance-beside-the-ledger step of cash application: deciding which open invoices one incoming payment settles and how the cash splits across them.

Audience

Whoever is deciding whether to put a model in front of a cash-application queue. The answer this kit gives them is a qualified NO on the headline and a specific YES on one slice, and it is the per-case table rather than the accuracy figure that carries it. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual payments, each with the open ledger put in front of it

The corpus is 60 payments, each with the open ledger put in front of it, 0.14 MB (json 2). Because the alternative is nothing. Cash application is a job whose input is two private documents at once, so a public corpus does not exist -- and a kit that waited for one would not exist either. What a generated corpus buys is a stated case mix, published per case, that a reader can weigh against their own queue: lump_sum_po 5 of 5 (free rules floor 0, null floor 0); discount_not_earned 4 of 4 (free rules floor 0, null floor 0); lump_sum_statement 3 of 4 (free rules floor 0, null floor 0); clean_single_cited 7 of 7 (free rules floor 7, null floor 7); credit_note 4 of 4 (free rules floor 4, null floor 0); credit_note_absent 2 of 2 (free rules floor 2, null floor 0); short_pay 3 of 3 (free rules floor 3, null floor 0); overpayment 3 of 3 (free rules floor 3, null floor 0); cited_settled 3 of 3 (free rules floor 3, null floor 3); ambiguous_amount 3 of 3 (free rules floor 3, null floor 3); no_match 3 of 3 (free rules floor 3, null floor 3); lump_sum_cited 4 of 5 (free rules floor 5, null floor 0); clean_single_uncited 4 of 6 (free rules floor 6, null floor 6); discount_taken 3 of 5 (free rules floor 5, null floor 0); cited_other_customer 0 of 3 (free rules floor 3, null floor 3). What it costs is that the hard cases were chosen by whoever wrote the generator, which is stated wherever a score is.

The corpus

  • The 60 payments, each with the open ledger put in front of itgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.

Swap this folder for your own material and the kit is pointed at your payments, each with the open ledger put in front of it. That is the whole change — there is no database to migrate.

One payments, each with the open ledger put in front of it, as the model receives itfields.json · 1 of 60
{
  "_comment": [
    "THE ANSWER a cash application has to produce, and the RULEBOOK it is produced under. This",
    "file is the single source: src/prompt.py builds the model's schema and its rulebook block",
    "from it, src/apply.py reads the caps and the draw order out of it, evals/scoring.py grades",
    "the four fields it marks `graded`, evals/baseline.py runs both free floors against the same",
    "rulebook, and tools/build_corpus.py writes both the ledger and the answer key from it.",
    "",
    "⚑ THIS KIT PROPOSES AN APPLICATION. IT NEVER POSTS ONE. Nothing here writes to a ledger,",
    "clears an invoice or moves money, and no output of this kit is a posting instruction. The",
    "unit of work is a PROPOSAL with its evidence attached, for a cash-application clerk to",
    "accept or reject. That boundary is why rule 3 below exists: when the evidence does not",
    "decide, the honest output is an unapplied payment sitting on account, not a guess that",
    "reconciles.",
    "",
    "⚑ EVERY AMOUNT IN THIS KIT IS AN INTEGER NUMBER OF CENTS, END TO END. There is no float",
    "anywhere in the pipeline -- not in the corpus, not in the rulebook, not in the prompt, not in",
    "the scorer. A discount rate is basis points and the allowance is an integer floor division,",
    "so two implementations of the same rule land on the same cent rather than within a cent of",
    "each other. `split` and `unapplied_cents` are scored on EXACT equality; a kit that scored",
    "them to the nearest dollar could not see the defect this rule exists to prevent.",
    "",
    "⚑ FOUR THINGS ARE GRADED AND THE SET OF IDS IS ONE OF THEM. Getting the money right against",
    "the wrong invoice is not a partial success: the cash clears, the aged debt report balances,",

Abridged — the file continues.

The outcomeWhat a good result looks like

A proposed application per payment with its evidence attached -- the invoices, the cents each receives, the cash left unapplied, the receivable left open, and one sentence per invoice saying why. A person accepts or rejects; nothing posts.

And when it cannot

It leaves the payment unapplied and says so. On this run that is what 7 of the 9 misses were: the model read the note correctly, correctly refused to apply an invoice the note named wrongly, and then would not fall through to the amount rule that settles the payment unambiguously. That is the cheap direction and it is not free -- it is exactly the pile the kit exists to shrink.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A remittance stream that mostly carries a clean invoice number -- an EDI 820, a portal, a customer whose AP system always writes the reference the same way — the free rules floor, and nothing else
    There is nothing to read, and that is exactly where the floor is perfect: 36 of 36 clean payments on this corpus, in no time at all, with no key and no network. The model scores 31 of the same 36.
  • A queue where payers routinely cite purchase orders, statement periods or nothing, and take deductions the terms may or may not permit — the model, on the payments the regex cannot resolve
    This is the whole of what it buys: 5 of 5 on purchase orders, 3 of 4 on statement months and 4 of 4 on discounts taken outside their window, against 0 for the floor on every one of them.
  • Anywhere the cost of applying cash to the wrong invoice is high — either arm, with the four-field scorer and the arithmetic in code
    No arm here ever misapplied cash -- 0 of 9 payments the key leaves alone, on every arm. That is a property of the conservative selection rule (two invoices matching equally means apply nothing), not of the model.

At a glanceHow the whole thing runs

42%all four pct
1,359 msp50, end to end
$0.00per 1,000 incoming payments · Gemini 3 Flash

Run once, for real, on 2026-09-10. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/invoices.json and data/payments.json with your own rows in the same shape -- invoice_id, customer_id, issue_date, due_date, amount_cents, po_number, terms_discount_bps, terms_discount_days, credit_note_cents, status; and payment_id, customer_id, payer_name, value_date, amount_cents, remittance_note. ⚠︎ WHAT STOPS BEING TRUE the moment you swap the corpus: every accuracy figure on this page, and the comparison between the arms. Corpus lens →
When is this the wrong choice?Avoid: Paying per payment for a job a regex and an equality test already do better. That is the case against the best-fitting scenario (“A remittance stream that mostly carries a clean invoice number -- an EDI 820, a portal, a customer whose AP system always writes the reference the same way”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A payer with a large open balance. The whole visible ledger goes into the prompt with no blocking step, no index and no shortlist -- deliberately, because pre-selecting candidate invoices would do the judging being measured. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Repeatability. ONE scored run, one day, one model. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-10 — r001-cash-application. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Observed on the committed kit this session, with the model key blanked for the process: API_KEY= python3 -m src.app starts the board on its assigned port and every panel renders -- the payment, the ledger, the answer key, both free floors and the whole arithmetic station -- with the model button saying plainly that no key is configured and nothing was called. python3 tools/build_corpus.py --check, python3 -m evals.check_labels, both floors and python3 -m evals.router all completed with no network. requirements.txt names nothing: there is no install step, because the kit is the Python standard library and nothing else.

A living map of modern AI — kept current every morning