Home › Use Cases › Disposition each invoice three-way match exception and name the desk that owns it
Use caseUC0255
🧪 Use-case kit · runnable

Disposition each invoice three-way match exception and name the desk that owns it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A purchase order, a goods receipt and a supplier invoice are supposed to agree. When they do not, accounts payable assembles an EXCEPTION PACK and somebody has to say why they disagree before anybody can say what to do about it. In an oilfield services spend file the packs arrive faster than anybody can read them and most of the disagreements are not what they look like: an invoice written in pallets against a ticket written in sacks reads as an 80 pct short-ship until the packaging conversion three lines further down is applied, and a partial delivery invoiced for what it delivered reads as one too. Today a clerk works the exception queue off the ERP's variance screen, which shows the printed numbers side by side with the purchase order as the quantity authority — and that screen is exactly the reading that gets these wrong. The exception then goes to the desk the wrong disposition routes it to, and a unit difference an AP clerk could have cleared in two days becomes a field supervisor driving to a wellsite to re-count a delivery that was correct. The manual pass over an AP three-way-match exception queue — the dispositioning and the routing, plus the discount-lapse forecast that nobody computes today. It does not replace the payment decision, which stays with the desk it is routed to, and it cannot: the only action this kit can take is to record and route.

Audience

A procurement or AP operations lead deciding whether an exception queue can be dispositioned automatically, and a controller who wants to know which exceptions are about to forfeit an early-payment discount and who approved the ones that were overridden. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual exception packs

The corpus is 72 exception packs, 0.07 MB (txt 72). A real three-way-match exception file cannot be published by anybody. It is the supplier's prices, the operator's contracts and the wellsite's activity in one document, and no redaction leaves anything worth measuring behind. That would matter less if the prose were what is measured. It is not: what is measured is a disposition, the sentence behind it, the desk it routes to and the arithmetic that follows — and every one of those needs facts that are KNOWN. Which disagreement a pack is really about, what it LOOKS like before you convert anything, and whether an override actually carries an approver and a reason are not things you can label reliably by reading somebody else's paperwork. So the corpus is generated from numbers, and the key is what the ladder returns when it is run over them: nobody typed a label, and evals/check_labels.py re-derives all of it from a retyped copy of the same rules.

The corpus

  • The 72 exception packsunder its source's terms — Generated. tools/build_corpus.py, seed 20260901, standard library only, re-runnable to the byte with --check.
  • Where each came fromnowhere — every one of the 72 packs was generated here from a fixed seed (20260901), so there is no source URL to map to and no publisher whose copyright page could be checked. python3 tools/build_corpus.py --check rebuilds all 72 byte for byte and diffs them against disk.

Swap this folder for your own material and the kit is pointed at your exception packs. That is the whole change — there is no database to migrate.

The outcomeWhat a good result looks like

One disposition per pack with the sentence that establishes it, the hold code and the owning desk derived from it in pure code, the override state reported as tracked or silent with no third option that means 'cleared', and a per-period list of which exceptions will age past their discount date with the arithmetic beside each one — so a reader can check a forfeit in ten seconds instead of trusting it.

And when it cannot

It files a pack at its face reading and the exception goes to a nine-day desk instead of a two-day one; or it dispositions a pack that does not determine one; or it reports an override the register flags as though the pack evidenced it. On the scored run NONE of the three happened: 0 of 14 divergent packs filed at the face, 4 of 4 undetermined packs answered needs-buyer-review, 6 of 6 silent overrides caught. WHAT DID HAPPEN is the quote: 8 of 68 evidence sentences were the wrong sentence, and on 7 of those the model quoted the invoice's own unit-price line where the key names the contract's tolerance sentence. That is a real disagreement with the key and it is published as a miss rather than re-labelled — and it is why the FREE FLOOR beats the model on the 40 clean packs, 40 of 40 against 35 of 40.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • An exception file whose packs are regular, whose units never differ, and where the receipt quantity is always recorded — the free rules floor alone
    Measured: on the 54 DIRECT packs the two arms are IDENTICAL at 54 of 54, and on the 40 CLEAN packs the floor is BETTER — 40 of 40 all-correct against the model's 35 of 40, for $0.00. Every one of the model's five losses there is the evidence quote.
  • A file where the invoice and the ticket are written in different units, or where deliveries are split across releases — the model
    14 of 14 divergent packs against the floor's 7 of 14, and 0 filed at the face against the floor's 7. This is the entire margin, and it arrives in dollars through the desk: DL-1 prices the floor's two mis-routings at $38.79 of forfeited discount nobody had to lose.
  • You need the abstention to be real — packs that do not determine a disposition must not be dispositioned anyway — the model
    4 of 4 on the undetermined packs against the floor's 0 of 4. Every branch of a rules ladder terminates in an answer; that is what a ladder is.
  • You need to know whether an override was properly approved — the model
    6 of 6 silent overrides caught against the floor's 1 of 6. The floor's keyword test asks whether somebody signed it; TW-9 asks whether somebody signed it AND said why, and the five packs that carry a name or an automatic rule with no reason are exactly the ones a control review would care about.

And where nothing here is good enough:

  • You only need to know which exception is biggest — neither — sort the register by invoice value
    A value ranking answers that with no model and no rules at all. It is also the wrong question: on this corpus the top three by value are EP-0044, EP-0014 and EP-0024, and only one of those three is on DL-1's list of what is about to cost something.

At a glanceHow the whole thing runs

100%disposition accuracy pct
16,926 msp50, end to end
$11.38per 1,000 exception packs · Gemini 3 Flash

Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt and data/gold.jsonl, and edit data/policy.json — the nine dispositions, the hold-code register, the desk table, every service level, the review service level, the period list and TW-2026's rule text all live there and nowhere else. Every measured number stops being true, and two of them stop being true in a way worth naming. Corpus lens →
When is this the wrong choice?Avoid: The paid call — on this slice it buys nothing and it is not free. That is the case against the best-fitting scenario (“An exception file whose packs are regular, whose units never differ, and where the receipt quantity is always recorded”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack whose period is unknown. DL-1 totals by period and nothing here guesses one from a date in the text — the local board makes you pick one for a pasted pack and says why. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled evidence sentence is the sentence a buyer would have quoted. The key names one per pack and scores a defensible alternative at zero — 7 of the model's 8 quote misses are exactly that, all of them the invoice's own price line on a price-variance pack. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-01 — r001-threeway-match. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, scores both free floors offline, re-derives the answer key independently (579 checks, 0 problems) and replays the committed scored run from its result file. Nothing in that path reaches a network. The only two things that spend are --yes on the r001 run id and the ‘check with the model’ button, which is disabled and says why when no key is set.

A living map of modern AI — kept current every morning