Home › Use Cases › Reconcile a contractor invoice against the contract and the buyer's own service log
Use caseUC0312
🧪 Use-case kit · runnable

Reconcile a contractor invoice against the contract and the buyer's own service log

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A buyer gets a month's invoice from a service contractor: sixty of them at once, five to nine lines each, every line a single run on a single date. Most of the check is arithmetic the buyer can already do — was the run in our own dispatch log, has it been billed twice, is the base the rate card's rate, is the accessorial one this contract permits, does the figure match what the log's own minutes come to. Two of the questions are not arithmetic. The service-type column is the CONTRACTOR'S CODING, not a fact about the run, and the rate card prices what was actually performed: a short repeated link between two points is a shuttle loop wherever the line says <code>Charter run</code>. And the service log records minutes stationary and never says why, while Clause 6.3 makes a wait the contractor itself caused unchargeable — so the only place that cause appears is the line's own narrative. Joining every invoice line to a service log by run reference and date, to a rate card by service type, and to an accessorial schedule by both — then reading each line's narrative to decide whether the coding is right and who caused any delay, and totalling the difference. One invoice at a time, for a month of them.

Audience

The reconciler in a buyer's finance or contract-management function who works the month's contractor invoices, and the contract manager who has to take whatever comes back to the contractor. The decision they are making is which lines to query — not whether to pay. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual contractor service invoices

The corpus is 60 contractor service invoices, 0.21 MB (txt 60). It is generated because a real one cannot be published. A contractor's invoice names a real buyer, a real supplier, real rates that are usually commercially confidential, and a dispute over it is a live commercial matter. And a corpus of real invoices would carry no answer key: the whole point here is a verdict per line that something can be right or wrong against, so the generator holds two facts the invoice never prints — what each run ACTUALLY was, and whether a recorded wait was the contractor's own — and derives the key from them by running the same engine every arm is graded against. No verdict in gold.jsonl was typed.

The corpus

  • The 60 contractor service invoicesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 60 invoices, data/contracts.json and the whole answer key are generated in-process by the file that sits beside them. Every buyer, contractor, service area, contract number, run reference, rate, date, narrative and signature is invented.

Swap this folder for your own material and the kit is pointed at your contractor service invoices. That is the whole change — there is no database to migrate.

One contractor service invoice, as the model receives itCI-0001.txt · 1 of 60
CONTRACTOR SERVICE INVOICE - RUN BILLING

INVOICE HEADER
  Invoice            CI-0001
  Contract           SVC-3100-01
  Buyer              Ardsley Unified District
  Contractor         Stonebridge Vehicle Hire
  Service area       Hill and valley sector
  Billing period     2026-01-01 to 2026-01-31
  Submitted          2026-02-04

CONTRACT TERMS
  Rate card version       RC-2026-A, in force for the whole of this billing period
  Contracted base rate by service type
    Charter run                 $369.00 per run
    Extended route              $264.00 per run
    Standard route              $214.00 per run
    Shuttle loop                $137.00 per run
  Accessorials permitted, by service type
    Charter run            Waiting time; Fuel adjustment
    Extended route         Waiting time; Extra stop; Fuel adjustment
    Standard route         Waiting time; Extra stop
    Shuttle loop           none
  How a permitted accessorial is computed
    Waiting time           $0.90 per minute beyond 10 free minutes, on the minutes the service log records
    Extra stop             $21.00 per stop beyond the scheduled count, on the stops the service log records
    Fuel adjustment        4 pct of the contracted base rate for the service performed

SERVICE LOG (runs the buyer's dispatch system recorded in this period)
   Run       Date        Wait min  Stops over
   R-4001    2026-01-02  1         0
   R-4004    2026-01-05  0         0
   R-4007    2026-01-08  34        0
   R-4010    2026-01-11  0         0
   R-4013    2026-01-14  0         0

INVOICE LINES
   #  Run      Date        Service billed    Base        Accessorial      Units  Amount      Line total  Narrative

Abridged — the file continues.

The outcomeWhat a good result looks like

One invoice in, and per line: one verdict from a closed set of six, the contract term it rests on, the amount at issue to the cent, and one row of the invoice copied verbatim as evidence. Then for the invoice as a whole: PASS or DISPUTE, the list of lines that do not tie, and the total at issue.

And when it cannot

And what it does when it cannot. On this corpus it got 398 of 400 line verdicts and 58 of 60 invoices completely right, and both failures point the SAME WAY: it named money that is not at issue. CI-0015 line 6 ties and it answered RATE-NOT-CONTRACTED, $85.00; CI-0050 line 1 is ACCESSORIAL-NOT-PERMITTED at $7.70 and it answered RATE-NOT-CONTRACTED at $73.70. Across the whole corpus it OVERSTATED the amount at issue by $151.00 and understated it by $0.00. That is the direction that writes a dispute letter to a contractor who billed correctly, and it is the opposite of how the free floors fail — the column floor understates by $1,229.89 and never overstates by a cent.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your contractors code the service type correctly and your disputes are about arithmetic, duplicates and runs that never happened. — the free column floor alone — evals/baseline.py
    Every one of those is a lookup, an equality or one multiplication. The floor is 384 of 400 lines on this corpus and it never overstates a cent.
  • Your invoice narratives come from a small template pool — a dispatch system writes them, not a person. — the keyword floor, and measure it the way evals/baseline.py does
    A word list over a template pool is memorisation and memorisation is free. Run memorisation_check() on your own corpus first; if your narratives repeat, the paid call is buying nothing.
  • Your contractors write the narrative by hand and the disputes turn on what the run actually was, or on who caused a delay. — the paid call, and keep the recheck station
    That is the whole measured difference here: the defanged keyword floor gets 0 of the 7 contractor-caused-wait rows and the call gets them. The station then re-derives the term, the amount and the totals so only the READING is bought.
  • A wrong DISPUTE costs you more than a missed one — a contractor relationship, or a dispute process with a fee. — the free column floor, or the paid call with a human on every DISPUTE
    The failure directions are opposite and measured: the column floor understates by $1,229.89 and overstates by nothing; the paid call overstated by $151.00 on this run. Two of its 60 invoices named money that was not at issue.

At a glanceHow the whole thing runs

97%invoice all correct pct
65,823 msp50, end to end
$12.70per 1,000 contractor service invoices · openai/gpt-5-6-luna

Run once, for real, on 2026-09-05. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own invoices in the same columnar shape, data/contracts.json with your own rate cards, permitted schedules and dispatch log, and data/policy.md + data/policy.json with your own rulebook. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying for a reading you do not need. That is the case against the best-fitting scenario (“Your contractors code the service type correctly and your disputes are about arithmetic, duplicates and runs that never happened.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An invoice whose line rows are not # run date service base accessorial units amount total narrative separated by runs of spaces. src/rules.py matches one regex and a row it cannot match is silently skipped — an invoice in a different layout parses to zero rows and every arm answers nothing. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?One scored run. Every percentage here is a single measurement of a stochastic system, and a repeat at the same settings was not bought. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 6 models on the fast tier, reasoning left to the provider (the published run) and the free column floor — the whole lookup pass and the free keyword floor as fitted to this corpus and the free keyword floor, memorising patterns mechanically removed and the modal floor — TIES on every line, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-05 — r001-route-invoice-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9312, scores all four free floors over every invoice offline, rebuilds the corpus and the key from the seed, runs the independent label gate, and replays the committed scored run at $0.00. Nothing to install — Python standard library end to end.

A living map of modern AI — kept current every morning