Home › Use Cases › Reconcile an executed order form's commercial terms against every invoice raised under it
Use caseUC0450
🧪 Use-case kit · runnable

Reconcile an executed order form's commercial terms against every invoice raised under it

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A software vendor raises invoices under an executed subscription order for months, sometimes years, and before month-end the revenue-operations desk has to know whether what was billed matches what the customer signed: the product, the seat count, the unit price, the discount, the billing frequency, the start date, the ramp and the renewal uplift, net of any credit memo that really stands. The billing system's own order-to-cash check COUNTS periods and invoiced value; it never reads a quantity, a price, a discount, an amendment or a credit memo. What decides the answer is often a sentence: an amendment withdrawn before the customer countersigned it, a signed page that came back after the extract, an invoice voided and reissued, a credit memo the controller reversed or approved. The billing system's check disagrees with the procedure on 37 of the 64 order packs in this corpus. Opening one order pack, deciding which amendment rows the customer actually countersigned, which invoices stand after every void and reissue, and which credit memos the controller approved and did not reverse, then laying the billing schedule out period by period and comparing every invoice line with the order value of the period it bills, clause by clause.

Audience

The revenue-operations desk and billing operations at a subscription software vendor, working a period's executed orders before the month-end close — and the controller, who alone approves any customer credit an over-billing finding proposes. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual order packs

The corpus is 64 order packs, 0.26 MB (txt 64). It is generated because it has to be. A real order book is a vendor's own commercial record: named customers, negotiated prices and discounts that are confidential, and deal-desk and billing notes written by named people about named accounts. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the sentence that decides the reading — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies.

The corpus

  • The 64 order packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every customer is an invented code and a sector, every order form, amendment, invoice, line and credit memo is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note speaks for the deal desk, billing operations, the controller, the account team or a customer remittance note. evals/check_labels.py sweeps all 64 files for a person-shaped name and an honorific on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your order packs. That is the whole change — there is no database to migrate.

One order pack, as the model receives itOIR-0001.txt · 1 of 64
====================================================================================================
ORDER-TO-INVOICE RECONCILIATION -- ONE EXECUTED ORDER, ONE AS-AT DATE
====================================================================================================
FILE              OIR-0001
CUSTOMER          CU-2200  a regional grocer
ORDER FORM        OF-31000
ORDER EXECUTED    2024-12-20
INITIAL TERM      2025-01-01 to 2025-12-31
RENEWAL TERM      2026-01-01 to 2026-12-31
AS AT             2026-07-05
EXTRACT TAKEN     2026-07-02
PACK COMPILED     2026-07-09
PAYMENT TERMS     net 30 days from invoice date
CURRENCY          USD, every amount before tax
CATALOGUE CODES   PLT-SEAT platform seat, SUP-PRM premium support, API-BLK API capacity block, ANL-MOD analytics module, SBX-ENV sandbox environment

-- COMMERCIAL TERMS ON FILE (the order form and every amendment printed with it) -------------------
ROW        PRODUCT   FROM        TO           QTY    UNIT/MO  DISC  BILLING    SIGNATURE          SOURCE
OT-3100-1  PLT-SEAT  2025-01-01  2025-12-31    40      40.00   10%  QUARTERLY  EXECUTED           OF-31000 cl. 2.1
OT-3100-2  PLT-SEAT  2025-07-01  2025-12-31    60      40.00   10%  QUARTERLY  EXECUTED           OF-31000 cl. 2.2 ramp
OT-3100-3  PLT-SEAT  2026-01-01  2026-12-31    60      42.00   10%  QUARTERLY  EXECUTED           OF-31000 cl. 5.1 renewal +5%
OT-3100-4  SUP-PRM   2025-01-01  2026-12-31     1     400.00    0%  QUARTERLY  EXECUTED           OF-31000 cl. 2.3
OT-3100-5  ANL-MOD   2025-01-01  2025-12-31     1   1,200.00    0%  ANNUAL     EXECUTED           OF-31000 cl. 2.4
OT-3100-6  ANL-MOD   2026-01-01  2026-12-31     1   1,260.00    0%  ANNUAL     EXECUTED           OF-31000 cl. 5.2 renewal +5%

Abridged — the file continues.

The outcomeWhat a good result looks like

One order pack in, one row out: which term rows are in force at AS AT, which invoices stand, which credit memos count, every line's over- or under-billing with the clause and invoice line behind it, both totals to the cent, both exceptions and one of three OIR-2026 verdicts. 57 of 64 packs come back with all seven graded fields right, against 44 for the best free floor and 0 for the billing system's own check.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. The 7 packs it got wrong are named in the kit README with what it answered, and every one is a reading: 3 countersignatures dated before AS AT that it left out, 2 voided invoices it kept, 1 void dated AFTER AS AT that it believed — the one pack in 64 where a real over-billing of 18,420.00 came back MATCHED — and 2 renewal rows it left out on a pack whose money it got right. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your quoting and billing systems record a withdrawn amendment, a void and reissue and a reversed credit as a STATUS on the extract — the free modal floor, and do not buy a call at all
    44 of 64 packs for $0.00. Seat counts, unit prices, discounts, billing frequency, ramps, uplifts, duplicates, unbilled periods and posted credits are all decidable from columns and dates, and on those 39 packs the floor gets 39.
  • Countersignatures, voids, reversals and controller approvals arrive as free text — a deal-desk note, a billing operations note, an email pasted into the pack — the paid call
    This is the whole product. On the 25 packs where a sentence decides a reading the paid arm is 19, the modal floor 5 and the best vocabulary floor 12. Withdrawn amendments 5 of 5, reversed credits 5 of 5 and approved credits 5 of 5, where the modal floor gets 0, 0 and 2.
  • You want real over-billing and under-billing kept from being reported MATCHED above all — the paid call
    The station returned MATCHED on a flagged order once in 64 for the paid arm (OIR-0038, a void dated after AS AT it believed), 9 times for the modal floor and 6 and 10 for the vocabulary floors.
  • You want the billing system's own order-to-cash check audited — either paid or free — both beat it comprehensively
    The check disagrees with OIR-2026 on 37 of 64 packs and gets 0 whole: it returns no reading at all. It is published as an arm so the comparison is against what is running today rather than against nothing.

And where nothing here is good enough:

  • Amendments are routinely countersigned after the quoting extract is taken — neither alone — re-export the terms at AS AT
    The paid arm TIES the modal floor at 2 of 5 late countersignatures: it never admitted a countersignature dated before AS AT. A fresh extract makes the reading a column again.
  • Your orders carry usage overages, mid-term co-terming, proration or a materiality threshold — neither, yet
    No file in this corpus does. The unit of work is fixed-quantity subscription lines billed in advance, and every percentage on this page is against that unit.

At a glanceHow the whole thing runs

89%rechecked all correct pct
1,849 msp50, end to end
$1.48per 1,000 order packs · the fast tier

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own order packs in the same block shape and data/orders.json with your own order register, then run python3 -m evals.run --run-id b000-<yours>-modal --floor modal — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per order for arithmetic you already have. That is the case against the best-fitting scenario (“Your quoting and billing systems record a withdrawn amendment, a void and reissue and a reversed credit as a STATUS on the extract”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?Usage-metered or overage lines. OIR-2026 bills a fixed quantity at a fixed unit price per period; a line whose amount depends on consumption has no order value to compare with, and nothing in this corpus is metered. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-order-invoice. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning