Home › Use Cases › QC an issued SaaS credit memo against the credit policy and its own evidence
Use caseUC0400
🧪 Use-case kit · runnable

QC an issued SaaS credit memo against the credit policy and its own evidence

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A billing analyst issues a credit memo and it lands on an approver's queue. Six things have to hold at once on every line before anybody signs it: the invoice it credits has to be in a state a credit may attach to at all; the attachment filed against the line has to exist AND be about that invoice; the period being credited has to fall inside the invoice's own service period; the reason coded on the line has to be the reason the attachment actually establishes, because the basis changes with it — full unit price for a billing error, a fraction of it for a service credit, the price less a restocking deduction for a return, a capped discretionary amount for a gesture, nothing of the net at all for a tax adjustment; the tax has to follow the reason, and two of the reasons are exceptions to the ordinary rate; and the approver's tier has to reach the line's amount band — except that a discretionary credit is the controller's whatever the amount. An approver does that by eye, at month end, against an invoice register and an evidence log, on every line of every memo. Joining every credit line to an invoice register by id, to an evidence log by attachment id, and to a schedule of three rates and two tier ceilings — then READING the attachment to decide which basis applies, whether it is even about the right invoice, and when the period credited actually closed. It does not replace the approver: nothing here approves, posts, reverses, re-issues or writes off a memo, and nothing touches the referenced invoice.

Audience

The approver a credit memo is routed to — a billing manager inside the limits, a controller above them — and the revenue-operations desk that owns the credit queue. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual issued credit memo packets

The corpus is 62 issued credit memo packets, 0.23 MB (json 3 · jsonl 1 · md 2 · txt 62). Because the job needs a document that can be WRONG in eight different ways at once and a key that knows which of them reached each line first, and because the two facts the measurement turns on — what an attachment establishes, and when the period credited closed — have to be absent from every column. A real credit queue cannot supply that: the memos are somebody's customers, the attachments are somebody's evidence, and nobody holds a per-line label saying which rule reached the line first. Generating it means the key is DERIVED from the same record the packet is rendered from, the family mix is stated rather than sampled, both directions of the money are built on purpose, and the whole thing rebuilds byte-identically. What it costs is that a keyword floor written with the corpus open can solve it completely, and data/SOURCES.md measures exactly that rather than hoping nobody checks.

The corpus

  • The 62 issued credit memo packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 62 memo packets, data/invoices.json and the whole answer key are generated in-process by the file that sits beside them. Every customer, vendor, account, memo number, invoice number, SKU, item name, date, price, rate, evidence entry and signature is invented. CREDITQC-2026 is invented too and is not an accounting standard, a revenue-recognition rule, a tax rule or anybody's credit policy. See data/SOURCES.md.

Swap this folder for your own material and the kit is pointed at your issued credit memo packets. That is the whole change — there is no database to migrate.

One issued credit memo packet, as the model receives itCM-0001.txt · 1 of 62
CREDIT MEMO QC PACKET - AN ISSUED CREDIT MEMO AGAINST THE CREDIT POLICY AND ITS OWN EVIDENCE

MEMO HEADER
  Memo               CM-0001
  Customer           Aldenmere Logistics Group
  Vendor             Quarrenden Systems
  Account            ACC-3100
  Currency           USD
  Memo date          2026-03-08
  Issued by          M. Hallberg, billing analyst

APPROVAL BLOCK
  Approved by        W. Okeke-Lyall
  Approver role      controller
  Approver tier      tier-3
  Approved on        2026-03-08
  Tier ceilings      Clause 8: tier-1 up to $1,000.00, tier-2 up to $25,000.00, tier-3 above that
  Goodwill rule      Clause 8: a goodwill credit needs tier-3 whatever the amount

INVOICE RECORD (the vendor's own billing record for every invoice this memo references)
   Invoice     Issued       Service period             Status        Tax rate   Invoice net    Invoice tax   Invoice gross
   INV-20000   2025-12-31   2026-01-01 to 2026-01-31   open          8.25 pct   $62,552.00     $5,160.54     $67,712.54
   INV-20003   2026-01-31   2026-02-01 to 2026-02-28   open          8.25 pct   $249,518.00    $20,585.23    $270,103.23

INVOICE LINES
   Invoice     Ln  SKU       Item                         Qty    Unit price    Line net
   INV-20000   1   PLT-STD   Platform Standard seat       205    $95.00        $19,475.00
   INV-20000   2   WKF-AUT   Workflow Automation module   249    $173.00       $43,077.00
   INV-20003   1   ANL-PRO   Analytics Pro seat           108    $596.00       $64,368.00
   INV-20003   2   INT-HUB   Integration Hub connector    229    $410.00       $93,890.00
   INV-20003   3   PLT-STD   Platform Standard seat       270    $338.00       $91,260.00

CREDIT SCHEDULE (Schedule C, in force for the whole of this period)

Abridged — the file continues.

The outcomeWhat a good result looks like

Every line of the memo carries a verdict, the thing it rests on, the amount at issue to the cent and signed, and one row copied verbatim as evidence — and the memo carries PASS or HOLD with the lines named and the gap stated. On the published run the pure-code station reaches the right verdict on 241 of 247 lines and the right recommendation on 59 of 62 memos.

And when it cannot

⚠︎ THE MEMO-LEVEL NUMBER IS 0 OF 62 AND THAT IS NOT A TYPO. memo_all_correct requires every one of the six graded fields on every line, and the arm returned a QUOTED ROW on 152 of the 200 lines that CLEAR, where the answer contract says the citation must be null. Every one of those quotes is locatable in the packet and every one scores zero. The station recomputes the verdict and the amount and deliberately does not repair the citation, because rewriting it would erase the evidence of the failure. Read the verdict and recommendation columns for what the product does and the all-correct column for what it costs to trust the reply as written.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Every fact the verdict turns on is in a column — the pure-column floor, evals/baseline.py --floor rules
    It is free, deterministic, offline, and it is right on 226 of the 226 lines this corpus settles by column. Nothing a model adds can improve a comparison.
  • The answer is in a short attachment written in somebody's own words — the paid call, with the rulebook re-applied in code to its two readings
    It reads 20 of 21 lines the columns cannot settle, against the column floor's 0, exact two-sided p = 1.9e-06.
  • The prose is templated, or you can write the patterns yourself — a keyword floor
    On THIS corpus a keyword floor written with the key open solves it completely for $0.00. That is a fact about generated prose and it is exactly the case a real queue is not — but if your attachments come from a form with fixed phrasings, it is your case too.
  • You need the output to go to a person who will act on it — the paid call PLUS the citation scorer, and read the all-six-fields column
    A verdict with no locatable row behind it is not workable evidence, and this kit measures that separately rather than folding it into accuracy.

At a glanceHow the whole thing runs

98%line verdict pct (rechecked)
2,612 msp50, end to end
$1.97per 1,000 issued credit memo packets · OpenAI GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own packets into data/corpus/ in the same panel layout and the whole free half runs on them: the parser, the engine, all four floors, the refusal reader and the board. The key is the boundary. Corpus lens →
When is this the wrong choice?Avoid: Do not call a model for a date test and two multiplications. That is the case against the best-fitting scenario (“Every fact the verdict turns on is in a column”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A memo packet whose credit lines do not match the fixed column layout. src/rules.py's row expression is anchored on it and a row that does not match is silently not a line — the parser reads a memo with fewer lines rather than raising. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the margin over the de-memorised keyword floor is real. It is five lines, exact two-sided p = 0.14, and this kit calls that NOT SIGNIFICANT rather than a win. 5 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-credit-memo-qc. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone, python3 -m src.app, open the port. No install, no key, no network: the corpus, the register, the key, the four floors' committed results and the paid run's committed result all ship, and the board replays them. The first thing that needs a credential is the one button that calls a model, and it is disabled with the reason printed beside it.

A living map of modern AI — kept current every morning