Home › Use Cases › Group and evidence every open GR/IR item before the close
Use caseUC0253
🧪 Use-case kit · runnable

Group and evidence every open GR/IR item before the close

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

GR/IR is a clearing account: a goods receipt credits it, the matching invoice debits it, and a matched pair nets to nothing. What is left over is the residue, and it accumulates all year because every line of it needs somebody to open the documents and work out what happened. Once a year finance clears the residue in a cleanup that is painful for one reason -- by then nobody can say what any line of it WAS, so a write-off memo goes to the auditors with a number on it and no explanation behind it. Today the aging report is the only tool, and it groups by BALANCE SIGNATURE, which is the one thing that does not tell you the cause: a debit residue with a receipt and no invoice attached reads as an unmatched receipt whether the vendor never invoiced, invoiced against the wrong order line, or was sent the whole delivery back. The pre-close pass over the GR/IR aging report -- opening each open line's documents, working out why it is still there and who can close it -- and the cause section of the annual write-off memo. Not the clearing, the write-off posting or the invoice release, which stay with a person.

Audience

A plant controller preparing a close, deciding which open GR/IR lines get chased this month and by whom -- and whether the aged list they are reading points at the lines that will actually be written off in January. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual open GR/IR items

The corpus is 64 open GR/IR items, 0.09 MB (txt 64). A real GR/IR account cannot be published by anybody: the line items name vendors, prices and terms, and no redaction leaves anything worth reading. That would matter less if the prose were what is measured. It is not: what is measured is a cause assignment, the document line behind it, and a replay against what a cleanup later did -- and every one of those needs facts that are KNOWN. Which cause a line is really open for; what its balance signature makes it look like; exactly which trail line establishes the cause; which lines genuinely do not determine one; and what the annual cleanup eventually did with each. That last one is not a label anybody writes down, and without it there is nothing to replay against at all. So the facts are injected and the key is derived. The distribution is the design: invoice-not-received is the largest group by count (15) and by open balance (249,350 USD) and contributed 3.0 pct of the write-off; quantity-variance (29.1 pct), po-line-mismatch (27.2 pct) and trail-unavailable (21.4 pct) contributed 77.7 pct of it between them. That gap between the balance ranking and the write-off ranking is the whole finding, and it is only a finding because both were authored to be checkable.

The corpus

  • The 64 open GR/IR itemsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your open GR/IR items. That is the whole change — there is no database to migrate.

One open GR/IR item, as the model receives itGR-0001.txt · 1 of 64
GR/IR OPEN ITEM GR-0001 -- Receipt with no invoice against it

ITEM FACTS
  Company code                  1000
  Plant                         2400 Vidalia
  GL account                    191100 GR/IR clearing
  Vendor                        V-47119 Selby Valve & Fitting
  Purchase order                4500060137 line 20
  Material                      M-11203 tapered roller bearing, 45 mm
  Posting date                  2025-05-19
  Age at close (days)           42
  Balance                       18,400.00 USD (debit)
  Balance signature             debit residue -- a goods receipt with no invoice document attached
  Close cycle read as at        2025-06-30

DOCUMENT TRAIL
  2025-05-19  goods receipt 5011200941 booked 60 EA of M-11203 at 25.25 USD per EA.
  2025-06-09  AP ran the unmatched-receipt report and no invoice document was found on this line.
  2025-07-06  vendor statement request SR-077219 sent by AP; the vendor confirms in writing that no invoice has yet been raised against purchase order 4500060137 line 20.

NOTES FROM THE PLANT AND AP

The receipt is signed and the material is in stock at the plant.
The buyer holds the vendor's written confirmation and is chasing it on the monthly supplier call.

VENDOR AND TERMS
  Payment terms                 net 45 from invoice date
  Incoterms                     DAP Vidalia
  Vendor contact                accounts receivable desk, Selby Valve & Fitting

The outcomeWhat a good result looks like

One cause per open line with the document-trail line that establishes it, the team that can close it and the document that would -- plus, across the whole close cycle, which cause GROUPS the eventual write-off came from, with the arithmetic beside each one so a reader can check a claim in ten seconds instead of trusting it.

And when it cannot

It groups a line by its balance signature and the wrong team spends the quarter on it -- a buyer chasing a vendor who already invoiced, while the line that gets written off is on nobody's list. Or it force-matches a migrated legacy balance to a plausible cause and sends somebody after a document that does not exist. On the scored run neither happened: 0 of 15 divergent items filed at the signature and 0 of 4 legacy items force-matched. THE FREE KEYWORD FLOOR DID BOTH -- it filed 5 of the 15 at the signature, and it named invoice-not-received as a write-off driver, which is the largest group on the account by open balance and 3.0 pct of the write-off.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your GR/IR account is groomed by keyword or by a spreadsheet filter on the balance signature today. — The free keyword floor, and read this kit's numbers as the argument for stopping there or not.
    It is right about the cause on 84.4 pct of this account for $0.00. What it is wrong about is the 15 items whose signature is not their cause -- 8 of 15 -- and that is where the write-off is.
  • You want the January write-off memo to say WHY, per group, with a document line behind each claim. — The model arm, and read the evidence_quote column before anything else.
    The quote is what turns a grouping into something a controller can check in a minute against the delivery note. Measured: mean coverage 0.9673, mean precision 0.9836, 0 unlocatable of 61 returned.
  • You need to choose an aging threshold for the account and have never had one. — The RP-1 sweep, whole. Not one row of it.
    Six candidates from 30 to 365 days, each with the list size it produces, the share of the eventual write-off it would have surfaced, and the two reasons a write-off is missed. Between 30 and 90 days the answer does not move at all on this account.
  • Your account carries pre-ERP balances nobody can explain. — GR-7 and the trail-unavailable group.
    They are listed at their true age with the trail marked unavailable, not force-matched and not dropped. The model answered 4 of 4 correctly; the null floor force-matched all 4.

At a glanceHow the whole thing runs

100%cause accuracy pct
11,945 msp50, end to end
$8.13per 1,000 open GR/IR items · Gemini 3 Flash

Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own open items in the same four-block shape (ITEM FACTS, DOCUMENT TRAIL, NOTES FROM THE PLANT AND AP, VENDOR AND TERMS), and data/records.json with one row per line carrying its age at the close, its signed balance and -- if you want RP-1 -- what your last cleanup did with it. The measured figures do not travel with the corpus. Corpus lens →
When is this the wrong choice?Avoid: Do not read its aggregate as adequacy. It named invoice-not-received as a write-off driver, which contributed 3.0 pct of the write-off, and missed po-line-mismatch, which contributed 27.2 pct. That is the case against the best-fitting scenario (“Your GR/IR account is groomed by keyword or by a spreadsheet filter on the balance signature today.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An item whose age at the close is unknown. RP-1's sweep is over ages and nothing here guesses one from a posting date typed into the prose -- the local board makes you supply it for a pasted item and says why a pasted item gets no replay. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled trail line is the line a plant controller would have quoted. The key names ONE evidence line per item and scores a neighbouring one carrying the same fact at zero -- GR-0018 is exactly that and it is counted as a miss. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-01 — r001-grir-clear. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on port 9253, scores both free floors offline in under a second, and replays every one of the 64 committed model answers off results/eval-r001-grir-clear.json. The ‘check with the model’ control is disabled and prints why on the page. python3 tools/build_corpus.py --check re-renders the corpus and the key from the table and reports 0 files differing; python3 -m evals.check_labels re-derives all 64 gold rows and the RP-1 driver set from retyped tables and reports no problems. Nothing in that sequence reaches a network.

A living map of modern AI — kept current every morning