Home › Use Cases › Reconcile a loyalty partner settlement statement against the partner agreement
Use caseUC0340
🧪 Use-case kit · runnable

Reconcile a loyalty partner settlement statement against the partner agreement

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A loyalty programme sells points to its partners and buys them back when a member redeems. Once a period closes the partner sends a settlement statement — line after line of points and money, one net position at the bottom — and somebody has to decide whether to believe it. The agreement's own earn rate, its elite tier bonus rate, its redemption cost, its volume rebate and its minimum-volume charge are all knowable; what is not is whether every line on the statement is activity the agreement supports. On 21 of these 62 periods it is not: a redemption billed for an award that was cancelled and the points returned, a batch run billed twice, a line covering activity already settled on an earlier statement, a points column entered in batches. Somebody has to open the file, read the note printed under every line, decide which lines really settled, notice whether the rate was amended, and say where this partner's period really stands. Opening one partner's settlement statement, multiplying the agreement's rate schedule out against the programme's own recorded activity by hand, adding the billed lines up, reading each note under a line to decide whether the activity really settled, checking the partner notes for a rate amendment, and deciding whether the difference breaks either investigate threshold.

Audience

A loyalty programme's settlement desk working a monthly partner list, and the finance analyst behind it. Whoever decides that this partner's period gets escalated, gets a note, or gets nothing at all this month. The position they produce is one a programme manager approves. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual settlement file

The corpus is 62 settlement file, 0.20 MB (txt 62). It is generated because it has to be. A real partner settlement statement is a commercial agreement's own trading record, and the exact shapes this kit measures — a redemption billed for an award that was cancelled, a batch run billed twice, a programme manager asking for a statement to be short-paid — are the rows a loyalty programme would least want published. Generating it also makes the key DERIVED rather than written: each file is built as a structure, the panels are rendered from it, and PSR-2026 is applied to the same structure by src/policy.py. There is no second place the answer lives.

The corpus

  • The 62 settlement filegenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every programme and partner is an invented trading name, every rate code, batch reference and line id is arithmetic on the file index, and there is NO MEMBER anywhere in the corpus — no name, no number, no address — which matters more here than on most kits, because a loyalty programme's real settlement data is member data. evals/check_labels.py sweeps all 62 files for six families of identifier on every run and reports 0.

Swap this folder for your own material and the kit is pointed at your settlement file. That is the whole change — there is no database to migrate.

One settlement file, as the model receives itPSF-0001.txt · 1 of 62
==============================================================================================================
PARTNER SETTLEMENT FILE                                                    PSF-0001
Programme: LOY-3100 - Cascade Rewards (invented)
Partner: PTR-AIR-04   Meridian Air   Period: 2026-03-01 to 2026-03-20   Procedure: PSR-2026
==============================================================================================================

AGREEMENT AND PERIOD AS THE PROGRAMME HOLDS IT
  currency unit                             USD
  partner kind                          airline
  minimum issuance commitment           8936000   points in the period
  tolerance pct                            0.75   pct of the agreement amount
  threshold amount                       500.00   or more
  threshold pct                            5.00   pct or more
  period                             part-month

RATE SCHEDULE AS PRINTED
  RATE       WHAT IT COVERS                 DIRECTION            PER 1K   RECORDED POINTS
  RC-4000    Base issuance                  partner-owes          15.70           9820000
  RC-4013    Elite tier bonus issuance      partner-owes          19.30            491000
  RC-4026    Redemption cost                programme-owes         8.20           5302000
  RC-4039    Volume rebate                  programme-owes         1.35           9820000

SETTLEMENT STATEMENT AS BILLED BY THE PARTNER
  LINE       DATE             POINTS  TYPE           PER 1K  STATUS        AMOUNT  REF          MEMO
  LN-0100    2026-03-03      7918547  issuance        15.70  BILLED     124321.19  BX-4001      395 batches at 20000 pts per batch plus 18547 part-batch, batch run BX-4001

Abridged — the file continues.

The outcomeWhat a good result looks like

One partner period in, one row out: which statement lines this procedure treats differently from the statement that printed them, whether a rate amendment was in force, and one verdict from a closed set of five — STATEMENT-BROKEN, ESCALATE, OVERBILLED, UNDERBILLED or IN-LINE. The supported settlement, the agreement amount, the variance and its percentage are re-derived in pure code from those readings and the partner register, so they follow PSR-2026 whatever the reply said.

And when it cannot

And what it does when it cannot. On the scored run 62 of 62 replies parsed and nothing stopped at the ceiling, so there is no unparsed row to report. What it gets WRONG is published by name and it is the headline: the arm returned 88 citations against a key of 29, and on the 21 periods where the statement is RIGHT it invents lines to remove. 2 periods came back carrying a line the agreement does not support (billed_line_the_agreement_does_not_support) and the pure-code station cannot see either — the line is real, the amount multiplies out, and the partner register knows nothing about any statement line.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your partner statements are mostly clean — the billing is right and you are reconciling to confirm it — the incumbent, and do not buy a call at all
    On the 21 periods here where the statement is right, copying the back office's own position scores 5 of 5 on the clean family and 8 of 8 on the threshold splits; the paid arm scores 3 and 4. It invents citations where there is nothing to cite, and every spurious one moves the derived settlement.
  • Your bad lines announce themselves — every reversal says "cancelled", every duplicate says "duplicate" — the free rules floor
    A keyword list over the notes reaches everything a keyword can reach for $0.00. It ties the paid arm on the statement_disagrees family at 8 of 21 each.
  • Your partner notes describe rate moves that were quoted, modelled or proposed and not all of them were executed — buy the call, and grade it on rate_amendment alone
    62 of 62 against a regex's 51 and the incumbent's 56, with 0 of the eleven refused rates applied and 0 invented. A regex cannot tell a signed variation from a costed proposal, because both state a date, a code and a figure in full.
  • Your notes talk about reversals and credits OF SOMETHING ELSE — an earlier batch, another partner's file, a duplicate already backed out — buy the call
    24 of the 62 periods carry one and the best free code is 0 of 24 on them: its keyword list removes money that really settled on every single one. The paid arm is 12 of 24 and the incumbent 10.
  • You need the periods whose net position is right BECAUSE two errors cancel — buy the call — nothing free finds these
    Both free floors score 0 of 4 on the offsetting family and the paid arm scores 2. No amount on the answer can carry that finding and neither can the verdict; cites is the only field that can.

At a glanceHow the whole thing runs

42%all five correct pct
2,092 msp50, end to end
$0.00per 1,000 settlement file · google/gemini-3-flash

Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own settlement files in the same shape and data/partners.json with your own agreement master, rate schedule, recorded activity and thresholds, then rebuild the key by labelling them. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: AVOID if any of your partners bills lines that are wrong without saying so in a note. The incumbent copies a statement, so it is right only while the statement is right, and it is 0 of 21 on the periods here where it is not. That is the case against the best-fitting scenario (“Your partner statements are mostly clean — the billing is right and you are reconciling to confirm it”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A statement that is not fixed-width columns. src/statement.py's row regex is the shape these files print; a CSV or a partner portal export needs a different parser and nothing above it changes. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed — which matters more than usual here, because the headline is a loss by five periods and five periods is not a wide margin. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-09 — r001-partner-settle. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all four free floors, all six committed runs and every screenshot. python3 -m evals.check_labels and python3 tools/build_corpus.py --check both run on a machine with nothing installed.

A living map of modern AI — kept current every morning