Home › Use Cases › Reconcile a licensing partner's safety-case exchange cycle against its agreement
Use caseUC0443
🧪 Use-case kit · runnable

Reconcile a licensing partner's safety-case exchange cycle against its agreement

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A marketing authorisation holder exchanges safety cases with each licensing partner under a written exchange agreement: which products are in scope, how many calendar days each side has to send a serious or non-serious case, how quickly the other side acknowledges or logs it. Once a cycle closes, four records have to agree — the cases sent, the partner's acknowledgements, the cases received and the intake log rows entered for them. The columns are the easy half: a reference, a version, a product, two dates. What decides the answer is often a sentence — an acknowledgement the partner's gateway generated in error and withdrew, a transmission nullified, a follow-up entered into an existing log row, an amendment whose signature page never came back. Read literally, the columns disagree with the procedure on 23 of the 64 packs in this corpus. Opening one partner's cycle, pairing every acknowledgement with the transmission it references and every intake log row with the receipt it logs, reading each dated note to see whether it withdraws, nullifies, re-points, merges or cancels something before AS AT, applying the agreement's scope on each case's day 0, and counting how many days past the agreement's window each gap surfaced.

Audience

The pharmacovigilance operations desk and the partner liaison desk at a marketing authorisation holder, reconciling a closed cycle with each licensing partner before anyone decides what, if anything, to do about it — and whoever has to choose between a free reconciliation script and a paid call. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual reconciliation packs

The corpus is 64 reconciliation packs, 0.21 MB (txt 64). It is generated because it has to be. A real partner exchange record is a safety database: case narratives about named patients, partner correspondence written by named people, and agreements that are confidential commercial contracts. None of that can be published, and a corpus that could be published would have had the one thing this kit measures — the dated sentence that withdraws, nullifies or re-points a row — stripped out of it first. So the whole thing is invented, declared, and generated from one seed with the key DERIVED by the same rulebook the kit applies. The row is BLOCKED-PENDING-ANCHOR, so no window anywhere is any authority's reporting clock: every window is the fictional agreement's own term, printed in section 1.

The corpus

  • The 64 reconciliation packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every partner is a PTR code, every product a PRD code, every agreement, amendment, case, acknowledgement, receipt and log row is arithmetic on the file index, and THERE ARE NO PEOPLE IN THIS CORPUS AT ALL — a note speaks for the PV operations desk, the partner liaison desk, the intake desk or the partner safety desk. evals/check_labels.py sweeps all 64 files for a person-shaped name and for an authority word on every run and reports 0 of each.

Swap this folder for your own material and the kit is pointed at your reconciliation packs. That is the whole change — there is no database to migrate.

One reconciliation pack, as the model receives itCXR-0001.txt · 1 of 64
PARTNER CASE EXCHANGE RECONCILIATION PACK  CXR-0001
cycle             2026-06-01 to 2026-06-30
AS AT             2026-07-01
compiled          2026-07-06
partner           PTR-11, a licensing partner for territory TR-9
procedure         PXR-2026

== 1. EXCHANGE AGREEMENT — contractual terms agreed between the two parties ==
agreement         XA-6415 version 4, effective 2026-04-15
schedule A        PRD-1377, PRD-2468, PRD-5031, PRD-5719
serious           sent within 4 calendar days of day 0
non-serious       sent within 21 calendar days of day 0
acknowledgement   within 4 calendar days of sent
intake logging    within 3 calendar days of receipt
formats           XF-2, NF-1
amendments
(none)

== 2. OUTBOUND — cases sent to the partner this cycle ==
TX ID     CASE          VERSION  PRODUCT   SERIOUSNESS  DAY 0       SENT        FORMAT
TX-26107  MAH-26-26818  FU1      PRD-1377  SERIOUS      2026-06-02  2026-06-02  XF-2
TX-26109  MAH-26-24649  INITIAL  PRD-5719  SERIOUS      2026-06-09  2026-06-11  NF-1
TX-26113  MAH-26-21353  INITIAL  PRD-2468  SERIOUS      2026-06-10  2026-06-11  NF-1
TX-26115  MAH-26-32127  INITIAL  PRD-2468  SERIOUS      2026-06-11  2026-06-12  NF-1
TX-26116  MAH-26-16850  INITIAL  PRD-5031  SERIOUS      2026-06-14  2026-06-16  NF-1
TX-26117  MAH-26-34882  INITIAL  PRD-2468  NON-SERIOUS  2026-06-16  2026-06-17  XF-2

== 3. PARTNER ACKNOWLEDGEMENTS — received from the partner ==
ACK ID     REFERENCE AS RECORDED  VERSION  PRODUCT   DATE        STATUS
ACK-72971  mah-26-26818           FU1      PRD-1377  2026-06-06  ACCEPTED
ACK-72972  MAH-26-21353           INITIAL  PRD-2468  2026-06-11  ACCEPTED
ACK-72974  MAH-26-24649           INITIAL  PRD-5719  2026-06-15  ACCEPTED
ACK-72975  MAH 26 32127           INITIAL  PRD-2468  2026-06-16  ACCEPTED

Abridged — the file continues.

The outcomeWhat a good result looks like

One reconciliation pack in, one row out: the acknowledgement pairs, the intake log pairs, the rows a dated note takes out, the amendments in force, every row finding as ROW:KIND with the days it surfaced past its window, the exception set and one of six PXR-2026 verdicts. 42 of 64 packs come back with all seven graded fields right, against 41 for the best free floor — a tie, p = 1.00.

And when it cannot

And what it does when it cannot. On the scored run 64 of 64 replies parsed, 0 stopped at the ceiling and no call failed. 22 packs came back with something wrong: 8 lose a finding the key raises, 9 add one the key does not, and 7 get every finding right while a reading is wrong. Only one goes the costly way — CXR-0001, returned RECONCILED for a cycle the key reports MISSING by 15 days, because the reading kept an acknowledgement a note withdrew. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your gateway records a withdrawn, nullified or re-pointed acknowledgement as a STATUS, and your intake log records merged follow-ups as a column — the free columns floor, and do not buy a call at all
    41 of 64 packs whole for $0.00 — one fewer than the paid call, p = 1.00. Late sends, pending and missing acknowledgements, duplicates, out-of-scope products, mistyped references and amendments taking effect are all decidable from columns and dates, and on the 27 packs no note decides the floor gets 27.
  • Withdrawals, nullifications, re-pointed acknowledgements and merged follow-ups arrive as free text — a desk's dated note — the paid call, with the columns floor's acknowledgement pairing read beside it
    This is the product. On the 23 packs where a dated note decides a reading the paid arm is 16, the columns floor 0 and the best free floor 9. The verdict margin over the columns floor is significant, 57 against 43, p = 0.0026.
  • You want no cycle with a missing, late, out-of-scope or duplicated exchange returned RECONCILED — the paid call
    It returned RECONCILED on a flagged cycle once in 64 (CXR-0001, a withdrawn acknowledgement it kept); the columns floor did it 4 times and the leftover floor 17.
  • You want the verdict right and nothing else — either the paid call or the free vocabulary floor
    The paid call gets 57 verdicts and the vocabulary floor 52; the difference is 8 packs against 3, p = 0.23, not significant.
  • No agreement is on file for the partner — any arm — the station returns UNSCOPED from section 1 alone
    All three no-agreement packs came back UNSCOPED on the paid arm and on every floor that reads section 1. The paid reply returned no readings on them, so it still fails whole-pack on all three.

And where nothing here is good enough:

  • Your agreements count business days, reach back across cycles or cover several partners in one pack — neither, yet
    No file in this corpus does. The unit is one partner, one cycle, calendar days, and every percentage on this page is against that unit.

At a glanceHow the whole thing runs

66%rechecked all correct pct
1,945 msp50, end to end
$0.86per 1,000 reconciliation packs · GPT-5.6 Luna

Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own reconciliation packs in the same six-block shape, then run python3 -m evals.run --run-id b000-<yours>-columns --floor columns — it needs no key and costs nothing. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per cycle for pairing you already have — the paid call gets 18 of those 27. That is the case against the best-fitting scenario (“Your gateway records a withdrawn, nullified or re-pointed acknowledgement as a STATUS, and your intake log records merged follow-ups as a column”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An agreement whose windows are counted in BUSINESS days, or from a time of day. PXR-2026 counts calendar days between two printed dates with no holiday list and no time zone; a real agreement that does neither changes every LATE and MISSING day count. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?NOT SHOWN TO BEAT FREE CODE: whole pack 42 vs the columns floor's 41 (p = 1.00); verdict 57 vs 43 (p = 0.0026) but not vs vocab's 52 (p = 0.23); ack pairs 43 vs the floor's 55 (p = 0.012). 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-13 — r001-case-exchange. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, all six free floors and every committed run. pip install -r requirements.txt installs nothing — the kit is standard library only. The only thing a key buys is the ASK THE MODEL button and a new scored run.

A living map of modern AI — kept current every morning