The business caseThe problem this solves
A sponsor issues a safety letter, a distribution system logs what went out, and somebody has to say whether the obligation to each investigational site was discharged. Subtracting dates is arithmetic and a spreadsheet does it perfectly, so the row worth finding is never the one with a missing date -- it is the one that already looks finished. A site that acknowledged inside its window, by somebody who had been removed from the site's delegation log a fortnight earlier. A site that acknowledged inside its window, against a version of the letter that was reissued four days later. A site the register shows as CLOSED, whose closure was rescinded, which is still receiving safety information and which nobody sent the letter to at all. Every printed field is clean in all three. And the trap runs the other way too: a site the register shows as OPEN, in scope on every field, with no log entry, which had actually held its close-out visit before the letter was issued -- pure code calls that a breach and it is not one. somebody reading a distribution extract against a site register before a reconciliation is signed: checking which sites the letter was owed to at all, whether each was sent inside the send window, whether each acknowledgement came back inside its own window counted from a different date, whether the person who acknowledged was entitled to, and whether any sentence in the distribution notes moves a row that the dates say is finished.
Audience
a clinical operations or safety distribution associate deciding which sites to chase before a reconciliation is signed off, and the people who build tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual safety letter reconciliation pack (one SYNTHETIC letter, its sponsor's window table, its site register, its distribution log and its distribution notes)
The corpus is 36 safety letter reconciliation pack (one SYNTHETIC letter, its sponsor's window table, its site register, its distribution log and its distribution notes), 0.23 MB (txt 36). A real distribution extract names an investigator, a site and a safety report that was actually filed; it is not publishable and there is no public corpus of them. And the thing being measured has to be planted to be measured: to count how often a reconciler ticks off an obligation nobody discharged, you have to know which cells were gaps, and a real archive does not come labelled.
The corpus
- The 36 safety letter reconciliation pack (one SYNTHETIC letter, its sponsor's window table, its site register, its distribution log and its distribution notes)generated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your safety letter reconciliation pack (one SYNTHETIC letter, its sponsor's window table, its site register, its distribution log and its distribution notes). That is the whole change — there is no database to migrate.
Safety Letter
----------------------------------------------------------------
Pack SLD-0001
Letter SL-2026-0101
Letter class EXPEDITED
Issue date 2026-06-10
Revision Revision 1 (original issue)
Product Velorimab (VLR-201)
Protocols in scope VLR-201-302
Regions in scope Northern Europe, North America
Reconciliation date 2026-07-16
Prepared by H. Duarte, Clinical Safety Operations
Distribution Standard
----------------------------------------------------------------
Sponsor distribution procedure SOP-CS-014, revision 6, in force from 2026-01-05.
Class Send within Acknowledge within
URGENT 1 days 3 days
EXPEDITED 5 days 14 days
AGGREGATE 15 days 30 days
The SEND window is counted in calendar days from this letter's issue date. The
ACKNOWLEDGEMENT window is counted in calendar days from the date the letter was SENT to
that site, not from the issue date.
An acknowledgement is valid only when it is returned by the investigator of record for
that site, or by the delegated coordinator named for that site, and only when that person
held the delegated duty on the date it was returned.
A site is in scope for this letter when all four hold: its protocol is one of the
protocols in scope; the region printed on its register entry is one of the regions in
scope; it was activated on or before the issue date; and it had not closed before the
issue date.
Nothing is overdue before its window closes. Where a window is still open at the
reconciliation date, the site is neither discharged nor in breach.
Site Register
----------------------------------------------------------------
Site S-2101
Protocol VLR-201-302Abridged — the file continues.
The outcomeWhat a good result looks like
one row per site on the register, in register order: DISCHARGED, NOT_REQUIRED, NOT_DUE, or GAP with one of six named grounds -- plus a letter-level status folded from the cells in pure code. Beside every row, what the strongest free floor alone would have written, computed on every render because it costs nothing.
And when it cannot
⚠︎ THE FREE FLOOR IS PERFECT ON THE 187 GENUINELY DISCHARGED CELLS AND THE PAID ARM IS NOT: 86.1 pct against 100.0 pct. The model wrongly raised 15 of them as gaps and called another 11 NOT_DUE -- 26 sites where nothing was wrong that a person would have to clear by hand. It also names the ground correctly on 95.89 pct of the gaps it catches where the floor manages 100.0 pct, and it loses one of the 39 structured gaps that pure code takes for nothing. Every one of those 30 verdict misses is traced, by letter, in Eval.taxonomy; none of them is a hallucinated site or an invented date.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your distribution log is clean and your site register is current -- the dates, the statuses and the delegated names all say what is actually true — the free floor.
window-gatetakes 100.0 pct of the 39 structured gaps and 100.0 pct of the 44 out-of-scope calls the register's own fields settle, for $0.00, and it beats the paid arm on the 187 discharged cells (100.0 pct against 86.1).
Everything the model adds on this corpus is decided by a sentence in the Distribution Notes. Where your fields are already true, there is no sentence to read and the model call buys nothing you did not already have -- at $0.0456 a letter and a p50 of 101032 ms against milliseconds. - Your register lags reality and the truth arrives as comments, emails and free-text notes — the model call. That is the only column it buys and it buys essentially all of it -- 92.11 pct of the gaps only a note reveals and 100.0 pct of the out-of-scope calls only a note settles, against the free floor's 0.0 pct on both.
No regular expression reaches 'the coordinator had been removed from the delegation log', 'the letter was reissued four days later' or 'the closure was rescinded'. On this corpus the paid arm reached 46 of those 49 cells and stated 0 of 29 complete-looking gap rows as discharged. - You are reconciling thousands of letters a cycle and the bill matters — run the free floor over all of them and send only the letters whose Distribution Notes are non-empty. The floor costs $0.00, is right on 85.33 pct of cells, and cannot be wrong about anything a note does not decide.
The cost is linear in letters and nothing amortises -- no index, no cache, no cross-letter state. 95.83 pct of this run's output tokens are provider-side reasoning, so the bill is the thinking and the only lever is how many letters you send. - A false chase is expensive -- your sites are already over-contacted and a wrong query costs goodwill — the free floor, or the model with a person on every gap it raises before anything leaves the building. The paid arm raised 15 false gaps against the floor's 11, on the same 257 non-gap cells.
This is the one direction where free code is measurably better here, and the page leads with it: 5.84 pct false gaps against 4.28 pct, plus 11 genuinely discharged sites the model parked as NOT_DUE that the floor called correctly. - A missed gap is the thing you cannot afford — the model, with the free floor computed beside it on every row -- which the UI does for nothing. It caught 94.81 pct of gaps against the floor's 50.65 pct, left 0 of 77 gap rows off the worksheet, and stated 0 of 29 complete-looking gap rows as discharged where every free floor states all of them.
The two arms fail in opposite directions, so running both and reading the disagreements is strictly better than either -- and the second column is free.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-27. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own packs into data/corpus/ in the layout data/SOURCES.md describes, write data/gold.jsonl to match, and everything on this page recomputes. ⚠︎ WHAT DOES NOT TRANSFER. Corpus lens → |
| When is this the wrong choice? | Avoid: Do not read 'my register is current' off a policy. It is current if somebody checked; if closures, delegation changes and letter reissues reach you as email rather than as columns, this scenario is not yours and the floor will state as DISCHARGED every one of the 29 complete-looking gap rows in this corpus -- 100.0 pct of them. The floor also cannot see a single one of the 38 gaps a note decides. If you cannot name the last time your register's Status column was reconciled against the actual close-out filings, assume it lags. That is the case against the best-fitting scenario (“Your distribution log is clean and your site register is current -- the dates, the statuses and the delegated names all say what is actually true”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A distribution extract whose sections are not underlined headings: src/segment.py returns one unnamed section, src/pack.py parses nothing, and the UI renders an empty table rather than an invented one. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A SECOND PAID TIER. One model was called against this corpus. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-27 — r001-safety-distribution. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — git clone, then python3 -m src.app -- no install, no key, no network. The corpus, the answer key and every recorded result ship in the repo, and all three free floors are pure Python. requirements.txt names nothing.



