Home › Use Cases › Reconcile a transfer log against the authorizations on file, and name the governing one
Use caseUC0307
🧪 Use-case kit · runnable

Reconcile a transfer log against the authorizations on file, and name the governing one

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

⚠︎ This kit has not been graded

0 completed runs, none scored.

Everything below is coverage, latency and cost. No figure on any of these pages says whether a brief is any good.

The business caseThe problem this solves

A holder logs what actually moved — who sent what, to whom, on which day, under which authorization — and keeps the authorizations themselves in a file. The two are supposed to agree and the log is the weaker record, because the cited id is what somebody typed at the time. The instinct is that checking it is a join: fetch the cited instrument, compare. Five of the six verdicts are why it is not. A memorandum on file amends an effective period or an allowance. A later instrument supersedes an earlier one and the log still cites the old id. A schedule of covered items continues in an annex. A recipient is recorded under a legal name and logged under its trading name. None of that is visible to a lookup on the cited id, and all four are ordinary paperwork. Worse, a transfer can be COVERED and correct while the log row beside it names an instrument that does not govern it — and no verdict can say so. Opening one period's file, and for every logged transfer: finding the instrument that actually governs it rather than the one the row cites, applying whatever memorandum on file amends its period or its allowance, checking the recipient against every name the instrument records for that party, checking the item against a schedule that may continue into an annex, counting that transfer's position among transfers of the same item under the same instrument in date order, and then quoting the one line that decides the verdict — by hand, eight times per packet.

Audience

A records officer or compliance reviewer closing a review period — the person who reads the row and decides whether a transfer needs looking into. It is not the person who grants an authorization, and nothing here grants one. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual review packets (one holder, one review period, eight logged transfers, and every authorization, memorandum and annex on file for it)

The corpus is 60 review packets (one holder, one review period, eight logged transfers, and every authorization, memorandum and annex on file for it), 0.25 MB (txt 60). It is generated because the answer has to be DERIVABLE rather than typed — each packet is built as a STRUCTURE of instruments, memoranda and a transfer log laid over them, the packet is rendered from that structure, and the key is TLR-2026 applied to the same structure. Every event is also laid out with an INTENDED verdict and the generator asserts that the rules actually produce it, so a generator that meant to plant an expiry and planted a covered transfer fails the build rather than shipping a fictional case label. ⚑ AND IT WAS ATTACKED AND REBUILT TWICE, WHICH IS THIS KIT'S METHODOLOGICAL CONTRIBUTION. v1 gave each of the five reading traps exactly ONE wording — five fixed strings, and a dozen lines of regex read all five, so the corpus was measuring a parser. v2 gave each construct FOUR wordings, balanced so every phrasing appears equally often within every case, which means the wording carries no information about the verdict; the free floor still scored 96.7 pct, because the POSITION gave it away — any line printed under AUTH-4110 that mentioned a date was AUTH-4110's amendment. v3 moved the memoranda into their own section, present on EVERY packet including the ones with nothing to amend, so a memo has to be ATTACHED to an instrument by reading it — one of its four subject phrasings names no authorization id at all and identifies its subject only by the date the instrument was issued — and added two inert distractor memoranda per packet. The free floor is 88.3 pct, and that number is the honest measure of how hard this corpus is.

The corpus

  • The 60 review packets (one holder, one review period, eight logged transfers, and every authorization, memorandum and annex on file for it)generated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 60 packets, the transfer register and the whole answer key are generated. Nothing is fetched, scraped, licensed or derived from anything that was. There is no personal data of any kind: every party in the corpus is an ORGANISATION or an internal unit, never a person, and evals/check_labels.py sweeps for person-shaped vocabulary and reports 0 packets. ⚠︎ Controlled-information transfer is governed in the real world by real regimes with real obligations and NONE of them is quoted, paraphrased or relied on here.

Swap this folder for your own material and the kit is pointed at your review packets (one holder, one review period, eight logged transfers, and every authorization, memorandum and annex on file for it). That is the whole change — there is no database to migrate.

One review packets (one holder, one review period, eight logged transfers, and every authorization, memorandum and annex on file for it), as the model receives itTLR-0001.txt · 1 of 60
================================================================================================
CONTROLLED-INFORMATION TRANSFER REVIEW PACKET                          TLR-0001
Holder: Ridgeline Components Ltd (invented)
Transfer register: Register 03 - Ridgeline Components Ltd
Review period: 2026-02        Transfers logged: 8        Material: design packs        Procedure: TLR-2026
================================================================================================

TRANSFER LOG AS FILED
  Each row is a completed transfer of controlled material out of the holder. The last column is the
  authorization the log CITES, which is what the record says and not a finding.
  EVENT  DATE        ITEM      SENDING UNIT           RECIPIENT                   MEDIUM                   AUTHORIZATION CITED
  E-04   2026-02-03  CI-1080   Design Authority       Prosper Bay Machining       physical courier         AUTH-4194
  E-05   2026-02-04  CI-1047   Quality Records        Vantry Structures SARL      secure transfer service  AUTH-4097
  E-06   2026-02-06  CI-1094   Configuration Control  Prosper Bay Machining       data room                AUTH-4194
  E-07   2026-02-09  CI-1040   Technical Library      Wren & Dalt Engineering     secure link              AUTH-4097
  E-08   2026-02-12  CI-1087   Export Records         Prosper Bay Machining       encrypted archive        AUTH-4194
  E-01   2026-02-13  CI-1800   Engineering Records    Sundown Composites Pty      secure link              (none cited)
  E-03   2026-02-24  CI-1054   Test and Validation    Sundown Composites Pty      controlled portal        AUTH-4097
  E-02   2026-02-27  CI-1087   Programme Office       Prosper Bay Machining       encrypted archive        AUTH-4194

AUTHORIZATIONS ON FILE

Abridged — the file continues.

The outcomeWhat a good result looks like

One packet in, eight reviewer's rows out: for each logged transfer the authorization on file that GOVERNS it, the two effective dates as amended, whether the instrument names this recipient and lists this item, the allowance as amended, one verdict from TLR-2026's closed six reached by the FIRST rule the transfer matches, and — on the four verdicts that rest on a document — the single line of the governing instrument that decides it, copied verbatim. 5 of 60 packets came back with the authorization, the verdict and the clause all right on every one of their eight transfers on the scored run; 368 of the 480 individual transfers were completely right.

And when it cannot

And what it does when it cannot. 60 of 60 calls returned a parseable reply, 0 were recorded at_ceiling, 0 event ids were invented, and every miss stays in every denominator. The harms are published apart and never averaged: 27 transfers the key says are findings came back COVERED — a transfer nothing on file covers, settled as a clean row that nobody opens again — and 55 covered transfers were reported as findings, which costs a reviewer a look. The dominant single failure is an invented LIMIT-EXCEEDED: 57 of the arm's 100 verdict errors. 36 of those 100 are one named failure — counting the allowance in event-id order rather than date order — in two halves: 29 invented and 7 missed. ⚑ AND MOST OF IT IS FREE TO UNDO: the pure-code station takes missed findings to 8 and invented findings to 3 without making a call.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your file is clean: every log row cites the instrument that governs it, no memorandum amends anything, no schedule continues into an annex — the free column join alone — evals/baseline.py, --arm rules-columns
    on the eight clean packets of this corpus the join is the whole job. It settles 26.7 pct of ALL packets, and the 61.6 points it gives away to the prose floor are entirely the four constructs a clean file does not contain.
  • Your file is ordinary paperwork — supersessions, amending memoranda, annexes, trading names — the free PROSE floor, and read its 88.3 pct as the number to beat
    it settles 53 of 60 packets for $0.00, misses 3 findings of 148, invents 2 of 332 and names all 12 transfers whose log row cites the wrong instrument. Its patterns were written with the corpus open, which is the only honest way to build a floor — and it is still the best arm here.
  • You are going to buy a call anyway — the call for the READINGS, and src/recheck.py for the verdict
    the arm reads at 96.2-99.4 pct on every column and applies the ordered table at 79.2. Throwing its verdict away and re-deriving it from its own readings costs nothing, moves the verdict to 96.2 pct, takes invented findings from 55 to 3 and missed findings from 27 to 8, and fixes EXPIRED and LIMIT-EXCEEDED completely.
  • You need the transfer that is COVERED, correct, and cites an instrument that does not govern it — the authorization column — never the verdict
    there are 12 such transfers and every one of them is COVERED under every rule. A review that publishes only verdicts hands back a clean sheet over a log that cites the wrong documents. The free prose floor names 12 of 12, the paid arm 11 of 12, and both arms that believe the record name 0.

And where nothing here is good enough:

  • The packets come from a party with a reason to want a transfer quiet — neither arm unsupervised — and read the security block first
    an instruction planted in the register notes deleted a true finding on 4 of the 7 packets it was tried on. The station recovers only what the register itself settles; it cannot restore a reading it never had.

At a glanceHow the whole thing runs

7–8%packet all correct
4,000 msp50, end to end
$4.29per 1,000 transfer review packets · google/gemini-3-flash

Run once, for real, on 2026-09-06. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt and data/registers.json with your own packets and register in the same shapes; src/packet.py's layout is a handful of named regexes bound to the section headings at the top of that file. You will have no answer key, so nothing can be scored and every accuracy figure on this page stops applying to you — none of them transfers. Corpus lens →
When is this the wrong choice?Avoid: Paying for any of it. A join, a date comparison and an ordinal are things code does not get wrong. That is the case against the best-fitting scenario (“Your file is clean: every log row cites the instrument that governs it, no memorandum amends anything, no schedule continues into an annex”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A packet too large for one call. Chunking separates a memorandum from the instrument it amends and a supersession sentence from the transfer it reaches — R-2 and R-3 have no defence against it. 4 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled line is the line a records officer would have quoted. The key names one line per finding — the line that decides the verdict under TLR-2026 — and a reviewer might reasonably cite the instrument's header, or the memorandum AND the clause it amends. 4 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, full rule text — the shipped arm, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-06 — r001-transfer-log-recon. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone, run python3 tools/build_corpus.py, open the board. No key, no install, no index build, no network: requirements.txt names no package, every recorded run ships in results/, and all three free floors, the stub, the second derivation of the key and the whole grader run offline.

A living map of modern AI — kept current every morning