Home › Use Cases › Check each production item of a law-enforcement request before it enters the index
Use caseUC0485
🧪 Use-case kit · runnable

Check each production item of a law-enforcement request before it enters the index

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A casino operator's compliance office receives a law-enforcement request for records — marker ledgers, cage transactions, player activity — and works it one production item at a time before anything is released. For each item somebody has to find the extract on file that answers it, check the chain of custody from the desk that pulled it into the compliance hold, read what counsel recorded on any scope flag, see whether the days of the requested range are covered or a gap memo documents them, and confirm the compliance officer has been told of the request. The desk keeps a tracker line per item, and it is wrong on 36 of the 64 items in this corpus, because the deciding fact is often a note typed after the columns were filled in. Opening one production item, finding the extract that answers it against the index line and any re-pull note, walking the custody hops into the compliance hold and reading every later verification by date, reading counsel's latest recorded decision on a scope flag, counting the uncovered days of the requested range, and checking the compliance officer notices for this request.

Audience

A gaming compliance office and the counsel it works with, deciding whether a model is worth paying to read production items before they enter the production index. This report's own answer is NO on this corpus: the scored run ties the best free reader (43 v 42 of 64, p = 1), and on the 27 items a note decides it gets 19 against free code's 21. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual production items

The corpus is 64 production items, 0.09 MB (json 4 · jsonl 1 · md 2 · txt 64). It is generated because it has to be. A real law-enforcement request file is a live matter: named patrons, a named requester, counsel's privileged notes and real ledgers. None of that can be published, and a redacted version would destroy the very thing being measured — the notes and their dates, which decide the reading on 27 of the 64 items. Generating it also buys something a real file could not: the key is DERIVED from the same structure the files are rendered from, and the failure modes are planted by name and counted.

The corpus

  • The 64 production itemsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement, including the failures named before any call was bought. No law, agency rule, court, operator or firm is named anywhere in the corpus, the code or the write-up, and no response clock or reporting duty is asserted.

Swap this folder for your own material and the kit is pointed at your production items. That is the whole change — there is no database to migrate.

One production item, as the model receives itLE-0001.txt · 1 of 64
LAW-ENFORCEMENT REQUEST PRODUCTION ITEM
  DOC                 LE-0001
  REQUEST             LER-3143
  REQUEST CLASS       FINANCIAL-RECORDS
  AS AT               2026-08-10
  SUBJECT ACCOUNT     PA-40148
  RECORD SET          Marker ledger for patron account PA-40148.
  RANGE REQUESTED     2026-01-01..2026-02-14
  INDEX NAMES         EX-4229
  INDEXED ON          2026-07-21
  TRACKER STATUS      CUSTODY-HOLD

EXTRACTS ON FILE
  EXTRACT  SYSTEM      PULLED-BY            PULLED-ON   COVERS                   ROWS  DIGEST
  EX-4229  SYS-MARKER  CREDIT-DESK          2026-07-15  2026-01-01..2026-02-14  7,156  c30f3acf

CUSTODY LOG
  HOP      EXTRACT  FROM                 TO                ON          CHECK
  CH-1031  EX-4229  CREDIT-DESK          REQUEST-DESK      2026-07-16  MATCH
  CH-1114  EX-4229  REQUEST-DESK         COMPLIANCE-HOLD   2026-07-20  MISMATCH

SCOPE FLAGS (raised on extracts for this item)
  FLAG     EXTRACT  REASON               COUNSEL   DECIDED-ON
  SF-0231  EX-4229  OTHER-ACCOUNT-ROWS   INCLUDED  2026-07-23

GAP RECORD
  MEMO     EXTRACT  NOT-COVERED             DOCUMENTED-ON
  (none)

COMPLIANCE OFFICER NOTICES (this subject account)
  NOTICE   REQUEST   SENT-ON     SENT-BY
  NT-0354  LER-3143  2026-07-12  REQUEST-DESK

NOTES
  - 2026-07-22 (REQUEST-DESK): the digest of EX-4229 was re-run at COMPLIANCE-HOLD on 2026-07-22 and read c3f03acf.
  - 2026-07-26 (COUNSEL-OFFICE): counsel recorded SF-0231 as included on 2026-07-26.

The outcomeWhat a good result looks like

One production item in, one row out: which extract answers it, whether custody stands, counsel's recorded decision, whether the compliance officer was notified, the uncovered days, the exception set and one LRA-2026 status — with the desk's own tracker line printed beside it so a reader can see where they disagree.

And when it cannot

And what it does when it cannot. On the scored run all 64 replies parsed, none stopped at the ceiling and the reply stop fired 0 times; the 21 items it gets wrong are named, and every one of them turns on which of two records is the later one — an older seal note read over a later MATCH hop, a re-pull dated before the index line followed anyway. Its confidence does not flag them: median 0.93 on the whole items and 0.90 on the others.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your production items are decided by printed columns — index line, hop checks, flag decisions, notice lines — and notes rarely change a reading — the free columns_asat floor at $0.00
    on the 37 column-decided items it gets all 37 whole, against the paid call's 24.
  • Your desk writes notes, and a rule 'the latest in-date note naming a record governs' matches how your office works — the free last_note floor at $0.00 — the floor of record
    42 of 64 whole, level with the paid call (p = 1) and ahead of it on the note-decided items, 21 v 19.
  • Your notes carry row counts and digests that disagree with the printed ROWS and DIGEST by a row or a character — the scored run, on that family only
    amount_note is the one family the paid call wins: 19 of 23 against the best shipped free arm's 16.
  • You hold this kit's own corpus builder and can write patterns against its sentences — the tuned regex — and understand what it measures
    it reaches 64 of 64 because it was compiled from the generator's own templates.

And where nothing here is good enough:

  • You want a model to decide what goes to the requester, or to answer the requester — nothing on this page
    the kit has no field in which a reply could release, communicate, determine scope or state a response clock, and 0 of 64 replies and 0 of 16 attacked calls did.

At a glanceHow the whole thing runs

67%rechecked all correct pct
1,522 msp50, end to end
$1.35per 1,000 production items · the fast tier

Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own production items in the same seven-block shape and data/requests.json with your own request register, then run python3 -m evals.run --run-id b000-<yours>-last_note --floor last_note — no key, no network, no spend — to see what free code already gets on your material before you buy a single call. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: The moment a note decides a reading. On the 27 note-decided items it gets 0. That is the case against the best-fitting scenario (“Your production items are decided by printed columns — index line, hop checks, flag decisions, notice lines — and notes rarely change a reading”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A production that is not seven fixed blocks per item — a spreadsheet export, an email thread, a scanned PDF. src/packets.py binds to the literal block headings, so anything else parses to nothing and every reading is empty. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether a second scored run would move the tie. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-16 — r001-lawenf-request. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A copy of the kit with no results/, no docs/, no key configured and nothing installed scored the floor of record offline — 42 of 64 production items whole after the station, 0 calls, 0.66 seconds of wall clock (measured 2026-09-16). The kit is Python standard library only; requirements.txt names nothing to install for the pipeline. The board needs the committed result files to replay; the only thing a missing key removes from it is the button that would buy a call, which is disabled and says why.

A living map of modern AI — kept current every morning