Home › Use Cases › Check a collection-step request against five protection sources before anything proceeds
Use caseUC0278
🧪 Use-case kit · runnable

Check a collection-step request against five protection sources before anything proceeds

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A collections floor raises thousands of steps a month — disconnect notices, field disconnect orders, service limiters, agency referrals — and every one of them is supposed to be checked against five protection sources first. Two of those five are not in any system: they are paper forms at a district office and a shared mailbox at Regulatory Affairs. So the check log that comes back reads CLEAR on all five, and nobody downstream can tell which of those five lookups actually ran. That is a silent skip, and it is the failure with a person on the other end of it. the line in a collections procedure that says the desk checks five sources, and the spreadsheet column that records CLEAR against all five whether or not anybody could look them up. It does not replace the protection review itself: every output is a referral.

Audience

the collections manager who has to answer 'prove a protection check ran on every single one, with no exceptions, and show me the named human who disposed every flagged case' Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual collection-step request packet

The corpus is 64 collection-step request packet, 0.10 MB (txt 64). Because the row asks for a NEGATIVE to be proved — that no request went through unchecked — and no real collections extract can prove one: the whole failure mode is that the record looks complete. A generated corpus is the only way to know how many protections were really there (31 check cells across 64 requests), how many sources genuinely could not be looked up (128, two per request, on every one) and therefore what a silent skip would have cost. ⚠︎ AND IT IS THE ONLY HONEST WAY TO SHIP MEDICAL MATERIAL AT ALL. A protected-account check reads life-support registry rows and medical certificates; every one of these is fabricated, and the kit asserts nothing about how a real one should be handled.

The corpus

  • The 64 collection-step request packetgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md carries the provenance, the licence pointer and a section titled 'what the generator costs the measurement' in which the corpus attacks itself with measured numbers rather than caveats.

Swap this folder for your own material and the kit is pointed at your collection-step request packet. That is the whole change — there is no database to migrate.

One collection-step request packet, as the model receives itPCR-4001.txt · 1 of 64
COLLECTION-STEP PROTECTION CHECK -- request PCR-4001
Raised 2026-07-18 by M. Tsuruoka (Recoveries)

REQUEST
  Account            8802-445-7288
  Customer           Osric Amrouche
  Premises           15 Tallow Lane, Harrowfield
  District           Harrowfield
  Step requested     FIELD DISCONNECT ORDER
  Balance stage      Stage 2

PROTECTION REVIEW ROSTER
  Officer on duty    C. Vandermeer
  Deputy             P. Oyelaran-Vance

LIFE-SUPPORT REGISTRY -- surname within district, searched AMROUCHE (system export)
  No rows returned on surname AMROUCHE.

MEDICAL CERTIFICATE FILE -- surname within district, searched AMROUCHE (system export)
  No rows returned on surname AMROUCHE.

PAYMENT ARRANGEMENT LEDGER -- account lookup on 8802-445-7288
  No arrangement rows on this account.

VULNERABLE-CUSTOMER DECLARATION FORMS -- NOT DIGITIZED
  Held as paper forms at the district office. No lookup is available from this desk.

COMPLAINT AND INQUIRY DOCKET -- NOT DIGITIZED
  Held as a shared mailbox at Regulatory Affairs. No lookup is available from this desk.

CASE NOTES
  The team reference card states that a source which cannot be looked up from this desk is never recorded as CLEAR.
  No field visit has been booked against this request.

The outcomeWhat a good result looks like

Every request comes back with a declared result for all five sources — CLEAR, PROTECTED or UNVERIFIABLE — the sentence in the notes that alleges something, the referral DPP-2026 produces, and the officer on that shift's roster it now sits with. On the 64-request corpus the check log is complete on 64 of 64 requests on EVERY arm including both free floors, with 0 silent skips out of 128 paper check cells, and the paid call gets all nine fields right on 62 of 64 (96.9 pct) against the free regex floor's 44 of 64 (68.8 pct).

And when it cannot

⚠︎ THE TWO IT GOT WRONG ARE THE SAME SENTENCE, AND THE KEY'S LINE ON IT IS CONTESTABLE. PCR-4027 and PCR-4029 carry "The customer says she has been in hospital for most of the quarter and has only just got home to this address." The key labels that a vulnerable indicator; the call answered none and referred both as REFER-UNVERIFIED rather than REFER-INDICATOR. src/recheck.py does NOT rescue either one — the indicator is a field the station trusts, so a missed reading goes straight through. DP-6 defines vulnerable as age, disability, a young child or an inability to manage the account; a recent hospital stay is arguable in either direction. The case is left in the corpus and scored against, and it is named here, because rewriting a case the model failed is how a corpus stops measuring anything.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to prove no collection step went out unchecked — the free floors alone — either of them
    check-log completeness is 64 of 64 and 0 silent skips on EVERY arm, because src/recheck.py writes DP-0's result for the two off-system sources out of data/policy.json and never out of the packet. The paid call buys none of it.
  • Your five panels already arrive as structured rows — evals/baseline.py's rules floor for the checks, one call for the case notes
    the five check results are 64 of 64 on the regex floor and 64 of 64 on the paid call — a tie — because DP-4 makes them an account-number join. The paid call's whole margin is the indicator (62 of 64 against 44) and what moves with it.
  • Your case notes are free prose written by whoever picked up the phone — the paid call
    this is the whole product. The floor turns "there is no complaint open on this account" into indicator complaint, and a Billing note about an arrangement the ledger already shows ACTIVE into indicator arrangement on all four requests where it appears. The call invented no indicator anywhere (0 against the floor's 15).

And where nothing here is good enough:

  • You want a confidence you can route on — nothing here
    median confidence is 0.99 on the 62 requests the call got fully right and 0.99 on the 2 it did not, floor 0.88 across all 64. The model is exactly as sure when it is wrong.

At a glanceHow the whole thing runs

62%request all correct
16,661 msp50, end to end
$11.27per 1,000 collection-step request packet · Gemini 3 Flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point tools/build_corpus.py at nothing and drop your own packets into data/corpus/ with a row each in data/requests.json (id, account, requester, officer, deputy). ⚠︎ YOU CANNOT BRING YOUR OWN ANSWER KEY FOR FREE. Corpus lens →
When is this the wrong choice?Avoid: Paying for completeness. It is an architecture property, and a prompt is the wrong place to enforce it. That is the case against the best-fitting scenario (“You want to prove no collection step went out unchecked”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a packet whose panels are shaped differently. evals/baseline.py binds to literal headings and an account-number pattern; a lookup export that renamed its columns silently returns CLEAR on every row — and CLEAR is the direction nobody queries. 9 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE PAID CALL IS WORTH BUYING FOR THE CHECK LOG. It is not measured to be: completeness, the five check results and the named human are ties with a free floor. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-disconnect-protect. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — git clone, no pip install, no key: python3 tools/build_corpus.py --check rebuilds all 64 packets byte for byte, python3 -m evals.check_labels re-derives and grades the whole answer key and exits 0, and python3 -m evals.run --run-id b000-disconnect-protect-rules --floor rules answers and scores all 576 graded cells. Nothing in that sequence touches a network.

A living map of modern AI — kept current every morning