Home › Use Cases › Read every preliminary notice logged, and say where the table carries no deadline rule
Use caseUC0496
🧪 Use-case kit · runnable

Read every preliminary notice logged, and say where the table carries no deadline rule

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A general contractor's contracts office logs every preliminary notice a claimant serves and works one intake item at a time. For each item somebody has to find which document on file answers it against the index line and any supersede note, walk the receipt trail hop by hop and read every later dated reading of it, read what the office RECORDED as the claimant's tier rather than what the claimant asserted, check whether the delivery confirmation still stands, and then look the deadline up in the operator's table by jurisdiction and notice type. The desk keeps one INTAKE STATUS line per item, and it disagrees with the approved procedure on 40 of the 64 items in this corpus. Opening one intake item and reading four things out of it by hand: which document answers it, whether the receipt trail stands, the tier the office recorded, and whether the confirmation stands. Everything after that is a lookup and arithmetic the contracts office already does in a spreadsheet — the channel, the received date, the contracting party, the lower-tier flag, the missing fields, the deadline from the operator's table by jurisdiction and notice type, the exception set and the PNI-2026 status.

Audience

A contracts office, and the contracts manager it reports to, deciding whether a model is worth paying to read intake items before a status is set. This report's own answer is NO on this corpus on ACCURACY and a QUALIFIED YES on one thing only: the scored run ties the best free reader (45 v 46, discordant 12/13, p = 1) and is BEHIND it on four of the five named slices, but it passes 0 broken receipt trails against that arm's 5 and 0 disagreeing amounts against its 3. What the money buys here is caution in the expensive direction, paid for with false alarms. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual intake items

The corpus is 64 intake items, 0.11 MB (json 5 · jsonl 1 · md 2 · txt 64). The rule this kit exists for cannot be tested on real notices. It needs a deadline table with pairs DELIBERATELY LEFT EMPTY — here 13 of 18 jurisdiction/notice-type pairs carry a row and 5 carry none, putting 4 items in a position where the only correct answer is the words no rule on file. It also needs the deciding fact to be a dated note rather than a column, on a known number of items: 24 of the 64 items are decided by a note and 40 by the printed columns alone, which is the split the money was bought to win.

The corpus

  • The 64 intake itemsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your intake items. That is the whole change — there is no database to migrate.

One intake item, as the model receives itPN-0001.txt · 1 of 64
PRELIMINARY NOTICE INTAKE ITEM
  DOC                 PN-0001
  NOTICE              PNX-3198
  AS AT               2026-08-10
  PROJECT             PRJ-4824
  JURISDICTION        JUR-DUNMARROW
  NOTICE TYPE         TYPE-SUPPLY
  CLAIMANT            CLM-2158
  SUPPLY STATED       Curtain-wall glazing supply and installation.
  INDEX NAMES         RD-4206
  INDEXED ON          2026-07-17
  INTAKE STATUS       READY

RECEIVED DOCUMENTS
  DOCUMENT  CHANNEL        LOGGED-BY       RECEIVED-ON  PROJECT   WITH      PAGES  RECEIPT-REF  FIELDS-SUPPLIED
  RD-4206   HAND-DELIVERY  FRONT-DESK      2026-07-11   PRJ-4824  PTY-0498     33  012d60fe     F1,F2,F3,F4,F5,F6,F7,F8,F9,F10,F11

RECEIPT TRAIL
  HOP       DOCUMENT  FROM              TO               ON           CHECK
  RH-1327   RD-4206   FRONT-DESK        INTAKE-DESK      2026-07-12   MATCH
  RH-1436   RD-4206   INTAKE-DESK       CONTRACTS-FILE   2026-07-16   MATCH

TIER RECORD (what the claimant claimed, and what the contracts office recorded)
  RECORD    DOCUMENT  CLAIMED  OFFICE      RECORDED-ON
  TR-0379   RD-4206   TIER-1   TIER-1      2026-07-19

SUPPLEMENTS RECEIVED
  SUPP      DOCUMENT  FIELDS-SUPPLIED   RECEIVED-ON
  (none)

DELIVERY CONFIRMATIONS (this claimant)
  CONF      NOTICE     CONFIRMED-ON  CONFIRMED-BY
  DC-0510   PNX-3198   2026-08-11    INTAKE-DESK

KNOWN PARTIES ON THIS PROJECT
  PTY-0498  PTY-0397  PTY-0302

NOTES
  - 2026-08-12 (INTAKE-DESK): on 2026-08-12 INTAKE-DESK logged delivery of PNX-3198 by telephone and recorded it here.

The outcomeWhat a good result looks like

One intake item in, one row out: which document answers it, whether the receipt trail stands, the tier the office recorded, whether the confirmation stands, the channel, the received date, the contracting party, the lower-tier flag, the missing-field set, the deadline or the words no rule on file, the exception set and one PNI-2026 status — with the desk's own INTAKE STATUS line printed beside it so a reader can see where they disagree.

And when it cannot

And what it does when it cannot. On the scored run all 64 replies parsed, 0 stopped at the 1,000-token ceiling and the largest reply drew 310 output tokens (31 pct of it), so no rung of the ceiling ladder was bought. The 19 items it gets wrong are named, and all but a handful turn on which of two dated records is the later one: an envelope-opened note read over a later MATCH hop, a supersede note dated before the index line followed anyway. Its own confidence does not flag them — median 0.90 on the whole items and 0.85 on the others, minimum 0.50.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know whether a model is worth paying to read your intake items — run the ten free floors first, then buy one scored run of the same items
    the floors cost $0.00 and took 0.70 seconds for all ten here, and the floor of record came out at 46 of 64 -- one item ABOVE the $0.031523 call. Buying first and comparing afterwards is how a kit publishes a lift that was never there.
  • You care most about a broken receipt trail getting through — the paid call, and watch model.rechecked_passed_a_broken_trail
    it passes 0 against the floor of record's 5 and 0 disagreeing amounts against its 3. That is the one thing the money measurably buys here.
  • You care most about not being told a deadline that no rule supports — either -- and read the SHAPE, not the score
    no arm invented a date, because deadline can only hold a value from data/rules.json or the literal no rule on file. The cap is a schema before it is a phrase list, and that is what held under four attack framings.
  • Your items are decided by the printed columns rather than by a note — free code, and nothing else
    on the 40 column-decided items the columns as at AS AT get all 40 and the paid call 28, p = 4.88e-04 AGAINST. Free code is PERFECT on that slice.
  • Your notes are prose nobody templated — re-run the whole comparison on your own corpus before deciding anything
    the word-list floors are strong HERE because the notes come from a phrase library, which is also why a regex tuned to that library reaches 64 of 64. Neither figure transfers.

At a glanceHow the whole thing runs

70%rechecked all correct pct
1,617 msp50, end to end
$2.05per 1,000 intake items · the fast tier

Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own intake items in the same eight blocks, put your own procedure in data/pni.json and your own deadline table in data/rules.json, then run python3 tools/build_corpus.py --check to re-render data/pni.md and re-derive data/gold.jsonl through src/intake.py. What stops being true is every number on this page. Corpus lens →
When is this the wrong choice?Avoid: Quoting the paid figure alone. 70.3 pct sounds like a result until the free arm beside it reads 71.9 pct. That is the case against the best-fitting scenario (“You want to know whether a model is worth paying to read your intake items”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?An intake item whose blocks are not the eight fixed headings this generator renders — a scanned notice, an email thread, a spreadsheet export. src/notices.py binds to the literal block headings, so anything else parses to nothing and the arm answers from an empty item. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?The paid call is BEHIND the best free arm this kit SHIPS on four of the five named slices, and level on the fifth. Measured over all ten free arms on the same items rather than against one: on the 24 note-decided items it scores 17 against vocab_asat's 19 (p = 0.625), and the earlier claim that "reading the desk notes is worth paying for -- 17 against 0" was scored against an arm that gets 0 on that slice BY CONSTRUCTION, because the slice is defined as the thing it cannot do. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-17 — r001-prelim-notice. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board and scores all ten free floors offline: measured at 0.70 seconds for the full sweep, with python3 -m evals.baseline --port-check reporting 10 of 10 arms matching the generator's own figures. results/cache-r001-prelim-notice.jsonl and results/cache-x001-prelim-notice.jsonl ship, so the paid arm's replies replay for free and every grader, every p-value and every panel re-derives at $0.00. Nothing in this repository is ignored by git except the shared call ledger.

A living map of modern AI — kept current every morning