Home › Use Cases › Practice-rule matrix entry drafting
Use caseUC0507
🧪 Use-case kit · runnable

Practice-rule matrix entry drafting

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A multistate telehealth group keeps a practice-rule matrix: one row per territory, one column per practice question, each cell recording what the group's own file says. Somebody has to draft each cell out of an operator-issued source pack — four to seven dated items written in prose by the desks that filed them, some superseded, some carved out, some filed after the drafting day, some reciting the entry already on the matrix, and some wrapping the desk's own wording inside a question. Five parts have to come out of it, and on a fifth of the packs the honest answer is that nothing in the pack settles the cell at all. The first pass over a source pack: reading the items, deciding which sentence governs the cell, and writing the verdict, the value, the item id, the verbatim sentence and the effective date — or saying the pack settles nothing.

Audience

Anyone putting an automated reader in front of a table a person later publishes, where the dangerous answer is not a wrong value but a confident value on a pack that settles nothing. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual source packs

The corpus is 64 source packs, 0.20 MB (md 64). It is invented on purpose, and twice over. The matrix axis is an INVENTED TERRITORY — Alder, Birch, Cedar, Larkspur — not a US state, so no cell can be read as a claim about a real place; and every item is an operator-issued filing on the group's own file, so no outside text is reproduced and nothing states what any outside body requires. Build 1 was measured and thrown away: a free regex over the filing forms read four of the five cells at 64/64 because the FRAME decided everything. The rebuild added a recital family (a value-shaped sentence that proposes nothing, decided only by its header) and a queried family (the desk's own value wording wrapped in a question, with the same question words drawn as lead-ins in front of spans that DO govern). That arm now reads 20/64.

The corpus

  • The 64 source packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your source packs. That is the whole change — there is no database to migrate.

One source pack, as the model receives itpacks/SP-2645-0001.md · 1 of 64
# Source pack -- SP-2645-0001

Synthetic pack. Every territory, desk, date, reference and filing below was generated by tools/build_corpus.py from a fixed seed. The territories and desks are invented words; they are not real places and name no real organisation. **No text of any regulator is reproduced here and nothing below states what any regulator requires.** Every line is an operator-issued item on this group's own file.

Pack reference: SP-2645-0001
Territory: Larkspur
Practice question: identity check point -- at what point in the visit the entry records identity as confirmed
Draft prepared: 2026-03-10
Registered desks in Larkspur: Lamplight, Orchard Gate, Riverbank, Southbank
Matrix of record on that day: before_booking, effective 2025-03-25 (entry M-2017)

## The closed practice-question list a cell key may be drawn from

- first-visit form -- whether the entry records a first visit as needing to be synchronous
- identity check point -- at what point in the visit the entry records identity as confirmed
- interpreter arrangement -- what the entry records about interpreter provision
- recognised modalities -- which consultation modalities the entry recognises for a first visit
- supervision pairing -- the supervision arrangement the entry records for an associate clinician

## The closed value list this question's column may carry

- at_visit_start
- before_booking
- either_point

A proposed cell is a draft. Counsel validates every entry; this pack extracts and drafts and never publishes, and nothing on it is a finding that any entry is correct.

## Source items on file

### SI-01
filed: 2025-07-22
kind: counsel memo on file
issued by: clinical operations desk

Abridged — the file continues.

The outcomeWhat a good result looks like

One proposed matrix cell — verdict, value, source item, the sentence verbatim and the effective date — or an honest null body where the pack settles nothing.

And when it cannot

It recites the entry already on the matrix back as a new filing on 3 of the 18 packs that carry one, proposes a value on 1 of the 14 packs where nothing governs, and gets the later-span rule wrong: 2 of 20 on later_span_governs against the bar's 9.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know whether paying a model beats a rules engine at drafting this cell — proposal-row-all-correct, read with evals/paired.py
    It is the question, it costs nothing to re-run, and the paired test ships with it so the gap cannot be quoted without its p-value.
  • You want to know what the money actually bought — The pre-registered slice table — eight slices, each scored against all six free arms
    One slice survives being scored against every arm: winner_is_a_paraphrase, 8 of 21 against 0 of 21 for all five shipped free arms, p = 0.0078 against each. That is the claim, and it is narrow and real.
  • You need to know whether the cap holds under pressure — anchor-scan plus the pressure grader
    The cap is structural — five cells, none of which could hold a published status, a determination or a deadline — and both graders count rather than assert: 0 and 0 over 88 calls.
  • You want to know how hard the corpus really is — The ceiling arm, b000-practice-rule-cell-tuned
    Free Python that has read the generator gets 64 of 64. That is the honest statement of how templated this corpus is, and it is why the 36-pack gap between the paid arm and the ceiling is a corpus fact and not only a model one.

And where nothing here is good enough:

  • You are worried about a confident answer on a pack that settles nothing — nothing-governs, read beside rulebook-guardrails
    It counts both directions — the honest null called right (9 of 14) and a value proposed anyway (1 of 14) — and names the bar on each.

At a glanceHow the whole thing runs

44%proposal row all correct pct
1,030 msp50, end to end
$20.48per 1,000 proposed matrix cells · Claude Fable 5

Run once, for real, on 2026-09-18. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/packs/*.md with your own source packs in the same shape — a header naming the territory, the practice question, the drafting day, the registered desks and the entry of record; the two closed lists a key and a value may be drawn from; then the dated items. The shape is portable and the labels are not. Corpus lens →
When is this the wrong choice?Avoid: Quoting 43.8% on its own. Against the bar's 60.9% it is a loss at p = 0.0708. That is the case against the best-fitting scenario (“You want to know whether paying a model beats a rules engine at drafting this cell”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A pack whose header names a cell key or a value outside the two closed lists it prints. The contract is a closed enumeration on both axes; an open value has nowhere to go and the station drops it to a sentinel rather than inventing a bucket. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?THE HEADLINE IS A LOSS AND NOTHING HERE RESCUES IT: 28 of 64 against the bar's 39, exact McNemar p = 0.0708, not significant and the point estimate AGAINST the paid call — and it also fails to beat the word list and the frame regex. The one claim this kit may make is the 21-pack paraphrase slice, 8 of 21 against 0 of 21 for all five shipped free arms, p = 0.0078; it never appears on any surface without this sentence. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-18 — r001-practice-rule-cell — 64 source packs, 64 billed model calls, the fast tier, reasoning off. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Clone and run: python3 -m src.app serves the board on 127.0.0.1:9507 with no key, no install and no network — requirements.txt has no third-party entry. Every free command in the README runs in under a second on a laptop; python3 -m evals.floors rebuilds all six free arms from the packs in 0.16 s. The reply caches ship, and the measurement that decided that is in .gitignore: the board is byte-identical without them, and the harness re-buys all 64 calls of r001 on --resume without them.

A living map of modern AI — kept current every morning