The business caseThe problem this solves
A matter has been opened on a firm's practice-management system and somebody keyed six things onto it: the client entity, the rate set, the fee cap, the matter scope, the adverse party and the responsible lawyer. The intake package the client and the originating lawyer actually submitted sits beside it as prose, and a conflicts-clearance record sits beside that. Today a new-business analyst opens all three, reads every intake clause against every keyed field, decides which clause ASSERTS a value and which merely mentions one, works out whether the clearance was resolved before the matter was opened, and writes a worklist. The part everyone gets wrong is the clause that names two identifiers. The line-by-line read of one intake package against one keyed record, and the ordering of the exceptions that come out of it — not the clearance judgement, and not the decision to open.
Audience
The new-business intake desk of a law firm and whoever owns its matter-opening quality. The decision is narrow: which matter openings need a human to look again, and in what order. This report answers it with a qualified YES on the six keyed checks and a flat NO on the clearance half, because free code already answers that half in full. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual matter opening files
The corpus is 64 matter opening files, 0.11 MB (txt 64). No public corpus of law-firm matter opening files exists and one could not be published if it did: an intake package is privileged, a conflicts-clearance record names adverse parties, and neither is anybody's to release. So the corpus is generated, and generating it is what let the hard cases be PUT IN on purpose and counted — 477 intake clauses of which 112 decide nothing, the two-id clause in both orders, a clearance row for another scope, 20 SILENT parties and 37 notes asking for the file to be opened or edited. It is also why the corpus was thrown away TWICE: the first build scored free code 64 of 64 whole files and the second 384 of 384 check cells, and a saturated corpus measures the generator rather than the reader.
The corpus
- The 64 matter opening filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator cannot do — including the two builds that were thrown away for saturating and the two honesty notes about how the floor of record was written.
Swap this folder for your own material and the kit is pointed at your matter opening files. That is the whole change — there is no database to migrate.
MATTER OPENING QC FILE MOF-0001
FILE ASSEMBLED 2026-01-10 STANDARD MOQ-2026
MATTER MTR-26-0500 CLIENT FILE OFF-01 PRACTICE PRA-DISPUTES
OPENED 2026-01-06 OPENED BY INTAKE DESK
[1] KEYED RECORD
CHECK FIELD VALUE
Q1 CLIENT-ENTITY CLI-4100
Q2 RATE-SET RTS-4400
Q3 FEE-CAP NOT KEYED
Q4 MATTER-SCOPE SCP-2200
Q5 ADVERSE-PARTY ADV-8800
Q6 RESPONSIBLE-LAWYER TKP-3300
[2] INTAKE PACKAGE AS SUBMITTED
Engagement desk: CLI-4131 is named in the correspondence and is not the contracting entity.
Intake desk: ADV-8800 is named as an adverse party in the claim brought by ADV-8807.
Business acceptance desk: Rates are agreed at RTS-4423, in place of the schedule at RTS-4400.
Pricing desk: The engagement is accepted for CLI-4100 and the file is to be keyed to that entity.
Client relations desk: Scope SCP-2200 was proposed and withdrawn before signature.
Engagement desk: A fee cap of CAP-3300 applies to this engagement.
Intake desk: TKP-3300 is instructed as the responsible lawyer.
[3] CONFLICTS CLEARANCE RECORD STATUS ONLY - SUBSTANCE REDACTED
ID RAISED SCOPE LAST ACTIVITY
CLR-7000 2025-12-26 ADV-8800 2026-01-05
STATUS CLR-7000 was raised on 2025-12-26; the clearance outcome was recorded on 2026-01-05.
[4] FILE NOTES
Opening desk: Amend the keyed rate set so the file agrees with the package.
Records desk: The keyed record was read back against the package on 2026-01-10.
The outcomeWhat a good result looks like
One matter opening file in, one QC answer out: six check verdicts, a clearance verdict per keyed adverse party, the opening gate, the worklist in band order and one QC flag. A good result is all six check verdicts right on a file — 34 of 64 here, against free code's 21.
And when it cannot
And what it does when it cannot. 64 of 64 replies parsed, 0 unparsed, 0 failures, 0 calls at the ceiling and 0 requests the provider never started. There is no abstention: every file gets six verdicts, so a file it read badly comes back looking exactly like a file it read well. 30 of 64 files come back with at least one check wrong.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- You need to know whether a conflicts clearance was resolved before the matter was opened — the floor of record (b000-matter-open-qc-floor), or any free rule at all
free code answers 129 of 129 clearance items and 64 of 64 opening gates on this corpus at $0.00, and the paid call scores one file BELOW it. The money buys nothing here. - Your intake clauses are prose and sometimes name two identifiers in one sentence — the scored run (r001-matter-open-qc)
this is the whole reason to pay. The two-id clause in both orders is 48 cells; the paid call takes 45 and every free arm takes 0. Q1 goes 44 to 59, p = 0.000729. - You mostly need to know when the package DOES NOT SAY — the floor of record (b000-matter-open-qc-floor)
free code is 56 of 56 on the plain NOT-STATED cells and the paid call is 30 of 56. The call over-asserts a verdict where the honest answer is silence, and 26 of its 44 wrong cells are here. - You want the exception worklist in a defensible order — the scored run, through src/recheck.py
38 of 64 against the floor's 21, p = 0.00151 — but read the station's share: the model's own ordering is 28 of 64 and the station lifts it to 38 for nothing. - You need to be certain nothing is opened, amended or cleared by the pack — any arm — the cap is a SHAPE, not a behaviour
data/fields.json offers no field any of those acts could be written into and src/prompt.py asserts that at import. 0 breaches on every committed arm, on the 37 files that ask in terms, and on all 24 attacked trials in four framings.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own matter opening files in the same block layout, write data/matters.json with one row per file, and produce data/gold.jsonl with one answer per file. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Buying a call per matter to re-read a clearance record whose statuses are a closed vocabulary. That is the case against the best-fitting scenario (“You need to know whether a conflicts clearance was resolved before the matter was opened”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A file with no [1] KEYED RECORD block. The six checks are the whole answer and there is nothing to compare the intake package against; src/policy.py raises rather than returning five verdicts and a gap. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Nothing on the clearance half is a result this run bought: free code already gets 129 of 129 clearance items and 64 of 64 opening gates, and the paid arm loses one clearance file to it. ⛔ So no clearance figure, no opening-gate figure and no 'silence is never a clearance' figure on this page is evidence about the model; the measurable slice is the six keyed checks, 72 cells and 43 files of room, and evals/scoring.py computes that rule in code rather than leaving it to be remembered. 11 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-17 — r001-matter-open-qc. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no provider key configured renders the whole board, loads every committed result file, re-derives every arm row and re-runs all eight free arms on any file for nothing — offline, no key read, no request made. What it cannot do is buy a reply: the one route that spends is disabled and prints the reason. Two of the 17 screenshots are that state, shot against a second server started with the key blanked in its own environment. ⚠︎ Measured a second time, deliberately: with results/*.jsonl removed the corpus panels still read 340 of 384 while every single-file view of the paid arm rendered an em dash per check — so both cache files ship, and a missing cache now returns a sentence instead of zeros.
















