The business caseThe problem this solves
A telehealth group captures a consent for every visit, and every one of them exists: signed, attested, filed. Whether the EDITION that was signed still covers the visit it is filed against is a different question, and it is the one nobody has time to ask. The consent form has been revised four times in eighteen months. Most revisions are editorial and the old consent carries forward; one of them widened what the patient agrees to, and every consent taken before it stopped covering anything. On top of that a signature can predate its own edition, expire on its own terms, name an edition that is not in the register, or be captured after the visit had already started -- and the consent record arrives in whichever habit the intake team uses: a portal export, a coordinator's narrative, or a transcript of what was said out loud. A person opening every completed visit's consent record, finding which edition of the form was signed and when, and walking the register's revision history to decide whether that edition still covers the visit.
Audience
A consent operations reviewer working a post-visit sweep, and the compliance lead who reads what the sweep found. Neither is being asked to decide anything by this pack: it names the edition, says why it does or does not cover the visit, and quotes the line. A person decides whether to go back to the patient, escalate, or record the visit as delivered without current consent. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual consent verification packets
The corpus is 60 consent verification packets, 0.08 MB (txt 60). It had to be generated, because the real version of this document is a patient's signed consent and there is no public corpus of those there could ever be. So the corpus is built to exercise the thing being measured and nothing else: three intake habits at 20 packets each, so the same fact appears as a labelled field, as a sentence and as something somebody said out loud; seven verdicts in a deliberately uneven mix, so the modal answer is worth only 30%; 18 packets where the intake system's own edition field disagrees with the document; 13 where a second edition is named but not signed; 15 where a second date sits above the signature date; and 10 that ask in writing for the two things this pack may never do. The register itself is four editions with two carry-forward revisions and one that re-consents, which is the smallest shape in which an old edition is not by itself a finding is a real distinction rather than a slogan.
The corpus
- The 60 consent verification packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your consent verification packets. That is the whole change — there is no database to migrate.
CONSENT VERIFICATION PACKET CVP-0001
Prepared 2026-09-10 under Vantera Virtual Health -- Telehealth Consent Form Register | visits 2026-01-15 to 2026-08-31
VISIT RECORD
Visit reference CVP-0001
Patient Tobias Renn
Visit date 2026-01-19
Visit start time 18:20
Service line behavioural health follow-up
Modality delivered video
Patient age class adult
Consent interpreted no
Clinician Dr A. Lindqvist
Intake coordinator M. Prideaux
Form edition logged by intake TC v2.0
CONSENT RECORD (portal export, Northgate portal team)
Patient signature e-signature TR-10000 (patient)
Signature captured 2025-01-14 09:55
Form edition signed TC v2.0
Scope of services behavioural health follow-up, delivered by video
Taken by M. Prideaux, intake coordinator
PRIOR FLAGS
a prior completeness finding was raised on this patient and remains open
INTAKE NOTES
Records has not been asked to retrieve the original scan.
The outcomeWhat a good result looks like
One packet in, four graded answers out plus the two readings they are derived from. On the scored run the paid call read BOTH readings correctly on 60 of 60 packets -- every edition, however it was written, and every signature date, however it was spelled -- and the pure-code station then derived the version verdict correctly on 60 of 60 (100.0%), against 50 of 60 (83.3%) for a free rules floor built from the same register and run through the same station. McNemar exact over the paired packets: 10 the paid call gets and the floor does not, 0 the other way, two-sided p = 0.0020.
And when it cannot
AND ON THE WHOLE ROW A PERSON WORKS, THE FREE FLOOR IS AHEAD. All four cells right at once: 43 of 60 for the paid call against 46 of 60 for the floor, and that difference is inside the noise (McNemar b=11, c=14, two-sided p = 0.69). It loses the completeness half outright -- the items map 49 of 60 against the floor's 57, and the first missing item 50 against 58, which is the one cell where the FLOOR's margin is significant (p = 0.039). Split by intake habit it is starker: on the 20 labelled portal exports both arms verify every edition, and the free floor takes the whole packet 20 times to the paid call's 17. Two more findings sit underneath. Without the pure-code station the paid call's own stated verdict is 48 of 60 against the floor's 50 -- indistinguishable (p = 0.82) -- so the entire measured win comes from doing the arithmetic in code over the model's readings. And the station repairs a verdict but cannot repair a citation: a reply that concluded LAPSED quoted the signature-date line, which is correct for THAT verdict, and when the station recomputes SUPERSEDED the quote is kept rather than invented, so it scores zero.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Consent records that arrive as a LABELLED PORTAL EXPORT, with the edition in a field and the date in another — the free rules floor. Do not pay for this.
MEASURED AT A TIE ON THE VERDICT -- 20 of 20 against 20 of 20 over the 20 form-shaped packets -- and the floor is AHEAD on the whole packet, 20 against 17. Reading a labelled field is a lookup, and a regular expression already does it for $0.00. - Consent records captured as a coordinator's NARRATIVE or a verbal attestation TRANSCRIPT, where the edition is named or spoken rather than printed — the fast tier, WITH the pure-code station behind it
This is where the whole measured margin lives: the verdict is 20 of 20 and 20 of 20 against the floor's 15 and 15. The edition has to be recognised in prose before the register can be applied, and pattern matching collapses. - You need the COMPLETENESS check -- which required items are present, missing or illegible — the free rules floor, on this evidence
It beats the paid call 57 to 49 on the item map and 58 to 50 on the first missing item, and the second of those is significant at p = 0.039. Nothing measured here supports paying for it. - You need a defensible audit trail rather than a triage signal — the fast tier plus the station, and read the citation cell before you rely on it
The verdict is perfect and the quoted line is 53 of 60. A verdict with no locatable line behind it is a finding a reviewer cannot check in one glance.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-10. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/register.json with your own form register -- the editions, their windows, their succession rules, their validity periods and the items each requires -- and point data/corpus/ at your own consent records with data/visits.json carrying the visit facts. THE MEASURED RESULT DOES NOT TRAVEL. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per call for a third of a corpus that discriminates nothing. Route on shape first -- the packet's own CONSENT RECORD heading names the habit. That is the case against the best-fitting scenario (“Consent records that arrive as a LABELLED PORTAL EXPORT, with the edition in a field and the date in another”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet holding more than one consent. One signature, one edition, one packet; a real file with a superseded consent and a half-finished re-consent in it has no representation here. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE LABELLED LINE IS THE LINE A REVIEWER WOULD HAVE QUOTED. The key names one line per flagged verdict -- the signature-date line for LAPSED and SIGNED-AFTER-VISIT, the edition line otherwise. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-10 — r001-consent-version. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, scores both free floors offline, replays every packet of the committed paid run, and rebuilds the corpus byte-identically -- measured, not asserted. The only control that needs a credential is the one that spends, and it is disabled and says so.





