The business caseThe problem this solves
A control states what has to happen, how often, who has to do it and what it has to produce. What somebody actually has to show it with is a pile: tickets, approvals, exported listings, meeting notes, text off a screenshot. Assembling that into an evidence pack is two different jobs wearing one name. The first is reading -- does this ticket show the March review was performed, or does it record that the role matrix was amended, which is a different event that names the same control and sits in the same window. The second is arithmetic, and it is the half every count check gets wrong: twelve artefacts against a monthly control over twelve months passes every count ever written, and three of them can be the same March review with July empty. The count is twelve. The evidence is ten months. Reading a pile of control artefacts and deciding, one at a time, which required attribute each one evidences -- and then counting by hand whether what is left covers every period the control required, at the frequency it names.
Audience
Whoever has to say that a control operated over a period, and whoever inherits the consequence of a period that was signed off empty. The reader is the person holding the pile, not the person who wrote the control -- so the answer they need is not a score, it is a verdict per required attribute with the artefacts beside it and a named list of the periods with nothing in them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual control evidence bundles
The corpus is 60 control evidence bundles, 0.35 MB (txt 60). The defect mix an evidence pack actually fails on, in a proportion nobody publishes and nobody could scrape. Fifteen named constructions, one planted per bundle, 50 hard and 10 plain: clustered_uncovered 8, fully_evidenced 7, wrong_performer 6, wrong_attribute_mention 5, noise_only 4, duplicate_occurrence 4, attribute_absent 4, out_of_window 3, event_driven_clean 3, output_wrong_kind 3, overevidenced_clean 3, performer_partial 3, short_count 3, event_driven_gap 2, nothing_filed 2. Twenty bundles have nothing wrong with them at all, because a corpus that is all defects measures a different job from the one a reviewer has. ⚑ THE NUMBER IT EXISTS FOR: on 12 of the 27 NOT_COVERED bundles the artefact count REACHES the required period total -- enough artefacts, a period still empty. A checker that compares a count to a frequency is wrong on all twelve by construction, and the free floor is: 12 for 12. The other 15 uncovered bundles are noise traps an arm can only call covered by attaching something that evidences nothing.
The corpus
- The 60 control evidence bundlesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your control evidence bundles. That is the whole change — there is no database to migrate.
CONTROL EVIDENCE BUNDLE CE-0001
Register Engineering control register
Control C-4423
Period under review 2025-07-01 to 2025-12-31
Assembled 2026-01-28
CONTROL C-4423 Approval of production changes before release
Requirement Every production release is approved before it is deployed, and the approval is recorded against
the release.
Frequency monthly
Periods in window 6
Performed by Release Manager
Must produce signed_approval -- a recorded approval carrying the approver and the date
Approved by Director of Platform Delivery
REQUIRED ATTRIBUTES
occurrence The control activity was performed, at the frequency the control states, across the whole period
under review.
performer Each performance on file was carried out by the role the control names.
output Each performance on file produced the record the control names, in the form it names.
approval Each performance on file was approved by the role the control names.
CANDIDATE ARTEFACTS 17
ARTEFACT A01
Kind approval_record
Dated 2025-08-07
Actor B. Oyelaran
Role Release Manager
Reference REL-69351
Body Release approval for August 2025: 15 releases cut; 14 approved before deployment, 1 held.
Performed by B. Oyelaran, Release Manager. Control C-4423.
ARTEFACT A02
Kind approval_record
Dated 2025-07-18
Actor L. Vasquez-Roth
Role Director of Platform Delivery
Reference REL-77309Abridged — the file continues.
The outcomeWhat a good result looks like
One evidence pack per bundle: the control it is about, one row per required attribute carrying a verdict from a closed three-word vocabulary and the identifiers of the artefacts that support it, a COVERED / NOT_COVERED answer for the whole window, and periods_uncovered -- the required periods with nothing in them, named, so the remaining work is a list rather than a re-read. On the measured run every one of the 60 bundles got the coverage answer right and 45 of 60 were right in every field.
And when it cannot
A period signed off as covered when it is empty. It is the one error the rest of the process does not recover: the sign-off is already given, and whoever finds the hole later finds it with somebody's name against it. The free count check does exactly this on 25 of the corpus's 27 NOT_COVERED bundles -- 92.6 per cent false-covered -- and on all 12 of the bundles where the artefact count reaches the required period total. The cheaper failure in the other direction is a complete file reopened and chased for artefacts that were always in it.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- A pile whose artefacts already carry a clean, parseable date and a kind, and where what you need is the periodicity answer rather than the attribution -- 'does what we have cover every month' over a set somebody has already triaged. — the free rules floor, plus the kit's own recheck. No model at all.
That combination is free and it scores 54 of 60 on the coverage answer, 391 of 411 required periods and 6 false-covered bundles. The paid arm's 60 of 60 is six bundles better. src/coverage.py is the whole of the arithmetic and it never needed a model. - A pile where the hard part is ATTRIBUTION -- which artefact evidences which required attribute, with distractors that name the control and record something else, and roles whose names differ by one word. — the paid arm, and read the attribute columns rather than the headline
This is where the gap is real and structural. PARTIAL: 29 of 29 against the floor's 0 of 29, on both floor columns, because PARTIAL needs a denominator a one-bit checker never computes.performer: 60 of 60 verdicts and 98.6 per cent attachment recall against a floor that attaches nothing at all. Attribute artefact sets exactly right: 185 of 208 against 34. And 0 of 12 count traps waved through against 12 of 12.
And where nothing here is good enough:
populationandapprovalquestions specifically -- 'was it done over the population the control names', 'was it approved by the named role'. — neither arm on the strength of this run
The denominators are 10 and 18 attributes. Onpopulationthe FREE FLOOR WINS, 10 of 10 against the model's 9 of 10, and the model's attachment precision there is 84.3 per cent -- its worst column, with 8 false positives. One row moves that figure by ten points. Nothing here supports a claim in either direction.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own bundles as .txt into data/corpus/ in the fielded layout src/bundle.py parses -- a CONTROL EVIDENCE BUNDLE header, the control block (Control, Frequency, Period under review, Periods in window), a REQUIRED ATTRIBUTES block, and one fielded record per artefact carrying an id, a kind from data/policy.json::artefact_kinds, a readable dated and a body. ⚠︎ THE WHOLE BUNDLE REACHES YOUR CONFIGURED PROVIDER VERBATIM, AND A CONTROL BUNDLE IS THE WORST-CASE DOCUMENT FOR THAT. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per bundle for a periodicity answer that twenty lines of Python already get right nine times in ten. That is the case against the best-fitting scenario (“A pile whose artefacts already carry a clean, parseable date and a kind, and where what you need is the periodicity answer rather than the attribution -- 'does what we have cover every month' over a set somebody has already triaged.”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | REAL PILES, BECAUSE OF THE DATES. Every artefact prints a parseable date in a fielded header and src/coverage.py reads a field rather than guessing. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER ANY OF IT REPEATS. ONE paid run, ONE model, ONE day, never re-fired. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-control-evidence. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured this capture pass on a copy of the kit with no .env and no API key, in a scratch directory so the committed tree was never written to. ⚑ THE REPRODUCIBILITY CLAIM WAS TESTED RATHER THAN REPEATED, ACROSS FOUR PYTHONHASHSEED VALUES (0, 1, 97531, 424242). Every one rebuilt data/corpus/ (60 files), data/gold.jsonl, data/bundles.json and data/corpus-stats.json BYTE-IDENTICAL to the committed set -- zero differing files on all four. A sibling in this same batch claims a byte-identical rebuild and does not deliver one, because it slices a set; this kit does. python3 -m evals.check_labels then re-derived the whole key from data/policy.json with no src/ import in 0.03 s: 60 bundles, 208 attributes, 1,208 attachments, 411 periods, 669 artefact dates, 0 disagreements. Both are free and neither needs a key.



