Home › Use Cases › The evidence a control needs, and the periods nothing covers
Use caseUC0234
🧪 Use-case kit · runnable

The evidence a control needs, and the periods nothing covers

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A control states what has to happen, how often, who has to do it and what it has to produce. What somebody actually has to show it with is a pile: tickets, approvals, exported listings, meeting notes, text off a screenshot. Assembling that into an evidence pack is two different jobs wearing one name. The first is reading -- does this ticket show the March review was performed, or does it record that the role matrix was amended, which is a different event that names the same control and sits in the same window. The second is arithmetic, and it is the half every count check gets wrong: twelve artefacts against a monthly control over twelve months passes every count ever written, and three of them can be the same March review with July empty. The count is twelve. The evidence is ten months. Reading a pile of control artefacts and deciding, one at a time, which required attribute each one evidences -- and then counting by hand whether what is left covers every period the control required, at the frequency it names.

Audience

Whoever has to say that a control operated over a period, and whoever inherits the consequence of a period that was signed off empty. The reader is the person holding the pile, not the person who wrote the control -- so the answer they need is not a score, it is a verdict per required attribute with the artefacts beside it and a named list of the periods with nothing in them. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual control evidence bundles

The corpus is 60 control evidence bundles, 0.35 MB (txt 60). The defect mix an evidence pack actually fails on, in a proportion nobody publishes and nobody could scrape. Fifteen named constructions, one planted per bundle, 50 hard and 10 plain: clustered_uncovered 8, fully_evidenced 7, wrong_performer 6, wrong_attribute_mention 5, noise_only 4, duplicate_occurrence 4, attribute_absent 4, out_of_window 3, event_driven_clean 3, output_wrong_kind 3, overevidenced_clean 3, performer_partial 3, short_count 3, event_driven_gap 2, nothing_filed 2. Twenty bundles have nothing wrong with them at all, because a corpus that is all defects measures a different job from the one a reviewer has. ⚑ THE NUMBER IT EXISTS FOR: on 12 of the 27 NOT_COVERED bundles the artefact count REACHES the required period total -- enough artefacts, a period still empty. A checker that compares a count to a frequency is wrong on all twelve by construction, and the free floor is: 12 for 12. The other 15 uncovered bundles are noise traps an arm can only call covered by attaching something that evidences nothing.

The corpus

  • The 60 control evidence bundlesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.

Swap this folder for your own material and the kit is pointed at your control evidence bundles. That is the whole change — there is no database to migrate.

One control evidence bundle, as the model receives itCE-0001.txt · 1 of 60
CONTROL EVIDENCE BUNDLE  CE-0001
  Register              Engineering control register
  Control               C-4423
  Period under review   2025-07-01 to 2025-12-31
  Assembled             2026-01-28

CONTROL C-4423  Approval of production changes before release
  Requirement           Every production release is approved before it is deployed, and the approval is recorded against
                        the release.
  Frequency             monthly
  Periods in window     6
  Performed by          Release Manager
  Must produce          signed_approval -- a recorded approval carrying the approver and the date
  Approved by           Director of Platform Delivery

REQUIRED ATTRIBUTES
  occurrence            The control activity was performed, at the frequency the control states, across the whole period
                        under review.
  performer             Each performance on file was carried out by the role the control names.
  output                Each performance on file produced the record the control names, in the form it names.
  approval              Each performance on file was approved by the role the control names.

CANDIDATE ARTEFACTS  17

ARTEFACT A01
  Kind                  approval_record
  Dated                 2025-08-07
  Actor                 B. Oyelaran
  Role                  Release Manager
  Reference             REL-69351
  Body                  Release approval for August 2025: 15 releases cut; 14 approved before deployment, 1 held.
                        Performed by B. Oyelaran, Release Manager. Control C-4423.

ARTEFACT A02
  Kind                  approval_record
  Dated                 2025-07-18
  Actor                 L. Vasquez-Roth
  Role                  Director of Platform Delivery
  Reference             REL-77309

Abridged — the file continues.

The outcomeWhat a good result looks like

One evidence pack per bundle: the control it is about, one row per required attribute carrying a verdict from a closed three-word vocabulary and the identifiers of the artefacts that support it, a COVERED / NOT_COVERED answer for the whole window, and periods_uncovered -- the required periods with nothing in them, named, so the remaining work is a list rather than a re-read. On the measured run every one of the 60 bundles got the coverage answer right and 45 of 60 were right in every field.

And when it cannot

A period signed off as covered when it is empty. It is the one error the rest of the process does not recover: the sign-off is already given, and whoever finds the hole later finds it with somebody's name against it. The free count check does exactly this on 25 of the corpus's 27 NOT_COVERED bundles -- 92.6 per cent false-covered -- and on all 12 of the bundles where the artefact count reaches the required period total. The cheaper failure in the other direction is a complete file reopened and chased for artefacts that were always in it.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A pile whose artefacts already carry a clean, parseable date and a kind, and where what you need is the periodicity answer rather than the attribution -- 'does what we have cover every month' over a set somebody has already triaged. — the free rules floor, plus the kit's own recheck. No model at all.
    That combination is free and it scores 54 of 60 on the coverage answer, 391 of 411 required periods and 6 false-covered bundles. The paid arm's 60 of 60 is six bundles better. src/coverage.py is the whole of the arithmetic and it never needed a model.
  • A pile where the hard part is ATTRIBUTION -- which artefact evidences which required attribute, with distractors that name the control and record something else, and roles whose names differ by one word. — the paid arm, and read the attribute columns rather than the headline
    This is where the gap is real and structural. PARTIAL: 29 of 29 against the floor's 0 of 29, on both floor columns, because PARTIAL needs a denominator a one-bit checker never computes. performer: 60 of 60 verdicts and 98.6 per cent attachment recall against a floor that attaches nothing at all. Attribute artefact sets exactly right: 185 of 208 against 34. And 0 of 12 count traps waved through against 12 of 12.

And where nothing here is good enough:

  • population and approval questions specifically -- 'was it done over the population the control names', 'was it approved by the named role'. — neither arm on the strength of this run
    The denominators are 10 and 18 attributes. On population the FREE FLOOR WINS, 10 of 10 against the model's 9 of 10, and the model's attachment precision there is 84.3 per cent -- its worst column, with 8 false positives. One row moves that figure by ten points. Nothing here supports a claim in either direction.

At a glanceHow the whole thing runs

100%coverage accuracy pct
50,444 msp50, end to end
$22.68per 1,000 control evidence bundles · Gemini 3 Flash

Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own bundles as .txt into data/corpus/ in the fielded layout src/bundle.py parses -- a CONTROL EVIDENCE BUNDLE header, the control block (Control, Frequency, Period under review, Periods in window), a REQUIRED ATTRIBUTES block, and one fielded record per artefact carrying an id, a kind from data/policy.json::artefact_kinds, a readable dated and a body. ⚠︎ THE WHOLE BUNDLE REACHES YOUR CONFIGURED PROVIDER VERBATIM, AND A CONTROL BUNDLE IS THE WORST-CASE DOCUMENT FOR THAT. Corpus lens →
When is this the wrong choice?Avoid: Paying per bundle for a periodicity answer that twenty lines of Python already get right nine times in ten. That is the case against the best-fitting scenario (“A pile whose artefacts already carry a clean, parseable date and a kind, and where what you need is the periodicity answer rather than the attribution -- 'does what we have cover every month' over a set somebody has already triaged.”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?REAL PILES, BECAUSE OF THE DATES. Every artefact prints a parseable date in a fielded header and src/coverage.py reads a field rather than guessing. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER ANY OF IT REPEATS. ONE paid run, ONE model, ONE day, never re-fired. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-31 — r001-control-evidence. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured this capture pass on a copy of the kit with no .env and no API key, in a scratch directory so the committed tree was never written to. ⚑ THE REPRODUCIBILITY CLAIM WAS TESTED RATHER THAN REPEATED, ACROSS FOUR PYTHONHASHSEED VALUES (0, 1, 97531, 424242). Every one rebuilt data/corpus/ (60 files), data/gold.jsonl, data/bundles.json and data/corpus-stats.json BYTE-IDENTICAL to the committed set -- zero differing files on all four. A sibling in this same batch claims a byte-identical rebuild and does not deliver one, because it slices a set; this kit does. python3 -m evals.check_labels then re-derived the whole key from data/policy.json with no src/ import in 0.03 s: 60 bundles, 208 attributes, 1,208 attachments, 411 periods, 669 artefact dates, 0 disagreements. Both are free and neither needs a key.

A living map of modern AI — kept current every morning