Home › Use Cases › Check that a work package carries every control its declared hazards require
Use caseUC0213
🧪 Use-case kit · runnable

Check that a work package carries every control its declared hazards require

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A work authorisation package lands on a permit desk before a shift -- a permit form with labels down the left margin, a system export, a set of CSV blocks, or a shift handover written in sentences -- and somebody has to satisfy themselves that every control the declared hazards require is actually in place before a crew goes to work. Almost none of that is a judgement. It is a join (which controls does this hazard require), a subtraction (is this gas test still fresh at the stated start), a comparison (is this certificate still valid) and a set-membership test (does this signature cover this hazard) -- in front of prose that says "spaded and the lock is on" and "still passing and sits under the flushing certificate". The supervisor's manual walk of the package against the rulebook -- deriving the checklist from the declared hazards, subtracting each gas test's age against the stated start, comparing each certificate's expiry, reading each isolation's state and testing each signature's scope -- before a person confirms the report.

Audience

The permit office or area authority who confirms the package before work starts, and the HSE owner who answers for a job that went ahead against a control nobody applied. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual work authorisation packages

The corpus is 60 work authorisation packages, 0.06 MB (txt 60). A defect mix a permit office would recognise and no public dataset can provide: an isolation listed and left open under a flushing certificate; a package signed by two people whose authorisations cover hot work and work at height on a job that also declares energised electrical work; a valve found open on the walk-down and not re-set; a flammable reading of 8.6 pctLEL on a confined-space entry presented as ready to start against a 5.0 limit; an electrical isolation on a package that never declared electrical work. Five sixths are planted because a clean package teaches nothing about the arithmetic, and a sixth are clean because a corpus that is all traps measures a different job. ⚑ AND THE CORPUS WAS DELIBERATELY MADE HARDER MID-BUILD -- this is the clearest recorded case of it in the estate. The FIRST version of the prose renderer wrote the handover note out of the rulebook's OWN hazard labels and its three state words, and the free rules floor then scored 100.0 pct of control verdicts on all fifteen prose packages and 99.4 pct OVERALL. A corpus whose "sentences" are field labels with punctuation around them cannot tell a reader from a regex, and a kit built on one has not measured a reader. The notes were rewritten into ordinary language -- "a crane lift", "we are going inside the vessel", "spaded and the lock is on", "still passing and sits under the flushing certificate", none of it in data/rulebook.json -- and the floor fell to 23.4 pct on the prose format and 80.2 pct overall. BOTH NUMBERS ARE PART OF THIS KIT'S EVIDENCE: 99.4 pct is what the floor scores against a corpus that separates nothing, and 80.2 pct is what it scores against this one. The separation this kit reports was manufactured on purpose, and a reader is owed that before they read the 23.4-to-98.6 headline. ⚠︎ data/SOURCES.md records the post-rewrite floor as 78.2 pct in its prose while the shipped result file and the README both measure 80.2 pct (449 of 560). The 80.2 pct is the measured figure and is what this spec publishes; the 78.2 in SOURCES.md is stale and is recorded here rather than silently reconciled.

The corpus

  • The 60 work authorisation packagesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your work authorisation packages. That is the whole change — there is no database to migrate.

One work authorisation package, as the model receives itPK-0001.txt · 1 of 60
PRE-JOB BRIEFING NOTE  --  PK-0001

Job on Alderoak Gas Plant, crude storage tank T-104: internal shell inspection after the wash. On the hazard side we have a break-in on a system that holds pressure and work in the amine. The authorisation runs from 08:00 on the 17th of September 2026 and it closes at 18:00 on the 17th of September 2026; all times in this note are UTC.

On isolations: the pressure point at PR-474 is isolated and the lock is hung.

No gas test is written on the sheet.

Paperwork on file: the depressurisation record is DP-3516; the substance data sheet is SD-2224.

Signatures: H. Dumas (responsible electrical person) signed at 07:50 on the 17th of September 2026 and is authorised for a break-in on a system that holds pressure and there is a hazardous substance in the line.

Crew are due on the job at 10:00 on the 17th of September 2026.

The outcomeWhat a good result looks like

A pre-start report a person confirms instead of a permit a person re-reads: one row per control the declared hazards require, in the rulebook's fixed order, each SATISFIED, MISSING, EXPIRED or INCONSISTENT, each failure naming the rule id it breaks -- R-ISO-02, not "isolation problem" -- and one sentence saying what decided it. may_start is false when any control is not SATISFIED.

And when it cannot

A FALSE CLEAR: a control that is missing, stale or contradicted, reported SATISFIED. Work starts against a control nobody applied. Every other mistake this kit can make delays a job; this one releases it, which is why it is scored on its own denominator and never averaged. On this corpus it is 0 of 67 in ALL THREE arms -- model RAW, model RECHECKED and the free rules floor -- and 0 of 50 at the package level (FALSE RELEASE). The most expensive error direction is not where the arms differ, and a page that leads with the 97.3 pct is leading with the wrong number.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Packages that come out of a permit system as labelled fields -- a permit form, a CSV export, a checklist export — the free rules floor alone
    415 of 415 control rows and 45 of 45 packages, for $0.00, in 0.1 s. The model's own report on those same formats is 405 of 415 -- WORSE -- and reaches 415 only after pure code re-derives its verdicts. The best the paid arm managed on three quarters of this corpus was a tie.
  • Shift handovers, pre-job briefings, anything where the hazards and the isolation states are sentences rather than fields — the model, rechecked
    23.4 pct to 98.6 pct on 145 control rows and 0 to 11 of 15 packages. The floor recovers 7 of 35 declared hazards across the fifteen notes and never builds the rows those hazards require: 93 rows go unreported, 16.6 pct of the corpus and 64.1 pct of the rows the prose packages carry -- including seven failing controls nobody is ever told about. That row, seven packages wide, is this kit's whole case for spending anything.
  • A mixed queue -- some packages from the system, some typed up by a supervisor — the floor first, the model only on what the floor cannot declare
    The floor's failure mode is legible and cheap to detect: its DEC-CONS row fails on 13 of the 15 prose packages (11 of them wrongly) and on none of the 45 labelled ones, and on the 9 where it recovered no hazards at all it says so in as many words -- "the package declares no hazards at all, so there is no checklist to compare it against". That is the signature of a reader that failed rather than a package that is wrong. Route those to the model and pay for 15 packages instead of 60.

At a glanceHow the whole thing runs

97–99.6%control-verdict accuracy over the 560 checklist rows
57,235 msp50, end to end
$29.25per 1,000 work authorisation packages · Google Gemini 3 Flash

Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own packages as .txt files into data/corpus/ and add one desk-record row per package to data/packages.json -- the package id, the asset, the job and THE MOMENT WORK IS DUE TO BEGIN, which is the one thing that must not be read out of the document (a package whose own window has closed still says it is valid). ⚠︎ A REAL WORK AUTHORISATION PACKAGE REACHES YOUR CONFIGURED PROVIDER VERBATIM. Corpus lens →
When is this the wrong choice?Avoid: Paying $0.006189 a package for a job a csv reader and a join already do perfectly. This kit said that before the run and the run did not soften it. That is the case against the best-fitting scenario (“Packages that come out of a permit system as labelled fields -- a permit form, a CSV export, a checklist export”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A REAL HANDOVER NOTE. The fifteen prose packages come from ONE renderer with small phrase pools, and their difficulty is a decision this kit made -- the first version of that renderer used the rulebook's own words and the floor scored 100 pct on all fifteen. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE PAID ARM IS WORTH ANYTHING ON LABELLED FORMATS. It is not, on this corpus: the floor is 415 of 415 and the model ties it at best. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-31 — r001-worksafe-precheck. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — ⚠︎ NOT RE-MEASURED FOR THIS SPEC -- stated from what ships in the repo, not from a fresh keyless clone timed on this machine. What is on disk and verifiable without a key: data/corpus/ (60 packages), data/packages.json, data/rulebook.json, data/gold.jsonl, data/corpus-stats.json, the free rules-floor run (results/eval-b000-worksafe-precheck-rules.json, 449 of 560 control rows, 0.1 s wall, usd null), the stub run (results/eval-t000-worksafe-precheck-stub.json), the scored paid run and its per-call cache (results/eval-r001-worksafe-precheck.json and cache-r001-worksafe-precheck.jsonl). requirements.txt names NOTHING -- the kit is Python standard library end to end -- and .env is gitignored from the first commit, so the repo has never held a credential. python3 -m tools.build_corpus, python3 -m evals.check_labels, python3 -m evals.run --floor rules and python3 -m src.app are the four free commands the README publishes. ⚠︎ .env.example still says "NO PAID ARM IS COMMITTED FOR THIS KIT"; that line predates r001 and is stale, and it is recorded here rather than repaired from this side of the fence.

A living map of modern AI — kept current every morning