Home › Use Cases › Verify a store-week of safety checklists against the required daypart schedule
Use caseUC0461
🧪 Use-case kit · runnable

Verify a store-week of safety checklists against the required daypart schedule

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A brand standard places required food-safety checks in named dayparts of every trading day, each with a declared clock window and a declared acceptable range, and a store logs them on a terminal through the week. Somebody then has to say whether the week is actually complete -- and the store's own weekly completion report cannot, because it counts entries rather than reading the schedule: it disagrees with the standard on 30 of these 64 store-weeks. Today that verification is a person opening the log, the schedule and the notes side by side, per store, per week. The manual read-through of one store-week: expanding the schedule over the trading days, joining each required slot to the entries logged under it, comparing a printed clock time with its window, a reading with the range row in force, and reading the week's notes for what was taken out of the requirement. It replaces NONE of it safely on this corpus, and the page says so: free code that reads dates does the same job better.

Audience

The area desk that receives a week of completed checklists and has to decide which stores to ask about what. The decision is which findings to raise, not whether anybody is in trouble -- this kit never gets near that. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual store-weeks

The corpus is 64 store-weeks, 0.52 MB (json 4 · jsonl 1 · md 2 · txt 64). A store-week is the smallest unit on which the question can be asked at all: a daypart that did not trade, a unit off the line for two days and a run of identical readings continued from last week are all facts about a WEEK, not about an entry. The distribution is designed rather than sampled -- sixteen planted cases, four weeks carrying two findings so the ladder's order is evidenced, three weeks whose readings sit exactly on an inclusive bound, and four date traps whose notes name a day outside the week. 30 of the 64 weeks have a store completion status word that disagrees with BSS-2026, so copying the incumbent is measurably not enough; 53 carry at least one note that decides nothing, so reading every note is not enough either.

The corpus

  • The 64 store-weeksgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdoc_count and bytes are the 64 graded store-week packs under data/corpus/ (549,353 bytes, p50 8,566, p95 8,857, max 8,910). formats counts every file the kit ships under data/, which is those 64 .txt packs plus the store register, the rulebook, the answer contract and the corpus statistics (.json), the answer key (.jsonl), and the rendered rulebook and the provenance note (.md). ⚠︎ THE CORPUS WAS REBUILT ONCE, because free code solved 61 of 64 on the first draft: every note named its unit by register code and its day by ISO date. Three of the five note kinds now share ONE phrasing library between the weeks that decide something and the weeks that do not, and every note refers to a unit, day, daypart or terminal in one of three styles drawn from its own counter, so style correlates with nothing. data/SOURCES.md §2 is the write-up.

Swap this folder for your own material and the kit is pointed at your store-weeks. That is the whole change — there is no database to migrate.

One store-week, as the model receives itFSW-0001.txt · 1 of 64
================================================================================
FOOD-SAFETY CHECKLIST VERIFICATION -- ONE STORE, ONE TRADING WEEK
================================================================================
FILE              FSW-0001
BRAND             Coppergate Kitchens
STORE             ST-0400  drive-thru, area R-1
WEEK              2026-06-01 to 2026-06-07
STANDARD          BSS-2026 rev 4
LOG SOURCE        line terminal LT-1 (export taken 2026-06-08 07:05)
PACK COMPILED     2026-06-09

-- DECLARED DAYPART SCHEDULE (BSS-2026, this store's trading pattern) ----------
DAYPART   OPENS  CLOSES  REQUIRED CHECKS
OPENING   05:30  10:30   SANITISER-PPM, COLD-HOLD-TEMP
MIDDAY    10:30  15:30   HOT-HOLD-TEMP, COLD-HOLD-TEMP
CLOSING   19:00  23:30   HOT-HOLD-TEMP, SANITISER-PPM

-- STORE UNIT REGISTER (the unit each check is taken on) -----------------------
UNIT    DESCRIPTION                          SERVES
EQ-S1   sanitiser dispenser, dish station    SANITISER-PPM
EQ-C1   cold rail, make line                 COLD-HOLD-TEMP
EQ-H1   hot well, front counter              HOT-HOLD-TEMP

-- DECLARED ACCEPTABLE RANGES (BSS-2026 table 3 and every bulletin printed on it) ---
ROW         CHECK            SCALE      LOW     HIGH  EFFECTIVE   SOURCE
RG-0400-1   HOT-HOLD-TEMP    degF     135.0    175.0  2026-01-01  BSS-2026 table 3
RG-0400-2   COLD-HOLD-TEMP   degF      33.0     41.0  2026-01-01  BSS-2026 table 3
RG-0400-3   SANITISER-PPM    ppm      200.0    400.0  2026-01-01  BSS-2026 table 3

-- REQUIRED SLOTS FOR THIS WEEK (the schedule above, expanded) -----------------
SLOT      DAY         DAYPART   CHECK             WINDOW
RQ-40000  2026-06-01  OPENING   SANITISER-PPM     05:30-10:30
RQ-40001  2026-06-01  OPENING   COLD-HOLD-TEMP    05:30-10:30

Abridged — the file continues.

The outcomeWhat a good result looks like

Every required slot of the week either satisfied, waived by a note the pack itself carries, or named as a finding with the schedule line it fails; and one BSS-2026 verdict, the first rule in W-1..W-5 order the week matches.

And when it cannot

When it cannot, it is wrong quietly rather than loudly: it returns an empty waiver list and the pure-code station faithfully reports slots NOT-RECORDED with full confidence. There is no 'I am unsure' state in the answer contract, and the honest reading of that is in Eval.could_not_verify rather than in a confidence number -- the arm's own median confidence is 0.93 where it is right and 0.90 where it is wrong, a separation of three hundredths.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your terminal exports a status field for waivers, re-takes and rescinded bulletins — free code -- the vocab_dates arm's shape, or less
    all three readings become column lookups and the model has nothing left to read. That is close to what this corpus already shows: 53 of 64 free against 42 paid.
  • Your waivers and re-takes arrive only as free-text notes, in wordings that change — the call, and measure it against a word-list floor before you believe it
    the one column this kit's call wins is the prose-decided one: a note that confirms a bulletin while containing the word 'rescinded'. A word list cannot do that and a regex written against today's wordings will break on tomorrow's.
  • You need the arithmetic right every time -- the join, the window to the minute, the range to the tenth, the ladder — src/rules.py, on every arm, free
    it is deterministic, it costs nothing, it never has a bad day, and on this corpus it beats the call on every structured field.
  • You want one number for the board — whole store-weeks, both arms, side by side
    a conjunction over six fields is the only figure that cannot be gamed by picking a flattering column, and it is what says 42 against 53.
  • You want to know whether a finding is safe to send to a store — the RAW and RECHECKED columns together
    the station can drop an id that is not on the pack and cannot notice that a plausible id list is the wrong one. 6 store-weeks were reported NOT-RECORDED with full confidence because a waiver was missed.
  • You want to re-check a published figure — API_KEY= python3 -m evals.run --run-id r001-safety-checklist --resume --rescore
    every grader is pure code, so every percentage on this page re-derives from the committed replies at $0.00, with no key and no network.

At a glanceHow the whole thing runs

4–6%rechecked whole pack correct, of 64 store-weeks
1,822 msp50, end to end
$4.01per 1,000 store-weeks · the fast tier

Run once, for real, on 2026-09-14. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own store-weeks in the same seven printed blocks, data/stores.json with your register, and data/standard.json with your own standard -- then regenerate data/standard.md and data/corpus-stats.json with python3 tools/build_corpus.py. EVERY NUMBER ON THIS PAGE STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Buying a call per store-week. That is the case against the best-fitting scenario (“Your terminal exports a status field for waivers, re-takes and rescinded bulletins”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A store with more than one unit serving the same check. The register lists one per check and the slot-to-entry join assumes it; a second unit's entry reads as off-register. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether this result holds on a real chain's checklists. The 64 store-weeks are a designed distribution from one seed, and the floors are what these five free arms score on THIS corpus. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-14 — r001-safety-checklist. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board and scores all five free arms offline: python3 -m evals.baseline re-derives every floor over the 64 packs in 0.3 seconds with no network and no install, and agrees with the generator's own --check-fillers line at 0 differences. The board on port 9461 renders every recorded arm from the committed result files and disables the live-call button. requirements.txt installs nothing.

A living map of modern AI — kept current every morning