The business caseThe problem this solves
A food brand's approved copy specification says, line by line, what a label must print: the brand, the statement of identity, the net quantity, the ingredient and allergen statements, the warnings, the preparation, storage and disposal copy — each on its panel, some verbatim and some adaptable. The press proof comes back from the artwork house with that copy retyped, reflowed and occasionally reworded, and a reviewer reads the two side by side before anyone signs. Most differences are harmless — a capital letter, a dropped full stop, 'Store in a cool, dry place.' for 'Store somewhere cool and dry.' — and a few are not: an allergen list reordered, a net quantity restated as 1 lb for 16 oz, 'use by' for 'best by', 'refrigerate before opening' for 'once opened, keep refrigerated'. A missed one is a relabel or a recall. Reading every line of a press proof against its approved copy specification by eye — finding the printed line for each element, checking it is on the right panel and in order, comparing verbatim copy character by character once case and spacing are set aside, deciding whether reworded adaptable copy still says the same thing, and spotting copy nobody approved. It does not replace the sign-off: every HOLD goes back to the artwork owner and a person owns it.
Audience
An artwork or label-compliance reviewer reading a press proof against its copy specification before the proof goes to sign-off. And the brand owner deciding whether a model call belongs in that step at all — on this corpus the honest answer is that free code does as well. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual label-copy packets
The corpus is 62 label-copy packets, 0.67 MB (json 3 · jsonl 1 · md 2 · txt 62). Because the shape of a label-copy failure is not the diff, it is the REWORDING. 584 of the 673 lines are settled by a locator and a string comparison once LCS-2026's five allowances are applied; the 89 that matter are adaptable lines whose wording differs from the approved copy — 71 harmless and 18 changed — and a corpus without both halves measures a diff tool. data/SOURCES.md attacks it: the reworded share (35 pct of adaptable lines) is generous to the paid call, and the adaptable copy comes from a closed pool of 16 approved meanings.
The corpus
- The 62 label-copy packetsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — all 62 packets, data/specs.json and the whole answer key are generated in-process by the file that renders them, so there is no third-party data in this kit and no third-party licence to honour.
Swap this folder for your own material and the kit is pointed at your label-copy packets. That is the whole change — there is no database to migrate.
LABEL COPY CONFORMANCE PACKET
PACKET HEADER
Packet LBL-0001
Company Quillmere Foods
Product Quillmere Kitchen Penne Rigate
Package 32 oz carton
SKU QF-20027
Copy spec of record QCS-20027 revision 5, approved 2026-03-03
Proof Press proof round 1, copy transcribed from artwork file QF-20027-AW-R5-P1, one printed line per row
Received 2026-07-01
Routed by R. Tallisker, Label & Artwork Operations
APPROVED COPY SPECIFICATION
Line Panel Requirement Element Approved copy
S-01 PDP MANDATORY Brand Quillmere Kitchen
S-02 PDP MANDATORY Statement of identity Penne Rigate
S-03 PDP MANDATORY Claim Made with durum wheat semolina
S-04 PDP DECLARED Net quantity Net Wt 32 oz (907 g)
S-05 INFO DECLARED Ingredient statement Ingredients: Semolina (wheat), durum wheat flour, niacin, iron (ferrous sulfate), thiamine mononitrate, riboflavin, folic acid.
S-06 INFO DECLARED Allergen statement Contains: Wheat.
S-07 INFO ADAPTABLE Preparation Cook in boiling salted water for 10 minutes, then drain.
S-08 INFO ADAPTABLE Storage Store in a cool, dry place.
S-09 INFO ADAPTABLE Date-code location See the end of the box for the best-by date.
S-10 INFO ADAPTABLE Disposal The box is recyclable.
S-11 INFO MANDATORY Distributor statement Distributed by Quillmere Foods, 40 Millrace Lane, Ashgrove.
PRESS PROOF COPY
PRINCIPAL DISPLAY PANEL
Quillmere Kitchen
Penne Rigate
Made with durum wheat semolina
Net Wt 32oz (907 g)Abridged — the file continues.
The outcomeWhat a good result looks like
Every specification line carries a finding, the rule of LCS-2026 it rests on and, where the line does not conform, the printed line quoted as evidence; the proof carries CONSISTENT or HOLD, the held lines and any unapproved copy. After the station: 655 of 673 specification lines and 45 of 62 whole proofs on the published run.
And when it cannot
⚠︎ THE PAID CALL DOES NOT BEAT FREE CODE ON THIS CORPUS, AND THAT IS THE FIRST THING THIS PAGE SAYS. After the station it gets 655 of 673 lines; the free keyword floor gets 656 (McNemar exact p = 1.000000) and the free always-carried floor 655 (p = 1.000000). On whole proofs it gets 45 of 62; the always-carried floor gets 53 (p = 0.115318) and the keyword floor 47 (p = 0.855536). None of the four differences is significant, in either direction. The call reads 16 of the 18 real substance changes, where the always-carried floor reads none, and pays for it with 16 harmless rewordings held as changed on 15 proofs, which that floor never raises.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your artwork house pastes approved copy and your disputes are about typos, dropped allergens, restated net quantities, wrong panels and unapproved copy. — the free diff floor alone — python3 -m evals.run --floor rules
Every one of those is a locator and a string comparison once L-5 is applied. All four floors get all 584 code-settled lines right, deterministically and for $0.00. - Your proofs reword adaptable copy — storage, preparation, disposal, serving and date-location lines — and you want every reworded line looked at. — the free keyword floor, with a person on every hold
It catches all 18 real substance changes and holds 17 harmless rewordings; the paid call catches 16 and holds 16. The difference is not significant (655 lines to 656, p = 1.000000). - You want the fewest proofs sent back. — the free always-carried floor
It gets 53 of 62 whole proofs right against the call's 45, because it never holds a harmless rewording — and it passes every real substance change, so it is only safe where a person reads the adaptable copy anyway. - Your rewordings are harder than these — paraphrases that keep every number, negation and condition word. — measure the call on a corpus of them before buying it
That is the population where a reading could beat keyword code, and this corpus does not contain enough of it to show it. The kit's ask is a harder corpus, not a purchase.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-12. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point data/corpus/ at your own exported proofs and data/specs.json at your own approved copy specifications. The boundary is the ANSWER KEY, not the documents. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a reading of copy that was not reworded. That is the case against the best-fitting scenario (“Your artwork house pastes approved copy and your disputes are about typos, dropped allergens, restated net quantities, wrong panels and unapproved copy.”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A packet whose sections are not the six this parser knows. src/rules.py splits on PACKET HEADER, APPROVED COPY SPECIFICATION, PRESS PROOF COPY, PROOF ROUTING NOTES, SIGN-OFF and END OF PACKET; a missing heading yields an empty section, so a differently-shaped packet parses to zero lines and is scored as zero lines. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | A second scored run at the same tier. One run, one model; no repeat was bought, so nothing here separates run-to-run variance from the one-line and eight-proof differences against the free floors. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one provider, one key, configured in .env; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-12 — r001-label-copy-check. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured scores all four free floors over all 62 packets in pure code, re-derives the whole key from a hand-retyped rulebook with evals/check_labels.py (0 disagreements), rebuilds the corpus byte-identically, and serves the board with the model button disabled and the recorded run replayed. All of it was $0.00.





