Home › Use Cases › Does the container label agree with its safety data sheet
Use caseUC0175
🧪 Use-case kit · runnable

Does the container label agree with its safety data sheet

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A container label and its safety data sheet are two renderings of one decision, and they drift. The sheet is revised when a component changes and the artwork is not reprinted; the artwork is reprinted for a brand refresh and the hazard panel is carried over unchanged; a translated label loses a statement in the reflow; a reformulation crosses a concentration cut-off and moves the classification while both documents sit still. Somebody downstream reads one of them and acts on information the other one contradicts. ⚑ MOST OF THAT IS A SET DIFFERENCE AND CODE IS EXCELLENT AT IT -- this kit's own free floor scores 91.11 pct on the whole decision for nothing. The part that is not a set difference is the pair where BOTH DOCUMENTS AGREE WITH EACH OTHER AND BOTH ARE WRONG, and the pair where the newer document is the stale one. A person putting a label proof and a safety data sheet side by side and reading six things across, then reading the revision register to work out which of the two is behind. ⚠︎ IT REPLACES THAT ONLY WHERE A SCRIPT DOES NOT ALREADY. Four of this kit's six elements are code lists, and differencing them is free.

Audience

The person who signs the artwork release before a print run -- a product steward, a regulatory affairs coordinator, a packaging technologist. Also anybody deciding whether to build this at all, because the honest answer on most of these elements is a Python script. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual label / safety data sheet pairs

The corpus is 45 label / safety data sheet pairs, 0.45 MB (txt 45). The task needed a labelling scheme, and there were two ways to get one. Assert a real one as fact -- which measures how well a model remembers that system, publishes regulatory claims nobody here can stand behind, and lets a model that memorised a real register pass without reading anything. Or invent one, print it in the pack, and measure whether a STATED scheme is applied consistently across two documents. This kit takes the second, and MHCS-4 is the result: three signal words, 16 hazard statements, 11 precautionary statements, 7 pictograms and 9 concentration cut-off rules, all made up, all printed in every pack. The distribution is a decision, not a seed: 45 explicit recipes, 79 of the 270 cells DISAGREE, 10 are INSUFFICIENT_EVIDENCE, and 24 of the disagreements are a classification drift declared only in prose. 5 pairs are fully consistent, because a checker that cannot pass a clean pair is a queue rather than a check.

The corpus

  • The 45 label / safety data sheet pairsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your label / safety data sheet pairs. That is the whole change — there is no database to migrate.

One label / safety data sheet pair, as the model receives itLS-0001.txt · 1 of 45
Label SDS Consistency Check
---------------------------
  Pack                        LS-0001
  Product                     Torvane Degreaser 75
  Product code                PRD-4364
  Site                        Ardsley Works
  Checked                     2026-08-26
  Scheme                      MHCS-4 (synthetic -- invented for this kit)

Scheme Extract
--------------
  Meridian Hazard Communication Scheme, Edition 4. INVENTED FOR THIS KIT. It is not a
  real standard and none of it is anybody's published labelling rule.

  Signal tiers, lowest to highest       NOTICE  <  ALERT  <  SEVERE

  Hazard statement register
  MH-224  Extremely flammable liquid and vapour                   HC-FLM  1  SEVERE   SC-4.1.1
  MH-225  Highly flammable liquid and vapour                      HC-FLM  2  SEVERE   SC-4.1.2
  MH-226  Flammable liquid and vapour                             HC-FLM  3  ALERT    SC-4.1.3
  MH-271  May cause fire or explosion, strong oxidiser            HC-OXI  1  SEVERE   SC-4.2.1
  MH-272  May intensify fire, oxidiser                            HC-OXI  2  ALERT    SC-4.2.2
  MH-301  Fatal if swallowed                                      HC-ATX  1  SEVERE   SC-4.4.1
  MH-302  Toxic if swallowed                                      HC-ATX  2  SEVERE   SC-4.4.2
  MH-305  Harmful if swallowed                                    HC-ATX  4  ALERT    SC-4.4.4
  MH-314  Causes severe skin burns and eye damage                 HC-COR  1  SEVERE   SC-4.3.1
  MH-315  Causes skin irritation                                  HC-COR  2  ALERT    SC-4.3.2
  MH-318  Causes serious eye damage                               HC-COR  1  SEVERE   SC-4.3.3  withdrawn 2026-02-16
  MH-334  May cause allergy or asthma symptoms if inhaled         HC-SEN  1  ALERT    SC-4.5.1

Abridged — the file continues.

The outcomeWhat a good result looks like

One finding per element, six elements per pair: AGREE, DISAGREE with a named ground, or INSUFFICIENT_EVIDENCE naming what the pack does not settle -- and on every disagreement, WHICH DOCUMENT IS OUT OF DATE, the scheme clause that carries the requirement, and the revision or component identifier that evidences it. Beside every one of them, what pure code alone would have decided.

And when it cannot

A FALSE CLEAR: an element that disagrees, called AGREE. On this use case that is a print run of containers carrying a panel the sheet contradicts. It is counted apart from every other kind of wrong and never averaged into anything: the fast tier recorded 3 of 79 (3.80 pct) on r001-label-sds; the strongest free floor recorded 24 of 79 (30.38 pct).

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Both of your documents are structured and every element you compare is a code list. A statement register, a pictogram list, a product code, a supplier block. — the free floor -- b002 scheme-recompute, $0.00
    It scores 91.11 pct on the whole finding, catches 100.00 pct of every disagreement the printed lists reveal, attributes the stale document correctly on 69.62 pct of them and invents nothing. Paying a model for this is paying for a set difference.
  • Your composition of record and your hazard panel can drift apart, and the change is recorded in a sentence somebody typed rather than in a field. — the paid arm -- but write the regex first and measure it
    This is the band a set difference cannot reach: 0.00 pct for the free floor. ⚠︎ AND IT IS ALSO THE BAND ONE TEMPLATE-MATCHING REGEX TOOK ENTIRELY ON THIS CORPUS (100.00 pct). Before paying, write the regex for your own change notes and score it. If your notes are written to a form, you have a free answer; if they are written by people, you do not, and that is the case this kit cannot measure.
  • You need to know WHICH document to reissue, not just that they differ. — either arm, but read the register rule first
    Ordering the two revision dates gets a whole class of pair backwards -- a label reprinted after the sheet, carrying its hazard panel over unchanged, is newer in date and stale in content. dated-diff scores 34.18 pct on stale attribution; reading the change note takes it to 69.62 pct for nothing; the fast tier reaches 91.14 pct.
  • You want a queue you can actually work through. — look at the over-flag rate, not the accuracy
    The fast tier flags 0 of 181 agreeing elements (0.00 pct); the free floor flags 0 of 181 (0.00 pct). An arm that raises clean pairs turns a check into a backlog, and the accuracy headline hides it.

At a glanceHow the whole thing runs

97%label sds consistency accuracy pct
106,415 msp50, end to end
$42.53per 1,000 label / safety data sheet pairs · Google Gemini 3 Flash

Run once, for real, on 2026-08-26. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own packs into data/corpus/ in the layout data/SOURCES.md describes and re-run. The scheme goes in the pack. Corpus lens →
When is this the wrong choice?Avoid: ⚠︎ DO NOT USE IT WHERE A DECISIVE FACT IS A SENTENCE. It passes every one of the 24 prose-only cells as AGREE -- a 30.38 pct false-clear rate overall -- and each of those is a pair where both documents agree with each other and neither carries what the composition requires. Nothing on the printed page looks wrong, which is exactly why nobody catches it downstream either. That is the case against the best-fitting scenario (“Both of your documents are structured and every element you compare is a code list. A statement register, a pictogram list, a product code, a supplier block.”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A real artwork proof. src/pack.py is regular expressions written for THIS corpus's fixed-column layout; point it at a PDF or an .ai file and it parses nothing and returns empty registers rather than guessing. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?ONE MODEL, ONE RUN, ONE CEILING, ONE CORPUS. No second tier was called and no arm was repeated, so there is no variance estimate here and none is implied. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-26 — r001-label-sds. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — git clone, python3 -m src.app, open the URL. No pip install, no key, no network. The corpus, the answer key and eight committed result files ship in the repo, so a fresh clone shows the real recorded run replayed off disk and all four free arms recomputed live.

A living map of modern AI — kept current every morning