The business caseThe problem this solves
One product is sold into several regions and each region gets its own safety data sheet, maintained by its own office, revised on its own schedule. They are supposed to describe one product. Over a few revisions they stop agreeing — a category on one sheet and not the other, a hazard statement present in one printing and absent in another, an exposure limit stated in a different unit. Nobody notices until somebody reads two sheets side by side, and the reading is sixteen numbered sections against sixteen numbered sections, three times over. Reading one product's regional safety data sheets side by side, numbered section against numbered section, and writing down which ones disagree and how. It does not replace deciding what to do about a divergence, and it never says which sheet is right.
Audience
A product-stewardship author or a regulatory-affairs coordinator keeping one product's regional sheets in step, and the person who has to decide what to do about a divergence once it is found. This kit produces the FILE they read. It never makes the decision. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual regional variant sets
The corpus is 63 regional variant sets, 0.21 MB (txt 63). It is generated because it has to be. A real safety data sheet is a supplier's controlled document carrying a named author, an emergency contact and a substance's actual classification, and publishing three regional printings of one would be republishing somebody's regulatory filing. More than that: a wrong hazard number that LOOKS real is the failure mode this whole subject has, so the corpus is built so that the SHAPES are real and every identifier is in a namespace nothing uses. That is also what makes the answer key derivable — the generator gives every cell a value id before any text exists, so equivalence is structural rather than a judgement somebody typed.
The corpus
- The 63 regional variant setsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromNowhere — every one of the 63 variant sets and the whole answer key is generated. Nothing here is fetched, scraped, licensed or derived from anything that was.
Swap this folder for your own material and the kit is pointed at your regional variant sets. That is the whole change — there is no database to migrate.
PRODUCT SAFETY DATA SHEET - REGIONAL VARIANT SET
SET HEADER
Product record PR-0001
Product name Selmack CR chain lubricant
Variants CA, EU, US
VARIANT CA
Region CA
Sheet reference SDS-SEL-3885-CA
Product code SEL-3885
Supplier Perrindale Coatings
Revision date 2026-03-07
SECTION 2 Signal word DANGER
SECTION 2 Hazard classes Hazardous to the aquatic environment chronic, Category 3; Skin irritation, Category 2
SECTION 2 Hazard statements HZ-214, HZ-227, HZ-228, HZ-239
SECTION 2 Pictograms PG-04
SECTION 3 Component concentration SIN-83341-5 at 20 - 40 %
SECTION 8 Exposure limit SIN-83341-5 at 20 ppm (15 min STEL)
SECTION 9 Flash point 41 degC
SECTION 9 Boiling point 41 degC
SECTION 9 pH 10.2
SECTION 14 Transport hazard class Class 8
SECTION 14 Packing group I
SECTION 14 Proper shipping name Flammable liquid, not otherwise specified
SECTION 15 Inventory status Listed on the regional inventory with a use restriction
VARIANT EU
Region EU
Sheet reference SDS-SEL-3885-EU
Product code SEL-3885
Supplier Perrindale Coatings
Revision date 2026-03-07
SECTION 2 Signal word DANGER
SECTION 2 Hazard classes Hazardous to the aquatic environment chronic, Category 3; Skin irritation, Category 2
SECTION 2 Hazard statements HZ-214, HZ-227, HZ-228, HZ-239
SECTION 2 Pictograms PG-04
SECTION 3 Component concentration SIN-83341-5 at 20 - 40 %Abridged — the file continues.
The outcomeWhat a good result looks like
One product's variant set in, one comparison report out: every comparable row with the first rule of SDSC-2026 that reaches it, the list of variants that stated a value, both sides quoted verbatim out of their own sheets, the sections in review and one status. 62 of 63 reports came back completely correct after the pure-code station, and 44 of the 45 rows whose answer rests entirely on a reading.
And when it cannot
And what it does when it cannot. Zero of the 63 calls failed to return a report — every reply parsed and none reached the ceiling — but 12 of the 205 quotations SDSC-2026 owes were simply not returned, and one report of 63 is still wrong after the station: a section 9 pH row printing 10.2 in two variants and pH 10.2 in the third, read as a content divergence, which turned an ALIGNED set into a REVIEW one. Confidence does not separate right from wrong here: the median on a correct report is 1.00 and the median on a wrong one is also 1.00.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your regional sheets are already in a structured system and differ only in punctuation, spacing or list order — src/rules.py's S-4 normaliser, with no model at all
743 of the 788 comparable rows here are settled by it, at $0.00 and with no network. There is no measured case for spending anything on them. - Your offices write the same classification two ways — a long form here, the abbreviation there — the paid call, and it is the one case this kit demonstrates
44 of the 45 rows where S-4 cannot decide came back right, against 30 for the strongest free arm. McNemar's exact test over the 63 paired sets: 12 the model gets and the floor does not, 0 the other way, two-sided p = 0.00049. - You maintain a controlled vocabulary of the printings your own offices use — a lookup table, and measure it before buying anything
data/SOURCES.md measures exactly this on this corpus: a table of the generator's own 30 printing pairs settles 41 of the 45 reading rows. It was not built and is not shipped, but if YOUR wordings come from a list, that list is the cheapest arm you have. - Your variants are in different languages — unknown from this kit — the paid arm is the candidate and it is NOT demonstrated here
every sheet in this corpus is in English. A translated classification is the same equivalence judgement in a harder form and nothing here measures it.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-08. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own variant sets in the same panel shape — a SET HEADER, one VARIANT <REGION> panel per sheet with SECTION <n> <field> <value> rows, a COMPARABLE ROWS list and SET NOTES — and re-run both floors, which need no key. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO, AND IT IS THIS KIT'S HEADLINE. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call whose answer a normaliser already produced. That is the case against the best-fitting scenario (“Your regional sheets are already in a structured system and differ only in punctuation, spacing or list order”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A real safety data sheet, which arrives as a PDF. There is no OCR, no PDF reader and no layout model in this kit; the corpus is already columnar text and getting a real sheet into that shape is work this kit does not do and does not measure. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the margin survives on real safety data sheets. Every number here is measured on 63 generated variant sets whose two printings of one classification were drawn from a 30-entry table, and data/SOURCES.md measures what that costs: a lookup built from that table settles 41 of the 45 reading rows and the field name alone determines the finding on 36 of them. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-08 — r001-sds-variant. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9327 — all 63 variant sets, both free floors computed live, SDSC-2026, and every committed run replayed at $0.00. The one control that would spend is disabled and says why. That state is one of the five published frames.




