The business caseThe problem this solves
A submittal review starts with a manufacturer's data sheet on one side and a specification section on the other, and the reviewer joins them by hand. The section numbers its requirements; the sheet is the manufacturer's own table, under the manufacturer's own property labels, often covering a family of models the project is not buying, with units the section does not use and a test standard's edition in a footnote. Every requirement has to be found on that sheet, converted, compared, and written down with the row it turns on — and the ones that are NOT on it have to be written down too, because a requirement nobody answered is the one that reaches the field. Sixty submittals into a package, at the end of a week, the requirement that quietly does not get looked up is the deviation that gets built. The manual join between a specification section and a manufacturer's data sheet, and the typing of the requirement table that comes out of it. It does not replace the review: the engineer of record still reads every deviation, still rules on every substitution, and still stamps.
Audience
the engineer of record and the technical reviewers who prepare the submittal register for them, on a design team or in a CM office Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual product data submittals
The corpus is 62 product data submittals, 0.14 MB (json 3 · jsonl 1 · md 2 · txt 62). Because the reading has to be separable from the arithmetic, and on a real submittal corpus it is not: you cannot tell whether a reviewer found a datum because they read the sheet or because the label happened to match. Here the generator holds both facts and the key is derived from them, so reading_required is MEASURED — 24 of 309 requirements, the difference between the key and the same engine run over the exact-label pass — rather than asserted. That is also its limit: it is this generator's idea of how a manufacturer renames a property.
The corpus
- The 62 product data submittalsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md carries the whole attack on this corpus: what a synonym table written with the answer key open scores on it (309 of 309 requirement verdicts, which is everything), how much of that is memorisation (11 of the 19 relabelled requirements carry a sheet label seen ONCE, and 4 of 40 synonym pairs survive a mechanical de-memorisation), and the three constructions this generator makes that a real corpus would not — exact-only unit conversions, an identification panel that always carries the applicability field, and one planted row per citation.
Swap this folder for your own material and the kit is pointed at your product data submittals. That is the whole change — there is no database to migrate.
PRODUCT DATA SUBMITTAL - SPECIFICATION SECTION COMPARISON
SUBMITTAL HEADER
Submittal SB-0001
Project Harrowgate Civic Centre
Specification Section 07 21 16 - Blanket Insulation
Revision Rev 0, submitted for review
Transmitted 2026-01-01
Contractor Meridian Construction Company
PRODUCT IDENTIFICATION (what is being submitted)
Manufacturer Alpha Insulation Systems
Product Thermaline AX series
Model submitted AX-20
Location exterior
Substrate gypsum board
SPECIFICATION REQUIREMENTS (Section 07 21 16, as issued)
# Ref Requirement
1 2.01 A Acceptable manufacturers: Alpha Insulation Systems, Bravo Thermal Products, Corveth Mineral Wool. Substitutions: a written request naming the specified product it replaces and stating a basis of equivalence.
2 2.02 B.1 Thermal conductivity: maximum 0.040 W/m.K
3 2.02 B.2 Board density: minimum 38 kg/m3
4 2.02 B.3 Facing tensile strength: minimum 0.55 kN
Unit conversions for this section: 1 mm = 1000 micron; 1 cm = 10 mm; 1 MPa = 1000 kPa; 1 kPa = 1000 Pa; 1 kg/m3 = 1 g/L; 1 kN = 1000 N; 1 L/s = 1000 mL/s; 1 kg = 1000 g; 1 W/m.K = 1000 mW/m.K; 1 h = 60 min.
MANUFACTURER PUBLISHED DATA (Alpha Insulation Systems, Thermaline AX series)
Property AX-20 AX-50 Test method
Thermal conductivity 0.038 W/m.K 0.038 W/m.K ASTM C518 2021
Board density 40 kg/m3 46 kg/m3 ASTM C303 2020
Facing tensile strength 0.57 kN 0.69 kN ASTM C1136 2019Abridged — the file continues.
The outcomeWhat a good result looks like
One requirement table per submittal: every numbered requirement of the section with a verdict, the clause it rests on, the signed margin in the section's own unit where there is one, and one row copied verbatim out of the submittal as evidence — plus the deviation list and, separately, the requirements the sheet never answered. It is what the engineer of record reads FIRST. It never approves, rejects or stamps anything.
And when it cannot
The two directions cost different things and this kit never averages them. A requirement the sheet answers, reported as a deviation, is a FALSE ALARM: a contractor sent back for data that is already in front of the reviewer, and a review cycle nobody needed — measured at 7 of 309 here, and at 18 of 309 on the free exact-label floor, which is the direction free code fails in. A requirement that is short, silent or off-list reported as MET is the opposite and the expensive one, because it puts a non-compliant product on the approved pile — measured at 0 of 309 on every arm of the scored run.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- a package of submittals against sections you have already numbered — the pure-code floor first, then decide whether to pay
the exact-label pass settles 285 of 309 requirements here for nothing, and it is the half of this job that is genuinely free - sheets that use the manufacturer's own vocabulary — the model, with the pure-code station on top
this is the only thing on the page a label match cannot do — 15 of 19 relabelled requirements rechecked, against the exact-label floor's 0 - a decision about whether the product may be installed — a person — the engineer of record
S-10 is a clause, a schema shape and a phrase list, and none of them is a substitute for a review. This produces the table the review starts from.
And where nothing here is good enough:
- units that are not exact decimal shifts of each other — neither arm, until you have widened the conversion table
S-6 refuses an inexact conversion and sends the requirement to NOT-ADDRESSED. That is correct and it will fire on every inch, every Fahrenheit and every R-value.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace tools/sections_data.py with your own sections and rebuild, or point src/packet.py at your own text files. The key does not transfer and neither does the phrase floor's synonym table — it was written against these forty-eight labels. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for the whole package before measuring how many of your requirements a label match already reaches. That is the case against the best-fitting scenario (“a package of submittals against sections you have already numbered”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A submittal that is a PDF. Everything here assumes the text is already text; the published-data table is a fixed-width layout and a scanned table is a picture. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | The USD of x001-submittal-compare. The escalated probe recorded output tokens only, so its cost is not measured and is not in any cost table here; evals/injection.py records both sides from now on. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one OpenAI-compatible endpoint, configured in .env and reached only through src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, reasoning disabled (THE PUBLISHED RUN). Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-09 — r001-submittal-compare. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Clone, python3 -m src.app, open the printed URL. No key, no install, no index, no network: the 62 submittals, the section requirements, the manufacturer tables, all four floors and the committed run are all in the repository and all render. python3 -m evals.check_labels and python3 tools/build_corpus.py --check both run offline and both pass.




