The business caseThe problem this solves
A customer sends a specification and asks whether you can make it. Answering means putting two documents written by two different organisations onto one scale -- what the customer will accept, property by property with a tolerance either side, against what the process has actually been demonstrated to deliver, property by property with the grades each range holds on -- and then saying yes, no, yes on these grades, or 'we have never measured that'. Almost none of it is judgement: it is a unit conversion, an interval containment test and a set intersection. What IS judgement is small and specific -- is 'Water by Karl Fischer' the same determination as 'Water content' (yes) and is 'Particle size D90' the same as 'Particle size D50' (no); and does a band the table prints as holding on 'all' grades actually hold on all of them, when a sentence in the technical notes says otherwise. Reading the two documents side by side in a spreadsheet: paste the enquiry into a column, VLOOKUP each property name against the capability sheet, compare the two pairs of numbers with an IF. That is what this job is actually done with, and it is the kit's b000 floor rather than a strawman -- it scores 122 of 300 properties and 29 of 60 enquiries on this corpus, and it fails in the two directions above 38 and 20 times.
Audience
Technical service and quality in a process plant -- whoever answers a customer's specification enquiry and has to defend the answer afterwards, plus the commercial lead who signs the acknowledgement. The reader who matters is the one who has to say no to a customer and show why, and the one who has to say yes and be right. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual capability reviews
The corpus is 60 capability reviews, 0.17 MB (txt 60). Because the two failures this use case has are opposite and both invisible, and only a labelled set separates them. An order refused for no reason leaves no trace anywhere -- nobody audits the enquiries that were turned away -- and a promise the plant cannot keep does not surface until the first rejected delivery. The corpus therefore plants both directions deliberately, in counted numbers, and plants the class the whole kit turns on -- NOT_CHARACTERISED, 28 cells, with 15 near-miss determinations printed beside them so that 'we have never measured it' and 'we measure something that reads like it' can be told apart at all.
The corpus
- The 60 capability reviewsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your capability reviews. That is the whole change — there is no database to migrate.
Specification Capability Review
-------------------------------
SYNTHETIC RECORD. Every customer, product family, grade, capability band and technical note
below was invented for this kit. No real product, company, published standard, test method
designation or capability sheet was used, reproduced or approximated.
Row formats. A Requested Specification line is `RS-nn <property, as the customer writes it>
target <value> tol -<value> / +<value> unit <unit>`; `open` on a side means the customer
sets no limit there. A Capability Sheet line is `CB-nn <property, as this plant names it>
<low> to <high> <unit> grades <grade identifiers, or all>`, and that range is what has been
demonstrated in production -- not a target and not a single best batch.
Enquiry
-------
Enquiry SCR-0001
Customer CUST-MIRENDAL
Product family PFM-EPXRES Epoxy resin, liquid
Capability sheet CAP-EPXRES-R2
Received 2026-01-09
Requested Specification
-----------------------
RS-01 Flash point, closed cup target 318.25 tol -8.40 / +8.40 unit K
RS-02 Chloride content target 61667 tol -open / +12333 unit ppb
RS-03 Acid number, mg KOH per g target 1.24 tol -open / +0.24 unit mgKOH/g
RS-04 Density at 20 C target 1.1490 tol -0.0034 / +0.0034 unit g/cm3
RS-05 Viscosity at 25 C target 2252 tol -313 / +313 unit mPa.s
Capability Sheet
----------------
CB-01 Acid value 0.16 to 1.14 mgKOH/g grades all
CB-02 Water content 0.264 to 0.316 % grades allAbridged — the file continues.
The outcomeWhat a good result looks like
A determination a person confirms rather than trusts. Every requested property carries one of four verdicts, the capability entries that decided it named by their printed identifiers, and the grades the answer holds on; the whole enquiry carries its own answer, the reason for it and the binding grade where there is one. The arithmetic is re-runnable off the printed page without the model, and three guardrails convict a misreading with no answer key at all.
And when it cannot
Two failures cost money and they cost it in opposite directions. A band called WITHIN when it is OUTSIDE is a promise the plant cannot keep -- the order is taken, the acknowledgement is in writing, and the bill arrives months later as rejected deliveries, sorting or a claim. A band called OUTSIDE when it is WITHIN is an order refused for no reason, and it is INVISIBLE: nobody audits the enquiries that were turned away. The third, quieter one is a condition dropped -- a specification quoted on a grade that cannot hold it, which reads as a yes right up to the first delivery.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Sheets and enquiries that already share one vocabulary and one unit system -- a customer quoting against your own published datasheet, or a framework where the property names are fixed before anyone asks — the free
unitsfloor alone
With no aliases and no notes-channel conditions in play, the units floor is doing the whole job: exact conversion, containment, every matching entry, the printed grades column and the grade intersection, in 0.1 s for sixty reviews at $0.00. On this corpus it already scores 171 of the 192 same-name properties. - Two organisations writing the same determination in their own words -- a customer's house standard against a plant's own sheet — the model, and this is the clearest case for it
Both shipped floors score 0 of 80 aliased properties and the model scores 79. ⚠︎ READ THE CAVEAT WITH IT: a lookup built from the generator's own 51-string synonym table also scores 77, so what is measured here is 'resolves a synonym from a closed trade vocabulary', not 'resolves a synonym nobody has ever tabulated'. On a real plant's sheet, whose wording nobody outside it uses, this is the row that would move most. - Sheets whose grade restrictions live in prose -- a grades column printing
allwith a technical note saying otherwise — the model, with a free notes reader as the control
Both shipped floors score 0 of 24 and the model scores 24. But a five-line regex -- one CB- entry plus at least one GRD- grade in the same sentence -- scores 21 of 24 free, and lifts grade conflicts from 1 to 6 of 10. Build that first and measure what is left before buying the call.
And where nothing here is good enough:
- An enquiry whose value turns on ONE property, where the difference between 'we have never measured it' and 'your window is tighter than our spread' decides whether a trial is commissioned — neither arm unsupervised
The kit's single miss is exactly this shape, and the guardrails passed it. A wrong match that is internally consistent is invisible without an answer key, and on a one-hard-property enquiry it flips the whole answer.
At a glanceHow the whole thing runs
Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point it at your own capability sheet and one real enquiry, in the row formats the documents themselves state at the top of every review. ⚠︎ THE WHOLE REVIEW LESS ONE SECTION GOES TO YOUR CONFIGURED PROVIDER VERBATIM, and on this use case that includes the capability sheet -- every demonstrated band, every grade restriction, every technical note. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call and waiting 31 seconds to re-derive a comparison a Fraction and an IF already give you correctly. That is the case against the best-fitting scenario (“Sheets and enquiries that already share one vocabulary and one unit system -- a customer quoting against your own published datasheet, or a framework where the property names are fixed before anyone asks”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A real datasheet, which is the honest first entry. Every line here prints its unit in a labelled unit <x> field, every band has two ends, and every property is spelled one way per document. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | The adversarial arm. evals/injection.py ships complete and HAS NEVER BEEN FIRED. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-31 — r001-spec-capability. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Everything except the paid arm runs on a clean checkout with no key and no network, and was timed that way this turn: rebuild the whole corpus and answer key 0.05 s, gate the key with evals/check_labels.py 0.04 s (KEY CLEAN, 60 reviews, 300 properties), score the b000 rules floor end to end 0.10 s. The board serves without a key and disables the one button that spends money. Total cost of everything a forker can verify: $0.00.



