The business caseThe problem this solves
A nonconformance report arrives from the shop floor or from receiving inspection. Nine panels: the part and the assembly it was found on, the register extract the part should match, whatever was measured, whatever coded finding the inspector keyed, a free-text note typed at the station, the conditions the report explicitly rules out, and the prior closed report on the same part. Someone has to say WHAT KIND of nonconformance this is, how critical the part is, and whose desk reviews it — before any of it reaches a review board. Today the coded finding is whatever the inspector keyed, and on this corpus that code names a different failure mode from the card on 26 of the 64 reports read at face value. the intake reviewer's read of nine panels per report to decide the failure mode, the criticality and the desk — it does not replace the review board, and it never proposes, records or implies a disposition.
Audience
A quality engineering lead deciding whether to put a model in front of the intake queue. This report's answer is: the classification is worth automating, and pure code written from the card alone does it as well as the paid call on this corpus — better on the complete row. The call earns its place on one family only, and that family is named. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual nonconformance reports
The corpus is 64 nonconformance reports, 0.17 MB (json 4 · jsonl 1 · md 2 · txt 64). A real nonconformance report cannot be published: it carries a manufacturer, a programme, part numbers, lot numbers and the names of the people who raised and inspected it. So the corpus is generated, and generated to exercise the thing the card is actually hard at — not volume. 16 reports state the condition only in an inspector's own sentence, 6 turn on the card's ORDER rather than on the loudest finding, 16 are unclassifiable for three different structural reasons, 39 print an assembly criticality class that DIFFERS from the part's, and every one of the 64 carries an excluded-condition panel and a prior closed report that decide the answer on some reports and decide nothing on others.
The corpus
- The 64 nonconformance reportsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your nonconformance reports. That is the whole change — there is no database to migrate.
NONCONFORMANCE REPORT NCR-0001
================================================================================================
PANEL 1 — REPORT HEADER
Manufacturer Wrenhall Aerostructures (synthetic manufacturer)
Site Site 7 — receiving
Work centre WC-405 stores intake
Raised 2026-06-12 (2026-Q2, 1 April to 30 June 2026)
Raised by Inspector, receiving inspection
Quantity affected 3 of 15
Lot LOT-26-5716
PANEL 2 — PART AS CITED ON THE REPORT
Part number as written WA-43170-005
Description as written bushing, hinge rev B
PANEL 3 — PART REGISTER EXTRACT (this manufacturer's own register)
Part Description Assembly Part class Assembly class
WA-43170-001 bushing, hinge ASM-3520 CLASS-3 CLASS-1
WA-43170-003 bushing, hinge ASM-3520 CLASS-1 CLASS-1
WA-43170-005 bushing, hinge rev B ASM-3520 CLASS-2 CLASS-1
Declared characteristic limits, by register row
WA-43170-001 hole position from nominal 0.000 to 0.250 mm
WA-43170-001 edge radius 0.60 to 1.20 mm
WA-43170-001 surface roughness 0.00 to 1.60 um
WA-43170-003 coating thickness 18.0 to 30.0 um
WA-43170-003 bore diameter 12.480 to 12.520 mm
WA-43170-003 web height 27.40 to 27.90 mm
WA-43170-005 flange thickness 4.12 to 4.38 mm
WA-43170-005 surface roughness 0.00 to 1.60 um
WA-43170-005 fastener grip length 9.50 to 10.10 mm
PANEL 4 — MEASURED CHARACTERISTICS AT INSPECTIONAbridged — the file continues.
The outcomeWhat a good result looks like
One row per report an intake reviewer could file as written: the failure mode under NCC-2026, the part as the register extract writes it, the criticality class that register row declares, the review route under NCR-2026, and one line of the report quoted verbatim as the evidence.
And when it cannot
It separates the report. A part number that is not a row of the register extract, a report with nothing measured and no evidence attached, or a note that never states a condition all return UNCLASSIFIABLE with the card's own reason and route to intake review — 16 of 64 reports, and it is a real answer here, not a failure. The paid arm got 13 of those 16 and forced 3 into a class; it also abstained on 1 report the key classifies.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- the classes and their ORDER are a written card and you have the card — the cards floor — pure code, evals/baseline.py
on this corpus it takes 50 of 64 complete rows for $0.00 and the paid call takes 47. Writing the card down as code is the cheapest thing on this page and it is also the best arm a team could ship. - the condition is stated only in an inspector's own sentence and no word list reaches it — the paid call
this is the one family where it earns its bill: prose goes 2 of 16 free to 9 of 16 paid, and all 8 reports the call wins outright are prose. - the report must be separated rather than classified — the pure-code station, src/recheck.py
the three structural overrides are exact on every arm — a part that is not a register row, and a report with nothing measured and nothing attached. All 11 of those reports are right on every arm including the modal control.
And where nothing here is good enough:
- you need the quoted evidence line to be exactly the reviewer's line — neither, yet
both arms sit at 34 of the 48 reports where the key has a line to quote — the SAME number, arrived at differently. The line is where the complete row is lost seven times with the mode right. - you want to know whether a bigger model would help — nothing here answers that
one tier was run. The ceiling arm says the corpus is solvable, but by vocabulary, not by scale; no second tier was bought.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own reports in the same nine-panel shape, data/register.json with your own part register extract and its criticality classes, and data/policy.md plus src/taxonomy.py with your own classification card — the modes, their ORDER, the unclassifiable reasons and the finding codes. The measured result does not travel with the data. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying for a call before the floor exists. Without it every paid figure reads as good. That is the case against the best-fitting scenario (“the classes and their ORDER are a written card and you have the card”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A PART NUMBER THAT IS NOT A ROW OF THE REGISTER EXTRACT. The match is exact by design, so a number typed with a different separator, a trailing revision or an extra digit group returns UNCLASSIFIABLE / PART-NOT-IN-REGISTER — 3 of the 6 such reports here sit one line from an almost-matching row and free code gets all 6, including the two where the paid arm's raw answer had to be overridden by the station. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether a reply that refuses in words would also refuse in an application that ACTED on it. Nothing in this kit acts, so a refusal phrase in a note is recorded and not scored. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is one OpenAI-compatible endpoint, reached over urllib in src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier, PEAK tariff. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-nonconformance-intake. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — a clean checkout with no key configured renders the whole board on 127.0.0.1, replays the committed run, re-derives the answer key from the rendered reports and scores all four free floors offline — measured at 0 calls and $0.00. Only the paid arm needs a key, and its replies are already committed to results/cache-r001-nonconformance-intake.jsonl, so re-scoring it also reaches no provider.






