The business caseThe problem this solves
A field action opens against a component, and somebody now has to decide, unit by unit, which vehicles are actually in the population. The columns look like they answer it: a build date, a part number at the action's position, a lot code. They do not. An as-built parts list records what was BUILT, and it does not change when a part is replaced in service — so a unit whose suspect component came off at a dealer two years ago still prints the suspect part number and its lot code, and looks exactly like an affected unit. A revised part fitted at rework lives in a build note the parts list never sees. And a lot code keyed by hand off a carton label establishes nothing about which lot went onto THIS unit. Today a recall engineer opens the record and reads it. Reading a unit traceability record end to end to place one candidate unit in or out of a field action's population, and writing down the line that decides it.
Audience
A recall or product-safety engineer scoping a field action's candidate population, and whoever has to defend the list afterwards. The decision they are making is which units to look at again — not who gets a notice. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual unit traceability records
The corpus is 60 unit traceability records, 0.09 MB (txt 60). A field action's candidate population is the one place where a table that looks complete is not. The corpus was built so the columns are RIGHT about 43 of 60 units and WRONG about 17 — and so the 17 it is wrong about are wrong in three different ways, each of which lives in a different prose panel. That is the whole shape of the job, and it is why a column-only arm cannot be the answer however good the columns look.
The corpus
- The 60 unit traceability recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your unit traceability records. That is the whole change — there is no database to migrate.
UNIT TRACEABILITY RECORD - BUILD AND SERVICE HISTORY
RECORD HEADER
Record UT-0001
Unit serial SN-A0E-638097
Product line Cadence long-wheelbase, 2025 build
Build plant Fenmore Assembly, line 2
Build date 2025-01-11
Field action FA-2026-014
Record compiled 2026-07-26 by the Product Safety Office
FIELD ACTION SCOPE (as issued)
Action FA-2026-014, opened 2026-07-06
Component position Front harness connector assembly, position C1
Suspect part PN-44810-A
Superseding part PN-44810-C
Suspect lot range L-104500 to L-105100 inclusive
Build window 2025-01-06 to 2025-06-27 inclusive
Remedy Renew the connector assembly with PN-44810-C
AS-BUILT PARTS LIST (the action's component position and its neighbours)
# Position Part number Lot code Installed Station Description
1 C1 PN-44810-A L-104776 2025-01-11 40 Front harness connector assembly
2 C2 PN-44902-B L-287736 2025-01-11 41 Harness retainer clip
3 C3 PN-45118-D L-807869 2025-01-11 41 Body-side grommet
SERVICE HISTORY
2025-07-10 Dealer 4000: scheduled service. No work on the connector assembly.
TRACEABILITY NOTES
The code above was read from the component barcode at station 40 and matched the shipment docket.
END OF RECORD
The outcomeWhat a good result looks like
One unit record in, four graded answers out: where the unit sits in the population from a closed set of seven, which KIND of record proves it from a closed set of six, the suspect component's lot code or an explicit null, and one line of the record the placement turns on.
And when it cannot
And what it does when it cannot. Three of 60 units end wrong after every check this kit has — all three superseded-narrative, where a revised part went on at rework and only a build note records it. The model did not read the note, answered IN-SCOPE at confidence 0.95, and the station confirmed it, because src/recheck.py keeps a reading only where the arm CLAIMS one of the three verdicts no column can prove. The same three fail in both scored runs. A notice goes out for a part that is already the fixed one.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- The columns are complete and trustworthy — every replacement is written back to the as-built list — the modal floor's rechecked column: pure code, $0.00, no provider
it scores 71.7% here with no reading at all. If your build system really does carry service history back into the parts list, that is the whole job and a model adds cost and variance for nothing. - The columns are right most of the time and wrong in ways that live in prose — the paid arm WITH src/recheck.py — the shipped configuration
95.0% against the columns' 71.7%, and 0 affected units dropped. The station is what makes it reproducible: two independent runs disagreed on 8 of 60 units as answered and landed on the IDENTICAL published figure after it. - Your service history is written by people in their own words — the paid arm, and re-measure — do not carry these numbers
the rules floor's 100%% here is a property of a generated corpus. On real prose a phrase list degrades and the model is the only arm that does not.
And where nothing here is good enough:
- A unit can carry two prose facts at once — remedied AND an unreliable lot code — nothing measured here
the engine handles it (RS-9's order decides) and this corpus never exercises it. Every record carries exactly one prose fact.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-05. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/actions.json with your own field actions (suspect part, superseding part, lot range, build window, component position — all typed terms) and drop your unit records into data/corpus/ in the five-panel layout src/records.py parses. THE MEASURED NUMBERS DO NOT COME WITH YOUR CORPUS, and two of them travel worse than the rest. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per unit for a call whose answer the columns already contain. That is the case against the best-fitting scenario (“The columns are complete and trustworthy — every replacement is written back to the as-built list”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A record with no parts-list row for the action's component position. src/records.py returns None and the engine refuses rather than guessing — a unit that cannot be joined to the action is not a unit the columns place. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE RULES FLOOR'S 100%% MEANS ANYTHING OUTSIDE THIS CORPUS. It does not, and the kit says so everywhere it publishes it — but the size of the gap on real prose is not measured here and cannot be, because real service history cannot be published. 6 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 3 models on the fast tier and pure Python, no key and the majority class, reads nothing, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-05 — r001-recall-vin-scope. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board — all 60 units, the parsed parts list, what the columns alone prove, the three prose panels, the rulebook and the committed run replayed — and scores both free floors offline at $0.00. python3 -m evals.check_labels and python3 -m evals.injection both run with no provider. Only evals/run.py without --floor or --stub needs a key, and it refuses to start without one rather than failing at the HTTP layer.




