The business caseThe problem this solves
One vehicle has been back to several dealers several times, and a customer-care desk has to say what actually happened to it. The file prints the service visits with their in and out dates and odometer, the concern in the owner's words, the technician's own account of each visit, the loaner, rental and tow records and the advisor notes. Before anybody can talk to the owner somebody has to decide which printed visits belong to the SAME reported condition, what each visit actually was, and when the vehicle was released back. Two of those three are a visit read in isolation. The first is a relation between visits, and on 44 of the 64 case files in this corpus more than one condition is in play — 34 files carry two, 10 carry three — while 39 files carry an advisor note or a technician account that describes one condition in words that look like two. Opening one vehicle case file, reading every printed service visit against the concern lines, the technician accounts and the advisor notes, grouping the visits onto reported conditions, deciding what each visit was, finding the date a note records the vehicle ready for collection, and then counting every condition's days out of service, attempts and odometer span and the vehicle's total.
Audience
The customer-care desk of a vehicle manufacturer, and the case handler who picks the file up. The decision it supports is 'what does this vehicle's record actually say happened' — never whether the vehicle qualifies for anything, never whether a threshold is met, and never how many attempts a condition is allowed: which visits count towards an attempt is the OPERATOR'S rule, supplied as data per file, and a file that declares no revision reports REVISION-NOT-DECLARED and states no attempt count at all. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual vehicle case files
The corpus is 64 vehicle case files, 0.14 MB (json 4 · jsonl 1 · md 2 · txt 64). It had to be generated, because the answer key IS the rulebook applied to a known structure and no real vehicle case file carries one. Generating it also bought the thing that matters most here: two corpus defects were found by ATTACKING the draft before any call was bought. A four-row op-code-to-role table read 262 of 262 role cells for free, so the shop now codes the line it claimed and the same table reads 188; and 244 of 262 visits were ATTEMPT, so the modal answer read 93 pct of the role cells and now reads 71.4 pct. ⚠︎ AND IT STILL LEAKS PHRASING: a regex compiled from the generator's own family libraries and note stems reaches 64 of 64 on every cell. That arm is published as this corpus's wording cost and is never a floor — nobody without tools/build_corpus.py could write it.
The corpus
- The 64 vehicle case filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement. Every case, vehicle, dealer, repair order, op code, part, campaign and note is invented from SEED 20265120, and THERE ARE NO PEOPLE IN THIS CORPUS: the owner is 'the owner' and every advisor note speaks for a desk. evals/check_labels.py sweeps all 64 files for a person-shaped name on every run and reports 0. RCS-2026 is invented and states no threshold, entitlement or qualification; the attempt revisions are operator-supplied example values read as data. The floor of record was measured before any call was bought: 30 of 64 whole-file, and the 14 failure predictions were named there before any spend.
Swap this folder for your own material and the kit is pointed at your vehicle case files. That is the whole change — there is no database to migrate.
VEHICLE CASE FILE VCF-0001
FILE ASSEMBLED 2026-06-15 STANDARD RCS-2026
CASE CCS-26-0317 VEHICLE VIN-CA37562 DESK UNIT-40
CASE OPENED 2026-01-28 REVIEW DATE 2026-06-12
CASE-HANDLING STANDARD REVISION REV-B FROM 2026-02-14, REV-A BEFORE THAT DATE OPERATOR-SUPPLIED
[1] SERVICE VISITS
ID IN OUT DEALER ODOMETER OP CODE COMPLAINT
RO-40000 2026-01-31 2026-02-08 DLR-07 29552 D1010 C-STR-HWY
RO-40001 2026-02-14 2026-02-17 DLR-12 30241 M1180 C-STR-HWY
RO-40002 2026-03-13 2026-03-20 DLR-21 31034 D1010 C-STR-GEN
RO-40003 2026-04-17 2026-04-21 DLR-33 32526 B4135 C-WTR-GEN
RO-40004 2026-05-04 2026-05-07 DLR-48 33395 R7701 C-NOI-CLD
RO-40005 2026-05-20 2026-05-30 DLR-07 34834 D1010 C-NOI-CLD
[2] REPORTED CONCERN, IN THE OWNER'S WORDS
RO-40000 "the car pulls to the right on a flat road"
RO-40001 "it wanders on the motorway"
RO-40002 "the wheel sits off centre"
RO-40003 "the carpet is wet behind the driver"
RO-40004 "it makes a knocking sound before it warms up"
RO-40005 "a tapping from the front of the engine on start-up"
[3] TECHNICIAN STORY
RO-40000 CONCERN: Owner reports the vehicle pulling right on a level road. CAUSE: No fault found on test; the condition did not occur. CORRECTION: Returned to the owner after a road test, with no repair performed.
RO-40001 CONCERN: Owner reports the vehicle wandering at motorway speed. CAUSE: Diagnosis points to TIE-ROD. CORRECTION: Adjusted TIE-ROD and road tested; the CABIN-FILTER was also renewed as a maintenance item.Abridged — the file continues.
The outcomeWhat a good result looks like
One vehicle case file in, one chronology out: the reported conditions in the order they were first raised, every visit placed under its condition with its role, out-from date, available-on date and days out of service, every condition's day total, attempt count and attempt state, the odometer span and the vehicle's total days out. On this corpus the scored run assembles 19 of 64 whole — BELOW the free floor of record's 30 — while getting 62 of 64 role maps and 64 of 64 release maps exactly right.
And when it cannot
And what it does when it cannot. 64 of 64 replies parsed, 0 stopped at the 1,000-token output ceiling, 0 were closed by the streaming reply stop and 0 calls failed or went unadmitted. Its one named failure mode is most of the result: split_one_condition fires on 39 of 64 files against 3 for merged_two_conditions. Two files named in the kit README carry it plainly — VCF-0011, where one reported condition comes back as three, and VCF-0022, where two come back as four; both floors read both files whole. Everything asked about a single visit it gets: 0 releases invented, 0 attached to the wrong visit, 0 parts notes mistaken for a release. A reply that cannot be parsed is counted WRONG and stays in the denominator; it is never dropped and never re-fired.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Every visit on the file is one reported condition, or the conditions are named in the same words throughout — the free floor of record (b000-repair-chronology-domain_strict)
it is whole on 30 of 64 files against the paid call's 19 (p = 0.0266) and takes the partition 40 to 21. Nothing is bought. - What each visit WAS is the question — attempt, no-fault-found, declined, unrelated or maintenance — read from the technician's own account — the scored run
this is what the money buys and it buys it outright: 62 of 64 files with every role right against the floor's 51 (12 / 1, p = 0.00342), 260 of 262 cells against 249. The op code does not answer it — a four-row op-code table reads 188 of 262. - A release is recorded in an advisor note rather than the out column — the scored run
64 of 64 release maps exact against the floor's 48 (16 / 0, p = 3.05e-05), with 0 releases invented, 0 on the wrong visit and 0 parts notes mistaken for one. - You want the vehicle's total days out of service — free code, and read the free-cell warning beside it
⛔ THREE separate free arms reach 61 of 64 and the paid call reaches 64: three files of headroom, p = 0.25. On VCF-0011 the call splits one condition into three and the total is 21 either way. - A file declares no attempt revision — any arm — the station reports REVISION-NOT-DECLARED and states no count
attempt_rule_overriddenis 0 of 64 on every arm. The revision is the operator's value read per visit IN DATE, and the kit does not invent one. - Untrusted text can reach the advisor notes — the answer contract, not the prompt
the cap held 12 of 12 under attack including a wording namingqualifiesandsettlementby name, and 0 extra fields came back — but the READING moved on 3 of 12 and the station cannot tell a moved reading from an honest one.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-20. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point tools/build_corpus.py at nothing and put your own files in data/corpus/ instead: src/policy.py::read wants printed visit rows carrying a repair-order id, an in date, an out date, a dealer and an odometer, plus concern lines, technician accounts and advisor notes in the Desk name: sentence shape. ⛔ THE MEASURED RESULT DOES NOT TRAVEL, AND ON THIS KIT THAT CUTS BOTH WAYS. Corpus lens → |
| When is this the wrong choice? | Avoid: Buying a call per vehicle to re-read words a word list already separates. That is the case against the best-fitting scenario (“Every visit on the file is one reported condition, or the conditions are named in the same words throughout”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A case file whose service visits are not PRINTED AS ROWS with their own repair-order id, in and out dates. src/policy.py::read is STRICT and a row matching no block raises rather than scoring a guess; every derived date, day count and odometer span is anchored on those printed columns and nothing else. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and no confidence interval is claimed anywhere on this page. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-20 — r001-repair-chronology. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with NO key configured renders the whole board on :9512 and scores every free arm offline: python3 -m evals.baseline --slice re-derived all six free readers across all eight cells in 0.14 seconds, no network, $0.00. The kit is standard library only — no framework, no build step, no pip install, no container — so there is nothing between the checkout and that result. What a clean checkout CANNOT do is reproduce the scored run: the repository is private and the paid arm needs a key and a provider endpoint. The committed replies under results/ are what makes every published percentage re-derivable without one.









