The business caseThe problem this solves
A dealer asks the vehicle distributor to cover a repair the warranty no longer covers. The request arrives as a file: the repair order lines with parts and labour, every record on file for the vehicle — technician diagnostic reports, prior repair orders, service visits, bulletin and service-program references, earlier goodwill decisions, customer-care contacts, each dated, FINAL or DRAFT and tied to a vehicle id — and a notes block where the service desk, the warranty desk and a field engineer have written to each other. Nobody disputes the records. What the warranty manager does not have is the EVIDENCE PACKET: for each fact category the distributor's standard asks about, which records actually count once the notes that withdraw a record have been read against the dates they carry, and — for every category with nothing that counts — a notice saying why. Today that is somebody on the warranty desk with the standard open in one window and the file in the other. Reading one goodwill request file end to end: testing every record against four qualification tests (is it for this vehicle, is it inside the 730-day lookback and on or before the request date, does its status read FINAL, is it named by an EFFECTIVE withdrawal), where effective means joining each withdrawal note's own date to the record's date and — for a standing bulletin or program — to the failure date; then reporting each of the seven categories PRESENT with every qualifying record or MISSING with the first reason of the declared ladder; and only then summing the repair order and routing it.
Audience
A distributor's warranty desk assembling goodwill evidence packets, and the manager who has to decide whether a model call is worth buying for it. This report's own answer for this corpus is NO — the call does not beat free code — and the numbers are laid out so that answer can be checked rather than taken. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual goodwill request files
The corpus is 64 goodwill request files, 0.10 MB (json 4 · jsonl 1 · md 2 · txt 64). It is generated because it has to be. A real goodwill file carries customer names, vehicle identification numbers, dealer economics and the distributor's own decisions, and no distributor publishes one. Generating it buys the one thing a scraped corpus could never give: an answer key derived by APPLYING the standard to the rendered text and re-derived by a second implementation that imports nothing from src/. 16 cases were planted deliberately — a clean file, a DRAFT diagnosis, another vehicle's record, the 730/731-day lookback edge, a contact dated after the request, a withdrawal in force on a point record and on a standing reference, late withdrawals of each, a partial withdrawal, a confirming decoy, the reason ladder, a file with no record at all, mixed noise, a cap probe and two traps at once — each written as two fact patterns and each pattern twice, so the atlas row's own question, whether two similar requests surface the same facts, is measurable.
The corpus
- The 64 goodwill request filesgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from and — section 4 — thirteen cases the generator was built to make the model fail on, written BEFORE any call was bought. The two it named first were confirmed: a note dated before the record it names was read as a withdrawal on 33 of the 34 records where that happens, and a late withdrawal of a standing bulletin or program was applied 18 of 18 times. Four it named were NOT observed and that is published too: the mirror decoys, the near-miss vehicle, the DRAFT diagnosis and the cap-probe echo all scored 0.
Swap this folder for your own material and the kit is pointed at your goodwill request files. That is the whole change — there is no database to migrate.
GOODWILL REQUEST FILE GWR-0001
FILE ASSEMBLED 2025-12-06 STANDARD GEAS-2026
DEALER DLR-142 MODEL LINE MDL-11 VEHICLE VEH-20404
IN SERVICE 2023-04-17 FAILURE DATE 2025-11-03 ODOMETER AT FAILURE 107777
REQUEST DATE 2025-11-23
DECLARED SCOPE LOOKBACK 730 DAYS REGIONAL THRESHOLD 2500.00
[1] REPAIR ORDER LINES
LINE OP CODE COMPONENT PARTS LABOUR
1 OPC-1196 WATER-PUMP 372.52 152.17
[2] RECORDS ON FILE
ID KIND VEHICLE DATED STATUS
REC-20001 DIAG-REPORT VEH-20404 2025-11-22 DRAFT
REC-20002 GW-HISTORY VEH-20404 2025-04-29 FINAL
REC-20003 SERVICE-RECORD VEH-20404 2024-03-10 FINAL
REC-20004 CONTACT-LOG VEH-20404 2025-11-01 FINAL
REC-20005 SERVICE-RECORD VEH-20404 2025-05-06 FINAL
REC-20006 PROGRAM-REF VEH-20404 2025-03-19 FINAL
[3] NOTES
Customer care desk: the finding of 2025-11-17 that REC-20006 was invalidated for VEH-20404 was itself upheld; REC-20006 no longer stands and nothing replaced it.
Warranty desk: the declared lookback and regional threshold are reprinted on the scope line above.
Customer care desk: the finding of 2025-06-24 that REC-20003 was invalidated for VEH-20404 was itself withdrawn; REC-20003 still stands and nothing replaced it.
Customer care desk: the finding of 2025-02-13 that REC-20002 was invalidated for VEH-20404 was itself upheld; REC-20002 no longer stands and nothing replaced it.
The outcomeWhat a good result looks like
One request file in, one evidence packet out: seven category rows in the declared order, each PRESENT with every qualifying record cited or MISSING with exactly one reason; then — in code, for every arm alike — the requested amount (parts plus labour, in integer minor units), whether it is strictly above the declared regional threshold, which manager decides, the present count, one notice per missing category and the list of records in front of the decider. The packet never grants, denies, prices or commits.
And when it cannot
And what it does when it cannot. On the scored run all 64 replies parsed, none stopped at the ceiling and none produced a shape defect — no category outside the standard, no status or reason outside the vocabulary, no record cited that is not on the file or is of another kind, no category missing from an answer. What it produces instead is a CONFIDENT WRONG READING: 49 of the 64 files have something wrong in them, at a median self-reported confidence of 0.90 against 0.93 on the 15 it gets completely right. There is no signal in that number and a reader must not route on it. The station then does honest arithmetic over the wrong reading and returns an internally consistent packet with a real fact taken away.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your records carry a withdrawal as a STRUCTURED COLUMN — a status of WITHDRAWN or a superseded-by id — rather than as a sentence in a notes block. — the free columns floor, and do not buy a call at all
22 of 64 files whole for $0.00, and on the 22 files the columns settle it is 22 of 22, where the call is 4. - Your withdrawals live in the notes, and you can write the date joins down. — a free note-reading rule with both date joins (reinst_last_dated)
38 of 64 files right on status and reason against the call's 23 — significantly better, p = 0.0167 — for $0.00. - Your notes are free text a rule cannot parse, and a missed withdrawal is costlier than a wrongly dropped record. — the call, with a person reviewing every MISSING / record-withdrawn row
it is the only arm that scores on the 42 files a note decides (11 whole, the floor 0), and it over-withdraws: 33 of 34 predecessor-named records and 18 of 18 late standing references dropped. - You want the routing and the notices, not the reading. — src/policy.py::rollup(), free, with any reading you like
the amount, the threshold comparison and the decider are right on 64 of 64 for every arm; none of it needs a model. - You want to know whether pressure in the notes can move the packet. — evals/injection.py, 32 live calls, $0.012411
0 breaches and 0 categories filled in under all four framings — and 2 regressions under firm and authority framing, a withdrawn record reported PRESENT.
And where nothing here is good enough:
- You are choosing a confidence threshold to route on. — none — do not route on this arm's confidence
the median is 0.93 where the file is completely right and 0.90 where it is not.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own request files in the same shape — header, DECLARED SCOPE line, repair order lines, records on file, notes — and rebuild the key with python3 tools/build_corpus.py. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per file for a reading the columns already give you. That is the case against the best-fitting scenario (“Your records carry a withdrawal as a STRUCTURED COLUMN — a status of WITHDRAWN or a superseded-by id — rather than as a sentence in a notes block.”). 6 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A file with no DECLARED SCOPE line. The lookback and the regional threshold are operator settings the file must carry; with neither printed there is no lookback test and no routing, for any arm. 7 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | WHETHER THE CALL IS REALLY WORSE THAN FREE CODE. It is behind: 15 of 64 whole files against 22 for free code that reads no note, p = 0.265, not significant on 64 files. 10 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-goodwill-evidence. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, re-runs all six free arms live, replays the scored run from results/, asserts every arm on every file equal to its result file (448 of 448) and serves all nine screenshots. What it cannot do is buy a call: 'Ask the model live' disables itself and says why. Measured from the kit folder with no key set; nothing else is required beyond Python 3 — requirements.txt pins no package.








