Home › Use Cases › Assemble a goodwill request's evidence packet, every fact category with its records
Use caseUC0481
🧪 Use-case kit · runnable

Assemble a goodwill request's evidence packet, every fact category with its records

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A dealer asks the vehicle distributor to cover a repair the warranty no longer covers. The request arrives as a file: the repair order lines with parts and labour, every record on file for the vehicle — technician diagnostic reports, prior repair orders, service visits, bulletin and service-program references, earlier goodwill decisions, customer-care contacts, each dated, FINAL or DRAFT and tied to a vehicle id — and a notes block where the service desk, the warranty desk and a field engineer have written to each other. Nobody disputes the records. What the warranty manager does not have is the EVIDENCE PACKET: for each fact category the distributor's standard asks about, which records actually count once the notes that withdraw a record have been read against the dates they carry, and — for every category with nothing that counts — a notice saying why. Today that is somebody on the warranty desk with the standard open in one window and the file in the other. Reading one goodwill request file end to end: testing every record against four qualification tests (is it for this vehicle, is it inside the 730-day lookback and on or before the request date, does its status read FINAL, is it named by an EFFECTIVE withdrawal), where effective means joining each withdrawal note's own date to the record's date and — for a standing bulletin or program — to the failure date; then reporting each of the seven categories PRESENT with every qualifying record or MISSING with the first reason of the declared ladder; and only then summing the repair order and routing it.

Audience

A distributor's warranty desk assembling goodwill evidence packets, and the manager who has to decide whether a model call is worth buying for it. This report's own answer for this corpus is NO — the call does not beat free code — and the numbers are laid out so that answer can be checked rather than taken. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual goodwill request files

The corpus is 64 goodwill request files, 0.10 MB (json 4 · jsonl 1 · md 2 · txt 64). It is generated because it has to be. A real goodwill file carries customer names, vehicle identification numbers, dealer economics and the distributor's own decisions, and no distributor publishes one. Generating it buys the one thing a scraped corpus could never give: an answer key derived by APPLYING the standard to the rendered text and re-derived by a second implementation that imports nothing from src/. 16 cases were planted deliberately — a clean file, a DRAFT diagnosis, another vehicle's record, the 730/731-day lookback edge, a contact dated after the request, a withdrawal in force on a point record and on a standing reference, late withdrawals of each, a partial withdrawal, a confirming decoy, the reason ladder, a file with no record at all, mixed noise, a cap probe and two traps at once — each written as two fact patterns and each pattern twice, so the atlas row's own question, whether two similar requests surface the same facts, is measurable.

The corpus

  • The 64 goodwill request filesgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromdata/SOURCES.md states where every byte came from and — section 4 — thirteen cases the generator was built to make the model fail on, written BEFORE any call was bought. The two it named first were confirmed: a note dated before the record it names was read as a withdrawal on 33 of the 34 records where that happens, and a late withdrawal of a standing bulletin or program was applied 18 of 18 times. Four it named were NOT observed and that is published too: the mirror decoys, the near-miss vehicle, the DRAFT diagnosis and the cap-probe echo all scored 0.

Swap this folder for your own material and the kit is pointed at your goodwill request files. That is the whole change — there is no database to migrate.

One goodwill request file, as the model receives itGWR-0001.txt · 1 of 64
GOODWILL REQUEST FILE  GWR-0001
FILE ASSEMBLED  2025-12-06   STANDARD  GEAS-2026
DEALER  DLR-142   MODEL LINE  MDL-11   VEHICLE  VEH-20404
IN SERVICE  2023-04-17   FAILURE DATE  2025-11-03   ODOMETER AT FAILURE  107777
REQUEST DATE  2025-11-23
DECLARED SCOPE  LOOKBACK 730 DAYS   REGIONAL THRESHOLD 2500.00

[1] REPAIR ORDER LINES
LINE  OP CODE   COMPONENT            PARTS       LABOUR
1     OPC-1196  WATER-PUMP           372.52      152.17

[2] RECORDS ON FILE
ID          KIND            VEHICLE    DATED       STATUS
REC-20001   DIAG-REPORT     VEH-20404  2025-11-22  DRAFT
REC-20002   GW-HISTORY      VEH-20404  2025-04-29  FINAL
REC-20003   SERVICE-RECORD  VEH-20404  2024-03-10  FINAL
REC-20004   CONTACT-LOG     VEH-20404  2025-11-01  FINAL
REC-20005   SERVICE-RECORD  VEH-20404  2025-05-06  FINAL
REC-20006   PROGRAM-REF     VEH-20404  2025-03-19  FINAL

[3] NOTES
Customer care desk: the finding of 2025-11-17 that REC-20006 was invalidated for VEH-20404 was itself upheld; REC-20006 no longer stands and nothing replaced it.
Warranty desk: the declared lookback and regional threshold are reprinted on the scope line above.
Customer care desk: the finding of 2025-06-24 that REC-20003 was invalidated for VEH-20404 was itself withdrawn; REC-20003 still stands and nothing replaced it.
Customer care desk: the finding of 2025-02-13 that REC-20002 was invalidated for VEH-20404 was itself upheld; REC-20002 no longer stands and nothing replaced it.

The outcomeWhat a good result looks like

One request file in, one evidence packet out: seven category rows in the declared order, each PRESENT with every qualifying record cited or MISSING with exactly one reason; then — in code, for every arm alike — the requested amount (parts plus labour, in integer minor units), whether it is strictly above the declared regional threshold, which manager decides, the present count, one notice per missing category and the list of records in front of the decider. The packet never grants, denies, prices or commits.

And when it cannot

And what it does when it cannot. On the scored run all 64 replies parsed, none stopped at the ceiling and none produced a shape defect — no category outside the standard, no status or reason outside the vocabulary, no record cited that is not on the file or is of another kind, no category missing from an answer. What it produces instead is a CONFIDENT WRONG READING: 49 of the 64 files have something wrong in them, at a median self-reported confidence of 0.90 against 0.93 on the 15 it gets completely right. There is no signal in that number and a reader must not route on it. The station then does honest arithmetic over the wrong reading and returns an internally consistent packet with a real fact taken away.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your records carry a withdrawal as a STRUCTURED COLUMN — a status of WITHDRAWN or a superseded-by id — rather than as a sentence in a notes block. — the free columns floor, and do not buy a call at all
    22 of 64 files whole for $0.00, and on the 22 files the columns settle it is 22 of 22, where the call is 4.
  • Your withdrawals live in the notes, and you can write the date joins down. — a free note-reading rule with both date joins (reinst_last_dated)
    38 of 64 files right on status and reason against the call's 23 — significantly better, p = 0.0167 — for $0.00.
  • Your notes are free text a rule cannot parse, and a missed withdrawal is costlier than a wrongly dropped record. — the call, with a person reviewing every MISSING / record-withdrawn row
    it is the only arm that scores on the 42 files a note decides (11 whole, the floor 0), and it over-withdraws: 33 of 34 predecessor-named records and 18 of 18 late standing references dropped.
  • You want the routing and the notices, not the reading. — src/policy.py::rollup(), free, with any reading you like
    the amount, the threshold comparison and the decider are right on 64 of 64 for every arm; none of it needs a model.
  • You want to know whether pressure in the notes can move the packet. — evals/injection.py, 32 live calls, $0.012411
    0 breaches and 0 categories filled in under all four framings — and 2 regressions under firm and authority framing, a withdrawn record reported PRESENT.

And where nothing here is good enough:

  • You are choosing a confidence threshold to route on. — none — do not route on this arm's confidence
    the median is 0.93 where the file is completely right and 0.90 where it is not.

At a glanceHow the whole thing runs

23%rechecked all correct pct
2,342 msp50, end to end
$0.83per 1,000 goodwill request files · the cheapest published long-context card

Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own request files in the same shape — header, DECLARED SCOPE line, repair order lines, records on file, notes — and rebuild the key with python3 tools/build_corpus.py. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying per file for a reading the columns already give you. That is the case against the best-fitting scenario (“Your records carry a withdrawal as a STRUCTURED COLUMN — a status of WITHDRAWN or a superseded-by id — rather than as a sentence in a notes block.”). 6 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A file with no DECLARED SCOPE line. The lookback and the regional threshold are operator settings the file must carry; with neither printed there is no lookback test and no routing, for any arm. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE CALL IS REALLY WORSE THAN FREE CODE. It is behind: 15 of 64 whole files against 22 for free code that reads no note, p = 0.265, not significant on 64 files. 10 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-16 — r001-goodwill-evidence. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, re-runs all six free arms live, replays the scored run from results/, asserts every arm on every file equal to its result file (448 of 448) and serves all nine screenshots. What it cannot do is buy a call: 'Ask the model live' disables itself and says why. Measured from the kit folder with no key set; nothing else is required beyond Python 3 — requirements.txt pins no package.

A living map of modern AI — kept current every morning