Home › Use Cases › Disposition of a returned unit from its paperwork
Use caseUC0216
🧪 Use-case kit · runnable

Disposition of a returned unit from its paperwork

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A unit comes back and somebody has to say what happens to it -- scrap it, repair it, send it to the supplier, or hold it -- and then say what decided that. Almost none of it is a judgement call. It is a month count from purchase to receipt against the model's warranty, three tests against a clause somebody signed, a join onto the desk's own closed repair orders inside a 36-month lookback, and one percentage against a replacement value. What makes it slow and what makes it wrong are different things: it is slow because those four lookups sit in four places, and it is wrong because the loudest sentence in the box is the customer's account of the fault and the fact that governs is the inspection finding. On 23 of these 60 units the two point in opposite directions. The desk's manual pass over one returned unit: reading the dossier, counting the months from the purchase date to the receipt date, checking that count against the model's warranty and against each of its supplier's clauses, counting the days the claim clock has run, looking the serial up in the repair history and deciding which of those repairs fall inside the lookback, comparing the estimate with 60 pct of the model's replacement value, and then applying the four-way precedence. It does not replace the person: the record is what they confirm.

Audience

The returns-desk supervisor who confirms what happens to a unit and releases the money to repair it, and the supplier-claims clerk who has to file inside a claim window that starts the day the unit lands. Also the person who has to explain, months later, why a unit was scrapped. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual returned units

The corpus is 60 returned units, 0.02 MB (txt 60). The cases are the ones that make this decision hard, and none of them can be scraped. A customer who writes 'I dropped it' on a unit whose main-board joint has a void in it. A clause wide open on age and already shut on a claim clock that started when the unit landed. Three repairs on a serial, two of them four years ago and outside the lookback. An estimate of 53.40 against a threshold of 53.40. A unit nobody has opened yet, whose bench note says 'not stripped yet -- it is queued behind the batch from last week'. Every one of those is a boundary you can only plant, because the paperwork that would show it names a customer, a serial and a supplier's agreement. The mix is CHOSEN -- 17 clean and 43 planted -- so a rate here is a statement about this corpus and about nothing else.

The corpus

  • The 60 returned unitsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your returned units. That is the whole change — there is no database to migrate.

One returned unit, as the model receives itRT-0001.txt · 1 of 60
From: returns@meridian-service.example
To: returns.desk@calderfield-works.example
Subject: [RMA RT-0001] returned unit, disposition required

RETURNED UNIT
  RMA reference        RT-0001
  Received at the desk 28 Jun 2026
  Model                CW-522  Countertop blender
  Serial               SN-522-713257
  Date of purchase     28 Feb 2024
  Reported fault       leaking
  Repair estimate      USD 49.02

  Inspection result    On the bench, liquid has got inside the unit and tracked across the board.

  Customer packed it in the original carton with the accessories.

The outcomeWhat a good result looks like

One disposition record a person confirms instead of a dossier they read: the disposition, the clause id that decided it, the repair money to release to the cent, whether the unit is inside its warranty, whether a supplier term is open, and one sentence naming the clause. Every arithmetic field is re-derived in pure code from data/rulebook.json and data/history.json, so the record can be checked against the rulebook rather than believed.

And when it cannot

Two directions and they cost different things. A FALSE AUTHORISE is money released on a unit nobody should have repaired -- a unit over the scrap threshold, past the repair cap, on the non-repairable list, or one a supplier owed. It is invisible: the repair happens and nothing downstream reports it. A FALSE WITHHOLD is the reverse, a repairable unit scrapped or held or shipped to a vendor who rejects it, which destroys an asset and closes the claim window on every other term while it travels. MEASURED ON THIS CORPUS: the free rules floor false-authorises 7 of 39 units (17.9 pct, USD 854.60) and false-withholds 1 of 21 (USD 97.11); the paid arm does neither, on either the raw or the rechecked column. The third direction is a MISSED SAFETY HOLD -- a hazard finding sent into a repair queue -- and it is the one mistake here that is not about money. Both arms catch all 4, which is every opportunity this corpus offers.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Dossiers whose inspection finding arrives as a machine CODE -- the CSV feed row and the returns-portal export — the free rules floor alone
    On all 30 coded units the floor reads the finding 30 of 30, exactly as the paid arm does, and every number after that is the same engine on both arms. It also gets in_warranty 60 of 60, all four safety holds, the three lookback units and the other five read fields 60 of 60 -- nine measurements where the paid call buys a tie, and on one of them (in_warranty) the raw paid arm is a row BEHIND.
  • Dossiers where a person WROTE what the bench found -- the RMA email and the bench notes — the model, and then the recheck
    30 of 30 against the floor's 16 of 30, and the floor's 14 misses are structural rather than tunable: where the finding is a sentence it falls back to the customer's reported fault through the rulebook's own suggests table, and four findings have no fault that points at them. Ten of those fourteen change a graded field, which is the entire measured case for the call on this corpus.
  • The unit where the customer's account and the bench point in opposite directions — the model -- and it is the only place this corpus separates the arms at all
    23 of 23 against the floor's 14 of 23, and 3 of 3 against the floor's 0 of 3 on the fault_disagreement_rtv family, where the customer says 'I dropped it' and the bench found a void under a main-board joint on a clause with years to run. Every one of the floor's seven false authorises and its one false withhold is here or in its neighbours.

At a glanceHow the whole thing runs

98–100%record all correct over the 60 returned units -- all five graded fields right on one unit
15,831 msp50, end to end
$10.36per 1,000 returned units · Google Gemini 3 Flash

Run once, for real, on 2026-08-31. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Drop your own dossiers as .txt into data/corpus/, add one row per unit to data/returns.json (id, format, received_on, decided_on), put your models, warranties, replacement values, findings, supplier clauses, threshold, cap and lookback into data/rulebook.json, and your own closed repair orders into data/history.json. ⚠︎ A REAL RETURNED UNIT'S FILE IS PERSONAL DATA AND A REAL SUPPLY AGREEMENT IS SOMEBODY ELSE'S CONFIDENTIAL CONTRACT, AND THE WHOLE DOSSIER GOES TO A PROVIDER VERBATIM. Corpus lens →
When is this the wrong choice?Avoid: Paying per unit for arithmetic. The month count, the claim clock, the clause tests, the history join, the threshold comparison and the four-way precedence are pure code on every arm and cost nothing. That is the case against the best-fitting scenario (“Dossiers whose inspection finding arrives as a machine CODE -- the CSV feed row and the returns-portal export”). 3 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A REAL RETURNS LINE. Sixty dossiers from a generator with four format writers and small phrase pools. 7 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?WHETHER THE CORPUS COULD SEPARATE A WEAKER MODEL. Every reading bucket is empty on the paid arm and the rechecked column is 60 of 60, so the set has no headroom left. 8 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-08-31 — r001-return-disposition. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — ⚠︎ PARTIALLY EVIDENCED, AND THE REST IS NOT MEASURED. What is on disk and checkable: the 60 dossiers, data/returns.json, data/history.json, data/rulebook.json, data/fields.json, the 60-row data/gold.jsonl, data/corpus-stats.json, and three committed run records -- the free floor, the stub arm and the paid run. requirements.txt pulls nothing at runtime and src/config.py reads the .env by hand, so a clone needs Python 3 and nothing else to render the UI, rebuild the corpus, grade the key and re-score every arm. VERIFIED THIS CAPTURE: the corpus rebuilds byte-identically out-of-tree under two different PYTHONHASHSEEDs. NOT VERIFIED: nobody has cloned this kit onto a machine that has never seen it and run the whole sequence from scratch.

A living map of modern AI — kept current every morning