Home › Use Cases › Read inbound customer contacts and flag the ones describing a product hazard
Use caseUC0382
🧪 Use-case kit · runnable

Read inbound customer contacts and flag the ones describing a product hazard

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A retailer's contact centre receives everything through the same door: a late delivery, a refund argument, a colour that does not match the photograph — and, a few times a week, somebody describing a child who stopped breathing. The ones that matter do not announce themselves. Half of them never use a hazard word at all: a parent writes that the skin came up bright red and blistered, or that his lips went dusky, because that is what they saw. Meanwhile the ANGRIEST contacts — 'your returns policy is a death trap', 'the running cost has burned a hole in my wallet' — are full of hazard words and describe nothing physical. A keyword rule reads exactly the wrong half. the first-pass read a contact-centre agent gives every inbound message before deciding whether to tick the 'product safety' box — done today either by eye, at volume, or by a keyword rule that this kit measures

Audience

The product safety function that decides which contacts get read by a person, and the operations lead deciding whether a model call per contact is worth $0.00025. On this corpus the answer is yes for the READING and no for the TYPING: the free code already pulls the batch code, the product and the date out perfectly, and it cannot tell a hazard from a metaphor. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual inbound customer contacts

The corpus is 60 inbound customer contacts, 0.04 MB (txt 60). ⚠ EVERY BRAND, PRODUCT, MODEL NUMBER, BATCH CODE, CUSTOMER AND INCIDENT IS INVENTED, and on this kit that is not a formality. The input to this system is somebody describing an injury to a child; publishing a worked example of that over real contacts would be the defect rather than the product, and there is no consent under which it could be done. There is no personal data either — a contact here has no name, no address and no contact detail, only an invented customer reference. What the corpus is BUILT to do is separate two things a keyword rule cannot: 15 contacts that describe a hazard and NAME it with an ordinary hazard word, and 15 that describe the same six kinds of harm and never name one. Against them sit 12 angry contacts with no hazard, 12 that use hazard words about something that is not a physical event ('a choking hazard for my patience'), and 6 where the product has genuinely failed and nobody was hurt. evals/check_labels.py FAILS THE BUILD if a single lexicon word appears in the unnamed band or if a metaphor contact carries none.

The corpus

  • The 60 inbound customer contactsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your inbound customer contacts. That is the whole change — there is no database to migrate.

One inbound customer contact, as the model receives itCT-0001.txt · 1 of 60
INBOUND CUSTOMER CONTACT  CT-0001
Channel        : marketplace message
Received       : 2026-08-09
Customer ref   : CX-1250-8349
Subject        : Kettle lid released boiling water over my arm

I filled the Kettleworth Rapid-Boil Kettle K7 to the mark and pressed the lid catch to pour, and the whole lid released at once. Boiling water went across my left forearm and I was scalded badly enough that the skin came up in blisters. I went to A&E the same evening and they dressed it and told me to come back in three days.

The batch code on the base is KTW-2291-HB. It happened on 3 August 2026. I am not writing for a refund. I want somebody to look at that lid catch, because if this had been my son holding it we would be having a different conversation.

--- end of contact ---

The outcomeWhat a good result looks like

Every contact carries a hazard class off a closed list, the five incident facts a safety reviewer would otherwise re-read the contact to find, the customer's own words quoted beside each one, and one of two routes. A reviewer opens the packet rather than the inbox, and can disagree with the reading because the quotation is there to disagree with.

And when it cannot

⚠ THE FAILURE THAT MATTERS IS A MISSED HAZARD, AND THIS RUN HAD 0 OF THEM ON 30 HAZARD CONTACTS. That is a 60-contact synthetic corpus with a deliberately clean key, not an inbox: the genuinely borderline contact — 'the handle feels like it is getting looser', no incident, no harm, a plausible mechanism — was deliberately kept OUT of this corpus because two safety reviewers would file it differently, and that excluded band is exactly where a real false-negative rate would come from. The three field disagreements the run did produce are all arguable rather than clear: two are harm_level NEAR_MISS read as PROPERTY_DAMAGE_ONLY on contacts that describe both, and one is medical_attention NOT_STATED read as NOT_SOUGHT on a contact that says nobody was hurt.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • A missed hazard is the thing you cannot accept — the paid call, and keep a free floor beside it
    it caught 30 of 30 hazard contacts including all 15 that never name the hazard, where every free floor caught 0. Paired on recall that is 15 to 0, p = 0.000061.
  • You only need the batch code, the product and the date off contacts somebody has already triaged — free code — evals/baseline.py, $0.00
    on those three fields the floor scores 60, 60 and 60 of 60 and the paid call scores the same, p = 1.000000 on each. Paying a model to copy a batch code out of a sentence is paying for a regular expression.
  • You want the whole job and will run one arm — the paid call, with the strict free floor running beside it for $0.00
    the floor cannot be talked out of anything by a sentence typed into a form, because it does not read; disagreement between the two is the cheapest triage signal in the kit and costs nothing to compute.

And where nothing here is good enough:

  • You are about to ship 'escalate anything angry' because it is cheap and cautious — neither — measure it first, because it is measured here and it is the worst arm in the kit
    it routes 8 contacts and catches 0 of the 30 that describe a hazard: recall 0.0%. The angriest contacts in this corpus are about delivery and price, and the ones describing a child who stopped breathing are written calmly.

At a glanceHow the whole thing runs

0%missed hazards
1,783 msp50, end to end
$0.65per 1,000 inbound customer contacts · GPT-5.6 Luna

Run once, for real, on 2026-09-11. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/products.json with your own catalogue — id, brand, name, aliases, category and the shape of that product's batch codes — and drop your contacts into data/corpus/ with a matching row in data/contacts.json. ⚠ THE MEASURED ACCURACY DOES NOT TRAVEL WITH YOUR CONTACTS, and on this kit the gap is likely to be large in one direction. Corpus lens →
When is this the wrong choice?Avoid: Do not read 0 missed hazards as a guarantee — the borderline band is absent from this corpus by design. That is the case against the best-fitting scenario (“A missed hazard is the thing you cannot accept”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?a contact that names no catalogued product. The kit answers NOT_IDENTIFIED and STILL ROUTES on the hazard — which is deliberate, because a hazard described about a product we cannot name is the contact a safety team most needs to see. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?whether the kit finds hazards in contacts that a safety reviewer would argue about. The genuinely borderline band — a mechanism described with no incident and no harm — was deliberately kept out of this corpus, and it is where a real false-negative rate lives. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is the runtime provider is not named on this page; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier (THE PUBLISHED RUN). Prompt lens →
And if it fits — what do I stand up?5 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-11 — r001-recall-signal. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board — the corpus, the catalogue, the answer key, the recorded run and all three free floors — and scores every free floor offline. evals/check_labels.py, evals/baseline.py and evals.run --rescore all run with no key, no network and no spend; re-scoring the recorded run reproduces 60 of 60 route decisions to the contact, because every grader is code. The two commands that need a key are evals.run without a floor and evals.injection, and both say so before they spend.

A living map of modern AI — kept current every morning