The business caseThe problem this solves
A medical information desk receives inbound questions from prescribers, pharmacists, nurses and the public, on the phone, by email and through a web form. Most are routine. Some of them CONTAIN something that is not the question: a patient who came out in a rash, a bottle that arrived two-thirds full, a question about a use the product is not approved for. The sender is not reporting any of those — they are asking about a dose and mentioning it in passing. Today an intake officer reads every message and decides, and the failure is silent: the dosing answer goes out, the sender is satisfied, and nothing anywhere records that a harm was described. The intake officer's read-and-decide pass over an inbound inquiry — the classification and the routing, not the medical answer, which stays with a qualified professional and is the kit's hard cap.
Audience
An intake or medical information operations lead deciding whether an automated classification step can be trusted in front of a human queue, and specifically whether it detects the content that must be split out regardless of how the message is phrased. The honest answer this board gives is a qualified yes on the detection and a firm no on the desk's own precedence: the model got 64 of 66 destinations and pure code got 66. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual inquiries
The corpus is 66 inquiries, 0.03 MB (txt 66). A medical information inquiry file cannot be published by anybody, ever — the narrative IS the sensitive part, and the sentence that makes an inquiry interesting to this kit is precisely the sentence describing a named person's harm. No redaction changes that. That would matter more if the claim depended on the corpus being real; it does not. What is measured is whether a written intake procedure can be applied consistently and whether a harm buried inside a routine question is detected regardless of the phrasing around it — properties of the ENGINE and the READING, which a generated corpus with a derived answer key measures more honestly, because every gold label traces to a rule rather than to somebody's judgement.
The corpus
- The 66 inquiriesunder its source's terms — Generated. tools/build_corpus.py, seed 20260901, standard library only, re-runnable to the byte —
--checkregenerates into a temp directory and diffs against. - Where each came fromnothing was taken from anywhere, so there is no source URL to map to and no publisher whose copyright page could be checked. Every inquiry, product, indication, prescriber and site is written by tools/build_corpus.py from seed 20260901 and reproduces to the byte; data/SOURCES.md records why a real medical-information inquiry file cannot be published by anybody.
Swap this folder for your own material and the kit is pointed at your inquiries. That is the whole change — there is no database to migrate.
The outcomeWhat a good result looks like
One destination per inquiry with the rule that produced it, plus three independent flags — adverse-event content, product-complaint content, off-label question — each with the sentence that establishes it, so the two questions this desk is actually asked can be answered from the run file: show me every inquiry flagged for adverse-event or complaint content and confirm each reached its destination, and show me every off-label-flagged inquiry.
And when it cannot
It files an inquiry at its surface question and the buried harm never reaches safety intake — the expensive direction, because nothing downstream reports that it happened. Or it flags harm that nobody experienced, which is cheap per inquiry and expensive in aggregate: the free floor on this corpus raises 16 false alarms against 14 real ones.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- An inquiry stream where the sender reliably says what kind of message they are sending, and adverse-event content arrives on its own form — the free keyword floor alone
Measured: on the 21 plain inquiries of this corpus the floor gets 15, and the 6 it misses are all cases where its word lists over-fire on vocabulary rather than cases it cannot read. If nothing is buried, the precedence rules never fire and there is nothing for a paid call to buy. - The job this kit is actually built for: a routine question with something else inside it — the model for the READING, pure code for the decision
The model detected 16 of 16 buried harms with zero false alarms on the nine inquiries planted to cause them, where the free floor managed 14 with sixteen false alarms. Then re-deriving the destination from that reading takes 64 of 66 to 66 of 66, for no additional call. - Deciding whether a question is about an approved use — the register, always
Both the model and the keyword floor score 12 of 12 with zero false flags — because neither is asked. They report a product string and a use string; src/register.py does the lookup. The measurement is of the READING of two strings, which is the only part of an off-label determination the kit's cap allows it to touch. - Deciding how much of the intake queue to automate on this evidence — not this, yet
Nothing here was attacked. evals/injection.py names the surface — an addendum asserting the reaction is already logged elsewhere — and has never been fired. A detection rate measured only on cooperative text is not a detection rate for an intake queue.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-01. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt and data/gold.jsonl, and edit data/intake.json and data/products.json — the inquiry types, their queues, which of them are clinical, the precedence order, the confidence floor, your products and their approved uses. Every measured number stops being true, and two of them stop being true in a specific direction. Corpus lens → |
| When is this the wrong choice? | Avoid: Anything where a routine question can contain a harm. The floor raises 16 false adverse-event alarms on this corpus — more than the 14 real ones it finds. That is the case against the best-fitting scenario (“An inquiry stream where the sender reliably says what kind of message they are sending, and adverse-event content arrives on its own form”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | An inquiry naming a product the register does not carry. src/register.on_label() returns None, the off-label rule declines to fire, and the inquiry routes on its type — which is the safe direction but is silently NOT a flag. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Whether the model would still detect a buried harm worded by somebody other than this generator. The thirteen harm sentences come from thirteen templates, and the 100 pct is a measurement of those. 9 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-01 — r001-inquiry-classify. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, scores all three free floors offline, re-derives the answer key independently, regenerates the corpus byte-for-byte and replays the committed model run's actual answers. What it cannot do without a key is make a new call — the live Classify button is disabled and says why. Observed on this kit at capture.




