Home › Use Cases › Screen a pharmacovigilance literature hit at the abstract stage and file its disposition
Use caseUC0275
🧪 Use-case kit · runnable

Screen a pharmacovigilance literature hit at the abstract stage and file its disposition

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

Every week a literature search returns hits for a company's monitored products, and somebody reads the ABSTRACT of each one — the full text has not been requested yet — and files it: to the case queue, the special-situation queue, the signal team, or to nobody. The damage runs one way. A record routed to the wrong queue costs a PV scientist a few minutes. A record filed LOG-ONLY that belonged in a queue is SILENTLY DROPPED: nobody downstream learns it existed, and there is no later step that catches it. The expensive version of getting this wrong is not a misfiled case — it is the report where nothing went wrong. A pregnancy with a healthy baby, an overdose with no symptoms, a dispensing error caught in time: an identifiable patient, a real exposure, and no adverse word anywhere in the text for a keyword screen to catch. Reading a week of literature-search abstracts by hand and deciding, record by record, which ones a PV scientist must now see. It does not replace the PV scientist, the full text, the causality assessment, the seriousness determination or any reporting decision — none of the five exists anywhere in this kit, and the answer contract has no field that could express any of them.

Audience

A PV scientist working a weekly literature-screening queue, and the safety lead who decides how much of that reading a machine may do. The decision this report is for is narrower than it looks: not 'should we buy a model', but 'what is a recall number worth when a free twelve-line script can score 100 pct on it'. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual literature-search records

The corpus is 66 literature-search records, 0.07 MB (txt 66). It is generated because it has to be. A real pharmacovigilance literature screening set is a week of third-party papers ABOUT NAMED MEDICINES, and every record in it is either copyrighted text somebody else owns or a case report describing an identifiable patient — there is no public version of one that is both usable and safe to publish. So the corpus is invented from a fixed seed, and everything about it is arranged to make the READING hard rather than the parsing: 66 records, 43 of them marked hard, one disposition each, and a patient who is an age band and a sex and nothing else. ⚑ THE HARD PART IS DELIBERATELY THE ABSENCE OF A CUE. Sixteen records describe an identifiable patient exposed in a special situation with NOTHING GOING WRONG — a pregnancy delivered at term, an overdose with unremarkable tests, a dispensing error caught in time — and the word adverse appears in 0 of the 17 adverse-event records and 0 of the 9 aggregate ones, so no arm can answer the safety reading by looking for a word. Four records are homonyms, where a registry, a frailty score and a trial acronym are spelled exactly like three of the brands, so the substance reading cannot be a dictionary lookup either — and those four are the only cell in the corpus where matching and reading come apart. ⚠︎ AND IT IS EASIER THAN THE JOB, WHICH IS THE FINDING RATHER THAN A DISCLAIMER. Abstracts here are 3 to 5 sentences, single-topic and uniform; real ones are longer, structured, duplicated across databases and sometimes contradict their own full texts. The cited sentence sits third in 32 of the 38 cited records, so an arm that returns sentence three and reads nothing scores 84.2 pct of the citation credit. The measured run answered 66 of 66, and data/SOURCES.md named six families this kit expected to lose BEFORE the first call was bought — all six were refuted. A corpus is finished when the model can miss something, and this one is not.

The corpus

  • The 66 literature-search recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — all 66 records and the whole answer key are generated in-process from seed 20260902 by tools/build_corpus.py, and re-running it rebuilds them byte for byte (--check reports 0 differences under PYTHONHASHSEED 0 and 12345). Nothing is fetched, scraped or licensed from anywhere, so there is no source URL to map and no third-party dedication to verify.

Swap this folder for your own material and the kit is pointed at your literature-search records. That is the whole change — there is no database to migrate.

One literature-search record, as the model receives itLIT-0001.txt · 1 of 66
LITERATURE SEARCH HIT

Record id                 LIT-0001
Search cycle              2026-W35, search run 2026-08-31
Search profile            ALD-PV-07 - every monitored INN, brand and class term
Database                  synthetic bibliographic index (this record is invented)

CITATION

Authors                   Cazorla O, Eberhardt J, Delacroix S
Title                     A case of acute interstitial nephritis in a patient treated with tavrelimod
Journal                   International Review of Metabolic Practice
Year / volume / pages     2026;36(8):451-458
Publication type          Journal Article; Case Reports

ABSTRACT

Tavrelimod is an oral TRK-9 modulator licensed for plaque psoriasis. A 65-year-old woman
developed acute interstitial nephritis 20 weeks after starting Vexira for plaque psoriasis. The
event settled after withdrawal and no rechallenge was attempted. Clinicians should be aware of
this presentation in patients receiving tavrelimod.

SCREENING RECORD

Screened by               not yet screened
Disposition               not yet assigned

The outcomeWhat a good result looks like

One record in, five graded answers out: the three readings, the disposition it is filed under after LSP-2026 is applied in code, and the sentence that establishes the safety content quoted verbatim and locatable in the record at character offsets. On this corpus the paid call kept 38 of 38 records that belong in a queue, over-included 0 of 28 that do not, and returned 66 of 66 dispositions exactly.

And when it cannot

⚠︎ AND THE HONEST HEADLINE IS THAT THE CORPUS SATURATED. 66 of 66 is a statement about lit-screen-v1-66abstracts, not about literature. data/SOURCES.md named six families this kit expected to lose — special situations with nothing wrong, sibling-product case reports, a patient singled out inside a trial, null aggregate findings, homonyms, bare tolerability sentences — IN WRITING, BEFORE THE FIRST CALL WAS BOUGHT, and the model refuted all six. The single non-exact cell in 330 is an AMBIGUOUS LABEL rather than a miss the model earned: on LIT-0020 the abstract says “an accidental 8-fold DOSING ERROR resulted in ingestion of 8 days of mirodapt in one day”, the key says the situation is overdose, the call said medication_error, and LSP-2026 states no precedence between the two for a record that is both. The disposition was right either way. ⚑ THE ADVERSARIAL ARM ALSO TOOK NOTHING: 12 of 12 held against a sentence asserting the authors had found the event unrelated and no follow-up required. That is a weaker result than 'defended' — the attack never reached the reading, which is the thing no rule could have restored.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want to know whether anything was silently dropped — not_dropped_pct, ALWAYS beside over_included_pct
    it is the only metric that distinguishes a record a person will see from one nobody will, which is the only distinction the job turns on
  • You want to know whether an arm is quietly turning into flag-all — over_included_pct
    it is the price of recall, and it is the half a recall-first kit is tempted to leave out
  • You want an early warning before a drop happens — safety_content_correct
    the disposition is DERIVED from the readings, so a drifting reading shows up here first; it can only stay hidden while wrong readings happen to map to the right queue
  • You want to know whether the arm read the record at all — substance_match_correct on the 4 homonym records
    a whole-word dictionary answers the other 62 and cannot answer these four — they are the only cell in the corpus where matching and reading come apart

And where nothing here is good enough:

  • You want to know whether the citation means anything — nothing in this kit
    there is no metric here that separates reading from position: 32 of the 38 cited records carry their sentence third in the abstract and an arm that returns the third sentence scores 84.2 pct

At a glanceHow the whole thing runs

100%disposition exact pct
4,580 msp50, end to end
$5.27per 1,000 literature-search records · google/gemini-3-flash

Run once, for real, on 2026-09-02. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own records, data/products.json with your own monitored list (INN, brand, class, indication — plus, for the generator only, your class siblings and the words that are homonyms of your brands), and data/gold.jsonl with your own key: one row per record carrying substance_match, safety_content, special_situation_type, disposition, citation_line and citation_span. Corpus lens →
When is this the wrong choice?Avoid: Quoting it alone. The free flag-all floor scores 100 pct on it for $0.00 by sending every record to a queue, so on its own it is a number a regex owns. That is the case against the best-fitting scenario (“You want to know whether anything was silently dropped”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A real abstract. These are 3 to 5 sentences, single-topic and uniform; real ones are longer, structured Background/Methods/Results/Conclusions, and much noisier. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled sentence is the sentence a PV scientist would have quoted. The key names ONE per record and the scorer has no partial credit. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-02 — r001-lit-screen. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board on 127.0.0.1:9275 and scores every graded cell offline: the corpus rebuild from seed 20260902, the independent label gate, all three free floors over all 330 graded cells, the --stub pass through the entire pipeline, and the replay of both committed model runs off their own result files. Observed here: python3 tools/build_corpus.py --check exits 0 with 0 differences; python3 -m evals.check_labels exits 0 with 0 problems over 66 records, 6 rules and 25 shipped text files; the rules floor scores 68.2 pct of dispositions exactly. There is nothing to pip install — the harness is standard library — and node plus Chrome are needed only for tools/shoot_ui.mjs, never for the eval.

A living map of modern AI — kept current every morning