Catch serious drug side-effect reports, even when they sound mild
Side-effect reports say "mild" or "severe", and neither tells you whether the case is serious under the rules. This app reads each report, decides whether a seriousness rule was met and flags the urgent ones for a case processor.
PresenterOpens the private repo. Visible to admins only.
For drug safety case intakeCross-domain · Pharma & Life Sciences
Why it matters
Today's manual process, and the same job with the app
A drug safety team at a drug maker, sorting incoming side-effect reports before a case processor works them.
✕Today's manual process
1Read every report to find the drug, the patient's age, the outcome and who reported it.
2Check each outcome against the seriousness rules, looking past words like mild or severe.
3Note the reporter's view on whether the drug caused it, then decide which cases go first.
4One slip means a serious case waits in the ordinary queue.
Every report read and judged manually
✓With the app
1Each report is read, and ten details are filled in from it.
2Seriousness is decided against the rules, not the report's own words like mild.
3The reporter's view is kept separate, and serious cases they link to the drug are flagged.
4Flagged cases reach a case processor first. A person still works and files every case.
People start with the urgent cases
See it work
One real case, read by the app, step by step
A consumer reports mild stomach pain in a child taking Orbadanil, with no hospital stay, yet the story describes a life-threatening episode.
Catch serious drug side-effect reports, even when they sound mildReference appBuilt to be shaped to your process
5
1The case report AE-0011: a child taking Orbadanil, reported by a consumer.
2The report's own word mild, the word most readers would go by.
3What does not count no hospital stay and a full recovery do not rule the case out.
4Serious under the rules a seriousness rule was met, whatever the report's wording.
5Sent to a case processor first serious, and the reporter says the drug possibly caused it.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Catch serious drug side-effect reports, even when they sound mild
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Triaging an adverse-event case means reading the narrative to decide whether the outcome met a regulatory seriousness criterion — a defined list, not an impression — and separately recording the reporter's own causality view, before the case reaches a case processor. The two judgments are easy to conflate, and the report's own severity wording ("severe", "mild") is the most available signal and the wrong one. Someone reading each adverse-event case narrative to decide whether the outcome actually met one of the regulatory seriousness criteria — as opposed to merely sounding bad — and then checking the reporter's own causality assessment before the case is routed to a case processor.
Audience
Pharmacovigilance case intake and safety operations teams who triage incoming case reports before a case processor works them, and the people who build tooling for them. The mechanic is generic: any structured extraction where one field is a defined classification the surrounding prose keeps arguing against. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual case reports
The corpus is 55 case reports, 0.03 MB (txt 55). Plain text, one format, invented rather than fetched — a real case narrative cannot be published, and there is no patient or reporter in this one by construction: no names, no dates of birth, no exact ages, only a bucketed age band and a role word. Ten fields are chosen because the same ones drive a real triage decision: which drug, what happened, whether the patient was admitted, the outcome, the reporter's own causality view, and — the field this kit exists to test — whether a regulatory seriousness criterion was actually met.
The corpus
The 55 case reportsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your case reports. That is the whole change — there is no database to migrate.
One case report, as the model receives itAE-0001.txt · 1 of 55
Case ID
-------
AE-2026-0001
Patient Age Range
-----------------
30-44
Suspect Drug
------------
Orbadanil
Event Description
-----------------
severe swelling of the lips and tongue
Hospitalization
---------------
no
Event Outcome
-------------
recovering
Reporter Causality Assessment
-----------------------------
not-assessed
Reporter Type
-------------
consumer
Case Narrative
--------------
Symptoms were still present at the time of reporting but the patient continues normal daily activity. Nothing beyond stopping the drug has been done and no admission or emergency visit has taken place.
The outcomeWhat a good result looks like
A ten-field extracted record per case, plus one pure-code computed flag folding two checks together — regulatorily serious AND a causality assessment this kit routes on — informational only, never a filed report and never a reporting clock.
And when it cannot
This run found 3 classification errors across both tiers and 1100 combined cells, all on the same seriousness criterion, and one review-flag false negative on the deliberating tier. What the run did not test: a follow-up report that changes the seriousness of a case already assessed, a narrative that describes a qualifying outcome without naming it, a report where the stated criterion is contradicted by an attached summary, or a case where the deciding fact is simply absent — hospitalization: unknown and event_outcome: unknown are allowed values this corpus never exercises. See Business.not_good_enough.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Screening an intake queue for which cases a case processor should work first — either tier over the free floor — but read Business.not_good_enough before trusting either flag figure 96% and 98% seriousness accuracy against the free severity-register floor's 49%, and 23 of 25 against 0 of 25 on the confusable reports the floor is built to fail.
Deciding the two tiers on cost, speed or accuracy — the fast tier, with a caveat rather than a clean win Roughly 11% cheaper per case ($0.0020372 vs $0.0022987), half the p50 latency (3538 ms vs 7084 ms), and 1.00 flag recall against 0.93. But the deliberating tier classified one MORE case correctly (54/55 vs 53/55), so the cheaper tier is not simply the better classifier — it is the one whose errors happened to land where the flag drops them.
At a glanceHow the whole thing runs
99.6%extraction accuracy
3,538 msp50, end to end
$2.04per 1,000 case reports · Google Gemini 3 Flash
Run once, for real, on 2026-08-21. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Catch serious drug side-effect reports, even when they sound mild14 steps · 4 questions · run once, for real · 2026-08-21
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace data/corpus/*.txt, write data/fields.json, and supply a gold record per case. Corpus lens →
When is this the wrong choice?
Avoid: The severity-register floor for the seriousness judgment specifically — it gets 25 of 25 confusable reports wrong, which is the whole population this kit exists to handle. That is the case against the best-fitting scenario (“Screening an intake queue for which cases a case processor should work first”). 2 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Scanned or image-only reports — there is no OCR step. 4 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether either tier's flag recall means anything at this sample size. 15 positive cases is a small denominator, and this run showed exactly how fragile the figure is: the fast tier's 1.00 recall survived two classification errors only because both landed on cases whose causality assessment the flag drops. 5 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 2 models on the fast tier and the deliberating tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-21 — r001-ae-seriousness. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Verified from a cold clone with no key configured: 55 reports segment into 495 sections in 0.009 seconds, and the assembled prompt for AE-0011 replays byte for byte — its three part sizes match what run r001 recorded, and its measured total of 1313 input tokens matches the capture call's own reported prompt_tokens exactly.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
3,538 msp50, end to end
9,961 msp95
2 minclone to first result
What the clock covers. model call only, one per case
Current processWhat it replaces
Someone reading each adverse-event case narrative to decide whether the outcome actually met one of the regulatory seriousness criteria — as opposed to merely sounding bad — and then checking the reporter's own causality assessment before the case is routed to a case processor.
Where it is not good enough
Three classification errors across both tiers, and every one is the same criterion. Both models missed only cases turning on "medically important enough to require an intervention to prevent one of the other outcomes" — the one criterion of the six that is a judgment rather than a fact. Both read "no admission was required" and stopped there. Worse: the fast tier's perfect 1.00 review-flag recall is luck, not safety. Its two misclassifications happened to land on cases whose reporter had said unrelated and not-assessed, which the flag drops anyway; had either fallen on a triggering case, recall would have been 13 of 15. The deliberating tier classified MORE cases correctly (54/55 against 53/55) and scored WORSE on the flag (0.93), because its single error did land on a related case. 15 positive cases is a small denominator for a rule this consequential.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free severity-register floor (evals/baseline.py, asks only whether the report says "severe") scores 94.91 pct extraction accuracy and 55 of 55 on every structured field — and 0 of 25 on the reports whose severity word points against the truth, missing 8 of the 15 cases that should be flagged. Both model tiers read the narrative instead and got 23 and 24 of those 25 right. But the fast tier's perfect 1.00 flag recall is luck, not safety: both of its classification errors happened to land on cases the flag drops anyway. No red-team run exists for this kit — this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
SECTION_HINTS
src/select.py
map fields to your own report's headings; unmatched falls back to the whole document
PROVIDERS
src/adapters/__init__.py
any OpenAI-compatible host, or Anthropic's Messages API
CAUSALITY_TRIGGERS
src/extract.py
this kit's own flat routing rule — replace with your own programme's actual documented criteria before trusting the flag for anything real. It is not a reporting clock and does not model one.
the field schema
data/fields.json
a different set of fields entirely, with its own types and allowed values
Components
Component
File
Role
segment
src/segment.py
cut the case report into addressable sections, pure code
select
src/select.py
pick which sections carry each field, pure code
prompt
src/prompt.py
assemble one call for all ten fields, with the six seriousness criteria stated in full
extract
src/extract.py
the AI layer, one provider one key — plus the pure-code routing computation downstream
judge
evals/judge.py
score field accuracy, the flag's recall/precision, and seriousness accuracy by register — three separate figures, pure code
Where it breaks at scale
One call per case, no concurrency and nothing shared between calls: 55 cases took 219 seconds wall clock on the fast tier and 497 on the deliberating tier. A day's intake at a real safety database needs batching and a rate-limit strategy this kit does not have. It also reads one report in isolation — a real intake queue is full of follow-up reports that amend a case already assessed, and nothing here reconciles them.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
Before anything is asked of the model. Ten named fields with their own types and allowed values, plus a second panel for the three figures computed afterwards in pure code.successOpen full size →AE-0011, extracted live, and every available shortcut gets it wrong. The report calls the event "mild", the Hospitalization field says no and the Event Outcome says recovered — and the narrative records an episode the treating clinician called immediately life-threatening. The model answers is_serious: yes, and the pure code routes it for expedited review because the reporter also called the drug possibly-related.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
The same button with no API_KEY configured. A calm 200 and a plain sentence, not a stack trace: nothing was called, nothing was spent, and the field table stays browsable.failureOpen full size →
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
55case reports
0.03 MiBtxt 55
495sections · p50 40 chars
$0.00setup · 0.009s
How it is cutWhat one section is
cut on underlined section headings; a report with none falls back to one whole-document segment so a span still resolves
SetupWhat the setup figure measured
There is no index. Preparation is segmentation only — 55 reports cut into 495 sections by src/segment.py, pure code, no model and no key.
LicenceLicence
MIT — this repository's own licence. Every case identifier, drug name and narrative is invented; a real adverse-event narrative cannot be published either — it is among the most identifying text a company holds.
Bring your ownBring your own case reports
Replace data/corpus/*.txt, write data/fields.json, and supply a gold record per case. SECTION_HINTS in src/select.py maps fields to headings and will need editing for a different report layout; when it does not match, selection falls back to the whole document — slower, more expensive, always correct.
What breaks it
Scanned or image-only reports — there is no OCR step.
A report whose sections are not headed — segment() falls back to one whole-document segment, so a span names "document" and locates nothing finer.
Follow-up reports are never reconciled against the case they amend — this kit reads one report in isolation, and a real intake queue is full of follow-ups that change a case's seriousness after it was first assessed.
A case where the deciding fact is absent — hospitalization: unknown and event_outcome: unknown are allowed values this corpus never exercises, so nothing measured here says how the classification behaves without them.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
2,287
582
field schema
2,442
535
report sections
678
196
Total
1,313
This is the cost lesson as arithmetic: of the 1,313 tokens assembled, 582 are instructions — 44% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Measured token-for-token via evals/prompt_tokens.py's nested-prefix subtraction against the provider's own tokenizer, not estimated from characters — see results/tokens-p001-ae-seriousness.json.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You extract structured fields from an adverse event case report. You return JSON and nothing else.
RULES, in order of importance:
1. If the report does not state a field, return null for it. Do not infer it, do not compute it, and do not use what you know about the world.
2. `is_serious` is a REGULATORY classification -- the seriousness criteria used across pharmacovigilance reporting -- and NOT a description of how bad the event sounds. Answer `yes` when the report describes ANY of the following having actually happened: the patient died; the event was life-threatening; the patient was admitted to hospital, or an existing admission was prolonged, because of the event; the event left a persistent or significant disability or incapacity; a congenital anomaly or birth defect followed exposure; or the event was medically important enough to require an intervention to prevent one of the outcomes just listed. Answer `no` when none of those is described.
3. DO NOT decide `is_serious` from the report's own severity wording. 'Severe', 'moderate' and 'mild' describe how the event felt and are NOT the regulatory test. A severe-sounding event that fully resolved at home with no medical attention is NOT serious. A mild-sounding event that led to a hospital admission IS serious. Read what actually happened to the patient and match it against the list in rule 2. Read the whole narrative before deciding: four of the six criteria are stated only in the narrative and never in a field of their own.
4. `causality_assessment` is a DIFFERENT judgment from `is_serious` and must not be conflated with it. Causality is the reporter's own stated view of whether the drug caused the event; seriousness is about how bad the outcome was. A case can be serious and unrelated at the same time. Report the reporter's own stated causality; do not form your own.
5. `narrative_severity_word` is the colloquial severity word the report itself uses, copied exactly as written. Return null when the report uses no such word -- do not supply one, and do not translate the classification back into a word.
6. Copy values verbatim from the report wherever possible.
7. Use the exact allowed value for a field that lists them.
8. Return every field named in the schema, even when the answer is null.
---
Extract these fields:
- case_id (string) -- the report's own case identifier, verbatim
- patient_age_range (string) -- the bucketed age range as stated (for example "45-54"). Never an exact age and never any other patient identifier -- if the report states a range, copy the range
- suspect_drug (string) -- the drug named as the suspect product
- event_description (string) -- the adverse event as described, copied verbatim
- narrative_severity_word (string) -- the colloquial severity word the report itself uses for the event -- "severe", "moderate", "mild" -- copied exactly as written, or null when the report uses no such word. It is recorded so a reader can see the report's own wording separately from the classification; it must NEVER be used to decide is_serious
- hospitalization (enum) one of: yes, no, unknown -- was the patient admitted to hospital, or was an existing admission prolonged, because of this event
- event_outcome (enum) one of: recovered, recovering, not-recovered, fatal, unknown -- the outcome of the event as stated in the report
- causality_assessment (enum) one of: related, possibly-related, unrelated, not-assessed -- the REPORTER's own stated view of whether the suspect drug caused the event. This is a DIFFERENT judgment from is_serious and must not be conflated with it -- causality is about whether the drug caused the event, seriousness is about how bad the outcome was. A case can be both serious and unrelated. Report the reporter's stated causality; do not form your own
- reporter_type (enum) one of: physician, pharmacist, consumer, other-healthcare-professional -- who reported the case
- is_serious (enum) one of: yes, no -- the REGULATORY seriousness classification, not a description of how bad the event sounds. Answer yes when the report describes ANY of these having actually happened: the patient died; the event was life-threatening; the patient was admitted to hospital or an existing admission was prolonged; the event left a persistent or significant disability or incapacity; a congenital anomaly or birth defect followed exposure; or the event was medically important enough to require an intervention to prevent one of the outcomes just listed. Answer no when none of those is described. Do NOT judge this by the report's own severity wording: a severe-sounding event that resolved at home with no medical attention is NOT serious, and a mild-sounding event that led to a hospital admission IS serious
Return a JSON object with exactly these keys: case_id, patient_age_range, suspect_drug, event_description, narrative_severity_word, hospitalization, event_outcome, causality_assessment, reporter_type, is_serious
Use null for any field the report does not state.
CASE REPORT
-----------
Case ID
-------
AE-2026-0011
Patient Age Range
-----------------
0-17
Suspect Drug
------------
Orbadanil
Event Description
-----------------
mild abdominal pain
Hospitalization
---------------
no
Event Outcome
-------------
recovered
Reporter Causality Assessment
-----------------------------
possibly-related
Reporter Type
-------------
consumer
Case Narrative
--------------
Within twenty minutes of the dose the patient deteriorated to the point that the treating clinician recorded the episode as immediately life-threatening and called for resuscitation support. The patient was stabilised in the treatment room over the following hour. No admission was required.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Catch serious drug side-effect reports, even when they sound mild — 55 case reports. Two tiers of one model family answered, and every answer was then graded Three different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
55case reports
55source documents
2model tiers
110graded answers
3grading methods
MeasurementsWhat was measured
COUNTED548 · 549 / 550extraction accuracy — cellsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED53 · 54 / 55seriousness accuracy — case reports classifiedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED15 · 14 / 15review-flag recall — cases that should have been flaggedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer (evals/judge.py) is pure code, exact match with light normalisation against a mechanically-derived gold — there is no judgement to validate, only comparison. What WAS validated: tools/build_corpus.py derives gold from the criterion branch that produced each narrative, never from a reading of its surface words, and its own _verify() pass asserts every gold value appears verbatim in its document AND that no severity word or form of "serious" leaked into an input that should not carry one. evals/check_labels.py asserts field completeness, enum membership, that the one nullable field exercises both states, and that the confusable slice is non-empty, before any run is allowed to spend. The TRUE review flag is derived by running the same pure-code compute() over gold's own values, never separately typed.
461.75output tokens · the fast tier · 3,538 ms p50
548.93output tokens · the deliberating tier · 7,084 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 2.0× as long, and lands one row apart on 55. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate — a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One case report
1,000 case reports
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.002037
$2.04
32%
Same work, 1× the bill
The same case reports, the same tokens — only the rate card changed. And on that card about 32% of what you pay is the prompt this pipeline sends, not the answer it writes.
which tier is called — and here the lever genuinely buys something and costs something at the same time. The deliberating tier classified one more case correctly and produced one more review-flag false negative, for about 13% more per case and double the latency. On this corpus the fast tier is the better buy; on 15 positive cases that is a preference, not a finding.
Rates checked 2026-08-18. The provider that actually ran r001 and r002 publishes no rate card this repo commits, so nothing here is what was actually paid — the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page.
The gradersThree ways to grade
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Per-field exact match, light normalisation Does the model's value for each field match gold, after trimming whitespace and punctuation — and, on the one nullable field, is a correct null scored as a hit rather than a miss?
$0.00
no
yes
the fast tier 99.6% · the deliberating tier 99.8%
Review-flag confusion matrix Does the run's own pure-code needs_review (regulatorily serious AND a causality assessment this kit routes on) match the same computation run over gold's own true values?
$0.00
no
yes
the fast tier 100.0% · the deliberating tier 93.3%
Seriousness accuracy, split by the report's own register On the one judgment field, is the model right BECAUSE it read the outcome, or only because the report's severity word happened to agree? Scored separately on the reports whose severity word points the wrong way.
$0.00
no
yes
the fast tier 96.4% · the deliberating tier 98.2%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, decisively between the models and the free floor, and only marginally between the two tiers. The floor separates cleanly: 0 of 25 against 23 of 25 and 24 of 25 on the confusable reports, and 0.47 flag recall against 1.00 and 0.93. Between the tiers the difference is a single case in each direction — the deliberating tier classified one more correctly, the fast tier's errors landed where the flag drops them — which is not a separation a 55-report corpus can resolve.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
If your situation is
Use
Because
And avoid
Screening an intake queue for which cases a case processor should work first
either tier over the free floor — but read Business.not_good_enough before trusting either flag figure
96% and 98% seriousness accuracy against the free severity-register floor's 49%, and 23 of 25 against 0 of 25 on the confusable reports the floor is built to fail.
the severity-register floor for the seriousness judgment specifically — it gets 25 of 25 confusable reports wrong, which is the whole population this kit exists to handle.
Deciding the two tiers on cost, speed or accuracy
the fast tier, with a caveat rather than a clean win
Roughly 11% cheaper per case ($0.0020372 vs $0.0022987), half the p50 latency (3538 ms vs 7084 ms), and 1.00 flag recall against 0.93. But the deliberating tier classified one MORE case correctly (54/55 vs 53/55), so the cheaper tier is not simply the better classifier — it is the one whose errors happened to land where the flag drops them.
reading either tier's flag score as a safety margin on 15 positive cases — see Business.not_good_enough.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
medically-important-criterion-missed
The one seriousness criterion that is a judgment rather than a fact
3
Every classification error either tier made is this criterion — "medically important enough to require an intervention to prevent one of the other outcomes". The fast tier missed AE-0027 and AE-0043; the deliberating tier missed AE-0028. All three narratives…
severity-register-shortcut
Reading the report's severity word instead of the outcome
28
Illustrated by the free severity-register floor (evals/baseline.py), NOT by either scored tier — neither model made this error on any report. The floor answers "does the report say severe?" and gets 25 of 25 confusable reports wrong in both directions: a…
What we could NOT verify
Whether either tier's flag recall means anything at this sample size. 15 positive cases is a small denominator, and this run showed exactly how fragile the figure is: the fast tier's 1.00 recall survived two classification errors only because both landed on cases whose causality assessment the flag drops. See Business.not_good_enough.
Whether the "medically important enough to require an intervention to prevent one of the other outcomes" criterion is genuinely harder for a model or merely harder as this generator words it — all three of this run's errors are that criterion, but it is also the criterion whose narratives were written to state an intervention and a non-admission in the same sentence.
How either tier behaves when the deciding fact is absent. hospitalization: unknown and event_outcome: unknown are allowed values this corpus never exercises, by construction.
Whether the criterion mix generalises. The six criteria are sampled UNIFORMLY here for coverage, so that a systematic miss on any one of them is visible; a real safety database is dominated by hospitalisation, which is one of the two criteria stated in a structured field and therefore among the easiest.
Whether CAUSALITY_TRIGGERS resembles any real pharmacovigilance programme's case-processing criteria — it does not, by design (see README). A real expedited reporting obligation is a clock with a defined start event and a defined day count, and none of that is modelled.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
1,303.93
461.75
3,538 ms
$0.002037
the deliberating tier
1,303.93
548.93
7,084 ms
$0.002299
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The severity-register floor (evals/baseline.py) and the scorer (evals/judge.py) are both pure code and cost $0.00 to run against either result set. The figure above is both runs' own token counts (the fast-tier and deliberating-tier runs, 55 reports each) priced at Google Gemini 3 Flash's published rate — the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The fixed system prompt and field schema (1117 of 1313 tokens on the example call, 85%) outweigh the case report sent (196 tokens) many times over. Stating all six seriousness criteria in full is what makes it that large — and the free floor that skips them gets 28 of 55 classifications wrong, so this is the floor every call pays to get the judgment right.
Output length: the model returns a full ten-key JSON record every call regardless of the narrative's own length.
Your volumeWhat it costs at your volume
Linear in cases: each call is independent and self-contained, with no shared context or index to amortise. This run's 55 reports cost about $0.112 projected onto Gemini 3 Flash's rate, so ten times the set is about $1.12 on the same rate and the same prompt — arithmetic on the measured per-call rate, not a second run.
Where pricing changes shape
Your return, with your numbers
Volume
What it replaces
Time saved per item
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
71,716input tokens · this run
25,396output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 55 adverse-event case reports extracted, each scored by pure code against a mechanically-derived gold set. This run (r001-ae-seriousness, the fast tier) answered all 55 of 55 reports with no truncation -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.045
$0.045
$0.81
2026-09-12
gemini-3-flash
Google
$0.112
$0.112
$2.04
2026-09-18
gemini-3-8-flash
Google
$0.149
$0.149
$2.71
2026-09-18
llama-5
Meta
$0.198
$0.198
$3.59
2026-09-18
claude-haiku-4-5
Anthropic
$0.199
$0.199
$3.61
2026-09-12
grok-4-5
xAI
$0.296
$0.296
$5.38
2026-09-18
grok-4-6
xAI
$0.296
$0.296
$5.38
2026-09-18
claude-sonnet-5
Anthropic
$0.397
$0.397
$7.23
2026-09-12
gemini-3-1-pro
Google
$0.448
$0.448
$8.15
2026-09-18
gpt-5-6-terra
OpenAI
$0.448
$0.448
$8.15
2026-09-12
gpt-5-6-sol
OpenAI
$0.795
$0.795
$14.45
2026-09-12
claude-opus-4-8
Anthropic
$0.993
$0.993
$18.06
2026-09-12
claude-opus-5
Anthropic
$0.993
$0.993
$18.06
2026-09-12
claude-fable-5
Anthropic
$1.987
$1.987
$36.13
2026-09-18
claude-fable-5-1
Anthropic
$1.987
$1.987
$36.13
2026-09-18
gpt-6-astra
OpenAI
$1.987
$1.987
$36.13
2026-09-17
Read this against the numbers above
Every row below prices the FAST TIER's own 55-call run (r001-ae-seriousness) -- the deliberating tier's own token counts are on Cost.cost_by_model[1] and are not separately projected here.
Neither tier's registered run left anything to disable -- src/adapters/__init__.py's thinking parameter is only sent when a caller passes one, and this kit's own harness never does (see LLM.settings) -- so unlike several sibling kits, there is no reasoning-on/reasoning-off discrepancy to caveat here.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Five modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Four of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/segment.pysegment
cut the case report into addressable sections, pure code
src/segment.py
# Cut a document into addressable sections. Pure code — no model, no network.
def sections(text):
def locate(text, value):
def span_label(secs, start):
src/select.pyselect — a swap seam
pick which sections carry each field, pure code
You change it to: map fields to your own report's headings; unmatched falls back to the whole document
src/select.py
# Pick which sections plausibly carry each field. Pure code -- the last deterministic step
SECTION_HINTS = {
def for_field(secs, field):
def plan(secs, fields):
src/prompt.pyprompt
assemble one call for all ten fields, with the six seriousness criteria stated in full
src/prompt.py
# Assemble the extraction prompt. One prompt per case report, all ten fields in it.
SYSTEM = (
def field_schema(fields):
def build(doc_text, secs, fields, selector):
def parse(raw, fields):
src/extract.pyextract — a swap seam
the AI layer, one provider one key — plus the pure-code routing computation downstream
You change it to: this kit's own flat routing rule — replace with your own programme's actual documented criteria before trusting the flag for anything real. It is not a reporting clock and does not model one.
src/extract.py
# Extract one case report's fields: segment, select, prompt, one model call, then a pure-code
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
FIELDS = os.path.join(HERE, "data", "fields.json")
CORPUS = os.path.join(HERE, "data", "corpus")
MAX_TOKENS = 4000
CAUSALITY_TRIGGERS = ("related", "possibly-related")
def load_fields():
def load_doc(stmt_id):
def documents():
def compute(values):
evals/judge.pyjudge
score field accuracy, the flag's recall/precision, and seriousness accuracy by register — three separate figures, pure code
evals/judge.py
# Score an extraction run. PURE CODE -- gold is exact and the answer is one value per cell, so
def norm(v):
def _num(v):
def equal(field, got, want):
def score(fields, records, golds):
def score_flags(flags, golds):
def score_seriousness(records, golds):
Start hereThe shortest path into it
src/segment.pycut the case report into addressable sections, pure code
src/select.pypick which sections carry each field, pure code A swap seam.
src/prompt.pyassemble one call for all ten fields, with the six seriousness criteria stated in full
src/extract.pythe AI layer, one provider one key — plus the pure-code routing computation downstream A swap seam.
evals/judge.pyscore field accuracy, the flag's recall/precision, and seriousness accuracy by register — three separate figures, pure code
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1303 input and 461 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's reports are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote any narrative. In a real deployment a case narrative arrives from a physician, a pharmacist, a consumer or a call-centre transcription -- exactly the kind of externally-authored input this kit's architecture reads verbatim and trusts, with no verification step before its text reaches the model. No attack has been tried against this kit; see redteam.why.
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/extract handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
The experimentWe did not attack it — and it failed on ordinary input anyway
An indirect prompt injection needs a field an outside party controls that reaches the prompt. In THIS corpus every report is generated by tools/build_corpus.py, so there is no live untrusted text in the run this report measures -- but a real deployment's narratives arrive from reporters, and that surface has not been attacked. The four gates below are boundaries confirmed by reading the code, not payloads run through it. The fourth is the one to read: it did not need an attacker. Confirmed by reading the code, not by a run, on 2026-08-21 -- no attack run was fired; see redteam.why.
Boundary checked
What could go wrong
What the code guarantees
Does an extraction ever file a report, start a reporting clock, or trigger a downstream action?
An extraction could plausibly submit a case, start an expedited clock, or write a seriousness determination.
No code path does. src/extract.py::extract() and src/app.py's /api/extract both return a record only; neither writes to any file or store -- confirmed by reading every call site.
Could a misconfigured provider key leak into a UI-visible error?
An exception raised from a bad key or base URL could echo the secret back to the browser.
src/app.py's /api/extract handler strips api_key and base_url out of any exception message before returning it -- confirmed by reading the handler (see key_handling).
Is the routing rule something a prompt or a reply can move?
A crafted narrative could plausibly shift which causality values trigger, or talk the arithmetic into a different answer.
No. CAUSALITY_TRIGGERS in src/extract.py is a module-level constant read once at import time -- nothing the model returns is consulted when deciding what triggers; needs_review is computed after the model call, in pure code.
Could a narrative talk the model into is_serious='no' despite a qualifying outcome being described?
A narrative describing a real qualifying outcome in routine, self-limiting language could move is_serious to 'no', clearing a case that should have been routed.
MEASURED, AND IT HAPPENED -- without any attacker. Three of this corpus's ordinary, non-adversarial reports did exactly this on the two tiers (AE-0027 and AE-0043 on the fast tier, AE-0028 on the deliberating tier), all three describing an intervention given to prevent a worse outcome alongside the words 'no admission'. is_serious is an enum with no span at all (see Eval.scores.non_spannable_fields), so nothing in this kit can check it against the narrative. What is unmeasured is how much WORSE this gets against a narrative written deliberately; see could_not_verify.
The first three boundaries hold, confirmed by reading the code, not by an attack trial. The fourth does not hold even on ordinary input: is_serious is the field the whole guardrail depends on, it is the one field this kit's span mechanism cannot check because it is an enum rather than a copied value, and three plain reports moved it the wrong way on this run. No crafted input has ever been tried against it.
The result0 attack trials, and three of four boundaries hold in code: no write path exists, the routing rule is a fixed constant the model cannot move, and a misconfigured key cannot leak into a UI error. The fourth failed on ordinary input -- three plain reports moved is_serious the wrong way with no attacker involved.
1externally-authored field a live deployment would carry (the case narrative itself, from a physician, pharmacist, consumer or transcription) -- synthetic on this run's corpus
0 of 0attack trials run
3 of 110ordinary, non-adversarial reports that moved is_serious the wrong way anyway
This run's corpus is entirely generated -- no report's text was authored by an outside party. A real deployment's narratives arrive from reporters and transcriptions, exactly the kind of externally-supplied text this kit reads verbatim and trusts. The number worth reading is the third: before any attacker is involved, three plain reports already talked both tiers out of a correct 'yes' by pairing an intervention with the words 'no admission'. Whether a narrative crafted on purpose could do this reliably is unmeasured.
Read this twice
The review flag is exactly as good as is_serious, and this run measured what that costs. src/extract.py::compute() never re-derives that field from the narrative -- it trusts the model's own answer and only checks the causality condition alongside it. Both tiers got is_serious wrong on this corpus: the fast tier twice, the deliberating tier once. The fast tier still scored 1.00 flag recall, and the ONLY reason is that both of its errors landed on cases whose reporter had said 'unrelated' and 'not-assessed' -- values the flag drops regardless. The deliberating tier's single error landed on a 'related' case and became a false negative. Do not read the perfect number as a safety margin. A future version should recompute is_serious independently before trusting a live flag, and a red-team pass targeting exactly that field is the natural next measurement.
HonestyWhat this does not prove
Whether a narrative crafted on purpose could move is_serious reliably -- no red-team run exists for this kit, and the three natural failures this run found suggest the surface is real rather than theoretical.
Whether the live app's own /api/extract behaves identically to the registered run under an adversarial report -- both use the same src/extract.py::extract(), but neither has been tested against one.
Whether a code-level consistency check on is_serious (re-deriving the six criteria from the narrative by pure code) would catch a model's wrong answer in practice -- none has been built or exercised.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineThe guardrail is enforced in code
Not a prompt instruction the model can ignore — a comparison applied after the reply is parsed, before anything downstream sees it.
the guardrail, verbatim
A case is routed for expedited review when the model's own is_serious is 'yes' AND its own causality_assessment is one of CAUSALITY_TRIGGERS ('related', 'possibly-related') — computed in pure code from the model's own extracted fields, never asked of the model directly, never overridden by anything it returns.
src/extract.py::compute() -- CAUSALITY_TRIGGERS is a module-level constant read once at import time. Unlike a prompt instruction the model could ignore, this guardrail cannot be talked out of firing by anything in the reply text; it only depends on the model's own is_serious and causality_assessment being right in the first place -- a DIFFERENT, prompt-level dependency the flag inherits, and the one this run found errors in (see is_not).
EvidenceDoes it hold?
What
Measured
The routing rule is not something a prompt or a reply can move
CAUSALITY_TRIGGERS in src/extract.py is a module-level constant, read once at import time -- confirmed by reading compute(); nothing the model returns is consulted when deciding what triggers.
A missing or out-of-vocabulary input returns None, never a False
compute() returns None when is_serious is not yes/no or causality_assessment is outside its four allowed values -- 'we could not decide' and 'we decided no' are different states and route differently. Confirmed by reading the function; the UI renders the third state as its own sentence.
The flag never fired on a case that should not be flagged
0 false positives among the 40 cases that should not have been flagged, on both tiers (1.00 precision).
No code path files a report, starts a clock, or writes a determination based on the flag
src/extract.py::extract() and src/app.py's /api/extract both return a record only; neither writes to any file or store -- confirmed by reading every call site.
The limitWhat a guardrail is not
IT IS NOT A CHECK ON is_serious ITSELF, AND THIS RUN SHOWED WHY THAT MATTERS. The flag trusts the model's own answer completely; nothing re-derives it from the narrative before computing. Both tiers made classification errors on this corpus -- two on the fast tier, one on the deliberating tier, all three the same criterion -- and the ONLY reason the fast tier still scored 1.00 recall is that both of its errors landed on cases whose causality assessment the flag drops anyway. The deliberating tier's single error did land on a triggering case and became a false negative. A guardrail that inherits an upstream field's accuracy is exactly as good as that field, and on this run that was visible.
It does not check whether the reporter's stated causality is REASONABLE -- it copies it. A case whose reporter wrongly recorded 'unrelated' is dropped by this rule and nothing here notices.
It is not a regulatory reporting deadline and does not model one. A real expedited reporting obligation is a clock with a defined start event and a defined day count; CAUSALITY_TRIGGERS is this kit's own flat simplification. See README and data/SOURCES.md.
It has not been attacked. Whether a narrative crafted to describe a qualifying outcome in routine, self-limiting language could suppress is_serious or needs_review is unmeasured -- see Security.could_not_verify.
WatchedWhat is watched, and why that one
2runs recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 17 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
8 measured by the latest run9 need the model half
Metric
Owner
Role
Why this one
field-exact-match-with-normalisation
Per-field exact match, light normalisation
alarm
extraction_accuracy specifically on is_serious, since that is the one field this corpus is built to test — the other nine are near-solved by regex, see Eval.baseline; refusal_accuracy on narrative_severity_word, since 7 of 55 reports use no severity word and inventing one there is a hallucination the aggregate would hide; span_rate on the five spannable fields, since a value with no span is an assertion rather than a located citation — alarm on Any drop in extraction_accuracy below the measured figures on either tier, or any non-zero hallucination count on narrative_severity_word.
flag-confusion-matrix
Review-flag confusion matrix
alarm
false_negative count specifically — a missed serious, plausibly-related case is the failure a safety team cares about most; WHERE a classification error lands, not just how many there are: on this run the fast tier made two errors and scored perfect recall while the deliberating tier made one and scored 0.93, because only the deliberating tier's error fell on a triggering case; whether recall holds as the corpus grows — 15 positive cases is a small sample for a rule this consequential — alarm on Any false negative — a case gold says should route for expedited review that the run's own computed flag did not catch.
seriousness-by-register
Seriousness accuracy, split by the report's own register
alarm
accuracy on the confusable slice specifically — an aggregate that mixes it with the matching slice hides the entire effect being measured; the gap between the confusable slice and the free floor's 0 of 25 there, which is the whole demonstration — alarm on Confusable-slice accuracy falling towards the matching-slice figure's mirror image — that is what reading the severity word instead of the outcome looks like.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
55
different corpus — nothing is comparable
corpus.bytes
33,163
case reports edited — the count held, the bytes did not
split.count
495
the sections count moved — a different set was scored
split.size_p50
40
the median size of one section moved
split.size_p95
244
the 95th-percentile size of one section moved
dataset.rows
55
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.009
no index is built here — a change is in corpus preparation, not an index rebuild
all shared guards held between the recorded runs (documents 55, extraction_cells 550, failures 0, refusal_cells 7, thinking True) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
Extraction accuracy
99.64 pct on the fast tier, 99.82 pct on the deliberating tier
550 cells
two independent tiers (r001-ae-seriousness, r002-ae-seriousness); every miss on both tiers is the is_serious field, never one of the other nine.
Seriousness accuracy, confusable register
23 of 25 on the fast tier, 24 of 25 on the deliberating tier, against 0 of 25 on the free floor
25 reports whose severity word points against their true classification
two independent tiers plus the free severity-register floor, scored by the same function.
Review-flag recall/precision
1.00 / 1.00 on the fast tier, 0.93 / 1.00 on the deliberating tier -- the tiers DISAGREE here, on one case
15 cases that should be flagged, 40 that should not
two independent tiers; the deliberating tier's one false negative is AE-0028, the same criterion as every other error this run found.
Hallucinations
exact match at 0 on both tiers, including on the 7 reports where the correct answer is null
550 cells
two independent tiers reproduced zero to the digit; refusal accuracy is 1.00 on both.
Correct abstention
exact match at 1.00 on both tiers -- null returned on every report that states no severity word, and a value invented on none
7 cells where the report states nothing and null is the right answer
two independent tiers reproduced the identical figure. Banded on its own rather than folded into extraction accuracy because it is the branch the sibling kits' scorer never exercised -- it marked every null a miss before looking at gold, which on this corpus would have cost the model 7 cells it got right.
Span rate
100 pct on both tiers
spannable extracted values (non-enum fields with a non-null answer)
two independent tiers reproduced the identical figure.
Latency
3538 ms / 9961 ms p50/p95 on the fast tier, 7084 ms / 19770 ms on the deliberating tier -- roughly double, because the two tiers are different models
55 calls per tier
measured directly on both tiers.
Token totals
71,716 input tokens on both tiers (identical prompt); 25,396 output on the fast tier vs 30,191 on the deliberating tier, about 19 pct more
55 calls per tier
measured directly on both tiers.
HistoryRun history
2 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-ae-seriousness 2026-08-21
r002-ae-seriousness 2026-08-21
extraction accuracy
0.9964
0.9982
refusal accuracy
1.000
1.000
invented values
0
0
values with a span
1.000
1.000
input tokens, whole run
71716
71716
model latency p50 ms
3538.00
7084.00
model latency p95 ms
9961.00
19770.00
output tokens, whole run
25396
30191
not a time series No two of these 2 runs measured the same system — they differ on the model — so nothing here may be differenced and the Verdict column is left off. Read this block DOWN a column, as a comparison between those runs, not along a row as a history.
DeviationsWhat deviated
0 breaches across 2 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
which tier is called (fast vs deliberating)
output tokens (+19 pct) and latency (roughly double) move together; hallucinations and span rate do not move at all; seriousness accuracy improves by one case while review-flag recall gets WORSE by one case -- the two move in opposite directions
measured
r001-ae-seriousness vs r002-ae-seriousness: 53/55 vs 54/55 seriousness, flag recall 1.00 -> 0.93, output tokens 25,396 -> 30,191, p50 latency 3538 ms -> 7084 ms.
a case the reporter called unrelated or not-assessed
needs_review resolves false however serious the case is -- and, on this run, that is also what rescued the fast tier's recall figure, because both of its classification errors landed there
measured
src/extract.py::compute() sets needs_review True only when BOTH is_serious=='yes' AND causality_assessment is in CAUSALITY_TRIGGERS. 16 of this corpus's 31 serious cases are serious-but-not-triggering, and AE-0027 (unrelated) and AE-0043 (not-assessed) are two of them.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Nothing fires yet. Not one of the 8 bands on this kit names a consequence, so there is no row to print. Which is the honest answer to the heading: a band that nothing is wired to is a number somebody has to notice, and this kit has only those. An empty table would have said the same thing while looking like a rendering failure.
NextThe three you would add first
Re-derive is_serious from the narrative by pure code and compare it to the model's own answer before trusting the flagthis run found three classification errors and the flag caught none of them -- it cannot, because it never looks at the narrative. A pattern check for the six criteria would at least raise a disagreement, and it would need to be stricter than the free severity-register floor this kit already ships (which gets 25 of 25 confusable reports wrong) to add anything.
Treat the "medically important enough to require an intervention to prevent one of the other outcomes" criterion as its own check rather than one clause in a listevery classification error either tier made on this corpus is that one criterion, and both models failed it the same way -- reading 'no admission was required' and stopping. See Eval.taxonomy.
Reconcile a follow-up report against the case it amends before trusting a stored classificationthe flag is computed from one report's own narrative alone; a real intake queue is full of follow-ups that change a case's seriousness after it was first assessed, and nothing here models that.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-check the flag's evidence on any change to CAUSALITY_TRIGGERS in src/extract.py, or to the is_serious rule in src/prompt.py -- either changes what the flag is computed FROM. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
What this cannot tell you
Whether the flag would hold its recall on a larger set. 15 positive cases is a small denominator, and this run demonstrated the figure's fragility directly rather than in theory -- see is_not.
Whether a crafted narrative could move is_serious without the flag noticing -- no red-team run exists for this kit, see Security.could_not_verify.
Whether CAUSALITY_TRIGGERS resembles any real programme's documented criteria -- it does not, by design; see README/SOURCES.md.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is empty, with a comment explaining that the emptiness is load-bearing. The whole extraction decision is three files: src/segment.py, src/select.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
55 case reports, generated from a fixed seed, never fetched. Gold comes from the criterion branch that produced each narrative, never from a reading of its words -- and _verify() refuses to write a corpus that leaks a register word or the label's own vocabulary into an input; see data/SOURCES.md.
segmentation and selection
src/segment.py, src/select.py
text splitters / retrievers
a heading-based cut and a fixed field-to-heading map (SECTION_HINTS) -- no embeddings, no index, no ranking. Two fields deliberately get more than one section: narrative_severity_word and is_serious are the pair this kit keeps apart, and narrowing either to one section would hand the model the half-view that produces the mistake being measured.
prompt assembly
src/prompt.py
prompt templates
the ten-field schema and all six seriousness criteria are one declaration, and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. MAX_TOKENS=4000 is a plain constant, not a client-library setting.
evaluation
evals/judge.py
eval harnesses
per-field exact match, a review-flag confusion matrix, and a seriousness accuracy split by register -- three loops and a handful of counters, not a platform. The one non-obvious line is that a correct null scores as a hit; the sibling kits' scorer marked every null a miss, a branch no sibling corpus ever exercised.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per report -- segment, select, prompt, call, parse, compute -- with no branching and no state carried between reports. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A different field schema (data/fields.json) or a different routing rule needs its own gold set and its own evals/check_labels.py pass -- a framework's own schema layer would not remove that work, only relocate it.
SECTION_HINTS is a five-minute edit for a new report layout because it is a plain dict, not a configured retriever -- a framework's chunking/retrieval abstraction would need its own re-tuning pass instead, with its own failure modes to learn.
Swapping providers is one function and one entry in PROVIDERS -- a framework's model abstraction would add a dependency and a version to track for the same one-line change this file already gives away free.
What we could NOT verify
Whether a framework's retrieval or agent abstraction would resolve the one error mode this run found -- all three misses are a criterion the prompt already states in full, so the gap is in reading the narrative, not in fetching more of it. Nothing here tests whether a second pass or a self-check step would catch them.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-ae-seriousness on the fast tier, 2026-08-21. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
3,538 ms
3538 ms / 9961 ms p50/p95 on the fast tier, 7084 ms / 19770 ms on the deliberating tier -- roughly double, because the two tiers are different models
—
Model, p95
9,961 ms
3538 ms / 9961 ms p50/p95 on the fast tier, 7084 ms / 19770 ms on the deliberating tier -- roughly double, because the two tiers are different models
—
Input tokens
71,716
71,716 input tokens on both tiers (identical prompt); 25,396 output on the fast tier vs 30,191 on the deliberating tier, about 19 pct more
—
Output tokens
25,396
71,716 input tokens on both tiers (identical prompt); 25,396 output on the fast tier vs 30,191 on the deliberating tier, about 19 pct more
—
No movement column. Not one of the 1 earlier run on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
DriftMedian call, run over run
r001-ae-seriousness3,538 ms
r002-ae-seriousness7,084 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — each document goes whole into one call.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-21, across 2 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
case reports
data/corpus/*.txt — 55 files, generated once from a fixed seed (SEED = 20260821) by tools/build_corpus.py
read whole by src/segment.py and src/select.py; never modified after generation
gold labels
data/gold.jsonl — 55 rows, derived from the criterion branch that produced each narrative (never re-read off its surface words), and asserted to appear verbatim in the document by tools/build_corpus.py's own _verify() pass
never — evals/judge.py is pure code, no model, no key
the field schema
data/fields.json — the ten-field record src/prompt.py assembles the user message from
read by src/prompt.py and src/extract.py only
the key
.env — never committed (see .gitignore)
only inside the Authorization header, to the configured BASE_URL (src/adapters/__init__.py)
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never requested from a reader on any surface. src/app.py's /api/extract handler strips api_key and base_url out of any exception message before it reaches the browser, so a misconfigured key cannot leak into a UI error.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
model
one call per CASE REPORT, carrying the selected sections plus the ten-field schema, behind src/adapters/__init__.py — OpenAI-compatible wire format over raw HTTP, so a forker runs this on whichever key they hold. MAX_TOKENS is fixed at 4000 (src/extract.py), set on the sibling extraction kits' experience with a similarly-shaped JSON record rather than a ceiling this kit has ever hit — 0 of 55 calls truncated on either tier.
3538 ms p50 / 9961 ms p95 on the fast tier, 7084 ms p50 / 19770 ms p95 on the deliberating tier (the fast-tier and deliberating-tier runs, 2026-08-21 -- see Cost.cost_by_model)
one call per report, no concurrency and nothing shared between calls — see Architecture.breaks_at_scale. A day's real intake needs batching and a rate-limit strategy this kit does not have; MAX_CALLS_PER_DAY in src/budget.py caps the shared key across every kit on this machine, not this kit's own throughput.
point src/adapters/__init__.py at a different provider or model and every published accuracy/cost/latency figure is void until re-run — this run's numbers are this model's numbers, not a property of the prompt.
corpus refresh
tools/build_corpus.py regenerates the whole corpus — 55 case reports across six seriousness criteria sampled uniformly — byte-identically from a fixed seed (SEED = 20260821) every time. There is no incremental refresh. Gold comes from the criterion branch that produced each narrative, and the script's own _verify() pass refuses to write a corpus where a gold value is not stated verbatim, where a severity word leaked into a report that should carry none, or where the label's own vocabulary ("serious") appears in an input at all.
under 0.05s wall time to regenerate all 55 reports and their gold labels (measured directly, 2026-08-21 (python3 tools/build_corpus.py, timed) -- see Data.index for the separate segmentation figure, a different step)
a real safety database's intake volume, product roster and narrative vocabulary do not come from a fixed seed. How a genuinely varied set of real reports would change what src/segment.py (heading-based) and src/select.py (SECTION_HINTS) can resolve was not measured; see Architecture.breaks_at_scale and Data.breaks_on.
point tools/build_corpus.py at your own case reports and every published accuracy figure is void — they are this corpus's own planted ambiguity, not a property of the model.
labels
evals/check_labels.py asserts gold/document consistency, enum membership, that the one nullable field exercises BOTH states, and that the confusable slice is non-empty, before evals/run.py is allowed to spend anything. Scoring (evals/judge.py) is per-(case,field) exact match with light normalisation — with a correct null scored as a hit rather than a miss — plus a separately-scored review-flag confusion matrix and a separately-scored seriousness accuracy split by register, both computed by pure code, never by a second model call.
550 of 550 possible (case, field) cells scored (55 reports x 10 fields), plus 55 review-flag verdicts and 55 seriousness verdicts scored separately (lenses.Eval.dataset, the fast-tier and deliberating-tier runs, 2026-08-21)
the gold set stops at 55 reports and covers one planted ambiguity (a severity word pointing against the true classification) at a 45 pct rate — a real queue's mix of confusable phrasing is unmeasured, and a follow-up report is never reconciled against the case it amends; see Data.breaks_on.
a different field schema (data/fields.json) or a different routing rule needs its own gold set and its own evals/check_labels.py pass before any published figure can be trusted again.
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
is_serious answered 'yes' on a report whose own severity word is "mild" or "moderate", with Hospitalization 'no' and a non-fatal outcome
the model read the narrative for which seriousness criterion was actually met rather than taking the three structured signals at face value — this is the corpus's planted ambiguity (AMBIGUOUS_FRACTION=0.40 in tools/build_corpus.py), and on this run the fast tier got 23 of 25 such reports right and the deliberating tier 24 of 25
read the narrative's own middle sentence before assuming a 'yes' next to "mild" is a mistake — four of the six seriousness criteria are stated only in the narrative, under a hospitalization: no and a non-fatal outcome, and the free floor that reads only the severity word gets every one of these reports wrong (the fast-tier and deliberating-tier result files, both runs, 2026-08-21 -- see Eval.taxonomy and Eval.baseline)
is_serious answered 'no' on a narrative that says an intervention was given specifically to stop the reaction progressing, and that no admission followed
this is the one error mode this run actually found, on BOTH tiers — the "medically important enough to require an intervention to prevent one of the other outcomes" criterion. All three classification errors across 1100 combined cells are this criterion and no other.
check whether the narrative names an intervention given to PREVENT one of the other five outcomes; that is itself a seriousness criterion, and 'no admission was required' in the following sentence does not cancel it. See Eval.taxonomy. (AE-0027 and AE-0043 on the fast tier, AE-0028 on the deliberating tier, 2026-08-21)
No machine symptom — this failure leaves no trace in any output.
no path in src/extract.py::compute() or src/app.py files a report, starts a reporting clock, or writes a seriousness determination anywhere — needs_review is returned to the caller as a field on the response record and nothing downstream of this kit acts on it. A flipped flag on a real deployment would show up only in whatever safety system consumes this kit's output, which this kit does not have and does not simulate — so there is no committed artifact naming that failure, on purpose: it is out of this kit's boundary, not unmeasured.
Whether either tier's flag figure survives a larger sample was not tested — see Eval.could_not_verify, and note that the fast tier's perfect recall survived two classification errors purely because of where they landed. Concurrency (every run in this series is one call at a time, sequential), CAUSALITY_TRIGGERS against any real programme's documented criteria, behaviour when the deciding fact is absent (unknown is an allowed value this corpus never exercises), and how either tier performs on a real intake feed with follow-up reports are all unmeasured — see Eval.could_not_verify and Data.breaks_on.
The corpus licence, from the Data lens: MIT — this repository's own licence. Every case identifier, drug name and narrative is invented; a real adverse-event narrative cannot be published either — it is among the most identifying text a company holds. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one linePer-field exact match, light normalisation
Does the model's value for each field match gold, after trimming whitespace and punctuation — and, on the one nullable field, is a correct null scored as a hit rather than a miss?
$0.00per 1,000 case reports
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score, in-process, no key and no model. The same function evals/run.py calls after every call, and the same field-match logic the free floor is scored by — a baseline and a real run scored by two different scorers cannot be compared honestly.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
doc
AE-0011
field
is_serious
narrative severity word
mild
hospitalization
no
event outcome
recovered
causality assessment
possibly-related
is serious model
yes
gold is serious
yes
needs review
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
AE-0011 states "mild abdominal pain" as the event and its narrative reads "Within twenty minutes of the dose the patient deteriorated to the point that the treating clinician recorded the episode as immediately life-threatening and called for resuscitation support. The patient was stabilised in the treatment room over the following hour. No admission was required." Gold: is_serious=yes. Both tiers answered yes, and all ten fields matched exactly.
Review-flag confusion matrix
correct
Gold needs_review=true (is_serious=yes, causality_assessment=possibly-related). Both models' own extracted values fed src/extract.py::compute(), which computed needs_review=true on both tiers — routed for expedited case-processor review, as it should be.
Seriousness accuracy, split by the report's own register
correct
AE-0011 is in the confusable register: the report calls the event "mild", the Hospitalization field says no and the Event Outcome says recovered — every structured signal points at not serious, and only the narrative's life-threatening episode says otherwise. The free severity-register floor calls this one not serious and records it as a false negative; both model tiers got it right.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 99.6%
the deliberating tier
scored 99.8%
In operationWhat to monitor
Reference standard: tools/build_corpus.py's gold, derived from the criterion branch that produced each narrative and asserted to appear verbatim in the document; evals/check_labels.py asserts field completeness, enum membership and both states of the nullable field before any run is allowed to spend.
These rates are UNKNOWN, on purpose
This grader IS the reference standard for field values — it cannot be scored against itself. What can go wrong is the corpus's own generation logic, which tools/build_corpus.py's _verify() pass checks by confirming every gold value appears verbatim and that no register word leaked into a report that should not carry one.
Watch these
extraction_accuracy specifically on is_serious, since that is the one field this corpus is built to test — the other nine are near-solved by regex, see Eval.baseline
refusal_accuracy on narrative_severity_word, since 7 of 55 reports use no severity word and inventing one there is a hallucination the aggregate would hide
span_rate on the five spannable fields, since a value with no span is an assertion rather than a located citation
Alarm on
Any drop in extraction_accuracy below the measured figures on either tier, or any non-zero hallucination count on narrative_severity_word.
How tight can the band be? There is no tolerance band on the field grade — it is exact match after trimming whitespace and punctuation, never a continuous score to round.
Cadence: Re-score on any change to src/prompt.py, data/fields.json or tools/build_corpus.py — the first two change what is asked, the third changes what is asked about. Re-run evals/run.py (paid) on any change to MAX_TOKENS or the provider/model.
The decisionWhen to reach for it
Use it
Gold is derived from the generator's own criterion branch and asserted against the document text — true of every kit corpus, never true of a real safety database's own case history.
Do not use it
The true field values are not known in advance — the normal state of a real intake queue, and the reason this corpus is generated rather than captured.
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineReview-flag confusion matrix
Does the run's own pure-code needs_review (regulatorily serious AND a causality assessment this kit routes on) match the same computation run over gold's own true values?
$0.00per 1,000 case reports
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_flags, in-process, no key and no model. Runs src/extract.py::compute() over both the run's own extracted values and gold's own values, then compares the two booleans per case.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
doc
AE-0011
field
is_serious
narrative severity word
mild
hospitalization
no
event outcome
recovered
causality assessment
possibly-related
is serious model
yes
gold is serious
yes
needs review
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
AE-0011 states "mild abdominal pain" as the event and its narrative reads "Within twenty minutes of the dose the patient deteriorated to the point that the treating clinician recorded the episode as immediately life-threatening and called for resuscitation support. The patient was stabilised in the treatment room over the following hour. No admission was required." Gold: is_serious=yes. Both tiers answered yes, and all ten fields matched exactly.
Review-flag confusion matrix
correct
Gold needs_review=true (is_serious=yes, causality_assessment=possibly-related). Both models' own extracted values fed src/extract.py::compute(), which computed needs_review=true on both tiers — routed for expedited case-processor review, as it should be.
Seriousness accuracy, split by the report's own register
correct
AE-0011 is in the confusable register: the report calls the event "mild", the Hospitalization field says no and the Event Outcome says recovered — every structured signal points at not serious, and only the narrative's life-threatening episode says otherwise. The free severity-register floor calls this one not serious and records it as a false negative; both model tiers got it right.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 93.3%
In operationWhat to monitor
Reference standard: The same pure-code compute() in src/extract.py, run over gold's own is_serious and causality_assessment values rather than the model's — the TRUE needs_review is never separately typed, only derived.
These rates are UNKNOWN, on purpose
Whether the flag generalises to cases this corpus did not plant. The confusion matrix is exact against these 55 reports and says nothing about a register mismatch harder than the one built here. See Eval.could_not_verify.
Watch these
false_negative count specifically — a missed serious, plausibly-related case is the failure a safety team cares about most
WHERE a classification error lands, not just how many there are: on this run the fast tier made two errors and scored perfect recall while the deliberating tier made one and scored 0.93, because only the deliberating tier's error fell on a triggering case
whether recall holds as the corpus grows — 15 positive cases is a small sample for a rule this consequential
Alarm on
Any false negative — a case gold says should route for expedited review that the run's own computed flag did not catch.
How tight can the band be? No tolerance band — a binary flag either matches gold's derived flag or it does not, scored per case.
Cadence: Re-score on any change to CAUSALITY_TRIGGERS, src/extract.py::compute, or the prompt's is_serious instruction — all three change what the flag means.
The decisionWhen to reach for it
Use it
Gold's true needs_review is derived by running the same compute() the kit itself uses — never a separately-typed truth that could drift from the rule the kit actually applies.
Do not use it
A real queue's true urgency is not known until a case processor has worked the case; this corpus's gold is a construction, not an observed outcome.
Seriousness accuracy, split by the report's own register
Catch serious drug side-effect reports, even when they sound mild
PresenterOpens the private repo. Visible to admins only.
In one lineSeriousness accuracy, split by the report's own register
On the one judgment field, is the model right BECAUSE it read the outcome, or only because the report's severity word happened to agree? Scored separately on the reports whose severity word points the wrong way.
$0.00per 1,000 case reports
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/judge.py::score_seriousness, in-process, no key and no model. Groups every case's is_serious verdict by the register tag the generator recorded.
Every grader on these pages scored the same 110 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
doc
AE-0011
field
is_serious
narrative severity word
mild
hospitalization
no
event outcome
recovered
causality assessment
possibly-related
is serious model
yes
gold is serious
yes
needs review
yes
Grader
Verdict
Why
Per-field exact match, light normalisation
correct
AE-0011 states "mild abdominal pain" as the event and its narrative reads "Within twenty minutes of the dose the patient deteriorated to the point that the treating clinician recorded the episode as immediately life-threatening and called for resuscitation support. The patient was stabilised in the treatment room over the following hour. No admission was required." Gold: is_serious=yes. Both tiers answered yes, and all ten fields matched exactly.
Review-flag confusion matrix
correct
Gold needs_review=true (is_serious=yes, causality_assessment=possibly-related). Both models' own extracted values fed src/extract.py::compute(), which computed needs_review=true on both tiers — routed for expedited case-processor review, as it should be.
Seriousness accuracy, split by the report's own register
correct
AE-0011 is in the confusable register: the report calls the event "mild", the Hospitalization field says no and the Event Outcome says recovered — every structured signal points at not serious, and only the narrative's life-threatening episode says otherwise. The free severity-register floor calls this one not serious and records it as a false negative; both model tiers got it right.
The formulaWhat it computes
The analysisWhat it actually did
Model
Result
the fast tier
scored 96.4%
the deliberating tier
scored 98.2%
In operationWhat to monitor
Reference standard: gold's is_serious, together with the register tag tools/build_corpus.py recorded when it generated the report — matching, confusable, or no severity word at all.
These rates are UNKNOWN, on purpose
The register split is a property of THIS generator. A real report's wording is not drawn from three buckets, and nothing here measures how a genuinely varied vocabulary would land.
Watch these
accuracy on the confusable slice specifically — an aggregate that mixes it with the matching slice hides the entire effect being measured
the gap between the confusable slice and the free floor's 0 of 25 there, which is the whole demonstration
Alarm on
Confusable-slice accuracy falling towards the matching-slice figure's mirror image — that is what reading the severity word instead of the outcome looks like.
How tight can the band be? No tolerance band; each report is right or wrong, counted within its own register.
Cadence: Re-score on any change to AMBIGUOUS_FRACTION or the criterion narratives in tools/build_corpus.py.
The decisionWhen to reach for it
Use it
The corpus records, per report, whether its severity word agrees with its true classification — knowable here only because the corpus was generated.
Do not use it
A captured corpus carries no register tag, so this split cannot be computed on real reports without labelling them by hand first.
A living map of modern AI — kept current every morning