Home › Use Cases › Match each promotional claim to the reference that supports it, and quote the line
Use caseUC0304
🧪 Use-case kit · runnable

Match each promotional claim to the reference that supports it, and quote the line

A small, forkable project that does one job end to end. Run twice for real over the same set, and every figure on these pages captured from those runs.

The business caseThe problem this solves

A promotional piece makes numbered claims and a reference pack is assembled to substantiate them. Somebody then has to read every claim against every reference and answer one question per claim: which line states this, and does it state it as the piece states it? The pack is not indexed to the claims — nothing in it says which study supports which sentence — so the work is a read-through of the whole pack per claim, and it is repeated on every version of every piece. The four things that get missed are always the same: a study that reports the result in a SUBGROUP, a figure that belongs to the COMPARATOR, a reference the register marks RETRACTED, and a claim written entirely in the label's own words that quietly drops one of its qualifiers. The read-through: one person opening a reference pack, finding the line that speaks to each claim somebody else wrote, deciding whether it states the claim as written, and — the step that actually gets skipped — deciding whether the difference is a narrower population, a different endpoint, a different figure, or the result reported the other way round. It does not replace the review. Nothing here clears a piece, rejects one, proposes replacement wording or names who may.

Audience

The person who works a promotional review queue and the medical reviewer who has to see the evidence behind every flag. The decision this report is for is narrower than it looks: not 'is our copy compliant' — nothing here answers that — but 'which of these claims can I evidence from this pack, and exactly which line do I look at'. More than half of that is a token matcher, and this report says which half is not. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual promotional claim review packs

The corpus is 62 promotional claim review packs, 0.21 MB (txt 62). Because claim-to-reference matching is a RETRIEVAL problem wearing a compliance hat, and a corpus has to make both halves separable. The conditions are invented eponyms — 'moderate to severe Kelbrand disease' — deliberately: a generated corpus about a real indication would read as being about real medicines whatever the disclaimer said, and a reader could not tell an invented claim from a remembered one. evals/check_labels.py ENFORCES that rather than asserting it: it retypes the ten invented brands, the ten comparators, the ten conditions and the fifteen trial acronyms and fails if a medicine-shaped name appears that is not on the list.

The corpus

  • The 62 promotional claim review packsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromNowhere — every one of the 62 packs, the reference register and the whole answer key are generated in-process from seed 20260903, and tools/build_corpus.py --check fails if a single byte moved. Verified byte-identical under two PYTHONHASHSEEDs.

Swap this folder for your own material and the kit is pointed at your promotional claim review packs. That is the whole change — there is no database to migrate.

One promotional claim review pack, as the model receives itMLR-0001.txt · 1 of 62
PROMOTIONAL CLAIM REVIEW PACK                                        MLR-0001
Prepared 2026-09-03 under PRS-2026 | Northgate Brand Team

PACKET FACTS
  Product                      Velmoretin 40 mg film-coated tablets
  Piece                        Sales aid VEL-SA-100, version 1
  Audience                     Prescribers
  Reference pack               RPK-1000
  Pack reference               MLR-0001

REFERENCE REGISTER
  ref    source                                          status        year
  LBL    Prescribing information extract                 ON-FILE       2024
  R1     HELIOS-2 primary publication                    ON-FILE       2021
  R2     Pooled safety analysis                          ON-FILE       2022
  R3     Prespecified subgroup analysis                  ON-FILE       2023
  R4     Pharmacokinetic substudy report                 RETRACTED     2024

THE CLAIMS IN THE PIECE
  claim  claim as written
  C1     In HELIOS-2, Velmoretin reduced major events of Kelbrand disease by 32.7% compared with Dartizole.
  C2     Velmoretin lowered the risk of hospitalisation for Kelbrand disease in HELIOS-2.
  C3     Response to Velmoretin began within the first week of treatment.
  C4     Velmoretin does not require routine laboratory monitoring.
  C5     Velmoretin improved the Merrow symptom score by 3.2 points in adults with moderate to severe Kelbrand disease.
  C6     Velmoretin reduces the need for rescue medication.
  C7     Discontinuation for adverse events occurred in 5.3% of patients treated with Velmoretin.

THE LABEL EXCERPT
  LBL-1    INDICATIONS AND USAGE. Velmoretin is indicated in adults for the treatment of moderate to severe Kelbrand disease.
  LBL-2    DOSAGE AND ADMINISTRATION. The recommended dose of Velmoretin is one 40 mg tablet once daily.

Abridged — the file continues.

The outcomeWhat a good result looks like

One row per printed claim: a verdict from five (SUPPORTED, PARTIALLY-SUPPORTED, CONTRADICTED, UNSUPPORTED, OFF-LABEL), the line id it was read from, whether that line states the claim as written, on which single axis it differs, whether the claim is inside the label excerpt, the line quoted verbatim, a confidence and one sentence of basis. 355 of 372 claims correct on the verdict (95.4 pct) with the line quoted right on 363 of them, and 50 of 62 pieces entirely correct on both graded fields.

And when it cannot

And what it does when it cannot. 17 rows of 372 were wrong and they are two things, not one. 10 are partial_read_as_contradicted: every one is the claim 'fewer than N pct of patients discontinued' against a pack line reporting a HIGHER rate, which the arm read as a direction reversal and the key calls a figure difference. ⚠︎ BOTH READINGS ARE DEFENSIBLE UNDER PRS-2026 AND THE KEY WAS NOT CHANGED TO AGREE WITH THE ARM. The other 7 are one piece, MLR-0013, whose reply was CUT OFF AT THE 32,000-TOKEN CEILING and recorded as a failure — not a partial score, and not re-fired. ⚑ THE DIRECTION OF EVERY ERROR IS THE SAFE ONE: 0 of 216 flagged claims were signed off, in either column, and 0 supported claims were flagged.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • Your reference lines repeat the claim's own words and its own figures, and your register is trustworthy — the free rules floor alone — evals/baseline.py
    It scores 55.4 pct for $0.00 and gets 100 pct of five of the fifteen construction cases, including every retracted and draft trap, because the pack prints its own register and the floor is given it. More than half of this job is a token matcher with a document-frequency filter.
  • Your pack states results in a subgroup, on a composite endpoint, as two absolute rates, or for the comparator — and your copy paraphrases — the paid call
    That is the whole measured margin: 150 rows in 6 families where the floor scores 0 and the call scores 147. Paired over all 372 claims: 197 both right, 158 the call only, 9 THE FLOOR ONLY, 8 neither — exact two-sided McNemar p = 2.5e-36.

And where nothing here is good enough:

  • You need the AXIS of the difference to be right, not just the flag — neither, yet
    differs_on is the arm's weakest reading at 77.8 pct, and it is what tells a copywriter whether to narrow the population, change the endpoint or restate the figure. The verdict is 95.4 pct and the axis is not.
  • Your pieces carry twenty claims each — neither, unmeasured
    One reply on this run already drew the whole 32000-token ceiling on a seven-claim piece and was recorded as a failure. Nothing here batches claims across calls.

At a glanceHow the whole thing runs

95–96%verdict pct · 2 runs, no ordering
92,870 msp50, end to end
$42.80per 1,000 promotional claim review packs · Gemini 3 Flash

Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/corpus/*.txt with your own packs in the same six-panel shape, replace data/packs.json with your own reference register, and write data/gold.jsonl with one row per claim carrying the four readings and the labelled line. Every number on this page stops being true the moment you do. Corpus lens →
When is this the wrong choice?Avoid: Paying for the half that is free. It is most of the volume and none of the value. That is the case against the best-fitting scenario (“Your reference lines repeat the claim's own words and its own figures, and your register is trustworthy”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A scanned or PDF reference pack. Every arm here reads printed, panelled text with a line id at the head of every line; there is no OCR, no layout model and no page geometry. 6 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the labelled line is the line a reviewer would quote. The key names one governing line per claim under CR-10; a second result sentence in the same reference, or the label's indication line beside a study line, scores zero here and a reviewer might accept either. 6 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is one OpenAI-compatible endpoint, reached only through src/adapters/; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?7 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-claim-reference. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured rebuilds all 62 packs from seed 20260903 in 0.1 s, passes the independent label gate in 0.08 s, scores both free floors in under a second, and renders the whole board — the packs, the register, PRS-2026, the citation scoring and every committed run replayed — with the model control disabled and the reason printed beside it. Nothing on that path touches the network or spends anything.

A living map of modern AI — kept current every morning