Home › Use Cases › Reportability assessment support
Use caseUC0499
🧪 Use-case kit · runnable

Reportability assessment support

A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.

The business caseThe problem this solves

A product complaint arrives and somebody has to fill in a reportability worksheet from it. The record is a header, a product line and a dated intake log written in prose by whoever took the call - 6 to 11 entries, some declined, some hedged, some corrected later, some superseded, some recorded after the worksheet was even asked for. Four facts have to come out of it, each with the entry it came from, and the honest answer is often that the log does not state one. The first pass over a complaint record: reading the intake log, deciding which entry states each element, and writing the four cells with their source entry ids - or marking a cell unavailable when nothing admitted states it.

Audience

Anyone putting an automated reader over a regulated worksheet, where the dangerous answer is not a wrong value but a value where the log has none. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual complaint records

The corpus is 64 complaint records, 0.13 MB (md 64). The smallest corpus that makes the interesting mistake unavoidable. Two earlier builds were measured and thrown away: the first let every decoy be dropped whole on a keyword found nowhere else, and the second left the attributed arm at 51 of 64 with E2 at 62 - two worksheets of room - because every value carried a unique keyword. The shipped build adds a paraphrase library: 117 of 509 log entries write their value two ways SHARING NO KEYWORD, which is what pushed the bar down to 28 of 64 and left every cell with room.

The corpus

  • The 64 complaint recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your complaint records. That is the whole change — there is no database to migrate.

One complaint record, as the model receives itrecords/CX-2611-0001.md · 1 of 64
# Complaint record -- CX-2611-0001

Synthetic record. Every reference, name, date and note below was generated by tools/build_corpus.py from a fixed seed. Nothing here is a real complaint, a real person or a real product.

Complaint reference: CX-2611-0001
Product family: nasal spray
Lot reference: LOT-1037
Record opened: 2026-02-15T17:00Z
Worksheet requested: 2026-02-28T16:08Z
Specialist assigned: 2026-03-07T02:20Z

## Decision-tree elements on the worksheet

- E1 -- outcome recorded for the person concerned
- E2 -- whether the product was in use at the time of the event
- E3 -- whether the product is described as not performing as intended
- E4 -- date the organisation first became aware of the event

The element list above is this organisation's own current practice. It is not a regulator's list and nothing on this record states a reporting clock.

## Intake and correspondence log

### C01
recorded: 2026-02-15T18:43Z
by: S. Obuya, complaint desk
note: Reporter asked which address to return the unit to.

### C02
recorded: 2026-02-17T06:37Z
by: N. Delacroix, complaint desk
note: The patient on this record had lot LOT-1037 of the nasal spray in use at the time.

### C03
recorded: 2026-02-20T23:58Z
by: N. Delacroix, complaint desk
note: The reporter describes the event as having happened on 2026-02-02.

### C04
recorded: 2026-02-22T04:32Z
by: N. Delacroix, complaint desk
note: The person named on this complaint first said they said the nasal spray worked exactly as it should, then said on the call back that they said the nasal spray let them down.

### C05
recorded: 2026-02-24T15:31Z
by: N. Delacroix, complaint desk
note: Query about remaining stock from the same lot.

### C06
recorded: 2026-02-27T02:46Z
by: K. Ferraro, complaint desk

Abridged — the file continues.

The outcomeWhat a good result looks like

Four cells, each with the log entry it came from, and an unavailable wherever nothing admitted states the fact.

And when it cannot

It takes the awareness date off the wrong entry - 24 of 64 records - and it fills in 10 of the 86 cells nobody recorded. The second is the failure this pack exists to prevent.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You need the four cells and the entry each came from, and a wrong cell is caught by a person downstream — the free attributed arm (evals/floors.py::arm_attributed)
    It is LEVEL with the paid call on the row of four - 28 against 29, p = 1 - and it costs nothing.
  • The facts in your logs are paraphrased rather than stated in the words your rules expect — the paid call
    This is the one thing it decisively buys: family_stated 66 of 66 against 43, p = 2.384e-07, and family_correction 22 of 22 against 15, p = 0.01562 - both LEVEL WITH THE CEILING ARM, and the only slice claim in this batch to survive being scored against every free arm.
  • The cost of an invented cell is higher than the cost of an empty one — free code - the attributed arm, or in the limit the all-unavailable constant
    The paid arm invents a value on 10 of the 86 cells nobody recorded; the bar invents 1 (p = 0.01172) and the constant invents 0 (p = 0.001953). Free code is safer here and it is not close.

And where nothing here is good enough:

  • You want a number to put in front of somebody — nothing on this page, on its own
    The headline is level and the safety number is a loss. The publishable finding is the SHAPE: the money buys the paraphrase reading and gives back the date arithmetic and the guardrail.

At a glanceHow the whole thing runs

45%worksheet all correct pct
1,000 msp50, end to end
$19.38per 1,000 complaint records · Claude Fable 5

Run once, for real, on 2026-09-17. Every figure on these pages was captured from that run — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Point data/records at your own complaint records and rewrite src/record.py::parse to return the same three things: a header with the two timestamps, a product line, and a list of dated log entries each with an id. Nothing measured here transfers to your corpus. Corpus lens →
When is this the wrong choice?Avoid: Do not expect it to read a paraphrased fact: family_stated is 43 of 66 for it against the paid arm's 66. That is the case against the best-fitting scenario (“You need the four cells and the entry each came from, and a wrong cell is caught by a person downstream”). 4 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A log entry with no date. Every reading rule in the prompt is ordered on the entry timestamps, and an undated entry has no place in that order. 5 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Two of this kit's numbers are won outright by free code and are never a result on their own. The queue wait is the subtraction of two header timestamps, which an all-unavailable constant computes exactly, and both guardrail counts go to that same constant at 86 of 86 called right and 0 inferred. 7 items this kit says it could not check. Eval lens →
Can I run this on a model I control?The shipped adapter is OpenAI-compatible endpoint; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens →
And if it fits — what do I stand up?6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-17 — r001-reportability-screen - 64 complaint records, 64 billed model calls, the fast tier, reasoning off. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — A clean checkout with no key configured renders the whole board, rebuilds all 66 corpus files byte-identically under two PYTHONHASHSEEDs, re-derives all 256 labels with 0 disagreements and scores all five free arms plus the re-score of the paid run offline, at $0.00. The one live control on the board is disabled without a key, and the committed reply cache is what makes the paid arm re-scorable.

A living map of modern AI — kept current every morning