Home › Use Cases › Classify an activity record for lobbying under the organisation's own card
Use caseUC0303
🧪 Use-case kit · runnable

Classify an activity record for lobbying under the organisation's own card

A small, forkable project that does one job end to end. Run twice for real over the same set, and every figure on these pages captured from those runs.

The business caseThe problem this solves

Every quarter somebody at a nonprofit reads the activity log — staff time entries, event descriptions, excerpts of things that went out — and codes each row for lobbying tracking. Today that is a person with the organisation's definitions card open beside them, reading five panels a row and deciding which of eight classes fits. Two rows can look identical in the body text and belong in different columns, because the difference is one line about who received the piece and one line about whether it asked anyone to do anything. The first pass over a quarter's activity log: which of eight classes each row is, and the line that establishes it. It replaces nothing downstream — no total, no limit, no determination and no charge is moved.

Audience

The person who codes the log, and the compliance coordinator who has to defend a coding six months later. The decision they are making is not 'should we use AI' but 'is a free keyword table enough', and on this corpus the honest answer for three quarters of the log is yes. Every number on these pages came from one real run of this code, not from a vendor page.

The inputThe actual activity records

The corpus is 64 activity records, 0.07 MB (txt 64). A lobbying activity log is the right shape for this question and the wrong shape to obtain: a real one names staff, names the people they met, and belongs to an organisation with a lobbying position it did not publish. So it is generated — and generated so that the answer never sits in a column. The audience and the measure ARE columns, deliberately, because a real log has both and a free floor denied them would be a strawman; what no column carries is whether a view was expressed, whose view it was, and whether anything was asked of the reader. Three of the ten measures are about the organisation itself and appear on records of five different classes, so a bill id never decides an answer. Twenty-one case families run from a clean control group to two forms of the member-exception trap.

The corpus

  • The 64 activity recordsgenerated from a fixed seed, so no real record, person or institution appears in it.
  • Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.

Swap this folder for your own material and the kit is pointed at your activity records. That is the whole change — there is no database to migrate.

One activity record, as the model receives itACT-0001.txt · 1 of 64
LOBBYING ACTIVITY RECORD                                  ACT-0001
Prepared 2026-09-03 under LDC-2026 | reporting period 2026-04-01 to 2026-06-30

RECORD FACTS
  Record type                  staff time entry
  Logged by                    Executive Director
  Activity date                2026-06-22
  Funding source               GRT-3188

AUDIENCE AND DISTRIBUTION
  Audience                     the Senate Housing Committee chair and two committee staff
  Channel                      in-person meeting at the Assembly building
  Reach                        2 recipients

MEASURE REFERENCED
  Bill                         V.A. 1206 - Childcare Subsidy Continuation Act

ACTIVITY DESCRIPTION
  Executive Director provided technical assistance on V.A. 1206 - Childcare Subsidy Continuation Act - cost tables and a coverage model.
  The request came verbally from the chair during the hearing on 2026-06-14 and nothing was received in writing.
  The coalition's objection to V.A. 1206 was stated in the meeting and repeated in writing.
  The measure was one of four the programme is following this session.

ACTIVITY NOTES
  No materials beyond those described were produced for this activity.
  Logged on the standard quarterly cadence for this programme.

The outcomeWhat a good result looks like

Each record carries one class, one line of the record quoted verbatim as the evidence for it, the tracking column it belongs in, and a flag where a lobbying activity was charged to a grant that forbids it. A coordinator confirms a row instead of re-reading it.

And when it cannot

When it cannot, it says UNCLEAR-NEEDS-REVIEW and the record goes back to whoever wrote the entry — which is a real answer and the only honest one on a record that does not state its audience. Three records in the corpus are exactly that, and every arm answered all three correctly. What it must never do is round that to NOT-LOBBYING, which is how a lobbying cost leaves a tally with nobody deciding that it should.

Where it fitsWhat did work

Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.

  • You want the restricted-funds flag and nothing else — the free modal floor plus the funding register — evals/baseline.py mode modal
    92.2 pct on the flag for $0.00, and 9 of 9 on the records that need it, because the flag is a property of the register and the rule table rather than of any reading. The paid call scores 85.9 raw on the same field.
  • You want a first pass over the whole log and your card is short — the free keyword rule table — evals/baseline.py mode rules
    48 of 64 classes and 37 of 64 records fully right for $0.00 with no network. On a log this regular it answers three quarters of the job.
  • Your log contains implicit calls to action, dual distribution, or somebody else's view quoted — the paid call, and read the classification field on its own
    63 of 64 against the floor's 48, and every one of the 16 it wins is one of three mechanisms a phrase list structurally cannot handle. On the named failure mode it is 3 of 3 where the floor is 0 of 3.
  • You want the tracking column and the flag to be right — the paid call WITH src/recheck.py bolted to it, never without
    64 of 64 on both, rechecked, against 55 and 55 raw. All nine restricted charges are missed by the raw arm and returned by the station, on both scored runs.

And where nothing here is good enough:

  • You need to know a lobbying cost will never leave the tally — neither arm, and this is a limit rather than a recommendation
    The paid call is 0 of 26 on both scored runs and 0 of 12 under an adversarial sentence, which is as clean as this corpus can measure. But the funding register knows nothing about any activity, so nothing in this kit can catch a reading that went wrong — register_cannot_defend is 12 of 12 on the adversarial arm, published as a measured zero.

At a glanceHow the whole thing runs

84%record all correct pct · 2 runs, no ordering
12,880 msp50, end to end
$5.01per 1,000 activity records · openai/gpt-5-6-luna

Run twice over the same set, for real, the last on 2026-09-03. Every figure on these pages was captured from those runs — nothing is written from intent.

14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.

Should you use this?What you bring, where it stops, and when not to use it

Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.

What do I have to bring?Replace data/policy.md and data/policy.json with your own definitions card — the same rules in both, stated in the same order, which evals/check_labels.py enforces. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens →
When is this the wrong choice?Avoid: Paying for it. And avoid reading the modal floor's 92.2 as competence — it flags 5 records that need no flag, and it gets every class wrong. That is the case against the best-fitting scenario (“You want the restricted-funds flag and nothing else”). 5 scenarios scored in all, each with its own. Eval lens →
Where does it stop working?A log that is not five fixed panels. The free floor's audience classifier and the label gate's structural checks are both bound to literal headings; a free-text note field returns UNCLEAR-NEEDS-REVIEW on everything, which fails safe and fails completely. 8 recorded failure modes, each from a run rather than a guess. Corpus lens →
What was never verified?Whether the key's class is the only defensible one on the two named_no_view public explainers. It is not: the card's test 4 comes before test 8 and both records read as full and fair public expositions with no call to action. 9 items this kit says it could not check. Eval lens →
Can I run this on a model I control?Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?8 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment →

Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).

Last verified 2026-09-03 — r001-lobbying-classify. Every figure on these pages was captured from that run.

Run itHow this reaches your data

Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.

Run this on your own data

  • The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
  • The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.

Talk to us →

Checked before this shipped — Measured on 2026-09-03 from a clean checkout with no key configured: tools/build_corpus.py rebuilt all 64 records, the funding register and the answer key in under two seconds and --check confirmed them byte-identical under two PYTHONHASHSEEDs; evals/check_labels.py re-derived the whole key and exited 0 on eleven checks; both free floors scored all 256 graded cells with no network; the --stub pass exercised the prompt assembly, the JSON parse, the normaliser, the recheck, the citation locator and the scorer end to end; and the board rendered every panel on 127.0.0.1:9303 with the model control disabled and the reason printed beside it. Nothing needed installing. What a clean checkout CANNOT do is call a provider, and the two scored runs and the adversarial arm replay from their committed result files instead.

A living map of modern AI — kept current every morning