The business caseThe problem this solves
A council's minutes have to say who moved a motion, who seconded it and how each member voted. The recording holds all of it and a clerk listens back to write it down, scrubbing to the roll call and replaying it to work out which voice said which word. the clerk's replay of the roll call when writing the minutes
Audience
A clerk deciding whether to buy transcription, and the officer who signs off the bill. The answer this kit gives them is a qualified no, and where the money would actually have to go. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual recorded council meetings
The corpus is 40 recorded council meetings, 74.82 MB (json 2 · txt 40 · wav 40). Synthetic for two reasons, and the second is the important one. A real recording would carry licence risk, and it would need its speaker turns HAND-LABELLED — and a hand label is an opinion. Here the boundary between two speakers is the sample offset where one say call ends and the next begins, so the ground truth is exact by construction and the grader is arithmetic. It also makes the result CONSERVATIVE: synthetic voices are spectrally consistent and the cuts are clean, with no crosstalk, no overlap and no room tone — all of which make diarization easier than a real chamber.
The corpus
- The 40 recorded council meetingsunder its source's terms — generated by tools/build_corpus.py from seed 20260906; each turn spoken by macOS
sayin that member's own voice and concatenated with ffmpeg to 16 kHz mono.
Swap this folder for your own material and the kit is pointed at your recorded council meetings. That is the whole change — there is no database to migrate.
The outcomeWhat a good result looks like
a motion record with every vote attributed to the member who cast it
And when it cannot
It leaves the vote out. On the winning arm 68 of 208 votes are attributed and the rest are absent rather than guessed — a gap a clerk can fill in thirty seconds, where a wrong name is a corrected minute.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- the fields you need are speaker identities — a diarizing station, and budget for checking
it is 2.2x the free floor where the alternative is AT the floor - you only need what was decided, not who decided it — the cheapest station with an acceptable word error rate
outcome scored 1.000 on BOTH arms and item within 0.025 — the chair says both out loud, so no speaker information is needed
And where nothing here is good enough:
- you need a filed minute rather than a draft — neither, yet
the best available station attributes 33% of votes; the rest are absent and a clerk still opens the recording
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-06. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop 16 kHz mono WAV files in data/audio/ and a gold.jsonl beside them with one line per meeting: the item, the mover, the seconder, and a votes map. The measured numbers do NOT carry onto a real chamber. Corpus lens → |
| When is this the wrong choice? | Avoid: Choosing on word error rate — it points the wrong way here. That is the case against the best-fitting scenario (“the fields you need are speaker identities”). 3 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | a meeting where two members share a surname — the resolver refuses rather than guessing, and every vote from both becomes unattributable. 4 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | whether any of this holds on a real chamber with crosstalk — see Data.bring_your_own_boundary. 3 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | The shipped adapter is openai-compatible; the Prompt lens states what swapping it costs. The published figures come from 1 model on the fast tier. Prompt lens → |
| And if it fits — what do I stand up? | 4 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-06 — r001-council-motions-aws. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured regenerates the whole corpus and re-scores every committed run offline — the graders are pure code, so re-scoring costs nothing and returns the same answer. Only the two transcription arms and the model stage need credentials.

