The business caseThe problem this solves
A small business misses calls and the voicemail box fills up. Somebody plays each message back, writes down who rang, what about and — the part that has to be right — the number to ring back on. A digit written down wrong is a call to a stranger and a customer who is never called at all. playing each voicemail back to write down the callback details
Audience
Whoever is deciding whether to buy transcription for a voicemail box, and whoever signs off the bill. The answer this kit gives them is a qualified yes for the phone number and a clear no for the name. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual voicemail recordings
The corpus is 60 voicemail recordings, 13.49 MB (jsonl 1 · md 1 · txt 60 · wav 60). A real voicemail carries a real person's name, a number they answer and their business with somebody — publishing one is a privacy decision no licence fixes. It would also need its transcript hand-corrected before it could be graded, and a hand correction is an opinion that every published number would then inherit. Generating the message means the words are known because we wrote them, so the grader is string comparison rather than judgement.
The corpus
- The 60 voicemail recordingsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your voicemail recordings. That is the whole change — there is no database to migrate.
Hello, this is Priya Quennell, calling about Ridgeway Autos. Something's gone wrong with the job you did. I need this sorted today if possible. You can get me on 5 5 5, 0 1 3, 5 6 6 5. Thanks very much.
The outcomeWhat a good result looks like
a callback record: who rang, the ten digits to ring back on, what it is about, how urgent, and which business they meant
And when it cannot
It records the wrong name against the right number. On the best path 43 of 60 records carry a caller name that does not match what was said, while 59 of 60 carry the correct ten digits — so the person gets called back and greeted by the wrong name.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- you need the callback number and nothing else — either engine — and on this evidence you cannot justify the dearer one on this field alone
59 of 60 against 53 of 60 looks like a clear win, and the paired test says p=0.0703: the two engines disagree on only 8 voicemails and 60 cases is not enough to call it. Worse for the dearer engine: a FOUR-LINE REGEX that takes the last ten digits out of its transcript gets 59 of 60 — the same as the model — so on this field the model is not earning its bill either. - the record must carry the caller's NAME — AWS Transcribe batch
17 of 60 names against 8 of 60, and this one IS separable (p=0.0225). It is twice the price per minute and the whole corpus cost $0.09 — but 28% is still a name you have to check, so the honest advice is that the dearer engine buys you less checking, not no checking. - you only need routing — which of six buckets, and how urgent — either engine, and take the cheaper one
reason and urgency are 1.000 on both paths and on the perfect transcript. Intent survives transcription noise that destroys a surname.
And where nothing here is good enough:
- the record must be filed without a human reading it — neither, on this evidence
the best path gets all five fields right on 75.0% of records and the floor this kit declares is 0.97. Nothing clears it, and the Tracks panel marks no winner.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-13. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Point tools/run_*.py at your own WAVs and write a gold.jsonl beside them with the true record for each. ⚠︎ THIS CORPUS IS SYNTHETIC AND ITS NUMBERS ARE A CEILING, NOT A PREDICTION. Corpus lens → |
| When is this the wrong choice? | Avoid: Reading the 59-against-53 gap as settled. It is the kind of gap that turns out to be real on two hundred voicemails and vanishes on the next sixty, and this kit ran sixty. That is the case against the best-fitting scenario (“you need the callback number and nothing else”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | the 'doubles' reading — "double five, five oh one nine, double one, double one". It is the worst style on every arm including the perfect transcript (16/17, 12/17 and 14/17), and it is how people actually say a number aloud. 4 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Anything about real voicemail audio. This corpus has no background noise, no crosstalk, no clipping and no accent beyond what macOS ships, so every number here is a ceiling — see Data.bring_your_own_boundary. 5 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 5 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-13 — r001-voicemail-intake-aws. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Measured on this machine, not written as an instruction: from a clean checkout python3 -m evals.score --self-test and python3 -m evals.significance --self-test both run green with NOTHING configured, and python3 -m src.app serves the whole board on port 9454 with no API key and no cloud credential — every transcript and every record is already on disk in results/. The only pip install in the kit is google-auth, and it is needed only to RE-RUN the Google station. ffmpeg and macOS say are needed only to regenerate the audio, never to score it.





