The business caseThe problem this solves
A health plan's marketing-compliance team has to know that a broker said the things a recorded enrollment call requires — that the call is recorded, that the broker does not work for the government, how many carriers they represent, that enrolling will end the current plan — and did not say the things a broker may not say. The recordings exist. Reading them does not scale: a reviewer plays the call, keeps eleven items in their head, and decides for each one whether it was said, by WHOM, and whether it came before the enrollment started. a reviewer playing a recorded enrollment call end to end and ticking eleven disclosure items by hand
Audience
Whoever is deciding whether a transcription engine is worth buying in front of a model, and whoever signs the bill for it. The answer here is specific and it is not the one the price list suggests: the engine that hears WORSE is the one this kit recommends, and the derived winner on the panel is still the cheap one. Both of those are true and the page says why. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual recorded enrollment calls
The corpus is 64 recorded enrollment calls, 176.31 MB (txt 64 · wav 64). Compliance on a recorded call is not a transcription problem and this corpus is built to prove it. 11 of the 64 calls plant a disclosure in the BENEFICIARY's mouth, and 30 more have the caller repeat a sentence the agent also said — two families a transcript can only tell apart by who spoke. 33 calls meet an obligation in words the checklist does not use, and 19 name a prohibited claim in order to REFUSE it, which is the trap for any keyword check: every content word of the breach is present and nothing was breached.
The corpus
- The 64 recorded enrollment callsgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your recorded enrollment calls. That is the whole change — there is no database to migrate.
AGENT: Good morning, this is Corin Ashvale calling from Cedarline Benefit Advisors about your Medicare options.
CALLER: Yes, hello. This is Desmond Achebe.
CALLER: I should say up front, I already have a plan and I have carried it for 4 years now.
AGENT: This call is being recorded for quality and compliance purposes.
AGENT: I do not work for Medicare or the federal government.
AGENT: There are eight different companies whose products I am able to offer where you live.
AGENT: There is no obligation to enroll in anything today.
AGENT: I am licensed in Arizona and my producer number is 1 8 9 7 9 7 7.
AGENT: Enrolling in this plan will end your current coverage.
AGENT: We are here to look at a Medicare Advantage plan, and the one I have in front of me is Bluewater Complete.
CALLER: That sounds reasonable so far.
AGENT: Your monthly premium on this one is 24 dollars.
CALLER: I appreciate you explaining it slowly.
AGENT: The out of pocket maximum for the year is 4900 dollars.
AGENT: I am not allowed to ask what conditions you are being treated for.
CALLER: Go on, I am listening.
AGENT: There is a drug list, and it is reissued at the start of each year.
AGENT: All right, let me go ahead and start your enrollment application now.
AGENT: That is everything on my side. A confirmation letter goes out this week.
CALLER: Thank you for your time.
The outcomeWhat a good result looks like
eleven item verdicts per call — met, breached or not_required — the quoted turn each one rests on, and one call outcome of clear or refer. The outcome is never asked of the model: it is derived in code from the readings by the six rulebook rules, so two people reading the same record get the same answer.
And when it cannot
It refers a clean call, or clears a call it should refer. On the diarized track 6 of 64 calls come out with the wrong outcome; the direction that costs is the second one, and on the undiarized track it is the common direction — a breach spoken by the BENEFICIARY reads as a disclosure met, so the call clears. This kit flags a segment for a human reviewer and never determines that an agent is in violation.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- your recordings are two-party and your checklist has a who-said-it rule in it — the diarized track
it is the only priced path on this stage's frontier that returns speaker labels, and the rule that a disclosure in the caller's mouth does not count fired 17 times on it against 6 on the undiarized text. It buys 9.4 points of call-outcome accuracy for roughly twice the transcription price. - your checklist has no rule that turns on who spoke — the undiarized track, or the free regexes
with attribution out of the question the two paths converge, and the cheap one is half the price. Note what the free floor already gives you on this corpus: 0.7812 of call outcomes against the cheap track's 0.8125, and 0.7080 F1 against 0.7027 — a tie on the finding. - you want to know whether to spend on the transcriber or on the model — run the perfect-reading arm first — it is free once the reference text exists
it separates the two questions completely. Here it scores 0.9688, so 3 wrong item verdicts are the model's own ceiling before any engine is bought, and everything between that and a purchased path is what the purchase cost. - you are pricing this against a human reviewer — count the queue, not the accuracy
the free floor has recall 1.0000 and precision 0.5479, so it refers almost everything; the diarized track has precision 0.6964. The difference is how many clean calls a reviewer opens, which is the cost that scales.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-16. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Drop your own 16 kHz mono WAVs in data/corpus/ named <case_id>.wav, and one line per call into data/gold.jsonl carrying items (the eleven verdicts), breached, outcome and spans (the turn text and speaker each item rests on). reference_text is the field a real corpus cannot honestly supply, and every stage-1 error rate on this page depends on it. Corpus lens → |
| When is this the wrong choice? | Avoid: Buying it on its transcription quality. It hears WORSE than the cheap path on every error measure this kit takes, and its own headline failure is a welded turn — a diarization error, not a hearing one. That is the case against the best-fitting scenario (“your recordings are two-party and your checklist has a who-said-it rule in it”). 4 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A call long enough that the transcript does not fit one prompt. Rule R3 is an ORDERING question — was this said before the enrollment started — and it is not answerable inside a window that does not hold both turns, so chunking is not a fix, it is a different kit. 5 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | Any of this on a REAL recording. Every call is synthesised speech reading a generated script, so every transcription error rate here is a floor and the diarization result is measured on audio with no crosstalk, no hold music and no background at all. 7 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-16 — r001-call-disclosure-check-aws. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Observed on this machine: a checkout with no .env and no key runs python3 -m evals.baseline and prints the whole floor table in about two seconds, because the floors call nothing. python3 -m evals.check_labels re-derives all 704 item verdicts from the committed turn structure and exits clean in under a second. python3 -m src.app serves the committed runs on port 9487 with no credential at all.





