The business caseThe problem this solves
A cue sheet is the document a performing-right society registers a program from, and it is cut by hand from a music editor's log after picture lock. Every cue in the locked picture has to be on it once, with the timings the picture actually carries, the usage code the scene actually supports, the writers and publishers who actually wrote and published it, and shares that total exactly 100.00. A cue that is missing is never registered and nobody is paid for it; a cue that is on the sheet and not in the picture takes an allocation from the ones that are. Checking it is reading two documents side by side, one cue at a time, and the delivery vendor's own intake check compares titles and timecodes and stops there. Opening one cue slot's two documents side by side, matching the record by eye, subtracting two timecodes at the production's frame rate, adding two columns of shares, comparing two lists of names and walking a seven-rule ladder — and doing it again for the next slot.
Audience
A music rights administrator or a music supervisor's assistant working a delivery before it is filed, and the operator deciding whether a call is worth buying against a rule they could write themselves. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual cue slot
The corpus is 62 cue slot, 0.11 MB (txt 62). It is generated because it has to be. A real cue sheet is a production's own delivery document naming real writers, real publishers and real shares, and a real music editor's log names the scenes of an unreleased picture. Neither is ours to publish. What matters for the measurement is not that the titles are real but that the READINGS are hard in the ways a real delivery is hard, and data/SOURCES.md is explicit about the three ways this corpus is easier than one.
The corpus
- The 62 cue slotgenerated from a fixed seed, so no real record, person or institution appears in it.
- Where each came fromdata/SOURCES.md states where every byte came from AND what the generator costs the measurement, including the two cases it makes decidable without reading anything.
Swap this folder for your own material and the kit is pointed at your cue slot. That is the whole change — there is no database to migrate.
==============================================================================
CUE SHEET QC FILE CSQ-0001
Production: PRD-2201 - Harbor Lights (invented) Episode: 1x04
Slot: SLOT-0001 log-led subject LOG-0011 frame rate 24 Standard: CSQ-2026
==============================================================================
PRODUCTION AND SLOT AS THE MASTER HOLDS IT
frame rate 24
timing tolerance 6 frames
reels in program 6
delivery DLV-2201-104
FINAL-CUT MUSIC LOG ENTRIES FOR THIS SLOT
LOG-0011 MARRAM GRASS 01:38:01:03 01:38:41:05 reel 4
WRITERS AS LOGGED: S. Calloway; C. Mbeki
SCENE: Score only, running beneath the whole exchange; nobody in the scene is playing it and there are no words.
DELIVERED CUE SHEET ROWS FOR THIS SLOT
ROW TITLE IN OUT DUR USE REEL
CSR-0012 MARRAM GRASS 01:38:01:09 01:38:41:11 00:00:40:02 BI 4
WRITERS: S. Calloway 32.00; C. Mbeki 68.00
PUBLISHERS: Quiet Ward Music 100.00
CSR-0013 THE LONG FIELD 01:00:41:00 01:01:36:00 00:00:55:00 BI 1
WRITERS: D. Sorrell 100.00
PUBLISHERS: Quiet Ward Music 100.00
QC AS THE DELIVERY REPORTS IT
matched record CSR-0012
usage code BI
status TIMING-MISMATCH
MUSIC SUPERVISION NOTES
Reel breaks were re-numbered after the online; the reel column on this sheet is the new numbering.
==============================================================================
The outcomeWhat a good result looks like
One cue slot in, one row out: which record on the other side is the same cue, the usage code the scene supports, whether the notes declare a conform offset, the records this QC treats differently from the delivery's own report, and one verdict from a closed ladder of seven — with the arm's own answer and the same answer through the pure-code station published side by side.
And when it cannot
And what it does when it cannot. On the scored run 62 of 62 replies parsed, 0 stopped at the ceiling and the largest was 234 output tokens of a 4000 ceiling. A reply that does not parse is counted WRONG on every field and stays in the denominator; it is never re-fired. The first canary of 8 returned one such reply — 751 output tokens of working the comparison out inside a free-text field — and the fix was the prompt, not the scorer.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
- Your delivered sheets are cut from the same log by the same tool, so a title and a timecode always line up — the free
matchfloor, and do not buy a call at all
On the 48 slots that have a counterpart the free matcher is right 33 times, and on the whole corpus it beats the paid arm on all six fields at once, 34 to 25. - Your music supervision notes carry conform offsets, and some of them describe offsets that were proposed and never carried out — the paid arm
The paid arm read all 17 offset slots right, 62 of 62, including every one of the 8 wordings that refuse an offset while stating its reel, its direction and its figure. The free regex applies 6 of those. - Your usage codes are assigned from the scene rather than copied from a previous delivery — the paid arm
It read the scene into the right code 56 of 62 against the free floor's 44, and it copied the delivered row's own code 0 times. A free rule can only take the code off the row, so it can never find a miscoding at all. - You want the shares checked and the writers compared — pure code — it is already in the kit and it is free
Q-6 and Q-7 are integer arithmetic on hundredths and a set comparison over folded names. src/policy.py does both, every arm gets them, and the paid arm scores SPLIT-INVALID 9/9 and PARTY-MISMATCH 10/10 because the station derives them.
And where nothing here is good enough:
- The findings you are paying for are missing cues and extra cues — NEITHER ARM, and the honest answer is that this kit does not do it yet
On the 14 slots where the right answer is 'no record on this file is this cue' the paid arm is right 2 times and the free matcher 0. Both will always find a best candidate. This is the finding that costs money and it is the one neither arm reaches.
At a glanceHow the whole thing runs
Run once, for real, on 2026-09-09. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
| What do I have to bring? | Replace data/corpus/*.txt with your own slots in the same shape and data/productions.json with your own frame rates and tolerances. ⚠︎ WHAT STOPS BEING TRUE THE MOMENT YOU DO. Corpus lens → |
| When is this the wrong choice? | Avoid: Paying per slot for a string comparison you could write in an afternoon. That is the case against the best-fitting scenario (“Your delivered sheets are cut from the same log by the same tool, so a title and a timecode always line up”). 5 scenarios scored in all, each with its own. Eval lens → |
| Where does it stop working? | A delivered sheet that is not fixed-width columns. src/cuesheet.py's row regex is the shape these files print; a PDF, a spreadsheet export or a society's own XML needs a different parser and nothing above it changes. 6 recorded failure modes, each from a run rather than a guess. Corpus lens → |
| What was never verified? | NO SECOND SCORED RUN. One was fired, so the run-to-run spread on this corpus is unknown and unclaimed. 8 items this kit says it could not check. Eval lens → |
| Can I run this on a model I control? | Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens → |
| And if it fits — what do I stand up? | 6 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one. step 14 — Run it in your environment → |
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-09-09 — r001-cue-sheet-qc. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
- The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
- The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A clean checkout with no key configured renders the whole board, all three free floors and every committed run, and python3 -m evals.run --run-id t000-cue-sheet-qc-stub --stub exercises the whole pipeline end to end for nothing.







