Who said that? Speaker-attributed Survivor transcripts
Survivor’s captions often do not identify the speaker: some seasons prefix names, most mark only speaker changes, and none mark overlap. This project adds speaker names to the transcripts from audio and caption-derived labels, with uncertain cases sent for review by ear, so I can examine what individual players say. Speaker identification is 86% accurate on single-speaker segments of at least ten seconds. It runs on my own library of episodes, so there is no demo; this page is the method, the measurements, and the status.
In a businessCall and meeting analysis depends on knowing who said what. This project matches speech to speakers, reports accuracy beside the share of speech it can assign with confidence, and sends uncertain cases to a review screen.
Technical skills: Python · PyTorch · WhisperX / SpeechBrain · SQLite · FastAPI · Human review workflows
- My contribution
- I set the specification and milestones, run the pipeline and review labels on the M3 Ultra, and make the implementation decisions. These include dropping the text-only speaker model and grouping utterances into speaker runs. I also review the reports.
- Result and limit
- Speaker identification is 74% accurate on individual speech segments and 86% on grouped segments of at least ten seconds, with each test example excluded from its reference bank. Of 17 errors in solo interviews, 15 came from mixed-speaker segments or incorrect labels. On separate data, confident assignments agree with caption names 95.9% of the time, covering 55–60% of speech. Two episodes labelled; no later analysis uses these transcripts yet.
On this page
Technical methods and evaluation
Jump to the pipeline, measurements, or current status.
What it does, per episode
The pipeline needs to answer three questions: when was each word spoken, which stretches contain one speaker, and whose voice is it?
First, it parses captions, detects their naming conventions, and extracts the audio. A source separator isolates vocals, and WhisperX aligns each word with the audio. Caption cues are grouped into speech segments (utterances), then into longer single-speaker runs. Names in captions and the host’s question-and-answer pattern supply some initial labels.
Next, a speaker-embedding model turns each utterance into a numerical representation for comparing voices. Trusted labels build a reference bank for each castaway, with separate references for solo interviews and camp conversation. Technically, each bank uses a recency-weighted centroid and a set of far-apart exemplars to represent variation in that voice.
Finally, the system compares unlabelled runs with those banks. Uncertain cases go to a web review queue, where I listen, label, and update the bank before reassignment. Human labels are never overwritten. The audio processing takes about four and a half minutes per 64-minute episode on the M3 Ultra.
Measured so far (US season 47, episodes 1–2)
- Speaker identification, testing each example against a voice bank built without that example (leave-one-out), over 14 speakers: 74% per utterance, 86% on runs of ten seconds or more. Four audio variants (raw, centre channel, separated vocals, both) and a second embedding model all land on the same numbers, so those alternatives did not improve identification on this sample.
- What the errors are: of 17 wrong confessional runs, 11 were two speakers merged into one run and 4 were a wrong caption-derived label; only 2 were the voice being confused. On clean single-speaker confessional runs the classifier is in the mid-90s. The consequence for the design: check run purity with the embeddings themselves before a run is scored or banked, and treat inherited names as weak labels that the bank verifies.
- Held out: episode 2 scored with a bank fitted on episode 1 agrees with the captions’ explicit names on 95.9% of confidently placed runs, but only 55–60% of body speech is placed confidently, because a one-episode bank is thin. Filling it is the review loop’s job.
- An additional end-to-end check: confessional counts derived from the transcript against survivoR’s hand counts, 17 players: Spearman 0.85 on counts, 0.88 on time, mean absolute error 1.4 confessionals per player.
- The idea that did not work: a local language model asked to name the speaker from the text alone, on 60 explicitly named runs: 26% for a 35B mixture model, 37.5% for a 27B dense one. The bar was 80%. It is retired from the pipeline and kept as a research hook; these text-only models performed poorly on this sample, so current work focuses on the audio and segmentation.
Where it stands
Two episodes are segmented and partly labelled. The export of a fine-tuning corpus (RTTM, UEM, speaker manifest, verification trials) and a fine-tuning script for the embedding model, with a baseline equal-error rate, are built and tested. They wait on about three labelled seasons of data. The >>-era seasons (21–39), where captions carry no names at all, need an on-screen-caption OCR bootstrap that is scaffolded and not yet run. When the transcripts exist, the edit and language features the boot-model plan lists (confessional in the last segment before tribal, “next time on” appearances, topics, mention sentiment) join the boot model as another labelled feature group under the same evaluation.
Stack
Python, SQLite, pandas, PyTorch on Apple Silicon (Demucs, WhisperX, SpeechBrain ECAPA, pyannote for the diarizer to come), FastAPI for the review interface, uv. Sixty-five tests on synthetic fixtures. Episodes and captions from my own library; survivoR for the cast lists.