Maxwell Grody

Who said that? Speaker-attributed Survivor transcripts

Speech · speaker identification · human-in-the-loop labelling · work in progress

Survivor’s captions often do not identify the speaker: some seasons prefix names, most mark only speaker changes, and none mark overlap. This project adds speaker names to the transcripts from audio and caption-derived labels, with uncertain cases sent for review by ear, so I can examine what individual players say. Speaker identification is 86% accurate on single-speaker segments of at least ten seconds. It runs on my own library of episodes, so there is no demo; this page is the method, the measurements, and the status.

In a businessCall and meeting analysis depends on knowing who said what. This project matches speech to speakers, reports accuracy beside the share of speech it can assign with confidence, and sends uncertain cases to a review screen.

Technical skills: Python · PyTorch · WhisperX / SpeechBrain · SQLite · FastAPI · Human review workflows

My contribution
I set the specification and milestones, run the pipeline and review labels on the M3 Ultra, and make the implementation decisions. These include dropping the text-only speaker model and grouping utterances into speaker runs. I also review the reports.
Result and limit
Speaker identification is 74% accurate on individual speech segments and 86% on grouped segments of at least ten seconds, with each test example excluded from its reference bank. Of 17 errors in solo interviews, 15 came from mixed-speaker segments or incorrect labels. On separate data, confident assignments agree with caption names 95.9% of the time, covering 55–60% of speech. Two episodes labelled; no later analysis uses these transcripts yet.
On this page

Technical methods and evaluation

Jump to the pipeline, measurements, or current status.

What it does, per episode

The pipeline needs to answer three questions: when was each word spoken, which stretches contain one speaker, and whose voice is it?

First, it parses captions, detects their naming conventions, and extracts the audio. A source separator isolates vocals, and WhisperX aligns each word with the audio. Caption cues are grouped into speech segments (utterances), then into longer single-speaker runs. Names in captions and the host’s question-and-answer pattern supply some initial labels.

Next, a speaker-embedding model turns each utterance into a numerical representation for comparing voices. Trusted labels build a reference bank for each castaway, with separate references for solo interviews and camp conversation. Technically, each bank uses a recency-weighted centroid and a set of far-apart exemplars to represent variation in that voice.

Finally, the system compares unlabelled runs with those banks. Uncertain cases go to a web review queue, where I listen, label, and update the bank before reassignment. Human labels are never overwritten. The audio processing takes about four and a half minutes per 64-minute episode on the M3 Ultra.

Measured so far (US season 47, episodes 1–2)

Where it stands

Two episodes are segmented and partly labelled. The export of a fine-tuning corpus (RTTM, UEM, speaker manifest, verification trials) and a fine-tuning script for the embedding model, with a baseline equal-error rate, are built and tested. They wait on about three labelled seasons of data. The >>-era seasons (21–39), where captions carry no names at all, need an on-screen-caption OCR bootstrap that is scaffolded and not yet run. When the transcripts exist, the edit and language features the boot-model plan lists (confessional in the last segment before tribal, “next time on” appearances, topics, mention sentiment) join the boot model as another labelled feature group under the same evaluation.

Stack

Python, SQLite, pandas, PyTorch on Apple Silicon (Demucs, WhisperX, SpeechBrain ECAPA, pyannote for the diarizer to come), FastAPI for the review interface, uv. Sixty-five tests on synthetic fixtures. Episodes and captions from my own library; survivoR for the cast lists.