Maxwell Grody

Which series is this?

Text classification · masking ladder · 951 episodes of Star Trek

Play the guessing game ↗ (opens in a new tab) · Repository ↗ · Plan ↗

Precomputed model answers; nothing runs in the page. Passages can spoil their episode.

Can a simple word-counting model tell which Star Trek series a paragraph of dialogue comes from? Yes, about four times in five: 78% of 128-token windows, against 19% by always guessing the biggest show. Most of that comes from spotting names, and “Odo”, “Chakotay” and “the Enterprise” are a cast list, not a writing style. The masking ladder below removes the lore one category at a time to find out how much is left in the writing itself, and the game (opens in a new tab) lets you try the hardest rung against the model.

In a businessA classifier can look accurate because it keys on names or other shortcuts. This project removes identifying terms in stages, holds out whole episodes, and measures how much of the accuracy the shortcuts supplied.

Technical skills: Python · pandas · scikit-learn · TF-IDF / logistic regression · Ablation studies · JavaScript

My contribution
I built the corpus from subtitle tracks, wrote the plan and the masking ladder, built the entity lexicon with STAPI and two language models, and set the evaluation (episode-level splits and intervals) before the runs.
Result and limit
Unmasked text: 78% of 128-token windows named correctly (19% by chance), 97% of episodes. Masking character names alone costs the most; at the last rung, with every lexicon term replaced by one token, the model still names the series well above chance, from function words and sentence shape. Numbers on the page regenerate from the run outputs.
On this page

Technical methods and evaluation

Jump to the ladder, what the model reads, confusions, or limits.

The corpus and the split

Windows, not episodes. Subtitle cues for 951 episodes are parsed, cleaned of music, stage directions and credits, deduplicated, and concatenated into windows of about 128 tokens that never cross an episode boundary. That is 40,600 windows in all. The other 81 of the 1,032 media files have no usable English track yet. Episodes are split 70/15/15 into train, validation and test, stratified by series, so every test window comes from an episode the model never saw. Accuracy is reported on windows with a 95% interval from resampling episodes, since windows of one episode are not independent, and as episode accuracy: the majority vote of an episode’s windows.

The model. TF-IDF over words and word bigrams (up to 300,000 features, minimum document frequency 3) and a multinomial logistic regression (C = 4), fit once per rung on the same split with the same seed. The settings were fixed before the ladder was run and are the same at every rung; nothing is tuned on the test split.

The lexicon. Candidate terms are the corpus’s frequent words and capitalised bigrams. Each is matched against STAPI, which covers characters, species, organisations, locations, spacecraft, titles, technology and materials, and classified by two language models given two lines of context. A category is final when the models agree. When they disagree and STAPI knows the term, STAPI’s category stands. The rest are reviewed by hand. The current version has 10,500 terms, 2,560 of them ordinary words that are never masked. A masked term becomes its category token at B1–B5 ([CHARACTER], [SHIP]) and one uniform [ENTITY] at B5u, because at B5 the model started reading the category tokens themselves.

The ladder

rungwhat is maskedwindow accuracy (95% interval)macro F1episode accuracy
Majority classAlways guess the most common series.18.5% (12.2%–25.5%)0.02618.1%
B0 · original textNothing masked.78.1% (74.9%–81.1%)0.69997.2%
B1 · charactersCharacter names, including nicknames and forms of address.66.9% (62.8%–70.1%)0.57491.7%
B2 · named entitiesB1 plus ships, places, astronomical objects and organisations.63.4% (59.2%–66.6%)0.52188.2%
B3 · franchise termsB2 plus species, the franchise's technology and its substances.60.9% (56.8%–64.4%)0.49388.2%
B4 · technical vocabularyB3 plus real-world science and technology terms.60.9% (56.8%–64.5%)0.49288.2%
B5 · everything, by categoryB4 plus titles and ranks; each term becomes its category token.58.6% (54.5%–62.2%)0.47085.4%
B5u · everything, one tokenThe B5 terms, every one replaced by the same [ENTITY] token.57.1% (53.3%–60.6%)0.46084.7%

Test split: 6,266 windows of 128 tokens from 144 held-out episodes; intervals from 500 bootstrap resamples of episodes. Episode accuracy is the majority vote of an episode’s windows.

Reading it. Each rung removes what the rung above left. The first step, character names, costs the most, eleven points. Ships, places and organisations cost three and a half more. Species and the franchise’s technology cost two and a half. The real-world technical vocabulary costs nothing at all once the proper nouns are gone. Titles and ranks cost two more. The last step, from category tokens to one token, is the honest floor: what the shows sound like when nothing they are about is left. Episode accuracy stays high at every rung because forty windows vote, and a weak signal per window compounds.

variantwindow accuracy (95% interval)macro F1episode accuracy
character 2–5-grams, original text81.4% (78.6%–83.6%)0.74398.6%
character 2–5-grams at B562.1% (58.7%–65.2%)0.54494.4%
words and character n-grams together, original text82.3% (79.5%–84.6%)0.75698.6%
words, original text, all-caps subtitle tracks excluded78.1% (74.8%–81.0%)0.69997.2%
words at B5, all-caps subtitle tracks excluded58.7% (54.6%–62.3%)0.47285.4%

Same split, same logistic regression; one seed each. The ladder was also run with 3 seeds; the largest spread in test accuracy at any rung is 0.03 points, since the classifier is deterministic and the seed only moves the bootstrap.

Variants. Character n-grams beat words on the original text and lose less to masking: at B5 the character model keeps 62% of windows and 94% of episodes against the word model’s 59% and 85%. Part of a show’s voice is below the word: contractions, spellings, the subtitler’s habits. Excluding the episodes whose only subtitle track is in capitals changes nothing, so the capitals are not carrying the signal.

What the model reads

seriesB0 · original textB3 · franchise termsB5u · everything, one token
TOS The Original Seriesspock, jim, captain, mr, kirk, mr spockcaptain, mr, mr character, doctor, character here, sirentity entity, sir, entity here, gentlemen, planet, men
TAS The Animated Seriesl, spock, captain, kirk, jim, mccoyl, captain, planet, lt, drug, beaml, planet, drug, sir, entity entity, beam
TNG The Next Generationpicard, data, riker, sir, worf, counselorsir, counselor, commander, number one, heh, character captain'ssir, number one, heh, wanna, entity entity, q
DS9 Deep Space Nineodo, sisko, dax, quark, o'brien, kiramajor, chief, station, the station, ship 9, the organizationstation, the station, the entity, going to, entity 9, going
VOY Voyagertuvok, voyager, neelix, chakotay, paris, janewaythe doctor, sick place, tech, technology, sick, character ofsick entity, technology, sick, program, assimilated, i've
ENT Enterprisearcher, t'pol, tucker, phlox, trip, malcolmspecies, thanks, the weapon, it's, they're, biospecies, thanks, they're, it's, the weapon, warp
DSC Discoveryburnham, discovery, saru, michael, tilly, stametstech drive, jump, um, okay, drive, section 31entity drive, jump, um, aye entity, drive, okay
PIC Picardjack, picard, rios, borg, raffi, luccharacter character, admiral, organization, ée, character ée, goddamngoddamn, entity ée, ée, and, hell, she
LD Lower Decksmariner, oh, boimler, yeah, cerritos, tendioh, yeah, guys, chirp, ugh, whoaoh, yeah, guys, ugh, chirp, whoa
PRO Prodigyjankom, dal, protostar, gwyn, zero, janewaywhoa, ah, huh, ha, oh, ughwhoa, ah, huh, ugh, ha, our
SNW Strange New Worldspike, spock, la'an, una, uhura, chapelcomms, okay, chin character, just, chin, gamblecomms, okay, gamble, just, chin entity, entity chin
SFA Starfleet Academycaleb, cadets, sam, nahla, darem, mircadets, chancellor, okay, war organization, mm, hmmokay, mm, war entity, hmm, shit, organics

The six word or bigram features with the largest positive weight for each series in the logistic regression at that rung.

At B0 the model is a cast list. As the proper nouns go, it reads settings, ranks and each show’s favourite interjections (a station and a runabout; sir; guys and ugh), and it finds whatever names the lexicon missed, which is what each new lexicon version is for. At the last rung it is partly reading function words, contractions and the shape of sentences, which is the part of “voice” that is actually style, and partly reading two leaks the ladder does not close. Bigrams that straddle a mask survive it: entity entity is a two-word name (Mr Spock, Number One), and entity 9 and sick entity are neighbours the lexicon left behind. So the B5u floor is still an upper bound on style alone; the next rung would collapse runs of the token and drop every feature that contains it. The tables are regenerated from the run outputs, so they show the current lexicon version.

What gets confused with what

Rows: the true series; columns: the model’s guess at B5u; cells: percent of the row.
DS9DSCENTLDPICPROSFASNWTASTNGTOSVOY
DS977120000009111
DSC1150221004017112
ENT201500000007023
LD10706710030416
PIC23192718113015210
PRO1520270180307127
SFA20272100044020211
SNW22154320016017318
TAS400000002527404
TNG161100000068311
TOS131210000022537
VOY200400000013161

Largest confusions: TAS read as TOS 40% of the time (31 windows); TAS read as TNG 27% of the time (21 windows); PRO read as VOY 27% of the time (30 windows).

The game and the human baseline

The game (opens in a new tab) shows 240 held-out passages with the lore masked, 20 per series. They were chosen by a fixed seed from the test split, are at least 60 words each, exclude recaps, and are spread across each series’ test episodes. The model’s answer for each one was computed once on the test split and is stored with the passage. Every guess is logged as one row: passage, condition, guess, answer, the model’s guess, and a random session id. Once a condition has 200 guesses (a bar set before the page went up), this section reports human accuracy with a Wilson interval, the model’s accuracy on the same passages, the paired difference, and the confusions people make. Until then it is unscored.

What this does not show

One model family, TF-IDF and logistic regression; frozen-embedding and fine-tuned models are in the plan and not on this page until they run. Accuracy is on held-out episodes of the same twelve series and says nothing about a thirteenth. The ladder only removes what the lexicon knows: a term the lexicon missed is still in the text, so the floor is an upper bound on what style alone is worth. Series and era are confounded (the six streaming shows all began after 2017), and a decade classifier at each rung, to say how much of “series” is “calendar”, is planned but not yet run. Subtitles are not scripts: they carry transcription conventions, and hearing-impaired tracks carry sound descriptions, which the cleaner removes as well as it can.

Stack

Python, pandas, scikit-learn. Subtitle extraction from the media files with ffmpeg. STAPI for episode metadata and the entity lexicon. DeepSeek and GLM to classify lexicon candidates. A Cloudflare Worker and D1 for the guess log. The page is static JSON and vanilla JavaScript. The plan, the masking rules and the run log are in the repository.