Which series is this?
Precomputed model answers; nothing runs in the page. Passages can spoil their episode.
Can a simple word-counting model tell which Star Trek series a paragraph of dialogue comes from? Yes, about four times in five: 78% of 128-token windows, against 19% by always guessing the biggest show. Most of that comes from spotting names, and “Odo”, “Chakotay” and “the Enterprise” are a cast list, not a writing style. The masking ladder below removes the lore one category at a time to find out how much is left in the writing itself, and the game (opens in a new tab) lets you try the hardest rung against the model.
In a businessA classifier can look accurate because it keys on names or other shortcuts. This project removes identifying terms in stages, holds out whole episodes, and measures how much of the accuracy the shortcuts supplied.
Technical skills: Python · pandas · scikit-learn · TF-IDF / logistic regression · Ablation studies · JavaScript
- My contribution
- I built the corpus from subtitle tracks, wrote the plan and the masking ladder, built the entity lexicon with STAPI and two language models, and set the evaluation (episode-level splits and intervals) before the runs.
- Result and limit
- Unmasked text: 78% of 128-token windows named correctly (19% by chance), 97% of episodes. Masking character names alone costs the most; at the last rung, with every lexicon term replaced by one token, the model still names the series well above chance, from function words and sentence shape. Numbers on the page regenerate from the run outputs.
On this page
Technical methods and evaluation
Jump to the ladder, what the model reads, confusions, or limits.
The corpus and the split
Windows, not episodes. Subtitle cues for 951 episodes are parsed, cleaned of music, stage directions and credits, deduplicated, and concatenated into windows of about 128 tokens that never cross an episode boundary. That is 40,600 windows in all. The other 81 of the 1,032 media files have no usable English track yet. Episodes are split 70/15/15 into train, validation and test, stratified by series, so every test window comes from an episode the model never saw. Accuracy is reported on windows with a 95% interval from resampling episodes, since windows of one episode are not independent, and as episode accuracy: the majority vote of an episode’s windows.
The model. TF-IDF over words and word bigrams (up to 300,000 features, minimum document frequency 3) and a multinomial logistic regression (C = 4), fit once per rung on the same split with the same seed. The settings were fixed before the ladder was run and are the same at every rung; nothing is tuned on the test split.
The lexicon. Candidate terms are the corpus’s frequent words and capitalised bigrams. Each is matched against STAPI, which covers characters, species, organisations, locations, spacecraft, titles, technology and materials, and classified by two language models given two lines of context. A category is final when the models agree. When they disagree and STAPI knows the term, STAPI’s category stands. The rest are reviewed by hand. The current version has 10,500 terms, 2,560 of them ordinary words that are never masked. A masked term becomes its category token at B1–B5 ([CHARACTER], [SHIP]) and one uniform [ENTITY] at B5u, because at B5 the model started reading the category tokens themselves.
The ladder
| rung | what is masked | window accuracy (95% interval) | macro F1 | episode accuracy |
|---|---|---|---|---|
| Majority class | Always guess the most common series. | 18.5% (12.2%–25.5%) | 0.026 | 18.1% |
| B0 · original text | Nothing masked. | 78.1% (74.9%–81.1%) | 0.699 | 97.2% |
| B1 · characters | Character names, including nicknames and forms of address. | 66.9% (62.8%–70.1%) | 0.574 | 91.7% |
| B2 · named entities | B1 plus ships, places, astronomical objects and organisations. | 63.4% (59.2%–66.6%) | 0.521 | 88.2% |
| B3 · franchise terms | B2 plus species, the franchise's technology and its substances. | 60.9% (56.8%–64.4%) | 0.493 | 88.2% |
| B4 · technical vocabulary | B3 plus real-world science and technology terms. | 60.9% (56.8%–64.5%) | 0.492 | 88.2% |
| B5 · everything, by category | B4 plus titles and ranks; each term becomes its category token. | 58.6% (54.5%–62.2%) | 0.470 | 85.4% |
| B5u · everything, one token | The B5 terms, every one replaced by the same [ENTITY] token. | 57.1% (53.3%–60.6%) | 0.460 | 84.7% |
Test split: 6,266 windows of 128 tokens from 144 held-out episodes; intervals from 500 bootstrap resamples of episodes. Episode accuracy is the majority vote of an episode’s windows.
Reading it. Each rung removes what the rung above left. The first step, character names, costs the most, eleven points. Ships, places and organisations cost three and a half more. Species and the franchise’s technology cost two and a half. The real-world technical vocabulary costs nothing at all once the proper nouns are gone. Titles and ranks cost two more. The last step, from category tokens to one token, is the honest floor: what the shows sound like when nothing they are about is left. Episode accuracy stays high at every rung because forty windows vote, and a weak signal per window compounds.
| variant | window accuracy (95% interval) | macro F1 | episode accuracy |
|---|---|---|---|
| character 2–5-grams, original text | 81.4% (78.6%–83.6%) | 0.743 | 98.6% |
| character 2–5-grams at B5 | 62.1% (58.7%–65.2%) | 0.544 | 94.4% |
| words and character n-grams together, original text | 82.3% (79.5%–84.6%) | 0.756 | 98.6% |
| words, original text, all-caps subtitle tracks excluded | 78.1% (74.8%–81.0%) | 0.699 | 97.2% |
| words at B5, all-caps subtitle tracks excluded | 58.7% (54.6%–62.3%) | 0.472 | 85.4% |
Same split, same logistic regression; one seed each. The ladder was also run with 3 seeds; the largest spread in test accuracy at any rung is 0.03 points, since the classifier is deterministic and the seed only moves the bootstrap.
Variants. Character n-grams beat words on the original text and lose less to masking: at B5 the character model keeps 62% of windows and 94% of episodes against the word model’s 59% and 85%. Part of a show’s voice is below the word: contractions, spellings, the subtitler’s habits. Excluding the episodes whose only subtitle track is in capitals changes nothing, so the capitals are not carrying the signal.
What the model reads
| series | B0 · original text | B3 · franchise terms | B5u · everything, one token |
|---|---|---|---|
| TOS The Original Series | spock, jim, captain, mr, kirk, mr spock | captain, mr, mr character, doctor, character here, sir | entity entity, sir, entity here, gentlemen, planet, men |
| TAS The Animated Series | l, spock, captain, kirk, jim, mccoy | l, captain, planet, lt, drug, beam | l, planet, drug, sir, entity entity, beam |
| TNG The Next Generation | picard, data, riker, sir, worf, counselor | sir, counselor, commander, number one, heh, character captain's | sir, number one, heh, wanna, entity entity, q |
| DS9 Deep Space Nine | odo, sisko, dax, quark, o'brien, kira | major, chief, station, the station, ship 9, the organization | station, the station, the entity, going to, entity 9, going |
| VOY Voyager | tuvok, voyager, neelix, chakotay, paris, janeway | the doctor, sick place, tech, technology, sick, character of | sick entity, technology, sick, program, assimilated, i've |
| ENT Enterprise | archer, t'pol, tucker, phlox, trip, malcolm | species, thanks, the weapon, it's, they're, bio | species, thanks, they're, it's, the weapon, warp |
| DSC Discovery | burnham, discovery, saru, michael, tilly, stamets | tech drive, jump, um, okay, drive, section 31 | entity drive, jump, um, aye entity, drive, okay |
| PIC Picard | jack, picard, rios, borg, raffi, luc | character character, admiral, organization, ée, character ée, goddamn | goddamn, entity ée, ée, and, hell, she |
| LD Lower Decks | mariner, oh, boimler, yeah, cerritos, tendi | oh, yeah, guys, chirp, ugh, whoa | oh, yeah, guys, ugh, chirp, whoa |
| PRO Prodigy | jankom, dal, protostar, gwyn, zero, janeway | whoa, ah, huh, ha, oh, ugh | whoa, ah, huh, ugh, ha, our |
| SNW Strange New Worlds | pike, spock, la'an, una, uhura, chapel | comms, okay, chin character, just, chin, gamble | comms, okay, gamble, just, chin entity, entity chin |
| SFA Starfleet Academy | caleb, cadets, sam, nahla, darem, mir | cadets, chancellor, okay, war organization, mm, hmm | okay, mm, war entity, hmm, shit, organics |
The six word or bigram features with the largest positive weight for each series in the logistic regression at that rung.
At B0 the model is a cast list. As the proper nouns go, it reads settings, ranks and each show’s favourite interjections (a station and a runabout; sir; guys and ugh), and it finds whatever names the lexicon missed, which is what each new lexicon version is for. At the last rung it is partly reading function words, contractions and the shape of sentences, which is the part of “voice” that is actually style, and partly reading two leaks the ladder does not close. Bigrams that straddle a mask survive it: entity entity is a two-word name (Mr Spock, Number One), and entity 9 and sick entity are neighbours the lexicon left behind. So the B5u floor is still an upper bound on style alone; the next rung would collapse runs of the token and drop every feature that contains it. The tables are regenerated from the run outputs, so they show the current lexicon version.
What gets confused with what
| DS9 | DSC | ENT | LD | PIC | PRO | SFA | SNW | TAS | TNG | TOS | VOY | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DS9 | 77 | 1 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 9 | 1 | 11 |
| DSC | 11 | 50 | 2 | 2 | 1 | 0 | 0 | 4 | 0 | 17 | 1 | 12 |
| ENT | 20 | 1 | 50 | 0 | 0 | 0 | 0 | 0 | 0 | 7 | 0 | 23 |
| LD | 10 | 7 | 0 | 67 | 1 | 0 | 0 | 3 | 0 | 4 | 1 | 6 |
| PIC | 23 | 19 | 2 | 7 | 18 | 1 | 1 | 3 | 0 | 15 | 2 | 10 |
| PRO | 15 | 20 | 2 | 7 | 0 | 18 | 0 | 3 | 0 | 7 | 1 | 27 |
| SFA | 20 | 27 | 2 | 10 | 0 | 0 | 4 | 4 | 0 | 20 | 2 | 11 |
| SNW | 22 | 15 | 4 | 3 | 2 | 0 | 0 | 16 | 0 | 17 | 3 | 18 |
| TAS | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 25 | 27 | 40 | 4 |
| TNG | 16 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 68 | 3 | 11 |
| TOS | 13 | 1 | 2 | 1 | 0 | 0 | 0 | 0 | 0 | 22 | 53 | 7 |
| VOY | 20 | 0 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 13 | 1 | 61 |
Largest confusions: TAS read as TOS 40% of the time (31 windows); TAS read as TNG 27% of the time (21 windows); PRO read as VOY 27% of the time (30 windows).
The game and the human baseline
The game (opens in a new tab) shows 240 held-out passages with the lore masked, 20 per series. They were chosen by a fixed seed from the test split, are at least 60 words each, exclude recaps, and are spread across each series’ test episodes. The model’s answer for each one was computed once on the test split and is stored with the passage. Every guess is logged as one row: passage, condition, guess, answer, the model’s guess, and a random session id. Once a condition has 200 guesses (a bar set before the page went up), this section reports human accuracy with a Wilson interval, the model’s accuracy on the same passages, the paired difference, and the confusions people make. Until then it is unscored.
What this does not show
One model family, TF-IDF and logistic regression; frozen-embedding and fine-tuned models are in the plan and not on this page until they run. Accuracy is on held-out episodes of the same twelve series and says nothing about a thirteenth. The ladder only removes what the lexicon knows: a term the lexicon missed is still in the text, so the floor is an upper bound on what style alone is worth. Series and era are confounded (the six streaming shows all began after 2017), and a decade classifier at each rung, to say how much of “series” is “calendar”, is planned but not yet run. Subtitles are not scripts: they carry transcription conventions, and hearing-impaired tracks carry sound descriptions, which the cleaner removes as well as it can.
Stack
Python, pandas, scikit-learn. Subtitle extraction from the media files with ffmpeg. STAPI for episode metadata and the entity lexicon. DeepSeek and GLM to classify lexicon candidates. A Cloudflare Worker and D1 for the guess log. The page is static JSON and vanilla JavaScript. The plan, the masking rules and the run log are in the repository.