trek-nlp / stylometry
Which series is this?
A paragraph of Star Trek dialogue, taken from an episode the model never trained on, with every name the lore lexicon knows replaced by [ENTITY]. Twelve series. You guess; then you see the answer, the episode, and what a bag-of-words model guessed from the same masked text. Ten passages make a round.
Passages are a paragraph long and come from across each show’s run; a passage can give away a plot point of its episode.
Masked is the hardest rung of the masking ladder. Changing the condition starts a new round.
Round over
| # | answer | you | model |
|---|
How this works
The passages are 128-token windows of subtitle dialogue from the held-out test episodes of a corpus of 1,032 episodes across twelve series; 20 per series, chosen by a fixed seed. The masked text is the ladder’s B5u condition: every term in an entity lexicon built from STAPI and the corpus (characters, places, ships, species, organisations, titles, technology) becomes one token, so the model cannot read the category either. The model is TF-IDF over words and bigrams with a logistic regression; its answer for each passage was computed once, on the test split, and stored with the passage. Nothing runs in the page but the scoring. Each guess is sent to this site as one row: passage id, condition, your guess, the answer, the model’s guess, seconds, and a random session id chosen by the page. No name, address or identifier. Once there are 200 guesses in a condition, the project page reports how people compare with the model on the same passages. Dialogue excerpts are quoted for commentary; the transcripts themselves are not published.