Gamebot / Survivor analytics
Who goes home?
Pick a season of the CBS TV show Survivor and a tribal council. Two models, trained only on other seasons, rank the castaways who could be voted out that night using what the tables record before the episode: challenge records, votes cast and received, who has voted with whom, idols held, and how often each player appears in solo interviews. The castaway who left is marked, so every prediction can be checked against what happened.
Full-season spoilers. A 30% chance does not mean someone will leave. Across many predictions near 30%, we check whether roughly three in ten players actually leave.
Linear model (conditional logit): combines each player’s records using learned weights. Tree model (gradient boosting): combines decision trees that can learn more complex patterns. Each column gives the estimated chance of being voted out and totals 100%, apart from rounding. The reasons show which records most raise (↑) or lower (↓) the linear model’s score relative to the council average; parentheses give the recorded value. These explain the model’s prediction, not why the players voted as they did.
| castaway | went home | Vote-out chance Linear model | Vote-out chance Tree model | What influences the linear model |
|---|
This season, council by council
The model's first pick at each council, and whether that castaway went home. A hit rate over one season is a handful of councils; the results across all 49 seasons below give a fuller picture.
| episode | eligible | went home | Linear model’s pick | Tree model’s pick |
|---|
How good is it?
| model | First pick correct | Actual exit in top 3 | Probability error (log loss) |
|---|
The 49 US seasons are divided into five groups. Each group is tested using models trained on the other four, so no model trains on the season it predicts. The equal-chance baseline gives every eligible player the same probability. “First pick correct” shares credit evenly when players tie. Log loss measures probability error: lower is better, and confident mistakes cost more. The project page also tests predictions using only earlier seasons, checks whether the probabilities match observed rates, removes groups of inputs to see what helps, and examines how long players last.
Methods and data
Data: the survivoR project's tables (MIT licence, licence file), through Gamebot, my warehouse of Survivor statistics. A council is an instance when a real vote was cast; the eligible players are everyone who attended except anyone holding individual immunity. Every input record uses only episodes before the council. Solo interviews with players are called confessionals in the data. The probabilities shown are out-of-fold: the model that scored a season was fitted on the other seasons. Code and the pre-registered evaluation plan are in the repository.