Who goes home?
Precomputed predictions; no model wait. Full-season spoilers.
Who is most likely to be voted out at a Survivor tribal council? Two models trained only on other seasons rank the eligible castaways from earlier episodes’ records, across 49 US seasons. Their first pick is correct 23–29% of the time, compared with 17% by chance. I also test using only earlier seasons, check whether predicted chances match observed outcomes, and try the US-trained models on four other franchises. The explorer (opens in a new tab) shows every council with the models’ probabilities, the reasons, and who actually left.
In a businessChurn and conversion models need probabilities a team can act on, and a ranked list is not enough. This project builds features from earlier records only, validates on whole held-out seasons, and checks that predicted probabilities match what happened.
Technical skills: Python · pandas · scikit-learn / SciPy · Feature engineering · Classification / calibration · Survival analysis
- My contribution
- I set the design and evaluation plan before training, chose individual councils as the prediction unit, built the models, and reviewed the results. The model of how long players last is implemented in numpy and checked against simulated seasons.
- Result and limit
- When tested on seasons excluded from training, the models’ first pick is correct 23–29% of the time, compared with 17% by chance. The player who leaves is in their top three about 60% of the time. Predicted chances track observed exit rates closely overall. A separate analysis links prior immunity wins with a higher rate of elimination and holding an idol with a lower rate; these associations do not establish cause and effect.
On this page
Technical methods and evaluation
Jump to prediction results, how long players last, or limits.
The question, and its unit
A council, not an episode. The instance is a tribal council with a real vote, 1,109 across all franchises and 692 in the US. The choice set is everyone who attended minus anyone holding individual immunity, who cannot be voted out. Councils with no vote (final tribals, fire-making, quits, medical evacuations) are excluded and counted: 57. Ranking the whole cast before the episode airs, with the losing tribe unknown, is a different and noisier question and is not modelled here.
Only the past. Every feature at council t is built from episodes before t: challenge records, votes cast and received, pairwise voting agreement (the Pearl Islands denominators generalized to every season), advantages held at the end of the previous episode, and confessional counts. The only same-episode input is the choice set itself. The build enforces the rule by construction.
Two models and two baselines. The explorer calls these the linear model, which combines records using learned weights, and the tree model, which combines decision trees to learn more complex patterns. Technically, the linear model is a conditional logit, a linear score with a softmax within the council, with L2 chosen by season-grouped cross-validation. The tree model is gradient boosting from scikit-learn, binary, renormalized within the council, with its size chosen the same way. The two baselines are a uniform guess and “the castaway with the most votes received so far”. Data: the survivoR tables via Gamebot, my Survivor warehouse.
Result
Reading the table. “First pick right” is how often the highest-ranked player leaves; “in the top 3” is how often the player who leaves appears among the three highest-ranked. Log loss measures probability error: lower is better, with a larger penalty for confident mistakes. “Uniform” gives every eligible player an equal chance.
| model | first pick right | in the top 3 | log loss |
|---|---|---|---|
| uniform | 16.8% | 33.2% | 1.841 |
| most votes received so far | 21.4% | 32.7% | 2.942 |
| conditional logit | 23.4% | 60.8% | 1.792 |
| gradient boosting | 29.4% | 60.5% | 1.775 |
Five season-grouped folds over 49 US seasons, 692 councils, 4,589 eligible castaway-council rows. Top-1 splits ties evenly; log loss is of the probability given to the castaway who left. Which seasons share a fold moves top-1 by about two points either way (a different fold assignment gave the logit 24.9% and boosting 28.0%), which is the resolution these numbers have; the forward chain has no fold assignment.
Better than guessing, well short of knowing. The first pick is right a quarter of the time, against one in six by chance. The castaway who leaves is in the top three about 60% of the time, against a third. The forward chain trains on the seasons before s and scores s, over the last ten seasons. It gives the same picture, logit 24.4% and boosting 28.5% top-1, so the gain is not an artefact of pooling eras. Post-merge councils are more predictable than pre-merge ones, logit 27.0% against 19.3%. Before the merge the tables know little about a tribe that has not been to council yet.
The probabilities are well calibrated on this test. Expected calibration error is the average gap between predicted chances and observed exit rates, weighted by the number of predictions in each group. It is 0.012, or 1.2 percentage points, for both models. The castaways given 30% went home about 30% of the time. The probabilities are well calibrated in aggregate but rarely confident. The most confident decile sits near 35%.
Confessional counts add little in this test. Without the confessional features the logit’s top-1 moves by +0.2 points and boosting’s by −0.4, within the fold noise. Without the alliance features, +2.3 and −2.3, the same. Without the voting history, −0.7 and −4.9, and log loss worsens most for both: it is the one group whose removal shows on every measure. The tables’ strongest signals are the boring ones: votes received at the last council, votes cast for people who stayed, and whether the castaway holds an idol.
A second feature set, v1.1, gives no clear improvement. After the v1 run I added what else the tables can say: advantage history (found, played, for whom, correctly or not), journeys, a split-vote-aware majority count and the merge vote, original-tribe numbers at the council and its gender and returnee mix, and season format and twist flags. Under the same folds the logit moves from 23.4% to 24.6% top-1 with a slightly worse log loss (1.801 against 1.792) and boosting from 29.4% to 29.0% (1.778 against 1.775): nothing outside the fold noise, and removing any one of the new groups changes nothing either. Recency weighting, down-weighting old seasons with a half-life of 6, 12 or 24 seasons, makes the forward-chain numbers worse at every half-life for both models: these weighting schemes did not improve prediction in the tested periods. Both are reported in the fit report and the explorer keeps v1. The tested feature additions did not improve these models. A separate project on speaker-attributed transcripts will explore whether what people say adds useful information.
The two models disagree about why. The rank agreement of their permutation importances is Spearman ρ = 0.40. The logit leans on idols, wrong votes, the share of correct votes and being in the majority last time. Boosting leans on idols, votes at the last council and the confessional share, and hardly at all on the majority flag the logit rates highly. The models have similar accuracy but rely on different features. I report both explanations rather than treating either as the definitive account.
Transfer. Applied to the other franchises without refitting, the US models beat a uniform guess on top-1 everywhere. In Australia they score 18.6% and 17.2% against 14.1%. In South Africa they score 24.0% and 22.7% against 16.8%. The UK and New Zealand, with 29 and 39 councils, are too small to read. On Australian councils the logit’s log loss (2.09) is worse than uniform’s (2.04), so its probabilities are miscalibrated there. That format has 24 castaways and non-elimination episodes, which the US model never saw.
How long do they last?
The second question is longevity: episodes survived, with being voted out as the event. Winners, finalists, quits and evacuations are censored: their records end without counting them as vote-outs. The Kaplan–Meier curves estimate the share not yet voted out as the season progresses, accounting for those incomplete records. Grouped by returning-player status, gender or three equally sized age groups, the curves barely separate (log-rank p = 0.10, 0.22, 0.95). A Cox model on those three covariates, stratified by season, has a cross-validated concordance of 0.52 against 0.50 for chance. Concordance is how often the model correctly orders comparable pairs by who lasts longer. These three characteristics provide little predictive discrimination in this model.
A second Cox model includes changing game conditions, each measured through the previous episode:
| covariate | hazard ratio | 95% interval | proportional-hazards test p |
|---|---|---|---|
| individual immunity wins to date | 1.20 | 1.03–1.40 | 0.08 |
| votes received to date | 1.09 per vote | 1.05–1.13 | 0.33 |
| holds a hidden idol | 0.48 | 0.35–0.65 | 0.83 |
| returning player | 0.70 | 0.44–1.11 | 0.64 |
| woman | 1.08 | 0.92–1.26 | 0.02 |
| age (per SD) | 0.99 | 0.91–1.07 | 0.08 |
A hazard ratio compares the rate of being voted out among players still in the game: 1 means no difference, above 1 means a higher rate, and below 1 a lower rate. It is not a player’s probability of leaving at a particular council. “Per SD” means per one standard deviation of age, a measure of the cast’s age spread. The final column tests whether each ratio can reasonably be treated as constant over the season.
Each prior immunity win is associated with about a 20% higher elimination hazard. Holding an idol is associated with about half the hazard. Each prior vote received is associated with a 9% increase. These are conditional associations, not estimates of what would happen if a player won a challenge or acquired an idol. The proportional-hazards test (Grambsch–Therneau on scaled Schoenfeld residuals) rejects for gender, so that hazard ratio is an average of an effect that changes over the season and should not be read as a constant. For the others it does not reject, which with 690 events is not proof that they are constant either. The estimator is written out in numpy, with Breslow ties and one stratum per season so the risk set is the season’s own cast. It recovers known coefficients from simulated seasons before it sees the real ones.
What this does not show
The model sees what the tables record, not the game: no confessional text, no “next time on”, no tribal-council talk. A calibrated 30% means that about three in ten comparable predictions end in a vote-out; it does not tell us which individuals will leave. Roughly 28% accuracy describes the tested models, not a proven ceiling for these features. The season-by-season explorer shows both successful predictions and missed blindsides.
Plan and departures
The evaluation was written before fitting (SPEC.md). There were three departures. LightGBM lambdarank became scikit-learn boosting with within-council renormalization, so the same library serves both models and the probabilities can be scored. SHAP became permutation importance on held-out councils for both models, which measures the same thing on the same footing. The alliance ally-count feature was dropped for its share, the era-robust version. The survival model was stratified by season rather than pooled, so its ties are only double boots. The v1.1 feature set and the recency weights were added after the v1 numbers were seen and are reported beside them, not in their place.
Stack
Python, pandas, numpy, scipy, scikit-learn. Data from the survivoR project (Daniel Oehm, MIT) through Gamebot, the warehouse described on its own page. The explorer is a static page over one JSON file of out-of-fold probabilities; no server.