Maxwell Grody

Who goes home?

Classification · calibration · survival analysis · 49 seasons of Survivor

Open the explorer ↗ (opens in a new tab) · Repository ↗ · Plan ↗

Precomputed predictions; no model wait. Full-season spoilers.

Who is most likely to be voted out at a Survivor tribal council? Two models trained only on other seasons rank the eligible castaways from earlier episodes’ records, across 49 US seasons. Their first pick is correct 23–29% of the time, compared with 17% by chance. I also test using only earlier seasons, check whether predicted chances match observed outcomes, and try the US-trained models on four other franchises. The explorer (opens in a new tab) shows every council with the models’ probabilities, the reasons, and who actually left.

In a businessChurn and conversion models need probabilities a team can act on, and a ranked list is not enough. This project builds features from earlier records only, validates on whole held-out seasons, and checks that predicted probabilities match what happened.

Technical skills: Python · pandas · scikit-learn / SciPy · Feature engineering · Classification / calibration · Survival analysis

My contribution
I set the design and evaluation plan before training, chose individual councils as the prediction unit, built the models, and reviewed the results. The model of how long players last is implemented in numpy and checked against simulated seasons.
Result and limit
When tested on seasons excluded from training, the models’ first pick is correct 23–29% of the time, compared with 17% by chance. The player who leaves is in their top three about 60% of the time. Predicted chances track observed exit rates closely overall. A separate analysis links prior immunity wins with a higher rate of elimination and holding an idol with a lower rate; these associations do not establish cause and effect.
On this page

Technical methods and evaluation

Jump to prediction results, how long players last, or limits.

The question, and its unit

A council, not an episode. The instance is a tribal council with a real vote, 1,109 across all franchises and 692 in the US. The choice set is everyone who attended minus anyone holding individual immunity, who cannot be voted out. Councils with no vote (final tribals, fire-making, quits, medical evacuations) are excluded and counted: 57. Ranking the whole cast before the episode airs, with the losing tribe unknown, is a different and noisier question and is not modelled here.

Only the past. Every feature at council t is built from episodes before t: challenge records, votes cast and received, pairwise voting agreement (the Pearl Islands denominators generalized to every season), advantages held at the end of the previous episode, and confessional counts. The only same-episode input is the choice set itself. The build enforces the rule by construction.

Two models and two baselines. The explorer calls these the linear model, which combines records using learned weights, and the tree model, which combines decision trees to learn more complex patterns. Technically, the linear model is a conditional logit, a linear score with a softmax within the council, with L2 chosen by season-grouped cross-validation. The tree model is gradient boosting from scikit-learn, binary, renormalized within the council, with its size chosen the same way. The two baselines are a uniform guess and “the castaway with the most votes received so far”. Data: the survivoR tables via Gamebot, my Survivor warehouse.

Result

Reading the table. “First pick right” is how often the highest-ranked player leaves; “in the top 3” is how often the player who leaves appears among the three highest-ranked. Log loss measures probability error: lower is better, with a larger penalty for confident mistakes. “Uniform” gives every eligible player an equal chance.

model first pick right in the top 3 log loss
uniform 16.8% 33.2% 1.841
most votes received so far 21.4% 32.7% 2.942
conditional logit 23.4% 60.8% 1.792
gradient boosting 29.4% 60.5% 1.775

Five season-grouped folds over 49 US seasons, 692 councils, 4,589 eligible castaway-council rows. Top-1 splits ties evenly; log loss is of the probability given to the castaway who left. Which seasons share a fold moves top-1 by about two points either way (a different fold assignment gave the logit 24.9% and boosting 28.0%), which is the resolution these numbers have; the forward chain has no fold assignment.

Better than guessing, well short of knowing. The first pick is right a quarter of the time, against one in six by chance. The castaway who leaves is in the top three about 60% of the time, against a third. The forward chain trains on the seasons before s and scores s, over the last ten seasons. It gives the same picture, logit 24.4% and boosting 28.5% top-1, so the gain is not an artefact of pooling eras. Post-merge councils are more predictable than pre-merge ones, logit 27.0% against 19.3%. Before the merge the tables know little about a tribe that has not been to council yet.

Reliability of the boot probability (v1 models, season-grouped CV, deciles)0.00.00.10.10.20.20.30.30.40.40.50.5predicted probability of going home (decile mean)share who went homeconditional logit: predicted 0.049, observed 0.076 (n = 459)conditional logit: predicted 0.078, observed 0.081 (n = 459)conditional logit: predicted 0.095, observed 0.107 (n = 459)conditional logit: predicted 0.108, observed 0.107 (n = 459)conditional logit: predicted 0.121, observed 0.122 (n = 458)conditional logit: predicted 0.135, observed 0.142 (n = 459)conditional logit: predicted 0.156, observed 0.146 (n = 459)conditional logit: predicted 0.181, observed 0.179 (n = 459)conditional logit: predicted 0.223, observed 0.235 (n = 459)conditional logit: predicted 0.360, observed 0.314 (n = 459)conditional logit · ECE 0.012boosting: predicted 0.070, observed 0.087 (n = 459)boosting: predicted 0.092, observed 0.089 (n = 459)boosting: predicted 0.102, observed 0.107 (n = 459)boosting: predicted 0.112, observed 0.100 (n = 459)boosting: predicted 0.123, observed 0.096 (n = 458)boosting: predicted 0.136, observed 0.135 (n = 459)boosting: predicted 0.154, observed 0.116 (n = 458)boosting: predicted 0.178, observed 0.217 (n = 460)boosting: predicted 0.218, observed 0.205 (n = 459)boosting: predicted 0.323, observed 0.355 (n = 459)boosting · ECE 0.019
Reliability of the out-of-fold boot probabilities by decile. A point on the diagonal means a predicted 20% went home 20% of the time. Hover a point for the numbers.

The probabilities are well calibrated on this test. Expected calibration error is the average gap between predicted chances and observed exit rates, weighted by the number of predictions in each group. It is 0.012, or 1.2 percentage points, for both models. The castaways given 30% went home about 30% of the time. The probabilities are well calibrated in aggregate but rarely confident. The most confident decile sits near 35%.

Confessional counts add little in this test. Without the confessional features the logit’s top-1 moves by +0.2 points and boosting’s by −0.4, within the fold noise. Without the alliance features, +2.3 and −2.3, the same. Without the voting history, −0.7 and −4.9, and log loss worsens most for both: it is the one group whose removal shows on every measure. The tables’ strongest signals are the boring ones: votes received at the last council, votes cast for people who stayed, and whether the castaway holds an idol.

Permutation importance, top 14 features (held-out increase in log loss)holds a hidden idolholds a hidden idol · conditional logit: +0.0323holds a hidden idol · boosting: +0.0157votes cast for someone who stayedvotes cast for someone who stayed · conditional logit: +0.0339votes cast for someone who stayed · boosting: +0.0062share of votes cast for the person who leftshare of votes cast for the person who left · conditional logit: +0.0331share of votes cast for the person who left · boosting: +0.0048votes received at the last councilvotes received at the last council · conditional logit: +0.0178votes received at the last council · boosting: +0.0182in the majority at the last councilin the majority at the last council · conditional logit: +0.0339in the majority at the last council · boosting: +0.0011confessional share, season to dateconfessional share, season to date · conditional logit: +0.0127confessional share, season to date · boosting: +0.0168consecutive wrong votesconsecutive wrong votes · conditional logit: +0.0104consecutive wrong votes · boosting: +0.0124prior seasons playedprior seasons played · conditional logit: +0.0205prior seasons played · boosting: +0.0000tribe challenge win ratetribe challenge win rate · conditional logit: +0.0141tribe challenge win rate · boosting: +0.0041individual immunity winsindividual immunity wins · conditional logit: +0.0074individual immunity wins · boosting: +0.0107councils with a vote receivedcouncils with a vote received · conditional logit: +0.0174councils with a vote received · boosting: +0.0005tribe challenges wontribe challenges won · conditional logit: +0.0170tribe challenges won · boosting: +0.0006mean agreement with the councilmean agreement with the council · conditional logit: +0.0049mean agreement with the council · boosting: +0.0102tribe challenges enteredtribe challenges entered · conditional logit: +0.0105tribe challenges entered · boosting: -0.0002conditional logitboostingrank agreement between the models: Spearman ρ = 0.43
Permutation importance on held-out councils: the increase in log loss when one feature is shuffled (across councils for a feature shared by the whole council). Hover a bar for the value.

A second feature set, v1.1, gives no clear improvement. After the v1 run I added what else the tables can say: advantage history (found, played, for whom, correctly or not), journeys, a split-vote-aware majority count and the merge vote, original-tribe numbers at the council and its gender and returnee mix, and season format and twist flags. Under the same folds the logit moves from 23.4% to 24.6% top-1 with a slightly worse log loss (1.801 against 1.792) and boosting from 29.4% to 29.0% (1.778 against 1.775): nothing outside the fold noise, and removing any one of the new groups changes nothing either. Recency weighting, down-weighting old seasons with a half-life of 6, 12 or 24 seasons, makes the forward-chain numbers worse at every half-life for both models: these weighting schemes did not improve prediction in the tested periods. Both are reported in the fit report and the explorer keeps v1. The tested feature additions did not improve these models. A separate project on speaker-attributed transcripts will explore whether what people say adds useful information.

The two models disagree about why. The rank agreement of their permutation importances is Spearman ρ = 0.40. The logit leans on idols, wrong votes, the share of correct votes and being in the majority last time. Boosting leans on idols, votes at the last council and the confessional share, and hardly at all on the majority flag the logit rates highly. The models have similar accuracy but rely on different features. I report both explanations rather than treating either as the definitive account.

Transfer. Applied to the other franchises without refitting, the US models beat a uniform guess on top-1 everywhere. In Australia they score 18.6% and 17.2% against 14.1%. In South Africa they score 24.0% and 22.7% against 16.8%. The UK and New Zealand, with 29 and 39 councils, are too small to read. On Australian councils the logit’s log loss (2.09) is worse than uniform’s (2.04), so its probabilities are miscalibrated there. That format has 24 castaways and non-elimination episodes, which the US model never saw.

How long do they last?

Survival by returning-player status (Kaplan–Meier, 95% bands)02468101214160.000.250.500.751.00episodes survivedshare still in the gamefirst time: after episode 1, 94% still in (±2), 696 at riskfirst time: after episode 2, 88% still in (±2), 647 at riskfirst time: after episode 3, 83% still in (±3), 606 at riskfirst time: after episode 4, 77% still in (±3), 563 at riskfirst time: after episode 5, 72% still in (±3), 516 at riskfirst time: after episode 6, 66% still in (±3), 474 at riskfirst time: after episode 7, 60% still in (±4), 430 at riskfirst time: after episode 8, 55% still in (±4), 388 at riskfirst time: after episode 9, 49% still in (±4), 343 at riskfirst time: after episode 10, 43% still in (±4), 303 at riskfirst time: after episode 11, 38% still in (±4), 264 at riskfirst time: after episode 12, 32% still in (±3), 222 at riskfirst time: after episode 13, 25% still in (±3), 169 at riskfirst time: after episode 14, 16% still in (±3), 72 at riskfirst time: after episode 15, 11% still in (±4), 16 at riskfirst time (n = 739)returning: after episode 1, 94% still in (±4), 134 at riskreturning: after episode 2, 89% still in (±5), 127 at riskreturning: after episode 3, 85% still in (±6), 121 at riskreturning: after episode 4, 82% still in (±6), 114 at riskreturning: after episode 5, 77% still in (±7), 107 at riskreturning: after episode 6, 71% still in (±8), 98 at riskreturning: after episode 7, 66% still in (±8), 90 at riskreturning: after episode 8, 61% still in (±8), 82 at riskreturning: after episode 9, 55% still in (±8), 74 at riskreturning: after episode 10, 48% still in (±8), 65 at riskreturning: after episode 11, 42% still in (±8), 57 at riskreturning: after episode 12, 36% still in (±8), 49 at riskreturning: after episode 13, 31% still in (±8), 38 at riskreturning: after episode 14, 23% still in (±7), 27 at riskreturning: after episode 15, 16% still in (±9), 5 at riskreturning: after episode 16, 8% still in (±9), 2 at riskreturning (n = 143)log-rank p = 0.097
Kaplan–Meier survival by returning-player status, US seasons, with 95% bands. Hover a step for the numbers.

The second question is longevity: episodes survived, with being voted out as the event. Winners, finalists, quits and evacuations are censored: their records end without counting them as vote-outs. The Kaplan–Meier curves estimate the share not yet voted out as the season progresses, accounting for those incomplete records. Grouped by returning-player status, gender or three equally sized age groups, the curves barely separate (log-rank p = 0.10, 0.22, 0.95). A Cox model on those three covariates, stratified by season, has a cross-validated concordance of 0.52 against 0.50 for chance. Concordance is how often the model correctly orders comparable pairs by who lasts longer. These three characteristics provide little predictive discrimination in this model.

A second Cox model includes changing game conditions, each measured through the previous episode:

covariate hazard ratio 95% interval proportional-hazards test p
individual immunity wins to date 1.20 1.03–1.40 0.08
votes received to date 1.09 per vote 1.05–1.13 0.33
holds a hidden idol 0.48 0.35–0.65 0.83
returning player 0.70 0.44–1.11 0.64
woman 1.08 0.92–1.26 0.02
age (per SD) 0.99 0.91–1.07 0.08

A hazard ratio compares the rate of being voted out among players still in the game: 1 means no difference, above 1 means a higher rate, and below 1 a lower rate. It is not a player’s probability of leaving at a particular council. “Per SD” means per one standard deviation of age, a measure of the cast’s age spread. The final column tests whether each ratio can reasonably be treated as constant over the season.

Each prior immunity win is associated with about a 20% higher elimination hazard. Holding an idol is associated with about half the hazard. Each prior vote received is associated with a 9% increase. These are conditional associations, not estimates of what would happen if a player won a challenge or acquired an idol. The proportional-hazards test (Grambsch–Therneau on scaled Schoenfeld residuals) rejects for gender, so that hazard ratio is an average of an effect that changes over the season and should not be read as a constant. For the others it does not reject, which with 690 events is not proof that they are constant either. The estimator is written out in numpy, with Breslow ties and one stratum per season so the risk set is the season’s own cast. It recovers known coefficients from simulated seasons before it sees the real ones.

What this does not show

The model sees what the tables record, not the game: no confessional text, no “next time on”, no tribal-council talk. A calibrated 30% means that about three in ten comparable predictions end in a vote-out; it does not tell us which individuals will leave. Roughly 28% accuracy describes the tested models, not a proven ceiling for these features. The season-by-season explorer shows both successful predictions and missed blindsides.

Plan and departures

The evaluation was written before fitting (SPEC.md). There were three departures. LightGBM lambdarank became scikit-learn boosting with within-council renormalization, so the same library serves both models and the probabilities can be scored. SHAP became permutation importance on held-out councils for both models, which measures the same thing on the same footing. The alliance ally-count feature was dropped for its share, the era-robust version. The survival model was stratified by season rather than pooled, so its ties are only double boots. The v1.1 feature set and the recency weights were added after the v1 numbers were seen and are reported beside them, not in their place.

Stack

Python, pandas, numpy, scipy, scikit-learn. Data from the survivoR project (Daniel Oehm, MIT) through Gamebot, the warehouse described on its own page. The explorer is a static page over one JSON file of out-of-fold probabilities; no server.