Maxwell Grody

The Oscar bump

Causal inference · event study · pre-registered design

Repository ↗ · Plan (fixed version) ↗

This project asks whether winning Best Picture changes a film’s rating activity. I wrote the analysis plan before running the estimates. The difference-in-differences event study uses timestamped MovieLens ratings to compare winners with the other nominees from the same ceremony.

In a businessMarketing teams need to separate a campaign's effect from trends already underway. This project uses a comparison group, an event study, and forecast backtests, and it documents a leakage mistake I found and corrected.

Technical skills: Python · pandas / NumPy · scikit-learn · Event studies · Uncertainty estimation · Forecast backtesting

My contribution
I set the design and the pre-registered plan, built the pipeline, spot-checked the events table against the sources, and ran the estimates.
Result and limit
Winners show a brief rise in rating activity (about 20% over two weeks, a comparison added after seeing the data). The planned six-month estimate is too imprecise to establish a sustained effect: its uncertainty interval allows roughly a 13% decrease to a 20% increase. Winners were already gaining before the ceremony, limiting what this study can attribute to the award.
On this page

Technical methods and evaluation

Jump to the design, results, or forecasting.

Design

Treatment and control. The Best Picture winner of each ceremony against the other nominees of the same ceremony, week by week: 28 ceremonies (1996–2023), 196 films, 28 winners. The films share a category, nomination year, and ceremony. They can still differ in release timing, audience, and momentum before the award. Winning is not randomly assigned among nominees. If winners were already gaining attention before the award, then a later difference could reflect that earlier momentum. Checking the trends before the ceremony is therefore necessary before attributing a change to winning.

Outcome and window. log(1 + weekly ratings), and mean weekly rating, from 26 weeks before the ceremony to 26 after; week 0 contains the ceremony, the week before it is the reference. Film-weeks before a film’s first rating are left out: many nominees open in December, and a zero from before release would turn release timing into a pre-trend.

Estimator. Two-way fixed effects (film; ceremony × week, so each winner is compared only with its own year’s nominees in the same week), standard errors clustered by film. The average post-ceremony effect is the mean of the 26 post-week coefficients. Beside the clustered p-value is a within-ceremony relabelling benchmark: one nominee per ceremony is relabelled as the winner, 1,000 times, and the observed average is placed in that distribution. This benchmark describes how unusual the estimate is under uniform relabeling. Interpreting it as a randomization test would require an assignment assumption that this observational design does not establish. Placebo dates a year later and a year earlier, and a matched-pairs estimate (each winner against the nominee with the nearest pre-ceremony trajectory) sit beside it.

The plan, and where the analysis departed from it. The plan was written before any estimate was run and is fixed at the repository’s first commit, which holds it together with the code. Departures, all made after seeing data:

Result

Reading the estimates. Log points express changes on a logarithmic scale: +0.10 is roughly a 10% rise in the weekly count plus one, the quantity modelled here. SE means standard error, a measure of uncertainty; the 95% interval shows a range of effects compatible with the estimate under the model’s assumptions. A p-value measures how unusual the result would be under the stated no-effect model or relabelling benchmark. It is not the probability that winning caused a change.

Best Picture winners against the nominees they beat: weekly ratings-26-22-18-14-10-6-2+2+6+10+14+18+22+26-1.25-1.00-0.75-0.50-0.25+0.00+0.25+0.50+0.75+1.00weeks from the ceremony (week 0 contains it; −1 is the reference)winner − nominee, log pointsBest Picture winners against the nominees they beat: weekly ratingsweek -26: -0.38 (95% interval -1.24 to +0.47)week -25: -0.34 (95% interval -1.05 to +0.37)week -24: -0.33 (95% interval -1.08 to +0.43)week -23: +0.08 (95% interval -0.73 to +0.88)week -22: +0.15 (95% interval -0.74 to +1.04)week -21: +0.03 (95% interval -0.76 to +0.82)week -20: -0.49 (95% interval -1.38 to +0.39)week -19: -0.60 (95% interval -1.34 to +0.15)week -18: -0.63 (95% interval -1.35 to +0.09)week -17: +0.11 (95% interval -0.63 to +0.84)week -16: -0.12 (95% interval -0.79 to +0.55)week -15: +0.03 (95% interval -0.53 to +0.59)week -14: -0.43 (95% interval -1.00 to +0.14)week -13: -0.28 (95% interval -0.78 to +0.21)week -12: -0.34 (95% interval -0.77 to +0.09)week -11: -0.27 (95% interval -0.66 to +0.11)week -10: -0.36 (95% interval -0.68 to -0.04)week -9: -0.37 (95% interval -0.70 to -0.04)week -8: -0.20 (95% interval -0.48 to +0.09)week -7: -0.21 (95% interval -0.47 to +0.05)week -6: -0.17 (95% interval -0.40 to +0.06)week -5: -0.05 (95% interval -0.22 to +0.12)week -4: -0.12 (95% interval -0.25 to +0.02)week -3: -0.04 (95% interval -0.17 to +0.08)week -2: -0.08 (95% interval -0.18 to +0.03)week +0: -0.04 (95% interval -0.18 to +0.10)week +1: +0.17 (95% interval -0.00 to +0.34)week +2: +0.20 (95% interval +0.04 to +0.36)week +3: +0.07 (95% interval -0.12 to +0.27)week +4: +0.08 (95% interval -0.08 to +0.25)week +5: +0.15 (95% interval -0.01 to +0.31)week +6: +0.08 (95% interval -0.11 to +0.27)week +7: -0.00 (95% interval -0.26 to +0.25)week +8: +0.03 (95% interval -0.17 to +0.23)week +9: -0.01 (95% interval -0.21 to +0.18)week +10: +0.01 (95% interval -0.17 to +0.19)week +11: +0.06 (95% interval -0.15 to +0.27)week +12: -0.03 (95% interval -0.25 to +0.19)week +13: +0.04 (95% interval -0.16 to +0.24)week +14: -0.02 (95% interval -0.24 to +0.19)week +15: +0.01 (95% interval -0.22 to +0.23)week +16: -0.09 (95% interval -0.31 to +0.13)week +17: -0.05 (95% interval -0.25 to +0.15)week +18: +0.03 (95% interval -0.18 to +0.24)week +19: -0.04 (95% interval -0.29 to +0.21)week +20: -0.06 (95% interval -0.30 to +0.18)week +21: -0.01 (95% interval -0.19 to +0.16)week +22: +0.05 (95% interval -0.10 to +0.21)week +23: -0.02 (95% interval -0.21 to +0.18)week +24: +0.03 (95% interval -0.16 to +0.21)week +25: -0.04 (95% interval -0.22 to +0.14)week +26: -0.03 (95% interval -0.23 to +0.17)week −1: the reference week (fixed at 0)
Winner minus nominee, log weekly ratings, weeks −26 to +26 around the ceremony, with 95% intervals. Hover a week for its coefficient.

The pre-registered estimate does not establish a sustained effect. Over the 26 weeks after the ceremony, winners averaged +0.02 log points more weekly ratings than the nominees they beat (SE 0.08; cluster p = 0.78; relabelling benchmark p = 0.81). The 95% interval, −0.14 to +0.18, is compatible with a 13% loss and a 20% gain. The estimate is uninformative about a half-year effect, which is different from showing there is none.

There is a short rise. In the two weeks after the ceremony winners get +0.17 and +0.20 log points, about 20% more ratings than the losing nominees relative to the week before. Week 5 shows a smaller estimated lift. From week 7 onward there is no clear sustained increase. The four-week average is +0.13 (SE 0.08; relabelling p = 0.11). That contrast was not pre-registered and is reported as post hoc.

Matched pairs, two contrasts. Each winner against the nominee with the nearest pre-ceremony trajectory. The primary matched contrast, 26 post-weeks against weeks −4…−1: +0.13 (SE 0.07), winner above its match in 16 of 28 pairs (sign test p = 0.57). The like-for-like contrast with the post hoc four-week average, weeks +1…+4 against week −1: +0.15 (SE 0.08), winner above its match in 20 of 28 pairs (sign test p = 0.04). The short-window estimates are similar when their time windows and reference weeks match. An earlier draft incorrectly compared the 26-week matched estimate with the four-week event-study average.

Winners were already gaining before the ceremony. The pre-registered pre-trend test over weeks −26 to −2 rejects (p = 0.003). Winners were gaining on nominees in the months before the ceremony, which the figure shows as coefficients of −0.3 to −0.6 far from the event. Over the nearer window (−8 to −2) the test does not reject (p = 0.14). However, with 28 winners it may have limited power to detect a difference. This narrower window was chosen after seeing the earlier result. That gives no independent reason to assume the trends were parallel (Roth 2022). The estimate uses the week just before the ceremony as its baseline, by which point the winners had already gained relative attention. Since the analysis does not explain that earlier gain, it cannot confidently attribute the later changes to the award. Mean rating and the share of ratings ≥ 4 show no clear change (+0.02 stars; +0.001).

Placebos. A year after the ceremony, the same films show −0.02 (no detectable average post-placebo change). A year before, 75% of the films did not yet exist on the site, so there was insufficient history to run that comparison.

Heterogeneity (exploratory). Ceremonies after 2010: +0.08 (SE 0.09); before: −0.03 (0.14). The winner was the pre-ceremony volume leader among its nominees in only 6 of 28 years; in those six the post effect was +0.18 (0.15), in the other 22, −0.02 (0.10). With only six such ceremonies, the apparent difference is too uncertain to support a broader conclusion about whether prior popularity changes the effect of winning.

Relabelling benchmark: a random nominee relabelled as the winner within its ceremony, 1,000 timesRelabelling benchmark: a random nominee relabelled as the winner within its ceremony, 1,000 times0 draws with an average effect between -0.274 and -0.2612 draws with an average effect between -0.261 and -0.2481 draws with an average effect between -0.248 and -0.2353 draws with an average effect between -0.235 and -0.2222 draws with an average effect between -0.222 and -0.2093 draws with an average effect between -0.209 and -0.1965 draws with an average effect between -0.196 and -0.1838 draws with an average effect between -0.183 and -0.17110 draws with an average effect between -0.171 and -0.15811 draws with an average effect between -0.158 and -0.14518 draws with an average effect between -0.145 and -0.13220 draws with an average effect between -0.132 and -0.11919 draws with an average effect between -0.119 and -0.10636 draws with an average effect between -0.106 and -0.09339 draws with an average effect between -0.093 and -0.08041 draws with an average effect between -0.080 and -0.06740 draws with an average effect between -0.067 and -0.05458 draws with an average effect between -0.054 and -0.04161 draws with an average effect between -0.041 and -0.02959 draws with an average effect between -0.029 and -0.01658 draws with an average effect between -0.016 and -0.00350 draws with an average effect between -0.003 and +0.01050 draws with an average effect between +0.010 and +0.02376 draws with an average effect between +0.023 and +0.03652 draws with an average effect between +0.036 and +0.04952 draws with an average effect between +0.049 and +0.06233 draws with an average effect between +0.062 and +0.07538 draws with an average effect between +0.075 and +0.08844 draws with an average effect between +0.088 and +0.10132 draws with an average effect between +0.101 and +0.11315 draws with an average effect between +0.113 and +0.12620 draws with an average effect between +0.126 and +0.13916 draws with an average effect between +0.139 and +0.1527 draws with an average effect between +0.152 and +0.1658 draws with an average effect between +0.165 and +0.1786 draws with an average effect between +0.178 and +0.1915 draws with an average effect between +0.191 and +0.2040 draws with an average effect between +0.204 and +0.2172 draws with an average effect between +0.217 and +0.2300 draws with an average effect between +0.230 and +0.242the true winners: +0.023; 81% of relabellings are at least this far from zero-0.27-0.19-0.10-0.02+0.07+0.16+0.24average post-ceremony effect, log points, under a random winner
The 26-week average effect under 1,000 random relabellings of the winner within each ceremony; the observed value is the marked line. Hover a bar for its count.

What this does not show

MovieLens records when people rate a film on one site. It does not record box office revenue, streaming activity, or total viewership, and a rating after the ceremony may come from someone rewatching a film. The outcome is therefore rating activity among these users.

There are also two distinct limits to the result. With twenty-eight winners, the six-month estimate is too imprecise to rule out meaningful changes in either direction. Separately, winners were already gaining attention before the award, so even the brief rise after it could reflect influences the comparison does not account for. The study identifies a pattern consistent with an Oscar bump, but cannot establish how much of that pattern winning caused.

Appendix: forecasting the weekly count

The data can also address a prediction question: how well can next week’s rating count be forecast, and does knowing the ceremony date improve that forecast? This does not require treating winners and nominees as otherwise comparable. It requires testing forecasts against later observations that were excluded from training.

I use the 500 most-rated films, with 12 forecast starting points each across their last five years and predictions 1 to 8 weeks ahead, giving 6,000 starting points in all. The outcome is again log(1 + weekly ratings), so an error of 0.10 is roughly a 10% difference in count plus one, not a ten-percentage-point error. The boosting model is refit for each block of three origins on target weeks before the block’s earliest origin, so no training row lies after a forecast it is scored on.

Reading the forecast table. MAE is mean absolute error—the average size of the forecast’s miss on the log scale; lower is better. The horizon, h, is how many weeks ahead the forecast looks. Coverage is the share of actual outcomes inside the prediction band, and width is how broad that band is. “Native” bands come from the model itself; “conformal” bands are adjusted using errors on separate data.

model MAE h=1 h=4 h=8 mean 1–8 native 80% band: coverage · width conformal 80% band: coverage · width
naive (last week) 0.343 0.357 0.423 0.372 94% · 2.17 80% · 1.19
seasonal naive (52 weeks back) 0.403 0.439 0.470 0.427 92% · 2.17 80% · 1.33
damped Holt (grid-tuned per film) 0.260 0.263 0.343 0.292 94% · 1.72 80% · 0.94
gradient boosting (lags, seasonality, film age) 0.274 0.270 0.347 0.300 91% · 1.45 80% · 0.95
Forecast error by horizon (MAE, log points; lower is better)Forecast error by horizon (MAE, log points; lower is better)123456780.00.10.20.30.40.5naive, 1 week ahead: MAE 0.343naive, 2 weeks ahead: MAE 0.396naive, 3 weeks ahead: MAE 0.344naive, 4 weeks ahead: MAE 0.357naive, 5 weeks ahead: MAE 0.344naive, 6 weeks ahead: MAE 0.394naive, 7 weeks ahead: MAE 0.378naive, 8 weeks ahead: MAE 0.423naiveseasonal naive, 1 week ahead: MAE 0.403seasonal naive, 2 weeks ahead: MAE 0.385seasonal naive, 3 weeks ahead: MAE 0.445seasonal naive, 4 weeks ahead: MAE 0.439seasonal naive, 5 weeks ahead: MAE 0.431seasonal naive, 6 weeks ahead: MAE 0.435seasonal naive, 7 weeks ahead: MAE 0.412seasonal naive, 8 weeks ahead: MAE 0.470seasonal naivedamped Holt, 1 week ahead: MAE 0.260damped Holt, 2 weeks ahead: MAE 0.298damped Holt, 3 weeks ahead: MAE 0.279damped Holt, 4 weeks ahead: MAE 0.263damped Holt, 5 weeks ahead: MAE 0.266damped Holt, 6 weeks ahead: MAE 0.317damped Holt, 7 weeks ahead: MAE 0.309damped Holt, 8 weeks ahead: MAE 0.343damped Holtboosting, 1 week ahead: MAE 0.275boosting, 2 weeks ahead: MAE 0.312boosting, 3 weeks ahead: MAE 0.281boosting, 4 weeks ahead: MAE 0.271boosting, 5 weeks ahead: MAE 0.277boosting, 6 weeks ahead: MAE 0.333boosting, 7 weeks ahead: MAE 0.301boosting, 8 weeks ahead: MAE 0.345boostingweeks ahead
Mean absolute error by horizon for each model. Hover a point for the number.

Fitting each film’s history gives the lowest forecast error. Damped Holt, three parameters fit to each film’s own history, has the lowest error at every horizon. Boosting on lags and calendar features, trained across films, is close behind and no better. An earlier version of the boosting evaluation included training data from after the forecast starting points and reported an error of 0.286. Correcting that leakage increased its reported error by about 0.015 log points and changed which model performed best. The earlier comparison could not support the claimed advantage because the model had access to information unavailable at forecast time.

Two kinds of interval. Each model’s native 80% band is its own claim about its uncertainty. The random-walk bands cover 92–94% of outcomes, wider than needed for the nominal 80% target on this test. Boosting’s quantile band is the narrowest and covers 91%. The split-conformal band is calibrated on half of the films’ residuals at the same horizon and scored on the rest, and the halves are then swapped. Its coverage is measured rather than guaranteed, since films and origins are not exchangeable in the way the guarantee assumes (Angelopoulos and Bates). It lands at 80% for all four models, so width is the comparison. Holt’s band is 0.94 wide and boosting’s 0.95, against the naive model’s 1.19.

Seasonality does not help on its own. Repeating the count from 52 weeks earlier is worse than repeating last week at every horizon: this simple annual baseline does not capture the changes in rating activity well.

The ceremony features add little in this test. Two boosting models were compared, one with the features “weeks until this film’s ceremony” and “weeks since it” and one without. Both were trained on weeks before the 83rd ceremony (February 2011) and scored on the 114 nominees of the 83rd–95th ceremonies. Each forecast the ceremony week and the seven after it from the week before. Mean absolute error was 0.367 without the features and 0.363 with, a paired difference of −0.004 ± 0.003. The features helped 55 of 114 films and made the 13 winners slightly worse (0.342 to 0.349). Two nominees with under eight weeks of history were skipped and are named in the report. Knowing the ceremony date therefore adds little to this forecast under the tested configuration. That finding answers a different question from whether the award causes a change. Other forecast inputs may already capture related information, and a feature’s contribution depends on the model and comparison used.

Stack

Python, pandas, numpy, matplotlib, scikit-learn for the boosting baseline. The two-way fixed effects are absorbed by alternating demeaning and the cluster-robust covariance is computed directly, so the estimator has no dependencies beyond numpy and every number can be traced. The figures are hand-written SVG from the results JSON, inlined at build time, with a tooltip on every point. Data from GroupLens Research (MovieLens 32M).