The Oscar bump
This project asks whether winning Best Picture changes a film’s rating activity. I wrote the analysis plan before running the estimates. The difference-in-differences event study uses timestamped MovieLens ratings to compare winners with the other nominees from the same ceremony.
In a businessMarketing teams need to separate a campaign's effect from trends already underway. This project uses a comparison group, an event study, and forecast backtests, and it documents a leakage mistake I found and corrected.
Technical skills: Python · pandas / NumPy · scikit-learn · Event studies · Uncertainty estimation · Forecast backtesting
- My contribution
- I set the design and the pre-registered plan, built the pipeline, spot-checked the events table against the sources, and ran the estimates.
- Result and limit
- Winners show a brief rise in rating activity (about 20% over two weeks, a comparison added after seeing the data). The planned six-month estimate is too imprecise to establish a sustained effect: its uncertainty interval allows roughly a 13% decrease to a 20% increase. Winners were already gaining before the ceremony, limiting what this study can attribute to the award.
On this page
Technical methods and evaluation
Jump to the design, results, or forecasting.
Design
Treatment and control. The Best Picture winner of each ceremony against the other nominees of the same ceremony, week by week: 28 ceremonies (1996–2023), 196 films, 28 winners. The films share a category, nomination year, and ceremony. They can still differ in release timing, audience, and momentum before the award. Winning is not randomly assigned among nominees. If winners were already gaining attention before the award, then a later difference could reflect that earlier momentum. Checking the trends before the ceremony is therefore necessary before attributing a change to winning.
Outcome and window. log(1 + weekly ratings), and mean weekly rating, from 26 weeks before the ceremony to 26 after; week 0 contains the ceremony, the week before it is the reference. Film-weeks before a film’s first rating are left out: many nominees open in December, and a zero from before release would turn release timing into a pre-trend.
Estimator. Two-way fixed effects (film; ceremony × week, so each winner is compared only with its own year’s nominees in the same week), standard errors clustered by film. The average post-ceremony effect is the mean of the 26 post-week coefficients. Beside the clustered p-value is a within-ceremony relabelling benchmark: one nominee per ceremony is relabelled as the winner, 1,000 times, and the observed average is placed in that distribution. This benchmark describes how unusual the estimate is under uniform relabeling. Interpreting it as a randomization test would require an assignment assumption that this observational design does not establish. Placebo dates a year later and a year earlier, and a matched-pairs estimate (each winner against the nominee with the nearest pre-ceremony trajectory) sit beside it.
The plan, and where the analysis departed from it. The plan was written before any estimate was run and is fixed at the repository’s first commit, which holds it together with the code. Departures, all made after seeing data:
- Film-weeks before a film’s first rating are dropped. The plan treated every empty week as a zero; the first run showed release timing as a spurious pre-trend, and the rule changed.
- Matching uses the root-mean-square distance over the pre-weeks both films have (at least 4) rather than Euclidean distance over all 26, because most nominees do not have 26 weeks of history at the ceremony. The primary matched contrast averages 26 post-weeks against weeks −4…−1; a second contrast, weeks +1…+4 against week −1, was added after review so the matched estimate could be compared with the event study’s four-week average like for like.
- The four-week and two-week post-ceremony averages and the −8…−2 pre-trend test are post hoc and labelled so.
- The year-earlier placebo could not be run as planned (75% of the films had no ratings yet); a year-later placebo was added.
- Not done: the plan’s secondary categories (Director, Actor, Actress) and the nomination-announcement event.
- Forecasting: exponential smoothing became damped Holt; the boosting backtest was rebuilt after review because its first version trained on weeks after the forecast origins.
Result
Reading the estimates. Log points express changes on a logarithmic scale: +0.10 is roughly a 10% rise in the weekly count plus one, the quantity modelled here. SE means standard error, a measure of uncertainty; the 95% interval shows a range of effects compatible with the estimate under the model’s assumptions. A p-value measures how unusual the result would be under the stated no-effect model or relabelling benchmark. It is not the probability that winning caused a change.
The pre-registered estimate does not establish a sustained effect. Over the 26 weeks after the ceremony, winners averaged +0.02 log points more weekly ratings than the nominees they beat (SE 0.08; cluster p = 0.78; relabelling benchmark p = 0.81). The 95% interval, −0.14 to +0.18, is compatible with a 13% loss and a 20% gain. The estimate is uninformative about a half-year effect, which is different from showing there is none.
There is a short rise. In the two weeks after the ceremony winners get +0.17 and +0.20 log points, about 20% more ratings than the losing nominees relative to the week before. Week 5 shows a smaller estimated lift. From week 7 onward there is no clear sustained increase. The four-week average is +0.13 (SE 0.08; relabelling p = 0.11). That contrast was not pre-registered and is reported as post hoc.
Matched pairs, two contrasts. Each winner against the nominee with the nearest pre-ceremony trajectory. The primary matched contrast, 26 post-weeks against weeks −4…−1: +0.13 (SE 0.07), winner above its match in 16 of 28 pairs (sign test p = 0.57). The like-for-like contrast with the post hoc four-week average, weeks +1…+4 against week −1: +0.15 (SE 0.08), winner above its match in 20 of 28 pairs (sign test p = 0.04). The short-window estimates are similar when their time windows and reference weeks match. An earlier draft incorrectly compared the 26-week matched estimate with the four-week event-study average.
Winners were already gaining before the ceremony. The pre-registered pre-trend test over weeks −26 to −2 rejects (p = 0.003). Winners were gaining on nominees in the months before the ceremony, which the figure shows as coefficients of −0.3 to −0.6 far from the event. Over the nearer window (−8 to −2) the test does not reject (p = 0.14). However, with 28 winners it may have limited power to detect a difference. This narrower window was chosen after seeing the earlier result. That gives no independent reason to assume the trends were parallel (Roth 2022). The estimate uses the week just before the ceremony as its baseline, by which point the winners had already gained relative attention. Since the analysis does not explain that earlier gain, it cannot confidently attribute the later changes to the award. Mean rating and the share of ratings ≥ 4 show no clear change (+0.02 stars; +0.001).
Placebos. A year after the ceremony, the same films show −0.02 (no detectable average post-placebo change). A year before, 75% of the films did not yet exist on the site, so there was insufficient history to run that comparison.
Heterogeneity (exploratory). Ceremonies after 2010: +0.08 (SE 0.09); before: −0.03 (0.14). The winner was the pre-ceremony volume leader among its nominees in only 6 of 28 years; in those six the post effect was +0.18 (0.15), in the other 22, −0.02 (0.10). With only six such ceremonies, the apparent difference is too uncertain to support a broader conclusion about whether prior popularity changes the effect of winning.
What this does not show
MovieLens records when people rate a film on one site. It does not record box office revenue, streaming activity, or total viewership, and a rating after the ceremony may come from someone rewatching a film. The outcome is therefore rating activity among these users.
There are also two distinct limits to the result. With twenty-eight winners, the six-month estimate is too imprecise to rule out meaningful changes in either direction. Separately, winners were already gaining attention before the award, so even the brief rise after it could reflect influences the comparison does not account for. The study identifies a pattern consistent with an Oscar bump, but cannot establish how much of that pattern winning caused.
Appendix: forecasting the weekly count
The data can also address a prediction question: how well can next week’s rating count be forecast, and does knowing the ceremony date improve that forecast? This does not require treating winners and nominees as otherwise comparable. It requires testing forecasts against later observations that were excluded from training.
I use the 500 most-rated films, with 12 forecast starting points each across their last five years and predictions 1 to 8 weeks ahead, giving 6,000 starting points in all. The outcome is again log(1 + weekly ratings), so an error of 0.10 is roughly a 10% difference in count plus one, not a ten-percentage-point error. The boosting model is refit for each block of three origins on target weeks before the block’s earliest origin, so no training row lies after a forecast it is scored on.
Reading the forecast table. MAE is mean absolute error—the average size of the forecast’s miss on the log scale; lower is better. The horizon, h, is how many weeks ahead the forecast looks. Coverage is the share of actual outcomes inside the prediction band, and width is how broad that band is. “Native” bands come from the model itself; “conformal” bands are adjusted using errors on separate data.
| model | MAE h=1 | h=4 | h=8 | mean 1–8 | native 80% band: coverage · width | conformal 80% band: coverage · width |
|---|---|---|---|---|---|---|
| naive (last week) | 0.343 | 0.357 | 0.423 | 0.372 | 94% · 2.17 | 80% · 1.19 |
| seasonal naive (52 weeks back) | 0.403 | 0.439 | 0.470 | 0.427 | 92% · 2.17 | 80% · 1.33 |
| damped Holt (grid-tuned per film) | 0.260 | 0.263 | 0.343 | 0.292 | 94% · 1.72 | 80% · 0.94 |
| gradient boosting (lags, seasonality, film age) | 0.274 | 0.270 | 0.347 | 0.300 | 91% · 1.45 | 80% · 0.95 |
Fitting each film’s history gives the lowest forecast error. Damped Holt, three parameters fit to each film’s own history, has the lowest error at every horizon. Boosting on lags and calendar features, trained across films, is close behind and no better. An earlier version of the boosting evaluation included training data from after the forecast starting points and reported an error of 0.286. Correcting that leakage increased its reported error by about 0.015 log points and changed which model performed best. The earlier comparison could not support the claimed advantage because the model had access to information unavailable at forecast time.
Two kinds of interval. Each model’s native 80% band is its own claim about its uncertainty. The random-walk bands cover 92–94% of outcomes, wider than needed for the nominal 80% target on this test. Boosting’s quantile band is the narrowest and covers 91%. The split-conformal band is calibrated on half of the films’ residuals at the same horizon and scored on the rest, and the halves are then swapped. Its coverage is measured rather than guaranteed, since films and origins are not exchangeable in the way the guarantee assumes (Angelopoulos and Bates). It lands at 80% for all four models, so width is the comparison. Holt’s band is 0.94 wide and boosting’s 0.95, against the naive model’s 1.19.
Seasonality does not help on its own. Repeating the count from 52 weeks earlier is worse than repeating last week at every horizon: this simple annual baseline does not capture the changes in rating activity well.
The ceremony features add little in this test. Two boosting models were compared, one with the features “weeks until this film’s ceremony” and “weeks since it” and one without. Both were trained on weeks before the 83rd ceremony (February 2011) and scored on the 114 nominees of the 83rd–95th ceremonies. Each forecast the ceremony week and the seven after it from the week before. Mean absolute error was 0.367 without the features and 0.363 with, a paired difference of −0.004 ± 0.003. The features helped 55 of 114 films and made the 13 winners slightly worse (0.342 to 0.349). Two nominees with under eight weeks of history were skipped and are named in the report. Knowing the ceremony date therefore adds little to this forecast under the tested configuration. That finding answers a different question from whether the award causes a change. Other forecast inputs may already capture related information, and a feature’s contribution depends on the model and comparison used.
Stack
Python, pandas, numpy, matplotlib, scikit-learn for the boosting baseline. The two-way fixed effects are absorbed by alternating demeaning and the cluster-robust covariance is computed directly, so the estimator has no dependencies beyond numpy and every number can be traced. The figures are hand-written SVG from the results JSON, inlined at build time, with a tooltip on every point. Data from GroupLens Research (MovieLens 32M).