A recommender you can steer
On this page
I wanted more control over the recommendations I was getting. A system can learn that I like certain films without knowing what I want tonight. “Scarier” should mean something different for a Pixar fan than for someone who watches horror every week. I built a movie recommender that starts from a viewer’s ratings and lets them change the mood with dials or a sentence.
The first build predicted ratings reasonably well but put Toy Story next to Seven Samurai and The Shining. That led me to change the model and check recognizable film neighbours before running the larger evaluation. The rebuilt system produced more sensible lists, and its dials increased the requested quality consistently in 94% of tested viewer–dial pairs.
It might seem that this is enough to show the recommender works as intended. However, increasing the requested quality is only part of the question. A simpler way of reordering recommendations preserved more of a viewer’s original taste for the same change in mood. The dials can therefore work without being the better way to answer the request, and whether people prefer their results remains a question for a user study.
I wrote the specification and built the system. Try the recommender (opens in a new tab) or see the result tables.
The project now also has a book demo (opens in a new tab), with a separate model and evaluation. This write-up describes the movie version; the shared project page explains the book extension and its results.
Technical methods and evaluation
Jump to the first build, what steering measures, or the comparison after review.
The idea
For a request to start from someone’s taste, the system first needs a way to represent that taste. Collaborative filtering provides one by learning which films appeal to similar viewers from patterns in their ratings. This model factorizes the rating matrix into a vector—a list of learned numbers—for each viewer and film. Those vectors act as positions, and nearby films become recommendation candidates.
A position alone still does not say how to make a list more melancholy. For that, I need a direction associated with the quality. The Tag Genome supplies scores for how strongly films match that description. I label the highest-scoring films as positives and the lowest as negatives, then fit a logistic regression to separate their vectors. The normal to the separating plane gives a direction to move along. This is a concept activation vector, following work by Kim and colleagues on interpreting neural networks.
With a direction in hand, I can add some multiple of it to the viewer’s position and rank the catalog around the new point. The unit for a step is the catalog’s own spread along that direction, so the dials use a common rule for their scale. A genre filter can restrict the available films; this operation changes the position from which they are ranked.
There is a further question, though. A direction that separates high-scoring films from low-scoring films does not necessarily produce better recommendations when a viewer moves along it. The returned films also depend on the viewer’s starting position and the rest of the ranking procedure. I therefore test what changes in the actual lists, rather than treating the fitted direction as sufficient evidence that steering works.
The first build
I wrote a spec, and the first pipeline run came back with respectable numbers. Held-out error of 0.862 for the factorization, every one of the 1,128 genome tags encoded with a cross-validated AUC well above chance, dials selected. Then I read the lists. The Pixar household’s unsteered top ten was Sully, The Hurt Locker, and Zodiac. Steering did nothing to it. The test I had written into the spec as the kill test, two viewers given one move on the same dial, failed outright.
This failure does not mean the error metric was calculated incorrectly. It means that predicting ratings and producing useful neighbours are different tests, and a model can do reasonably well on the first while failing the second. The factorization was a nonnegative one without bias terms, and it had spent nearly all of its capacity on a single direction that says how much a film gets rated at all. The remaining directions were close to random. Toy Story’s nearest neighbours in that space were Seven Samurai and The Shining. Predicting a held-out rating is possible with a good bias model and a poor geometry, and RMSE cannot tell the two apart. What told them apart was printing nearest neighbours for five films, which the pipeline now does before any evaluation runs.
The fix had three parts, and each is logged with its reason. Biased alternating least squares replaced the nonnegative factorization, so that the “how much is this film rated” signal lives in the bias terms and the factors are free to describe taste. The item space is centered, so the average film sits at the origin and every vector is a lean away from it; directions are fit and steering happens in that centered space. And a viewer is placed at the centroid of the films they rated, weighted by rating minus three, so that a five pulls twice as hard as a four and a two pushes away. Toy Story’s neighbours became Toy Story 2, Toy Story 3, A Bug’s Life, and The Incredibles. Notting Hill’s became Sleepless in Seattle and While You Were Sleeping. The lists were readable.
The direction also needs to point somewhere the catalog actually varies. Fitting the regression on standardized coordinates and mapping the weights back, which is the textbook move, produced directions that aimed along the lowest-variance dimensions of the space, where the catalog barely spreads at all; fitting on the raw centered coordinates, with the regularization strength chosen per tag by the one-standard-error rule, gave directions the catalog actually varies along.
The same distinction matters when choosing which dials to show. Selecting by AUC alone from all 1,128 tags gave “predictable”, “oscar winner”, “pg-13”, “cartoon”, and “nudity (topless)”. These labels were well represented in the model, but that did not make them a useful set of mood controls. I restricted the selection to a pool of 73 mood and content tags. The evaluation can tell me how well a direction separates films; choosing which qualities a viewer should be able to adjust also requires a design decision, recorded in the log.
What steering measures
The test asks whether turning a dial further keeps increasing the requested quality. This is monotonicity. For 2,000 viewers placed from their own ratings, each dial was turned to −2, −1, 0, +1, and +2 steps, and the top twenty’s mean genome score for that tag was recorded at each. The five means rise in order for 94% of viewer–dial pairs. The melancholy dial takes a typical viewer’s list from a mean score of 0.05 to 0.59; stylish goes 0.10 to 0.65; crude humor 0.18 to 0.48. The dials that are strongest in the geometry are the ones that separate an arthouse audience from a mainstream one, and several tags name that same axis from different ends. “Predictable” and “criterion” both reach an AUC of 0.99, and their directions sit at cosine −0.97.
The Sixth Sense test is the one I built the thing to pass. Two synthetic viewers, one from seventeen five-star ratings on Pixar and Disney films and one from seventeen on horror, get the identical move on the dial nearest “scary”, which turns out to be eerie. The comparison asks whether the same move can retain a distinction between the two tastes. The dial gives the household The Iron Giant, Psycho, Coraline, Pan’s Labyrinth, Who Framed Roger Rabbit, and 12 Angry Men, and gives the devotee The Omen, The Shining, Carrie, Ringu, and Alien. The two lists share nothing. The household’s mean scary score rises from 0.17 to 0.44 and the devotee’s from 0.80 to 0.82, which is about right for someone already at the ceiling.
What the geometry says about “scarier”
The two lists support the idea that the same request can preserve a distinction between viewers. However, they required a four-step move, while the spec had called for one. At one step the household’s scary score barely changes (0.17 to 0.19). To understand how well the control answers the original request, I needed to account for that difference.
The two synthetic viewers represent unusually concentrated tastes. Seventeen five-star ratings on one cluster give a position with twice the radius of a typical viewer’s, and a step of fixed size is a smaller turn for a longer vector. That part was fixable, and is fixed. Every viewer’s position is scaled to the median radius of the training users before steering, so a step means the same thing for a concentrated taste as for a diffuse one. It was not enough on its own.
The rest is the direction itself. The global “scary” direction separates people who rate horror from people who do not. It is a taste-community axis, and Coraline is not on it. Coraline sits 48° from the household’s centroid with a small lean along scary; A Bug’s Life sits at 21° with almost none. No single step along “scary” short of leaving animation altogether puts Coraline above A Bug’s Life. A direction fit only among the household’s 600 nearest films did no better. What surfaces Coraline is the eerie dial, a different direction that happens to run through the atmospheric end of the family cluster. My reading is that a global concept direction is the right tool for “more of this quality, wherever it is” and the wrong tool for “more of this quality, near me”. The spec already had the second thing deferred, under the name cluster-local directions, and this is the result that argues for building it.
What was already known
None of the core idea is new. Göpfert and colleagues at Google published a framework in 2022, extended in 2024, that fits concept activation vectors for MovieLens tags on the item embeddings of a rating-only recommender, tests whether the recommender has learned each attribute, and critiques a user by adding a step along the vector to their embedding. That is the same update rule as a dial here, and their finding that a pairwise ranking loss beats logistic regression for fitting the vectors is one I should try, since the Tag Genome’s scores are continuous and my top-versus-bottom quartile labels throw away the middle half. Where they went further is subjectivity: a tag can mean a different threshold to different people, or a different thing entirely, and they recover several vectors per tag by clustering users. Where this build went instead is toward the step unit as a formula, the centered space and the geometry checks, and an evaluation of steering as movement. Their notion of a tag having several senses is a second candidate answer to the Coraline problem, alongside the cluster-local directions above.
After a review
A reviewer raised a further objection to the first version: showing that a dial increases its own quality leaves open what happens to the viewer’s other preferences. This is especially relevant here, since preserving a connection to the viewer’s taste was the reason for starting the move from their position. I added four evaluations to examine that objection, and one went against my expectation. Two hundred pairs of dissimilar viewers got the same +2σ on each dial through three controllers: steering, a top-quartile tag filter, and a user-independent re-rank that adds a weight times the tag score to the unsteered ranking, with the weight set so that its mean response matched steering’s. At equal response, the re-rank keeps the viewer nearer their own taste than steering does. The mean cosine between the returned films and the viewer’s unsteered position is 0.81 before any move; under the re-rank it is 0.72 to 0.79, under steering 0.50 to 0.73. The re-rank also keeps more of the viewer’s original list. I had assumed that moving the position would preserve personalization better than adding a content score to the ranking, because the move starts from the viewer. The comparison shows why that reasoning was insufficient: both methods can retain a dependence on the viewer, and the method that moves their position can still take them further from their original taste. On this build, the simpler re-rank does better on that comparison.
The remaining tests help explain what else a move changes. The cross-talk matrix, which records what every other tag does when one dial moves, has its largest entry on the diagonal in six of eight rows, the exceptions being the two weakest dials, eerie and bizarre, which move tense and dark more than themselves; the off-diagonal entries read as the space’s own correlations (stylish drags violent and dark, heartwarming drags touching and sentimental). Held-out relevance falls to about half by two steps. A stronger move therefore comes with a substantial loss on the relevance measure, even when it increases the requested quality. Response also depends on the starting position: viewers who already sit high on a concept move less than viewers who sit low. These results help describe the tradeoff a viewer is making, but a user study is still needed to judge whether that tradeoff is worthwhile.
Constants
A rule I work under is that a constant is either the output of a formula or a decision written down with its reason. The number of factors is chosen by the one-standard-error rule on held-out error over three seeds. A tag is encoded when its AUC minus two standard errors clears 0.5, which compares its discriminative performance with chance. The regularization on each direction is chosen by the same one-standard-error rule. The shrinkage that keeps thinly rated films from crowding the list is n/(n + n₀) with n₀ the median rating count, and the radius every viewer is scaled to is the median over training users. The dial pool, the eight-dial limit, and the half-orthogonality guard between dials are decisions, and the log says so. The one-standard-error rule also chose 256 factors, the largest in its grid, on differences in the fourth decimal; the grid should be extended before that number is final, and the report says that too.
What runs in the browser
The ranking calculations run locally after the offline fit. The page loads a 13.5 MB float array (13,174 films by 256 factors, centered) and a 3 MB bundle with titles, genres, the eight directions, and the presets. Placement, steering, cosine scoring, and the diversity re-rank are a few loops of JavaScript, and a re-rank over the whole catalog takes about 25 milliseconds. The demo needs no account, and your ratings stay in the browser. The typed-request interpreter and surprise-seat checks use a server; the page describes what they receive. You can start from one of five presets, the two Sixth Sense viewers and three built from clusters of real users, or rate some of twenty anchor films chosen to spread across the space.
Data from GroupLens Research, MovieLens 25M and the Tag Genome. The spec, its amendments, and the numbered decision log are in the repository.