Steerable Recommender
Two catalogs, separate models and evaluations. Dials run in your browser; typed requests use a server.
I wanted a recommender that changes the mood of a list while starting from the viewer’s or reader’s existing taste. The movie and book demos share that interaction: place yourself with ratings, then move a dial, type a request, or name a title. Each catalog uses its own data, learned taste space, and evaluation.
In a businessProduct teams need recommendations that follow both a customer's history and what they want right now. I measure whether a requested change takes effect and what relevance it costs against simpler re-ranking, so a team can pick an approach before running an A/B test.
Technical skills: Python · NumPy / SciPy · scikit-learn · Collaborative filtering · Ranking evaluation · JavaScript
- My contribution
- I wrote the specification, built the system, and reviewed the results. The first build’s poor recommendation lists led to replacing the factorization despite acceptable prediction error.
- Result and limit
- Movie evaluation: The dials increase the requested quality consistently in 94.3% of tested viewer–dial pairs. For the same change in mood, a simpler re-rank preserves more of the viewer’s original taste. Typed requests move in every expected direction in 45/47 author-written test phrasings, and toward the film they name in 11/11. The unexpected pick shows no statistically significant advantage over an equally distant, equally popular film on withheld ratings (0.9% vs 0.6%, p = 0.61). These tests do not establish user preference.
On this page
Movies and books
Try movies (opens in a new tab) · Try books (opens in a new tab)
The movie version learns from MovieLens ratings. The book version uses Goodreads data from the Tag Genome for Books. They share the placement and steering interface, but the movie model fits observed star ratings while the book model learns from which books readers rated. A dial is learned separately in each space; a result on one catalog is not evidence for the other.
Books
The book demo covers 9,374 titles in a 256-dimensional taste space, with 719 fitted tag directions and eight featured dials (love, magic, suspense, fun, coming of age, literary, erotic, philosophical). It uses implicit-feedback alternating least squares on a read-or-not matrix, treating the presence of a rating as evidence of interest. On Goodreads the stars carry little taste once the biases are removed, and which books a reader rated carries a lot. The ratings you enter in the demo still place your taste using a signed, rating-weighted average of book vectors.
Held-out recall@20 rose with every factor count tried, 0.168, 0.198, 0.209 and 0.2319 at 32, 64, 128 and 256 factors, so the top 20 recovered about 23% of the withheld liked books under this evaluation. The selected count is the largest tested and the curve had not flattened, where the movie model’s had by 64, so the search has not established whether a larger model would improve it. In the book evaluation, the requested tag score rose monotonically across −2 to +2 standardized steps in 78.9% of tested reader–dial pairs, across 2,000 readers (94.3% for the movie dials).
Three results run the other way from the movies. A two-step move costs the reader nothing measurable. Recall@20 stays between 0.209 and 0.219 across −2 to +2σ, where the movie lists lost half their relevance. In the shared-intent test, steering keeps 60–66% of the reader’s own list against the re-rank’s 68–80%, so the movie finding holds by a much smaller margin. And the model’s own user factor places a reader better than the rating-weighted centroid, recall@20 0.250 against 0.214, the reverse of the movie ordering. The demo keeps the centroid, since a visitor with a handful of ratings has no user factor and a fold-in scored worse than either. These measures do not establish that readers prefer the steered lists.
The book evaluation report contains the factorization, dial, and preset comparisons. The movie results and user study below apply to movies.
Movie methods and evaluation
Jump to the measurements, the comparison after review, or the user study.
A few terms used below. Collaborative filtering learns from patterns in people’s ratings: films rated similarly get related numerical representations, or vectors. The Tag Genome is a separate set of scores describing how strongly each film matches labels such as “eerie.” A step of 1σ is one standard deviation of the catalog along a direction, not a percentage change. AUC measures how well a tag detector separates high-scoring films from low-scoring ones; 0.5 is chance and 1 is perfect separation. Recall@20 is the share of a viewer’s withheld liked films recovered in the first 20 recommendations.
Data and factorization
I first learn film representations from ratings, then test prediction on ratings excluded from training. The data is MovieLens 25M restricted to the 13,174 genome-scored films with at least 50 ratings, and 50,000 users with at least 20 ratings on that catalog (6.84M training ratings, 737k held-out ratings, split within users). The rating matrix is factorized by biased alternating least squares over the observed entries only, r ≈ μ + b_u + b_i + w·h, 12 epochs, ridge 0.1 per rating.
Here, k is the number of learned dimensions. RMSE measures rating-prediction error in rating units, giving larger mistakes more weight; lower is better. “Held-out” ratings were excluded from training, and “3 seeds” means three runs with different random starting values.
| k | held-out RMSE (3 seeds) |
|---|---|
| 32 | 0.8000 |
| 64 | 0.7997 |
| 128 | 0.7996 |
| 256 | 0.7995 |
The one-standard-error rule chose k = 256, which is the edge of the grid; the curve is flat to the fourth decimal from k = 64 on. The item space is centered so the average film sits at the origin. Nearest neighbours by cosine serve as a geometry check. Toy Story’s are Toy Story 2 (0.98), Toy Story 3, A Bug’s Life and The Incredibles. The Shining’s are Full Metal Jacket, Apocalypse Now and A Clockwork Orange. Notting Hill’s are Sleepless in Seattle, While You Were Sleeping and The American President. The first build, masked nonnegative factorization without bias terms, reached RMSE 0.862 and put Toy Story next to Seven Samurai; RMSE said nothing about it.
Directions
To make a named dial, I look for a direction that separates films scoring high on that quality from films scoring low. For each of the 1,128 genome tags, a class-balanced logistic regression on the centered item coordinates separates the tag’s top-quartile films from its bottom quartile. The normalized coefficient vector is the tag’s direction, the linear separator the probe found rather than a gradient of the concept itself; whether moving along it raises the concept in retrieved lists is what E3 below tests. The catalog’s standard deviation along it is the unit for a step. A tag counts as encoded when its five-fold AUC minus two standard errors is above 0.5. All 1,128 are. The strongest are one arthouse-versus-mainstream axis under several names: predictable 0.99, criterion 0.99, bad plot 0.98, melancholy 0.99. Predictable and criterion sit at cosine −0.97. For those tags the regression’s direction agrees with the difference of class means at cosine 0.96.
Dials are chosen greedily by AUC from a pool of 73 mood and content tags, with a guard that keeps every dial at |cos| < 0.5 to every other. The eight are melancholy (0.99), eerie (0.93), irreverent (0.93), campy (0.93), crude humor (0.93), heartwarming (0.92), stylish (0.91) and bizarre (0.91). The pool is a logged decision. Chosen from all tags by AUC alone, the first build’s dials were predictable, oscar winner, pg-13, cartoon, and nudity (topless).
Placement and retrieval
A viewer’s position is the centroid of the centered vectors of the films they rated, weighted by rating minus three, so a 5 pulls twice as hard as a 4 and a 2 pushes away. The position is scaled to the median radius of the training users (0.342) before steering, so a step means the same thing for a concentrated taste as for a diffuse one. A dial at value v adds v·σ along its direction. The catalog is ranked by cosine to the moved position, multiplied by n/(n + 321) with 321 the median rating count, so thinly rated films with noisy vectors do not crowd the list. The same arithmetic runs in the browser over a 13.5 MB float array; a re-rank takes about 25 ms.
Measured
These results are from the movie model and MovieLens evaluation.
Steerability. For 2,000 training users placed from their own ratings, each dial was turned to −2, −1, 0, +1, +2 σ and the top-20’s mean genome score for that tag was recorded. The five means are non-decreasing in 94.3% of viewer–dial pairs.
| dial | monotonic | mean score at −2 … +2 σ |
|---|---|---|
| melancholy | 0.865 | 0.054, 0.077, 0.365, 0.537, 0.590 |
| eerie | 0.949 | 0.088, 0.126, 0.177, 0.237, 0.303 |
| irreverent | 0.968 | 0.044, 0.077, 0.148, 0.268, 0.412 |
| campy | 0.880 | 0.039, 0.046, 0.066, 0.113, 0.205 |
| crude humor | 0.981 | 0.176, 0.209, 0.258, 0.341, 0.483 |
| heartwarming | 0.962 | 0.079, 0.149, 0.278, 0.434, 0.554 |
| stylish | 0.973 | 0.102, 0.196, 0.352, 0.516, 0.651 |
| bizarre | 0.970 | 0.033, 0.049, 0.094, 0.168, 0.248 |
The Sixth Sense test. Two synthetic viewers, one built from seventeen five-star ratings on Pixar and Disney films and one from seventeen on horror, get the identical move on the dial nearest “scary”, which is eerie. Their positions sit at cosine −0.13. A genre filter alone restricts eligibility; this test asks how a change in mood interacts with each starting taste. At +4σ the two lists share no titles. The household’s list is The Iron Giant, Psycho, Coraline, Pan’s Labyrinth, A Hard Day’s Night, Who Framed Roger Rabbit, A Christmas Story, 12 Angry Men, L.A. Confidential, and The Adventures of Tintin, with the genome’s scary score rising from 0.17 to 0.44. The devotee’s is The Omen, The Shining, Carrie, Ringu, Alien, An American Werewolf in London, Night of the Living Dead, 28 Days Later, The Fly, and Let the Right One In, at 0.80 → 0.82. At +1σ, the step the spec named, the household barely moves (0.17 → 0.19); these two presets represent unusually concentrated tastes, and that finding is written up below.
Beyond accuracy. Over 200 users and ±1σ on every dial, cosine retrieval covers 17% of the catalog against 9% for dot product, with lower popularity bias. The diversity re-rank (λ = 0.8 over dial scores and genre flags) raises intra-list distance from 0.227 to 0.246 at no cost in coverage. Serendipity against each user’s held-out liked films is small everywhere (0.007 per slot) and does not separate the settings.
After review: controllability
An outside review of the first numbers asked what a dial does besides move its own concept. Four more evaluations answer that, over 400 of the same users, and two of them cut against the thesis.
The shared-intent test. 200 pairs of users at cosine ≤ 0 got the same +2σ on each dial, through three controllers. The first is steering. The second is an absolute filter that keeps the tag’s top quartile. The third is a user-independent re-rank that adds λ × tag score to the unsteered score, with λ set per dial by bisection so its mean response equals steering’s. At equal response, the re-rank keeps more of each user’s own unsteered list (23–60% against steering’s 9–17%) and stays nearer the user’s own direction. The mean cosine between the returned films and the viewer’s unsteered position is 0.81 before any move, 0.72–0.79 under the re-rank, and 0.50–0.73 under steering. The reviewer’s question was whether steering preserves personalization better than a re-rank that ignores the geometry. On this build, at equal response, it does not. The two users’ lists never overlap under any controller (Jaccard ≤ 0.005) with one exception: at +2σ on melancholy, the strongest axis in the space, steering pulls dissimilar users toward the same corner (0.125).
Cross-talk. The slope per σ of all 73 pool tags when each dial moves. In six of eight rows the dial’s own tag is the largest entry. The two exceptions are the weakest dials, eerie (own slope +0.05) and bizarre (+0.05), which move tense (+0.11) and dark (+0.13) more than themselves. The off-diagonal structure reads as the space’s own correlations. Stylish drags violent (+0.13) and dark (+0.13); heartwarming drags touching (+0.12) and sentimental (+0.11) and pushes stylish down (−0.08); crude humor drags funny (+0.12) and pushes tense (−0.10) and bleak (−0.08) down.
Relevance retention and saturation. Recall@20 of held-out films rated ≥ 4 falls from 0.122 unsteered to 0.091–0.095 at ±1σ and 0.054–0.072 at ±2σ. A two-step move costs about half the list’s relevance. Response depends on where the user starts. Users in the lowest third of unsteered melancholy score gain +0.37 at +2σ, the highest third +0.09; for stylish +0.38 against +0.19; for eerie the three thirds are nearly equal (+0.15, +0.12, +0.11).
Placement. Under the page’s own retrieval, the signed centroid gives recall@20 of 0.122, the ALS user factor 0.032, and a ridge fold-in with the item factors fixed 0.024. The model’s own rating prediction gives 0.043. The most-rated unseen films give 0.185. Steerability is the same for all three placements (monotonic 0.92–0.95, slope +0.090). The recall metric is dominated by popularity. About placements it says only that the factor the model learned is a poor position under cosine and the centroid is a usable one. It says more about the model, whose own top-20 by predicted rating is beaten four to one by a popularity list. That is an old result about RMSE-trained factorizations.
Two things a film can be. Because placement is a rating-weighted centroid, one five-star rating puts a viewer exactly at that film. So the page lets anyone start from a movie or two they love, searched by title, and watch the list move as they rate. A rating is where the viewer lives: it saves with the profile and is still there tomorrow. A mood is different, and the page keeps it separate. “Vibes” on any film, in the search box or on a result, adds a lean toward that film as a dial in σ units, next to the tag dials. It can be pushed further or negative, and reset wipes it. Like every dial it is kept for twelve hours in the browser and then cleared, so tonight’s mood does not become tomorrow’s. “I’m feeling Pulp Fiction vibes” typed into the word box does the same thing. The film is a fixed point in the space; the dial is how far toward it the viewer goes tonight, and the tag dials explore around it from there.
What the geometry says about “scarier”
The global scary direction is a taste-community axis. It separates people who rate horror from everyone else. Coraline sits 48° from the Pixar household’s centroid with a lean of 0.16 along that axis; A Bug’s Life sits at 21° with 0.06. No single linear step short of leaving animation puts Coraline above A Bug’s Life along “scary”, which is why the eerie dial, a different direction, is what surfaces it. A direction fit only among the household’s 600 nearest films did no better.
Two more attempts at a local direction followed, and both are in the repository as results. Partialling genre out of both the tag relevance and the item coordinates, and refitting every probe on the residuals, leaves all 1,128 tags encoded and the dials at cosine 0.80–0.97 to their global directions. Genre explains 11% of the space, and the community axis is not a genre axis. The within-genre scary direction at +4σ takes the household to The Ring and Poltergeist, with 0.00 of the list kept in its genres. The second attempt fits each user’s own direction on the films inside their neighborhood, with the ridge centered on the global direction and its strength chosen by the one-SE rule. It returns the global direction in 3,182 of 3,200 fits, at cosine 1.000 on every dial. The global direction already separates the tag inside each neighborhood at AUC 0.81–0.86. There is nothing local to learn. What does keep the household in animation is where the system looks: the same dial with eligibility limited to the viewer’s own top genres, which is what “still ours” now does on the page. In the shared-intent test that constraint keeps the genre by construction and pays for it in taste alignment at equal response. A pool limited to the viewer’s neighborhood saturates before it reaches the response at all. The re-rank inside that neighborhood matches the re-rank over the whole catalog on every column. That says the re-rank was a neighborhood method all along, and it is why the re-rank wins that test. On this space, “scary” does not mean something different near the household; the neighborhood is what differs.
Words, films, and a model at the controls
The dials are eight of 1,128 directions, one per Tag Genome tag. The other 1,120 are reachable three ways, all through the same arithmetic.
Search. The page’s dial search matches tag names by word and stem, and by meaning through a sentence embedding of the names (MiniLM), so “soothing” finds calm, relaxing, and feel-good though no name contains it. A meaning match says nothing about which way a tag points; in this space inspiring and courage lean with bleak (grim heroism), and the space’s opposite of bleak is happy ending (cosine −0.56), not “hopeful”, which it does not have. The search therefore splits its matches into the two poles they form in the taste space, anchored on the most opposed pair, and the viewer picks the side that means what they meant.
The word box. A sentence goes to a model and comes back as numbers. “Scarier, but nothing before 1990”, “like Children of Men but more hopeful”, “Hot Fuzz vibes” and “deeper cuts” are all requests it reads. The model is DeepSeek-Flash with reasoning by default. A local model on the Mac (Qwen3.6-35B-A3B) is the fallback and the other choice beside the box. The model sees about thirty candidate tags, made of stem matches for each content word, meaning neighbors, and the current dials. It answers only over that closed vocabulary. It can move at most four tags by up to ±2σ, move toward or away from at most two films, set genre and year constraints, keep or leave out named people, and step each list knob once. The server checks every field. A tag must be a candidate. A genre constraint stands only if the sentence named the genre. A word of a film’s title or of a person’s name is never a quality. The server then fits the stack to ±4σ and writes the note the page shows. No model text reaches the page.
The rules the server keeps. Six failures that were once the server’s are now rules rather than hopes for the prompt. A word of a named film’s title is never a tag. An empty reply counts as unmapped. The films a request names are found in the catalog before the model sees the sentence. A one-word title matches with its article set aside. The first word of a sentence counts as a title only when it is neither a request word nor a word of any tag. So “Godfather vibes” names The Godfather, and “Nothing like Transformers” names only Transformers. A tag with a same-lean neighbor that the space encodes more cleanly is steered as that neighbor, and the note says so. Every “scarier” phrasing now steers as creepy, whose AUC is 0.91 against 0.86 for scary, at a direction cosine of 0.73. Eerie, at cosine 0.57, is a different direction and is never substituted. People a request names are found the same way, so the model only chooses whether they are kept or left out.
The local model’s numbers. The evaluation runs 52 author-written phrasings at temperature 0. In the latest run the constraints were exact in 52/52. The film a phrasing names was moved toward, or away from for “nothing like Transformers”, in 11/11 of the phrasings that name one. The net tag move projected the expected way on every expected direction in 50/52. All three held in 50/52. The two misses are the model’s. For “like Children of Men but more hopeful” it moves toward the film but chooses heartwarming and dramatic over a negative move on bleak, so the net move does not leave bleakness. For “a scary movie but not a horror film” it adds a lean away from horror beside the genre exclusion the prompt asks for. Across eight runs of the growing phrasing set, under small prompt edits, the miss count stayed between one and four. A 33-phrasing set from before the knobs were added scored 33/33. Two cautions apply. The small model’s choices on the hardest phrasings are not stable. And these are the author’s phrasings, used during development, so the number is a development check. It is not evidence of generalization. Median latency is 3 s.
What the sign test does not see. The evaluation also asks whether every tag the model moved leans with something the sentence asked for. A moved tag passes when its direction has cosine at least 0.7 to an expected tag, to a quality the sentence licenses, or to the named film. On that column the local model is loose. It left 28/52 phrasings with no extra tag, added 1.06 extra tags per phrasing, and answered “no horror” with a genre exclusion and three tag moves besides. Two hosted models ran through the same server on the earlier 47 phrasings. Both scored 47/47 on the sign test, and they separated only on this column. DeepSeek-Flash with reasoning stayed within the request on 39/46 at a median of 3 s. GLM-5.3-Flash did on 38/47 at 14 s. DeepSeek-Flash without reasoning did on 13/47 at 1.2 s, reading every intent and stacking a tag and a half on each. DeepSeek-Flash with reasoning is what the page uses unless a visitor picks otherwise, so the numbers a visitor gets are its 47/47 and 39/46. The local model answers when the hosted one is over its daily allowance or fails, and the note says so. Every model stacks tags beside a film move. “Hot Fuzz vibes” gets irreverent, stylish, explosions and brutal from all four, so those phrasings separate nobody.
People and “still ours”. “Anything with Tom Hanks” was the model unmappable request until the catalog got its credits. The server resolves an actor, director or writer against TMDB’s credits for the 13,174 films, the way it resolves titles. The credits hold the top-ten cast, the directors and the writers; for books, the author. The person becomes a hard constraint beside genre and year. So “Tom Hanks but scarier” is a filter and a dial, and “nothing directed by Michael Bay” leaves his thirteen films out. A series or a studio is the same kind of constraint. Each film’s TMDB collection and production companies are indexed beside the credits. So “more Marvel movies” is the 60 films Marvel made, ranked by the viewer’s taste, and “the Harry Potter films” is eight. The first attempt at that sentence had filtered the list to Elizabeth Marvel, an actress with six films here. That is why a bare surname or a studio name that is also a word of the tag vocabulary now has to pass a test. The tag’s top films must be mostly that person’s or that studio’s. The taste space has no axis for a person and should not, and the model never names a film the request did not; a franchise the viewer wants comes from the catalog’s own data, not from the model’s memory. “Scarier, still ours” is a constraint too. The page takes the viewer’s own top-two genres, by count among the films they rated, and limits the list to them while the dial moves. The household’s list gets scarier inside animation that way, with Coraline third at +2σ on the eerie dial, and no direction tested did the same. The geometry section below has the two attempts at a local direction that did not.
A film as a direction. “Like The Godfather” is not a tag. Moving toward a film is a step along the unit direction from the viewer’s position to the film’s vector, with σ the catalog’s spread along that direction, so +1 toward a film means what +1 on a dial means. Several films can be active; the film itself is left out of the list. For the Pixar household, Children of Men at +2σ with bleak at −1 and a Sci-Fi constraint gives Ex Machina, Gravity, District 9, Looper, Arrival, 28 Days Later and Snowpiercer. The list’s alignment to the household’s own taste falls from 0.95 to 0.19. That is the price of leaving, and the page shows it.
Three list knobs. λ is diversity, maximal marginal relevance over the genome. β runs −0.5…0.5; below zero it favors the thinly known, with the shrink still keeping noisy vectors out. Explore runs 0–1 and penalizes similarity to the viewer’s own taste, so a steered list leaves their neighborhood in the steered direction. With nothing steered, explore has no direction and only adds diversity, and the page says so. Every list ends with its own numbers: intra-list diversity, popularity percentile, alignment with the viewer’s taste.
The MCP server. The same session logic powers MCP tools (rec_place, rec_tags, rec_steer, rec_toward, rec_inspect, rec_list, rec_knobs, rec_reveal, …) that a language model drives in conversation. The server has no model inside it. Grounding the viewer’s words in the vocabulary is the client’s job. The tools return the evidence the client needs to explain a list: each film’s genome scores on the active tags, its cosines to the toward-films, and the net movement along each toward-film’s direction. A stack of tags that cancels a film move is therefore visible. Seven fresh-session transcripts drove the design: the client inventing ratings, mood words the vocabulary lacks, opposed synonyms, warnings ignored three times until the scaling became automatic.
The surprise seat
Heart of Gold’s recipe, ported. Beside the twenty films nearest the viewer (the floor), one more seat is filled by a film from well outside that neighborhood which a certificate connects to the floor through a concept the floor shares. The 1,128 genome tags stand in for a trained dictionary, so nothing is learned. The floor’s shared concepts are its top eight tags by z-score against the catalog, genre names excluded. The outside pool begins one floor-width beyond the floor’s edge in cosine. Obviousness is sharing the floor’s dominant genre. The certificate is the strongest film on the concept within its cosine decile of the pool. The seat is unmarked and scores are withheld from any list that carries one, so it cannot be told apart until the reveal. For the Pixar household the seat is Jobs (2013), on pixar animation: the floor’s mean 0.49 on that tag, this film 0.31 and the strongest of 652 in its band, at cosine 0.00 to the household.
E12. Over 1,000 sampled viewers with no steer, a hit is a film among the viewer’s held-out likes (test rating ≥ 4, 7.8 per viewer). The seat is compared with two placebos from its own cosine band of the same outside pool. One is a random non-obvious film. The other is the non-obvious film nearest the seat in rating count. The first version of this test read as a win (16 seat hits against 3) and was wrong twice over. The seat and its placebo had been drawn from deciles of two different pools, which an outside review caught. Once they shared a pool, the seat’s edge (16 against 6) came with a median 5,409 ratings against the placebo’s 1,661. Every seat hit was a film with 24,000 or more ratings, and all but one sat on a certificate like oscar (best sound) or brilliant. A shared verdict is not a connection, so the 89 tags that grade a film (awards, lists, verdicts on craft) now sit outside the concept set with the genre names. With that change, the certificates are spielberg, exciting, coen brothers, intelligent, tarantino, kubrick, short-term memory loss and fighting the system. The seat hits 0.9% (9/995, Wilson 95% 0.5–1.7), the random placebo 0.2% (2/995), and the popularity-matched placebo 0.6% (6/995). Against the popularity-matched film the discordant pairs are 9 to 6, p = 0.61 by exact sign test. The ordinary recommendation list has a held-out hit rate of 3.8% per slot.
On held-out ratings the certified seat showed no statistically significant advantage over an equally distant, equally popular film; the test does not establish that the two are equivalent, only that it found no difference. The certificate measures salience on a concept the list shares, and the page claims nothing more for it. The seat stays, by decision rather than by this number. Whether a connection like tarantino or fighting the system pleased anyone is a question held-out ratings do not answer, since a film two neighborhoods away is one the viewer mostly never rated. Only feedback on the seats themselves can answer it. The reveal shows the certificate so the viewer can judge it, and now asks: under every revealed seat, on the page and through the MCP server, is a yes-or-no on whether it was worth seating, recorded with the certificate. The bar is a majority over thirty verdicts; below it the seat goes. That tally is the evaluation this build does not yet have, and the page will carry it when it exists.
A study, pre-registered
The shared-intent result above is a number about lists. Whether people agree is a different question, and it has its own study: Judge two lists. A participant rates a few anchor films and picks three requests. For each request they see two blinded ten-film lists, the steered one and a re-rank calibrated to the same mean response. They answer which list better answers the request and which still feels like theirs. The hypotheses, exclusions, tests, and power table were fixed in the plan (commit 6b65566, 2026-09-12) before recruitment: thirty participants detect a 0.75 preference four times in five; fewer see only a strong one. Results will appear here, whatever they say.
The closest published work is “Discovering Personalized Semantics for Soft Attributes in Recommender Systems using Concept Activation Vectors” (ACM Transactions on Recommender Systems, 2024; an earlier version at The Web Conference 2022). Its authors are Göpfert, Haig, Hsu, Chow, Vendrov, Lu, Ramachandran, Pham, Ghavamzadeh and Boutilier. They fit CAVs for MovieLens tags on the item embeddings of a rating-only collaborative-filtering model and test whether the model has learned each attribute. They use a CAV to critique a user by adding a step along it to the user’s embedding, which is the same update rule as a dial here. They go further on the labels. A tag can be subjective in degree, with a per-user threshold, and in sense, with several CAVs per tag found by an EM procedure over users. They find pairwise ranking losses beat logistic regression for CAV accuracy. This build differs in four ways. It uses the Tag Genome’s continuous relevance scores rather than users’ sparse tag applications. It fixes the step unit as the catalog’s spread along the direction rather than a tuned hyperparameter. It scores by cosine in a centered space with a popularity shrink. And it evaluates steering as movement, through monotonicity, cross-talk, relevance retention and the shared-intent test, rather than through a simulated critiquing session. It has no notion of subjectivity; their sense-clustering is one candidate answer to the “scarier, still ours” problem above, alongside cluster-local directions.
Stack
Python, numpy, scipy, scikit-learn; one browser page with no dependencies; static on Cloudflare. Spec, amendments, and a numbered decision log are in the repository. Data from GroupLens Research (MovieLens 25M and the Tag Genome).