Maxwell Grody

Player comps / hitters who operate the same way

Find the hitters who operate like this one

Type a hitter. You get the hitters whose seasons are built from the same habits, whatever their production. A topic model reads every pitch of nine hundred hitter-seasons and finds six ways of operating. Each comp's badge says in how many of twenty runs of that model the pair held up.

The project page says how the words are made, what the model found, and what it could not test yet. The warehouse page (opens in a new tab) has the bronze, silver and gold tables behind it, with SQL in the browser. Statcast pitch data from Baseball Savant, identities from the Chadwick Bureau register.

The comps

Pick a hitter and a season, then narrow the list with the filters. Comps match habits and ignore output, so a star's closest comp can be a much weaker hitter. The wOBA column shows how far apart two comps are in production.

How to read the list

A season is treated as a document and every pitch as a word, labeled by what was thrown, where it crossed the plate and what the hitter did with it. Two hitters are comps when their mixes of the six topics are close. A topic model gives a slightly different answer each time it runs, so this one was run twenty times, at different numbers of topics and from different random starts. Each run is a fit. The list is the fifty hitters who were his top-twenty neighbor in the most fits, filtered by what you set. The two numbers beside each comp answer two different questions. The fits count says how robust the pairing is. In how many of the twenty fits did these two land in each other's top twenty? A comp at 20 of 20 comes out the same however the model is run. A comp at 8 of 20 depends on how many topics you count. Similarity says how close the two mixes are in one fit, the six-topic one the page draws below. It is one minus the distance between the two mixes, so 1 is identical and 0 is no overlap. Two comps with the same fits count are ordered by similarity. wOBA is production, computed from the same pitches, and it is shown so you can see that the comps span it. No topic tracks production, so there is one similarity score. Everything runs in your browser from a third of a megabyte.

The fantasy question is "who is a cheaper version of this hitter?" A cheaper version operates the same way and has produced less, which is what the last switch shows. Production is not price, so a hitter who produced less may still be expensive, and a draft-position column is the next step.

How the topics are made

Every number here comes from the files the page reads, and the thresholds were written down before the runs. A word has three parts. The pitch class is fastball, breaking or offspeed. The attack zone is heart, shadow, chase or waste, rebuilt from the pitch coordinates with Savant's published rings. The action is a take, a whiff, a foul, or one of Savant's six contact classes. That is 108 words. Words with fewer than 500 appearances across the league merge into a neighbor. The counts are then rescaled so every hitter looks as if he was thrown the league-average mix of pitches. The first fit skipped that step and its topics described the pitchers instead of the hitters. The table shows both fits.

Known limits

Only hitters with 400 pitches in a season are in the corpus, so the topics describe the styles that survive in the majors. Styles that never reach 400 pitches are not in it. The topics are gradients more than types. Most hitters sit in a smooth cloud, and the sentences under the bars have less to say than they would if hitters came in kinds. Reliability within a season is about 0.7, so a hitter's own two halves differ a little. A comp at 20 of 20 survives the modeling choices. That is a different thing from being certain. The test of whether close hitters share a platoon split or a park sensitivity could not be run at one season, because a season's platoon split does not agree with itself between its own halves. It is queued for the pooled seasons. A season of a hitter is one document, so a hitter whose approach changed in July is one blended document here.

Two comparison tools exist already. Baseball Savant's comparison tool draws two hitters you pick side by side. Baseball Reference's similarity scores compare career counting stats. This page searches for the comps. It names what a comp is made of, and it says how much the answer depends on the modeling choices.

Data: Statcast pitch-level data through Baseball Savant's search export, read through pybaseball, regular season only, and the Chadwick Bureau register for names and ages. Contact classes, attack zones and wOBA are Savant's definitions. Built by Max Grody.