Maxwell Grody

Player comps

Analytics · dbt · PostgreSQL · a topic model over 1,373,138 pitches · 2025 and 2026 through September 15

Find the comps ↗ (opens in a new tab) Query the warehouse ↗ (opens in a new tab) · Query the warehouse ↗ (opens in a new tab)

Everything loads at once, about a third of a megabyte, and the ranking runs in the browser. Defaults to the most recent season.

Two comparison tools exist for hitters, and neither searches. Baseball Savant’s comparison tool draws two hitters you pick. Baseball Reference’s similarity scores compare career counting stats. This tool (opens in a new tab) takes one hitter and returns the hitters whose seasons are built from the same habits. It names what the comp is made of, and it says how much the answer depends on the choices a modeler makes. It was built in a week of evenings, and it is in flight. The version here has hitters and two seasons. Pitchers, Triple-A, and a model with uncertainty on every number are the next steps, and the design document in the repository says what each one waits on.

In a businessMany teams need to find the entities that behave like a given one without confusing behavior with quality, from customers who use a product alike to suppliers that fail alike. Each comp here reports how robust it is, and the page says what the model could not test.

Technical skills: Python · SQL · dbt · PostgreSQL · Topic modeling (NMF, LDA) · Model selection with stated rules · Pre-registered evaluation · Data modeling (medallion layers)

My contribution
I built the whole thing: the Statcast warehouse in PostgreSQL and dbt with tests, the 108-word vocabulary and its merge rule, the topic-count sweep with split-half reliability, the pre-registered gates that caught a fit describing the pitchers, the consensus layer behind every badge, the evaluation, and the page.
Result and limit
902 hitter-seasons from 2025 and 2026 through September 15, 6 topics at reliability 0.73, no topic tracking production (largest wOBA correlation 0.13), and a first fit thrown out because 5 of its topics tracked the pitch diet. The median comp appears in 50 percent of 20 fits.
On this page

A season as a document

The idea is borrowed from text. A topic model reads many documents, finds a few recurring themes, and describes each document as a mix of them. Here one hitter’s season is a document and one pitch is a word. The word has three parts. What was thrown, in three classes. Where it crossed the plate, in Savant’s four rings, heart, shadow, chase and waste, rebuilt from the pitch coordinates. And what the hitter did, a take, a whiff, a foul, or one of Savant’s six contact classes from weak to barrel. That is 108 words. The contact class replaces the outcome on purpose. A barrel a fielder catches is still a barrel, so the park and the defense stay out of the words.

The pipeline is the same shape as Gamebot, and the warehouse page (opens in a new tab) lets you query it. Statcast lands in PostgreSQL as delivered, dbt builds the labeled pitches and the count matrix with tests, and the model reads Parquet. Every pitch either gets a word or a stated reason for none, and 27 words under 500 league-wide appearances merge into a neighbor by a written rule. The median hitter reaches five occurrences on 46 of the 81 surviving words, which is the effective vocabulary, and it is the reason the topic count is small.

The fit that was thrown out

The first fit described the pitchers. Its topics tracked how often a hitter was thrown offspeed pitches, or breaking balls, and the largest such correlation was 0.93. That was a named risk in the design, with a trigger written down before the run. Any topic that correlates with the pitch diet above 0.5, and more than with production, means the words must be conditioned on the pitch. The trigger fired on 5 of 6 topics. The fix rescales each hitter’s counts so he looks as if he was thrown the league-average mix, with what he did with each kind of pitch unchanged. After that, the largest diet correlation is 0.47 and no topic triggers. Both tables are on the page.

What the model found

The topic count comes from the rule the Mindscape explorer uses, with one substitution. That rule scores each count by coherence plus restart stability and takes the smallest count within one standard error of the best. Coherence was built for English, so here it is replaced by split-half reliability, the correlation of a hitter’s topic share between two halves of his own season. The rule picked 6 topics, with 4 and 11 as the coarser and finer views. Stability is 0.73, reliability is 0.73 when the halves are drawn at random and 0.7 when they are the first and second halves of the season. The gap between those two is how much a hitter’s approach moves inside a season, and it is small.

No topic is production in disguise. The design expected the first topic to be “good hitter” wearing a style costume, and it set a threshold of 0.85 in advance for catching it. The largest correlation between any topic and wOBA is 0.13. That changes the tool. It was going to show two similarities, one with production and one without, and since there is nothing to remove, it shows one. It also means the comps span production. James Wood’s 2026 list has Roman Anthony at 20 of 20, and further down it has Edouard Julien, who operates the same way and produced far less.

The topics are gradients more than types. Every topic’s most common words are takes and fouls on fastballs, because that is what every hitter mostly does. The names on the page were read off each topic’s most distinctive words instead, the words it uses far more than the average topic does. Three of the six are “Takes strikes”, “Big swings in the zone, barrels and whiffs” and “Lifts fastballs over the heart”. The distinctive words sit beside each name so a reader can disagree.

The badge

A topic model gives a different answer from a different random start, or at a different number of topics. Each such run is a fit, and the tool does not hide that the fits disagree. It runs every count from four to ten topics, five times each, which is twenty fits. It records each hitter’s twenty nearest neighbors in each fit, and it ranks comps by the number of fits in which the pair appear. That share is the badge, “18 of 20”. The median pair in a top-twenty list appears in 50 percent of fits, and 1.1 percent of pairs appear in all of them, so the badge separates comps that survive the choices from comps that depend on them. A check was written to remove the badge if nearly every pair came out at 20 of 20. It did not.

What could not be tested yet

The strongest test in the design asks whether hitters who are close in topic space also match on two things the words never read, the platoon split and the home-away gap in production. It was run with the thresholds fixed first, and against fifteen plain Statcast rates at the same job. It came back inconclusive. A hitter’s platoon split in the first half of a season agrees with his second half at r = +0.20, over 118 hitters with enough plate appearances on each side, and the home-away gap at r = -0.02. Those are the ceilings, and no method can beat them. With that little signal in one season, neither topic distance nor the plain rates predicted either quantity. The test is queued for the pooled seasons, where the ceiling is higher, and the result goes on the page whichever way it comes out.

Stack

Python and uv, pybaseball for the Statcast export and the Chadwick register, PostgreSQL 17 and dbt with schema tests, scikit-learn for NMF and LDA, pytest with saved fixtures. The page is static, a third of a megabyte of square-rooted topic mixes and neighbor lists, and the distance runs in the browser. Every constant is logged with its reason in the repository’s decision log. No HTML was scraped. FanGraphs was dropped on the first day when the library’s route to it turned out to be a scrape, and wOBA is computed from Savant’s own per-pitch fields instead.