Maxwell Grody

I’m Max Grody, a data scientist and data engineer in Rockville, Maryland.

I build data systems end to end, from the pipeline to the model to the evaluation, and I test whether they work before anyone relies on them. Most recently, at Mathematica, that meant leading ETL design for Medicaid analytics. On this site it means a warehouse, a recommender, and two question-answering services, each with its evaluation fixed before any tuning.

I’m particularly interested in what an evaluation can miss when a system appears to work.

I’m looking for my next role in data science, applied ML, or data engineering.

Maxwell Grody

Try a live demo ↓Experience & skills

Try one in thirty seconds

No account required. Start with an example query, a recommendation dial, or a guessing round; typed requests use a hosted model.

Selected work

If you read two, read the first two. Each one names the business problem it shares its mechanics with.

A warehouse for 77 seasons of Survivor, queried in your browser

Data engineering · Airflow, dbt, PostgreSQL · medallion lakehouse

A data warehouse for 77 seasons of Survivor that you can query in your browser. Airflow and dbt turn the raw records into 8 feature tables and 2 model-ready tables, with tests at every layer. The data is also published as a Python package, gamebot-lite, and the other Survivor projects on this site are built on it.

In a business: Teams need reliable tables before they can trust a dashboard or train a model. This project builds that foundation with scheduled ingestion, SQL transformations, data tests, and documented lineage, then ships the data as a Python package analysts can use.

Technical skills: SQL · Python / pandas · Airflow · dbt · PostgreSQL · DuckDB / Parquet · Docker

DuckDB loads in the browser on the first query (a few MB, once). Full-season spoilers in the tables.

A warehouse of job openings, read every morning

Data engineering · DuckDB · LLM labeling · 115 companies watched daily

A warehouse of job postings from 115 companies, refreshed every morning. Three language models label each posting for role, seniority, remote policy, tools, and pay, and the page reports how often they agree. It uses the same warehouse design as Gamebot on business data, and I use it in my own job search.

In a business: Many teams watch a market through public records, whether the records are prices, filings, or job postings. This project keeps full history, publishes how often its classifiers agree, and stops a bad run before it reaches the page.

Technical skills: SQL · Python · DuckDB · Data modeling (medallion layers) · LLM labeling with measured agreement · Data quality checks · Scheduling and deployment · DuckDB-WASM

DuckDB loads in the browser on the first query (a few MB, once). The data is as old as this morning's run.

Hitters who operate the same way, from every pitch of two seasons

Analytics · dbt, PostgreSQL · topic modeling

Type an MLB hitter and get the hitters with the same approach at the plate, whether or not they produce like him. The model learns 6 styles of hitting from every pitch of the last two seasons. Each match shows how many of 20 model runs agree on it, so you can tell the solid matches from the shaky ones.

In a business: Many teams need to find the entities that behave like a given one without confusing behavior with quality, from customers who use a product alike to suppliers that fail alike. Each comp here reports how robust it is, and the page says what the model could not test.

Technical skills: Python · SQL · dbt · PostgreSQL · Topic modeling (NMF, LDA) · Model selection with stated rules · Pre-registered evaluation · Data modeling (medallion layers)

Everything loads at once, about a third of a megabyte, and the ranking runs in the browser. Defaults to the most recent season.

Question answering across a podcast archive and an encyclopedia

Retrieval and question answering · LangGraph · evaluated against recorded answers

Ask a question across 441 episodes of a science podcast and get an answer with its sources, a second model’s check of each claim, and one related passage you did not ask for. In a test of 50 listener questions against the original 436-episode archive, it found the reference passage 35 times and cited it 34 times. The same system answers questions from the Stanford Encyclopedia of Philosophy.

In a business: A team searching its own policies or records needs answers it can check against the source. This project evaluates finding the right source and writing the answer separately, because a found source does not make the answer right.

Technical skills: Python · FastAPI · LangGraph · MCP · Retrieval-augmented generation (RAG) · Source-recovery evaluation

Live answers may take a minute or longer. A recorded example is available in the demo.

Recommenders you can steer

Recommendation systems · collaborative filtering · user controls

Start from your taste in movies or books, ask for a change like “darker, but nothing scary,” and watch the list re-sort. In testing, the controls moved recommendations the requested way for 94% of viewers and dials, and typed requests worked in 45 of 47 test phrasings. The project page compares steering with a simpler re-rank.

In a business: Product teams need recommendations that follow both a customer's history and what they want right now. I measure whether a requested change takes effect and what relevance it costs against simpler re-ranking, so a team can pick an approach before running an A/B test.

Technical skills: Python · NumPy / SciPy · scikit-learn · Collaborative filtering · Ranking evaluation · JavaScript

Two catalogs, separate models and evaluations. Dials run in your browser; typed requests use a server.

Browse all projects →

Let’s talk

I’m open to data science, applied ML, and data engineering roles. My background spans production data systems, analytics, and applied AI.

Experience & skills →

Selected writing

All writing and essays →