Maxwell Grody

Experience & Skills

I’m a data scientist and data engineer with a master’s in data science from UMBC and six years of experience in healthcare and e-commerce data. I build data systems end to end, from the pipeline to the model to the evaluation, and I test whether they work before anyone relies on them. At work that has meant Medicaid claims pipelines and product classifiers. My independent projects apply the same habits to retrieval, recommendation, and language-model evaluation.

I’m looking for data science, applied machine learning, and data engineering roles, remote or hybrid, where I can define the problem, build the system, and find out whether the result is useful.

Resume (PDF) · LinkedIn · Email me

Professional experience

Mathematica · Data Engineer

July 2023–September 2026 · remote

Lead ETL architect for Medicaid analytics, building PySpark and Spark SQL pipelines in Databricks over millions of claims.

Automated Export Processing · Data Scientist Lead

May 2020–May 2023 · remote

Automated Export Processing sells software that automates the export paperwork for outdoor gear with complex export rules. Its customers are e-commerce retailers that had sold only in the US because those rules made selling abroad too hard.

Before data · Teaching

2017–2019

I taught English as a second language to Pre-K through 8th grade students at City Arts & Prep Public Charter School in Washington, DC, for the 2018–19 school year. Over two summers I also wrote and taught a criminal justice course for high school students at Explo at Yale. That is where I learned to explain a system to the person who has to use it.

How the projects map to business problems

The subject matter on this site is unusual on purpose. There is no ready-made notebook for Survivor tribal councils or Star Trek subtitles, so each project had to start with the questions a business project starts with. What is the unit of prediction, what data can I actually get, what would count as working, and what would fool me into thinking it works? The table gives the business problem each project shares its mechanics with.

The business problem Public example What the example shows
Building the tables analysts rely on, and letting people ask them questions in plain English Gamebot A medallion lakehouse on Airflow, dbt, and PostgreSQL with tests and lineage, a published package, and a question box with allow-listed SQL, spending caps, and an evaluation set
Watching a market from public records, and labeling what appears with a classifier whose agreement is measured Jobwatch A daily warehouse of job postings from public job-board endpoints, a jury of three models with a regex anchor, kappa and unanimity published from every run, checks that stop a bad run before it reaches the site, and a page used every morning
Finding the entities that behave like a given one, without confusing behavior with quality Player comps A topic model over every pitch of two MLB seasons, a vocabulary designed to measure the hitter and rescaled when a pre-registered gate showed it measuring the pitchers, a ranking that reports how many of twenty fits agree, and an evaluation published whichever way it came out
Predicting which customers, accounts, or claims will do something next Who goes home? One prediction per event, features built only from earlier records, an evaluation fixed before training, calibration checks, and a survival model for how long people last
Finding who acts together in event records, like accounts that transact together Pearl Islands Agreement rates over comparable opportunities, each traceable to the ballots behind it
Measuring whether a campaign or a launch changed behavior, and forecasting demand The Oscar bump An event study with matched comparisons and permutation checks, and a forecast backtested without leakage
Recommending products or content, and proving a change is an improvement Steerable recommender Two recommenders on separate datasets, a steering interface, a first factorization replaced because its lists were poor despite good prediction error, and a blinded user study with the comparison decided in advance
Answering questions from a company’s own documents without inventing sources Ask Mindscape and Ask the Encyclopedia Retrieval, a draft written blind, a visible review of every claim, an answer hashed before the sources are revealed, and an evaluation against recorded answers
Searching a large document collection reliably Heart of Gold and Mindscape Explorer Dense and keyword retrieval with reranking, readiness checks, time budgets, and fallbacks, and a way to browse when the right search terms are not known
Checking that a text classifier is reading the content rather than the names Which series is this? A ladder of masks that removes identifying terms step by step, with a human baseline for the model
Knowing who said what in recorded conversations Who said that? Speaker attribution from audio and captions, with the measurements reported before the modeling
Deciding whether a modified model is actually better than the one in use StargazerLabs Open-weight model experiments that report what an apparent improvement costs

The habits behind these are the same on every project: a written plan before the code, an evaluation fixed before any tuning, nothing dropped from the data without being counted, and every constant on a page traceable to a file. Those habits are what make a result reproducible by the next person, which in a team is most of the value.

Technical toolkit

Education

University of Maryland, Baltimore County · Master of Professional Studies in Data Science, December 2022.

Washington University in St. Louis · BA in International Relations, Philosophy, and Chinese, May 2018.

Get in touch

For a hiring conversation, email me with the role and the problems your team is working on. I’m based in Rockville, Maryland, and looking for remote or hybrid work. My resume provides a compact overview, and the project pages link to demos and technical write-ups.