Experience & Skills
I’m a data scientist and data engineer with a master’s in data science from UMBC and six years of experience in healthcare and e-commerce data. I build data systems end to end, from the pipeline to the model to the evaluation, and I test whether they work before anyone relies on them. At work that has meant Medicaid claims pipelines and product classifiers. My independent projects apply the same habits to retrieval, recommendation, and language-model evaluation.
I’m looking for data science, applied machine learning, and data engineering roles, remote or hybrid, where I can define the problem, build the system, and find out whether the result is useful.
Resume (PDF) · LinkedIn · Email me
Professional experience
Mathematica · Data Engineer
July 2023–September 2026 · remote
Lead ETL architect for Medicaid analytics, building PySpark and Spark SQL pipelines in Databricks over millions of claims.
- Analyzed medications for opioid use disorder (MOUD) across every Medicaid beneficiary in the country.
- Led ETL architecture for Utah’s Medicaid cost benchmark on a six-person team of analysts, developers, and a project director. The team’s work secured a $560,000 contract renewal.
- Worked with claims and healthcare spending data for New Jersey and Utah, and built data pipelines for North Carolina and Tennessee.
- Cut weekly pipeline runtime by 75% during a platform migration, without missing a delivery.
- Built a module that identifies co-occurring physical and behavioral health conditions. CMS now receives it as an analytics product.
- Carried up to seven client projects at once in 2025 and 2026, with weekly, quarterly, annual, and ad hoc deliveries.
- Onboarded new developers with live walkthroughs, written guides, and pipeline documentation.
Automated Export Processing · Data Scientist Lead
May 2020–May 2023 · remote
Automated Export Processing sells software that automates the export paperwork for outdoor gear with complex export rules. Its customers are e-commerce retailers that had sold only in the US because those rules made selling abroad too hard.
- Cut manual product labeling by 60% with a deep learning classifier that assigns categories to unstructured product data. That freed staff to take on more clients.
- Designed the company’s analytics platform in Python, SQL, Tableau, and Power BI for customer profiling, market analysis, and forecasting. It became a core product feature.
- Found systemic errors in the source data with automated quality checks and worked with the U.S. Census Bureau to resolve them.
Before data · Teaching
2017–2019
I taught English as a second language to Pre-K through 8th grade students at City Arts & Prep Public Charter School in Washington, DC, for the 2018–19 school year. Over two summers I also wrote and taught a criminal justice course for high school students at Explo at Yale. That is where I learned to explain a system to the person who has to use it.
How the projects map to business problems
The subject matter on this site is unusual on purpose. There is no ready-made notebook for Survivor tribal councils or Star Trek subtitles, so each project had to start with the questions a business project starts with. What is the unit of prediction, what data can I actually get, what would count as working, and what would fool me into thinking it works? The table gives the business problem each project shares its mechanics with.
| The business problem | Public example | What the example shows |
|---|---|---|
| Building the tables analysts rely on, and letting people ask them questions in plain English | Gamebot | A medallion lakehouse on Airflow, dbt, and PostgreSQL with tests and lineage, a published package, and a question box with allow-listed SQL, spending caps, and an evaluation set |
| Watching a market from public records, and labeling what appears with a classifier whose agreement is measured | Jobwatch | A daily warehouse of job postings from public job-board endpoints, a jury of three models with a regex anchor, kappa and unanimity published from every run, checks that stop a bad run before it reaches the site, and a page used every morning |
| Finding the entities that behave like a given one, without confusing behavior with quality | Player comps | A topic model over every pitch of two MLB seasons, a vocabulary designed to measure the hitter and rescaled when a pre-registered gate showed it measuring the pitchers, a ranking that reports how many of twenty fits agree, and an evaluation published whichever way it came out |
| Predicting which customers, accounts, or claims will do something next | Who goes home? | One prediction per event, features built only from earlier records, an evaluation fixed before training, calibration checks, and a survival model for how long people last |
| Finding who acts together in event records, like accounts that transact together | Pearl Islands | Agreement rates over comparable opportunities, each traceable to the ballots behind it |
| Measuring whether a campaign or a launch changed behavior, and forecasting demand | The Oscar bump | An event study with matched comparisons and permutation checks, and a forecast backtested without leakage |
| Recommending products or content, and proving a change is an improvement | Steerable recommender | Two recommenders on separate datasets, a steering interface, a first factorization replaced because its lists were poor despite good prediction error, and a blinded user study with the comparison decided in advance |
| Answering questions from a company’s own documents without inventing sources | Ask Mindscape and Ask the Encyclopedia | Retrieval, a draft written blind, a visible review of every claim, an answer hashed before the sources are revealed, and an evaluation against recorded answers |
| Searching a large document collection reliably | Heart of Gold and Mindscape Explorer | Dense and keyword retrieval with reranking, readiness checks, time budgets, and fallbacks, and a way to browse when the right search terms are not known |
| Checking that a text classifier is reading the content rather than the names | Which series is this? | A ladder of masks that removes identifying terms step by step, with a human baseline for the model |
| Knowing who said what in recorded conversations | Who said that? | Speaker attribution from audio and captions, with the measurements reported before the modeling |
| Deciding whether a modified model is actually better than the one in use | StargazerLabs | Open-weight model experiments that report what an apparent improvement costs |
The habits behind these are the same on every project: a written plan before the code, an evaluation fixed before any tuning, nothing dropped from the data without being counted, and every constant on a page traceable to a file. Those habits are what make a result reproducible by the next person, which in a team is most of the value.
Technical toolkit
- Data engineering and analytics: Python, SQL, PySpark, Spark SQL, Databricks, Airflow, dbt, PostgreSQL, DuckDB, Parquet, pandas, Docker, Tableau, and Power BI.
- Modeling and evaluation: scikit-learn, PyTorch, classification, forecasting, recommendation, NLP, cross-validation, and uncertainty estimates.
- Applied AI systems: embeddings, retrieval and reranking, LangGraph, FastAPI, MCP servers, local model serving, and MLX. My StargazerLabs work includes open-weight models and post-training experiments.
- Deployment and workflow: Git, Cloudflare Workers and D1, services on launchd behind a Cloudflare tunnel, and AI coding assistants (mostly Claude).
Education
University of Maryland, Baltimore County · Master of Professional Studies in Data Science, December 2022.
Washington University in St. Louis · BA in International Relations, Philosophy, and Chinese, May 2018.
Get in touch
For a hiring conversation, email me with the role and the problems your team is working on. I’m based in Rockville, Maryland, and looking for remote or hybrid work. My resume provides a compact overview, and the project pages link to demos and technical write-ups.