If you read two, read the first two. Each one names the business problem it shares its mechanics with.
Data engineering · Airflow, dbt, PostgreSQL · medallion lakehouse
A data warehouse for 77 seasons of Survivor that you can query in your browser. Airflow and dbt turn the raw records into 8 feature tables and 2 model-ready tables, with tests at every layer. The data is also published as a Python package, gamebot-lite, and the other Survivor projects on this site are built on it.
In a business: Teams need reliable tables before they can trust a dashboard or train a model. This project builds that foundation with scheduled ingestion, SQL transformations, data tests, and documented lineage, then ships the data as a Python package analysts can use.
Technical skills: SQL · Python / pandas · Airflow · dbt · PostgreSQL · DuckDB / Parquet · Docker
DuckDB loads in the browser on the first query (a few MB, once). Full-season spoilers in the tables.
Data engineering · DuckDB · LLM labeling · 115 companies watched daily
A warehouse of job postings from 115 companies, refreshed every morning. Three language models label each posting for role, seniority, remote policy, tools, and pay, and the page reports how often they agree. It uses the same warehouse design as Gamebot on business data, and I use it in my own job search.
In a business: Many teams watch a market through public records, whether the records are prices, filings, or job postings. This project keeps full history, publishes how often its classifiers agree, and stops a bad run before it reaches the page.
Technical skills: SQL · Python · DuckDB · Data modeling (medallion layers) · LLM labeling with measured agreement · Data quality checks · Scheduling and deployment · DuckDB-WASM
DuckDB loads in the browser on the first query (a few MB, once). The data is as old as this morning's run.
Analytics · dbt, PostgreSQL · topic modeling
Type an MLB hitter and get the hitters with the same approach at the plate, whether or not they produce like him. The model learns 6 styles of hitting from every pitch of the last two seasons. Each match shows how many of 20 model runs agree on it, so you can tell the solid matches from the shaky ones.
In a business: Many teams need to find the entities that behave like a given one without confusing behavior with quality, from customers who use a product alike to suppliers that fail alike. Each comp here reports how robust it is, and the page says what the model could not test.
Technical skills: Python · SQL · dbt · PostgreSQL · Topic modeling (NMF, LDA) · Model selection with stated rules · Pre-registered evaluation · Data modeling (medallion layers)
Everything loads at once, about a third of a megabyte, and the ranking runs in the browser. Defaults to the most recent season.
Retrieval and question answering · LangGraph · evaluated against recorded answers
Ask a question across 441 episodes of a science podcast and get an answer with its sources, a second model’s check of each claim, and one related passage you did not ask for. In a test of 50 listener questions against the original 436-episode archive, it found the reference passage 35 times and cited it 34 times. The same system answers questions from the Stanford Encyclopedia of Philosophy.
In a business: A team searching its own policies or records needs answers it can check against the source. This project evaluates finding the right source and writing the answer separately, because a found source does not make the answer right.
Technical skills: Python · FastAPI · LangGraph · MCP · Retrieval-augmented generation (RAG) · Source-recovery evaluation
Live answers may take a minute or longer. A recorded example is available in the demo.
Recommendation systems · collaborative filtering · user controls
Start from your taste in movies or books, ask for a change like “darker, but nothing scary,” and watch the list re-sort. In testing, the controls moved recommendations the requested way for 94% of viewers and dials, and typed requests worked in 45 of 47 test phrasings. The project page compares steering with a simpler re-rank.
In a business: Product teams need recommendations that follow both a customer's history and what they want right now. I measure whether a requested change takes effect and what relevance it costs against simpler re-ranking, so a team can pick an approach before running an A/B test.
Technical skills: Python · NumPy / SciPy · scikit-learn · Collaborative filtering · Ranking evaluation · JavaScript
Two catalogs, separate models and evaluations. Dials run in your browser; typed requests use a server.