Jobwatch
DuckDB loads in the browser on the first query (a few MB, once). The data is as old as this morning's run.
Jobwatch is a data warehouse for a job search. Every morning it reads the public job-board endpoints that Greenhouse, Lever and Ashby publish for 58 of the 115 companies on a list I chose. It keeps every posting it has seen and marks the day a posting disappears from its board. Three language models then read each new posting with the same prompt and answer from closed vocabularies. What the role is, how senior, where the work happens, which tools the text names, and what it pays when it says. The verdict is the majority. The page (opens in a new tab) is the tool I use to watch 115 careers pages without opening them. It is the same DuckDB-in-the-browser machinery as Gamebot, on a business dataset instead of a TV show.
In a businessMany teams watch a market through public records, whether the records are prices, filings, or job postings. This project keeps full history, publishes how often its classifiers agree, and stops a bad run before it reaches the page.
Technical skills: SQL · Python · DuckDB · Data modeling (medallion layers) · LLM labeling with measured agreement · Data quality checks · Scheduling and deployment · DuckDB-WASM
- My contribution
- I built the whole thing: the seed list and board finder, three ATS adapters with fixtures, the DuckDB warehouse in three layers, the three-model jury and its rules anchor, the gold SQL and the checks, the daily launchd run that redeploys the site, and the page.
- Result and limit
- 115 companies watched, 58 through public endpoints; 3,167 postings seen and 249 open data roles at 42 companies as of 2026-10-01. Juror agreement on the role family is kappa 0.85 to 0.92, unanimous 92% of the time, reported as agreement and not accuracy.
On this page
The list is the point. It holds health-tech companies, sports organizations, public-interest groups, product companies in the DC area, remote-first product companies, and a few consumer brands with small data teams. It leaves out consulting and contracting firms, defense contractors, and the tool vendors whose postings ask for experts in the tool. The warehouse does not cover the market and the page says so. It covers the places I would apply.
What it reads, and what it does not
Each of the three applicant-tracking systems publishes an unauthenticated JSON endpoint per company board, meant for embedding the board on a careers site. Jobwatch reads those and nothing else. No HTML is scraped, no aggregator is queried, and no people data is collected. A company on Workday, iCIMS or UKG has no clean endpoint. It is listed on the page as watched by hand, with its careers link, and nothing is fetched. The finder script tried each vendor’s endpoint with the obvious slugs for every company. Two slugs answered for the wrong company. A costar board belonged to an astrology app. The endpoint returns the board’s own name, and that name is how the mismatch gets caught.
The warehouse
One DuckDB file with the layers as schemas, and no orchestrator. The whole thing is one process a day, and Gamebot already shows Airflow and dbt. Bronze keeps every payload a board returned, one row per posting per change, with a hash so an unchanged posting costs nothing. Silver has one row per posting with its first and last sighting, its closed date, and the description as text. The converter’s version is a column. A second silver table has one row per posting per labeler. Gold is pure SQL, seven create or replace table files run in order. openings is the page’s table. verdicts turns three votes into one, with GPT-5.6 deciding a three-way split and a tool entering the stack when two jurors named it. company_summary, company_week and stack_month are the cards and the trends. label_agreement and label_summary are the agreement numbers, computed in SQL over every posting.
Checks run after gold and stop the run before export. They fail when a count moves the wrong way, when a verdict leaves its vocabulary, or when the verdicts on a fixed set of 200 postings differ from a snapshot taken when the prompt was final. That last one is a regression guard. The pipeline can change in a hundred ways without a label moving. If a label moves anyway, something changed that should have been a new prompt version. A launchd job runs the whole thing at seven each morning, then rebuilds and deploys this site. The page is as old as the last successful run. A failed check leaves yesterday’s data in place.
The jury, and how it is checked
There are no hand labels. I had two hours to label 200 postings and decided the hours were better spent on a design that did not need them. Three models label every posting instead. DeepSeek-Flash, GLM-5.3-Flash and GPT-5.6 Luna, through one OpenAI-compatible client at temperature zero, JSON out. The prompt is under 2,000 tokens, with the vocabularies and two examples per role family. A regex labeler reads the same text and is reported beside the vote without voting, as an anchor that cannot drift.
What is reported is agreement, and agreement is not accuracy. A label all three models get wrong the same way counts as unanimous. Over the 3,167 postings in the warehouse, Cohen’s kappa between pairs of jurors is 0.85 to 0.92 on the role family and 0.81 to 0.87 on seniority. It is 0.81 to 0.85 on workplace and 0.97 to 1.00 on the clearance flag. The jury is unanimous on the role family 92% of the time and on workplace 82% of the time. The majority matches the regex anchor 86% of the time on the role family. Workplace is the weakest label, as expected. The models read the location text and the description together and the regex sees only the location, and onsite against hybrid is where they split. On a fixed sample of 200 postings drawn before the prompt was final, the same numbers hold. The seventeen contested titles are ones a person would also split on. “Engineering Manager, Data” and “Staff Product Engineer, Enterprise AI” are two of them. The page’s last section shows the numbers from the latest run. gold.verdicts holds every juror’s answer beside the verdict, so a reader can pull the disagreements and judge them.
Two clarifications are queued for a second prompt version. A posting titled “Analytics Engineer” belongs to analytics engineering by definition, and “Manager, Data Analytics” is a manager before it is anything else. A prompt change means relabeling everything and re-snapshotting the 200, so it waits until the numbers are worth moving.
The page
The first section is the tool. Every open posting the jury put in a data family, newest first. Filters cover the role family, workplace, seniority, company group and how recently it opened. A search box and a row of stack chips, built from the tools the open data roles name, narrow it further. Each row links to the posting on the company’s own board. Pay appears only when the posting states it, from the board’s structured field first and from the text second, and never estimated. “Ask in words” sends a question to a hosted model that has read the tables, the vocabularies and the column notes. The model writes one query for the browser to run. The SQL panel runs it, with six ready queries. The lineage map shows the bronze and silver tables that stay in the warehouse beside the gold tables the page serves. A click on a gold table shows its columns and the SQL that builds it. The company cards carry what each company has open now and its careers link. Each card also has one link to my school’s LinkedIn alumni page with the company as the keyword, which is where a warm introduction starts. The alumni step stays a human step.
Stack
Python and uv, httpx, DuckDB, pyarrow, html2text, pytest. Three provider APIs through one client. A launchd agent on a Mac. DuckDB-WASM in the page. The same Cloudflare Worker route as Gamebot for the question box, with a second prompt. Keys live outside the repository. Postings are quoted on the page only as a short excerpt and linked to their source.