Data Jobs Observatory

Methodology and caveats

What this data is, how it is built, and the ways it can mislead.

Sources

Public job-board APIs of three applicant tracking systems. No scraping of HTML, no logins, no aggregators.

The board list (data/boards.csv) holds 2,898 boards: 1,506 ashby, 1,095 greenhouse, 297 lever. Requests are paced to one start every 0.3 seconds per API host with a descriptive User-Agent, once a day.

Coverage bias

Pipeline

  1. Ingest (Python): every board is fetched; postings whose title matches a data-role pattern are landed as Parquet, one file per board per day, with the raw ATS JSON kept verbatim (bronze). Runs are resumable and idempotent.
  2. History: each day is folded into four small Parquet files: postings with first and last sighting, daily sightings, board-crawl results, and SCD type 2 versions of each posting (title, location, pay range, description hash).
  3. Warehouse (DuckDB + dbt): staging, intermediate and mart models; tests on every key, relationship and enumerated column; a hand-labelled title set that the classifier must reproduce exactly. Browse the dbt docs and lineage graph.
  4. Dashboard: static pages rebuilt from the marts after every run.

Definitions

Keyword matching counts mentions, not requirements

Skills are 83 documented patterns (skill_taxonomy.csv) matched against the description. "Nice to have: Spark", "our stack includes Spark" and "you will never touch Spark" all count as a mention. Company boilerplate (benefits, values, "we use AI to...") also produces mentions, which inflates generic terms such as communication, LLM and dashboards. Shares are therefore upper bounds on how often a skill is actually required.

Pay

Degrees and years

A degree is "named" when the word appears anywhere, including "or equivalent experience" and "preferred". Years asked is the highest "N years ... experience" phrase, which is a proxy for the bar, not a hard cut-off.

Recent crawls