Data Jobs Observatory

Can a listing predict its own posted pay?

A model that predicts the posted pay midpoint from the role family, seniority, skills mentioned, experience and degree asks, and the description text. 2,416 priced listings from 847 companies, posted in USD, any location.

Verdict. The best model (Ridge: structured + skills + TF-IDF) modestly beats the best simple baseline (median of family x seniority): mean absolute error $36.2k against $41.2k, an improvement of $5.0k (12%, 95% CI $3.6k to $6.4k). For scale, the middle half of posted midpoints spans $170k to $250k. Title words (family and level) carry most of what is predictable; pay bands are set per company and per location policy, which a posting's text only partly reveals.
$36.2k
MAE, best model (company-grouped CV)
$41.2k
MAE, family x seniority median
39%
of predictions within $20k
$28.5k
MAE if CV ignored companies (leaky)

Error by model

Out-of-fold mean absolute error, GroupKFold(5) by company, with 95% intervals from 1000 company-cluster resamples. Lower is better.

Table view

Where the best model misses

MAE by family, best model against the baseline.

What the boosted model leans on

Permutation importance (drop in log-pay MAE when a feature is shuffled) on one held-out company fold, structured features only. Correlated features share credit, so a low score does not mean irrelevant.

Description terms associated with higher pay

Largest Ridge coefficients on TF-IDF terms, controlling for the structured features. Descriptive, not causal: many are company boilerplate.

... and with lower pay

Same model, most negative coefficients.

Largest misses

Out-of-fold predictions furthest from the posted midpoint. Usually a company paying far above or below the market for the title, or a range parsed from text that was not the base salary.

How it was built

Generated 2026-09-29T14:55:46+00:00 in 87.7s. Code: observatory/ml.py.