Can a listing predict its own posted pay?
A model that predicts the posted pay midpoint from the role family, seniority, skills mentioned, experience and degree asks, and the description text. 2,416 priced listings from 847 companies, posted in USD, any location.
Error by model
Out-of-fold mean absolute error, GroupKFold(5) by company, with 95% intervals from 1000 company-cluster resamples. Lower is better.
Table view
Where the best model misses
MAE by family, best model against the baseline.
What the boosted model leans on
Permutation importance (drop in log-pay MAE when a feature is shuffled) on one held-out company fold, structured features only. Correlated features share credit, so a low score does not mean irrelevant.
Description terms associated with higher pay
Largest Ridge coefficients on TF-IDF terms, controlling for the structured features. Descriptive, not causal: many are company boilerplate.
... and with lower pay
Same model, most negative coefficients.
Largest misses
Out-of-fold predictions furthest from the posted midpoint. Usually a company paying far above or below the market for the title, or a range parsed from text that was not the base salary.
How it was built
- Target: the midpoint of the posted range, annual USD, for canonical listings posted in USD. Models fit log(pay); errors are scored in dollars.
- Features: family, seniority, remote-US flag, years asked (with a missing flag), degree flags, 83 skill-mention flags; the text models add TF-IDF unigrams and bigrams (Ridge) or a 100-dimension SVD of them (gradient boosting).
- Validation: GroupKFold(5) by company. Every listing from a company sits in one fold, so the model is always scored on companies it has never seen. Scoring the same model with a random split gives $28.5k instead of $36.2k: the gap is how much a company's shared pay bands and boilerplate would leak.
- Uncertainty: listings inside a company are not independent, so confidence intervals resample whole companies.
- Limits: posted ranges are not offers; wide ranges make the midpoint a blunt target; the sample is venture-backed tech; roles with no posted range (about a quarter) are missing not at random.
Generated 2026-09-29T14:55:46+00:00 in 87.7s. Code: observatory/ml.py.