Not All Job Data Sources Are Created Equal
If you are building an HR tech product, training a labor market model, or analyzing workforce trends, you need job market data. But most vendors will not tell you this: the quality gap between the best and worst sources is enormous, and aggregate metrics hide the problem.
We ingest data from over 200,000 employer career portals and every major job board. Our pipeline processes 10,003,672,873 raw observations into 1,011,115,180 job postings across 100+ fields, which cross-source deduplication then resolves into 433,666,433 unique jobs. Along the way, we have learned exactly how data quality varies across source types. That knowledge shapes everything from our deduplication logic to our NLP enrichment priorities.
The Source Landscape: Three Tiers
Our pipeline ingests from three broad categories. Here is how volume breaks down in the final delivery table:
Tier 1: Major Job Boards. Indeed contributes approximately 225.9 million unique jobs, making it the single largest source. LinkedIn follows with 176.4 million. A major professional job feed (PJF) contributes 215.7 million.
Tier 2: Direct ATS Feeds. Career portals hosted on Workday, Greenhouse, Lever, and SmartRecruiter. Workday alone accounts for over 11.8 million records. Across 200,000+ employer portals, ATS feeds collectively represent a substantial and growing share.
Tier 3: Aggregators. Google Jobs, Jobs2Careers, Jora, and similar services re-aggregate data from other sources. The volume hierarchy is clear: major boards dominate by raw record count. But volume is not quality. In fact, there is often an inverse relationship. The sources with the highest raw counts also tend to have the highest duplication rates and the lowest per-record enrichment quality.
Company Name Coverage: Where Extraction Gaps Hide
Company name is a foundational field. Without it, you cannot do company-level analytics, match to firmographic databases, or build employer-specific trend data.
The quality gap across sources is stark. Indeed, LinkedIn, and most major boards deliver company names on the vast majority of records. Greenhouse, Lever, and established ATS platforms reliably include employer identity.
Then there are the gaps:
| Source | Company Name Null Rate | Postings |
|---|---|---|
| Jora | 12.51% | 79.7M |
| SimplyHired | 4.08% | 119.4M |
| CareerBuilder | 2.20% | 59.1M |
| 0.99% | 198.4M | |
| PJF | 0.29% | 249.2M |
| Indeed | 0.24% | 256.6M |
| iCIMS / Lever | 0.02% | 3.5M |
| MyWorkdayJobs | 0.00% | 15.4M |
| Overall delivery | 1.93% | 1,011.1M |
MyWorkdayJobs reads 0.00% because company identity is recovered from the URL subdomain rather than the structured field, which is exactly the point: the null rate you get depends on how many signals the vendor bothers to mine. The remaining gaps concentrate in regional aggregators that syndicate listings without attribution. The lesson for data buyers: always ask whether company names are extracted from all available signals, not just the structured field in the raw feed.
Description Length Drives Enrichment Quality
Job description length is the single best predictor of downstream enrichment quality. Longer descriptions give our NLP models more signal for SOC classification, skills extraction, and seniority detection.
| Description Length | Skills Extraction Rate |
|---|---|
| Under 50 words | 47.5% |
| 50 to 100 words | ~65% |
| 100 to 200 words | 81.6% |
| 200 to 500 words | ~95% |
| 500+ words | 99.5% |
Our overall average across 599.3 million processed descriptions is 486 words, comfortably above the 200-word threshold where NLP enrichment becomes reliable. But this average masks significant variation by source.
ATS feeds, particularly Greenhouse and Lever postings, tend to include detailed descriptions with bulleted requirements and qualifications. These structured descriptions are optimal for NLP extraction. Job board descriptions are more variable. Indeed reformats employer descriptions, sometimes truncating them. Aggregators often serve shortened versions.
The practical implication: a dataset of 100 million Greenhouse postings with 500-word average descriptions will yield better enrichment than 200 million aggregator postings averaging 150 words. For details on how our NLP enrichment system handles varying input quality, see the linked post.
Deduplication Rates Tell the Real Story
Our pipeline uses semantic deduplication combining vector similarity, MinHash company matching, title similarity modeling, and geo-clustering with graph-based transitive resolution. You can read more about the approach in our methodology. The dedup rates by source tell you a lot about what you are actually buying:
| Source | Job Postings | Unique Jobs | Cross-Source Collapse |
|---|---|---|---|
| PJF | 249.2M | 85.7M | 65.6% |
| SimplyHired | 119.4M | 41.8M | 65.0% |
| 198.4M | 78.3M | 60.5% | |
| Jobs2Careers | 15.1M | 6.4M | 57.7% |
| Indeed | 256.6M | 128.7M | 49.9% |
| MyWorkdayJobs | 15.4M | 11.1M | 27.8% |
| iCIMS | 2.4M | 2.4M | 2.6% |
| Dayforce | 670K | 661K | 1.3% |
| UltiPro | 952K | 942K | 1.1% |
Roughly 90% of raw observations are re-observations, and a further 57% of the surviving postings are cross-source duplicates. If your vendor counts "job postings" without disclosing dedup methodology, you have no idea how many unique jobs you are actually getting. Our own Indeed feed is a fair illustration: 256.6 million postings that resolve to 128.7 million unique jobs.
The split is clean: aggregators lose half to two thirds of their volume to a canonical record held elsewhere, while employer-operated ATS portals lose almost nothing because they are the origin. If you are paying per record, that ratio is the difference between buying jobs and buying copies. Workday sits in the middle at 27.8% precisely because its largest tenants also syndicate to the boards. This kind of source-specific behavior is invisible in aggregate metrics. Understanding it requires source-level quality tracking.
NLP Enrichment Is Not Uniform Across Sources
Classification coverage is now essentially flat across sources, which was not always true and is worth checking in any vendor you evaluate.
| Source | SOC Code | Seniority |
|---|---|---|
| CareerBuilder | 100.00% | 99.46% |
| Jora | 99.95% | 99.72% |
| Jobs2Careers | 99.95% | 99.78% |
| PJF | 99.93% | 99.69% |
| SimplyHired | 99.89% | 99.61% |
| MyWorkdayJobs | 99.89% | 99.74% |
| Lever | 99.86% | 99.61% |
| 99.85% | 99.69% | |
| Indeed | 99.82% | 99.59% |
| iCIMS | 99.81% | 99.61% |
Flat coverage is a property of running enrichment retroactively over the whole archive rather than only on new intake. A vendor whose classification coverage varies by source or by year is showing you a backfill gap: the model works, it just has not reached everything yet. That is a scheduling problem rather than a quality problem, but it is your problem if you are the one filtering on the field.
Where sources still differ sharply is extraction, which depends on the employer's text rather than on our schedule. Skills coverage runs from 74.9% on PJF to 95.6% on CareerBuilder.
The question to ask your vendor: what percentage of records have SOC codes, and how does that vary by source, vintage, and region? For a detailed breakdown, see our datasets page.
Salary Coverage: A Structural Transformation
Salary data availability has shifted dramatically since 2021, driven by US state transparency laws:
| Year | Salary Coverage (stated) | Key Events |
|---|---|---|
| 2022 | 31.4% | Colorado law in effect, NYC law takes effect Nov 2022 |
| 2023 | 20.4% | CA, WA, NY State laws take effect |
| 2024 | 20.3% | Hawaii law takes effect |
| 2025 | 23.0% | Illinois, Minnesota laws take effect |
The counterintuitive decline from 2022 to 2023 reflects a source composition shift, not a data quality failure. One of our largest feeds had near-100% salary coverage in 2022 but experienced a format change. When we isolate Indeed specifically, salary coverage improved from 51.1% in 2022 to 88.8% in 2023.
This is a critical lesson: aggregate metrics can obscure source-level trends. A single source's format change can swing overall coverage by 10+ percentage points. For sources without stated salary, our prediction model fills the gap when state, zip code, and SOC code are available.
The ATS Advantage
Across every quality metric, direct ATS feeds consistently outperform aggregated sources:
- Employer-verified content. No reformatting, no truncation, no re-aggregation errors.
- Near-zero internal duplication. Compare to Indeed's 89% or Google Jobs' pathological duplication.
- Structured metadata. ATS platforms enforce fields like location, department, and salary range.
- Freshness guarantees. ATS feeds update in near-real-time. Board scrapers may lag by days or weeks.
This is why coverage of 200,000+ employer ATS portals is a strategic asset, not just a volume play. These feeds provide ground truth that calibrates NLP models and validates deduplication logic.
What Data Buyers Should Ask
If you are evaluating job market data providers, these five questions will separate the serious vendors from the ones selling volume:
- Source composition. 500 million records from Google Jobs is not comparable to 500 million from ATS feeds.
- Source-level metrics. Null rates, description lengths, and dedup rates should be available per source, not just in aggregate.
- Dedup methodology. Hash-based dedup catches exact duplicates. Semantic dedup catches near-duplicates. The difference can be 20 to 30 percentage points. See our glossary for definitions.
- Freshness per source. A provider might have fresh Indeed data but stale ATS feeds.
- Enrichment depth. A fully enriched record with normalized title, SOC code, seniority, skills, and predicted salary is worth more than ten raw records with just title and location.
To see how this looks in practice, explore source-level quality metrics with a sample dataset, or compare raw and enriched records side by side in our comparison tool.
Want to see the data for yourself?
Get a free sample of 5,000 enriched job records.