The Dataset at a Glance
1,011,115,180 job postings. Each one is a unique link, distilled from over 10 billion raw scrape observations. Cross-source semantic deduplication then resolves those postings into 433,666,433 unique jobs, because one real opening is typically posted to several places at once. How that dedup works is a story in itself, covered in Deduplication Is Harder Than You Think.
22 sources. These range from the two largest global job boards (Indeed and LinkedIn) to employer ATS platforms (Workday, Greenhouse, Lever, iCIMS, SmartRecruiter) to regional aggregators and specialized feeds.
100+ fields per record. Coverage varies by field, source, and data vintage, but the schema is consistent. Every record passes through the same enrichment pipeline. Full field definitions are in the data schema.
Five years of continuous coverage: 2022 through 2026, with volume growing every year.
Source Composition
Not all job data is created equal. The top sources by volume:
| Source | Postings | Share | Unique Jobs |
|---|---|---|---|
| Indeed | 256.6M | 25.4% | 128.7M |
| PJF (aggregator) | 249.2M | 24.6% | 85.7M |
| 198.4M | 19.6% | 78.3M | |
| SimplyHired | 119.4M | 11.8% | 41.8M |
| Jora | 79.7M | 7.9% | 36.0M |
| CareerBuilder | 59.1M | 5.9% | 31.5M |
| MyWorkdayJobs | 15.4M | 1.5% | 11.1M |
| Jobs2Careers | 15.1M | 1.5% | 6.4M |
| Google Jobs | 6.6M | 0.7% | 3.9M |
| SmartRecruiter | 3.4M | 0.3% | 3.1M |
| iCIMS | 2.4M | 0.2% | 2.4M |
The third column is what a source contributes after cross-source deduplication, and it reorders the table. Indeed contributes 25.4% of postings but 29.7% of unique jobs, while PJF contributes 24.6% of postings and only 19.8% of unique jobs.
The dataset is not dependent on any single source. Indeed is the largest contributor at 25.4%, but the top three sources each contribute meaningful volume. If any single source degrades or restricts access, the dataset does not collapse.
The mix includes both aggregator platforms (Indeed, LinkedIn, SimplyHired) and direct employer ATS feeds (Greenhouse, Lever, Workday/MyWorkdayJobs, iCIMS). ATS feeds are smaller in volume but higher in data quality, representing first-party employer data rather than re-aggregated scrapes. For a detailed comparison of what that quality difference looks like in practice, see ATS vs. Job Boards: A Data Quality Comparison.
Year-Over-Year Growth
| Year | Postings | YoY Growth |
|---|---|---|
| 2022 | 172.7M | baseline |
| 2023 | 198.1M | +14.7% |
| 2024 | 226.8M | +14.5% |
| 2025 | 271.7M | +19.8% |
| 2026 | 141.7M | partial year |
Volume has grown every year, and 2025 grew fastest at +19.8%. This reflects both expanding source coverage and increasing labor market activity.
Field Coverage: The 85+ Field Anatomy
Identity and Core Fields
| Field | Coverage | Notes |
|---|---|---|
| Normalized title | 99.71% | Raw titles canonicalized via ML model |
| Company name | 98.07% | Gaps concentrated in two ATS sources |
| Country | 99.23% | Parsed from raw location strings |
| State | 97.19% | Parsed or inferred from the location string |
| City | 97.67% | Metro areas and multi-location strings are the main gap |
| Zipcode | 91.94% | Often inferred when not stated |
The 99.71% title normalization rate means about 2.9 million records out of 1,011,115,180 lack a normalized title. This is the single most impactful enrichment for downstream analysis. Without it, "Sr. Software Eng II," "Senior Software Engineer," and "Snr SW Developer" are three different things. With it, they are one. The methodology page covers how normalization works, and Inside Our NLP Enrichment System walks through the full stack of models behind classification and extraction.
Classification Fields
| Field | Coverage | What It Does |
|---|---|---|
| SOC code | 99.88% | 6-digit occupational classification (BLS standard) |
| Seniority | 99.64% | Entry / Mid / Senior / Lead / Executive |
| Employment type | 99.64% | Full-time / Part-time / Contract / Temporary |
| Remote/work mode | 99.64% | Remote / Hybrid / On-site |
Sources supply almost none of this natively; the classification layer produces it on essentially every record. Coverage is not accuracy: these figures say a field was populated, and the published accuracy benchmarks for each model live on the methodology page. Every classification also ships a per-field confidence score, so you can set your own threshold rather than inherit ours.
Remote/work mode coverage represents a structural shift that is worth understanding historically. Before 2020, virtually no postings carried a structured remote field because the concept barely existed. The 2020-2022 period was a messy transition where "remote" appeared in descriptions but not as structured data. By 2023+, work mode is a standard field. The full timeline is in Remote Work by the Numbers: 2020-2025.
Extraction Fields
| Field | Coverage | Scale |
|---|---|---|
| Skills (technical) | 86.63% | 40,000+ skills in taxonomy |
| Soft skills | 67.83% | 260+ categories |
| Benefits | 55.01% | Health, dental, 401k, PTO |
| Certifications | 21.54% | 3,400+ certifications tracked |
The 86.63% skills coverage means 876.0 million records have structured, taxonomy-mapped skill arrays. This is high-speed dictionary matching against a 40,000+ taxonomy followed by contextual relevance filtering, not keyword search. "Java" in a barista posting gets filtered out. "Java" in a backend engineer posting gets kept. Certification coverage at 21.54% reflects reality: most postings do not require specific certifications. Among roles that do (nursing, IT security, finance, trades), extraction rates are substantially higher. The full extraction pipeline is described in 100,000 Skills: Extracting Signal from Noise.
Salary
Employer-stated salary sits at 7.66% (77.5 million records) as a structured field, and description-text extraction recovers another 30.22% (305.6 million records) for 37.9% stated in total. This number is rising rapidly due to US salary transparency laws: Colorado (2021), NYC (2022), California and Washington (2023), with continued expansion through 2025. The year-by-year impact of these laws on data availability is tracked in The Salary Transparency Data Shift.
For records without stated salary, a prediction model trained on 50M+ observations provides estimates where sufficient context exists (valid state, zipcode, and SOC code required). The model returns no prediction rather than a bad one when prerequisites are missing. Details on how the model works are in Salary Prediction from 50 Million Observations.
What the Data Reveals
Healthcare Dominates Posting Volume
The top normalized titles are overwhelmingly healthcare:
| Rank | Title | Postings |
|---|---|---|
| 1 | Medical Surgical Travel Nurse | 9.8M |
| 2 | Travel Nurse | 7.8M |
| 3 | Telemetry Travel Registered Nurse | 7.5M |
| 4 | Delivery Driver | 6.8M |
| 5 | Travel Physical Therapist | 5.5M |
| 6 | Licensed Practical Nurse (LPN) | 5.5M |
| 7 | Registered Nurse (RN) | 5.2M |
| 8 | Assistant Manager | 5.0M |
| 9 | Labor and Delivery Travel Registered Nurse | 4.7M |
| 10 | Cashier | 4.3M |
Seven of the top ten titles are nursing or allied health roles. This reflects the structural healthcare labor shortage since 2020: aging population, pandemic-accelerated burnout, and geographic maldistribution of providers. Travel nursing dominates because facilities compete for a mobile workforce, generating high posting volumes as contracts turn over every 8-13 weeks.
Source Quality Is Not Uniform
| Source | Skills Coverage | Benefits Coverage | Salary Coverage |
|---|---|---|---|
| CareerBuilder | 95.6% | 51.9% | 97.3% |
| Indeed | 93.5% | 59.4% | 91.7% |
| SimplyHired | 93.4% | 62.2% | 97.9% |
| 91.3% | 54.2% | 90.7% | |
| MyWorkdayJobs | 89.7% | 53.0% | 63.0% |
| iCIMS | 87.2% | 62.1% | 79.5% |
| Lever | 85.9% | 51.1% | 56.7% |
| PJF | 74.9% | 50.8% | 78.9% |
Extraction quality tracks description quality, and description quality varies by 20 points across sources. PJF aggregator listings yield structured skills on 74.9% of postings while CareerBuilder reaches 95.6%, because aggregators frequently truncate or restate the employer's original text. Salary coverage splits differently: it is highest where the prediction model has the location and occupation context it needs, which is why direct ATS feeds like Lever (56.7%) trail the aggregators despite carrying richer descriptions.
Source provenance is preserved on every record, so source-aware analysis is straightforward. The data schema documents how source metadata is structured.
Coverage by Data Vintage
Coverage rates are not static, and the pattern is not the one most buyers expect. The classification layer is re-run across the whole archive, so a 2022 record gets the same SOC, seniority, and work-mode treatment as a 2026 one. What genuinely varies by vintage is what the employer wrote.
| Field | 2022 | 2023 | 2024 | 2025 | 2026 |
|---|---|---|---|---|---|
| Salary (stated) | 43.0% | 38.5% | 35.7% | 47.4% | 41.3% |
| Work mode | 99.6% | 99.7% | 99.7% | 99.7% | 99.4% |
| Seniority | 99.6% | 99.7% | 99.7% | 99.7% | 99.4% |
| SOC code | 100.0% | 100.0% | 99.9% | 100.0% | 99.3% |
| Skills | 93.5% | 91.7% | 71.0% | 93.0% | 84.0% |
The classification rows are flat because enrichment is retroactive. The two rows that move are the ones driven by source text: stated salary tracks the transparency-law rollout and the source mix posting in a given year, and skills coverage dips in 2024 where a larger share of that year's intake came from aggregators that truncate descriptions. If a vendor shows you rising classification coverage by year, they are showing you a backfill gap, not a data improvement. The glossary defines terms like "data vintage" and "structural break" that are relevant here.
The Quality Signal
Job market data quality is not a single number. It is a matrix: field by source by vintage by geography. A dataset that provides 97.19% state coverage also has corners where skills extraction drops to 74.9% (PJF) or salary coverage to 56.7% (Lever). A 2024 record has different expected skills coverage than a 2025 one.
The value of a deeply enriched dataset is that coverage is measured, source provenance is preserved, confidence scores are available, and the enrichment pipeline applies consistently across every record. That consistency is what makes the data usable for rigorous analysis. You can explore available datasets and compare what is included at datasets or see pricing at pricing.
1B+ job postings. 400M+ unique jobs. 22 sources. 100+ fields. Measured coverage on every one.
To see how this looks on actual records, request a sample.
Want to see the data for yourself?
Get a free sample of 5,000 enriched job records.