Working With Labor Statistics Data Isn't as Clean as You'd Think

Most people who come across the States Department Of Labor Bureau Of Labor Statistics first see the published tables and think they're looking at raw fact. They're not. They're looking at estimates derived from surveys that have been adjusted for seasonal patterns, chain-weighted, and sometimes revised multiple times before the data ever reaches the public. I spent about five years pulling and cleaning BLS data for workforce modeling projects. What I learned is that the surface-level numbers are only the starting point. The Bureau of Labor Statistics collects data on employment, unemployment, wages, pricing, and workplace injuries across the United States. It runs several major surveys simultaneously. The Current Population Survey, commonly called the household survey, captures the unemployment rate from roughly 60,000 households each month. The Current Employment Statistics program, often called the establishment survey, pulls payroll data from about 145,000 businesses and government agencies. Then there's the American Community Survey, the Occupational Employment and Wage Statistics program, and various price indexes including the Consumer Price Index. Each of these measures different things and produces different numbers. Confusing them is one of the most common mistakes I see. A lot of people read the unemployment rate from the household survey and the job growth number from the establishment survey in the same month and expect them to reconcile. They don't have to. The household survey counts people who identify as unemployed based on their own reporting. The establishment survey counts jobs based on employer reports. A person can hold two jobs and appear twice in the establishment data but once in the household data. Someone can be employed but not counted if they're not in the labor force according to their own definition. These surveys overlap in purpose but not in methodology, and they never will be designed to match each other perfectly.

How to Actually Use the Data Without Getting Tripped Up

If you need labor statistics for a project or report, the first thing you should do is decide which survey actually answers your question. The household survey is better for understanding labor force participation, unemployment duration, and demographic breakdowns. The establishment survey is better for tracking industry-level employment changes and average weekly hours. The OES program is the source you want for occupational wage data broken down by Metropolitan Statistical Area. Don't mix them. I ran into a specific problem last year where a client needed historical wage data for a set of occupations going back twenty-five years. I pulled from the OES tables and noticed something odd. The data showed a sharp structural break in 2001 where certain occupation codes disappeared entirely and reappeared under different titles. That was the census year when the BLS revised its SOC classification system. If you just concatenate the series without accounting for that revision, your trend lines will look like they dropped by half overnight. The workaround is straightforward but tedious. You have to use the BLS concordance files that map old SOC codes to new SOC codes. There's no automated tool that does this cleanly. You build a mapping table in a spreadsheet, merge on the code, and then reconstruct the time series manually. It took about three days to get the full twenty-five-year series right. Another issue that bites people frequently is seasonal adjustment. The BLS releases seasonally adjusted and non-seasonally adjusted data for most of its monthly indicators. Seasonally adjusted data removes predictable patterns like holiday hiring spikes or agricultural cycles. Non-seasonally adjusted data shows the raw count. If you're doing year-over-year comparisons, seasonally adjusted is usually what you want. If you're comparing January to February for retail employment, you absolutely need the unadjusted numbers because the seasonal adjustment factor itself is an estimate and can introduce noise in short windows. I've seen analysts use seasonally adjusted data for month-to-month change analysis and then wonder why their results looked jittery. It was the adjustment factor moving around, not actual employment shifting.

Where the Data Falls Short

The BLS is good at what it does, but it has real limitations. The household survey has a margin of error that becomes significant at the state level for smaller labor markets. If you're looking at Wyoming or Vermont, the unemployment rate can swing by half a percentage point or more purely from sampling variation. The establishment survey excludes self-employed workers, unpaid family workers, and employees of private households. That's a meaningful chunk of the economy in certain sectors. Gig economy work is still poorly captured in both surveys, though the BLS has been running a supplementary survey on alternative work arrangements since 2017 and the coverage is slowly improving. Revisions are another issue. The establishment survey gets three rounds of revision after the initial release. The first revision comes out a month later. The second comes out two months after that. After a year, the data is considered final. If you're building models that depend on labor market figures, you need to account for the fact that the headline number you see on release day is almost certainly wrong. I once saw a policy brief published with employment figures from the day of the initial release that had to be completely rewritten six months later when the revisions settled. The direction of the trend hadn't changed, but the magnitude was off by nearly twenty percent in a couple of industries. That's not unusual. The price indexes have their own set of problems. The CPI-U measures a fixed basket of goods and services, which means it struggles to capture quality improvements and substitution behavior in real time. The BLS uses hedonic adjustment for certain product categories like electronics, but that only covers a subset of items. Chain-weighted indexes like the CPI-U-RS attempt to address some of this, but they are less widely referenced and the methodology is harder to explain to a general audience.

Get the Full Details

The Bureau of Labor Statistics Must Be Adequately Funded To Preserve ...
The Bureau of Labor Statistics Must Be Adequately Funded To Preserve ...

Practical Steps for Getting the Data Yourself

The primary entry point is laborstats.gov. From there you can access the data through the BLS API, which is free and well-documented if you know how to use it. The API returns JSON or XML and lets you query by series ID, time range, and frequency. I typically write a Python script that pulls the series IDs I need, handles pagination, and writes the output to a CSV. The script takes about forty lines and runs in under a minute for most queries. If you're pulling data for hundreds of series across multiple years, it can take longer, but the API has generous rate limits compared to most government sources. For state-level data, the Quarterly Census of Employment and Wages is probably the most useful dataset if you need it. It covers all employers subject to unemployment insurance law and provides employment and wages by quarter, industry, and county. The catch is that it's released with a lag of about four months and gets revised annually when the QCEW syncs with administrative records. If you need real-time state data, you're better off with the Local Area Unemployment Statistics program, but those estimates have wider margins of error. I also recommend downloading the metadata documentation that accompanies each dataset. The technical notes explain how vacancies are counted in the JOLTS survey, how the CPS defines unemployed versus not in the labor force, and how the OES handles outliers. Without that context, you're making decisions about your data based on incomplete assumptions. I've reviewed analysis produced by consulting firms that mischaracterized the JOLTS quit rate as a measure of worker confidence when it's actually a measure of voluntary separations from employers who are still hiring. The distinction matters more than people realize.

If you're working on something large scale, consider using the BLS Integrated Microdata System. It gives you access to microdata from the CPS and OES surveys under controlled conditions. You don't get the raw person-level records, but you do get anonymized data that lets you run custom cross-tabulations and regressions without submitting a research application. The turnaround time is about two weeks. For most one-off projects it's overkill, but if you're doing repeated custom analysis it saves a lot of time compared to repeatedly querying the public API and trying to aggregate from published tables.