Narrative of the audit/repair work plus a practical usage guide for the client, query helpers, series builders, QCEW module, caching/errors/retry, testing, and coverage. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
12 KiB
BLS Data Library — Overview & Usage Guide
Repository: giteaadmin/bls-data on gitea.doesworks.net
Status: feature-complete · 96/96 query helpers verified live · 131 tests · pip-installable
Last updated: June 22, 2026
This document covers two things: what the library is and how it got here, and a
practical guide to using it. For the terse reference, see README.md,
USAGE.md, and series_id_formats.md.
1. What this is
A dependency-light Python client for the U.S. Bureau of Labor Statistics (BLS) API.
Its job is to make BLS data retrievable without hand-encoding series IDs — the
cryptic 11-to-25-character codes BLS uses as primary keys (e.g.
OEUN000000000000015125213 for "national median annual wage, software developers").
The library wraps two BLS data services:
| Service | What it serves | Key? | Quota? |
|---|---|---|---|
Timeseries API (api.bls.gov/publicAPI/v2) |
Most surveys — unemployment, payrolls, prices, wages, openings, productivity | yes (free) | 500/day |
QCEW Open Data (data.bls.gov/cew/data/api) |
County- and industry-level employment & wages | no | none |
It provides: a batching/caching HTTP client, ~96 pre-built "query helpers" for common series, low-level series-ID builders for everything else, typed errors, retry/backoff, and a test suite that locks the encodings in place.
2. What we did (engineering summary)
The library existed but was substantially broken — a full audit found only 58 of 82 query helpers returned live data. Whole survey modules emitted series IDs that BLS rejected. The work fell into four phases:
Phase 1 — Fix the broken encodings
Every dead helper traced back to a wrong series-ID encoding (mostly wrong field widths inherited from an internally-inconsistent reference doc). Fixed and verified against the live API:
| Survey | Bug | Fix |
|---|---|---|
| JOLTS | IDs were 18 chars (12 zeros) | correct format is 21 chars / 15 zeros (added state+area+sizeclass fields) |
| OES | national area code hard-coded 0000400 |
corrected to 0000000 |
| ECI | wages/benefits put as the trailing char | trailing char is the estimate code; component is a mid-string field; published unadjusted (CIU) |
| Productivity | 2-digit sector + 3-digit measure | real format is 4-digit each (e.g. PRS85006092) |
| PPI | "finished goods" index (discontinued by BLS) | repointed to final demand (WPUFD4) |
| SOC code | software developers 151132 (2010 SOC, retired) |
151252 (2018 SOC) |
Result: 82/82 helpers live.
Phase 2 — Make it shareable
- API key moved to the
BLS_API_KEYenvironment variable (config file gitignored; the key was never in git history). - Added
requirements.txt; rewroteREADME.mdas a proper project front door.
Phase 3 — Package + test
pyproject.toml— the library is nowpip install-able; builds a wheel and sdist.- A pytest suite that locks every series-ID encoding against known-good IDs (offline,
no quota) plus a live smoke test behind an opt-in
-m livemarker. This is the regression guard — it fails instantly if an encoding is ever re-broken.
Phase 4 — Complete the coverage
Three gaps closed in one efficient pass (total cost: ~2 of the 500 daily API queries, because the decode tables and QCEW data are off-quota and ECEC verification batched into a single call):
- QCEW county/industry detail — a new
qcewmodule over the separate CSV service, reaching the per-county, per-industry data the timeseries API can't. - ECEC benefit breakdown — health insurance, retirement, paid leave, etc., built
from BLS's authoritative decode table (not guessed). This also caught a latent bug:
ecec_total_benefitshad pointed at an education/health-industry series rather than the all-civilian total. - Caching + typed errors — an on-disk response cache (so reruns don't re-spend
quota) and a
BLSQuotaErrorthat distinguishes "you're throttled" from "your series ID is wrong" — the exact ambiguity that had previously stalled debugging.
Plus a retry/backoff layer for transient network failures (timeouts, 429/5xx),
which never retries quota rejections or 4xx.
End state: 96/96 helpers live, 131 tests, every survey at full useful depth.
3. Architecture
bls_client/
├── client.py BLSClient — batching, caching, retry, named fetches, row flattening
├── series.py low-level series-ID builders (every survey)
├── qcew.py QCEW Open Data CSV client (county/industry; no key, no quota)
├── cache.py on-disk response cache (FileCache)
├── retry.py exponential-backoff retry for transient network errors
├── errors.py BLSError / BLSQuotaError / BLSRequestError
└── queries/
├── employment.py payrolls, unemployment (LAUS), JOLTS, QCEW totals
├── prices.py CPI, PPI, average prices, import/export prices
├── wages.py OES occupational wages, ECI, ECEC breakdown
└── productivity.py major-sector productivity & costs
Two layers: query helpers (queries/…) return ready-to-use series IDs for common
asks; series builders (series.py) construct any valid ID from parameters when a
helper doesn't exist.
4. Install & configure
pip install -r requirements.txt # just `requests`
# or, to install the package itself:
pip install -e .
Get a free API key (instant): https://data.bls.gov/registrationEngine/ — 500 queries/day.
cp config.example.py config.py
export BLS_API_KEY="your-key" # config.py reads this env var; takes precedence
config.py is gitignored. You can also pass the key directly: BLSClient("your-key").
Note: on a PEP-668 "externally managed" Python (e.g. recent Ubuntu), install into a virtualenv:
python3 -m venv .venv && . .venv/bin/activate && pip install -e ".[dev]".
5. Usage
The client
from bls_client import BLSClient
client = BLSClient(API_KEY, cache=True) # caching on (recommended)
# Fetch by series ID(s) — auto-batches >50 series
data = client.fetch(["CES0000000001"], 2020, 2025)
# Most recent N years
data = client.fetch_latest("CES0000000001", years=1)
# Labeled dict in, labeled dict out
data = client.fetch_named({"Payrolls": "CES0000000001"}, 2024, 2025)
# Helpers for reading results
obs = BLSClient.latest_obs(next(iter(data.values()))) # most recent observation
rows = BLSClient.to_rows(data) # flat rows for CSV/pandas
Query helpers (the common case)
from bls_client.queries import employment, prices, wages, productivity
employment.nonfarm_payrolls() # headline jobs number
employment.dc_region_unemployment() # DC/MD/VA + 3 metros (labeled dict)
employment.jolts_dashboard() # openings/hires/quits/layoffs/separations
prices.cpi_dashboard() # CPI: all-items, core, food, energy, …
prices.ppi_final_demand()
wages.occupation_annual_median_wage("151252") # any SOC code (software devs here)
wages.ecec_dashboard() # full employer cost-per-hour breakdown
wages.eci_total_compensation()
productivity.productivity_dashboard() # output/hr, ULC, etc.
Putting it together:
client = BLSClient(API_KEY, cache=True)
result = client.fetch_named(wages.ecec_dashboard(), 2024, 2024)
for label, series in result.items():
obs = BLSClient.latest_obs(series)
print(f"{label:22} ${obs['value']}/hr")
# Total Compensation $47.20/hr
# Health Insurance $3.54/hr
# Retirement & Savings $2.45/hr ...
Series builders (anything not pre-built)
from bls_client import series
series.laus_state(48, "rate") # Texas unemployment rate
series.ces_national("manufacturing") # manufacturing payrolls
series.cpi("gasoline", seasonal=True)
series.oes_national("291141", data_type="annual_median") # registered nurses
series.ecec("health_insurance", owner="private")
QCEW — county & industry detail
The timeseries helpers give QCEW national/state totals; the qcew module reaches the
full detail via the separate CSV service (no key, no quota):
from bls_client import qcew
rows = qcew.area("11000", 2024, "a") # everything for DC, annual averages
dc_private = qcew.filter_rows(rows, own_code="private", industry_code="10", agglvl_code="51")
# -> establishments, employment, total wages, avg weekly wage for DC private sector
hospitals = qcew.industry("622", 2024, "a") # one NAICS industry across all areas
Each call returns a list of dict rows (the CSV columns). Quarter is 1-4 or "a" for
annual. Area FIPS: US000 national, 11000 = DC, 11001 = a county, C#### = metro.
Caching, errors, and retry
client = BLSClient(API_KEY, cache=True, retries=2, backoff=0.5)
client.fetch(...) # identical requests served from ~/.cache/bls (1-day TTL)
client.queries_used # network calls this session (cache hits excluded)
from bls_client import BLSQuotaError, BLSRequestError
try:
client.fetch(ids, 2024, 2025)
except BLSQuotaError: # daily 500-query limit hit (resets midnight ET)
...
except BLSRequestError: # malformed request / bad series ID
...
Transient failures (timeouts, dropped connections, 429/5xx) retry automatically with
exponential backoff and honor a Retry-After header. Quota rejections and 4xx are never
retried. Set retries=0 to disable.
6. Coverage
| Survey | What's available |
|---|---|
| LAUS — local area unemployment | state / metro / county rates; DC-region dashboard |
| CES — payroll employment | national by supersector; state/metro |
| CPS — household survey | national unemployment rate, participation |
| JOLTS — job openings & turnover | openings, hires, quits, layoffs, separations; dashboard |
| CPI — consumer prices | all-items, core, food, energy, gasoline, shelter, …; dashboard |
| PPI — producer prices | all commodities, final demand, food, energy |
| OES — occupational wages | employment + wage percentiles by SOC; 15 named occupations |
| ECI — employment cost index | total comp / wages / benefits × civilian/private/gov |
| ECEC — employer cost levels | full breakdown: comp, wages, benefits, paid leave, supplemental, health insurance, retirement, legally-required |
| Productivity & costs | output/hr, unit labor cost, compensation, hours × business/nonfarm/manufacturing |
| QCEW — quarterly census | national/state totals plus full county/industry detail (CSV module) |
7. Testing
pip install -e ".[dev]"
pytest # 125 offline tests — lock every encoding; no network, no quota
pytest -m live # +live smoke tests against the real API (needs network + key)
The offline tests are the regression guard: they assert each builder/helper produces its exact known-good series ID, so a re-broken encoding fails immediately and for free. The live tests catch the other failure mode — BLS silently retiring a code (as happened with the 2010→2018 SOC change).
8. Known limitations
- Coverage is the headline cut of each survey, not an exhaustive mirror (CES is national supersectors + a couple of states; CPI is the common items; OES ships 15 named occupations). Any SOC/industry still works through the parameterized builders — the pre-built helper dicts just cover the common asks. Widen them as needed.
- No rate-limit pacing beyond caching + retry; a heavy batch can still approach the
500/day quota (the
queries_usedcounter andBLSQuotaErrormake that visible).
9. Change history
| Commit | What |
|---|---|
4fe7343 |
Fixed all broken series-ID encodings (JOLTS/OES/ECI/productivity/PPI/SOC); env-var config; project README |
881d62c |
pyproject.toml packaging + pytest suite |
9e2c55c |
QCEW CSV module, ECEC benefit breakdown, response caching + typed errors |
0d0d6da |
Transient-network retry with exponential backoff |