Files
bls-data/docs/BLS_Library_Overview.md
Dave Boyd 8f0122011a Add project overview & usage guide (docs/BLS_Library_Overview.md)
Narrative of the audit/repair work plus a practical usage guide for the
client, query helpers, series builders, QCEW module, caching/errors/retry,
testing, and coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 10:51:57 -04:00

294 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# BLS Data Library — Overview & Usage Guide
**Repository:** `giteaadmin/bls-data` on `gitea.doesworks.net`
**Status:** feature-complete · 96/96 query helpers verified live · 131 tests · pip-installable
**Last updated:** June 22, 2026
This document covers two things: **what the library is and how it got here**, and a
**practical guide to using it**. For the terse reference, see `README.md`,
`USAGE.md`, and `series_id_formats.md`.
---
## 1. What this is
A dependency-light Python client for the U.S. Bureau of Labor Statistics (BLS) API.
Its job is to make BLS data retrievable **without hand-encoding series IDs** — the
cryptic 11-to-25-character codes BLS uses as primary keys (e.g.
`OEUN000000000000015125213` for "national median annual wage, software developers").
The library wraps two BLS data services:
| Service | What it serves | Key? | Quota? |
|---|---|---|---|
| **Timeseries API** (`api.bls.gov/publicAPI/v2`) | Most surveys — unemployment, payrolls, prices, wages, openings, productivity | yes (free) | 500/day |
| **QCEW Open Data** (`data.bls.gov/cew/data/api`) | County- and industry-level employment & wages | no | none |
It provides: a batching/caching HTTP client, ~96 pre-built "query helpers" for common
series, low-level series-ID builders for everything else, typed errors, retry/backoff,
and a test suite that locks the encodings in place.
---
## 2. What we did (engineering summary)
The library existed but was **substantially broken** — a full audit found only 58 of 82
query helpers returned live data. Whole survey modules emitted series IDs that BLS
rejected. The work fell into four phases:
### Phase 1 — Fix the broken encodings
Every dead helper traced back to a wrong series-ID encoding (mostly wrong field widths
inherited from an internally-inconsistent reference doc). Fixed and verified against the
live API:
| Survey | Bug | Fix |
|---|---|---|
| **JOLTS** | IDs were 18 chars (12 zeros) | correct format is 21 chars / 15 zeros (added state+area+sizeclass fields) |
| **OES** | national area code hard-coded `0000400` | corrected to `0000000` |
| **ECI** | wages/benefits put as the trailing char | trailing char is the *estimate* code; component is a mid-string field; published unadjusted (CIU) |
| **Productivity** | 2-digit sector + 3-digit measure | real format is 4-digit each (e.g. `PRS85006092`) |
| **PPI** | "finished goods" index (discontinued by BLS) | repointed to **final demand** (`WPUFD4`) |
| **SOC code** | software developers `151132` (2010 SOC, retired) | `151252` (2018 SOC) |
Result: **82/82 helpers live.**
### Phase 2 — Make it shareable
- API key moved to the `BLS_API_KEY` environment variable (config file gitignored; the
key was never in git history).
- Added `requirements.txt`; rewrote `README.md` as a proper project front door.
### Phase 3 — Package + test
- `pyproject.toml` — the library is now `pip install`-able; builds a wheel and sdist.
- A **pytest suite** that locks every series-ID encoding against known-good IDs (offline,
no quota) plus a live smoke test behind an opt-in `-m live` marker. This is the
regression guard — it fails instantly if an encoding is ever re-broken.
### Phase 4 — Complete the coverage
Three gaps closed in one efficient pass (total cost: ~2 of the 500 daily API queries,
because the decode tables and QCEW data are off-quota and ECEC verification batched into
a single call):
1. **QCEW county/industry detail** — a new `qcew` module over the separate CSV service,
reaching the per-county, per-industry data the timeseries API can't.
2. **ECEC benefit breakdown** — health insurance, retirement, paid leave, etc., built
from BLS's authoritative decode table (not guessed). This also caught a latent bug:
`ecec_total_benefits` had pointed at an education/health-industry series rather than
the all-civilian total.
3. **Caching + typed errors** — an on-disk response cache (so reruns don't re-spend
quota) and a `BLSQuotaError` that distinguishes "you're throttled" from "your series
ID is wrong" — the exact ambiguity that had previously stalled debugging.
Plus a **retry/backoff layer** for transient network failures (timeouts, `429`/`5xx`),
which never retries quota rejections or 4xx.
**End state:** 96/96 helpers live, 131 tests, every survey at full useful depth.
---
## 3. Architecture
```
bls_client/
├── client.py BLSClient — batching, caching, retry, named fetches, row flattening
├── series.py low-level series-ID builders (every survey)
├── qcew.py QCEW Open Data CSV client (county/industry; no key, no quota)
├── cache.py on-disk response cache (FileCache)
├── retry.py exponential-backoff retry for transient network errors
├── errors.py BLSError / BLSQuotaError / BLSRequestError
└── queries/
├── employment.py payrolls, unemployment (LAUS), JOLTS, QCEW totals
├── prices.py CPI, PPI, average prices, import/export prices
├── wages.py OES occupational wages, ECI, ECEC breakdown
└── productivity.py major-sector productivity & costs
```
Two layers: **query helpers** (`queries/…`) return ready-to-use series IDs for common
asks; **series builders** (`series.py`) construct any valid ID from parameters when a
helper doesn't exist.
---
## 4. Install & configure
```bash
pip install -r requirements.txt # just `requests`
# or, to install the package itself:
pip install -e .
```
Get a free API key (instant): https://data.bls.gov/registrationEngine/ — 500 queries/day.
```bash
cp config.example.py config.py
export BLS_API_KEY="your-key" # config.py reads this env var; takes precedence
```
`config.py` is gitignored. You can also pass the key directly: `BLSClient("your-key")`.
> **Note:** on a PEP-668 "externally managed" Python (e.g. recent Ubuntu), install into a
> virtualenv: `python3 -m venv .venv && . .venv/bin/activate && pip install -e ".[dev]"`.
---
## 5. Usage
### The client
```python
from bls_client import BLSClient
client = BLSClient(API_KEY, cache=True) # caching on (recommended)
# Fetch by series ID(s) — auto-batches >50 series
data = client.fetch(["CES0000000001"], 2020, 2025)
# Most recent N years
data = client.fetch_latest("CES0000000001", years=1)
# Labeled dict in, labeled dict out
data = client.fetch_named({"Payrolls": "CES0000000001"}, 2024, 2025)
# Helpers for reading results
obs = BLSClient.latest_obs(next(iter(data.values()))) # most recent observation
rows = BLSClient.to_rows(data) # flat rows for CSV/pandas
```
### Query helpers (the common case)
```python
from bls_client.queries import employment, prices, wages, productivity
employment.nonfarm_payrolls() # headline jobs number
employment.dc_region_unemployment() # DC/MD/VA + 3 metros (labeled dict)
employment.jolts_dashboard() # openings/hires/quits/layoffs/separations
prices.cpi_dashboard() # CPI: all-items, core, food, energy, …
prices.ppi_final_demand()
wages.occupation_annual_median_wage("151252") # any SOC code (software devs here)
wages.ecec_dashboard() # full employer cost-per-hour breakdown
wages.eci_total_compensation()
productivity.productivity_dashboard() # output/hr, ULC, etc.
```
Putting it together:
```python
client = BLSClient(API_KEY, cache=True)
result = client.fetch_named(wages.ecec_dashboard(), 2024, 2024)
for label, series in result.items():
obs = BLSClient.latest_obs(series)
print(f"{label:22} ${obs['value']}/hr")
# Total Compensation $47.20/hr
# Health Insurance $3.54/hr
# Retirement & Savings $2.45/hr ...
```
### Series builders (anything not pre-built)
```python
from bls_client import series
series.laus_state(48, "rate") # Texas unemployment rate
series.ces_national("manufacturing") # manufacturing payrolls
series.cpi("gasoline", seasonal=True)
series.oes_national("291141", data_type="annual_median") # registered nurses
series.ecec("health_insurance", owner="private")
```
### QCEW — county & industry detail
The timeseries helpers give QCEW national/state totals; the `qcew` module reaches the
full detail via the separate CSV service (no key, no quota):
```python
from bls_client import qcew
rows = qcew.area("11000", 2024, "a") # everything for DC, annual averages
dc_private = qcew.filter_rows(rows, own_code="private", industry_code="10", agglvl_code="51")
# -> establishments, employment, total wages, avg weekly wage for DC private sector
hospitals = qcew.industry("622", 2024, "a") # one NAICS industry across all areas
```
Each call returns a list of dict rows (the CSV columns). Quarter is `1`-`4` or `"a"` for
annual. Area FIPS: `US000` national, `11000` = DC, `11001` = a county, `C####` = metro.
### Caching, errors, and retry
```python
client = BLSClient(API_KEY, cache=True, retries=2, backoff=0.5)
client.fetch(...) # identical requests served from ~/.cache/bls (1-day TTL)
client.queries_used # network calls this session (cache hits excluded)
```
```python
from bls_client import BLSQuotaError, BLSRequestError
try:
client.fetch(ids, 2024, 2025)
except BLSQuotaError: # daily 500-query limit hit (resets midnight ET)
...
except BLSRequestError: # malformed request / bad series ID
...
```
Transient failures (timeouts, dropped connections, `429`/`5xx`) retry automatically with
exponential backoff and honor a `Retry-After` header. Quota rejections and 4xx are never
retried. Set `retries=0` to disable.
---
## 6. Coverage
| Survey | What's available |
|---|---|
| **LAUS** — local area unemployment | state / metro / county rates; DC-region dashboard |
| **CES** — payroll employment | national by supersector; state/metro |
| **CPS** — household survey | national unemployment rate, participation |
| **JOLTS** — job openings & turnover | openings, hires, quits, layoffs, separations; dashboard |
| **CPI** — consumer prices | all-items, core, food, energy, gasoline, shelter, …; dashboard |
| **PPI** — producer prices | all commodities, final demand, food, energy |
| **OES** — occupational wages | employment + wage percentiles by SOC; 15 named occupations |
| **ECI** — employment cost index | total comp / wages / benefits × civilian/private/gov |
| **ECEC** — employer cost levels | full breakdown: comp, wages, benefits, paid leave, supplemental, health insurance, retirement, legally-required |
| **Productivity & costs** | output/hr, unit labor cost, compensation, hours × business/nonfarm/manufacturing |
| **QCEW** — quarterly census | national/state totals **plus** full county/industry detail (CSV module) |
---
## 7. Testing
```bash
pip install -e ".[dev]"
pytest # 125 offline tests — lock every encoding; no network, no quota
pytest -m live # +live smoke tests against the real API (needs network + key)
```
The offline tests are the regression guard: they assert each builder/helper produces its
exact known-good series ID, so a re-broken encoding fails immediately and for free. The
live tests catch the other failure mode — BLS silently retiring a code (as happened with
the 2010→2018 SOC change).
---
## 8. Known limitations
- **Coverage is the headline cut** of each survey, not an exhaustive mirror (CES is
national supersectors + a couple of states; CPI is the common items; OES ships 15 named
occupations). Any SOC/industry still works through the parameterized builders — the
pre-built helper dicts just cover the common asks. Widen them as needed.
- No rate-limit *pacing* beyond caching + retry; a heavy batch can still approach the
500/day quota (the `queries_used` counter and `BLSQuotaError` make that visible).
---
## 9. Change history
| Commit | What |
|---|---|
| `4fe7343` | Fixed all broken series-ID encodings (JOLTS/OES/ECI/productivity/PPI/SOC); env-var config; project README |
| `881d62c` | `pyproject.toml` packaging + pytest suite |
| `9e2c55c` | QCEW CSV module, ECEC benefit breakdown, response caching + typed errors |
| `0d0d6da` | Transient-network retry with exponential backoff |