From 8f0122011a210cbabba5fface2f78ec83e02ec93 Mon Sep 17 00:00:00 2001 From: Dave Boyd Date: Mon, 22 Jun 2026 10:51:57 -0400 Subject: [PATCH] Add project overview & usage guide (docs/BLS_Library_Overview.md) Narrative of the audit/repair work plus a practical usage guide for the client, query helpers, series builders, QCEW module, caching/errors/retry, testing, and coverage. Co-Authored-By: Claude Opus 4.8 (1M context) --- docs/BLS_Library_Overview.md | 293 +++++++++++++++++++++++++++++++++++ 1 file changed, 293 insertions(+) create mode 100644 docs/BLS_Library_Overview.md diff --git a/docs/BLS_Library_Overview.md b/docs/BLS_Library_Overview.md new file mode 100644 index 0000000..aa5b935 --- /dev/null +++ b/docs/BLS_Library_Overview.md @@ -0,0 +1,293 @@ +# BLS Data Library — Overview & Usage Guide + +**Repository:** `giteaadmin/bls-data` on `gitea.doesworks.net` +**Status:** feature-complete · 96/96 query helpers verified live · 131 tests · pip-installable +**Last updated:** June 22, 2026 + +This document covers two things: **what the library is and how it got here**, and a +**practical guide to using it**. For the terse reference, see `README.md`, +`USAGE.md`, and `series_id_formats.md`. + +--- + +## 1. What this is + +A dependency-light Python client for the U.S. Bureau of Labor Statistics (BLS) API. +Its job is to make BLS data retrievable **without hand-encoding series IDs** — the +cryptic 11-to-25-character codes BLS uses as primary keys (e.g. +`OEUN000000000000015125213` for "national median annual wage, software developers"). + +The library wraps two BLS data services: + +| Service | What it serves | Key? | Quota? | +|---|---|---|---| +| **Timeseries API** (`api.bls.gov/publicAPI/v2`) | Most surveys — unemployment, payrolls, prices, wages, openings, productivity | yes (free) | 500/day | +| **QCEW Open Data** (`data.bls.gov/cew/data/api`) | County- and industry-level employment & wages | no | none | + +It provides: a batching/caching HTTP client, ~96 pre-built "query helpers" for common +series, low-level series-ID builders for everything else, typed errors, retry/backoff, +and a test suite that locks the encodings in place. + +--- + +## 2. What we did (engineering summary) + +The library existed but was **substantially broken** — a full audit found only 58 of 82 +query helpers returned live data. Whole survey modules emitted series IDs that BLS +rejected. The work fell into four phases: + +### Phase 1 — Fix the broken encodings +Every dead helper traced back to a wrong series-ID encoding (mostly wrong field widths +inherited from an internally-inconsistent reference doc). Fixed and verified against the +live API: + +| Survey | Bug | Fix | +|---|---|---| +| **JOLTS** | IDs were 18 chars (12 zeros) | correct format is 21 chars / 15 zeros (added state+area+sizeclass fields) | +| **OES** | national area code hard-coded `0000400` | corrected to `0000000` | +| **ECI** | wages/benefits put as the trailing char | trailing char is the *estimate* code; component is a mid-string field; published unadjusted (CIU) | +| **Productivity** | 2-digit sector + 3-digit measure | real format is 4-digit each (e.g. `PRS85006092`) | +| **PPI** | "finished goods" index (discontinued by BLS) | repointed to **final demand** (`WPUFD4`) | +| **SOC code** | software developers `151132` (2010 SOC, retired) | `151252` (2018 SOC) | + +Result: **82/82 helpers live.** + +### Phase 2 — Make it shareable +- API key moved to the `BLS_API_KEY` environment variable (config file gitignored; the + key was never in git history). +- Added `requirements.txt`; rewrote `README.md` as a proper project front door. + +### Phase 3 — Package + test +- `pyproject.toml` — the library is now `pip install`-able; builds a wheel and sdist. +- A **pytest suite** that locks every series-ID encoding against known-good IDs (offline, + no quota) plus a live smoke test behind an opt-in `-m live` marker. This is the + regression guard — it fails instantly if an encoding is ever re-broken. + +### Phase 4 — Complete the coverage +Three gaps closed in one efficient pass (total cost: ~2 of the 500 daily API queries, +because the decode tables and QCEW data are off-quota and ECEC verification batched into +a single call): + +1. **QCEW county/industry detail** — a new `qcew` module over the separate CSV service, + reaching the per-county, per-industry data the timeseries API can't. +2. **ECEC benefit breakdown** — health insurance, retirement, paid leave, etc., built + from BLS's authoritative decode table (not guessed). This also caught a latent bug: + `ecec_total_benefits` had pointed at an education/health-industry series rather than + the all-civilian total. +3. **Caching + typed errors** — an on-disk response cache (so reruns don't re-spend + quota) and a `BLSQuotaError` that distinguishes "you're throttled" from "your series + ID is wrong" — the exact ambiguity that had previously stalled debugging. + +Plus a **retry/backoff layer** for transient network failures (timeouts, `429`/`5xx`), +which never retries quota rejections or 4xx. + +**End state:** 96/96 helpers live, 131 tests, every survey at full useful depth. + +--- + +## 3. Architecture + +``` +bls_client/ +├── client.py BLSClient — batching, caching, retry, named fetches, row flattening +├── series.py low-level series-ID builders (every survey) +├── qcew.py QCEW Open Data CSV client (county/industry; no key, no quota) +├── cache.py on-disk response cache (FileCache) +├── retry.py exponential-backoff retry for transient network errors +├── errors.py BLSError / BLSQuotaError / BLSRequestError +└── queries/ + ├── employment.py payrolls, unemployment (LAUS), JOLTS, QCEW totals + ├── prices.py CPI, PPI, average prices, import/export prices + ├── wages.py OES occupational wages, ECI, ECEC breakdown + └── productivity.py major-sector productivity & costs +``` + +Two layers: **query helpers** (`queries/…`) return ready-to-use series IDs for common +asks; **series builders** (`series.py`) construct any valid ID from parameters when a +helper doesn't exist. + +--- + +## 4. Install & configure + +```bash +pip install -r requirements.txt # just `requests` +# or, to install the package itself: +pip install -e . +``` + +Get a free API key (instant): https://data.bls.gov/registrationEngine/ — 500 queries/day. + +```bash +cp config.example.py config.py +export BLS_API_KEY="your-key" # config.py reads this env var; takes precedence +``` + +`config.py` is gitignored. You can also pass the key directly: `BLSClient("your-key")`. + +> **Note:** on a PEP-668 "externally managed" Python (e.g. recent Ubuntu), install into a +> virtualenv: `python3 -m venv .venv && . .venv/bin/activate && pip install -e ".[dev]"`. + +--- + +## 5. Usage + +### The client + +```python +from bls_client import BLSClient + +client = BLSClient(API_KEY, cache=True) # caching on (recommended) + +# Fetch by series ID(s) — auto-batches >50 series +data = client.fetch(["CES0000000001"], 2020, 2025) + +# Most recent N years +data = client.fetch_latest("CES0000000001", years=1) + +# Labeled dict in, labeled dict out +data = client.fetch_named({"Payrolls": "CES0000000001"}, 2024, 2025) + +# Helpers for reading results +obs = BLSClient.latest_obs(next(iter(data.values()))) # most recent observation +rows = BLSClient.to_rows(data) # flat rows for CSV/pandas +``` + +### Query helpers (the common case) + +```python +from bls_client.queries import employment, prices, wages, productivity + +employment.nonfarm_payrolls() # headline jobs number +employment.dc_region_unemployment() # DC/MD/VA + 3 metros (labeled dict) +employment.jolts_dashboard() # openings/hires/quits/layoffs/separations + +prices.cpi_dashboard() # CPI: all-items, core, food, energy, … +prices.ppi_final_demand() + +wages.occupation_annual_median_wage("151252") # any SOC code (software devs here) +wages.ecec_dashboard() # full employer cost-per-hour breakdown +wages.eci_total_compensation() + +productivity.productivity_dashboard() # output/hr, ULC, etc. +``` + +Putting it together: + +```python +client = BLSClient(API_KEY, cache=True) +result = client.fetch_named(wages.ecec_dashboard(), 2024, 2024) +for label, series in result.items(): + obs = BLSClient.latest_obs(series) + print(f"{label:22} ${obs['value']}/hr") +# Total Compensation $47.20/hr +# Health Insurance $3.54/hr +# Retirement & Savings $2.45/hr ... +``` + +### Series builders (anything not pre-built) + +```python +from bls_client import series + +series.laus_state(48, "rate") # Texas unemployment rate +series.ces_national("manufacturing") # manufacturing payrolls +series.cpi("gasoline", seasonal=True) +series.oes_national("291141", data_type="annual_median") # registered nurses +series.ecec("health_insurance", owner="private") +``` + +### QCEW — county & industry detail + +The timeseries helpers give QCEW national/state totals; the `qcew` module reaches the +full detail via the separate CSV service (no key, no quota): + +```python +from bls_client import qcew + +rows = qcew.area("11000", 2024, "a") # everything for DC, annual averages +dc_private = qcew.filter_rows(rows, own_code="private", industry_code="10", agglvl_code="51") +# -> establishments, employment, total wages, avg weekly wage for DC private sector + +hospitals = qcew.industry("622", 2024, "a") # one NAICS industry across all areas +``` + +Each call returns a list of dict rows (the CSV columns). Quarter is `1`-`4` or `"a"` for +annual. Area FIPS: `US000` national, `11000` = DC, `11001` = a county, `C####` = metro. + +### Caching, errors, and retry + +```python +client = BLSClient(API_KEY, cache=True, retries=2, backoff=0.5) +client.fetch(...) # identical requests served from ~/.cache/bls (1-day TTL) +client.queries_used # network calls this session (cache hits excluded) +``` + +```python +from bls_client import BLSQuotaError, BLSRequestError +try: + client.fetch(ids, 2024, 2025) +except BLSQuotaError: # daily 500-query limit hit (resets midnight ET) + ... +except BLSRequestError: # malformed request / bad series ID + ... +``` + +Transient failures (timeouts, dropped connections, `429`/`5xx`) retry automatically with +exponential backoff and honor a `Retry-After` header. Quota rejections and 4xx are never +retried. Set `retries=0` to disable. + +--- + +## 6. Coverage + +| Survey | What's available | +|---|---| +| **LAUS** — local area unemployment | state / metro / county rates; DC-region dashboard | +| **CES** — payroll employment | national by supersector; state/metro | +| **CPS** — household survey | national unemployment rate, participation | +| **JOLTS** — job openings & turnover | openings, hires, quits, layoffs, separations; dashboard | +| **CPI** — consumer prices | all-items, core, food, energy, gasoline, shelter, …; dashboard | +| **PPI** — producer prices | all commodities, final demand, food, energy | +| **OES** — occupational wages | employment + wage percentiles by SOC; 15 named occupations | +| **ECI** — employment cost index | total comp / wages / benefits × civilian/private/gov | +| **ECEC** — employer cost levels | full breakdown: comp, wages, benefits, paid leave, supplemental, health insurance, retirement, legally-required | +| **Productivity & costs** | output/hr, unit labor cost, compensation, hours × business/nonfarm/manufacturing | +| **QCEW** — quarterly census | national/state totals **plus** full county/industry detail (CSV module) | + +--- + +## 7. Testing + +```bash +pip install -e ".[dev]" +pytest # 125 offline tests — lock every encoding; no network, no quota +pytest -m live # +live smoke tests against the real API (needs network + key) +``` + +The offline tests are the regression guard: they assert each builder/helper produces its +exact known-good series ID, so a re-broken encoding fails immediately and for free. The +live tests catch the other failure mode — BLS silently retiring a code (as happened with +the 2010→2018 SOC change). + +--- + +## 8. Known limitations + +- **Coverage is the headline cut** of each survey, not an exhaustive mirror (CES is + national supersectors + a couple of states; CPI is the common items; OES ships 15 named + occupations). Any SOC/industry still works through the parameterized builders — the + pre-built helper dicts just cover the common asks. Widen them as needed. +- No rate-limit *pacing* beyond caching + retry; a heavy batch can still approach the + 500/day quota (the `queries_used` counter and `BLSQuotaError` make that visible). + +--- + +## 9. Change history + +| Commit | What | +|---|---| +| `4fe7343` | Fixed all broken series-ID encodings (JOLTS/OES/ECI/productivity/PPI/SOC); env-var config; project README | +| `881d62c` | `pyproject.toml` packaging + pytest suite | +| `9e2c55c` | QCEW CSV module, ECEC benefit breakdown, response caching + typed errors | +| `0d0d6da` | Transient-network retry with exponential backoff |