Files
bls-data/README.md
Dave Boyd 9e2c55c568 Add QCEW CSV module, ECEC breakdown, response caching + typed errors
Item 1 — QCEW county/industry detail:
- new bls_client/qcew.py: thin client over the QCEW Open Data CSV service
  (area/industry/size slices; no API key, no quota), with filter_rows helper
- covers the county/industry detail the timeseries API can't reach

Item 2 — ECEC benefit breakdown:
- new series.ecec() builder from the authoritative cm.estimate decode table
- real helpers: wages, total benefits, paid leave, supplemental pay,
  health insurance, retirement & savings, legally required + ecec_dashboard()
- FIX: ecec_total_benefits was CMU1036... (education/health industries only);
  correct all-civilian total benefits is CMU1030000000000D

Item 3 — caching + typed errors:
- bls_client/cache.py FileCache (on-disk, TTL); BLSClient(cache=True),
  queries_used counter (cache hits don't spend quota)
- bls_client/errors.py: BLSError / BLSQuotaError / BLSRequestError; quota
  exhaustion now raises a clear BLSQuotaError instead of a generic RuntimeError

Tests: 119 offline (+24) incl. ECEC goldens, QCEW fixture parse, cache/quota
behavior; live smoke extended to ECEC + QCEW. Helper sweep: 96/96 live.
README: caching/QCEW usage, updated coverage + limitations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 10:38:12 -04:00

196 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# BLS Data Library
A small, dependency-light Python library for pulling U.S. Bureau of Labor Statistics
(BLS) time series — unemployment, payrolls, inflation, wages, job openings, and
productivity — with pre-built series-ID helpers so you never have to hand-encode a
series ID.
```python
from bls_client import BLSClient
from bls_client.queries import employment, prices
client = BLSClient(API_KEY)
# Latest national nonfarm payrolls
client.fetch_latest(employment.nonfarm_payrolls(), years=1)
# A labeled CPI dashboard in one call
client.fetch_named(prices.cpi_dashboard(), 2024, 2025)
```
All 82 zero-argument query helpers are verified against the live API.
---
## Install
```bash
pip install -r requirements.txt # just `requests`
# or install the package itself (editable):
pip install -e .
```
## Configure your API key
Register for a free key (instant, no approval): https://data.bls.gov/registrationEngine/
The free v2 key allows 500 queries/day, 50 series/query, 20 years/query.
```bash
cp config.example.py config.py
export BLS_API_KEY="your-key" # preferred — config.py reads this env var
```
`config.py` is gitignored; the env var takes precedence over anything written in the file.
You can also pass the key directly: `BLSClient("your-key")`.
## Quick start
```bash
python3 examples/basic_pull.py # payrolls, CPI dashboard, DC-region unemployment
python3 examples/custom_series.py # building custom series IDs
```
## Tests
```bash
pip install -e ".[dev]" # pytest
pytest # offline: 95 tests locking the series-ID encodings
pytest -m live # also hit the live API (needs network + BLS_API_KEY)
```
---
## What's in the box
```
bls_client/
├── client.py BLSClient — batching, caching, named fetches, catalog metadata, row flattening
├── series.py low-level series-ID builders (LAUS, CES, CPI, PPI, OES, JOLTS, ECI, ECEC, QCEW, productivity)
├── qcew.py QCEW Open Data CSV client (county/industry detail; no key, no quota)
├── cache.py on-disk response cache (FileCache)
├── errors.py BLSError / BLSQuotaError / BLSRequestError
└── queries/
├── employment.py payrolls, unemployment (LAUS), JOLTS, QCEW totals
├── prices.py CPI, PPI, average prices, import/export prices
├── wages.py OES occupational wages, ECI, ECEC breakdown
└── productivity.py major-sector productivity & costs
```
`BLSClient` highlights:
- `fetch(ids, start, end)` — auto-batches >50 series, returns `{series_id: {data, catalog}}`
- `fetch_latest(ids, years=N)` — most recent N years
- `fetch_named({label: id})` — returns results keyed by your labels
- `latest_obs(series)` / `to_rows(results)` — convenience for the most recent value / CSV-ready rows
See **[USAGE.md](USAGE.md)** for the full API and **[series_id_formats.md](series_id_formats.md)**
for the series-ID decode tables.
---
## Coverage
| Survey | Helpers | Notes |
|---|---|---|
| LAUS — local area unemployment | state / metro / county rates, DC-region dashboard | ✅ |
| CES — payroll employment | national by supersector; state/metro | ✅ |
| CPS — household survey | national unemployment rate, participation | ✅ |
| JOLTS — job openings & turnover | openings/hires/quits/layoffs, dashboard | ✅ |
| CPI — consumer prices | all-items, core, food, energy, gasoline, …; dashboard | ✅ |
| PPI — producer prices | all commodities, final demand, food, energy | ✅ |
| OES — occupational wages | employment + wage percentiles by SOC; 15 common occupations | ✅ |
| ECI — employment cost index | total comp / wages / benefits × civilian/private/gov | ✅ |
| ECEC — employer cost levels | full breakdown: comp, wages, benefits, paid leave, supplemental, health insurance, retirement, legally-required | ✅ |
| Productivity & costs | output/hr, ULC, comp, hours × business/nonfarm/manufacturing | ✅ |
| QCEW — quarterly census | national/state totals (timeseries) **plus** full county/industry detail via the CSV module (`bls_client.qcew`) | ✅ |
## Caching & quota handling
```python
client = BLSClient(API_KEY, cache=True) # on-disk cache at ~/.cache/bls (1-day TTL)
client.fetch(...) # repeated identical requests don't re-spend quota
client.queries_used # network calls made this session (cache hits excluded)
```
A blown daily quota raises `BLSQuotaError` (vs `BLSRequestError` for a bad request), so
"am I throttled or is my series ID wrong?" is no longer ambiguous.
## QCEW county/industry detail
The timeseries helpers give QCEW national/state totals; the `qcew` module reaches the full
county- and industry-level detail via BLS's separate CSV service (no key, no quota):
```python
from bls_client import qcew
rows = qcew.area("11000", 2024, "a") # everything for DC, annual
dc_private = qcew.filter_rows(rows, own_code="private", industry_code="10", agglvl_code="51")
hospitals = qcew.industry("622", 2024, "a") # one NAICS across all areas
```
## Known limitations
- **No automatic retry/backoff** on transient network errors (a `requests` failure surfaces
directly). Caching mitigates repeat load but there's no rate-limit pacing.
- **Coverage is the headline cut** of each survey, not an exhaustive mirror — e.g. CES is
national supersectors + a couple of states; CPI is the common items; OES ships 15 named
occupations (any SOC works via `occupation_*`). Broaden the helper dicts as needed.
---
## Appendix: BLS API reference
Reference material for working directly with the API (the library wraps all of this).
### API versions
| Feature | v1 (no key) | v2 (registered) |
|---------------------------|-------------|-----------------|
| Daily query limit | 25 | 500 |
| Series per query | 25 | 50 |
| Years of history | 10 | 20 |
| Net/percent changes | No | Yes |
| Series descriptions | No | Yes |
| Calculations | No | Yes |
### Endpoints
```
GET /v2/surveys # all survey codes and names (see surveys.json)
GET /v2/surveys/{abbr} # one survey
GET /v2/timeseries/popular?survey={abbr} # 25 most-requested series IDs for a survey
POST /v2/timeseries/data/ # the time-series data endpoint
```
Base URL: `https://api.bls.gov/publicAPI/v2`
**POST body (registered, all v2 features):**
```json
{
"seriesid": ["LAUST110000000000003", "CES0000000001"],
"startyear": "2020", "endyear": "2025",
"registrationkey": "YOUR_KEY",
"catalog": true, "calculations": true, "annualaverage": true
}
```
**Period codes:** monthly `M01``M12` (`M13` = annual avg); quarterly `Q01``Q04` (`Q05` = annual avg); annual `A01`.
**Status codes:** `REQUEST_SUCCEEDED`, `REQUEST_FAILED`, `REQUEST_NOT_PROCESSED` (often = daily quota hit).
**Footnote codes:** `R` revised, `P` preliminary, `X`/`N` unavailable.
A full real response is saved in `api_response_example.json`.
### Bulk flat files
Base: `https://download.bls.gov/pub/time.series/` — each survey folder has
`{prefix}.series` (master list), `{prefix}.data.*` (observations), and `{prefix}.{dimension}`
decode tables (area, industry, measure). Tab-delimited; handy for bulk PostgreSQL ingest.
Key prefixes: `la/` LAUS, `ce/` CES, `sm/` state-metro, `en/` QCEW, `oe/` OES, `jt/` JOLTS,
`cu/` CPI-U, `wp/` PPI.
### Other files in this repo
- `series_id_formats.md` — series-ID decode tables for each survey
- `qcew_field_schema.md` — QCEW quarterly/annual CSV field layouts
- `surveys.json` — complete survey list from the API
- `bls_dataset_explorer.py` / `.html` — browsable survey catalog
- `dc_md_va_unemployment.py` / `generate_report.py` — example report generators