Files
bls-data/README.md
Dave Boyd 4fe734381e Fix broken series-ID builders; env-var config; project README
Repair every dead series-ID encoding (library now 82/82 query helpers
return live data, verified against the BLS API):

- JOLTS: 21-char format (was 18) — add state/area/sizeclass fields
- OES: national area code 0000000 (was invalid 0000400)
- ECI: correct owner/component/estimate encoding, default unadjusted (CIU)
- Productivity: 4-digit sector + 4-digit measure codes (was 2+3)
- QCEW: 13-char timeseries-API form (ENUUS00010510 / ENU{fips}00010{own}10)
- PPI: repoint finished-goods -> final demand (WPUFD4); keep alias
- ECEC: drop fabricated health-insurance/retirement helpers; add total benefits
- wages SOC: software developers 151132 -> 151252 (2018 SOC)

Tooling/docs:
- config reads BLS_API_KEY env var (takes precedence; config.py gitignored)
- add requirements.txt
- rewrite README as project front door + coverage table + limitations
- correct JOLTS/OES tables in series_id_formats.md, USAGE.md, dataset explorer

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 09:45:27 -04:00

167 lines
6.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# BLS Data Library
A small, dependency-light Python library for pulling U.S. Bureau of Labor Statistics
(BLS) time series — unemployment, payrolls, inflation, wages, job openings, and
productivity — with pre-built series-ID helpers so you never have to hand-encode a
series ID.
```python
from bls_client import BLSClient
from bls_client.queries import employment, prices
client = BLSClient(API_KEY)
# Latest national nonfarm payrolls
client.fetch_latest(employment.nonfarm_payrolls(), years=1)
# A labeled CPI dashboard in one call
client.fetch_named(prices.cpi_dashboard(), 2024, 2025)
```
All 82 zero-argument query helpers are verified against the live API.
---
## Install
```bash
pip install -r requirements.txt # just `requests`
```
## Configure your API key
Register for a free key (instant, no approval): https://data.bls.gov/registrationEngine/
The free v2 key allows 500 queries/day, 50 series/query, 20 years/query.
```bash
cp config.example.py config.py
export BLS_API_KEY="your-key" # preferred — config.py reads this env var
```
`config.py` is gitignored; the env var takes precedence over anything written in the file.
You can also pass the key directly: `BLSClient("your-key")`.
## Quick start
```bash
python3 examples/basic_pull.py # payrolls, CPI dashboard, DC-region unemployment
python3 examples/custom_series.py # building custom series IDs
```
---
## What's in the box
```
bls_client/
├── client.py BLSClient — batching, named fetches, catalog metadata, row flattening
├── series.py low-level series-ID builders (LAUS, CES, CPI, PPI, OES, JOLTS, ECI, QCEW, productivity)
└── queries/
├── employment.py payrolls, unemployment (LAUS), JOLTS, QCEW
├── prices.py CPI, PPI, average prices, import/export prices
├── wages.py OES occupational wages, ECI, ECEC
└── productivity.py major-sector productivity & costs
```
`BLSClient` highlights:
- `fetch(ids, start, end)` — auto-batches >50 series, returns `{series_id: {data, catalog}}`
- `fetch_latest(ids, years=N)` — most recent N years
- `fetch_named({label: id})` — returns results keyed by your labels
- `latest_obs(series)` / `to_rows(results)` — convenience for the most recent value / CSV-ready rows
See **[USAGE.md](USAGE.md)** for the full API and **[series_id_formats.md](series_id_formats.md)**
for the series-ID decode tables.
---
## Coverage
| Survey | Helpers | Notes |
|---|---|---|
| LAUS — local area unemployment | state / metro / county rates, DC-region dashboard | ✅ |
| CES — payroll employment | national by supersector; state/metro | ✅ |
| CPS — household survey | national unemployment rate, participation | ✅ |
| JOLTS — job openings & turnover | openings/hires/quits/layoffs, dashboard | ✅ |
| CPI — consumer prices | all-items, core, food, energy, gasoline, …; dashboard | ✅ |
| PPI — producer prices | all commodities, final demand, food, energy | ✅ |
| OES — occupational wages | employment + wage percentiles by SOC; 15 common occupations | ✅ |
| ECI — employment cost index | total comp / wages / benefits × civilian/private/gov | ✅ |
| ECEC — employer cost levels | total compensation, total benefits | ✅ (totals only) |
| Productivity & costs | output/hr, ULC, comp, hours × business/nonfarm/manufacturing | ✅ |
| QCEW — quarterly census | national & state private totals | ⚠️ totals only via this API |
## Known limitations / what "complete" would add
- **QCEW** is only partially served by the BLS *timeseries* API used here (national and
state-level totals work). County- and industry-level QCEW detail requires the separate
**QCEW Open Data API** (CSV: `https://data.bls.gov/cew/data/api/...`). Not yet wired up.
- **ECEC benefit subcomponents** (health insurance, retirement & savings, etc.) need
specific benefit-subcell codes from the ECEC component list; only the compensation and
benefits totals are currently exposed.
- **No caching / rate-limit handling.** Repeated runs spend against the 500/day quota; a
small on-disk cache and a friendly error on `REQUEST_NOT_PROCESSED` (quota hit) would help.
- **No packaging.** Importable in-tree but not `pip install`-able; add a `pyproject.toml`
to ship it as a real package.
- **No automated tests.** A pytest suite asserting builder outputs against known-good IDs
(and a live smoke test behind a marker) would lock the series-ID encodings in place.
---
## Appendix: BLS API reference
Reference material for working directly with the API (the library wraps all of this).
### API versions
| Feature | v1 (no key) | v2 (registered) |
|---------------------------|-------------|-----------------|
| Daily query limit | 25 | 500 |
| Series per query | 25 | 50 |
| Years of history | 10 | 20 |
| Net/percent changes | No | Yes |
| Series descriptions | No | Yes |
| Calculations | No | Yes |
### Endpoints
```
GET /v2/surveys # all survey codes and names (see surveys.json)
GET /v2/surveys/{abbr} # one survey
GET /v2/timeseries/popular?survey={abbr} # 25 most-requested series IDs for a survey
POST /v2/timeseries/data/ # the time-series data endpoint
```
Base URL: `https://api.bls.gov/publicAPI/v2`
**POST body (registered, all v2 features):**
```json
{
"seriesid": ["LAUST110000000000003", "CES0000000001"],
"startyear": "2020", "endyear": "2025",
"registrationkey": "YOUR_KEY",
"catalog": true, "calculations": true, "annualaverage": true
}
```
**Period codes:** monthly `M01``M12` (`M13` = annual avg); quarterly `Q01``Q04` (`Q05` = annual avg); annual `A01`.
**Status codes:** `REQUEST_SUCCEEDED`, `REQUEST_FAILED`, `REQUEST_NOT_PROCESSED` (often = daily quota hit).
**Footnote codes:** `R` revised, `P` preliminary, `X`/`N` unavailable.
A full real response is saved in `api_response_example.json`.
### Bulk flat files
Base: `https://download.bls.gov/pub/time.series/` — each survey folder has
`{prefix}.series` (master list), `{prefix}.data.*` (observations), and `{prefix}.{dimension}`
decode tables (area, industry, measure). Tab-delimited; handy for bulk PostgreSQL ingest.
Key prefixes: `la/` LAUS, `ce/` CES, `sm/` state-metro, `en/` QCEW, `oe/` OES, `jt/` JOLTS,
`cu/` CPI-U, `wp/` PPI.
### Other files in this repo
- `series_id_formats.md` — series-ID decode tables for each survey
- `qcew_field_schema.md` — QCEW quarterly/annual CSV field layouts
- `surveys.json` — complete survey list from the API
- `bls_dataset_explorer.py` / `.html` — browsable survey catalog
- `dc_md_va_unemployment.py` / `generate_report.py` — example report generators