Files
bls-data/README.md
Dave Boyd 881d62cf32 Add pyproject.toml packaging and pytest suite
- pyproject.toml: setuptools build, deps (requests), dev extra (pytest),
  pytest config with opt-in `live` marker
- tests/: 95 offline tests locking every series-ID encoding against
  known-good IDs (regression guard for the JOLTS/OES/ECI/productivity/QCEW
  fixes), plus a live API smoke test behind `-m live`
- README: install-as-package + test instructions; drop the now-resolved
  "no packaging"/"no tests" limitations
- gitignore build/test artifacts

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 09:53:37 -04:00

173 lines
6.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# BLS Data Library
A small, dependency-light Python library for pulling U.S. Bureau of Labor Statistics
(BLS) time series — unemployment, payrolls, inflation, wages, job openings, and
productivity — with pre-built series-ID helpers so you never have to hand-encode a
series ID.
```python
from bls_client import BLSClient
from bls_client.queries import employment, prices
client = BLSClient(API_KEY)
# Latest national nonfarm payrolls
client.fetch_latest(employment.nonfarm_payrolls(), years=1)
# A labeled CPI dashboard in one call
client.fetch_named(prices.cpi_dashboard(), 2024, 2025)
```
All 82 zero-argument query helpers are verified against the live API.
---
## Install
```bash
pip install -r requirements.txt # just `requests`
# or install the package itself (editable):
pip install -e .
```
## Configure your API key
Register for a free key (instant, no approval): https://data.bls.gov/registrationEngine/
The free v2 key allows 500 queries/day, 50 series/query, 20 years/query.
```bash
cp config.example.py config.py
export BLS_API_KEY="your-key" # preferred — config.py reads this env var
```
`config.py` is gitignored; the env var takes precedence over anything written in the file.
You can also pass the key directly: `BLSClient("your-key")`.
## Quick start
```bash
python3 examples/basic_pull.py # payrolls, CPI dashboard, DC-region unemployment
python3 examples/custom_series.py # building custom series IDs
```
## Tests
```bash
pip install -e ".[dev]" # pytest
pytest # offline: 95 tests locking the series-ID encodings
pytest -m live # also hit the live API (needs network + BLS_API_KEY)
```
---
## What's in the box
```
bls_client/
├── client.py BLSClient — batching, named fetches, catalog metadata, row flattening
├── series.py low-level series-ID builders (LAUS, CES, CPI, PPI, OES, JOLTS, ECI, QCEW, productivity)
└── queries/
├── employment.py payrolls, unemployment (LAUS), JOLTS, QCEW
├── prices.py CPI, PPI, average prices, import/export prices
├── wages.py OES occupational wages, ECI, ECEC
└── productivity.py major-sector productivity & costs
```
`BLSClient` highlights:
- `fetch(ids, start, end)` — auto-batches >50 series, returns `{series_id: {data, catalog}}`
- `fetch_latest(ids, years=N)` — most recent N years
- `fetch_named({label: id})` — returns results keyed by your labels
- `latest_obs(series)` / `to_rows(results)` — convenience for the most recent value / CSV-ready rows
See **[USAGE.md](USAGE.md)** for the full API and **[series_id_formats.md](series_id_formats.md)**
for the series-ID decode tables.
---
## Coverage
| Survey | Helpers | Notes |
|---|---|---|
| LAUS — local area unemployment | state / metro / county rates, DC-region dashboard | ✅ |
| CES — payroll employment | national by supersector; state/metro | ✅ |
| CPS — household survey | national unemployment rate, participation | ✅ |
| JOLTS — job openings & turnover | openings/hires/quits/layoffs, dashboard | ✅ |
| CPI — consumer prices | all-items, core, food, energy, gasoline, …; dashboard | ✅ |
| PPI — producer prices | all commodities, final demand, food, energy | ✅ |
| OES — occupational wages | employment + wage percentiles by SOC; 15 common occupations | ✅ |
| ECI — employment cost index | total comp / wages / benefits × civilian/private/gov | ✅ |
| ECEC — employer cost levels | total compensation, total benefits | ✅ (totals only) |
| Productivity & costs | output/hr, ULC, comp, hours × business/nonfarm/manufacturing | ✅ |
| QCEW — quarterly census | national & state private totals | ⚠️ totals only via this API |
## Known limitations / what "complete" would add
- **QCEW** is only partially served by the BLS *timeseries* API used here (national and
state-level totals work). County- and industry-level QCEW detail requires the separate
**QCEW Open Data API** (CSV: `https://data.bls.gov/cew/data/api/...`). Not yet wired up.
- **ECEC benefit subcomponents** (health insurance, retirement & savings, etc.) need
specific benefit-subcell codes from the ECEC component list; only the compensation and
benefits totals are currently exposed.
- **No caching / rate-limit handling.** Repeated runs spend against the 500/day quota; a
small on-disk cache and a friendly error on `REQUEST_NOT_PROCESSED` (quota hit) would help.
---
## Appendix: BLS API reference
Reference material for working directly with the API (the library wraps all of this).
### API versions
| Feature | v1 (no key) | v2 (registered) |
|---------------------------|-------------|-----------------|
| Daily query limit | 25 | 500 |
| Series per query | 25 | 50 |
| Years of history | 10 | 20 |
| Net/percent changes | No | Yes |
| Series descriptions | No | Yes |
| Calculations | No | Yes |
### Endpoints
```
GET /v2/surveys # all survey codes and names (see surveys.json)
GET /v2/surveys/{abbr} # one survey
GET /v2/timeseries/popular?survey={abbr} # 25 most-requested series IDs for a survey
POST /v2/timeseries/data/ # the time-series data endpoint
```
Base URL: `https://api.bls.gov/publicAPI/v2`
**POST body (registered, all v2 features):**
```json
{
"seriesid": ["LAUST110000000000003", "CES0000000001"],
"startyear": "2020", "endyear": "2025",
"registrationkey": "YOUR_KEY",
"catalog": true, "calculations": true, "annualaverage": true
}
```
**Period codes:** monthly `M01``M12` (`M13` = annual avg); quarterly `Q01``Q04` (`Q05` = annual avg); annual `A01`.
**Status codes:** `REQUEST_SUCCEEDED`, `REQUEST_FAILED`, `REQUEST_NOT_PROCESSED` (often = daily quota hit).
**Footnote codes:** `R` revised, `P` preliminary, `X`/`N` unavailable.
A full real response is saved in `api_response_example.json`.
### Bulk flat files
Base: `https://download.bls.gov/pub/time.series/` — each survey folder has
`{prefix}.series` (master list), `{prefix}.data.*` (observations), and `{prefix}.{dimension}`
decode tables (area, industry, measure). Tab-delimited; handy for bulk PostgreSQL ingest.
Key prefixes: `la/` LAUS, `ce/` CES, `sm/` state-metro, `en/` QCEW, `oe/` OES, `jt/` JOLTS,
`cu/` CPI-U, `wp/` PPI.
### Other files in this repo
- `series_id_formats.md` — series-ID decode tables for each survey
- `qcew_field_schema.md` — QCEW quarterly/annual CSV field layouts
- `surveys.json` — complete survey list from the API
- `bls_dataset_explorer.py` / `.html` — browsable survey catalog
- `dc_md_va_unemployment.py` / `generate_report.py` — example report generators