Skip to content
All library documents

Using Survivor-Free Historical US Equity Data

Article Machine Learning for Trading

Summary

This document describes a daily US equity dataset covering roughly 3,200 companies from 1962 through March 2018. It includes delisted issuers, making it useful for historical research where excluding failed or delisted firms could distort results. The data contains standard OHLCV fields and adjusted close, and can be loaded for selected symbols, dates, or a smaller random sample for prototyping.

The dataset is frozen because its provider stopped updating it in March 2018. Researchers should therefore treat it as a fixed historical source rather than a current market feed. The document explains how to download and load the parquet data and identifies example research uses, but presents no trading analysis or performance evidence. Its main research value is the stated historical coverage and survivorship-bias-free universe; it does not explain how delisted securities are represented, how corporate actions are handled, or what other data quality checks may be needed.

Key ideas

  • The dataset contains daily OHLCV history for roughly 3,200 US companies, including delisted issuers.
  • The historical series runs from 1962 to March 2018 and is no longer updated.
  • Researchers can load the full universe or filter by symbol, date range, or sample size.
  • The document describes data access and loading, but provides no strategy results or data-quality analysis.

Tags

Full text
# US Equities (NASDAQ Data Link WIKI Prices)


# US Equities (NASDAQ Data Link WIKI Prices)

Historical US equity OHLCV for ~3,200 companies from 1962-01-02 to
2018-03-27. Survivorship-bias free — includes delisted issuers. The
dataset is frozen (NASDAQ Data Link stopped updating it in March 2018),
so the downloader writes once and you never re-pull.

## Dataset

- **Source**: NASDAQ Data Link (formerly Quandl) `WIKI/PRICES`
- **Coverage**: 1962-01-02 → 2018-03-27 daily OHLCV, ~3,199 US tickers
- **Size on disk**: ~650 MB parquet
- **Access**: Free with a NASDAQ Data Link API key
  (https://data.nasdaq.com/sign-up)
- **Canonical schema**: `symbol`, `timestamp` (Date), `open`, `high`,
  `low`, `close`, `volume`, `adj_close`, …

## Download

```bash
# One-off: ~several minutes, 650 MB
uv run python data/equities/market/us_equities/download.py

# Preview without hitting the API
uv run python data/equities/market/us_equities/download.py --dry-run

# Re-download even if the parquet is already on disk
uv run python data/equities/market/us_equities/download.py --force
```

API key resolution order (first match wins):

1. `--api-key <key>` CLI flag
2. `QUANDL_API_KEY` environment variable
3. `NASDAQ_DATA_LINK_API_KEY` environment variable
4. Same keys in a `.env` file at the repo root

Output: `$ML4T_DATA_PATH/equities/us_equities/us_equities.parquet`.

## Loading

```python
from data import load_us_equities

df = load_us_equities()                        # full universe
df = load_us_equities(symbols=["AAPL", "MSFT"])
df = load_us_equities(start_date="2000-01-01", end_date="2010-12-31")
df = load_us_equities(max_symbols=50)          # random 50 for prototyping
```

If the parquet is missing, the loader raises `DataNotFoundError` with
the exact download command.

## Consumers

- `case_studies/us_equities_panel/` — the dedicated case study (Ch7 →
  Ch20 pipeline) operates on this universe.
- Chapter 2 EDA notebook `01_us_equities_eda.py` — coverage survey.
- Any chapter that needs a long, survivor-free US equities history.

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.