跳至正文
返回文库全部文档

Chen-Pelger-Zhu 公司特征与资产定价

文章 《交易机器学习》

总结

本文档介绍一个可复现的 US 股票基准数据集,取自 Chen、Pelger 和 Zhu 的资产定价复现档案。数据集包含约 1.2 百万条月度股票观测,覆盖 1967 至 2016,包括 46 项公司特征、宏观经济指标、前向收益,以及预先定义的训练、验证和测试分区。公司标识符在各自发布的张量中匿名且保持一致,但不同标识符命名空间使研究者无法跨张量关联公司。

该数据集支持资产定价机器学习研究,适用于一致的基准数据比解释个别证券更重要的场景。其公开的时间分区将 1967–1986 用于训练,1987–1991 用于验证,1992–2016 用于测试。文档介绍了如何下载、将作者提供的张量文件转换为 Parquet,并加载完整面板或单独分区,也包括载入宏观数据的选项。这是数据集说明,而非实证结果:表现取决于所用模型和评估方法,匿名标识符也限制了跨分区的公司层面分析。

核心观点

  • 该基准包含 1967 至 2016 年的月度股票观测、公司特征、宏观指标和前向收益。
  • 预先定义的时间分区将训练、验证和测试时段彼此分开。
  • 匿名公司标识符在各张量内保持一致,但无法跨张量命名空间关联。
  • 该数据集用于可复现的线性和非线性资产定价实验。
  • 转换后的 Parquet 文件保留了从原始张量数据中恢复的标识符。

标签

全文
# Firm Characteristics (Chen-Pelger-Zhu 2020)


# Firm Characteristics (Chen-Pelger-Zhu 2020)

Anonymized panel of ~1.2M stock-month observations with 46 firm
characteristics and forward returns, spanning 1967-2016. Built from the
replication archive of Chen, Pelger, and Zhu (2020), *Deep Learning in
Asset Pricing*. Used throughout the book for ML-based asset-pricing
examples where a standard, reproducible benchmark matters more than
symbol-level interpretation.

## Dataset

- **Source**: GitHub replication repo
  (https://github.com/jasonzy121/Deep_Learning_Asset_Pricing), which
  itself ships the published dataset via Google Drive
- **Coverage**: 1967-01 to 2016-12, monthly observations, ~1.2M rows
- **Features**: 46 firm characteristics (accounting ratios, price-based
  measures, momentum variants), 178 macro indicators, forward returns
- **Size on disk**: ~1.1 GB raw CSV; converted to ~500 MB parquet
- **Access**: Public, no API key required
- **Canonical schema**: `symbol` (anonymous integer id), `timestamp`
  (monthly Date), 46 characteristic columns, `ret`, `split`
- **Identity scope**: `symbol` is persistent within each tensor released by
  the authors. Separate numeric namespaces prevent accidental linking across
  tensors because the archive does not publish a cross-split map.

## Pre-defined Splits

The dataset ships with deterministic train/valid/test splits aligned to
the authors' released tensors:

| Split | Period    |
| ----- | --------- |
| train | 1967-1986 |
| valid | 1987-1991 |
| test  | 1992-2016 |

## Download

This is the largest free dataset in the book: ~1.5 GB pulled from the
authors' Google Drive folder (RetChar.csv alone is ~1.1 GB), followed by
a tensor-to-Parquet conversion. The tensor conversion recovers persistent
anonymous firm identifiers that the flattened CSV omits. How long it takes
depends on your bandwidth and disk speed. Per-file progress is printed as it
runs. A single command downloads and converts; no separate `--convert` pass
is needed.

```bash
# Download the ~1.5 GB folder and convert to parquet in one step
uv run python data/equities/firm_characteristics/download.py

# Verify what's already on disk, do not refetch
uv run python data/equities/firm_characteristics/download.py --check

# Force a re-download even if files exist
uv run python data/equities/firm_characteristics/download.py --force

# Re-run only the tensor-to-Parquet conversion (files already downloaded)
uv run python data/equities/firm_characteristics/download.py --convert
```

It is also fetched automatically as part of `data/download_all.py`. To
skip it there (e.g. on a metered or space-constrained connection), pass
`--skip-firm-characteristics`.

Output layout under `$ML4T_DATA_PATH/equities/firm_characteristics/`:

```
firm_characteristics_all.parquet      # full panel
firm_characteristics_train.parquet    # 1967-1986
firm_characteristics_valid.parquet    # 1987-1991
firm_characteristics_test.parquet     # 1992-2016
dl_asset_pricing/                     # raw CSV + NPZ staging
    RetChar.csv
    Macro.csv
    char/Char_{train,valid,test}.npz
    macro/macro_{train,valid,test}.npz
    RF/RF_{train,valid,test}_normalized_task_1.npz
```

The staging directory is kept so experiments that need the pre-split
NPZ arrays (the original Chen-Pelger-Zhu format) can read them
directly.

## Loading

```python
from data import load_firm_characteristics

df = load_firm_characteristics()                        # full panel
df = load_firm_characteristics(split="train")
df = load_firm_characteristics(split="test", include_macro=True)
```

If the canonical parquets are missing, the loader raises
`DataNotFoundError` pointing at the download command.

## Consumers

- Chapters 10-16 - standard benchmark for linear and nonlinear
  asset-pricing models.
- `case_studies/us_firm_characteristics/` - Chen-Pelger-Zhu replication
  pipeline (CV, GBM, latent factor, deep models).

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。