Chen-Pelger-Zhu 公司特征与资产定价
文章 《交易机器学习》
总结
本文档介绍一个可复现的 US 股票基准数据集,取自 Chen、Pelger 和 Zhu 的资产定价复现档案。数据集包含约 1.2 百万条月度股票观测,覆盖 1967 至 2016,包括 46 项公司特征、宏观经济指标、前向收益,以及预先定义的训练、验证和测试分区。公司标识符在各自发布的张量中匿名且保持一致,但不同标识符命名空间使研究者无法跨张量关联公司。
该数据集支持资产定价机器学习研究,适用于一致的基准数据比解释个别证券更重要的场景。其公开的时间分区将 1967–1986 用于训练,1987–1991 用于验证,1992–2016 用于测试。文档介绍了如何下载、将作者提供的张量文件转换为 Parquet,并加载完整面板或单独分区,也包括载入宏观数据的选项。这是数据集说明,而非实证结果:表现取决于所用模型和评估方法,匿名标识符也限制了跨分区的公司层面分析。
核心观点
- 该基准包含 1967 至 2016 年的月度股票观测、公司特征、宏观指标和前向收益。
- 预先定义的时间分区将训练、验证和测试时段彼此分开。
- 匿名公司标识符在各张量内保持一致,但无法跨张量命名空间关联。
- 该数据集用于可复现的线性和非线性资产定价实验。
- 转换后的 Parquet 文件保留了从原始张量数据中恢复的标识符。
标签
全文
# Firm Characteristics (Chen-Pelger-Zhu 2020)
# Firm Characteristics (Chen-Pelger-Zhu 2020)
Anonymized panel of ~1.2M stock-month observations with 46 firm
characteristics and forward returns, spanning 1967-2016. Built from the
replication archive of Chen, Pelger, and Zhu (2020), *Deep Learning in
Asset Pricing*. Used throughout the book for ML-based asset-pricing
examples where a standard, reproducible benchmark matters more than
symbol-level interpretation.
## Dataset
- **Source**: GitHub replication repo
(https://github.com/jasonzy121/Deep_Learning_Asset_Pricing), which
itself ships the published dataset via Google Drive
- **Coverage**: 1967-01 to 2016-12, monthly observations, ~1.2M rows
- **Features**: 46 firm characteristics (accounting ratios, price-based
measures, momentum variants), 178 macro indicators, forward returns
- **Size on disk**: ~1.1 GB raw CSV; converted to ~500 MB parquet
- **Access**: Public, no API key required
- **Canonical schema**: `symbol` (anonymous integer id), `timestamp`
(monthly Date), 46 characteristic columns, `ret`, `split`
- **Identity scope**: `symbol` is persistent within each tensor released by
the authors. Separate numeric namespaces prevent accidental linking across
tensors because the archive does not publish a cross-split map.
## Pre-defined Splits
The dataset ships with deterministic train/valid/test splits aligned to
the authors' released tensors:
| Split | Period |
| ----- | --------- |
| train | 1967-1986 |
| valid | 1987-1991 |
| test | 1992-2016 |
## Download
This is the largest free dataset in the book: ~1.5 GB pulled from the
authors' Google Drive folder (RetChar.csv alone is ~1.1 GB), followed by
a tensor-to-Parquet conversion. The tensor conversion recovers persistent
anonymous firm identifiers that the flattened CSV omits. How long it takes
depends on your bandwidth and disk speed. Per-file progress is printed as it
runs. A single command downloads and converts; no separate `--convert` pass
is needed.
```bash
# Download the ~1.5 GB folder and convert to parquet in one step
uv run python data/equities/firm_characteristics/download.py
# Verify what's already on disk, do not refetch
uv run python data/equities/firm_characteristics/download.py --check
# Force a re-download even if files exist
uv run python data/equities/firm_characteristics/download.py --force
# Re-run only the tensor-to-Parquet conversion (files already downloaded)
uv run python data/equities/firm_characteristics/download.py --convert
```
It is also fetched automatically as part of `data/download_all.py`. To
skip it there (e.g. on a metered or space-constrained connection), pass
`--skip-firm-characteristics`.
Output layout under `$ML4T_DATA_PATH/equities/firm_characteristics/`:
```
firm_characteristics_all.parquet # full panel
firm_characteristics_train.parquet # 1967-1986
firm_characteristics_valid.parquet # 1987-1991
firm_characteristics_test.parquet # 1992-2016
dl_asset_pricing/ # raw CSV + NPZ staging
RetChar.csv
Macro.csv
char/Char_{train,valid,test}.npz
macro/macro_{train,valid,test}.npz
RF/RF_{train,valid,test}_normalized_task_1.npz
```
The staging directory is kept so experiments that need the pre-split
NPZ arrays (the original Chen-Pelger-Zhu format) can read them
directly.
## Loading
```python
from data import load_firm_characteristics
df = load_firm_characteristics() # full panel
df = load_firm_characteristics(split="train")
df = load_firm_characteristics(split="test", include_macro=True)
```
If the canonical parquets are missing, the loader raises
`DataNotFoundError` pointing at the download command.
## Consumers
- Chapters 10-16 - standard benchmark for linear and nonlinear
asset-pricing models.
- `case_studies/us_firm_characteristics/` - Chen-Pelger-Zhu replication
pipeline (CV, GBM, latent factor, deep models).在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。