使用匿名公司特征开展资产定价研究
笔记本 《交易机器学习》
总结
本文介绍一个按月构建的 US 股票数据集,专为资产定价机器学习研究设计。数据集包含 94 个经过横截面排名变换的公司特征,包括会计指标和技术特征,以及收益数据。公司身份经过匿名化处理,数据也包含已退市公司。数据集划分为训练期和测试期,中间设置了间隔,以减少评估中的前视偏差问题。
本文说明,已发布的特征张量在每个数据块中保留匿名公司身份,而收益—特征 CSV 文件缺少加载器所需的身份结构。文中概述了受支持的下载和转换流程、如何加载完整数据集或预定义的数据划分,以及如何检查覆盖范围和数据概况。该数据集是静态的,止于 2016;若无其他数据源,无法用于此后时期的研究。匿名化限制了将观测值关联到具名证券的能力,部分特征也有缺失值。训练—测试结构适用于历史模型评估,但研究者仍需考虑数据集的间隔和固定的发布范围。
核心观点
- 该数据集提供匿名 US 家公司的月度收益和 94 个公司特征。
- 训练期与测试期之间留有间隔,以支持样本外研究。
- 已发布的张量保留了受支持加载器所需的匿名公司身份。
- 特征按横截面排名变换,部分公司特征存在缺失值。
- 静态数据止于 2016,且不通过股票代码识别公司。
标签
全文
# Firm Characteristics Dataset
# Firm Characteristics Dataset
Academic dataset of anonymized firm characteristics for ML-based asset pricing.
| Property | Value |
|----------|-------|
| **Provider** | GitHub (Chen, Pelger, Zhu 2020) |
| **Asset Class** | US Equities (anonymized) |
| **Frequency** | Monthly |
| **Firms** | Anonymized |
| **Coverage** | 1967-2016 |
| **Size** | ~258 MB |
| **API Key** | None (free) |
| **Loader** | `load_firm_characteristics()` |
**NOTE**: This is a **static academic dataset**. Firms are anonymized and not updateable.
```python
"""Firm Characteristics - download, explore, and update workflow."""
import polars as pl
```
## 1. Configuration
This is a **static academic dataset** from Chen, Pelger, and Zhu (2020)
"Deep Learning in Asset Pricing". No local configuration file.
### Dataset Characteristics
- **94 firm characteristics**: Accounting ratios, technical indicators, etc.
- **Anonymized firms**: No stock identifiers to prevent data mining
- **Pre-split periods**: Train (1967-1989), Test (2000-2016)
- **Gap period**: 1990-1999 excluded to prevent look-ahead bias
```python
print("=== Firm Characteristics Configuration ===")
print("Provider: GitHub (Chen, Pelger, Zhu 2020)")
print("Paper: 'Deep Learning in Asset Pricing'")
print("Coverage: 1967-1989 (train), 2000-2016 (test)")
print("Features: 94 firm characteristics")
print("Frequency: Monthly")
print("\nThis is a static academic dataset. Firms are anonymized.")
```
## 2. API Key Setup
**No API key required.** This dataset is freely available on GitHub.
```python
print("No API key required - data is hosted on GitHub.")
print("Source: https://github.com/jasonzy121/Deep_Learning_Asset_Pricing")
```
## 3. Download Data
`data/equities/firm_characteristics/download.py` is the only path that produces what
`load_firm_characteristics()` reads. It fetches the archive from the Google Drive folder
the paper's repository links, then converts the published `char/*.npz` tensors - not
`RetChar.csv` - into `equities/firm_characteristics/firm_characteristics_{train,valid,test,all}.parquet`.
The tensors are what carry firm identity. Each block has a fixed anonymous firm axis whose
positions are persistent within the block, so the converter can emit a `symbol` column; the
CSV drops that axis and can only emit `permno`, which the loader rejects. A split offset keeps
the three blocks' identifier namespaces disjoint, because the archive publishes no mapping
between them.
```bash
uv run python data/equities/firm_characteristics/download.py # fetch and convert
uv run python data/equities/firm_characteristics/download.py --check # verify what is there
uv run python data/equities/firm_characteristics/download.py --convert # convert an existing archive
```
This card used to carry a second downloader of its own, writing
`firm_characteristics_{all,train,test}.parquet` into an `academic/` directory from
`RetChar.csv`. Nothing read that directory and the loader rejects that schema, so the files
it produced were unreachable whichever way a reader arrived at them.
```python
from utils import ML4T_DATA_PATH
from utils.paths import display_path
parquet_dir = ML4T_DATA_PATH / "equities" / "firm_characteristics"
present = sorted(path.name for path in parquet_dir.glob("firm_characteristics_*.parquet"))
print("=== Firm Characteristics Download ===")
print("Downloader: data/equities/firm_characteristics/download.py")
print(f"Writes to: {display_path(parquet_dir)}")
print(f"Present: {present or 'nothing yet - run the downloader'}")
```
## 4. Load and Explore
Once downloaded, use the loader throughout the book:
```python
from data import load_firm_characteristics
# Load the full dataset
df = load_firm_characteristics()
print(f"Shape: {df.shape}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
# Count features (exclude timestamp, symbol, ret)
feature_cols = [c for c in df.columns if c not in ["timestamp", "symbol", "ret"]]
print(f"Features: {len(feature_cols)}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")
```
```python
# Schema
df.schema
```
```python
# Preview
df.head(10)
```
### Feature Overview
```python
# Feature statistics
feature_cols = [c for c in df.columns if c not in ["timestamp", "symbol", "split", "ret"]]
print(f"Number of features: {len(feature_cols)}")
print("\nFeature names (first 20):")
for col in feature_cols[:20]:
print(f" {col}")
if len(feature_cols) > 20:
print(f" ... and {len(feature_cols) - 20} more")
```
### Train/Test Split Coverage
```python
# Yearly coverage
yearly = (
df.with_columns(pl.col("timestamp").dt.year().alias("year"))
.group_by("year")
.agg(
pl.len().alias("n_observations"),
)
.sort("year")
)
print("Yearly coverage:")
yearly
```
## 5. Data Profile
```python
from ml4t.data.storage.data_profile import load_profile
from utils import ML4T_DATA_PATH
profile_path = (
ML4T_DATA_PATH / "equities" / "firm_characteristics" / "firm_characteristics_all_profile.json"
)
profile = load_profile(profile_path)
if profile is None:
print(f"No profile at {display_path(profile_path)}")
print(
"The downloader above writes it, next to the data, through\n"
"ml4t.data.storage.data_profile. Re-run it to produce one; there is no separate\n"
"profile-generating script."
)
else:
print("=== Firm Characteristics Profile ===")
print(f"Written by {profile.source}")
print(profile.summary())
```
## 6. Loader Options
The loader supports loading the full dataset or pre-defined splits:
```python
# Load train split (1967-1989)
train = load_firm_characteristics(split="train")
print(f"Train: {train.shape}, {train['timestamp'].min()} to {train['timestamp'].max()}")
```
```python
# Load test split (2000-2016)
test = load_firm_characteristics(split="test")
print(f"Test: {test.shape}, {test['timestamp'].min()} to {test['timestamp'].max()}")
```
```python
# Load full dataset (default)
full = load_firm_characteristics()
print(f"Full: {full.shape}")
```
## 7. Documentation
### Source
- **Paper**: Chen, Pelger, Zhu (2020) "Deep Learning in Asset Pricing"
- **GitHub**: https://github.com/jasonzy121/Deep_Learning_Asset_Pricing
- **Published**: Management Science, 2024
### Dataset Columns
| Column | Description |
|--------|-------------|
| `date` | Month-end date |
| `permno` | Anonymized firm identifier |
| `ret` | Monthly stock return |
| `me` | Market equity |
| `bm` | Book-to-market ratio |
| `mom12m` | 12-month momentum |
| `... (94 total)` | Various firm characteristics |
### Train/Test Split
| Split | Period | Purpose |
|-------|--------|---------|
| Train | 1967-1989 | Model training |
| Gap | 1990-1999 | Excluded (prevents look-ahead) |
| Test | 2000-2016 | Out-of-sample evaluation |
### Data Quality Notes
- **Anonymized firms**: No stock identifiers to prevent data mining
- **Cross-sectional ranking**: Features are rank-transformed
- **Missing values**: Some characteristics have gaps
- **Survivorship**: Includes delisted firms
## 8. Updating Data
**This dataset is NOT updateable.**
This is a static academic dataset published with a research paper.
The data cannot be extended beyond the original publication period (2016).
### Related Resources
For more recent firm characteristics data, consider:
| Resource | Coverage | Access |
|----------|----------|--------|
| WRDS/CRSP | 1926-present | Subscription |
| Open Source Asset Pricing | Varies | Free |
| Ken French Library | 1926-present | Free (factors only) |
## Summary
| Item | Value |
|------|-------|
| Features | 94 firm characteristics |
| Frequency | Monthly |
| Coverage | 1967-2016 (with gap 1990-1999) |
| Provider | GitHub (Chen, Pelger, Zhu 2020) |
| Loader | `load_firm_characteristics(split=None)` |
**Primary use**: Deep learning asset pricing research (Chapters 10-15).
**Limitation**: Anonymized firms, static dataset ends 2016.在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。