Caractéristiques anonymisées d’entreprises pour la recherche en valorisation
Résumé
Ce document décrit un jeu de données mensuel sur les actions US, conçu pour la recherche en apprentissage automatique sur la valorisation des actifs. Il comprend 94 caractéristiques d’entreprises transformées par rang, notamment des mesures comptables et des caractéristiques techniques, ainsi que les rendements. Les entreprises sont anonymisées et les données incluent celles qui ont été radiées de la cote. Le jeu de données est organisé en périodes d’entraînement et de test séparées par un intervalle destiné à réduire les risques d’anticipation lors de l’évaluation.
Le document explique que les tenseurs de caractéristiques publiés préservent les identités anonymes des entreprises dans chaque bloc de données, tandis qu’un fichier CSV de rendements et caractéristiques ne possède pas la structure d’identité attendue par le chargeur. Il décrit le processus pris en charge de téléchargement et de conversion, le chargement du jeu de données complet ou de découpages prédéfinis, ainsi que l’examen de la couverture et des profils de données. Le jeu de données est statique et s’arrête en 2016 ; il ne peut donc pas servir à étudier des périodes ultérieures sans autre source. L’anonymisation limite le rapprochement des observations avec des titres nommés, et certaines caractéristiques comportent des valeurs manquantes. La séparation entraînement-test est utile pour l’évaluation historique des modèles, mais les chercheurs doivent aussi tenir compte de l’intervalle et du périmètre de publication fixe du jeu de données.
Idées clés
- Le jeu de données fournit des rendements mensuels et 94 caractéristiques d’entreprises pour des entreprises anonymisées US.
- Les périodes d’entraînement et de test sont séparées par un intervalle pour faciliter la recherche hors échantillon.
- Les tenseurs publiés conservent les identités anonymes d’entreprises nécessaires au chargeur pris en charge.
- Les caractéristiques sont classées transversalement, et certaines contiennent des valeurs manquantes.
- Les données statiques s’arrêtent en 2016 et n’identifient pas les entreprises par leur symbole boursier.
Étiquettes
Texte intégral
# Firm Characteristics Dataset
# Firm Characteristics Dataset
Academic dataset of anonymized firm characteristics for ML-based asset pricing.
| Property | Value |
|----------|-------|
| **Provider** | GitHub (Chen, Pelger, Zhu 2020) |
| **Asset Class** | US Equities (anonymized) |
| **Frequency** | Monthly |
| **Firms** | Anonymized |
| **Coverage** | 1967-2016 |
| **Size** | ~258 MB |
| **API Key** | None (free) |
| **Loader** | `load_firm_characteristics()` |
**NOTE**: This is a **static academic dataset**. Firms are anonymized and not updateable.
```python
"""Firm Characteristics - download, explore, and update workflow."""
import polars as pl
```
## 1. Configuration
This is a **static academic dataset** from Chen, Pelger, and Zhu (2020)
"Deep Learning in Asset Pricing". No local configuration file.
### Dataset Characteristics
- **94 firm characteristics**: Accounting ratios, technical indicators, etc.
- **Anonymized firms**: No stock identifiers to prevent data mining
- **Pre-split periods**: Train (1967-1989), Test (2000-2016)
- **Gap period**: 1990-1999 excluded to prevent look-ahead bias
```python
print("=== Firm Characteristics Configuration ===")
print("Provider: GitHub (Chen, Pelger, Zhu 2020)")
print("Paper: 'Deep Learning in Asset Pricing'")
print("Coverage: 1967-1989 (train), 2000-2016 (test)")
print("Features: 94 firm characteristics")
print("Frequency: Monthly")
print("\nThis is a static academic dataset. Firms are anonymized.")
```
## 2. API Key Setup
**No API key required.** This dataset is freely available on GitHub.
```python
print("No API key required - data is hosted on GitHub.")
print("Source: https://github.com/jasonzy121/Deep_Learning_Asset_Pricing")
```
## 3. Download Data
`data/equities/firm_characteristics/download.py` is the only path that produces what
`load_firm_characteristics()` reads. It fetches the archive from the Google Drive folder
the paper's repository links, then converts the published `char/*.npz` tensors - not
`RetChar.csv` - into `equities/firm_characteristics/firm_characteristics_{train,valid,test,all}.parquet`.
The tensors are what carry firm identity. Each block has a fixed anonymous firm axis whose
positions are persistent within the block, so the converter can emit a `symbol` column; the
CSV drops that axis and can only emit `permno`, which the loader rejects. A split offset keeps
the three blocks' identifier namespaces disjoint, because the archive publishes no mapping
between them.
```bash
uv run python data/equities/firm_characteristics/download.py # fetch and convert
uv run python data/equities/firm_characteristics/download.py --check # verify what is there
uv run python data/equities/firm_characteristics/download.py --convert # convert an existing archive
```
This card used to carry a second downloader of its own, writing
`firm_characteristics_{all,train,test}.parquet` into an `academic/` directory from
`RetChar.csv`. Nothing read that directory and the loader rejects that schema, so the files
it produced were unreachable whichever way a reader arrived at them.
```python
from utils import ML4T_DATA_PATH
from utils.paths import display_path
parquet_dir = ML4T_DATA_PATH / "equities" / "firm_characteristics"
present = sorted(path.name for path in parquet_dir.glob("firm_characteristics_*.parquet"))
print("=== Firm Characteristics Download ===")
print("Downloader: data/equities/firm_characteristics/download.py")
print(f"Writes to: {display_path(parquet_dir)}")
print(f"Present: {present or 'nothing yet - run the downloader'}")
```
## 4. Load and Explore
Once downloaded, use the loader throughout the book:
```python
from data import load_firm_characteristics
# Load the full dataset
df = load_firm_characteristics()
print(f"Shape: {df.shape}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
# Count features (exclude timestamp, symbol, ret)
feature_cols = [c for c in df.columns if c not in ["timestamp", "symbol", "ret"]]
print(f"Features: {len(feature_cols)}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")
```
```python
# Schema
df.schema
```
```python
# Preview
df.head(10)
```
### Feature Overview
```python
# Feature statistics
feature_cols = [c for c in df.columns if c not in ["timestamp", "symbol", "split", "ret"]]
print(f"Number of features: {len(feature_cols)}")
print("\nFeature names (first 20):")
for col in feature_cols[:20]:
print(f" {col}")
if len(feature_cols) > 20:
print(f" ... and {len(feature_cols) - 20} more")
```
### Train/Test Split Coverage
```python
# Yearly coverage
yearly = (
df.with_columns(pl.col("timestamp").dt.year().alias("year"))
.group_by("year")
.agg(
pl.len().alias("n_observations"),
)
.sort("year")
)
print("Yearly coverage:")
yearly
```
## 5. Data Profile
```python
from ml4t.data.storage.data_profile import load_profile
from utils import ML4T_DATA_PATH
profile_path = (
ML4T_DATA_PATH / "equities" / "firm_characteristics" / "firm_characteristics_all_profile.json"
)
profile = load_profile(profile_path)
if profile is None:
print(f"No profile at {display_path(profile_path)}")
print(
"The downloader above writes it, next to the data, through\n"
"ml4t.data.storage.data_profile. Re-run it to produce one; there is no separate\n"
"profile-generating script."
)
else:
print("=== Firm Characteristics Profile ===")
print(f"Written by {profile.source}")
print(profile.summary())
```
## 6. Loader Options
The loader supports loading the full dataset or pre-defined splits:
```python
# Load train split (1967-1989)
train = load_firm_characteristics(split="train")
print(f"Train: {train.shape}, {train['timestamp'].min()} to {train['timestamp'].max()}")
```
```python
# Load test split (2000-2016)
test = load_firm_characteristics(split="test")
print(f"Test: {test.shape}, {test['timestamp'].min()} to {test['timestamp'].max()}")
```
```python
# Load full dataset (default)
full = load_firm_characteristics()
print(f"Full: {full.shape}")
```
## 7. Documentation
### Source
- **Paper**: Chen, Pelger, Zhu (2020) "Deep Learning in Asset Pricing"
- **GitHub**: https://github.com/jasonzy121/Deep_Learning_Asset_Pricing
- **Published**: Management Science, 2024
### Dataset Columns
| Column | Description |
|--------|-------------|
| `date` | Month-end date |
| `permno` | Anonymized firm identifier |
| `ret` | Monthly stock return |
| `me` | Market equity |
| `bm` | Book-to-market ratio |
| `mom12m` | 12-month momentum |
| `... (94 total)` | Various firm characteristics |
### Train/Test Split
| Split | Period | Purpose |
|-------|--------|---------|
| Train | 1967-1989 | Model training |
| Gap | 1990-1999 | Excluded (prevents look-ahead) |
| Test | 2000-2016 | Out-of-sample evaluation |
### Data Quality Notes
- **Anonymized firms**: No stock identifiers to prevent data mining
- **Cross-sectional ranking**: Features are rank-transformed
- **Missing values**: Some characteristics have gaps
- **Survivorship**: Includes delisted firms
## 8. Updating Data
**This dataset is NOT updateable.**
This is a static academic dataset published with a research paper.
The data cannot be extended beyond the original publication period (2016).
### Related Resources
For more recent firm characteristics data, consider:
| Resource | Coverage | Access |
|----------|----------|--------|
| WRDS/CRSP | 1926-present | Subscription |
| Open Source Asset Pricing | Varies | Free |
| Ken French Library | 1926-present | Free (factors only) |
## Summary
| Item | Value |
|------|-------|
| Features | 94 firm characteristics |
| Frequency | Monthly |
| Coverage | 1967-2016 (with gap 1990-1999) |
| Provider | GitHub (Chen, Pelger, Zhu 2020) |
| Loader | `load_firm_characteristics(split=None)` |
**Primary use**: Deep learning asset pricing research (Chapters 10-15).
**Limitation**: Anonymized firms, static dataset ends 2016.Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT
Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.