比较PCA与IPCA:US股票面板中的潜在因子
笔记本 《交易机器学习》
总结
本文介绍潜在因子,即股票之间的共同收益模式;每只股票的载荷表示其对某种模式的暴露。它对比主成分分析与工具变量主成分分析:前者仅从收益面板提取因子和股票载荷,后者将载荷建模为可观测股票特征的函数。这样根据特征进行条件化,使载荷能随特征变化而调整,并为此前未观测到的股票分配暴露。
笔记本本身是两个因子建模笔记本的结果索引;它检查两种模型对每个标签是否都有完整的验证结果,以及比较是否采用相同的交叉验证设计。它不拟合模型,也不判断哪种模型的预测性更强。IPCA依赖特征与载荷之间的关系在股票和时间维度上保持稳定。两种方法都会压缩宽泛的面板并省略个股特有信息;预先声明的因子数量也未与其他方案比较。
核心观点
- PCA不使用股票特征,直接提取共同收益模式。
- IPCA根据可观测特征进行条件化,使因子载荷能随这些特征变化。
- 进行比较要求两种模型均有完整结果,并采用相同的验证设计评分。
- 潜在因子压缩广泛的市场信息,并省略个股特有变动。
- 特征与载荷之间假定的关系未必能在不同股票或时期保持稳定。
标签
全文
# US equities panel: finding the few things three thousand stocks have in common
# US equities panel: finding the few things three thousand stocks have in common
Every model so far has predicted each stock from that stock's own features. But stocks do not
move independently - most of what a broad panel does on any day is one thing happening to all of
it, and a handful of further things happening to overlapping groups of it. A **latent factor**
is one of those common movements: not a column anybody computed, but a pattern extracted from
how the returns move together, with each stock carrying a **loading** saying how much of that
pattern it takes.
Two ways of extracting them are fitted here, and the difference between them is the whole
lesson:
- [`13a_pca`](13a_pca.ipynb) takes the factors from the return panel alone. Principal component
analysis asks which combinations of stocks account for the most common variation, and answers
without being told anything about the stocks. A loading is then a number attached to a stock,
fitted over the training window and carried forward.
- [`13b_ipca`](13b_ipca.ipynb) conditions the loadings on what the stocks *are*. Instrumented
principal components makes a stock's loading a function of its observable characteristics, so
two stocks with the same characteristics load the same way and a stock whose characteristics
change has its loading change with them.
**Why the second exists.** A loading attached to a stock says nothing about a stock that has not
been seen, and cannot move when the stock does. On a panel where names enter and leave and a
company's size and value change over a decade, that is a real limitation rather than a technical
one, and conditioning on characteristics is what removes it. What it costs is a stronger
assumption: that the relation between characteristics and loadings is stable, and is the same
for every stock.
**This notebook runs nothing.** It is the index over the two that do: it names them, opens what
they published, and shows that both are complete. Which of the two is worth more is a predictive
question, and it is answered in [`15_model_analysis`](15_model_analysis.ipynb).
**Learning objectives.** By the end of this notebook you will be able to:
- Say what a latent factor and a loading are, in terms of a panel of returns rather than of an
algorithm.
- State the difference between a loading attached to a stock and a loading conditioned on the
stock's characteristics, and name a situation in which only the second can answer.
- Say what the conditioned version assumes in exchange, and when that assumption would be
uncomfortable.
- Read a table of published latent-factor results and tell a complete one from an incomplete one.
**Book reference**: Chapter 13.
**Prerequisites**: [`13a_pca`](13a_pca.ipynb) and [`13b_ipca`](13b_ipca.ipynb) have published the
results this index reads.
**What it writes**: nothing. It reads.
```python
"""Reference index for the latent-factor execution notebooks."""
import os
from pathlib import Path
import polars as pl
from case_studies.research import Study, open_study
```
```python
CASE_STUDY_ID = "us_equities_panel"
EXECUTION_TIER = "canonical"
WORKSPACE = "experiments"
```
## What the two notebooks published
One row per label per factor model. Read it for two things.
**Both models present at every label.** A label carrying a PCA row and no IPCA row means the
second notebook did not finish there, and the comparison in
[`15_model_analysis`](15_model_analysis.ipynb) would then be measuring a difference between
labels rather than between factor models.
**`cv_identity` the same across the rows being compared.** It records which walk-forward design
a result was fitted and scored under. Two rows with different values measured themselves over
different windows, and ranking them is not a comparison.
Canonical execution reads the released study; preview execution reads an isolated workspace.
```python
if EXECUTION_TIER == "canonical":
# `Study.open` with no workspace, not `open_study`: this notebook writes nothing, and that is
# the call that opens the released study read-only. `open_study` opens it for regeneration.
study = Study.open(CASE_STUDY_ID)
elif EXECUTION_TIER == "preview":
study = open_study(
CASE_STUDY_ID,
execution_tier=EXECUTION_TIER,
workspace=Path(os.environ.get("ML4T_OUTPUT_DIR") or WORKSPACE),
)
else:
raise ValueError(f"Unsupported execution tier: {EXECUTION_TIER!r}")
latent_results = (
study.predictions.table(include_preview=EXECUTION_TIER == "preview")
.filter(
(pl.col("family") == "latent_factors")
& (pl.col("split") == "validation")
& (pl.col("execution_tier") == EXECUTION_TIER)
& pl.col("complete")
)
.select(
"label",
"config_name",
"checkpoint_kind",
"checkpoint_value",
"cv_identity",
"training_hash",
"prediction_hash",
)
.sort("label", "config_name", "checkpoint_kind", "checkpoint_value")
)
# This notebook indexes what `13a_pca` and `13b_ipca` register and computes nothing of its
# own, so it has to run after them. Nothing else enforces that: the filter returns an empty
# frame rather than raising, the frame is the notebook's only result, and a render whose
# single result cell is blank is indistinguishable from a clean run. Empty is a different
# condition from the partial one the prose below anticipates - "if a row is missing above,
# run the notebook that produces it" expects some rows and got none, which means no
# latent-factor model has been fitted at all.
if latent_results.is_empty():
raise ValueError(
"no latent_factors predictions are registered for this execution tier: run "
"13a_pca and 13b_ipca first. This notebook only indexes what they publish."
)
latent_results
```
## What happens next
If a row is missing above, run the notebook that produces it - the execution notebooks reuse an
identity that already exists rather than refitting it, so re-running is cheap and safe.
[`15_model_analysis`](15_model_analysis.ipynb) reads these results alongside the other model
families and asks which ranks the cross-section better.
[`16_backtest`](16_backtest.ipynb) backtests every one of them. This index chooses nothing, and
the number of factors is not tuned anywhere in this case study: each model's preset declares one
and a sweep over that count would be a different experiment.
## What to notice
**The two models answer the same question with different information.** PCA sees only how the
returns moved together; IPCA is additionally told what each stock is. Any difference between
them is what the characteristics were worth, on this panel, under the assumption that the
relation between characteristics and loadings holds across stocks and over time.
**A factor model is a compression, and a compression discards.** A handful of factors summarise
a three-thousand-name panel, so whatever is specific to one stock is by construction not in the
prediction. That is the trade being made rather than a defect: the models before this one are
where stock-specific information lives.
**Known limitations.** The factor count is declared, not searched, so nothing here says the
declared one is right - only what it gives. Both models are fitted on training windows only and
scored on validation folds that have been read many times over by the time a case study reaches
this notebook. And a latent factor has no name: it is a direction in the returns, and reading an
economic story into it is an interpretation this notebook does not support.在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。