PCAとニューラル割引ファクターによる先物シグナル
ノートブック Machine Learning for Trading
サマリー
この文書では、先物リターンに対する2つの潜在ファクター手法を紹介します。主成分分析はリターン変動の大部分を説明する無相関な方向を見つけます。一方、ニューラル確率的割引ファクターは期待リターンのクロスセクションに関連する組み合わせを探し、キャリーやボラティリティなどの商品特性を利用できます。したがって、手法の問いは異なります。一方は広範な共変動、もう一方は価格付けされている可能性のあるリターン構造を扱います。推定されたファクターの経済的意味は、さらに解釈する必要があります。
将来の共分散情報が過去のポートフォリオ形成に影響しないよう、両手法を各学習分割内で当てはめます。後続のバックテストに向けて両方の要求をまとめて宣言し、情報係数ではなく検証バックテストのシャープレシオで選択します。PCAの分散目的は、報酬につながらない変動を重視する場合があります。一方、割引ファクター法は使用する操作変数とモデルの選択に依存します。ファクター推定値は商品ユニバースと利用可能な学習履歴にも左右されるため、比較には一貫したユニバースと明示された設定が必要です。
主なアイデア
- PCAはリターンの分散を説明する方向を抽出しますが、予測可能性は対象にしません。
- 確率的割引ファクターはリターンのクロスセクションの価格付けを対象とし、商品特性を使えます。
- 将来情報がポートフォリオに影響しないよう、各学習分割内で潜在ファクターを当てはめます。
- 先物ユニバースが変わると、ファクターの定義も変わります。
- この手法群は、ICだけでなく他のモデルとともに検証バックテストで評価します。
タグ
全文
# CME Futures: Latent-Factor Requests
# CME Futures: Latent-Factor Requests
The latent-factor stage contains two declared configurations. `10a_pca` fits principal components
within each training fold. `10b_stochastic_discount_factor` estimates the neural stochastic
discount factor within the same fold contract. Neither notebook selects by IC.
This index exposes the complete request population without launching either computation. The two
execution notebooks publish disjoint official populations that `13_backtest` later combines with
the other predictive families.
## What a latent factor is, and how this stage differs from the ones before it
Every model up to this point was handed named predictors. Carry, momentum, the volatility
estimate, the regime probability - each is a quantity somebody decided to compute, and the
model's job was to weigh them. The choice of what to compute came from the researcher, and a
driver nobody thought to name was a driver no model could use.
A latent factor is inferred instead of specified. The starting observation is that futures
returns move together far more than thirty independent series would: energy contracts rise and
fall as a group, the metals do, the equity indices do, and there are days on which nearly
everything moves the same way. That co-movement is evidence of a small number of underlying
drivers acting on many contracts at once. A latent-factor method estimates those drivers from
the covariance of returns themselves, without being told in advance what they are or how many
there should be.
The appeal is that it can find structure nobody encoded. The cost is that what it finds has no
name and no economic interpretation attached - a factor is a direction in return space that
explains variance, and whether it corresponds to anything a reader would recognise is a
separate question the method does not answer.
## Why two configurations, and what separates them
The two are not variations on one method. They disagree about what a factor is *for*, and
that disagreement is the reason both are here.
**`10a_pca` maximizes explained variance.** Principal components find the directions along
which returns vary most, in order, each uncorrelated with the ones before it. It is linear,
it has a closed-form solution, and it makes no reference to returns being predictable at all.
Its first component on a futures panel is typically close to "everything moves together"; the
next few usually separate the sectors. It is the standard baseline for exactly the reasons
equal weight is one in the backtest stage: it is well understood, it estimates little, and
anything more elaborate has to beat it to justify itself.
The weakness is that variance and return are different quantities. The direction along which
a panel varies most is not necessarily the direction that pays, and PCA has no mechanism for
preferring one that does - a factor capturing a large, entirely unrewarded common movement is
exactly what it is built to find first.
**`10b_stochastic_discount_factor` starts from what prices assets.** Asset pricing theory says
that if markets are free of arbitrage there exists a single random variable - the stochastic
discount factor - whose covariance with any asset's return explains that asset's expected
return. Everything that is priced is priced by the same object. The SDF is not observable, but
it is a well-defined thing to estimate, and estimating it with a neural network means not
having to assume in advance which functional form it takes.
The difference from PCA is the objective and the inputs, not the architecture. PCA asks which
directions explain the most variation; the SDF asks which combination best explains the
cross-section of *returns*. A factor that moves a lot but earns nothing is a success for the
first and a failure for the second.
They also see different data, which is easy to miss and changes what each can find.
`run_pca_fold` takes the characteristics panel and discards it with `del`, so PCA is handed
returns alone. `run_sdf_fold` passes the characteristics through, and the number of
instruments it builds is derived from their width. So the SDF can express "products with high
carry and low volatility load on this factor" and PCA structurally cannot, because PCA never
sees carry.
So the comparison between the two is not "which fits better". It is a question about this
panel: whether the directions along which futures returns vary most are also the directions
along which they are compensated. The two configurations are run under the same fold contract
and the same universe precisely so the comparison isolates that.
## Why the factors are fitted inside each fold, and why that matters more here
Both configurations estimate their factors within the training portion of each fold, never
once over the whole panel. That is the same discipline every other family follows, but the
consequence of breaking it is worse here and easier to miss.
A supervised model that saw future data would be caught by its own validation score looking
implausible. A latent-factor model fitted on the full sample fails more quietly: the factors
are estimated from the covariance of returns, so a factor fitted over 2011 to 2025 encodes
which contracts moved together across the entire period. Using it to form a position in 2014
means holding a portfolio constructed from the knowledge that those contracts would go on
co-moving. Nothing about the resulting prediction looks impossible. The returns are real, the
weights are finite, and the backtest runs - it just reports a strategy that could not have
been held.
The cost of doing it correctly is visible in what the early folds can support. A covariance
matrix over thirty products needs a meaningful amount of history before its estimate means
anything, so the earliest training window supports fewer reliable factors than the latest,
and a factor count fixed across folds is a compromise rather than a free choice. That is the
tradeoff the declared configuration is making, and it is the reason the count is declared in
`setup.yaml` rather than selected per fold - selecting it per fold on validation performance
would choose the number that best suited each window's outcomes.
## What the two tables below show
The universe table is the set of products the factors are estimated across. It is worth
reading before the request catalog, because a latent factor is a property of the panel rather
than of any one contract: adding or removing products changes what the factors are, in a way
that changing the universe for a per-product model does not. Two runs over different universes
do not produce comparable factors even under identical settings.
The request catalog is the complete declared population - one row per label and configuration,
resolved but unfitted. Reading it here is what makes the count the execution notebooks produce
checkable against a declaration.
## Why neither selects by IC, and why this page launches nothing
Both fit within each training fold, and both publish predictions like any other family. They
are not privileged by being unsupervised: their rows enter `13_backtest` alongside the linear,
gradient-boosting and sequence families and are selected on validation backtest Sharpe like
everything else. A high IC here decides nothing, which is the same rule the whole case study
runs under.
This notebook computes neither. It exists so the declared request population can be read
before anything is fitted - the two execution notebooks publish disjoint official populations,
and seeing what they *will* contain is what makes a later count checkable against a
declaration rather than against whatever finished.
Disjoint is the part worth noticing. The two populations share no members, so `13_backtest`
combines rather than reconciles them, and a configuration missing from one is not covered by
the other being complete.
```python
"""Show the declared CME futures latent-factor requests."""
from case_studies.cme_futures.research_workflow import (
ALL_LABELS,
model_request_catalog,
product_universe_table,
)
```
```python
requests = model_request_catalog("latent_factors", labels=ALL_LABELS)
universe = product_universe_table()
universe
```
```python
requests.sort("label", "config_name")
```出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。