본문으로 건너뛰기
라이브러리 문서 전체

PCA과 신경망 할인인자: 폴드 안전 선물 신호

노트북 Machine Learning for Trading

요약

이 문서는 선물 수익률을 위한 두 가지 잠재 요인 접근법을 소개합니다. 주성분 분석은 수익률 변동을 가장 많이 설명하는 비상관 방향을 찾고, 신경망 확률적 할인인자는 기대수익률의 횡단면 가격결정과 관련된 조합을 찾으며 캐리나 변동성과 같은 상품 특성을 활용할 수 있습니다. 따라서 두 방법은 서로 다른 질문에 답합니다. 광범위한 동조 움직임과 가격에 반영될 수 있는 수익률 구조를 다루며, 추정된 요인의 경제적 의미는 추가 해석이 필요합니다.

두 접근법은 미래 공분산 정보가 과거 포트폴리오에 영향을 주지 않도록 각 학습 폴드 안에서 적합합니다. 두 방법을 후속 백테스팅에 함께 선언하며, 선택 기준은 정보계수가 아니라 검증 백테스트의 Sharpe입니다. PCA의 분산 목표는 보상되지 않는 움직임을 강조할 수 있고, 할인인자 방법은 사용 도구와 모델 선택에 좌우됩니다. 요인 추정치는 상품 유니버스와 가용 학습 이력에도 달려 있으므로 비교하려면 일관된 유니버스와 명시된 설정이 필요합니다.

핵심 아이디어

  • PCA은 예측 가능성을 목표로 하지 않고 수익률 분산을 설명하는 방향을 추출합니다.
  • 확률적 할인인자는 횡단면 수익률 가격결정을 대상으로 하며 상품 특성을 사용할 수 있습니다.
  • 미래 정보가 포트폴리오 구성에 영향을 주지 않도록 각 학습 폴드 안에서 잠재 요인을 적합합니다.
  • 선물 계약 유니버스가 바뀌면 요인 정의도 달라집니다.
  • 검증 백테스트를 사용해 이 모델 계열을 다른 모델과 함께 평가하고 IC만으로 판단하지 마세요.

태그

전문
# CME Futures: Latent-Factor Requests


# CME Futures: Latent-Factor Requests

The latent-factor stage contains two declared configurations. `10a_pca` fits principal components
within each training fold. `10b_stochastic_discount_factor` estimates the neural stochastic
discount factor within the same fold contract. Neither notebook selects by IC.

This index exposes the complete request population without launching either computation. The two
execution notebooks publish disjoint official populations that `13_backtest` later combines with
the other predictive families.

## What a latent factor is, and how this stage differs from the ones before it

Every model up to this point was handed named predictors. Carry, momentum, the volatility
estimate, the regime probability - each is a quantity somebody decided to compute, and the
model's job was to weigh them. The choice of what to compute came from the researcher, and a
driver nobody thought to name was a driver no model could use.

A latent factor is inferred instead of specified. The starting observation is that futures
returns move together far more than thirty independent series would: energy contracts rise and
fall as a group, the metals do, the equity indices do, and there are days on which nearly
everything moves the same way. That co-movement is evidence of a small number of underlying
drivers acting on many contracts at once. A latent-factor method estimates those drivers from
the covariance of returns themselves, without being told in advance what they are or how many
there should be.

The appeal is that it can find structure nobody encoded. The cost is that what it finds has no
name and no economic interpretation attached - a factor is a direction in return space that
explains variance, and whether it corresponds to anything a reader would recognise is a
separate question the method does not answer.

## Why two configurations, and what separates them

The two are not variations on one method. They disagree about what a factor is *for*, and
that disagreement is the reason both are here.

**`10a_pca` maximizes explained variance.** Principal components find the directions along
which returns vary most, in order, each uncorrelated with the ones before it. It is linear,
it has a closed-form solution, and it makes no reference to returns being predictable at all.
Its first component on a futures panel is typically close to "everything moves together"; the
next few usually separate the sectors. It is the standard baseline for exactly the reasons
equal weight is one in the backtest stage: it is well understood, it estimates little, and
anything more elaborate has to beat it to justify itself.

The weakness is that variance and return are different quantities. The direction along which
a panel varies most is not necessarily the direction that pays, and PCA has no mechanism for
preferring one that does - a factor capturing a large, entirely unrewarded common movement is
exactly what it is built to find first.

**`10b_stochastic_discount_factor` starts from what prices assets.** Asset pricing theory says
that if markets are free of arbitrage there exists a single random variable - the stochastic
discount factor - whose covariance with any asset's return explains that asset's expected
return. Everything that is priced is priced by the same object. The SDF is not observable, but
it is a well-defined thing to estimate, and estimating it with a neural network means not
having to assume in advance which functional form it takes.

The difference from PCA is the objective and the inputs, not the architecture. PCA asks which
directions explain the most variation; the SDF asks which combination best explains the
cross-section of *returns*. A factor that moves a lot but earns nothing is a success for the
first and a failure for the second.

They also see different data, which is easy to miss and changes what each can find.
`run_pca_fold` takes the characteristics panel and discards it with `del`, so PCA is handed
returns alone. `run_sdf_fold` passes the characteristics through, and the number of
instruments it builds is derived from their width. So the SDF can express "products with high
carry and low volatility load on this factor" and PCA structurally cannot, because PCA never
sees carry.

So the comparison between the two is not "which fits better". It is a question about this
panel: whether the directions along which futures returns vary most are also the directions
along which they are compensated. The two configurations are run under the same fold contract
and the same universe precisely so the comparison isolates that.

## Why the factors are fitted inside each fold, and why that matters more here

Both configurations estimate their factors within the training portion of each fold, never
once over the whole panel. That is the same discipline every other family follows, but the
consequence of breaking it is worse here and easier to miss.

A supervised model that saw future data would be caught by its own validation score looking
implausible. A latent-factor model fitted on the full sample fails more quietly: the factors
are estimated from the covariance of returns, so a factor fitted over 2011 to 2025 encodes
which contracts moved together across the entire period. Using it to form a position in 2014
means holding a portfolio constructed from the knowledge that those contracts would go on
co-moving. Nothing about the resulting prediction looks impossible. The returns are real, the
weights are finite, and the backtest runs - it just reports a strategy that could not have
been held.

The cost of doing it correctly is visible in what the early folds can support. A covariance
matrix over thirty products needs a meaningful amount of history before its estimate means
anything, so the earliest training window supports fewer reliable factors than the latest,
and a factor count fixed across folds is a compromise rather than a free choice. That is the
tradeoff the declared configuration is making, and it is the reason the count is declared in
`setup.yaml` rather than selected per fold - selecting it per fold on validation performance
would choose the number that best suited each window's outcomes.

## What the two tables below show

The universe table is the set of products the factors are estimated across. It is worth
reading before the request catalog, because a latent factor is a property of the panel rather
than of any one contract: adding or removing products changes what the factors are, in a way
that changing the universe for a per-product model does not. Two runs over different universes
do not produce comparable factors even under identical settings.

The request catalog is the complete declared population - one row per label and configuration,
resolved but unfitted. Reading it here is what makes the count the execution notebooks produce
checkable against a declaration.

## Why neither selects by IC, and why this page launches nothing

Both fit within each training fold, and both publish predictions like any other family. They
are not privileged by being unsupervised: their rows enter `13_backtest` alongside the linear,
gradient-boosting and sequence families and are selected on validation backtest Sharpe like
everything else. A high IC here decides nothing, which is the same rule the whole case study
runs under.

This notebook computes neither. It exists so the declared request population can be read
before anything is fitted - the two execution notebooks publish disjoint official populations,
and seeing what they *will* contain is what makes a later count checkable against a
declaration rather than against whatever finished.

Disjoint is the part worth noticing. The two populations share no members, so `13_backtest`
combines rather than reconciles them, and a configuration missing from one is not covered by
the other being complete.

```python
"""Show the declared CME futures latent-factor requests."""

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    model_request_catalog,
    product_universe_table,
)
```

```python
requests = model_request_catalog("latent_factors", labels=ALL_LABELS)
universe = product_universe_table()
universe
```

```python
requests.sort("label", "config_name")
```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.