본문으로 건너뛰기
라이브러리 문서 전체

CME 선물 시퀀스 모델: 설계, 누수 통제와 선택

노트북 Machine Learning for Trading

요약

이 노트북은 선물 예측을 위한 LSTM 및 NLinear 시퀀스 모델을 평가합니다. 의사결정마다 특성 행 하나를 사용하는 모델과 달리 시퀀스 모델은 과거 관측값의 순서가 있는 윈도우를 입력으로 받아, 고정 요약에서는 놓칠 수 있는 패턴을 학습할 가능성이 있습니다. NLinear는 각 윈도우의 마지막 값에 대한 변화를 예측하고, LSTM는 윈도우를 따라 정보가 어떻게 전달될지 학습합니다. 정교한 모델은 더 큰 용량을 가질 수 있지만, 선형 기준 모델도 예측 과제에서 경쟁력 있는 성과를 낸 바 있어 두 모델을 모두 시험한다고 설명합니다.

평가에서는 각 윈도우를 예측 시점 전에 끝내고, 윈도우가 제거 구간이나 폴드 경계를 넘지 않도록 하며, 상품과 폴드 사이에서 은닉 상태를 초기화해 누수를 방지합니다. 검증 백테스트를 통해 모든 선언된 에포크 체크포인트를 선택에 활용할 수 있도록 공개하며, 예측값을 바꾸는 MC 드롭아웃도 선언된 모델 구성의 일부로 취급합니다. 문서에는 모델 성과 결과가 없습니다. 주요 한계는 제한된 데이터입니다. 상품이 서른 개뿐이어서 대형 예측 벤치마크보다 시퀀스 사례가 적고 서로 많이 겹치므로, 상품이 많은 환경에서는 결과가 일반화되지 않을 수 있습니다.

핵심 아이디어

  • 시퀀스 모델은 순서가 있는 과거 데이터를 사용해 고정된 특성 요약에 없는 정보를 포착할 수 있습니다.
  • NLinear는 윈도우의 마지막 관측값 대비 변화를 예측하고, LSTM는 단계 간 정보 유지 방식을 학습합니다.
  • 퍼지 간격, 윈도우 종료 시점, 은닉 상태 초기화는 서로 다른 시간 누수 원인을 다룹니다.
  • 검증 기반 선택에 모델 선택 과정이 포함되도록 선언된 모든 체크포인트를 공개합니다.
  • 선물 상품 수가 적으면 독립적인 시퀀스 사례가 줄어 모델 성과가 제한될 수 있습니다.

태그

전문
# CME Futures: Sequence Models


# CME Futures: Sequence Models

This notebook evaluates the declared NLinear and LSTM sequence configurations. Each input window
contains observations from one product and ends before its prediction timestamp. Purge gaps and
fold boundaries prevent a sequence from crossing into another validation interval, and hidden
state does not pass between products or folds.

Every declared epoch checkpoint is published with fitted weights and exact chronological
eligibility. MC dropout is not an undeclared side experiment. Configuration selection remains the
validation backtest decision in `13_backtest`.

Prerequisites: `03_financial_features`, `04_model_based_features`, and `05_evaluation`.

## What a sequence model reads that the other families do not

Every family up to this point saw one row per product per decision: a vector of features
describing that product at that moment. Anything about how it got there had to be engineered
into a column - a 21-session volatility, a momentum composite, a carry z-score against a
rolling window. The model saw the summary, never the path.

A sequence model reads the path. Its input is a window of consecutive observations for one
product, and the architecture is built to make use of their order. The claim being tested is
that the shape of recent history carries information that no fixed set of summary statistics
captured - that a product whose carry rose steadily to its current level differs from one that
spiked and fell back, even where both end at the same value with the same 21-session
volatility.

For thirty futures products the windows are also the scarcest data in the case study. A
feature-row model gets one training example per product per session; a sequence model needs a
whole window per example, so the same history yields fewer independent examples and they
overlap heavily with each other. That is the structural reason to expect these models to
struggle here relative to a benchmark with millions of series, and it is worth holding
alongside whatever the backtest reports.

### Two architectures, and why both

**LSTM** processes the window one step at a time, carrying a hidden state forward and learning
what to keep and what to forget. It is the general answer, and its generality is the cost: it
has many parameters, it trains slowly, and on a short noisy series it has ample capacity to
memorize.

**NLinear** is close to the opposite. It is a linear map from the window to the forecast, with
a normalization step that subtracts the window's last value before the map and adds it back
afterwards. That subtraction is the whole idea: it makes the model predict the *change* from
where the series currently sits rather than the level, which removes the drift that otherwise
dominates a naive fit.

It is here because a body of recent work found that simple linear baselines matched or beat
elaborate sequence architectures on many forecasting benchmarks once evaluated carefully - a
finding that survived enough scrutiny to be worth designing around. Running both is what turns
"the sophisticated model should win" into something this case study measures rather than
assumes.

## Where a sequence model can leak, and what stops it

A window is a span of time rather than a point, which gives leakage more places to enter than
the other families have.

- **Each window ends before its prediction timestamp.** The last observation a window contains
  is strictly earlier than the moment being predicted, so a forecast never reads the bar it is
  forecasting.
- **Purge gaps and fold boundaries stop a window crossing into another interval.** Without them
  a window ending just after a fold boundary would extend back across it, and validation rows
  would be predicted from a window overlapping the training period. The failure would be
  invisible in the output: the prediction is dated correctly and the returns are real.
- **Hidden state does not pass between products or folds.** An LSTM's state accumulates
  whatever it has seen, so carrying it across a boundary carries information across that
  boundary too - and unlike a feature column, the state never appears in any frame, so nothing
  downstream could detect it.

The three are separate mechanisms rather than one guarantee stated three times, which is why
they are enforced separately rather than by a single check on the output.

```python
"""Fit the declared CME futures sequence-model population."""

import polars as pl

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    model_request_catalog,
    open_study,
    product_universe_table,
    resolve_model_requests,
    resolved_model_plan,
    run_official_model_catalog,
    run_resolved_model_requests,
)
from case_studies.research import population_supersedes
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_REDUCTIONS: dict = {}
# The population hash this run replaces, read from the registry and set by a person. A
# first population takes None; a re-run whose membership has changed is refused without
# the hash it supersedes, and the refusal names the value required.
SUPERSEDES_POPULATION: str | None = "8c2c87299a47"
# The device to fit on. Empty means the device this population was published on.
DEVICE: str = ""
# The population this run publishes into. Empty publishes the canonical one, which a run
# on another device may not do.
POPULATION_NAME: str = ""
```

## Declared requests

The request rows identify architecture, label, and published configuration. Sequence length,
checkpoint schedule, seed, gap policy, and device enter the resolved computation identity.

**The device is declared here rather than inherited.** With no override the shared sequence
adapter falls back to a literal `"cuda"` written in `case_studies/utils/deep_learning.py`, and
resolving the request raises `CUDA was requested for sequence training, but CUDA is unavailable`
rather than quietly moving the fit to the CPU. That refusal comes from resolving the request, so
it arrives before any fitting starts. Stating it in the request puts that requirement where a
reader meets it instead of two layers below. The resolved specification hash is the same with the
override as without, so this names what the published run already did.

The device is part of what the fitted model is, not a note beside it: the same architecture
trained on a GPU and on a CPU accumulates its sums in different orders and reaches different
weights. `PUBLISHED_DEVICE` is the device this population was fitted on, and the canonical
population accepts no other. A reader without an NVIDIA card sets `DEVICE="cpu"` and passes a
`POPULATION_NAME` to fit the same grid into a population of its own, which the backtest does
not read.

```python
PUBLISHED_DEVICE = "cuda"
device = DEVICE or PUBLISHED_DEVICE
if device != PUBLISHED_DEVICE and not POPULATION_NAME:
    raise ValueError(
        f"this run fits on device {device!r}, which is not the {PUBLISHED_DEVICE!r} this "
        f"population was published on, so it cannot publish the canonical population; pass "
        f"POPULATION_NAME to give it its own"
    )

population_name = POPULATION_NAME or "cme_futures-deep_learning-validation-v1"

study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
requests = model_request_catalog("deep_learning", labels=ALL_LABELS)
resolved = resolve_model_requests(
    study,
    requests,
    execution_tier=EXECUTION_TIER,
    overrides={"device": device},
    preview_reductions=PREVIEW_REDUCTIONS,
)
universe = product_universe_table()
universe
```

```python
resolved_model_plan(resolved)
```

## Execute and validate

The shared sequence adapter owns window construction, checkpoint reload, prediction coverage, and
restart. A failed configuration cannot remove itself from the population snapshot.

**"Cannot remove itself" is the load-bearing clause.** The natural way to write a sweep is to
catch a failure, log it, and carry on with what worked - which produces a population defined
by what happened to train rather than by what was declared. The leaderboard still looks
sensible, and the configuration that failed is indistinguishable from one that was never
requested. Sequence models make this more likely than the other families do, because they are
the ones that run out of memory or fail to converge on a thin product.

### What MC dropout is, and why it is declared rather than switched on

Dropout during training randomly disables units so the network cannot rely on any one path.
**MC dropout** leaves it enabled at prediction time and runs the forward pass several times, so
each pass gives a slightly different answer and their spread estimates the model's uncertainty
about that prediction.

That is a useful quantity - it is what an allocator sizing inversely to uncertainty would want
- but it changes what the model outputs. A prediction averaged over stochastic passes is not
the same number as the deterministic one, and a run that quietly enabled it would publish
different values under the same configuration name. So it is part of the declared
configuration and enters the identity, which is what the header means by "not an undeclared
side experiment": either the population says these predictions are MC-dropout predictions, or
they are not, and no run gets to decide that on its own.

### Why checkpoints are published rather than chosen

As in `08_tabular_dl`: a neural fit is a trajectory, and choosing the best epoch by validation
performance before reporting that model's validation performance is selection inside the
number being reported. Every declared checkpoint becomes a candidate row and `13_backtest`
selects among them on Sharpe, so the choice sits in the same funnel and the same trial count
as everything else.

```python
if EXECUTION_TIER == "canonical":
    execution, population = run_official_model_catalog(
        study,
        requests,
        population_name=population_name,
        resolved_requests=resolved,
        supersedes=population_supersedes(
            study,
            name=population_name,
            declared=SUPERSEDES_POPULATION,
        ),
    )
else:
    if WORKSPACE is None or not PREVIEW_REDUCTIONS:
        raise ValueError("preview execution requires WORKSPACE and PREVIEW_REDUCTIONS")
    execution = run_resolved_model_requests(study, resolved)
    population = None
```

```python
catalog = execution.catalog_rows.select(
    "family",
    "label",
    "config_name",
    "checkpoint_kind",
    "checkpoint_value",
    "execution_tier",
    "complete",
    "training_hash",
    "prediction_hash",
).sort("label", "config_name", "checkpoint_value")
if catalog.filter(~pl.col("complete")).height:
    raise RuntimeError("sequence execution returned a partial prediction")
catalog
```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.