본문으로 건너뛰기
라이브러리 문서 전체

S&P 500 옵션용 PatchTST와 통제된 모델 평가

노트북 Machine Learning for Trading

요약

이 노트북은 사전에 정의한 모델 집합의 하나로 PatchTST를 학습해 S&P 500 옵션을 연구합니다. 과거 시계열을 고정 길이 패치로 나누고 토큰으로 임베딩한 뒤, 어텐션을 사용해 구간 내 위치 간 관계를 학습합니다. LSTM의 단계별 순환과 비교하면 어텐션은 떨어진 시퀀스 구간을 직접 연결하고 구간 전체를 병렬 처리할 수 있습니다. 순서 정보는 위치 임베딩을 통해 학습해야 하며, 매개변수 수가 많아 데이터가 제한적일 때 학습이 더 어려울 수 있습니다.

이 노트북은 시퀀스 모델 계열 전반에 공통 프리셋, 폴드, 레이블, 적격성 규칙과 실행 러너를 사용합니다. 윈도우가 폴드 경계 안에 머물도록 하고, 예정된 학습 간격마다 체크포인트를 저장하며, 전체 모델 집합을 등록하기 전에 예측이 선언된 적격 행을 모두 포함하는지 검증합니다. 이를 통해 후보를 비교할 수 있게 하고 체크포인트 선택은 후속 백테스트에 맡깁니다. 이 노트북은 모델과 체크포인트가 빠짐없이 준비되었음을 확인하지만 순위를 매기거나 예측 우위를 주장하지 않으며, 어떤 모델도 수익성이 있다고 보여주지 않습니다.

핵심 아이디어

  • PatchTST는 고정 길이 패치를 임베딩하고 어텐션으로 시계열 구간의 여러 부분을 연결합니다.
  • 어텐션은 멀리 떨어진 세션을 직접 연결하지만 학습된 위치 정보에 의존합니다.
  • 공통 폴드, 레이블, 윈도우, 적격성 규칙으로 시퀀스 아키텍처를 비교할 수 있습니다.
  • 갭을 고려한 윈도우 구성으로 검증 기간 정보가 모델 입력에 들어가는 것을 방지합니다.
  • 후속 백테스트에서 순위를 매기기 전에 체크포인트가 빠짐없이 준비되었는지 확인합니다.

태그

전문
# S&P 500 Options: PatchTST


# S&P 500 Options: PatchTST

This notebook fits the declared PatchTST member of the sequence population snapshotted by
`09_deep_learning`. After publishing every PatchTST checkpoint, it verifies that the complete
NLinear, LSTM, and PatchTST population is present.

Prerequisites: `09_deep_learning` and `09a_lstm`.

**Why the population is declared in one notebook and filled by several.** The set of members is
a claim made once, before any of them is fitted, so that no family can be added or dropped after
its results are visible. This notebook fits the last declared member and then checks that all
three are present, which is the point at which the population becomes readable downstream.

## What this model is, and how it differs from the LSTM beside it

PatchTST cuts the lookback window into fixed-length patches, embeds each patch as a token, and
lets attention weigh every token against every other. Where the LSTM reads the window one
session at a time and carries a state forward, this model sees the whole window at once and
learns which parts of it to look at.

**What the difference buys.** Recurrence reaches a distant session only by carrying information
through every session between; attention reaches it directly, so a pattern that depends on two
separated stretches of the window is easier to represent. It also parallelizes across the
window, where recurrence is sequential by construction.

**What it costs, and why both are run.** Attention has no built-in notion that yesterday is
nearer than last month - the ordering has to be learned from position embeddings rather than
being structural - and it has more parameters to fit from the same data. Running both against
the same population and the same folds is what turns "sequence models on options data" from an
assertion into a comparison, and either can lose.

```python
"""Fit the declared S&P 500 options PatchTST request."""

import polars as pl

from case_studies.sp500_options.research_workflow import (
    ALL_LABELS,
    declared_dl_device,
    model_request_catalog,
    open_study,
    published_dl_device,
    resolve_model_requests,
    resolved_model_plan,
    run_official_model_subset,
    run_resolved_model_requests,
)
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
DEVICE: str = ""

POPULATION_NAME: str = ""
```

### The device the population was fitted on

A network trained on a GPU and the same network trained on a CPU accumulate their sums in a
different order and reach different weights, so the device is part of what the fitted model is
and sits inside the training identity rather than beside it. The device this population was
fitted on is declared once, in `modeling.dl.device` in `config/setup.yaml`, and read from there
by all four deep-learning notebooks rather than retyped in each. On a machine with no NVIDIA
card the run stops here rather than quietly training something else: set `DEVICE="cpu"` and pass
a `POPULATION_NAME` to fit the same requests there, under a name of their own.

```python
CANONICAL_POPULATION_NAME = "sp500-options-sequence-validation-v1"

published_device = published_dl_device()
device = declared_dl_device(DEVICE)
population_name = POPULATION_NAME or CANONICAL_POPULATION_NAME
if device != published_device and population_name == CANONICAL_POPULATION_NAME:
    raise ValueError(
        f"this run fits on {device!r}, not the published {published_device!r}, so its "
        f"identities are not the ones {CANONICAL_POPULATION_NAME!r} holds; pass "
        f"POPULATION_NAME to give them a population of their own"
    )
print(f"training device: {device} (declared: {published_device})")
```

## Declared request

**What the settings decide.** `lookback: 60` and `patch_size: 16` together set the tokens: a
sixty-session window becomes a handful of patches rather than sixty steps, which is what makes
attention affordable here and also what limits its resolution, since nothing inside a patch is
distinguished. `d_model: 64` is the width each patch is embedded into and `n_heads: 4` the
number of attention patterns learned in parallel, so the model can attend to several
relationships at once instead of averaging them into one. `n_layers: 2` stacks that twice.
`dropout: 0.1` is the same regularizer the LSTM uses, and deliberately so: two families whose
regularization differs are not being compared on architecture.

**The configuration is read from a preset, not written here**, so this architecture cannot
quietly differ from the same architecture in another chapter.

**Every label is fitted**, because selection downstream ranks across labels as well as across
configurations, and a label with no candidates cannot be chosen or ruled out.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
requests = model_request_catalog(
    "deep_learning",
    labels=ALL_LABELS,
    config_names=("patchtst",),
)
resolved = resolve_model_requests(
    study,
    requests,
    execution_tier=EXECUTION_TIER,
    overrides={"device": device},
    preview_reductions=PREVIEW_REDUCTIONS,
)
resolved_model_plan(resolved)
```

## Execute and validate

The shared sequence runner owns gap-safe window construction, fold fitting, fitted-state reload,
checkpoint publication, restart, and exact eligible-key validation.

**A checkpoint is part of a configuration, not a detail of how it was fitted.** Training runs for
100 epochs and publishes every fifth, so this one request becomes twenty scored candidates. A
network's validation performance is not monotone in training time, and the epoch at which it
peaks is a property of the fit a reader is entitled to see rather than a number chosen after the
fact. Picking the best epoch after seeing the results is selection, and selection happens once,
downstream, on backtests.

**Gap-safe means a window never reaches across a fold boundary.** A sequence handed to the model
has to end before the validation window opens, or its state carries information from the period
being scored. The runner owns that construction because it is exactly the kind of rule that gets
restated slightly differently in each notebook and is impossible to notice when it is.

```python
if EXECUTION_TIER == "canonical":
    execution, population = run_official_model_subset(
        study,
        resolved,
        population=population_name,
        require_population_complete=True,
    )
else:
    if not WORKSPACE or not PREVIEW_REDUCTIONS:
        raise ValueError("preview execution requires WORKSPACE and PREVIEW_REDUCTIONS")
    execution = run_resolved_model_requests(study, resolved)
    population = None
```

```python
catalog = execution.catalog_rows.select(
    "family",
    "label",
    "config_name",
    "checkpoint_kind",
    "checkpoint_value",
    "execution_tier",
    "complete",
    "training_hash",
    "prediction_hash",
).sort("checkpoint_value")
if catalog.filter(~pl.col("complete")).height:
    raise RuntimeError("PatchTST execution returned a partial checkpoint")
catalog
```

The official sequence population is complete and ready for model analysis and backtesting. This
notebook does not compare configurations or choose a checkpoint.

**What completeness means here, and why it is checked rather than assumed.** Every requested
checkpoint of all three families produced predictions on exactly the rows its eligibility
contract declared, not more and not fewer. A partial checkpoint is refused rather than
published: a downstream comparison against a model scored on part of the panel is not a
comparison, and by the time anyone reads the result the missing part is invisible.

**The three families share one eligibility group, and that is what makes them comparable.** All
of them need the same sixty-session window before a symbol can be scored, so they are eligible
on the same rows and the difference between their numbers is the models. That does not extend to
the cross-sectional families, which score a symbol from its first row; `11_model_analysis` groups
by eligibility for exactly that reason.

**What has and has not been established at this point.** Three architectures have been fitted on
identical folds, identical windows and identical labels, and every checkpoint each of them
published is registered and complete. Nothing has been ranked. A reader who wants to know which
sequence family works better on options data has the material for that question, and the answer
comes from backtests rather than from anything on this page.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.