본문으로 건너뛰기
라이브러리 문서 전체

주식·옵션 수익률 예측을 위한 패치 트랜스포머

코드 Machine Learning for Trading

요약

이 노트북은 S&P 500 주식 및 옵션 연구에서 주요 레이블을 예측하도록 패치 트랜스포머를 적합합니다. 각 주식의 과거 입력 구간을 연속 패치로 나누고 각 패치를 어텐션 토큰으로 취급합니다. 어텐션은 구간 내 멀리 떨어진 부분을 직접 비교할 수 있고, 패치 방식은 시퀀스 길이를 줄입니다. 패치 크기가 클수록 계산 부담은 줄지만 패치 안의 세부 정보가 압축됩니다. 이 모형은 연구의 다른 딥러닝 아키텍처와 동일한 관측 구간 및 학습 일정을 사용합니다.

학습에는 워크포워드 폴드, 종목별 구간, 학습 데이터에서 산출해 검증에 적용하는 특성 스케일링을 사용합니다. 검증 이력은 관측 가능한 특성으로 사용할 수 있지만 예측 구간의 레이블은 제외합니다. 노트북은 지정된 에포크 체크포인트에서 예측값을 기록하고 학습 과정의 검증 정보 계수를 추적합니다. 근거에는 한계가 있습니다. 패치 모형은 목표 하나에만 지정되며 패치 크기는 다른 값과 비교하지 않고 고정했습니다. 체크포인트는 이후 선택을 위해 등록되며, 이 노트북은 전략을 선택하거나 백테스트하지 않습니다.

핵심 아이디어

  • 패칭은 어텐션이 시퀀스 전반의 관계를 비교하기 전에 시간 구간을 토큰으로 압축합니다.
  • 패치가 크면 시퀀스가 짧아지지만 패치 안의 세밀한 구조가 흐려질 수 있습니다.
  • 워크포워드 학습은 종목별 이력을 사용하고 학습 데이터에서 산출한 스케일링을 검증 데이터에 적용합니다.
  • 저장된 학습 체크포인트마다 검증 정보 계수를 추적합니다.
  • 목표가 하나이고 패치 크기가 고정되어 있어 아키텍처에 대해 입증할 수 있는 범위가 제한됩니다.

태그

전문
# 10_dl_patchtst.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # S&P 500 Equity+Options: Patched Attention
#
# [`09_dl_lstm`](09_dl_lstm.ipynb) fitted the two sequence architectures that read a window one
# observation at a time: a linear map with no memory, and gated memory carried along the window.
# This notebook fits the third the menu declares, which reads the same window a different way.
#
# **A patched transformer cuts the window into contiguous blocks and treats each block as one
# token.** Sixty observations become a handful of patches, and attention runs over those patches
# rather than over the sixty steps. Two things follow. Attention compares every patch with every
# other directly, so a relationship between the start of the window and its end does not have to
# survive being carried step by step the way it does in a recurrent model. And the sequence the
# attention layers see is short, which is what makes attention affordable on a window of this length
# at all.
#
# **The patch size is the modelling choice, and it is a coarseness decision.** Observations inside
# one patch are summarized together before attention ever sees them, so structure finer than a patch
# has to survive that summary. A larger patch buys a shorter sequence and loses resolution; a
# smaller one keeps resolution and gives attention more to compare.
#
# **This architecture is declared on one label, and that limits what the notebook can say.** The
# menu declares `patchtst` for the primary label alone, so unlike its siblings this population has
# no companion rows on the other two targets, and nothing here separates what the architecture does
# from what this particular target rewards. Section 1 reports what the menu declares rather than
# assuming it.
#
# **Learning objectives.** By the end of this notebook you will be able to:
#
# - Say what patching does to a window before attention sees it, and what a larger patch trades
#   away.
# - Explain why attention over patches is affordable where attention over raw steps would not be.
# - Read a curve of out-of-sample ranking accuracy against training epoch and tell a model still
#   learning apart from one fitting its training window.
# - Say what a single-label population cannot establish, and where the comparison that needs more
#   than one label is made instead.
#
# **Book reference**: Chapter 13, Section 13.6 (attention and transformer architectures for time
# series), and Chapter 6, Section 6.7 (Search accounting and run logging) for the run log this
# notebook writes into.
#
# **Prerequisites**: [`05_evaluation`](05_evaluation.ipynb) established the walk-forward folds, and
# [`09_dl_lstm`](09_dl_lstm.ipynb) fitted the rest of the declared `deep_learning` family, whose
# lookback and epoch schedule this configuration shares.
#
# **What it writes**: one training run and one complete validation prediction set per epoch
# checkpoint, in `run_log/registry.db` and under `run_log/training/` and `run_log/predictions/`,
# grouped under a named population. Together with `09_dl_lstm`'s population it covers the declared
# family. **Selection happens in [`14_backtest`](14_backtest.ipynb), not here.**

# %%
"""Fit the declared option-analytics patched-transformer population on the walk-forward folds."""

import plotly.graph_objects as go
import polars as pl

from case_studies.research import (
    declared_labels,
    load_model_configs,
    model_requests,
    open_study,
    population_supersedes,
    primary_label,
    resolved_model_plan,
    run_model_population,
)
from utils.style import COLORS, show_plotly_with_alt

# %% tags=["parameters"]
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
POPULATION_NAME = ""
SUPERSEDES_POPULATION: str = "af58ce93d6d2"
DEVICE: str = ""

# %%
study = open_study(
    "sp500_equity_option_analytics",
    execution_tier=EXECUTION_TIER,
    workspace=WORKSPACE or None,
    entry_point="10_dl_patchtst",
)

# %% [markdown]
# ## 1. Which labels, and which model
#
# `declared_labels` reports every label whose training menu declares `deep_learning`. The
# configuration frame below is the subset of that menu naming `patchtst`, and the gap between the
# two is the fact this notebook has to live with: the other architectures in the family are declared
# on all three regression labels and this one is not.
#
# `patch_size` is how many consecutive observations become one token, `d_model` the width of the
# representation each token is embedded into, `n_heads` how many attention comparisons run in
# parallel, and `n_layers` how many times the sequence is passed through the block. `lookback`
# matches the rest of the family, so the input window is the same one `09_dl_lstm` read and a
# difference between the populations is a difference between architectures rather than samples.
#
# `n_epochs` and `checkpoint_interval` come from the preset rather than from here, because they
# decide how many prediction sets the configuration owes: 100 epochs saved every 5 is 20, and a run
# that quietly trained for fewer would publish a different population under the same name.

# %%
declared_labels(study, "deep_learning")

# %%
configs = load_model_configs(
    study,
    "deep_learning",
    labels=LABELS or None,
    config_names=["patchtst"],
)
declared = load_model_configs(study, "deep_learning", config_names=["patchtst"])
if not configs.height:
    raise RuntimeError("no declared label names patchtst under deep_learning")
configs

# %% [markdown]
# `LABELS` narrows the run below what the menu declares, and a narrowed run declares a different set
# of members than the published population does. A population is immutable once written, so such a
# run must publish under its own name rather than registering an incomplete snapshot under the
# published one.
#
# The device is checked in the same cell. A network trained on a GPU and the same network trained on
# a CPU accumulate their sums in different orders and reach different weights, so the device is part
# of what the fitted model is and is recorded inside the computation's identity rather than beside
# it. The runner refuses to substitute a CPU for a requested GPU rather than publishing a different
# model under the published name, so on a machine with no NVIDIA card this notebook stops at the
# next cell; set `DEVICE="cpu"` and pass a `POPULATION_NAME` to fit the same configuration there.

# %%
PUBLISHED_DEVICE = "cuda"
device = DEVICE or PUBLISHED_DEVICE
print(f"training device: {device}")

narrows = set(zip(configs["label"], configs["config_name"], strict=True)) != set(
    zip(declared["label"], declared["config_name"], strict=True)
)
if (narrows or device != PUBLISHED_DEVICE) and not POPULATION_NAME:
    raise ValueError(
        f"this run declares {configs.height} label-configuration pairs on device {device!r}, "
        f"which is not the declared patchtst catalog on {PUBLISHED_DEVICE!r}, so it cannot "
        f"publish the patched-attention population; pass POPULATION_NAME to give it its own"
    )

# %% [markdown]
# ## 2. Binding the declaration to the data
#
# A menu entry says which network to fit. It does not say which feature columns exist today, where
# the walk-forward folds fall, or which symbol-date pairs have both a feature row and a label - nor,
# for a sequence model, which of those pairs have a full lookback behind them. **Resolving** a
# request goes and finds all of that, and fits nothing, so the plan can be inspected before any
# training starts.
#
# - **`eligible_rows` is below what a tabular family reports on the same label.** A prediction needs
#   a full, gap-free window of prior observations behind it, so what drops out is a stock too new to
#   have accumulated one, or a stretch where the calendar has a hole inside the window. It should
#   match what `09_dl_lstm` reported for this label, because the two share a lookback - a difference
#   would mean the two populations are measured on different samples and their results are not
#   directly comparable.
# - **`folds` equals the number of walk-forward splits** `05_evaluation` established.
# - **`validation_start` and `validation_end` bracket the development sample.** The held-out tail
#   must not appear here: it is scored once, at the end of the case study.
# - **`checkpoints` is the epoch schedule**, and multiplying it by the number of rows gives the
#   number of candidate models this notebook is about to create.
#
# The `training_hash` on each row is the identity of that computation, derived from everything that
# can change its result - the device, the lookback and the patch size included.
# [`RUN_LOG.md`](../RUN_LOG.md#identity) sets out what goes into one.

# %%
requests = model_requests(
    study,
    configs,
    execution_tier=EXECUTION_TIER,
    overrides={"device": device},
    preview_reductions=PREVIEW_REDUCTIONS,
)
resolved = tuple(request.resolve() for request in requests)

plan = resolved_model_plan(resolved)
plan.select(
    "label",
    "config_name",
    "feature_count",
    "eligible_entities",
    "eligible_rows",
    "folds",
    "checkpoints",
    "validation_start",
    "validation_end",
)

# %% [markdown]
# ## 3. Fitting the population
#
# `run_model_population` walks the folds. On each one it cuts the training rows into overlapping
# windows belonging to one stock and ending before the timestamp they predict, replaces missing
# feature values with zero and standardizes each column - the mean and scale measured on the
# training rows and applied unchanged to validation - then trains for the declared epochs, writing
# weights every fifth, and predicts that fold's validation rows from each saved set.
#
# **A window never crosses a stock, and it reads only what was observable.** The observations behind
# a prediction are always that stock's own. They are not confined to the training window: each
# stock's validation history is primed with the rows immediately preceding it, and later validation
# dates read earlier validation rows - feature values already on the table at the timestamp being
# predicted, never a label from the interval that prediction covers. The purge gap the folds impose
# on labels is not crossed.
#
# What the call publishes is a **population**: a named, immutable list of the prediction sets it is
# going to produce, computed from the resolved specification before the first fit. Afterwards every
# member must exist and be complete, which is why a failure fails the whole call rather than
# publishing a population one member short. Everything that finished stays registered, and
# re-running trains only what is missing.
#
# `SUPERSEDES_POPULATION` names the population hash this run replaces, and is empty because this is
# the first generation under this name. Anything that moves a training identity - a changed epoch
# schedule, lookback, patch size or device - produces a different population under the same name,
# and the registry refuses to write it without being told which snapshot it supersedes.

# %%
population_name = POPULATION_NAME or "sp500_equity_option_analytics-patchtst-validation-v1"
execution, population = run_model_population(
    study,
    resolved,
    population_name=population_name,
    supersedes=population_supersedes(study, name=population_name, declared=SUPERSEDES_POPULATION),
)

reused = sum(1 for item in execution.diagnostics if item.get("reused"))
print(
    f"{len(execution.runs)} configurations: {len(execution.runs) - reused} trained, {reused} read"
)
print(f"population {population.name}: {len(population.members)} prediction sets")

# %% [markdown]
# ## 4. What came out
#
# One row per epoch checkpoint. `ic_mean` is the **information coefficient**: on each validation
# date, rank the stocks by the model's prediction, rank them by the return they went on to earn,
# correlate the two rankings, and average that daily correlation over the validation period. Zero is
# no relationship.
#
# `ic_n_days` is how many validation dates produced a defined correlation. A network that has
# settled into predicting nearly the same value for every stock on a date gives that date no spread
# to rank, and the date drops out of the average - so a checkpoint with fewer scored dates has its
# `ic_mean` taken over a different sample from its neighbours', and `full_coverage` marks the rows
# measured on every date this label offers.

# %% tags=["results"]
catalog = (
    execution.catalog_rows.select(
        "config_name",
        "label",
        "task",
        "complete",
        "checkpoint_value",
        "ic_mean",
        "ic_std",
        "ic_n_days",
        "auc_mean_daily",
        "direction_label",
        "n_folds",
        "training_hash",
        "prediction_hash",
    )
    .sort("label", "checkpoint_value")
    .join(configs.select("config_name", "label", "params"), on=["config_name", "label"], how="left")
)

if catalog.filter(~pl.col("complete")).height:
    raise RuntimeError("patched-attention execution returned a partial prediction set")
if set(catalog.get_column("prediction_hash")) != set(population.members):
    raise RuntimeError("the published catalog differs from the population declared before fitting")

catalog = catalog.with_columns(
    full_coverage=pl.col("ic_n_days") == pl.col("ic_n_days").max().over("label")
)

primary = primary_label(study)
present = sorted(set(catalog.get_column("label")))
print(f"{catalog.height} candidate models across {len(present)} label(s)")
print(f"primary label fitted here: {primary in present}")
catalog.select(
    "label",
    "params",
    "checkpoint_value",
    "ic_mean",
    "ic_std",
    "ic_n_days",
    "auc_mean_daily",
    "full_coverage",
)

# %% [markdown]
# ### What more training does
#
# The line traces out-of-sample IC as epochs are added. It separates two things a single
# end-of-training number cannot.
#
# A line that rises and then falls has an interior optimum: the network was still learning, then
# began fitting the training window at the expense of the validation folds. A line that wanders
# around zero without trend never had anything to learn in the first place, and its highest point is
# wherever the noise happened to peak. Both produce a respectable-looking maximum, which is why the
# maximum is not what a configuration is judged on, and why every checkpoint is registered rather
# than only the strongest.

# %%
curve = catalog.filter("full_coverage").sort("label", "checkpoint_value")
fig = go.Figure()
for label in sorted(set(curve.get_column("label"))):
    rows = curve.filter(pl.col("label") == label).sort("checkpoint_value")
    fig.add_trace(
        go.Scatter(
            x=rows.get_column("checkpoint_value").to_list(),
            y=rows.get_column("ic_mean").to_list(),
            mode="lines+markers",
            name=label,
            line=dict(color=COLORS["blue"] if label == primary else COLORS["recede"], width=2),
            marker=dict(size=5),
        )
    )
fig.add_hline(y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"])
fig.update_yaxes(title_text="Mean IC (validation)")
fig.update_xaxes(title_text="Training epochs completed")
fig.update_layout(
    title="Validation IC against training epoch",
    height=460,
    width=880,
    margin=dict(t=90),
)
# The span of the line and where it sits relative to zero are facts about the frame, so the
# description counts them rather than asserting a shape the next run may not reproduce.
spans = "; ".join(
    "{} spans {:+.4f} to {:+.4f} with {} of {} checkpoints above zero".format(
        row["label"], row["lo"], row["hi"], row["above"], row["total"]
    )
    for row in curve.group_by("label")
    .agg(
        lo=pl.col("ic_mean").min(),
        hi=pl.col("ic_mean").max(),
        above=(pl.col("ic_mean") > 0).sum(),
        total=pl.len(),
    )
    .sort("label")
    .iter_rows(named=True)
)
show_plotly_with_alt(
    fig,
    "Line chart of mean validation information coefficient against training epochs completed, one "
    f"line per fitted label with a dashed zero line. Counted from the frame: {spans}.",
)

# %% [markdown]
# ## 5. What to notice
#
# **A single-label population cannot separate the architecture from the target.** The rest of the
# family is fitted on three labels, so a pattern that holds across all three there is evidence about
# the architecture. Here there is one, and whatever this curve does is a statement about this
# architecture on this target, jointly. Reading it as a statement about patched attention in
# general takes more than one label's worth of evidence out of one label.
#
# **The patch size is not searched.** One value is declared and one is fitted, so nothing here says
# whether a coarser or finer patch reads this surface better. Adding a second preset and listing it
# under the label's `deep_learning` menu is what would make that a question this notebook answers;
# as it stands the patch size is a fixed assumption rather than a tested one.
#
# **Nothing here is selected.** Every checkpoint is registered, and what advances is decided in
# [`14_backtest`](14_backtest.ipynb) on validation backtest Sharpe after costs.
#
# **Next**: [`11_latent_factors`](11_latent_factors.ipynb) turns to the latent-factor families, and
# [`13_model_analysis`](13_model_analysis.ipynb) compares every family fitted so far.

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.