使用 PatchTST 进行 NASDAQ-100 分钟级微观结构预测
代码 《交易机器学习》
总结
本笔记介绍 PatchTST 如何利用固定长度的连续市场特征片段预测 NASDAQ-100 标签。注意力机制在过去一小时的数据中关联各个片段,相较于逐分钟进行注意力计算,可缩短序列长度。该设置在滚动折上拟合多个预测期限的回归目标;笔记也说明当前序列运行器为何排除了分类目标。
工作流会在拟合前确定数据资格和折边界,按照各标签期限设置训练窗口间隔以减少结果重叠,仅使用训练数据缩放特征,并在预先声明的检查点保存验证预测。它会检查已发布样本总体中的每个预测集是否完整,并绘制各训练轮次的验证集信息系数。笔记提醒,序列模型的可用样本行数少于表格模型,限制了直接比较;固定回看窗口也无法使用一小时之前的信息。信息系数图仅用于诊断;模型选择留待单独的验证回测。
核心观点
- PatchTST 将连续的短时间段作为标记,使注意力机制能以较低计算成本比较序列中相距较远的部分。
- 根据预测期限设置训练窗口间隔,以减少重复且重叠的标签。
- 滚动折、仅使用训练数据进行缩放,以及按股票划分窗口,有助于防止验证数据和不同股票之间的信息泄漏。
- 各检查点预测构成独立的验证候选项;信息系数曲线不决定最终模型选择。
- 固定回看窗口和较少的可用样本限制了模型可见的信息,也增加了与表格模型比较的难度。
标签
全文
# 11_dl_patchtst.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # PatchTST - NASDAQ-100 Microstructure
#
# PatchTST cuts the window into short consecutive segments - patches - and treats
# each patch as one item in a sequence. An attention mechanism then lets every
# patch weigh every other patch directly, so the model can relate two stretches
# of the window without passing information through everything in between.
#
# Patching is what makes that affordable. Attention costs grow with the square of
# the number of items it compares, so attending over 60 individual observations
# is far more expensive than attending over a handful of patches covering the
# same window. Patching also gives each item more content than a single
# observation, which matters when any one minute is mostly noise.
#
# The label looks 15 minutes ahead on a one-minute grid, so consecutive windows
# overlap heavily and neighbouring rows are far from independent. That shapes
# everything below: how windows are built, where folds are cut, and how much of
# the training set is sampled.
#
# **Learning Objectives**:
# - Fit an attention-based sequence model on a panel by declaring one request
# rather than assembling folds and windows in the notebook
# - Read a learning curve across training epochs and say what it shows about
# capacity and noise
# - Check that a fitted model produced predictions on every fold it was asked
# for, before any of those predictions are used
#
# **Book Reference**: Chapter 13
#
# **Prerequisites**: [`05_evaluation`](05_evaluation.ipynb)
#
# **What it writes**: one training run per label and one complete validation prediction set per
# label and checkpoint, in `run_log/registry.db` and under `run_log/training/` and
# `run_log/predictions/`, grouped under a population this notebook alone publishes.
# [`13_model_analysis`](13_model_analysis.ipynb) reads that population beside the other
# families. **It selects nothing**: selection is validation backtest Sharpe in
# [`14_backtest`](14_backtest.ipynb).
# %%
"""Fit the declared NASDAQ-100 microstructure PatchTST population on the walk-forward folds."""
import plotly.graph_objects as go
import polars as pl
from case_studies.research import (
declared_labels,
load_model_configs,
model_requests,
open_study,
resolved_model_plan,
run_model_population,
)
from utils.style import COLORS, show_plotly_with_alt
# %% tags=["parameters"]
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
CONFIG_NAMES: list[str] = []
POPULATION_NAME = ""
SUPERSEDES_POPULATION: str = ""
DEVICE: str = ""
# %%
study = open_study(
"nasdaq100_microstructure",
execution_tier=EXECUTION_TIER,
workspace=WORKSPACE or None,
entry_point="11_dl_patchtst",
)
# %% [markdown]
# ## 1. Which labels, and which model
#
# The labels are the ones whose training menu declares `deep_learning`, and fitting all of them
# in one run is what makes this population comparable against the linear and gradient boosting
# ones: the families differ, the targets do not. `fwd_ret_15m` is the return over the fifteen
# minutes after the decision minute and the horizon the strategy chapters trade; `fwd_ret_5m`
# and `fwd_ret_60m` are the same construction at shorter and longer horizons.
#
# The classification label `fwd_dir_15m` is absent, and not by oversight. The sequence runner
# refuses a non-regression label outright - `case_studies/utils/deep_learning.py`, "sequence
# runner currently supports regression labels only" - so that label declares `linear` and `gbm`
# and nothing else.
# %%
declared_labels(study, "deep_learning")
# %% [markdown]
# `patchtst` is this notebook's slice of the declared family. The menu declares four
# architectures and each has its own notebook, because each is a different claim about what
# structure in the window matters. They resolve against the same menu, the same folds and the
# same windows, so a difference between their results is a difference between architectures.
#
# `lookback` is how many prior one-minute observations enter a window - 60 gives the model the
# trailing hour - and it is the same across all four, so the sample they are measured on is the
# same. `n_epochs` and `checkpoint_interval` are declared with the architecture rather than
# passed in here, because together they decide how many prediction sets each configuration owes:
# 100 epochs saved every 5 is 20, and a run that quietly trained for fewer would publish a
# different population under the same name.
# %%
SEQUENCE_CONFIG = "patchtst"
declared = load_model_configs(study, "deep_learning", config_names=[SEQUENCE_CONFIG])
configs = load_model_configs(
study,
"deep_learning",
labels=LABELS or None,
config_names=CONFIG_NAMES or [SEQUENCE_CONFIG],
)
configs
# %% [markdown]
# `LABELS` and `CONFIG_NAMES` narrow the run below this notebook's own slice, and a narrowed run
# declares a different member set than the published population does. A population is immutable
# once written, so such a run must publish under its own name.
#
# The device is checked in the same cell, because it is inside the training identity rather than
# beside it: a network trained on a GPU and the same network trained on a CPU accumulate their
# sums in different orders and reach different weights. The runner refuses to substitute a CPU
# for a requested GPU rather than publishing a different model under the published name, so on a
# machine with no NVIDIA card this notebook stops at the next cell; set `DEVICE="cpu"` and pass a
# `POPULATION_NAME` to fit the same grid there.
# %%
PUBLISHED_DEVICE = "gpu"
device = DEVICE or PUBLISHED_DEVICE
print(f"training device: {device}")
narrows = set(zip(configs["label"], configs["config_name"], strict=True)) != set(
zip(declared["label"], declared["config_name"], strict=True)
)
if (narrows or device != PUBLISHED_DEVICE) and not POPULATION_NAME:
raise ValueError(
f"this run declares {configs.height} label-configuration pairs on device {device!r}, "
f"which is not this notebook's declared slice on {PUBLISHED_DEVICE!r}, so it cannot "
f"publish the {SEQUENCE_CONFIG} population; pass POPULATION_NAME to give it its own"
)
# %% [markdown]
# ## 2. Binding the declarations to the data
#
# A menu entry says which network to fit. It does not say which feature columns exist today,
# where the walk-forward folds fall, or which symbol-minute pairs have both a feature row and a
# label - nor, for a sequence model, which of those have sixty prior observations behind them.
# **Resolving** a request goes and finds all of that, and fits nothing, so the plan can be read
# before any training starts.
#
# Three things to check in it.
#
# - **`eligible_rows` is below what the linear and gradient boosting families report on the same
# label.** A prediction needs a full, gap-free window behind it, so what drops out is a name
# too new to have accumulated one, or a stretch where the session boundary falls inside the
# window. Comparing a sequence result against a tabular one is therefore comparing measurements
# on different samples, which [`13_model_analysis`](13_model_analysis.ipynb) has to account for.
# - **`folds` is the same everywhere** and equals the walk-forward splits `05_evaluation`
# established.
# - **`validation_start` and `validation_end` bracket the development sample.** The held-out tail
# must not appear; it is scored once, at the end of the case study.
#
# **How far apart the training windows sit is declared, not left to the tier.** Every row of this
# panel could start a window, and the panel is minute bars, so an unspaced fold would build about
# four million near-identical overlapping sequences - neighbouring windows would share 59 of their
# 60 observations. `modeling.dl.train_sequence_stride_horizons` in `config/setup.yaml` declares
# the spacing instead: one window per label horizon, so consecutive windows of a symbol carry
# labels that do not overlap. [`02_labels`](02_labels.ipynb) takes the decision on the bar
# closing at t, enters on the next bar's VWAP and holds for the horizon from that fill, so two
# decisions H observations apart have return intervals that abut rather than share bars - the
# exit of one is the entry of the next. It is declared in horizons rather than in windows because the
# horizon differs by label: `fwd_ret_5m` strides 5 observations, `fwd_ret_15m` 15 and
# `fwd_ret_60m` 60, each spacing its own labels exactly, where a single window count could only
# have been right for one of them. The spacing travels in the training identity; **how many
# windows it yields is derived at run time** from the eligible rows of each fold rather than
# declared, and is printed below.
# %%
requests = model_requests(
study,
configs,
execution_tier=EXECUTION_TIER,
overrides={"device": device},
preview_reductions=PREVIEW_REDUCTIONS,
)
resolved = tuple(request.resolve() for request in requests)
plan = resolved_model_plan(resolved)
plan.select(
"label",
"config_name",
"feature_count",
"eligible_rows",
"folds",
"checkpoints",
"validation_start",
"validation_end",
)
# %% [markdown]
# ## 3. Fitting the population
#
# `run_model_population` fits every resolved request. For one request it walks the folds, and on
# each one:
#
# 1. takes the rows inside that fold's training window and cuts them into overlapping windows of
# sixty one-minute observations, each belonging to one stock and ending before the minute it
# predicts, up to the declared cap,
# 2. standardizes each column on the training rows and applies that scale unchanged to the
# validation rows, so nothing measured on the validation window reaches the fit,
# 3. trains for the declared number of epochs, writing the weights to disk at each checkpoint,
# 4. predicts the fold's validation rows from each saved set of weights.
#
# **A window never crosses a stock, and it reads only what was observable at the minute it
# predicts.** Hidden state is reset between stocks and between folds. What a window carries is
# feature values already on the table at that minute, never a label from the interval the
# prediction covers, so the purge the folds impose is not crossed.
#
# Step 4 is what makes one training run produce twenty results. Each checkpoint's fold
# predictions are concatenated into one series covering the whole validation period, and each
# becomes its own registered prediction set with its own identity.
#
# **What the call publishes is a population**: a named, immutable list of the prediction sets it
# will produce, written down before the first fit. Afterwards every member must exist and be
# complete, which is why a configuration that raises fails the whole call rather than publishing
# a population one member short. Everything that finished stays registered, and re-running trains
# only what is missing.
#
# `SUPERSEDES_POPULATION` names the population hash this run replaces. Anything that moves a
# training identity - a changed epoch schedule, lookback, sequence cap or device as much as a
# changed menu - produces a different population under the same name, and the registry refuses to
# write it without being told which snapshot it supersedes. It refuses before the first fit, so
# the cost of forgetting it is seconds rather than the run.
# %%
population_name = POPULATION_NAME or "nasdaq100_microstructure-patchtst-validation-v1"
execution, population = run_model_population(
study,
resolved,
population_name=population_name,
supersedes=SUPERSEDES_POPULATION or None,
)
reused = sum(1 for item in execution.diagnostics if item.get("reused"))
print(
f"{len(execution.runs)} configurations: {len(execution.runs) - reused} trained, {reused} read"
)
print(f"population {population.name}: {len(population.members)} prediction sets")
# %% [markdown]
# `reused` is not zero on a second run. Every identity is re-derived from the inputs, the
# registry already holds the matching rows and the saved weights, and the runner returns the
# stored result rather than training again.
# %% [markdown]
# ## 4. What came out
#
# The learning curve is the information coefficient on the validation rows at each checkpoint,
# which is a rank correlation between the prediction and the realized return. It is read across
# checkpoints of one fit rather than across configurations, so it says what more training did to
# this model rather than what a different model would have done.
#
# The figure is drawn to show how the coefficient moves across checkpoints rather than what level
# it reaches, and [`13_model_analysis`](13_model_analysis.ipynb) compares the four architectures
# on the same rows.
# %% tags=["results"]
# Scoped to this population's own members. `catalog_rows` is the study's whole prediction
# table, so an unfiltered height check would compare this run against every sequence row the
# registry holds and pass or fail for reasons that have nothing to do with it.
catalog = execution.catalog_rows.filter(
pl.col("prediction_hash").is_in(list(population.members))
).sort("label", "checkpoint_value")
if catalog.height != len(population.members) or catalog.filter(~pl.col("complete")).height:
raise RuntimeError(
f"the {SEQUENCE_CONFIG} population declares {len(population.members)} members and the "
f"registry holds {catalog.height} complete ones"
)
catalog.select("label", "config_name", "checkpoint_value", "ic_mean", "complete")
# %%
fig = go.Figure()
for _label in catalog["label"].unique().sort().to_list():
_rows = catalog.filter(pl.col("label") == _label).sort("checkpoint_value")
fig.add_scatter(
x=_rows["checkpoint_value"].to_list(),
y=_rows["ic_mean"].to_list(),
mode="lines+markers",
name=_label,
)
fig.add_hline(y=0, line_width=1, line_dash="dash", line_color=COLORS["recede"])
fig.update_layout(
title=f"Validation IC by checkpoint - {SEQUENCE_CONFIG}",
xaxis_title="training epoch",
yaxis_title="information coefficient",
)
show_plotly_with_alt(
fig,
"Validation information coefficient against training epoch, one line per label, with a "
"dashed line at zero.",
)
# %% [markdown]
# ## Key takeaways
#
# 1. **A checkpoint is part of a configuration, not a detail of how it was fitted.** Scoring one
# fit at twenty points produces twenty candidates, and keeping whichever of them scores
# highest on this page would be a selection decision taken on the wrong statistic: the
# curve below is an information coefficient, and selection is validation backtest Sharpe in
# [`14_backtest`](14_backtest.ipynb), over the population published here.
#
# 2. **How far apart the training windows sit is part of the model.** On a minute panel the
# spacing decides what was fitted, so it is declared in `config/setup.yaml` as one window per
# label horizon and travels in the training identity rather than arriving with the invocation.
#
# 3. **A sequence family is measured on fewer rows than a tabular one.** A prediction needs a
# full window behind it, so the samples differ and the comparison has to say so.
#
# **Known limitations.** The window is fixed at sixty observations, so nothing earlier than the
# trailing hour reaches the model whatever the architecture can represent. The window spacing is a
# modelling choice rather than a compute budget - one window per label horizon is what stops two
# training examples carrying the same return twice - so the number of windows a fold yields
# follows from how many eligible rows it holds, and a panel of a different size changes the count
# without changing the declaration.
```在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。