時間軸間隔を空けたウィンドウで日中の注文フローをLSTM学習
ノートブック Machine Learning for Trading
サマリー
このノートブックでは、NASDAQ-100のマイクロストラクチャ特徴量を使い、数分単位の複数の期間で将来リターンを予測する長短期記憶ネットワークを学習させます。平坦化した線形モデルと異なり、LSTMは各観測値を順に処理し、過去の情報をどれだけ保持するかを学習します。モデルは1分データの過去ウィンドウを使い、銘柄やフォールドが変わると隠れ状態をリセットします。各ウィンドウは同一銘柄内に収まり、予測時点で利用可能な情報だけを使う必要があります。
隣接する分足ウィンドウは大きく重なり、ラベルも将来の期間にまたがるため、隣り合う学習ターゲットが重ならないよう、各ラベル期間につき1つの例を抽出します。スケーリングは学習行で適合し、検証時にも変更せず使います。ウォークフォワード分割は定めた時系列分割を維持します。各学習チェックポイントで検証用の予測セットを作り、各エポックの情報係数を図示します。モデル選択は後の検証用バックテストにおけるSharpe比較で行います。固定した直近1時間の入力ではそれ以前の履歴を使えず、系列モデルは表形式モデルより対象行が少ないため、性能比較ではサンプルの違いを考慮する必要があります。このノートブックは研究手順を示すもので、LSTMが他のモデル群より優れている証拠ではありません。
主なアイデア
- LSTMは観測値の系列を処理しながら、情報を保持するか取り込むかを学習します。
- 学習ターゲットの重複を減らすため、日中の例は各ラベル期間に応じた間隔で抽出します。
- 各ウィンドウは単一銘柄内に収まり、予測時点で観測可能な特徴量だけを使います。
- 学習データのみでのスケーリングとウォークフォワード分割は、検証情報が学習に入るのを防ぐのに役立ちます。
- 各チェックポイントを評価しますが、モデル選択は情報係数だけではなく、後の検証用バックテストで行います。
タグ
全文
# LSTM - NASDAQ-100 Microstructure
# LSTM - NASDAQ-100 Microstructure
A long short-term memory network reads the window one observation at a time
and carries a running summary forward, deciding at each step how much of that
summary to keep and how much of the new observation to admit. Those decisions
are learned rather than fixed, which is what lets the model hold on to
something that happened early in the window if it turns out to matter, and
discard it otherwise.
That is a real difference from a model that maps a flattened window through one
linear layer: the same feature carries different weight depending on what
preceded it. Whether it is a *useful* difference on order-flow features at this
horizon is the question the fit answers.
The label looks 15 minutes ahead on a one-minute grid, so consecutive windows
overlap heavily and neighbouring rows are far from independent. That shapes
everything below: how windows are built, where folds are cut, and how much of
the training set is sampled.
**Learning Objectives**:
- Fit a recurrent sequence model on a panel by declaring one request rather
than assembling folds and windows in the notebook
- Read a learning curve across training epochs and say what it shows about
capacity and noise
- Check that a fitted model produced predictions on every fold it was asked
for, before any of those predictions are used
**Book Reference**: Chapter 13
**Prerequisites**: [`05_evaluation`](05_evaluation.ipynb)
**What it writes**: one training run per label and one complete validation prediction set per
label and checkpoint, in `run_log/registry.db` and under `run_log/training/` and
`run_log/predictions/`, grouped under a population this notebook alone publishes.
[`13_model_analysis`](13_model_analysis.ipynb) reads that population beside the other
families. **It selects nothing**: selection is validation backtest Sharpe in
[`14_backtest`](14_backtest.ipynb).
```python
"""Fit the declared NASDAQ-100 microstructure LSTM population on the walk-forward folds."""
import plotly.graph_objects as go
import polars as pl
from case_studies.research import (
declared_labels,
load_model_configs,
model_requests,
open_study,
resolved_model_plan,
run_model_population,
)
from utils.style import COLORS, show_plotly_with_alt
```
```python
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
CONFIG_NAMES: list[str] = []
POPULATION_NAME = ""
SUPERSEDES_POPULATION: str = ""
DEVICE: str = ""
```
```python
study = open_study(
"nasdaq100_microstructure",
execution_tier=EXECUTION_TIER,
workspace=WORKSPACE or None,
entry_point="09_dl_lstm",
)
```
## 1. Which labels, and which model
The labels are the ones whose training menu declares `deep_learning`, and fitting all of them
in one run is what makes this population comparable against the linear and gradient boosting
ones: the families differ, the targets do not. `fwd_ret_15m` is the return over the fifteen
minutes after the decision minute and the horizon the strategy chapters trade; `fwd_ret_5m`
and `fwd_ret_60m` are the same construction at shorter and longer horizons.
The classification label `fwd_dir_15m` is absent, and not by oversight. The sequence runner
refuses a non-regression label outright - `case_studies/utils/deep_learning.py`, "sequence
runner currently supports regression labels only" - so that label declares `linear` and `gbm`
and nothing else.
```python
declared_labels(study, "deep_learning")
```
`lstm_h64` is this notebook's slice of the declared family. The menu declares four
architectures and each has its own notebook, because each is a different claim about what
structure in the window matters. They resolve against the same menu, the same folds and the
same windows, so a difference between their results is a difference between architectures.
`lookback` is how many prior one-minute observations enter a window - 60 gives the model the
trailing hour - and it is the same across all four, so the sample they are measured on is the
same. `n_epochs` and `checkpoint_interval` are declared with the architecture rather than
passed in here, because together they decide how many prediction sets each configuration owes:
100 epochs saved every 5 is 20, and a run that quietly trained for fewer would publish a
different population under the same name.
```python
SEQUENCE_CONFIG = "lstm_h64"
declared = load_model_configs(study, "deep_learning", config_names=[SEQUENCE_CONFIG])
configs = load_model_configs(
study,
"deep_learning",
labels=LABELS or None,
config_names=CONFIG_NAMES or [SEQUENCE_CONFIG],
)
configs
```
`LABELS` and `CONFIG_NAMES` narrow the run below this notebook's own slice, and a narrowed run
declares a different member set than the published population does. A population is immutable
once written, so such a run must publish under its own name.
The device is checked in the same cell, because it is inside the training identity rather than
beside it: a network trained on a GPU and the same network trained on a CPU accumulate their
sums in different orders and reach different weights. The runner refuses to substitute a CPU
for a requested GPU rather than publishing a different model under the published name, so on a
machine with no NVIDIA card this notebook stops at the next cell; set `DEVICE="cpu"` and pass a
`POPULATION_NAME` to fit the same grid there.
```python
PUBLISHED_DEVICE = "gpu"
device = DEVICE or PUBLISHED_DEVICE
print(f"training device: {device}")
narrows = set(zip(configs["label"], configs["config_name"], strict=True)) != set(
zip(declared["label"], declared["config_name"], strict=True)
)
if (narrows or device != PUBLISHED_DEVICE) and not POPULATION_NAME:
raise ValueError(
f"this run declares {configs.height} label-configuration pairs on device {device!r}, "
f"which is not this notebook's declared slice on {PUBLISHED_DEVICE!r}, so it cannot "
f"publish the {SEQUENCE_CONFIG} population; pass POPULATION_NAME to give it its own"
)
```
## 2. Binding the declarations to the data
A menu entry says which network to fit. It does not say which feature columns exist today,
where the walk-forward folds fall, or which symbol-minute pairs have both a feature row and a
label - nor, for a sequence model, which of those have sixty prior observations behind them.
**Resolving** a request goes and finds all of that, and fits nothing, so the plan can be read
before any training starts.
Three things to check in it.
- **`eligible_rows` is below what the linear and gradient boosting families report on the same
label.** A prediction needs a full, gap-free window behind it, so what drops out is a name
too new to have accumulated one, or a stretch where the session boundary falls inside the
window. Comparing a sequence result against a tabular one is therefore comparing measurements
on different samples, which [`13_model_analysis`](13_model_analysis.ipynb) has to account for.
- **`folds` is the same everywhere** and equals the walk-forward splits `05_evaluation`
established.
- **`validation_start` and `validation_end` bracket the development sample.** The held-out tail
must not appear; it is scored once, at the end of the case study.
**How far apart the training windows sit is declared, not left to the tier.** Every row of this
panel could start a window, and the panel is minute bars, so an unspaced fold would build about
four million near-identical overlapping sequences - neighbouring windows would share 59 of their
60 observations. `modeling.dl.train_sequence_stride_horizons` in `config/setup.yaml` declares
the spacing instead: one window per label horizon, so consecutive windows of a symbol carry
labels that do not overlap. [`02_labels`](02_labels.ipynb) takes the decision on the bar
closing at t, enters on the next bar's VWAP and holds for the horizon from that fill, so two
decisions H observations apart have return intervals that abut rather than share bars - the
exit of one is the entry of the next. It is declared in horizons rather than in windows because the
horizon differs by label: `fwd_ret_5m` strides 5 observations, `fwd_ret_15m` 15 and
`fwd_ret_60m` 60, each spacing its own labels exactly, where a single window count could only
have been right for one of them. The spacing travels in the training identity; **how many
windows it yields is derived at run time** from the eligible rows of each fold rather than
declared, and is printed below.
```python
requests = model_requests(
study,
configs,
execution_tier=EXECUTION_TIER,
overrides={"device": device},
preview_reductions=PREVIEW_REDUCTIONS,
)
resolved = tuple(request.resolve() for request in requests)
plan = resolved_model_plan(resolved)
plan.select(
"label",
"config_name",
"feature_count",
"eligible_rows",
"folds",
"checkpoints",
"validation_start",
"validation_end",
)
```
## 3. Fitting the population
`run_model_population` fits every resolved request. For one request it walks the folds, and on
each one:
1. takes the rows inside that fold's training window and cuts them into overlapping windows of
sixty one-minute observations, each belonging to one stock and ending before the minute it
predicts, up to the declared cap,
2. standardizes each column on the training rows and applies that scale unchanged to the
validation rows, so nothing measured on the validation window reaches the fit,
3. trains for the declared number of epochs, writing the weights to disk at each checkpoint,
4. predicts the fold's validation rows from each saved set of weights.
**A window never crosses a stock, and it reads only what was observable at the minute it
predicts.** Hidden state is reset between stocks and between folds. What a window carries is
feature values already on the table at that minute, never a label from the interval the
prediction covers, so the purge the folds impose is not crossed.
Step 4 is what makes one training run produce twenty results. Each checkpoint's fold
predictions are concatenated into one series covering the whole validation period, and each
becomes its own registered prediction set with its own identity.
**What the call publishes is a population**: a named, immutable list of the prediction sets it
will produce, written down before the first fit. Afterwards every member must exist and be
complete, which is why a configuration that raises fails the whole call rather than publishing
a population one member short. Everything that finished stays registered, and re-running trains
only what is missing.
`SUPERSEDES_POPULATION` names the population hash this run replaces. Anything that moves a
training identity - a changed epoch schedule, lookback, sequence cap or device as much as a
changed menu - produces a different population under the same name, and the registry refuses to
write it without being told which snapshot it supersedes. It refuses before the first fit, so
the cost of forgetting it is seconds rather than the run.
```python
population_name = POPULATION_NAME or "nasdaq100_microstructure-lstm_h64-validation-v1"
execution, population = run_model_population(
study,
resolved,
population_name=population_name,
supersedes=SUPERSEDES_POPULATION or None,
)
reused = sum(1 for item in execution.diagnostics if item.get("reused"))
print(
f"{len(execution.runs)} configurations: {len(execution.runs) - reused} trained, {reused} read"
)
print(f"population {population.name}: {len(population.members)} prediction sets")
```
`reused` is not zero on a second run. Every identity is re-derived from the inputs, the
registry already holds the matching rows and the saved weights, and the runner returns the
stored result rather than training again.
## 4. What came out
The learning curve is the information coefficient on the validation rows at each checkpoint,
which is a rank correlation between the prediction and the realized return. It is read across
checkpoints of one fit rather than across configurations, so it says what more training did to
this model rather than what a different model would have done.
On a fifteen-minute horizon almost all of the target is noise, so the curve to expect is not a
rising one. A curve that climbs and then falls is the model beginning to fit the training
window; one that never rises is the architecture finding nothing this target rewards, which is
a result rather than a failure.
```python
# Scoped to this population's own members. `catalog_rows` is the study's whole prediction
# table, so an unfiltered height check would compare this run against every sequence row the
# registry holds and pass or fail for reasons that have nothing to do with it.
catalog = execution.catalog_rows.filter(
pl.col("prediction_hash").is_in(list(population.members))
).sort("label", "checkpoint_value")
if catalog.height != len(population.members) or catalog.filter(~pl.col("complete")).height:
raise RuntimeError(
f"the {SEQUENCE_CONFIG} population declares {len(population.members)} members and the "
f"registry holds {catalog.height} complete ones"
)
catalog.select("label", "config_name", "checkpoint_value", "ic_mean", "complete")
```
```python
fig = go.Figure()
for _label in catalog["label"].unique().sort().to_list():
_rows = catalog.filter(pl.col("label") == _label).sort("checkpoint_value")
fig.add_scatter(
x=_rows["checkpoint_value"].to_list(),
y=_rows["ic_mean"].to_list(),
mode="lines+markers",
name=_label,
)
fig.add_hline(y=0, line_width=1, line_dash="dash", line_color=COLORS["recede"])
fig.update_layout(
title=f"Validation IC by checkpoint - {SEQUENCE_CONFIG}",
xaxis_title="training epoch",
yaxis_title="information coefficient",
)
show_plotly_with_alt(
fig,
"Validation information coefficient against training epoch, one line per label, with a "
"dashed line at zero.",
)
```
## Key takeaways
1. **A checkpoint is part of a configuration, not a detail of how it was fitted.** Scoring one
fit at twenty points produces twenty candidates, and keeping whichever of them scores
highest on this page would be a selection decision taken on the wrong statistic: the
curve below is an information coefficient, and selection is validation backtest Sharpe in
[`14_backtest`](14_backtest.ipynb), over the population published here.
2. **How far apart the training windows sit is part of the model.** On a minute panel the
spacing decides what was fitted, so it is declared in `config/setup.yaml` as one window per
label horizon and travels in the training identity rather than arriving with the invocation.
3. **A sequence family is measured on fewer rows than a tabular one.** A prediction needs a
full window behind it, so the samples differ and the comparison has to say so.
**Known limitations.** The window is fixed at sixty observations, so nothing earlier than the
trailing hour reaches the model whatever the architecture can represent. The window spacing is a
modelling choice rather than a compute budget - one window per label horizon is what stops two
training examples carrying the same return twice - so the number of windows a fold yields
follows from how many eligible rows it holds, and a panel of a different size changes the count
without changing the declaration.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。