Designing an Intraday Equity Microstructure Study
Summary
This configuration describes an intraday study of order-flow signals across a point-in-time NASDAQ-100 universe. It defines a cost-screened subset using liquidity information available before each evaluation period, specifies decision schedules for several forecast horizons, and distinguishes those schedules from the one-minute price feed used to monitor positions. Feature families cover quote liquidity, microprice, volatility, order flow, price impact, and other microstructure measures, with session-bounded lookbacks and declared timing conventions.
The setup also records choices intended to make the experiment reproducible and reduce leakage: float32 feature storage, horizon-based spacing for training labels, a one-bar treatment window, and walk-forward double machine learning with stated confounders. The cost-feasible universe is selected from pre-evaluation quote data, avoiding the use of future liquidity to choose names. The file is a specification rather than a report of results. It describes hypotheses and implementation assumptions, but does not by itself establish that order-flow imbalance predicts returns or that the screened strategy is profitable. Its conclusions would depend on the associated feature construction, model fitting, backtests, and execution assumptions.
Key ideas
- Use point-in-time index membership and pre-evaluation liquidity data to define eligible stocks.
- Keep decision cadence distinct from the finer price feed used to monitor positions.
- Build session-bounded microstructure features with explicit lookbacks, lags, and failure modes.
- Space training examples according to each label horizon to avoid repeatedly sampling overlapping outcomes.
- Treat the declared order-flow measure as a causal treatment and evaluate it with walk-forward double machine learning.
Tags
Full text
# setup.yaml
```yaml
strategy_id: nasdaq100_microstructure
setup_version: v1
universe:
# Point-in-time NASDAQ-100 membership over the sample, restricted to the names
# that contribute at least one session before evaluation.holdout_start. The
# published AlgoSeek archive carries 123 symbols; eight of them first quote
# after 2021-07-01, so nothing is ever fit on them, the pre-holdout liquidity
# profile has no spread for them, and 05_evaluation drops them before it
# screens anything. 01_feasibility_analysis asserts this list against the
# archive in both directions. Names that leave the index end where they left:
# AAL and WLTW stop at 2020-05-01, which is the membership, not a gap.
symbols:
- AAL
- AAPL
- ADBE
- ADI
- ADP
- ADSK
- AEP
- ALGN
- ALXN
- AMAT
- AMD
- AMGN
- AMZN
- ANSS
- ASML
- ATVI
- AVGO
- BIDU
- BIIB
- BKNG
- BMRN
- CDNS
- CDW
- CERN
- CHKP
- CHTR
- CMCSA
- COST
- CPRT
- CSCO
- CSGP
- CSX
- CTAS
- CTSH
- CTXS
- DLTR
- DOCU
- DXCM
- EA
- EBAY
- EXC
- EXPE
- FAST
- FB
- FISV
- FOX
- FOXA
- GILD
- GOOG
- GOOGL
- IDXX
- ILMN
- INCY
- INTC
- INTU
- ISRG
- JD
- KDP
- KHC
- KLAC
- LBTYA
- LBTYK
- LRCX
- LULU
- MAR
- MCHP
- MDLZ
- MELI
- MNST
- MRNA
- MRVL
- MSFT
- MTCH
- MU
- MXIM
- NFLX
- NTAP
- NTES
- NVDA
- NXPI
- OKTA
- ORLY
- PAYX
- PCAR
- PDD
- PEP
- PTON
- PYPL
- QCOM
- REGN
- ROST
- SBUX
- SGEN
- SIRI
- SNPS
- SPLK
- SWKS
- TCOM
- TEAM
- TMUS
- TSLA
- TTWO
- TXN
- UAL
- ULTA
- VRSK
- VRSN
- VRTX
- WBA
- WDAY
- WDC
- WLTW
- XEL
- XLNX
- ZM
n_assets: 115
eligibility_rule: nasdaq100_membership_with_pre_holdout_session
# Cost-feasible universe: the cheapest-to-trade names by round-trip cost,
# frozen per split. Provenance: _build_cost_feasible_universe.py
# (round-trip proxy 2*(per_share/mean_price)*1e4 + 2*median_half_spread_bps).
# The validation list is profiled on quote bars strictly before the
# validation window (2020-01-01 -> 2020-06-30); the holdout list from the
# pre-holdout liquidity profile (< 2021-07-01). No look-ahead. Selected via
# strategy.signal.universe_filter='cost_feasible'; the split-specific list is
# chosen at backtest time by resolving the prediction set's split (NOT a
# spec-hash input — like the 'liquid' filter, the hash carries only the
# filter name). The full 115 universe collapses out of sample (featured
# config holdout -0.42) while the cost-feasible universe is positive; the
# screen is load-bearing. See 17_costs.py for the full-vs-screened contrast.
cost_feasible:
validation:
- AAPL
- MSFT
- FB
- SBUX
- PEP
- ATVI
- INTC
- PYPL
- MDLZ
- AMD
- QCOM
- TXN
- GILD
- MU
- NVDA
- AMAT
- CSCO
- TMUS
- WBA
- COST
- JD
- CSX
- XEL
- CMCSA
- EXC
- EBAY
- MNST
- FISV
- FAST
- EA
- CTSH
- NFLX
- XLNX
- AMZN
- ADBE
- PAYX
- CDNS
- AMGN
- MXIM
- BIDU
- PCAR
- CERN
- UAL
- ADP
- KHC
- FOXA
- WDC
- TSLA
- FOX
- CTXS
holdout:
- AAPL
- MSFT
- FB
- SBUX
- PEP
- GILD
- AMD
- QCOM
- INTC
- MDLZ
- ATVI
- AMAT
- MU
- PYPL
- AEP
- TXN
- JD
- CSX
- MRVL
- COST
- TMUS
- CMCSA
- EBAY
- CSCO
- NVDA
- XEL
- FISV
- WBA
- EXC
- FAST
- TSLA
- CTSH
- MXIM
- ADI
- AMZN
- NFLX
- AMGN
- EA
- PAYX
- CERN
- ADBE
- ADP
- KHC
- CDNS
- XLNX
- MNST
- KDP
- BIDU
- DLTR
- UAL
decision:
# How often the strategy is allowed to act, NOT the size of a bar in the panel.
# The observations are one minute apart; 03_financial_features keeps every
# fifteenth of them to build the decision schedule. Nothing here declares the
# observation grid, so anything that needs it - an embargo, a permutation block,
# any count of periods spanning a duration - measures it from the data instead
# of dividing by this number. Reading this field as the bar size yields an
# answer fifteen times too small and nothing in the result shows it.
bar_frequency: 15_minute
# A model decides on its own horizon: a five-minute model every five minutes, a
# sixty-minute model every hour. `bar_frequency` above is the default for a label that
# declares nothing here, and stays the primary label's cadence.
#
# This is the decision schedule only. The engine's price feed is one minute
# (`config/backtest/base.yaml::calendar.data_frequency`), so a position entered on one
# decision is watched every minute until the next - which is what a stop or a
# take-profit is measured on, and what a fifteen-minute feed could not express.
cadence_by_label:
fwd_ret_5m: 5_minute
fwd_ret_15m: 15_minute
fwd_dir_15m: 15_minute
fwd_ret_60m: 60_minute
decision_snapshot: bar_close
execution_delay: 1_bar
execution_price_assumption: next_bar_open_or_vwap
# The feature specification read by 03_financial_features. Windows are counted
# in minute bars and every one of them is session-bounded: a window never spans
# an overnight gap, so each session restarts its own warmup.
features:
# Feature matrices are stored and fitted in single precision. 16,098,877 rows x 88 features
# in float64 is a 10.9 GB modeling dataset and a 58.9 GB peak through one gradient-boosting
# fold; in float32 those are 5.6 GB and 38.0 GB. LightGBM bins to uint8 regardless, and the
# linear families fit standardised columns where the extra mantissa carries nothing this
# case study can measure. Declared here rather than globally: the smaller case studies fit
# comfortably in double precision, and narrowing them would move their numbers for nothing.
storage_dtype: float32
windows:
fast: 5 # one fast aggregate
decision: 15 # one decision bar at the declared bar_frequency
slow: 30 # realized-volatility horizon and the Amihud average
hour: 60 # Kyle's lambda regression and the slow regime average
ewma_half_life: 30 # half-life of the EWMA volatility, in bars
stale_cap: 5 # consecutive stale-quote bars tolerated before nulling
edge_block: 30 # width of the open and close blocks the flags mark
# The quantity the thesis puts forward: the order-flow imbalance measured over
# one decision bar. Its one-bar form is this case study's causal treatment (see
# `causal.treatment` below). Read by 03_financial_features, which draws it in
# F2, F3 and F6.
carrier: signed_vol_share_15m
# One row per family. `pattern` claims a family's columns in both
# representations - the level and its cross-sectional z-score twin - because
# the two carry the same hypothesis on two scales. `lookback` and `lag` are
# minute bars, and they are what the warmup audit and the timing figure read.
families:
- name: quote_liquidity
pattern: rel_spread_*|depth_imb*|quote_rate*
role: state
hypothesis: >-
The cost and the depth of the book describe the environment a signal is read
in: an imbalance is worth acting on only where the round trip is cheap enough
to leave something behind.
inputs: NBBO close quotes, quote update count
lookback: 60
lag: 0
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
A carried-forward quote reports a stale spread as a live one, so the state
looks calm exactly where the book has stopped updating.
- name: microprice
pattern: microprice_dev*
role: signal
hypothesis: >-
The depth-weighted price leans toward the thin side of the book, so its distance
from the midpoint is where the next trade is more likely to print (Stoikov, 2018).
inputs: NBBO close prices and sizes
lookback: 15
lag: 0
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
The deviation is a few tenths of a cent on a wide book and a rounding artifact
on a one-tick one, so it is not comparable across names until it is ranked.
- name: volatility
pattern: r1m*|rv_*|quote_range*
role: state
hypothesis: >-
How much the quoted price is moving is what decides whether an imbalance is
information or noise, and it varies by name and by hour.
inputs: NBBO quote midpoint and quote extremes
lookback: 60
lag: 0
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
Every estimator here is a variance over a short window, so a single bad quote
dominates it and reports turbulence where there was a data error.
- name: order_flow
pattern: signed_vol_share*|tick_imb_share*|trade_to_mid*|trades_per_1k_shares*|cross_locked_share*
role: signal
hypothesis: >-
Aggressive volume arrives in runs, so the side that has been paying the spread
over the last few minutes keeps paying it over the next few.
inputs: AlgoSeek trade-location and tick buckets, total traded volume
lookback: 60
lag: 1
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
Trade location is assigned against the prevailing quote, so a bar whose quote
was stale attributes its volume to the wrong side rather than to neither.
- name: price_impact
pattern: dollar_vol*|illiq*|kyle_lambda*|trade_range*
role: state
hypothesis: >-
How far a given quantity of order flow moves the price is what separates a
crowded name from a thin one, and it is the cost side of every signal here.
inputs: consolidated trade prices, traded dollars across both venues
lookback: 60
lag: 1
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
Amihud is undefined on a bar that printed nothing, and its rolling average
carries that gap across the whole window rather than skipping the bar.
- name: hidden_liquidity
pattern: finra_share_60m*
role: state
hypothesis: >-
The share of volume printing away from the exchanges tracks how much of the
session's liquidity is not visible in the book, and it moves over hours.
inputs: FINRA/TRF reported volume, total traded volume
lookback: 60
lag: 1
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
TRF prints are reported with a delay, so the share attributed to a bar is not
fully public at that bar's close unless it is lagged.
- name: session_clock
pattern: time_since_open|time_to_close|is_first_30m|is_last_30m
role: state
hypothesis: >-
Intraday liquidity is U-shaped, so where in the session a bar sits conditions
every other family without predicting anything on its own.
inputs: exchange calendar
lookback: 0
lag: 0
frame: per session, from the scheduled open and close
representation: fraction of the session and block flags
failure_mode: >-
Measured against the session's realized bar count rather than its scheduled
close, it is a quantity no one had at the time.
# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study.
#
# Spec-hash inputs (each invalidates every backtest_hash on change):
# - execution.initial_cash
# - execution.share_type
# - execution.allocator_lookback (CS-level fallback for moment allocators)
# - per-allocator overrides in backtest.sweep.allocators
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator. It is counted in ROWS OF THE PRICE
# FRAME, and that frame is one-minute: ``config/backtest/base.yaml`` declares
# ``calendar.data_frequency: 1m``, ``load_backtest_prices_for`` returns bars
# measured at 60-second spacing, and ``_compute_rolling_vol`` applies
# ``rolling_std(vol_window)`` to those rows with no resampling. So 520 is 520
# MINUTES - about 1.3 regular sessions - and not the 20 trading days this comment
# claimed until 2026-09-06. The value is unchanged; what was wrong was the
# sentence describing it, which read the decision cadence as the bar size.
# ``periods_per_year`` annualizes Sharpe at daily-equivalent grain (252),
# independently of this window.
#
# Expensive allocators (RP/HRP/MVO_LW) are dropped via expensive_allocators_skip
# below — covariance estimation at 15-min cadence is prohibitive and
# empirically degenerate. MVO_LW carries no per-allocator lookback override
# here for that reason; revisit if expensive allocators are reintroduced
# (a 6-month equivalent at 15-min would need ~3,300 bars).
#
# ``initial_cash`` restored to 1_000_000 (2026-05-16) — the 2026-05-15 SSOT
# migration drop to 100k caused catastrophic degenerate output here
# (avg Sharpe -11 across 339 signal runs; integer-share rounding × the
# nasdaq-100 high-priced tail (BKNG≈$5k, MSTR/NFLX/AVGO $700-1.5k)). See
# memory/feedback_2026_05_15_equity_sizing_invalidated.md.
execution:
initial_cash: 1_000_000 # restores prior validated state
share_type: integer # US equities trade in whole shares
allocator_lookback: 520 # rows of the 1-minute price frame: ~1.3 sessions
mapping:
class: intraday_rank_and_trade
position_state_space: long_short
entry_logic: rank_or_threshold
sizing: dollar_neutral_or_beta_neutral
costs:
class: dominant
# Per-share commission plus measured per-asset half-spread slippage from
# AlgoSeek NBBO close quotes over the training window. Integer share
# sizing. The asset_spreads_source parquet is loaded by the backtest
# cost-preset builder and joined per symbol; default_half_spread_usd is
# the universe p75 fallback for symbols not in the profile.
model: per_share_plus_spread
per_share: 0.0035 # IBKR Pro Tiered top tier (<=300K shares/mo)
minimum: 0.0
asset_spreads_source: liquidity_profile.parquet
asset_spreads_column: median_half_spread_usd
default_half_spread_usd: 0.025
spread_convention: half_spread
source: AlgoSeek minute-bar NBBO close quotes, training window 2020-01-02 to 2021-07-01.
friction_floor_bps: 5
backtest:
rebalance:
# A rebalance is skipped when the per-asset weight change is below
# min_weight_change AND the resulting trade notional is below
# min_trade_value. The benchmark profile disables thresholds so that
# full-universe equal-weight (1/N per asset) rebalances at all.
default:
min_weight_change: 0.005
min_trade_value: 100.0
benchmark:
min_weight_change: 0.0
min_trade_value: 0.0
sweep:
# Canonical strategy universe. The 15-min round-trip cost on the
# cost-expensive tail of the 115-name panel consumes the intraday edge:
# the featured config collapses out of sample on the full universe
# (holdout -0.42) but is positive on the cost-feasible subset (the
# cheapest-to-trade names, frozen per split — see universe.cost_feasible).
# The full-universe variant is NOT a canonical rank-1 / cohort / DSR
# candidate; it lives only in the 17_costs.py full-vs-screened comparison,
# mirroring sp500_options' liquid screen. Read by
# case_studies.utils.sweep_config.get_universe_filters_for.
universe_filter: cost_feasible
# Iteration controls per stage. ``signal: 0`` means "all predictions";
# downstream stages take the top-N from the upstream stage's rank-1.
# Notebooks read these via get_top_n_predictions(case_study, stage).
top_n_predictions:
signal: 0 # all signal predictions (eq-weight baseline)
allocation: 10 # top-10 model configs by equal-weight baseline Sharpe
# top-1 of every pre-cost stage per label. `17_costs` pools {signal, allocation,
# risk_overlay} - `STAGE_SEQUENCE` minus the stage it writes - because an overlay is a
# candidate to carry the case study and pricing only its parent would put a cost curve
# in the chapter for a strategy this case study does not select.
cost_sensitivity: 1
risk_overlay: 1 # top-1 of {signal+allocation} per label
# No expensive allocators on nasdaq. At 15-min frequency the 1.3M-bar
# panel makes covariance estimation on the rebalance schedule
# prohibitive; empirically every expensive allocator (risk_parity,
# mvo_ledoit_wolf, hrp) produces deeply negative Sharpes (-7 to -10)
# on every nasdaq label and equal_weight ties at the top. The 15-min
# cadence is a costs-dominated regime where signal-stage selection
# carries the entire story.
expensive_allocators_skip: true
# Ch16 signal-stage selection. Long-short equal-weight top-k on the
# 15-min schedule across four labels (5m/15m/60m continuous + 15m
# direction). The classification label fwd_dir_15m is a first-class
# signal alongside the continuous labels and carries the same top_k
# grid.
top_k_grid:
fwd_ret_5m: [5, 10, 20]
fwd_ret_15m: [5, 10, 20]
fwd_ret_60m: [5, 10, 20]
fwd_dir_15m: [5, 10, 20]
# Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
# Nasdaq runs only the three fast allocators — the moment-based ones are
# too slow at intraday frequency (``expensive_allocators_skip: true``).
# No max_weight cap on inverse_vol — see
# memory/feedback_max_weight_caps_intentionally_absent.md (would force
# equal-weight at top_k=5 and defeat the moment-allocator comparison).
allocators:
- {name: equal_weight, method: equal_weight}
- {name: score_weighted, method: score_weighted}
- {name: inverse_vol, method: inverse_vol}
# Ch18 cost sensitivity (bps regime).
cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
# Companion per-share cost regime: dollars per share half-spread. Swept
# alongside cost_grid_bps for CSes whose declared cost model is
# per_share_plus_spread. Values: 0¢, 0.5¢, 1¢, 2.5¢, 5¢, 10¢.
cost_grid_half_spread_usd: [0.0, 0.005, 0.01, 0.025, 0.05, 0.10]
# Alternative rebalance cadences explored by Ch18's cadence × cost
# heatmap (17_costs.py Section 4). The first entry must be the CS's
# default cadence (``decision.bar_frequency``). Each cadence reuses the
# ``cost_grid_half_spread_usd`` grid above. Tokens are the same cadence
# vocabulary recognized by the engine; the notebook pulls the list from
# here, never hardcoded.
cadence_sweep: [15_minute, 30_minute, 1_hour, 4_hour]
# nasdaq100 v4 slot-mechanism sweep (Ch16 signal stage replacement).
#
# Adds two selection methods to the canonical Ch16 expansion:
# - slot_persistent_signal_exit: per-symbol rolling-percentile entry,
# fixed weight-per-slot allocation (slots ARE the allocation —
# Ch17 allocator stage is skipped for these rows), max-hold +
# optional signal-exit (rolling lower quantile crossover) for
# position lifecycle. See case_studies/utils/slot_strategy.py for
# the mechanism. Sub-params live under selection_method_config in
# the flattened scheme dict.
# - eq_w_topk: the canonical equal-weight top-K baseline, kept as
# a within-sweep comparison. Direction grid covers long_only,
# long_short.
#
# slot × long_short is intentionally dropped — slot books are
# single-direction by construction (one symbol cannot occupy a long
# AND short slot simultaneously). The cross-asset long-short story
# belongs to eq_w_topk / quintile_long_short.
#
# Per-prediction cardinality:
# slot: long_q(3) × max_slots(3) × hold_bars(2) × exit_signal_q(2)
# × bars_per_day_grid(1) = 36 (long_only only)
# eq_w_topk: top_k(3) × direction(2) = 6
# 42 per prediction, and `signal_passes` below says how many predictions
# each of them runs on.
#
# `hold_bars` and `exit_signal_q` were cut here, from 3 and 4 values to 2
# each, on 2026-09-11. Both keep their endpoints: hold_bars keeps 2h and 8h
# and drops the 4h midpoint, exit_signal_q keeps "off" and the median
# crossover and drops 0.30 and 0.70. A midpoint is what a grid can lose
# without losing the comparison it draws - the question the sweep asks is
# whether a longer hold and a signal exit move the result, and two values
# answer it. `long_q` and `max_slots` are untouched: `14_backtest.py`'s
# Act 2 selects the carrier by `slots == 10` and `entry_q == 0.9`
# literally, so removing either value empties the table that section
# prints.
signal_nasdaq100:
selection_method: [slot_persistent_signal_exit, eq_w_topk]
long_q: [0.90, 0.95, 0.99]
direction: [long_only, long_short]
max_slots: [5, 10, 20]
# hold_bars and bars_per_day count rows of the grid the slot simulator walks, which is
# the PRICE grid, now one minute. Restated from the fifteen-minute counts they carried
# (8/16/32 and 14) at the same wall-clock durations, so the sweep asks the same question
# of the same intervals rather than silently asking about 8, 16 and 32 minutes.
hold_bars: [120, 480] # 2h, 8h, in minute bars
exit_signal_q: [null, 0.50] # null disables signal-exit
pred_freshness_max_min: 14 # tolerated staleness of a prediction, in minutes
bars_per_day_grid: [210] # was 14 bars of 15 minutes; same 3.5h window
top_k_grid: [5, 10, 20]
lookback_days: 21
# How the signal stage walks the grid above. Two passes, because the whole
# grid over every prediction is not affordable here and nothing else in the
# book has this shape: every other case study declares 2 to 4 entry schemes
# and one universe, and no completed case study holds more than about 3,000
# signal backtests.
#
# Counted against this registry and measured on 2026-09-11. A backtest of
# this panel costs 31.8 s (four arms timed on the live 10.1M-row price
# frame, both universes, 29.2 s to 36.3 s), and the registry offers 741
# admissible prediction sets: 226 for each of the three continuous labels
# and 63 for `fwd_dir_15m`.
#
# pass 1 baseline 741 × 3 canonical arms × cost_feasible = 2,223
# pass 2 mechanism 32 × 42 arms × cost_feasible = 1,344
# pass 2 reference 32 × 3 canonical arms × full = 96
# total = 3,663 ≈ 32 h
#
# Crossing all 117 arms with both universes over all 741 predictions, which
# is what this block replaces, is 173,394 backtests and about 1,500 hours.
# The prediction count is not what was wrong with it: 741 sets is what three
# DL families at 20 epoch checkpoints, 15 gbm configs at 10 tree-count
# checkpoints and 16 linear configs come to, and every one of them is a
# configuration this case study fitted on purpose. The arm count was.
signal_passes:
# Pass 1 ranks every admissible prediction on three equal-weight
# concentrations, on the universe the carrier actually trades. Ranking on
# the full universe instead would order the predictions by a Sharpe no
# selected strategy ever earns, because the cost-expensive tail of the
# panel is exactly what `universe.cost_feasible` removes.
baseline_schemes: [ew_top5, ew_top10, ew_top20]
baseline_universe: cost_feasible
# Pass 2 runs the mechanism grid on the strongest predictions of pass 1,
# per label, ranked by pass-1 validation Sharpe.
#
# By Sharpe and not by IC. `top_n_predictions.signal` is the other lever
# that would cut this grid and it is the wrong one: `load_prediction_index`
# ends `ORDER BY m.ic_mean DESC NULLS LAST`, so any positive value there
# screens the signal stage on rank correlation and decides which models are
# backtested at all before a backtest has run. All nine case studies
# declare `signal: 0` for that reason, and it stays 0 here.
mechanism_top_n: 8
# The full universe is kept as a reference arm rather than a second copy of
# the sweep: the canonical concentrations only, on the pass-2 predictions.
# That is what `17_costs` reads for its full-vs-screened comparison - it
# prices the top-1 of each pre-cost stage, which is always inside pass 2 -
# and what Act 1 of `14_backtest` reads to show the unscreened baseline.
reference_schemes: [ew_top5, ew_top10, ew_top20]
reference_universe: full
# Ch19 risk overlays.
risk_controls:
position:
- {name: stop_loss_3pct, type: stop_loss, threshold: 0.03}
- {name: stop_loss_5pct, type: stop_loss, threshold: 0.05}
- {name: stop_loss_10pct, type: stop_loss, threshold: 0.10}
- {name: stop_loss_15pct, type: stop_loss, threshold: 0.15}
- {name: trailing_1pct, type: trailing_stop, threshold: 0.01}
- {name: trailing_2pct, type: trailing_stop, threshold: 0.02}
- {name: trailing_3pct, type: trailing_stop, threshold: 0.03}
- {name: trailing_5pct, type: trailing_stop, threshold: 0.05}
- {name: trailing_10pct, type: trailing_stop, threshold: 0.10}
- {name: trailing_15pct, type: trailing_stop, threshold: 0.15}
- {name: trailing_20pct, type: trailing_stop, threshold: 0.20}
- {name: time_exit_10, type: time_exit, bars: 10}
- {name: time_exit_20, type: time_exit, bars: 20}
- {name: time_exit_40, type: time_exit, bars: 40}
# The mean-forecast ensemble this case study features, and the only thing in the repository
# that produces `family: ensemble` rows. Declared rather than assembled in the notebook
# because the member set is what the ensemble *is*: change the rule and it is a different
# object under the same name.
#
# The lesson this case study ends on is that choosing among a family's configurations on a
# validation window does not carry to a later window, while the family itself does. An object
# that makes no choice is what shows it, and that object has to be registered like a model so
# the backtest stage, the cohort comparison and the holdout all read it the way they read one.
ensemble:
# The family averaged, and the rule that decides which of its configurations are in.
#
# 31 leaves is the boundary between the regularized presets and the one that is not:
# `leaves_7`, `leaves_15` and `leaves_31` carry `lambda_l1`, `lambda_l2`, `bagging_fraction`
# and `feature_fraction`; `default` carries none of them and trains at LightGBM's own
# default of 31 leaves; `leaves_63` is the only configuration above the line. The rule keeps
# twelve of the fifteen for each continuous label - three losses across four leaf counts.
#
# A count is not declared here on purpose. The rule is the declaration and the members are
# resolved from what is registered, so a configuration added to `config/training/*.yaml`
# joins the ensemble and a number written down here could only disagree with it.
member_family: gbm
max_num_leaves: 31
method: mean_forecast
config_name: gbm_mean_leaves31
# Each member enters at its last checkpoint. Taking the checkpoint with the best validation
# metric would select on the window the ensemble is then measured on, which is the thing
# the ensemble exists not to do.
checkpoint: last
# The entry scheme the ensemble is backtested on, by the name `get_entry_schemes_for`
# gives it: 10 slots, an 8-hour maximum hold, a 0.90 entry quantile, no signal exit, on the
# 210-bar execution window. `14_backtest.py` Act 2 reads the carrier by `slots == 10` and
# `entry_q == 0.9` literally, so this arm and that filter have to agree.
#
# One arm and not the grid: the ensemble is not a sweep participant. It is the object the
# chapter carries, backtested on the design the sweep already chose, and putting it through
# the sweep would make it one more configuration to select among.
featured_scheme: slot_l_lq90_s10_h480_noexit_b210
universe_filter: cost_feasible
# The model specification 04_model_based_features estimates under. Declared here rather than
# typed into the notebook so that a window is one number in one place, readable without opening
# the code that consumes it.
#
# This case study fits no GARCH, no regime chain and no ARIMA: its three procedures are a
# heterogeneous autoregression of realized variance, a rolling Fourier transform and a rolling
# depth-2 path signature. Only the first estimates anything, which is why only `har` carries a
# refit schedule.
model_based:
har:
# The three lengths of history the regression averages squared one-minute returns over,
# in bars. Corsi's day/week/month on a minute grid: minutes, a quarter hour, an hour.
components: [5, 15, 60]
# Bars of history each refit reads. Two hours - long enough for four coefficients to be
# identified, short enough that the fit follows the day rather than the quarter. It is
# also the burn-in: the first `fit_window + 1` bars of a symbol carry no HAR value,
# because they are what pays for the first estimate.
fit_window: 120
# Bars between refits. One, which is the finest cadence the panel has and the reason this
# artifact carries no fold column: the coefficients describing bar t come from a
# regression whose last observation is bar t-1, so no fold boundary enters the fit and a
# bar's value is the same value whichever fold selects it.
refit_every: 1
# Usable observations a refit window needs before it is fitted at all. Below this the four
# coefficients are estimated off too few points to identify them, and the bar keeps a null
# rather than a fit on whatever survived the warm-up.
min_train_obs: 20
spectrum:
# Bars each trailing window spans. Sixty resolves repetition periods up to an hour, which
# is the longest intraday rhythm this panel is asked about.
window: 60
# The period, in bars, above which power counts as low frequency. Twenty separates a slow
# drift in activity from minute-to-minute churn, and it is what `*_low_freq_ratio` is the
# share of.
low_frequency_period: 20
signature:
# Bars each signature path spans. Thirty is the horizon over which order flow and price are
# expected to lead one another, which is the quantity the cross terms measure.
window: 30
evaluation:
n_splits: 2
train_size: 6M
val_size: 6M
holdout_start: '2021-07-01'
holdout_end: '2021-12-31'
calendar: NYSE
periods_per_year: 252 # NYSE 5d/wk (daily MTM despite intraday signal cadence)
labels:
primary: fwd_ret_15m
# The outcome horizon each label's name states. Declared rather than left to the default,
# because `resolve_label_horizon` falls back to the BUFFER when no horizon is given - so
# widening a buffer below would silently lengthen the return the label measures, and
# `fwd_ret_15m` would be a sixteen-minute return that every consumer still called fifteen.
horizons:
fwd_ret_15m: 15min
fwd_ret_5m: 5min
fwd_ret_60m: 60min
fwd_dir_15m: 15min
# The purge, which is the horizon PLUS ONE BAR - wider than the horizon, and deliberately.
# The entry leg is the VWAP of the bar after the decision and the exit leg is H past the
# entry, so a label at t consumes a quote at t+H+1: one bar past the horizon its name
# states. A buffer of H would leave the last bar of every training window inside the first
# validation label's window. The extra minute is not conservatism; it is the width of the
# window that the horizon alone does not describe.
buffer: 16min
variants:
- fwd_ret_5m
- fwd_ret_60m
- fwd_dir_15m
variant_buffers:
fwd_ret_5m: 6min
fwd_ret_60m: 61min
fwd_dir_15m: 16min
# How many slots of the DECISION schedule to advance per trade, so a position is held for
# its label's horizon and no two holdings overlap. A slot is one interval of the declared
# cadence, so these are `ceil(horizon / cadence)` and nothing else.
#
# That reading is only true because `resolve_rebalance_timestamps` now builds the schedule
# from the declared cadence. It used to return the panel's own timestamps unchanged for
# every intraday cadence, and this panel carries every minute, so these values thinned
# minutes rather than fifteen-minute slots and three of the four labels traded every
# minute (ml4t/agent-workspace#187). The values did not change; what they count did.
# Every label now trades on a cadence equal to its own horizon (decision.cadence_by_label
# above), so one slot is one holding period and no label needs a multi-slot step. The
# `fwd_ret_60m: 4` this replaces was `ceil(60 / 15)` against a single fifteen-minute
# default; the cadence carries that now, and carries it for the Ch18 sweep arms too,
# which differ from the default in exactly this token.
rebalance_step:
fwd_ret_15m: 1
fwd_ret_5m: 1
fwd_ret_60m: 1
fwd_dir_15m: 1
# Continuous return that each classification label is derived from.
classification_eval_label:
fwd_dir_15m: fwd_ret_15m
modeling:
dl:
# The backend the sequence families are trained on in production. Declared rather
# than defaulted: the four deep-learning notebooks refuse to train on a device other
# than the one asked for, because the device is part of what the run is registered
# under, and a notebook that invents `gpu` from a missing section makes a
# CUDA-free environment fail on a requirement nobody wrote down. A run on other
# hardware sets the notebook's DEVICE parameter, which is then what gets registered.
device: gpu
# How far apart the training windows of one sequence configuration sit, counted in
# label horizons. This is the only case study that needs the key: every row of a panel
# starts a window, and this panel is minute bars, so an unspaced fold builds about 4.0
# million near-identical overlapping sequences where the daily case studies build tens
# of thousands. Measured on the canonical plan, 2026-09-02: fold 0 has 4,066,041
# training rows and fold 1 has 3,997,571, across 115 symbols.
#
# One window per label horizon. Consecutive windows of a symbol then carry labels that
# do not overlap: the window ending at t is scored on the return from t to t+H, and the
# next one starts at t+H. Every window is still an example the model has to fit; what
# changes is that no two of them are the same example seen twice.
#
# Declared in horizons rather than in windows because the horizon differs by label, and
# one line then covers all four: `fwd_ret_15m` and `fwd_dir_15m` stride 15 observations,
# `fwd_ret_5m` strides 5 and `fwd_ret_60m` strides 60, each spacing its own labels
# exactly. A single count could only be right for one of them. On fold 0 the primary
# label draws about 271,000 windows, and the count is derived at run time and printed,
# not written down here - it differs between folds because the folds differ in length.
#
# This replaces `max_train_sequences: 750000`, which had been carried unchanged since
# the original release and through the stage-06 rebuild that moved the folds underneath
# it, so nothing derived it. At 750,000 a window was drawn every 5.4 minutes, finer than
# the 15-minute horizon, so neighbouring training windows carried overlapping labels and
# the extra 2.8x of them were re-presentations of examples already in the batch. That is
# not wrong - correlated examples are not invalid ones - but it was a number with no
# argument behind it, and the number is part of the training identity
# (ml4t/agent-workspace#1015).
#
# `max_train_sequences` is still available and is what a case study declares when it
# wants the overlap: it fixes the count and lets the spacing follow. The two are
# mutually exclusive.
train_sequence_stride_horizons: 1
gbm:
libraries: [lightgbm]
preset: default
device: cpu
# LightGBM's own CPU default. 63 is the GPU default and was carried over with the
# device when these runs moved off the GPU, so every CPU fit was quantizing the design
# matrix into a quarter of the bins the library would have used. Coarser bins are
# faster and lose split points; the reader running this on a CPU gets what the
# documentation describes.
max_bin: 255
causal:
treatment: signed_vol_share
# Bars the treatment's own construction window spans, which is what the placebo block has
# to cover: permuting signed_vol_share in blocks shorter than this destroys the serial
# dependence the refutation exists to preserve, and the resulting p-value reads like a
# refutation without being one. Declared here rather than inferred, because guessing which
# element of a window list a column was built from puts a wrong number behind a right-looking
# one. Derived from the construction, not chosen:
#
# `signed / TRADED_VOLUME` in 03_financial_features, both from the same bar. The smoothed
# versions this feeds are separate columns; the treatment itself spans one bar.
treatment_window: 1
confounders: [rel_spread_close, rv_5m, r1m]
method: walk_forward_dml
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.