Diseño de un estudio intradía de microestructura de renta variable
Resumen
Esta configuración describe un estudio intradía de señales de flujo de órdenes en un universo NASDAQ-100 definido en cada momento. Define un subconjunto filtrado por costes usando información de liquidez disponible antes de cada periodo de evaluación, especifica calendarios de decisión para varios horizontes de predicción y los distingue del flujo de precios de un minuto usado para vigilar las posiciones. Las familias de características abarcan liquidez de cotizaciones, microprecio, volatilidad, flujo de órdenes, impacto de precios y otras medidas de microestructura, con retrospectivas limitadas a la sesión y convenciones temporales declaradas.
La configuración también registra decisiones destinadas a hacer reproducible el experimento y reducir filtraciones: almacenamiento de características en float32, espaciado de etiquetas de entrenamiento según el horizonte, una ventana de tratamiento de una barra y aprendizaje automático doble walk-forward con factores de confusión declarados. El universo viable en términos de costes se selecciona con datos de cotizaciones anteriores a la evaluación, evitando usar la liquidez futura para elegir valores. El archivo es una especificación, no un informe de resultados. Describe hipótesis y supuestos de implementación, pero por sí solo no demuestra que el desequilibrio del flujo de órdenes prediga rendimientos ni que la estrategia filtrada sea rentable. Sus conclusiones dependerían de la construcción de características asociada, el ajuste del modelo, los backtests y los supuestos de ejecución.
Ideas clave
- Usa la composición del índice en cada momento y los datos de liquidez anteriores a la evaluación para definir las acciones elegibles.
- Mantén separado el ritmo de decisión del flujo de precios más detallado que se usa para vigilar las posiciones.
- Crea características de microestructura limitadas a la sesión con retrospectivas, desfases y modos de fallo explícitos.
- Espacia los ejemplos de entrenamiento según el horizonte de cada etiqueta para evitar muestrear repetidamente resultados solapados.
- Trata la medida de flujo de órdenes declarada como un tratamiento causal y evalúala con aprendizaje automático doble walk-forward.
Etiquetas
Texto completo
# setup.yaml
```yaml
strategy_id: nasdaq100_microstructure
setup_version: v1
universe:
# Point-in-time NASDAQ-100 membership over the sample, restricted to the names
# that contribute at least one session before evaluation.holdout_start. The
# published AlgoSeek archive carries 123 symbols; eight of them first quote
# after 2021-07-01, so nothing is ever fit on them, the pre-holdout liquidity
# profile has no spread for them, and 05_evaluation drops them before it
# screens anything. 01_feasibility_analysis asserts this list against the
# archive in both directions. Names that leave the index end where they left:
# AAL and WLTW stop at 2020-05-01, which is the membership, not a gap.
symbols:
- AAL
- AAPL
- ADBE
- ADI
- ADP
- ADSK
- AEP
- ALGN
- ALXN
- AMAT
- AMD
- AMGN
- AMZN
- ANSS
- ASML
- ATVI
- AVGO
- BIDU
- BIIB
- BKNG
- BMRN
- CDNS
- CDW
- CERN
- CHKP
- CHTR
- CMCSA
- COST
- CPRT
- CSCO
- CSGP
- CSX
- CTAS
- CTSH
- CTXS
- DLTR
- DOCU
- DXCM
- EA
- EBAY
- EXC
- EXPE
- FAST
- FB
- FISV
- FOX
- FOXA
- GILD
- GOOG
- GOOGL
- IDXX
- ILMN
- INCY
- INTC
- INTU
- ISRG
- JD
- KDP
- KHC
- KLAC
- LBTYA
- LBTYK
- LRCX
- LULU
- MAR
- MCHP
- MDLZ
- MELI
- MNST
- MRNA
- MRVL
- MSFT
- MTCH
- MU
- MXIM
- NFLX
- NTAP
- NTES
- NVDA
- NXPI
- OKTA
- ORLY
- PAYX
- PCAR
- PDD
- PEP
- PTON
- PYPL
- QCOM
- REGN
- ROST
- SBUX
- SGEN
- SIRI
- SNPS
- SPLK
- SWKS
- TCOM
- TEAM
- TMUS
- TSLA
- TTWO
- TXN
- UAL
- ULTA
- VRSK
- VRSN
- VRTX
- WBA
- WDAY
- WDC
- WLTW
- XEL
- XLNX
- ZM
n_assets: 115
eligibility_rule: nasdaq100_membership_with_pre_holdout_session
# Cost-feasible universe: the cheapest-to-trade names by round-trip cost,
# frozen per split. Provenance: _build_cost_feasible_universe.py
# (round-trip proxy 2*(per_share/mean_price)*1e4 + 2*median_half_spread_bps).
# The validation list is profiled on quote bars strictly before the
# validation window (2020-01-01 -> 2020-06-30); the holdout list from the
# pre-holdout liquidity profile (< 2021-07-01). No look-ahead. Selected via
# strategy.signal.universe_filter='cost_feasible'; the split-specific list is
# chosen at backtest time by resolving the prediction set's split (NOT a
# spec-hash input — like the 'liquid' filter, the hash carries only the
# filter name). The full 115 universe collapses out of sample (featured
# config holdout -0.42) while the cost-feasible universe is positive; the
# screen is load-bearing. See 17_costs.py for the full-vs-screened contrast.
cost_feasible:
validation:
- AAPL
- MSFT
- FB
- SBUX
- PEP
- ATVI
- INTC
- PYPL
- MDLZ
- AMD
- QCOM
- TXN
- GILD
- MU
- NVDA
- AMAT
- CSCO
- TMUS
- WBA
- COST
- JD
- CSX
- XEL
- CMCSA
- EXC
- EBAY
- MNST
- FISV
- FAST
- EA
- CTSH
- NFLX
- XLNX
- AMZN
- ADBE
- PAYX
- CDNS
- AMGN
- MXIM
- BIDU
- PCAR
- CERN
- UAL
- ADP
- KHC
- FOXA
- WDC
- TSLA
- FOX
- CTXS
holdout:
- AAPL
- MSFT
- FB
- SBUX
- PEP
- GILD
- AMD
- QCOM
- INTC
- MDLZ
- ATVI
- AMAT
- MU
- PYPL
- AEP
- TXN
- JD
- CSX
- MRVL
- COST
- TMUS
- CMCSA
- EBAY
- CSCO
- NVDA
- XEL
- FISV
- WBA
- EXC
- FAST
- TSLA
- CTSH
- MXIM
- ADI
- AMZN
- NFLX
- AMGN
- EA
- PAYX
- CERN
- ADBE
- ADP
- KHC
- CDNS
- XLNX
- MNST
- KDP
- BIDU
- DLTR
- UAL
decision:
# How often the strategy is allowed to act, NOT the size of a bar in the panel.
# The observations are one minute apart; 03_financial_features keeps every
# fifteenth of them to build the decision schedule. Nothing here declares the
# observation grid, so anything that needs it - an embargo, a permutation block,
# any count of periods spanning a duration - measures it from the data instead
# of dividing by this number. Reading this field as the bar size yields an
# answer fifteen times too small and nothing in the result shows it.
bar_frequency: 15_minute
# A model decides on its own horizon: a five-minute model every five minutes, a
# sixty-minute model every hour. `bar_frequency` above is the default for a label that
# declares nothing here, and stays the primary label's cadence.
#
# This is the decision schedule only. The engine's price feed is one minute
# (`config/backtest/base.yaml::calendar.data_frequency`), so a position entered on one
# decision is watched every minute until the next - which is what a stop or a
# take-profit is measured on, and what a fifteen-minute feed could not express.
cadence_by_label:
fwd_ret_5m: 5_minute
fwd_ret_15m: 15_minute
fwd_dir_15m: 15_minute
fwd_ret_60m: 60_minute
decision_snapshot: bar_close
execution_delay: 1_bar
execution_price_assumption: next_bar_open_or_vwap
# The feature specification read by 03_financial_features. Windows are counted
# in minute bars and every one of them is session-bounded: a window never spans
# an overnight gap, so each session restarts its own warmup.
features:
# Feature matrices are stored and fitted in single precision. 16,098,877 rows x 88 features
# in float64 is a 10.9 GB modeling dataset and a 58.9 GB peak through one gradient-boosting
# fold; in float32 those are 5.6 GB and 38.0 GB. LightGBM bins to uint8 regardless, and the
# linear families fit standardised columns where the extra mantissa carries nothing this
# case study can measure. Declared here rather than globally: the smaller case studies fit
# comfortably in double precision, and narrowing them would move their numbers for nothing.
storage_dtype: float32
windows:
fast: 5 # one fast aggregate
decision: 15 # one decision bar at the declared bar_frequency
slow: 30 # realized-volatility horizon and the Amihud average
hour: 60 # Kyle's lambda regression and the slow regime average
ewma_half_life: 30 # half-life of the EWMA volatility, in bars
stale_cap: 5 # consecutive stale-quote bars tolerated before nulling
edge_block: 30 # width of the open and close blocks the flags mark
# The quantity the thesis puts forward: the order-flow imbalance measured over
# one decision bar. Its one-bar form is this case study's causal treatment (see
# `causal.treatment` below). Read by 03_financial_features, which draws it in
# F2, F3 and F6.
carrier: signed_vol_share_15m
# One row per family. `pattern` claims a family's columns in both
# representations - the level and its cross-sectional z-score twin - because
# the two carry the same hypothesis on two scales. `lookback` and `lag` are
# minute bars, and they are what the warmup audit and the timing figure read.
families:
- name: quote_liquidity
pattern: rel_spread_*|depth_imb*|quote_rate*
role: state
hypothesis: >-
The cost and the depth of the book describe the environment a signal is read
in: an imbalance is worth acting on only where the round trip is cheap enough
to leave something behind.
inputs: NBBO close quotes, quote update count
lookback: 60
lag: 0
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
A carried-forward quote reports a stale spread as a live one, so the state
looks calm exactly where the book has stopped updating.
- name: microprice
pattern: microprice_dev*
role: signal
hypothesis: >-
The depth-weighted price leans toward the thin side of the book, so its distance
from the midpoint is where the next trade is more likely to print (Stoikov, 2018).
inputs: NBBO close prices and sizes
lookback: 15
lag: 0
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
The deviation is a few tenths of a cent on a wide book and a rounding artifact
on a one-tick one, so it is not comparable across names until it is ranked.
- name: volatility
pattern: r1m*|rv_*|quote_range*
role: state
hypothesis: >-
How much the quoted price is moving is what decides whether an imbalance is
information or noise, and it varies by name and by hour.
inputs: NBBO quote midpoint and quote extremes
lookback: 60
lag: 0
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
Every estimator here is a variance over a short window, so a single bad quote
dominates it and reports turbulence where there was a data error.
- name: order_flow
pattern: signed_vol_share*|tick_imb_share*|trade_to_mid*|trades_per_1k_shares*|cross_locked_share*
role: signal
hypothesis: >-
Aggressive volume arrives in runs, so the side that has been paying the spread
over the last few minutes keeps paying it over the next few.
inputs: AlgoSeek trade-location and tick buckets, total traded volume
lookback: 60
lag: 1
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
Trade location is assigned against the prevailing quote, so a bar whose quote
was stale attributes its volume to the wrong side rather than to neither.
- name: price_impact
pattern: dollar_vol*|illiq*|kyle_lambda*|trade_range*
role: state
hypothesis: >-
How far a given quantity of order flow moves the price is what separates a
crowded name from a thin one, and it is the cost side of every signal here.
inputs: consolidated trade prices, traded dollars across both venues
lookback: 60
lag: 1
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
Amihud is undefined on a bar that printed nothing, and its rolling average
carries that gap across the whole window rather than skipping the bar.
- name: hidden_liquidity
pattern: finra_share_60m*
role: state
hypothesis: >-
The share of volume printing away from the exchanges tracks how much of the
session's liquidity is not visible in the book, and it moves over hours.
inputs: FINRA/TRF reported volume, total traded volume
lookback: 60
lag: 1
frame: per symbol-session, trailing
representation: level and cross-sectional z-score
failure_mode: >-
TRF prints are reported with a delay, so the share attributed to a bar is not
fully public at that bar's close unless it is lagged.
- name: session_clock
pattern: time_since_open|time_to_close|is_first_30m|is_last_30m
role: state
hypothesis: >-
Intraday liquidity is U-shaped, so where in the session a bar sits conditions
every other family without predicting anything on its own.
inputs: exchange calendar
lookback: 0
lag: 0
frame: per session, from the scheduled open and close
representation: fraction of the session and block flags
failure_mode: >-
Measured against the session's realized bar count rather than its scheduled
close, it is a quantity no one had at the time.
# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study.
#
# Spec-hash inputs (each invalidates every backtest_hash on change):
# - execution.initial_cash
# - execution.share_type
# - execution.allocator_lookback (CS-level fallback for moment allocators)
# - per-allocator overrides in backtest.sweep.allocators
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator. It is counted in ROWS OF THE PRICE
# FRAME, and that frame is one-minute: ``config/backtest/base.yaml`` declares
# ``calendar.data_frequency: 1m``, ``load_backtest_prices_for`` returns bars
# measured at 60-second spacing, and ``_compute_rolling_vol`` applies
# ``rolling_std(vol_window)`` to those rows with no resampling. So 520 is 520
# MINUTES - about 1.3 regular sessions - and not the 20 trading days this comment
# claimed until 2026-09-06. The value is unchanged; what was wrong was the
# sentence describing it, which read the decision cadence as the bar size.
# ``periods_per_year`` annualizes Sharpe at daily-equivalent grain (252),
# independently of this window.
#
# Expensive allocators (RP/HRP/MVO_LW) are dropped via expensive_allocators_skip
# below — covariance estimation at 15-min cadence is prohibitive and
# empirically degenerate. MVO_LW carries no per-allocator lookback override
# here for that reason; revisit if expensive allocators are reintroduced
# (a 6-month equivalent at 15-min would need ~3,300 bars).
#
# ``initial_cash`` restored to 1_000_000 (2026-05-16) — the 2026-05-15 SSOT
# migration drop to 100k caused catastrophic degenerate output here
# (avg Sharpe -11 across 339 signal runs; integer-share rounding × the
# nasdaq-100 high-priced tail (BKNG≈$5k, MSTR/NFLX/AVGO $700-1.5k)). See
# memory/feedback_2026_05_15_equity_sizing_invalidated.md.
execution:
initial_cash: 1_000_000 # restores prior validated state
share_type: integer # US equities trade in whole shares
allocator_lookback: 520 # rows of the 1-minute price frame: ~1.3 sessions
mapping:
class: intraday_rank_and_trade
position_state_space: long_short
entry_logic: rank_or_threshold
sizing: dollar_neutral_or_beta_neutral
costs:
class: dominant
# Per-share commission plus measured per-asset half-spread slippage from
# AlgoSeek NBBO close quotes over the training window. Integer share
# sizing. The asset_spreads_source parquet is loaded by the backtest
# cost-preset builder and joined per symbol; default_half_spread_usd is
# the universe p75 fallback for symbols not in the profile.
model: per_share_plus_spread
per_share: 0.0035 # IBKR Pro Tiered top tier (<=300K shares/mo)
minimum: 0.0
asset_spreads_source: liquidity_profile.parquet
asset_spreads_column: median_half_spread_usd
default_half_spread_usd: 0.025
spread_convention: half_spread
source: AlgoSeek minute-bar NBBO close quotes, training window 2020-01-02 to 2021-07-01.
friction_floor_bps: 5
backtest:
rebalance:
# A rebalance is skipped when the per-asset weight change is below
# min_weight_change AND the resulting trade notional is below
# min_trade_value. The benchmark profile disables thresholds so that
# full-universe equal-weight (1/N per asset) rebalances at all.
default:
min_weight_change: 0.005
min_trade_value: 100.0
benchmark:
min_weight_change: 0.0
min_trade_value: 0.0
sweep:
# Canonical strategy universe. The 15-min round-trip cost on the
# cost-expensive tail of the 115-name panel consumes the intraday edge:
# the featured config collapses out of sample on the full universe
# (holdout -0.42) but is positive on the cost-feasible subset (the
# cheapest-to-trade names, frozen per split — see universe.cost_feasible).
# The full-universe variant is NOT a canonical rank-1 / cohort / DSR
# candidate; it lives only in the 17_costs.py full-vs-screened comparison,
# mirroring sp500_options' liquid screen. Read by
# case_studies.utils.sweep_config.get_universe_filters_for.
universe_filter: cost_feasible
# Iteration controls per stage. ``signal: 0`` means "all predictions";
# downstream stages take the top-N from the upstream stage's rank-1.
# Notebooks read these via get_top_n_predictions(case_study, stage).
top_n_predictions:
signal: 0 # all signal predictions (eq-weight baseline)
allocation: 10 # top-10 model configs by equal-weight baseline Sharpe
# top-1 of every pre-cost stage per label. `17_costs` pools {signal, allocation,
# risk_overlay} - `STAGE_SEQUENCE` minus the stage it writes - because an overlay is a
# candidate to carry the case study and pricing only its parent would put a cost curve
# in the chapter for a strategy this case study does not select.
cost_sensitivity: 1
risk_overlay: 1 # top-1 of {signal+allocation} per label
# No expensive allocators on nasdaq. At 15-min frequency the 1.3M-bar
# panel makes covariance estimation on the rebalance schedule
# prohibitive; empirically every expensive allocator (risk_parity,
# mvo_ledoit_wolf, hrp) produces deeply negative Sharpes (-7 to -10)
# on every nasdaq label and equal_weight ties at the top. The 15-min
# cadence is a costs-dominated regime where signal-stage selection
# carries the entire story.
expensive_allocators_skip: true
# Ch16 signal-stage selection. Long-short equal-weight top-k on the
# 15-min schedule across four labels (5m/15m/60m continuous + 15m
# direction). The classification label fwd_dir_15m is a first-class
# signal alongside the continuous labels and carries the same top_k
# grid.
top_k_grid:
fwd_ret_5m: [5, 10, 20]
fwd_ret_15m: [5, 10, 20]
fwd_ret_60m: [5, 10, 20]
fwd_dir_15m: [5, 10, 20]
# Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
# Nasdaq runs only the three fast allocators — the moment-based ones are
# too slow at intraday frequency (``expensive_allocators_skip: true``).
# No max_weight cap on inverse_vol — see
# memory/feedback_max_weight_caps_intentionally_absent.md (would force
# equal-weight at top_k=5 and defeat the moment-allocator comparison).
allocators:
- {name: equal_weight, method: equal_weight}
- {name: score_weighted, method: score_weighted}
- {name: inverse_vol, method: inverse_vol}
# Ch18 cost sensitivity (bps regime).
cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
# Companion per-share cost regime: dollars per share half-spread. Swept
# alongside cost_grid_bps for CSes whose declared cost model is
# per_share_plus_spread. Values: 0¢, 0.5¢, 1¢, 2.5¢, 5¢, 10¢.
cost_grid_half_spread_usd: [0.0, 0.005, 0.01, 0.025, 0.05, 0.10]
# Alternative rebalance cadences explored by Ch18's cadence × cost
# heatmap (17_costs.py Section 4). The first entry must be the CS's
# default cadence (``decision.bar_frequency``). Each cadence reuses the
# ``cost_grid_half_spread_usd`` grid above. Tokens are the same cadence
# vocabulary recognized by the engine; the notebook pulls the list from
# here, never hardcoded.
cadence_sweep: [15_minute, 30_minute, 1_hour, 4_hour]
# nasdaq100 v4 slot-mechanism sweep (Ch16 signal stage replacement).
#
# Adds two selection methods to the canonical Ch16 expansion:
# - slot_persistent_signal_exit: per-symbol rolling-percentile entry,
# fixed weight-per-slot allocation (slots ARE the allocation —
# Ch17 allocator stage is skipped for these rows), max-hold +
# optional signal-exit (rolling lower quantile crossover) for
# position lifecycle. See case_studies/utils/slot_strategy.py for
# the mechanism. Sub-params live under selection_method_config in
# the flattened scheme dict.
# - eq_w_topk: the canonical equal-weight top-K baseline, kept as
# a within-sweep comparison. Direction grid covers long_only,
# long_short.
#
# slot × long_short is intentionally dropped — slot books are
# single-direction by construction (one symbol cannot occupy a long
# AND short slot simultaneously). The cross-asset long-short story
# belongs to eq_w_topk / quintile_long_short.
#
# Per-prediction cardinality:
# slot: long_q(3) × max_slots(3) × hold_bars(2) × exit_signal_q(2)
# × bars_per_day_grid(1) = 36 (long_only only)
# eq_w_topk: top_k(3) × direction(2) = 6
# 42 per prediction, and `signal_passes` below says how many predictions
# each of them runs on.
#
# `hold_bars` and `exit_signal_q` were cut here, from 3 and 4 values to 2
# each, on 2026-09-11. Both keep their endpoints: hold_bars keeps 2h and 8h
# and drops the 4h midpoint, exit_signal_q keeps "off" and the median
# crossover and drops 0.30 and 0.70. A midpoint is what a grid can lose
# without losing the comparison it draws - the question the sweep asks is
# whether a longer hold and a signal exit move the result, and two values
# answer it. `long_q` and `max_slots` are untouched: `14_backtest.py`'s
# Act 2 selects the carrier by `slots == 10` and `entry_q == 0.9`
# literally, so removing either value empties the table that section
# prints.
signal_nasdaq100:
selection_method: [slot_persistent_signal_exit, eq_w_topk]
long_q: [0.90, 0.95, 0.99]
direction: [long_only, long_short]
max_slots: [5, 10, 20]
# hold_bars and bars_per_day count rows of the grid the slot simulator walks, which is
# the PRICE grid, now one minute. Restated from the fifteen-minute counts they carried
# (8/16/32 and 14) at the same wall-clock durations, so the sweep asks the same question
# of the same intervals rather than silently asking about 8, 16 and 32 minutes.
hold_bars: [120, 480] # 2h, 8h, in minute bars
exit_signal_q: [null, 0.50] # null disables signal-exit
pred_freshness_max_min: 14 # tolerated staleness of a prediction, in minutes
bars_per_day_grid: [210] # was 14 bars of 15 minutes; same 3.5h window
top_k_grid: [5, 10, 20]
lookback_days: 21
# How the signal stage walks the grid above. Two passes, because the whole
# grid over every prediction is not affordable here and nothing else in the
# book has this shape: every other case study declares 2 to 4 entry schemes
# and one universe, and no completed case study holds more than about 3,000
# signal backtests.
#
# Counted against this registry and measured on 2026-09-11. A backtest of
# this panel costs 31.8 s (four arms timed on the live 10.1M-row price
# frame, both universes, 29.2 s to 36.3 s), and the registry offers 741
# admissible prediction sets: 226 for each of the three continuous labels
# and 63 for `fwd_dir_15m`.
#
# pass 1 baseline 741 × 3 canonical arms × cost_feasible = 2,223
# pass 2 mechanism 32 × 42 arms × cost_feasible = 1,344
# pass 2 reference 32 × 3 canonical arms × full = 96
# total = 3,663 ≈ 32 h
#
# Crossing all 117 arms with both universes over all 741 predictions, which
# is what this block replaces, is 173,394 backtests and about 1,500 hours.
# The prediction count is not what was wrong with it: 741 sets is what three
# DL families at 20 epoch checkpoints, 15 gbm configs at 10 tree-count
# checkpoints and 16 linear configs come to, and every one of them is a
# configuration this case study fitted on purpose. The arm count was.
signal_passes:
# Pass 1 ranks every admissible prediction on three equal-weight
# concentrations, on the universe the carrier actually trades. Ranking on
# the full universe instead would order the predictions by a Sharpe no
# selected strategy ever earns, because the cost-expensive tail of the
# panel is exactly what `universe.cost_feasible` removes.
baseline_schemes: [ew_top5, ew_top10, ew_top20]
baseline_universe: cost_feasible
# Pass 2 runs the mechanism grid on the strongest predictions of pass 1,
# per label, ranked by pass-1 validation Sharpe.
#
# By Sharpe and not by IC. `top_n_predictions.signal` is the other lever
# that would cut this grid and it is the wrong one: `load_prediction_index`
# ends `ORDER BY m.ic_mean DESC NULLS LAST`, so any positive value there
# screens the signal stage on rank correlation and decides which models are
# backtested at all before a backtest has run. All nine case studies
# declare `signal: 0` for that reason, and it stays 0 here.
mechanism_top_n: 8
# The full universe is kept as a reference arm rather than a second copy of
# the sweep: the canonical concentrations only, on the pass-2 predictions.
# That is what `17_costs` reads for its full-vs-screened comparison - it
# prices the top-1 of each pre-cost stage, which is always inside pass 2 -
# and what Act 1 of `14_backtest` reads to show the unscreened baseline.
reference_schemes: [ew_top5, ew_top10, ew_top20]
reference_universe: full
# Ch19 risk overlays.
risk_controls:
position:
- {name: stop_loss_3pct, type: stop_loss, threshold: 0.03}
- {name: stop_loss_5pct, type: stop_loss, threshold: 0.05}
- {name: stop_loss_10pct, type: stop_loss, threshold: 0.10}
- {name: stop_loss_15pct, type: stop_loss, threshold: 0.15}
- {name: trailing_1pct, type: trailing_stop, threshold: 0.01}
- {name: trailing_2pct, type: trailing_stop, threshold: 0.02}
- {name: trailing_3pct, type: trailing_stop, threshold: 0.03}
- {name: trailing_5pct, type: trailing_stop, threshold: 0.05}
- {name: trailing_10pct, type: trailing_stop, threshold: 0.10}
- {name: trailing_15pct, type: trailing_stop, threshold: 0.15}
- {name: trailing_20pct, type: trailing_stop, threshold: 0.20}
- {name: time_exit_10, type: time_exit, bars: 10}
- {name: time_exit_20, type: time_exit, bars: 20}
- {name: time_exit_40, type: time_exit, bars: 40}
# The mean-forecast ensemble this case study features, and the only thing in the repository
# that produces `family: ensemble` rows. Declared rather than assembled in the notebook
# because the member set is what the ensemble *is*: change the rule and it is a different
# object under the same name.
#
# The lesson this case study ends on is that choosing among a family's configurations on a
# validation window does not carry to a later window, while the family itself does. An object
# that makes no choice is what shows it, and that object has to be registered like a model so
# the backtest stage, the cohort comparison and the holdout all read it the way they read one.
ensemble:
# The family averaged, and the rule that decides which of its configurations are in.
#
# 31 leaves is the boundary between the regularized presets and the one that is not:
# `leaves_7`, `leaves_15` and `leaves_31` carry `lambda_l1`, `lambda_l2`, `bagging_fraction`
# and `feature_fraction`; `default` carries none of them and trains at LightGBM's own
# default of 31 leaves; `leaves_63` is the only configuration above the line. The rule keeps
# twelve of the fifteen for each continuous label - three losses across four leaf counts.
#
# A count is not declared here on purpose. The rule is the declaration and the members are
# resolved from what is registered, so a configuration added to `config/training/*.yaml`
# joins the ensemble and a number written down here could only disagree with it.
member_family: gbm
max_num_leaves: 31
method: mean_forecast
config_name: gbm_mean_leaves31
# Each member enters at its last checkpoint. Taking the checkpoint with the best validation
# metric would select on the window the ensemble is then measured on, which is the thing
# the ensemble exists not to do.
checkpoint: last
# The entry scheme the ensemble is backtested on, by the name `get_entry_schemes_for`
# gives it: 10 slots, an 8-hour maximum hold, a 0.90 entry quantile, no signal exit, on the
# 210-bar execution window. `14_backtest.py` Act 2 reads the carrier by `slots == 10` and
# `entry_q == 0.9` literally, so this arm and that filter have to agree.
#
# One arm and not the grid: the ensemble is not a sweep participant. It is the object the
# chapter carries, backtested on the design the sweep already chose, and putting it through
# the sweep would make it one more configuration to select among.
featured_scheme: slot_l_lq90_s10_h480_noexit_b210
universe_filter: cost_feasible
# The model specification 04_model_based_features estimates under. Declared here rather than
# typed into the notebook so that a window is one number in one place, readable without opening
# the code that consumes it.
#
# This case study fits no GARCH, no regime chain and no ARIMA: its three procedures are a
# heterogeneous autoregression of realized variance, a rolling Fourier transform and a rolling
# depth-2 path signature. Only the first estimates anything, which is why only `har` carries a
# refit schedule.
model_based:
har:
# The three lengths of history the regression averages squared one-minute returns over,
# in bars. Corsi's day/week/month on a minute grid: minutes, a quarter hour, an hour.
components: [5, 15, 60]
# Bars of history each refit reads. Two hours - long enough for four coefficients to be
# identified, short enough that the fit follows the day rather than the quarter. It is
# also the burn-in: the first `fit_window + 1` bars of a symbol carry no HAR value,
# because they are what pays for the first estimate.
fit_window: 120
# Bars between refits. One, which is the finest cadence the panel has and the reason this
# artifact carries no fold column: the coefficients describing bar t come from a
# regression whose last observation is bar t-1, so no fold boundary enters the fit and a
# bar's value is the same value whichever fold selects it.
refit_every: 1
# Usable observations a refit window needs before it is fitted at all. Below this the four
# coefficients are estimated off too few points to identify them, and the bar keeps a null
# rather than a fit on whatever survived the warm-up.
min_train_obs: 20
spectrum:
# Bars each trailing window spans. Sixty resolves repetition periods up to an hour, which
# is the longest intraday rhythm this panel is asked about.
window: 60
# The period, in bars, above which power counts as low frequency. Twenty separates a slow
# drift in activity from minute-to-minute churn, and it is what `*_low_freq_ratio` is the
# share of.
low_frequency_period: 20
signature:
# Bars each signature path spans. Thirty is the horizon over which order flow and price are
# expected to lead one another, which is the quantity the cross terms measure.
window: 30
evaluation:
n_splits: 2
train_size: 6M
val_size: 6M
holdout_start: '2021-07-01'
holdout_end: '2021-12-31'
calendar: NYSE
periods_per_year: 252 # NYSE 5d/wk (daily MTM despite intraday signal cadence)
labels:
primary: fwd_ret_15m
# The outcome horizon each label's name states. Declared rather than left to the default,
# because `resolve_label_horizon` falls back to the BUFFER when no horizon is given - so
# widening a buffer below would silently lengthen the return the label measures, and
# `fwd_ret_15m` would be a sixteen-minute return that every consumer still called fifteen.
horizons:
fwd_ret_15m: 15min
fwd_ret_5m: 5min
fwd_ret_60m: 60min
fwd_dir_15m: 15min
# The purge, which is the horizon PLUS ONE BAR - wider than the horizon, and deliberately.
# The entry leg is the VWAP of the bar after the decision and the exit leg is H past the
# entry, so a label at t consumes a quote at t+H+1: one bar past the horizon its name
# states. A buffer of H would leave the last bar of every training window inside the first
# validation label's window. The extra minute is not conservatism; it is the width of the
# window that the horizon alone does not describe.
buffer: 16min
variants:
- fwd_ret_5m
- fwd_ret_60m
- fwd_dir_15m
variant_buffers:
fwd_ret_5m: 6min
fwd_ret_60m: 61min
fwd_dir_15m: 16min
# How many slots of the DECISION schedule to advance per trade, so a position is held for
# its label's horizon and no two holdings overlap. A slot is one interval of the declared
# cadence, so these are `ceil(horizon / cadence)` and nothing else.
#
# That reading is only true because `resolve_rebalance_timestamps` now builds the schedule
# from the declared cadence. It used to return the panel's own timestamps unchanged for
# every intraday cadence, and this panel carries every minute, so these values thinned
# minutes rather than fifteen-minute slots and three of the four labels traded every
# minute (ml4t/agent-workspace#187). The values did not change; what they count did.
# Every label now trades on a cadence equal to its own horizon (decision.cadence_by_label
# above), so one slot is one holding period and no label needs a multi-slot step. The
# `fwd_ret_60m: 4` this replaces was `ceil(60 / 15)` against a single fifteen-minute
# default; the cadence carries that now, and carries it for the Ch18 sweep arms too,
# which differ from the default in exactly this token.
rebalance_step:
fwd_ret_15m: 1
fwd_ret_5m: 1
fwd_ret_60m: 1
fwd_dir_15m: 1
# Continuous return that each classification label is derived from.
classification_eval_label:
fwd_dir_15m: fwd_ret_15m
modeling:
dl:
# The backend the sequence families are trained on in production. Declared rather
# than defaulted: the four deep-learning notebooks refuse to train on a device other
# than the one asked for, because the device is part of what the run is registered
# under, and a notebook that invents `gpu` from a missing section makes a
# CUDA-free environment fail on a requirement nobody wrote down. A run on other
# hardware sets the notebook's DEVICE parameter, which is then what gets registered.
device: gpu
# How far apart the training windows of one sequence configuration sit, counted in
# label horizons. This is the only case study that needs the key: every row of a panel
# starts a window, and this panel is minute bars, so an unspaced fold builds about 4.0
# million near-identical overlapping sequences where the daily case studies build tens
# of thousands. Measured on the canonical plan, 2026-09-02: fold 0 has 4,066,041
# training rows and fold 1 has 3,997,571, across 115 symbols.
#
# One window per label horizon. Consecutive windows of a symbol then carry labels that
# do not overlap: the window ending at t is scored on the return from t to t+H, and the
# next one starts at t+H. Every window is still an example the model has to fit; what
# changes is that no two of them are the same example seen twice.
#
# Declared in horizons rather than in windows because the horizon differs by label, and
# one line then covers all four: `fwd_ret_15m` and `fwd_dir_15m` stride 15 observations,
# `fwd_ret_5m` strides 5 and `fwd_ret_60m` strides 60, each spacing its own labels
# exactly. A single count could only be right for one of them. On fold 0 the primary
# label draws about 271,000 windows, and the count is derived at run time and printed,
# not written down here - it differs between folds because the folds differ in length.
#
# This replaces `max_train_sequences: 750000`, which had been carried unchanged since
# the original release and through the stage-06 rebuild that moved the folds underneath
# it, so nothing derived it. At 750,000 a window was drawn every 5.4 minutes, finer than
# the 15-minute horizon, so neighbouring training windows carried overlapping labels and
# the extra 2.8x of them were re-presentations of examples already in the batch. That is
# not wrong - correlated examples are not invalid ones - but it was a number with no
# argument behind it, and the number is part of the training identity
# (ml4t/agent-workspace#1015).
#
# `max_train_sequences` is still available and is what a case study declares when it
# wants the overlap: it fixes the count and lets the spacing follow. The two are
# mutually exclusive.
train_sequence_stride_horizons: 1
gbm:
libraries: [lightgbm]
preset: default
device: cpu
# LightGBM's own CPU default. 63 is the GPU default and was carried over with the
# device when these runs moved off the GPU, so every CPU fit was quantizing the design
# matrix into a quarter of the bins the library would have used. Coarser bins are
# faster and lose split points; the reader running this on a CPU gets what the
# documentation describes.
max_bin: 255
causal:
treatment: signed_vol_share
# Bars the treatment's own construction window spans, which is what the placebo block has
# to cover: permuting signed_vol_share in blocks shorter than this destroys the serial
# dependence the refutation exists to preserve, and the resulting p-value reads like a
# refutation without being one. Declared here rather than inferred, because guessing which
# element of a window list a column was built from puts a wrong number behind a right-looking
# one. Derived from the construction, not chosen:
#
# `signed / TRADED_VOLUME` in 03_financial_features, both from the same bar. The smoothed
# versions this feeds are separate columns; the treatment itself spans one bar.
treatment_window: 1
confounders: [rel_spread_close, rv_5m, r1m]
method: walk_forward_dml
```Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.