Señales bursátiles con opciones y configuración del backtest
Resumen
Este documento de configuración define un proceso semanal de investigación de acciones del S&P 500 que combina características del mercado de opciones con señales basadas en precios. Especifica el universo elegible, los horarios de decisión y ejecución, y distintas frecuencias de rebalanceo para etiquetas con diferentes horizontes de previsión. Las familias de características incluyen niveles y dinámica de la volatilidad implícita, sesgo y estructura temporal, prima de riesgo de varianza, volatilidad realizada, momentum de las acciones, rangos transversales y calidad de la superficie. El documento también registra ventanas retrospectivas, retardos de información, representaciones, hipótesis y modos de fallo, como cotizaciones desactualizadas o diferencias entre la volatilidad implícita prospectiva y la volatilidad realizada retrospectiva. Además, establece posiciones long-only con ponderación igual basada en rangos, supuestos de ejecución, componentes de costes de transacción, umbrales de rebalanceo y pruebas por etapas de asignación y superposición de riesgo. Una sección aparte configura la estimación walk-forward de un modelo de volatilidad GJR-GARCH, incluidos el periodo de calentamiento y el calendario de reajuste, y un análisis causal mediante double machine learning walk-forward con un diseño placebo que tiene en cuenta bloques. Estas declaraciones permiten inspeccionar los datos, las características y las decisiones de análisis previstos. Es una especificación de investigación, no evidencia de que ninguna señal o estrategia sea rentable. Sus supuestos, estimaciones de costes, construcción de características y decisiones de modelado definen el experimento y requieren validación empírica.
Ideas clave
- La configuración combina características de la superficie de opciones con medidas de momentum bursátil y volatilidad realizada.
- Las ventanas y los retardos de las características codifican cuándo se observan los datos de entrada y cómo se construyen sus valores.
- La estrategia usa una clasificación long-only y ponderación igual, con costes de ejecución y reglas de rebalanceo declarados.
- Las etiquetas de previsión con horizontes más largos usan una frecuencia de rebalanceo menor para limitar las posiciones solapadas.
- La configuración especifica la estimación GJR-GARCH y un análisis causal walk-forward, pero no demuestra por sí misma que las señales sean rentables.
Etiquetas
Texto completo
# setup.yaml
```yaml
strategy_id: sp500_equity_option_analytics
setup_version: v1
universe:
n_assets: 633
eligibility_rule: sp500_with_options
decision:
cadence: weekly_friday_close
snapshot: friday_16:00_et
execution_delay: monday_open
iv_feature_lag: 1_day
# Per-label decision cadence; `cadence` above applies to every label not named here. The five
# 5-day labels trade the weekly grid, which is what they forecast. The two 10-day labels
# resolve after two weeks, so on the weekly grid a new position is opened while the previous
# one is still inside its own horizon - overlapping the very quantity being measured. They
# rebalance every other week instead.
cadence_by_label:
fwd_ret_10d: biweekly
fwd_dir_10d: biweekly
# The feature-specification register for 03_financial_features, and every window
# it computes over. It lives here rather than in the notebook for the same reason
# the label name and the holdout boundary do: it is the statement of what this
# case study's feature set is, and a statement only the notebook holds cannot be
# read by a test, by a later stage, or by anyone asking what changed.
#
# ``lookback`` is counted in daily bars back from the decision timestamp and is
# the floor the warmup audit holds each column to. ``lag`` is the delay with
# which the input becomes knowable: every option-derived family carries one
# session because the surface summary is stamped at the close it summarizes and
# is not read until the next session's decision, and every equity-price family
# carries zero because the close is the decision snapshot itself.
features:
# Surface contract selection, fixed by data/equities/market/sp500/materialize_options.py.
# Repeated here only so section B can state the observability of what it loads.
surface:
dte_buckets: {'7d': [5, 10], '30d': [25, 35], '90d': [80, 110]}
delta_targets: {atm: 0.50, '25d': 0.25, '10d': 0.10}
windows:
iv_zscore: [63, 252]
iv_percentile: 252
iv_momentum: [5, 21]
skew_zscore: 63
term_zscore: 63
vrp_zscore: 63
realized_vol: [20, 63]
garman_klass: 21
vol_of_vol: 21
realized_skew: 21
momentum: [5, 21, 63, 126, 252]
skip_recent: 21 # skip-month momentum runs t-252 to t-21
skip_start: 252
risk_adjusted: 63 # the momentum horizon its own volatility scales
# Annualized volatility floor in the risk-adjusted momentum denominator. A share
# that realized almost nothing over a quarter would otherwise carry an unbounded
# ratio, and one such row dominates the within-date percentile taken from it.
risk_adjusted_vol_floor: 0.01
iv_forward_fill: 5 # sessions a lagged surface value may be carried forward
# Source column -> the name its within-date percentile is written under. The
# names predate this register and later stages select by them, so the mapping
# is explicit rather than a suffix rule.
ranked:
iv_30_atm: iv_rank
skew_rr_30_25d: skew_rank
ivrv_spread: vrp_rank
mom_21d: mom_21d_rank
mom_63d: mom_63d_rank
rv_20: rv_rank
iv_mom_21d: iv_mom_rank
d_iv_30_atm: d_iv_rank
families:
- name: cross-sectional rank
pattern: '*_rank'
role: signal
hypothesis: A long-short book acts on relative standing, not on the level of a quantity
inputs: the eight level and dynamics columns named under features.ranked
lookback: 252
lag: 1
frame: cross section within the decision date
representation: percentile within the date, in (0, 100)
failure_mode: a thin cross-section makes the percentile a coarse ordering of few names
- name: implied volatility level
pattern: iv_30_atm|iv_7_atm|iv_90_atm|iv_30_put_25d|iv_30_call_25d
role: signal
hypothesis: What the option market charges for a name's coming month is priced against what its shares then do
inputs: daily IV surface summary, delta-selected within fixed DTE buckets
lookback: 1
lag: 1
frame: time series
representation: annualized implied volatility, in variance points
failure_mode: a name whose surface is quoted thinly carries a level set by one stale contract
- name: implied volatility dynamics
pattern: d_iv_30_atm|iv_mom_*|iv_30_atm_z_*|iv_30_atm_pct_*
role: signal
hypothesis: Where implied volatility sits against its own recent history says more than its level
inputs: at-the-money 30-day implied volatility
lookback: 252
lag: 1
frame: time series
representation: first difference, multi-session change, trailing z-score and trailing percentile
failure_mode: the z-score is unbounded as a name's own dispersion approaches zero
- name: skew and term structure
pattern: skew_*|term_*|d_skew_*|d_term_*
role: signal
hypothesis: The shape of the surface prices asymmetry and horizon that its level cannot express
inputs: 25-delta put and call IV, and the 7-, 30- and 90-day at-the-money points
lookback: 63
lag: 1
frame: term structure
representation: differences and ratios between surface points, and their trailing z-scores
failure_mode: the far segment reads a bucket whose contracts were interpolated rather than quoted
- name: variance risk premium
pattern: ivrv_spread|vrp_z_*
role: signal
hypothesis: A name whose implied volatility stands above what it goes on to realize is richly priced
inputs: at-the-money implied volatility and trailing realized volatility
lookback: 63
lag: 1
frame: time series
representation: spread in volatility points and its trailing z-score
failure_mode: it compares a forward-looking month against a backward-looking one, which differ around an event
- name: realized volatility
pattern: rv_20|rv_63|gk_vol_*|vol_of_vol_*|realized_skew_*
role: state
hypothesis: Dispersion and the asymmetry of the realized path set what a ranking can earn
inputs: split- and dividend-adjusted OHLC bars
lookback: 63
lag: 0
frame: time series
representation: annualized standard deviation, a range-based estimator, and third moments
failure_mode: close-to-close misses the overnight gap the range-based estimator is here to catch
- name: equity momentum
pattern: mom_*
role: signal
hypothesis: Option-derived information has to earn its place beside the price signal it would replace
inputs: split- and dividend-adjusted close
lookback: 252
lag: 0
frame: time series
representation: simple return at five horizons, its skip-month form, and its volatility-scaled twin
failure_mode: reverses over the most recent month, which skip-month momentum drops
- name: surface quality
pattern: spread_atm_*|qc_converged_share
role: state
hypothesis: A surface point solved from wide or unconverged quotes should not be read like one that was not
inputs: bid-ask spread and solver convergence flags of the selected contracts
lookback: 1
lag: 1
frame: time series
representation: relative spread, and the share of selected points that converged
failure_mode: it measures quotation quality, not liquidity, and the two part company around expiry
# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study.
#
# Spec-hash inputs (each invalidates every backtest_hash on change):
# - execution.initial_cash
# - execution.share_type
# - execution.allocator_lookback (CS-level fallback for moment allocators)
# - per-allocator overrides in backtest.sweep.allocators
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator. Daily S&P 500 equity bars
# → 63 ≈ 3 months. ``mvo_ledoit_wolf`` carries an explicit per-allocator
# override below (126 bars ≈ 6 months) so N/K shrinkage degeneracy doesn't
# collapse it to identity-target at top_k=20.
#
# ``initial_cash`` restored to 1_000_000 (2026-05-16) — the 2026-05-15 SSOT
# migration drop to 100k caused near-zero rebalancing here (min num_trades
# = 27 over 4y daily on 500-name × top_k=20: at $5k/name budget × integer
# rounding × the S&P high-priced tail, the engine almost never trades).
# See memory/feedback_2026_05_15_equity_sizing_invalidated.md.
execution:
initial_cash: 1_000_000 # restores prior validated state
share_type: integer # US equities trade in whole shares
allocator_lookback: 63 # 3 months of daily bars (IV/RP/HRP fallback)
mapping:
class: long_only_rank_and_rebalance
position_state_space: long_only
entry_logic: rank_by_iv_signal
sizing: equal_weight
costs:
class: material
model: percentage # Loader path: bps regime via per_leg_cost_bps_range midpoint.
components: [spread, commission, market_impact]
per_leg_cost_bps_range: [3, 10]
round_trip_cost_bps: 13 # Midpoint of per_leg_cost_bps_range (6.5 bps × 2 legs).
# per_share is the commission rate for the exploratory per-share
# cost-sensitivity regime (read by Ch18 17_costs.py and the run_sweep
# planner). IBKR Pro Tiered top tier. NOT used in the headline bps
# regime — bps regime is the production cost model.
per_share: 0.0035
note: Trades equities (not options); S&P 500 names are liquid and weekly rebalancing keeps turnover moderate.
backtest:
rebalance:
# A rebalance is skipped when the per-asset weight change is below
# min_weight_change AND the resulting trade notional is below
# min_trade_value. The benchmark profile disables thresholds so that
# full-universe equal-weight (1/N per asset) rebalances at all.
default:
min_weight_change: 0.005
min_trade_value: 100.0
benchmark:
min_weight_change: 0.0
min_trade_value: 0.0
sweep:
# Iteration controls per stage. ``signal: 0`` means "all predictions";
# downstream stages take the top-N from the upstream stage's rank-1.
# Notebooks read these via get_top_n_predictions(case_study, stage).
top_n_predictions:
signal: 0 # all signal predictions (eq-weight baseline)
allocation: 10 # top-10 model configs by equal-weight baseline Sharpe
cost_sensitivity: 1 # top-1 of {signal+allocation} per label
risk_overlay: 1 # top-1 of {signal+allocation} per label
# Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
expensive_allocators_skip: false
# Equal weight is the baseline above, not an allocation-stage method.
# Score-weighted and inverse-vol take the cheap path.
# Ch16 signal-stage selection. Long-only equal-weight top-k for every
# label.
top_k_grid:
fwd_ret_5d: [5, 10, 20]
fwd_ret_10d: [5, 10, 20]
fwd_ret_risk_adj_5d: [5, 10, 20]
fwd_dir_5d: [5, 10, 20]
fwd_dir_10d: [5, 10, 20]
# Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
# Moment-based allocators (IV/RP/HRP) use the CS-level
# ``execution.allocator_lookback`` (63 bars). ``mvo_ledoit_wolf`` gets
# an explicit 126-bar override (6 months) so N/K ≥ 2.5 at top_k=20 and
# Ledoit-Wolf shrinkage doesn't collapse to identity-target.
# No max_weight cap — see memory/feedback_max_weight_caps_intentionally_absent.md
# (the prior 0.40 cap was already the loosened-from-0.20 workaround;
# restoring it would push moment allocators back toward equal-weight).
allocators:
- {name: score_weighted, method: score_weighted}
- {name: inverse_vol, method: inverse_vol}
- {name: risk_parity, method: risk_parity}
- {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf, lookback: 126}
- {name: hrp, method: hrp}
# Sizes each position by the width of its conformal prediction interval rather than by a
# moment of returns, so it is the one allocator here that reads the model's own
# uncertainty. It was absent while etfs, cme_futures and fx_pairs declared it, which made
# this case study's sweep narrower than theirs for no stated reason. 13_model_analysis
# measures the coverage those widths come from, so a run of this allocator is also the
# test of whether that calibration is good enough to size with.
- {name: conformal_weighted, method: conformal_weighted}
# Ch18 cost sensitivity (bps regime — headline). A companion per-share
# regime is run from Ch18 cost notebooks for regime comparison; for
# sp500_eoa it is exploratory only because flat-default half-spread on
# split-adjusted prices conflates split adjustment with realized
# friction.
cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
# Companion per-share half-spread grid (USD per share). Swept alongside
# cost_grid_bps by the planner. Values: 0¢, 0.5¢, 1¢, 2.5¢, 5¢, 10¢.
cost_grid_half_spread_usd: [0.0, 0.005, 0.01, 0.025, 0.05, 0.10]
# Ch19 risk overlays.
risk_controls:
position:
- {name: stop_loss_3pct, type: stop_loss, threshold: 0.03}
- {name: stop_loss_5pct, type: stop_loss, threshold: 0.05}
- {name: stop_loss_10pct, type: stop_loss, threshold: 0.10}
- {name: stop_loss_15pct, type: stop_loss, threshold: 0.15}
- {name: trailing_1pct, type: trailing_stop, threshold: 0.01}
- {name: trailing_2pct, type: trailing_stop, threshold: 0.02}
- {name: trailing_3pct, type: trailing_stop, threshold: 0.03}
- {name: trailing_5pct, type: trailing_stop, threshold: 0.05}
- {name: trailing_10pct, type: trailing_stop, threshold: 0.10}
- {name: trailing_15pct, type: trailing_stop, threshold: 0.15}
- {name: trailing_20pct, type: trailing_stop, threshold: 0.20}
- {name: time_exit_10, type: time_exit, bars: 10}
- {name: time_exit_20, type: time_exit, bars: 20}
- {name: time_exit_40, type: time_exit, bars: 40}
evaluation:
n_splits: 2
train_size: 2Y
val_size: 1Y
holdout_start: '2021-01-01'
holdout_end: '2021-12-31'
calendar: NYSE
periods_per_year: 252 # NYSE 5d/wk
labels:
primary: fwd_ret_5d
buffer: 10D
# Outcome horizons seal validation before holdout. They are separate from
# the deliberately conservative primary train-to-validation buffer above.
horizons:
fwd_ret_5d: 5D
fwd_ret_10d: 10D
fwd_ret_risk_adj_5d: 5D
fwd_dir_5d: 5D
fwd_dir_10d: 10D
variants:
- fwd_ret_10d
- fwd_ret_risk_adj_5d
- fwd_dir_5d
- fwd_dir_10d
variant_buffers:
fwd_ret_10d: 10D
fwd_ret_risk_adj_5d: 5D
fwd_dir_5d: 5D
fwd_dir_10d: 10D
# Vectorized-backtest thinning step per label: number of schedule slots
# to advance per trade so holding periods don't overlap.
# The step is ceil(horizon / cadence), and `cadence` is now per label, so it must be read
# against decision.cadence_by_label rather than against the weekly default. The two 10-day
# labels moved to the biweekly grid, which already spaces their decisions 10 sessions apart;
# leaving their step at 2 would thin an already-thinned schedule and trade them every four
# weeks under a spec that says biweekly.
rebalance_step:
fwd_ret_5d: 1 # ceil(5 / 5) on the weekly_friday_close schedule
fwd_ret_10d: 1 # ceil(10 / 10) on the biweekly schedule
fwd_ret_risk_adj_5d: 1 # ceil(5 / 5)
fwd_dir_5d: 1 # ceil(5 / 5)
fwd_dir_10d: 1 # ceil(10 / 10) on the biweekly schedule
# Continuous return that each classification label is derived from.
classification_eval_label:
fwd_dir_5d: fwd_ret_5d
fwd_dir_10d: fwd_ret_10d
modeling:
gbm:
libraries: [lightgbm]
preset: default
device: cpu
# LightGBM's own CPU default. 63 is the GPU default and was carried over with the
# device when these runs moved off the GPU, so every CPU fit was quantizing the design
# matrix into a quarter of the bins the library would have used. Coarser bins are
# faster and lose split points; the reader running this on a CPU gets what the
# documentation describes.
max_bin: 255
latent_factors:
persistent_entities: true
# The device the family is fitted on. Declared rather than left out: it sits inside
# the hashed computation (case_studies/utils/latent_factors/adapter.py, computation
# .runtime and .numerical_runtime), and with no key here case_study.py falls back to
# preferred_latent_device(), which returns cuda or cpu depending on what the machine
# running it happens to have. That makes the training identity a property of the host.
# The three neural members are fitted here; 11a_pca and 11b_ipca override it to cpu,
# because run_pca_fold and run_ipca_fold are numpy and scipy and take no device at
# all, so recording cuda for them would describe a computation that did not happen.
device: cuda
model_kwargs:
ipca:
max_iter: 10000
factor_ridge: 0.01
gamma_ridge: 0.01
sdf:
checkpoint_epochs: [256, 512, 768, 1024] # conditional-relative; publishes global 256..1280
beta_checkpoint_epochs: [256]
beta_default_checkpoint: 256
sae:
# Rows per gradient step. Declared because the alternative is not a smaller batch but
# no batching at all: SAEConfig.batch_size defaults to None and the library reads that
# as one batch holding the whole training window, about 250,000 rows here, which does
# not fit a 24 GB card. Its sibling run_cae_fold has carried this same value as a
# runner default all along; only the SAE runner never passed one.
batch_size: 10000
# What `04_model_based_features` decides, declared here for the same reason the feature
# windows are: an estimation window is part of a fitted feature's definition, so it belongs
# where the definition lives rather than inside the notebook that runs it. Every count is in
# NYSE sessions.
#
# A fitted feature is bounded by the schedule below and not by a cross-validation fold. The
# parameters behind a value are estimated from sessions strictly before it, refreshed on the
# cadence given, and the same value comes out whichever fold later selects the row - so the
# artifact carries no fold column.
#
# The panel runs 2017-01-03 to 2021-12-31, 1,259 sessions over 624 securities carrying an
# option surface. Fold 0's training window opens 2018-01-04, so only about 250 sessions of
# history precede the first fold, and that run-up is what a burn-in has to be paid out of.
model_based:
gjr_garch:
# Sessions of a security's own adjusted returns before its volatility model is fitted.
# One year, and it is bounded above by the panel rather than chosen freely: 252 is very
# nearly the 250 sessions that precede fold 0, so it comes out of the run-up instead of
# out of training data. Measured against the alternative on the 624-security roster, a
# 504 burn-in takes emitted coverage from 76.5% of the panel to 54.7%, takes the
# securities that emit nothing at all from 57 to 103, and eats 159,463 of fold 0's
# training sessions across 592 securities rather than 22,130 across 121.
#
# It is not free. A GJR-GARCH fitted on 252 observations returns a degenerate parameter
# vector - alpha + gamma < 0, so a larger down-shock lowers next session's variance -
# for 19.0% of securities on the first block, against 1.8% of fits averaged over the
# whole walk as the expanding window grows. Section C reports both.
burnin: 252
# How often the parameters are re-estimated. A month. Each estimate is a
# quasi-maximum-likelihood fit over the whole expanding history, and the leverage
# parameter it is estimated for is not a fast-moving quantity.
refit_every: 21
# The specification itself, for the same reason the schedule is here: what was fitted is
# part of what the feature means, and a reader comparing this chapter's model-based
# features against another's needs to see where they differ without reading two notebooks.
#
# `o: 1` is this case study's declared deviation from the shared GARCH(1,1)-Normal default,
# and it is what makes this a GJR rather than a plain GARCH. The justification is a property
# of the data, not a preference: a negative return raises next session's variance by more
# than a positive return of the same size, and on single-name equity underlying an option
# book that asymmetry is what the option prices are quoted around. Section C of the notebook
# carries the measurement. `sp500_options` declares the same deviation for the same reason.
mean: Constant
vol: GARCH
p: 1
o: 1
q: 1
dist: Normal
causal:
treatment: ivrv_spread
# Bars the treatment's own construction window spans, which is what the placebo block has
# to cover: permuting ivrv_spread in blocks shorter than this destroys the serial
# dependence the refutation exists to preserve, and the resulting p-value reads like a
# refutation without being one. Declared here rather than inferred, because guessing which
# element of a window list a column was built from puts a wrong number behind a right-looking
# one. Derived from the construction, not chosen:
#
# `iv_30_atm - rv_{realized_vol[0]}` in 03_financial_features, and features.windows
# declares realized_vol as [20, 63]. The implied leg is a quote and rolls nothing; the
# realized leg is what makes the spread autocorrelated, over its own 20 sessions.
treatment_window: 20
confounders: [rv_20, mom_21d, skew_rr_30_25d]
method: walk_forward_dml
```Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.