Zum Inhalt springen
Alle Bibliotheksdokumente

Aktien- und Optionssignale: Backtest-Konfiguration

Code Machine Learning for Trading

Zusammenfassung

Dieses Setup-Dokument definiert eine wöchentliche Forschungs-Pipeline für 500-Aktien, die Optionsmarkt-Features mit preisbasierten Signalen kombiniert. Es legt das zulässige Universum, Entscheidungs- und Ausführungszeitpunkte sowie unterschiedliche Rebalancing-Takte für Labels mit verschiedenen Prognosehorizonten fest. Zu den Feature-Familien gehören Niveau und Dynamik der impliziten Volatilität, Skew und Laufzeitenstruktur, Varianzrisikoprämie, realisierte Volatilität, Aktienmomentum, querschnittliche Ränge und Oberflächenqualität. Das Dokument hält zudem Rückblickfenster, Informationsverzögerungen, Darstellungen, Hypothesen und Fehlermodi fest, etwa veraltete Quotes oder Unterschiede zwischen impliziter zukünftiger und realisierter vergangener Volatilität. Außerdem legt es eine Long-only-Positionierung nach Rang mit gleicher Gewichtung, Ausführungsannahmen, Transaktionskostenkomponenten, Rebalancing-Schwellen sowie abgestufte Tests der Allokation und Risiko-Overlays fest. Ein eigener Abschnitt konfiguriert die Walk-Forward-Schätzung eines GJR-GARCH-Volatilitätsmodells einschließlich Warm-up und Neu-Schätzungsplan sowie eine Kausalanalyse mit Walk-Forward Double Machine Learning und blockbewusstem Placebo-Design. Diese Festlegungen machen die vorgesehenen Daten, Features und Analyseentscheidungen nachvollziehbar. Dies ist eine Forschungsspezifikation und kein Beleg dafür, dass ein Signal oder eine Strategie profitabel ist. Annahmen, Kostenschätzungen, Feature-Konstruktion und Modellierungsentscheidungen definieren das Experiment und müssten empirisch validiert werden.

Kernaussagen

  • Das Setup kombiniert Optionsoberflächen-Features mit Aktienmomentum und Maßen der realisierten Volatilität.
  • Feature-Fenster und Verzögerungen legen fest, wann Eingaben beobachtbar werden und wie ihre Werte konstruiert werden.
  • Die Strategie verwendet Long-only-Ranking und gleiche Gewichtung mit festgelegten Ausführungskosten und Rebalancing-Regeln.
  • Für Prognose-Labels mit längeren Horizonten gilt ein langsamerer Rebalancing-Takt, um überlappende Positionen zu begrenzen.
  • Die Konfiguration legt GJR-GARCH-Schätzung und eine kausale Walk-Forward-Analyse fest, belegt aber für sich genommen keine Profitabilität des Signals.

Schlagwörter

Volltext
# setup.yaml


```yaml
strategy_id: sp500_equity_option_analytics
setup_version: v1

universe:
  n_assets: 633
  eligibility_rule: sp500_with_options

decision:
  cadence: weekly_friday_close
  snapshot: friday_16:00_et
  execution_delay: monday_open
  iv_feature_lag: 1_day
  # Per-label decision cadence; `cadence` above applies to every label not named here. The five
  # 5-day labels trade the weekly grid, which is what they forecast. The two 10-day labels
  # resolve after two weeks, so on the weekly grid a new position is opened while the previous
  # one is still inside its own horizon - overlapping the very quantity being measured. They
  # rebalance every other week instead.
  cadence_by_label:
    fwd_ret_10d: biweekly
    fwd_dir_10d: biweekly

# The feature-specification register for 03_financial_features, and every window
# it computes over. It lives here rather than in the notebook for the same reason
# the label name and the holdout boundary do: it is the statement of what this
# case study's feature set is, and a statement only the notebook holds cannot be
# read by a test, by a later stage, or by anyone asking what changed.
#
# ``lookback`` is counted in daily bars back from the decision timestamp and is
# the floor the warmup audit holds each column to. ``lag`` is the delay with
# which the input becomes knowable: every option-derived family carries one
# session because the surface summary is stamped at the close it summarizes and
# is not read until the next session's decision, and every equity-price family
# carries zero because the close is the decision snapshot itself.
features:
  # Surface contract selection, fixed by data/equities/market/sp500/materialize_options.py.
  # Repeated here only so section B can state the observability of what it loads.
  surface:
    dte_buckets: {'7d': [5, 10], '30d': [25, 35], '90d': [80, 110]}
    delta_targets: {atm: 0.50, '25d': 0.25, '10d': 0.10}
  windows:
    iv_zscore: [63, 252]
    iv_percentile: 252
    iv_momentum: [5, 21]
    skew_zscore: 63
    term_zscore: 63
    vrp_zscore: 63
    realized_vol: [20, 63]
    garman_klass: 21
    vol_of_vol: 21
    realized_skew: 21
    momentum: [5, 21, 63, 126, 252]
    skip_recent: 21          # skip-month momentum runs t-252 to t-21
    skip_start: 252
    risk_adjusted: 63        # the momentum horizon its own volatility scales
    # Annualized volatility floor in the risk-adjusted momentum denominator. A share
    # that realized almost nothing over a quarter would otherwise carry an unbounded
    # ratio, and one such row dominates the within-date percentile taken from it.
    risk_adjusted_vol_floor: 0.01
    iv_forward_fill: 5       # sessions a lagged surface value may be carried forward
  # Source column -> the name its within-date percentile is written under. The
  # names predate this register and later stages select by them, so the mapping
  # is explicit rather than a suffix rule.
  ranked:
    iv_30_atm: iv_rank
    skew_rr_30_25d: skew_rank
    ivrv_spread: vrp_rank
    mom_21d: mom_21d_rank
    mom_63d: mom_63d_rank
    rv_20: rv_rank
    iv_mom_21d: iv_mom_rank
    d_iv_30_atm: d_iv_rank
  families:
    - name: cross-sectional rank
      pattern: '*_rank'
      role: signal
      hypothesis: A long-short book acts on relative standing, not on the level of a quantity
      inputs: the eight level and dynamics columns named under features.ranked
      lookback: 252
      lag: 1
      frame: cross section within the decision date
      representation: percentile within the date, in (0, 100)
      failure_mode: a thin cross-section makes the percentile a coarse ordering of few names
    - name: implied volatility level
      pattern: iv_30_atm|iv_7_atm|iv_90_atm|iv_30_put_25d|iv_30_call_25d
      role: signal
      hypothesis: What the option market charges for a name's coming month is priced against what its shares then do
      inputs: daily IV surface summary, delta-selected within fixed DTE buckets
      lookback: 1
      lag: 1
      frame: time series
      representation: annualized implied volatility, in variance points
      failure_mode: a name whose surface is quoted thinly carries a level set by one stale contract
    - name: implied volatility dynamics
      pattern: d_iv_30_atm|iv_mom_*|iv_30_atm_z_*|iv_30_atm_pct_*
      role: signal
      hypothesis: Where implied volatility sits against its own recent history says more than its level
      inputs: at-the-money 30-day implied volatility
      lookback: 252
      lag: 1
      frame: time series
      representation: first difference, multi-session change, trailing z-score and trailing percentile
      failure_mode: the z-score is unbounded as a name's own dispersion approaches zero
    - name: skew and term structure
      pattern: skew_*|term_*|d_skew_*|d_term_*
      role: signal
      hypothesis: The shape of the surface prices asymmetry and horizon that its level cannot express
      inputs: 25-delta put and call IV, and the 7-, 30- and 90-day at-the-money points
      lookback: 63
      lag: 1
      frame: term structure
      representation: differences and ratios between surface points, and their trailing z-scores
      failure_mode: the far segment reads a bucket whose contracts were interpolated rather than quoted
    - name: variance risk premium
      pattern: ivrv_spread|vrp_z_*
      role: signal
      hypothesis: A name whose implied volatility stands above what it goes on to realize is richly priced
      inputs: at-the-money implied volatility and trailing realized volatility
      lookback: 63
      lag: 1
      frame: time series
      representation: spread in volatility points and its trailing z-score
      failure_mode: it compares a forward-looking month against a backward-looking one, which differ around an event
    - name: realized volatility
      pattern: rv_20|rv_63|gk_vol_*|vol_of_vol_*|realized_skew_*
      role: state
      hypothesis: Dispersion and the asymmetry of the realized path set what a ranking can earn
      inputs: split- and dividend-adjusted OHLC bars
      lookback: 63
      lag: 0
      frame: time series
      representation: annualized standard deviation, a range-based estimator, and third moments
      failure_mode: close-to-close misses the overnight gap the range-based estimator is here to catch
    - name: equity momentum
      pattern: mom_*
      role: signal
      hypothesis: Option-derived information has to earn its place beside the price signal it would replace
      inputs: split- and dividend-adjusted close
      lookback: 252
      lag: 0
      frame: time series
      representation: simple return at five horizons, its skip-month form, and its volatility-scaled twin
      failure_mode: reverses over the most recent month, which skip-month momentum drops
    - name: surface quality
      pattern: spread_atm_*|qc_converged_share
      role: state
      hypothesis: A surface point solved from wide or unconverged quotes should not be read like one that was not
      inputs: bid-ask spread and solver convergence flags of the selected contracts
      lookback: 1
      lag: 1
      frame: time series
      representation: relative spread, and the share of selected points that converged
      failure_mode: it measures quotation quality, not liquidity, and the two part company around expiry

# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study.
#
# Spec-hash inputs (each invalidates every backtest_hash on change):
#   - execution.initial_cash
#   - execution.share_type
#   - execution.allocator_lookback (CS-level fallback for moment allocators)
#   - per-allocator overrides in backtest.sweep.allocators
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator. Daily S&P 500 equity bars
# → 63 ≈ 3 months. ``mvo_ledoit_wolf`` carries an explicit per-allocator
# override below (126 bars ≈ 6 months) so N/K shrinkage degeneracy doesn't
# collapse it to identity-target at top_k=20.
#
# ``initial_cash`` restored to 1_000_000 (2026-05-16) — the 2026-05-15 SSOT
# migration drop to 100k caused near-zero rebalancing here (min num_trades
# = 27 over 4y daily on 500-name × top_k=20: at $5k/name budget × integer
# rounding × the S&P high-priced tail, the engine almost never trades).
# See memory/feedback_2026_05_15_equity_sizing_invalidated.md.
execution:
  initial_cash: 1_000_000        # restores prior validated state
  share_type: integer            # US equities trade in whole shares
  allocator_lookback: 63         # 3 months of daily bars (IV/RP/HRP fallback)

mapping:
  class: long_only_rank_and_rebalance
  position_state_space: long_only
  entry_logic: rank_by_iv_signal
  sizing: equal_weight

costs:
  class: material
  model: percentage          # Loader path: bps regime via per_leg_cost_bps_range midpoint.
  components: [spread, commission, market_impact]
  per_leg_cost_bps_range: [3, 10]
  round_trip_cost_bps: 13    # Midpoint of per_leg_cost_bps_range (6.5 bps × 2 legs).
  # per_share is the commission rate for the exploratory per-share
  # cost-sensitivity regime (read by Ch18 17_costs.py and the run_sweep
  # planner). IBKR Pro Tiered top tier. NOT used in the headline bps
  # regime — bps regime is the production cost model.
  per_share: 0.0035
  note: Trades equities (not options); S&P 500 names are liquid and weekly rebalancing keeps turnover moderate.

backtest:
  rebalance:
    # A rebalance is skipped when the per-asset weight change is below
    # min_weight_change AND the resulting trade notional is below
    # min_trade_value. The benchmark profile disables thresholds so that
    # full-universe equal-weight (1/N per asset) rebalances at all.
    default:
      min_weight_change: 0.005
      min_trade_value: 100.0
    benchmark:
      min_weight_change: 0.0
      min_trade_value: 0.0
  sweep:
    # Iteration controls per stage. ``signal: 0`` means "all predictions";
    # downstream stages take the top-N from the upstream stage's rank-1.
    # Notebooks read these via get_top_n_predictions(case_study, stage).
    top_n_predictions:
      signal: 0                 # all signal predictions (eq-weight baseline)
      allocation: 10            # top-10 model configs by equal-weight baseline Sharpe
      cost_sensitivity: 1       # top-1 of {signal+allocation} per label
      risk_overlay: 1           # top-1 of {signal+allocation} per label
    # Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
    expensive_allocators_skip: false
    # Equal weight is the baseline above, not an allocation-stage method.
    # Score-weighted and inverse-vol take the cheap path.
    # Ch16 signal-stage selection. Long-only equal-weight top-k for every
    # label.
    top_k_grid:
      fwd_ret_5d:           [5, 10, 20]
      fwd_ret_10d:          [5, 10, 20]
      fwd_ret_risk_adj_5d:  [5, 10, 20]
      fwd_dir_5d:           [5, 10, 20]
      fwd_dir_10d:          [5, 10, 20]
    # Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
    # Moment-based allocators (IV/RP/HRP) use the CS-level
    # ``execution.allocator_lookback`` (63 bars). ``mvo_ledoit_wolf`` gets
    # an explicit 126-bar override (6 months) so N/K ≥ 2.5 at top_k=20 and
    # Ledoit-Wolf shrinkage doesn't collapse to identity-target.
    # No max_weight cap — see memory/feedback_max_weight_caps_intentionally_absent.md
    # (the prior 0.40 cap was already the loosened-from-0.20 workaround;
    # restoring it would push moment allocators back toward equal-weight).
    allocators:
      - {name: score_weighted,  method: score_weighted}
      - {name: inverse_vol,     method: inverse_vol}
      - {name: risk_parity,     method: risk_parity}
      - {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf, lookback: 126}
      - {name: hrp,             method: hrp}
      # Sizes each position by the width of its conformal prediction interval rather than by a
      # moment of returns, so it is the one allocator here that reads the model's own
      # uncertainty. It was absent while etfs, cme_futures and fx_pairs declared it, which made
      # this case study's sweep narrower than theirs for no stated reason. 13_model_analysis
      # measures the coverage those widths come from, so a run of this allocator is also the
      # test of whether that calibration is good enough to size with.
      - {name: conformal_weighted, method: conformal_weighted}
    # Ch18 cost sensitivity (bps regime — headline). A companion per-share
    # regime is run from Ch18 cost notebooks for regime comparison; for
    # sp500_eoa it is exploratory only because flat-default half-spread on
    # split-adjusted prices conflates split adjustment with realized
    # friction.
    cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
    # Companion per-share half-spread grid (USD per share). Swept alongside
    # cost_grid_bps by the planner. Values: 0¢, 0.5¢, 1¢, 2.5¢, 5¢, 10¢.
    cost_grid_half_spread_usd: [0.0, 0.005, 0.01, 0.025, 0.05, 0.10]
    # Ch19 risk overlays.
    risk_controls:
      position:
        - {name: stop_loss_3pct,  type: stop_loss,     threshold: 0.03}
        - {name: stop_loss_5pct,  type: stop_loss,     threshold: 0.05}
        - {name: stop_loss_10pct, type: stop_loss,     threshold: 0.10}
        - {name: stop_loss_15pct, type: stop_loss,     threshold: 0.15}
        - {name: trailing_1pct,   type: trailing_stop, threshold: 0.01}
        - {name: trailing_2pct,   type: trailing_stop, threshold: 0.02}
        - {name: trailing_3pct,   type: trailing_stop, threshold: 0.03}
        - {name: trailing_5pct,   type: trailing_stop, threshold: 0.05}
        - {name: trailing_10pct,  type: trailing_stop, threshold: 0.10}
        - {name: trailing_15pct,  type: trailing_stop, threshold: 0.15}
        - {name: trailing_20pct,  type: trailing_stop, threshold: 0.20}
        - {name: time_exit_10,    type: time_exit,     bars: 10}
        - {name: time_exit_20,    type: time_exit,     bars: 20}
        - {name: time_exit_40,    type: time_exit,     bars: 40}

evaluation:
  n_splits: 2
  train_size: 2Y
  val_size: 1Y
  holdout_start: '2021-01-01'
  holdout_end: '2021-12-31'
  calendar: NYSE
  periods_per_year: 252  # NYSE 5d/wk

labels:
  primary: fwd_ret_5d
  buffer: 10D
  # Outcome horizons seal validation before holdout. They are separate from
  # the deliberately conservative primary train-to-validation buffer above.
  horizons:
    fwd_ret_5d: 5D
    fwd_ret_10d: 10D
    fwd_ret_risk_adj_5d: 5D
    fwd_dir_5d: 5D
    fwd_dir_10d: 10D
  variants:
    - fwd_ret_10d
    - fwd_ret_risk_adj_5d
    - fwd_dir_5d
    - fwd_dir_10d
  variant_buffers:
    fwd_ret_10d: 10D
    fwd_ret_risk_adj_5d: 5D
    fwd_dir_5d: 5D
    fwd_dir_10d: 10D
  # Vectorized-backtest thinning step per label: number of schedule slots
  # to advance per trade so holding periods don't overlap.
  # The step is ceil(horizon / cadence), and `cadence` is now per label, so it must be read
  # against decision.cadence_by_label rather than against the weekly default. The two 10-day
  # labels moved to the biweekly grid, which already spaces their decisions 10 sessions apart;
  # leaving their step at 2 would thin an already-thinned schedule and trade them every four
  # weeks under a spec that says biweekly.
  rebalance_step:
    fwd_ret_5d: 1            # ceil(5 / 5) on the weekly_friday_close schedule
    fwd_ret_10d: 1           # ceil(10 / 10) on the biweekly schedule
    fwd_ret_risk_adj_5d: 1   # ceil(5 / 5)
    fwd_dir_5d: 1            # ceil(5 / 5)
    fwd_dir_10d: 1           # ceil(10 / 10) on the biweekly schedule
  # Continuous return that each classification label is derived from.
  classification_eval_label:
    fwd_dir_5d: fwd_ret_5d
    fwd_dir_10d: fwd_ret_10d

modeling:
  gbm:
    libraries: [lightgbm]
    preset: default
    device: cpu
    # LightGBM's own CPU default. 63 is the GPU default and was carried over with the
    # device when these runs moved off the GPU, so every CPU fit was quantizing the design
    # matrix into a quarter of the bins the library would have used. Coarser bins are
    # faster and lose split points; the reader running this on a CPU gets what the
    # documentation describes.
    max_bin: 255
  latent_factors:
    persistent_entities: true
    # The device the family is fitted on. Declared rather than left out: it sits inside
    # the hashed computation (case_studies/utils/latent_factors/adapter.py, computation
    # .runtime and .numerical_runtime), and with no key here case_study.py falls back to
    # preferred_latent_device(), which returns cuda or cpu depending on what the machine
    # running it happens to have. That makes the training identity a property of the host.
    # The three neural members are fitted here; 11a_pca and 11b_ipca override it to cpu,
    # because run_pca_fold and run_ipca_fold are numpy and scipy and take no device at
    # all, so recording cuda for them would describe a computation that did not happen.
    device: cuda
    model_kwargs:
      ipca:
        max_iter: 10000
        factor_ridge: 0.01
        gamma_ridge: 0.01
      sdf:
        checkpoint_epochs: [256, 512, 768, 1024]  # conditional-relative; publishes global 256..1280
        beta_checkpoint_epochs: [256]
        beta_default_checkpoint: 256
      sae:
        # Rows per gradient step. Declared because the alternative is not a smaller batch but
        # no batching at all: SAEConfig.batch_size defaults to None and the library reads that
        # as one batch holding the whole training window, about 250,000 rows here, which does
        # not fit a 24 GB card. Its sibling run_cae_fold has carried this same value as a
        # runner default all along; only the SAE runner never passed one.
        batch_size: 10000

# What `04_model_based_features` decides, declared here for the same reason the feature
# windows are: an estimation window is part of a fitted feature's definition, so it belongs
# where the definition lives rather than inside the notebook that runs it. Every count is in
# NYSE sessions.
#
# A fitted feature is bounded by the schedule below and not by a cross-validation fold. The
# parameters behind a value are estimated from sessions strictly before it, refreshed on the
# cadence given, and the same value comes out whichever fold later selects the row - so the
# artifact carries no fold column.
#
# The panel runs 2017-01-03 to 2021-12-31, 1,259 sessions over 624 securities carrying an
# option surface. Fold 0's training window opens 2018-01-04, so only about 250 sessions of
# history precede the first fold, and that run-up is what a burn-in has to be paid out of.
model_based:
  gjr_garch:
    # Sessions of a security's own adjusted returns before its volatility model is fitted.
    # One year, and it is bounded above by the panel rather than chosen freely: 252 is very
    # nearly the 250 sessions that precede fold 0, so it comes out of the run-up instead of
    # out of training data. Measured against the alternative on the 624-security roster, a
    # 504 burn-in takes emitted coverage from 76.5% of the panel to 54.7%, takes the
    # securities that emit nothing at all from 57 to 103, and eats 159,463 of fold 0's
    # training sessions across 592 securities rather than 22,130 across 121.
    #
    # It is not free. A GJR-GARCH fitted on 252 observations returns a degenerate parameter
    # vector - alpha + gamma < 0, so a larger down-shock lowers next session's variance -
    # for 19.0% of securities on the first block, against 1.8% of fits averaged over the
    # whole walk as the expanding window grows. Section C reports both.
    burnin: 252
    # How often the parameters are re-estimated. A month. Each estimate is a
    # quasi-maximum-likelihood fit over the whole expanding history, and the leverage
    # parameter it is estimated for is not a fast-moving quantity.
    refit_every: 21
    # The specification itself, for the same reason the schedule is here: what was fitted is
    # part of what the feature means, and a reader comparing this chapter's model-based
    # features against another's needs to see where they differ without reading two notebooks.
    #
    # `o: 1` is this case study's declared deviation from the shared GARCH(1,1)-Normal default,
    # and it is what makes this a GJR rather than a plain GARCH. The justification is a property
    # of the data, not a preference: a negative return raises next session's variance by more
    # than a positive return of the same size, and on single-name equity underlying an option
    # book that asymmetry is what the option prices are quoted around. Section C of the notebook
    # carries the measurement. `sp500_options` declares the same deviation for the same reason.
    mean: Constant
    vol: GARCH
    p: 1
    o: 1
    q: 1
    dist: Normal

causal:
  treatment: ivrv_spread
  # Bars the treatment's own construction window spans, which is what the placebo block has
  # to cover: permuting ivrv_spread in blocks shorter than this destroys the serial
  # dependence the refutation exists to preserve, and the resulting p-value reads like a
  # refutation without being one. Declared here rather than inferred, because guessing which
  # element of a window list a column was built from puts a wrong number behind a right-looking
  # one. Derived from the construction, not chosen:
  #
  # `iv_30_atm - rv_{realized_vol[0]}` in 03_financial_features, and features.windows
  # declares realized_vol as [20, 63]. The implied leg is a quote and rolls nothing; the
  # realized leg is what makes the spread autocorrelated, over its own 20 sessions.
  treatment_window: 20
  confounders: [rv_20, mom_21d, skew_rr_30_25d]
  method: walk_forward_dml

```

Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT

Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.