Passer au contenu
Tous les documents de la bibliothèque

Configuration de stratégie de momentum et carry sur paires FX

Code Machine Learning for Trading

Résumé

Cette configuration décrit une stratégie quotidienne long-short sur 20 paires FX. Elle définit une heure de décision à la clôture de New York, une exécution à la barre suivante, une pondération égale et des signaux de classement fondés sur le momentum ou le carry. Elle précise aussi les coûts de spread et de points de swap, les seuils de rééquilibrage, les tailles de portefeuille candidates, d’autres méthodes d’allocation, des scénarios de coûts de transaction ainsi que des sorties par stop, stop suiveur et durée.

Le registre des caractéristiques réunit des mesures de momentum, de momentum ajusté du risque, de retour à la moyenne, de volatilité, d’amplitude, de tendance, de facteur dollar et de positionnement transversal. Il consigne l’hypothèse de chaque famille et un mode d’échec possible, comme un retournement du momentum ou une tendance persistante qui contrecarrerait les signaux de retour à la moyenne. La section causale définit un traitement de momentum excluant la période récente et explique pourquoi les permutations placebo exigent des blocs assez longs pour préserver la dépendance due au chevauchement des fenêtres. L’extrait est incomplet : il ne présente pas tous les détails de modélisation et d’évaluation, ni les performances de la stratégie ; la configuration seule ne prouve pas qu’un signal ou un allocateur soit rentable.

Idées clés

  • La stratégie classe FX paires pour des positions long-short à partir du momentum ou du carry et attribue initialement un poids égal aux positions.
  • Les caractéristiques couvrent la tendance, le retour à la moyenne, la volatilité, l’amplitude, l’exposition au dollar et les classements transversaux.
  • La configuration prend en compte les spreads et les points de swap, et examine plusieurs choix d’allocation et de contrôle du risque.
  • Les valeurs de momentum excluant la période récente se chevauchent fortement ; des blocs placebo courts peuvent sous-estimer la dépendance.
  • La configuration décrit une expérience, mais ne fournit aucun élément de preuve sur les performances.

Étiquettes

Texte intégral
# setup.yaml


```yaml
strategy_id: fx_pairs
setup_version: v1

universe:
  symbols:
    - AUD_JPY
    - AUD_NZD
    - AUD_USD
    - CAD_JPY
    - CHF_JPY
    - EUR_AUD
    - EUR_CAD
    - EUR_CHF
    - EUR_GBP
    - EUR_JPY
    - EUR_USD
    - GBP_AUD
    - GBP_CHF
    - GBP_JPY
    - GBP_USD
    - NZD_JPY
    - NZD_USD
    - USD_CAD
    - USD_CHF
    - USD_JPY
  n_assets: 20

decision:
  cadence: daily_ny_close
  snapshot: ny_5pm_close
  execution_delay: next_bar_open
  # The venue calendar that implements the 5PM rollover, and so the one that assigns a
  # four-hour bar to the session it was printed in. Distinct from `evaluation.calendar`,
  # which is the calendar the cross-validation splitter counts train and validation
  # windows on. `02_labels` and `04_model_based_features` aggregate on this same value.
  session_calendar: CME_FX

# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study (cash + share_type are spec-hash inputs).
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator (inverse_vol, risk_parity, hrp,
# mvo_ledoit_wolf). FX pairs has daily OANDA bars → 63 ≈ 3 months. Keeps
# allocators comparable — MVO is not granted a longer covariance window than
# IV by default.
execution:
  initial_cash: 100_000          # IBKR retail cohort default
  share_type: integer            # FX traded in units of base currency, integer share semantics
  allocator_lookback: 63         # 3 months of daily OANDA bars

mapping:
  class: long_short_rank_rebalance
  position_state_space: long_short
  entry_logic: rank_by_momentum_or_carry
  sizing: equal_weight

costs:
  class: material
  components: [spread, swap_points]
  spread_bps:
    major_pairs: [1, 3]
    cross_pairs: [3, 8]

backtest:
  rebalance:
    # A rebalance is skipped when the per-asset weight change is below
    # min_weight_change AND the resulting trade notional is below
    # min_trade_value. The benchmark profile disables thresholds so that
    # full-universe equal-weight (1/N per asset) rebalances at all.
    default:
      min_weight_change: 0.005
      min_trade_value: 100.0
    benchmark:
      min_weight_change: 0.0
      min_trade_value: 0.0
  sweep:
    # Iteration controls per stage. ``signal: 0`` means "all predictions";
    # downstream stages take the top-N from the upstream stage's rank-1.
    # Notebooks read these via get_top_n_predictions(case_study, stage).
    top_n_predictions:
      signal: 0                 # all signal predictions (eq-weight baseline)
      allocation: 10            # top-10 model configs by equal-weight baseline Sharpe
      cost_sensitivity: 1       # top-1 of {signal+allocation} per label
      risk_overlay: 1           # top-1 of {signal+allocation} per label
    # Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
    expensive_allocators_skip: false
    # All alternative allocators take the cheap path except the two moment-based
    # methods covered by expensive_allocators_skip. Equal weight is the baseline
    # and is intentionally absent here so the allocation stage cannot double-count it.
    # Ch16 baseline selection. Long-short equal-weight top-k on all
    # three label horizons; long-short is enforced at the backtest config
    # level (allow_short_selling=true).
    top_k_grid:
      fwd_ret_1d:  [5, 10]
      fwd_ret_5d:  [5, 10]
      fwd_ret_21d: [5, 10]
    # Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
    # Moment-based allocators (IV/RP/HRP/MVO_LW) all use the CS-level
    # ``execution.allocator_lookback`` — no per-allocator vol_window / lookback
    # fields. No max-weight cap: each allocator expresses its natural
    # concentration profile so the cross-method comparison is honest.
    allocators:
      - {name: score_weighted,  method: score_weighted}
      - {name: inverse_vol,     method: inverse_vol}
      - {name: risk_parity,     method: risk_parity}
      - {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf}
      - {name: hrp,             method: hrp}
      - {name: conformal_weighted, method: conformal_weighted}
    # Ch18 cost sensitivity (bps; spread is the dominant cost component,
    # captured here as combined commission + slippage in bps).
    cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
    # Ch19 risk overlays.
    risk_controls:
      position:
        - {name: stop_loss_3pct,  type: stop_loss,     threshold: 0.03}
        - {name: stop_loss_5pct,  type: stop_loss,     threshold: 0.05}
        - {name: stop_loss_10pct, type: stop_loss,     threshold: 0.10}
        - {name: stop_loss_15pct, type: stop_loss,     threshold: 0.15}
        - {name: trailing_1pct,   type: trailing_stop, threshold: 0.01}
        - {name: trailing_2pct,   type: trailing_stop, threshold: 0.02}
        - {name: trailing_3pct,   type: trailing_stop, threshold: 0.03}
        - {name: trailing_5pct,   type: trailing_stop, threshold: 0.05}
        - {name: trailing_10pct,  type: trailing_stop, threshold: 0.10}
        - {name: trailing_15pct,  type: trailing_stop, threshold: 0.15}
        - {name: trailing_20pct,  type: trailing_stop, threshold: 0.20}
        - {name: time_exit_10,    type: time_exit,     bars: 10}
        - {name: time_exit_20,    type: time_exit,     bars: 20}
        - {name: time_exit_40,    type: time_exit,     bars: 40}

# The feature-specification register. ``03_financial_features`` binds every window
# from here and renders ``families`` as the register table; nothing in that notebook
# retypes a number this block declares. ``lookback`` is the longest chain of trailing
# bars a family reads, counted from the decision timestamp, and it is the floor the
# warmup audit holds each column to. ``lag`` is zero throughout: every input is an
# OANDA spot bar, on the tape at the NY 5PM snapshot the decision is taken at.
features:
  windows:
    momentum: [5, 10, 21, 42, 63, 126, 252]
    close_to_close_volatility: [21, 63]
    garman_klass: [21, 63, 126, 252]
    zscore: 252              # trailing window the multi-horizon returns are standardized against
    zscore_horizons: [21, 63, 126]
    channel: [21, 63, 126]
    skip_recent: 21
    drawdown: [63]
    rsi: [14, 63]
    moving_average: [21, 63, 252]
    range: [21, 63]
    bollinger: 21
    dollar: [21, 63]         # horizons the dollar proxy is accumulated over
    dollar_exposure: 252     # window each pair's correlation with the proxy is taken over
  ranked:
    - ret_21d
    - ret_63d
    - ret_126d
    - vol_gk_21d
    - vol_gk_63d
    - sharpe_21d
    - sharpe_63d
    - sharpe_126d
  # Rows are kept from the bar this column can first hold a value, which is the
  # longest chain in the matrix and the point past which every family is dense.
  null_policy_carrier: zscore_126d
  persistence_horizon: 63    # bars F6 reads the feature autocorrelation out to
  # Rank correlation at which the redundancy tree is cut: two columns agreeing this
  # closely on their ordering are one column as far as a linear model is concerned.
  redundancy_cut: 0.7
  families:
    - name: momentum
      pattern: ret_*d|mom_skip_recent|accel_*
      role: signal
      hypothesis: A pair that has moved keeps moving over the following session
      inputs: daily close
      lookback: 252
      lag: 0
      frame: time series
      representation: simple return, and differences between horizons
      failure_mode: reverses at the shortest horizons
    - name: risk-adjusted momentum
      pattern: sharpe_*d
      role: signal
      hypothesis: Trend earned with less dispersion repeats more reliably
      inputs: daily log return
      lookback: 252
      lag: 0
      frame: time series
      representation: mean log return over its own dispersion, annualized
      failure_mode: unbounded when dispersion approaches zero
    - name: mean reversion
      pattern: zscore_*|channel_pos_*|bollinger_pctb_*
      role: signal
      hypothesis: A pair stretched against its own recent range comes back
      inputs: daily close
      lookback: 378
      lag: 0
      frame: time series
      representation: trailing z-score, position in the trailing range
      failure_mode: a trending pair sits at the edge of its channel for months
    - name: volatility
      pattern: vol_gk_*|vol_cc_*|vol_ratio_*
      role: state
      hypothesis: Dispersion sets how far apart the cross-section can spread
      inputs: daily OHLC
      lookback: 252
      lag: 0
      frame: time series
      representation: Garman-Klass and close-to-close deviation, and ratios of windows
      failure_mode: lags a shock by roughly half its window
    - name: range and drawdown
      pattern: max_dd_*|avg_range_*
      role: state
      hypothesis: How far a pair sits below its peak conditions what a signal is worth
      inputs: daily OHLC
      lookback: 63
      lag: 0
      frame: time series
      representation: share below the trailing peak, mean normalized daily range
      failure_mode: says nothing about direction
    - name: oscillator and trend
      pattern: rsi_*|price_to_ma_*
      role: signal
      hypothesis: Price against its own recent path separates trend from exhaustion
      inputs: daily close
      lookback: 252
      lag: 0
      frame: time series
      representation: Wilder-smoothed oscillator, ratio to a moving average
      failure_mode: saturates in a sustained trend
    - name: dollar factor
      pattern: usd_factor_*|usd_corr_*
      role: state
      hypothesis: A broad dollar move is common to seven pairs and is not pair-specific
      inputs: daily returns of the seven USD pairs
      lookback: 315
      lag: 0
      frame: cross-asset
      representation: signed average return, and each pair's rolling correlation with it
      failure_mode: a signed average is not an estimated factor
    - name: cross-sectional position
      pattern: rank_*
      role: signal
      hypothesis: Only relative standing is tradable in a long-short rank strategy
      inputs: ret_21d, ret_63d, ret_126d, vol_gk_21d, vol_gk_63d, sharpe_21d, sharpe_63d, sharpe_126d
      lookback: 252
      lag: 0
      frame: cross-section
      representation: percentile within the decision date
      failure_mode: discards the level the state families carry instead

# What `04_model_based_features` decides, declared here for the same reason the
# feature windows above are: an estimation window is part of a fitted feature's
# definition, so it belongs where the definition lives rather than inside the
# notebook that runs it. Every count is in trading sessions on the FX calendar.
#
# A fitted feature is bounded by the schedule below and not by a cross-validation
# fold. The parameters behind a value are estimated from sessions strictly before
# it, refreshed on the cadence given, and the same value comes out whichever fold
# later selects the row - so the artifact carries no fold column.
#
# The panel runs 2011-01-03 to 2025-12-31, 3,874 sessions, and every pair quotes
# on all of them. Only one session precedes the oldest fold's training start, so
# each burn-in below is paid out of that fold's training window rather than ahead
# of it. The sessions it costs are named against each entry; no validation window
# is touched, because the earliest one opens 2016-01-05.
model_based:
  kalman:
    # Sessions of a pair's own log price before its local linear trend model is
    # fitted. One year. The three noise variances are separated by how often the
    # level moves against how far a quote strays from it, and a shorter window
    # does not contain enough of either to tell them apart. It costs the oldest
    # fold 251 of its 1,289 training sessions and no validation session.
    burnin: 252
    # How often the three variances are re-estimated. A quarter. Each estimate is
    # a Nelder-Mead search over the whole expanding history, which is the most
    # expensive fit in the notebook, and quoting noise is not a fast-moving
    # quantity. The first value lands 2011-12-22.
    refit_every: 63
    # Iteration cap for that search.
    maxiter: 300
  hmm:
    # A calm dollar state and a turbulent one.
    n_states: 2
    # Sessions of the dollar factor before the regime chain is fitted. Two years.
    # A two-state chain has to see both states switch several times before its
    # transition matrix means anything, and this is one market-level series
    # rather than a panel, so the burn-in is paid once for the whole notebook.
    # It matches cme_futures, whose regime model reads a daily market-level
    # series of the same shape. The dollar factor itself starts 2011-02-01,
    # after the 21-session volatility window it carries, so the first regime
    # value lands 2013-01-15: 504 of the oldest fold's 1,269 dollar-factor
    # training sessions, and no validation session.
    burnin: 504
    # How often the chain is re-estimated. A quarter, as in etfs and cme_futures:
    # regime parameters are the slowest-moving thing this notebook fits and each
    # estimate costs `n_restarts` expectation-maximization searches.
    refit_every: 63
    # Expectation-maximization reaches a local optimum, so each fit is repeated
    # from this many starting points and the highest training likelihood is kept.
    n_restarts: 10
    # A restart is rejected when its final EM step falls by more than this
    # fraction of the log-likelihood's own magnitude. Real divergence moves
    # hundreds of nats; what this has to tolerate is single digits against a
    # likelihood of about 4.3e4.
    stability_rel_tol: 0.001
  arima:
    # Sessions of a pair's own returns before its short-memory return model is
    # fitted. One year, matching cme_futures' ARIMA, which is the same kind of
    # model on the same kind of daily per-entity series. A return needs the
    # session before it, so this series starts one session later than the price
    # panel and the first value lands 2011-12-23, costing the oldest fold 252 of
    # its 1,289 training sessions and no validation session.
    burnin: 252
    # Monthly. The coefficients of a one-lag return model are the fastest-moving
    # parameters here and the fit is cheap, so this is the shortest cadence in
    # the block.
    refit_every: 21
    # One autoregressive term and one moving-average term - the shortest memory
    # the family offers, which is all daily currency returns support.
    order: [1, 0, 1]

evaluation:
  n_splits: 8
  train_size: P5Y
  val_size: P1Y
  holdout_start: '2024-01-01'
  holdout_end: '2025-12-31'
  calendar: FX
  periods_per_year: 252  # FX 5d/wk

labels:
  primary: fwd_ret_1d
  buffer: 1D
  variants:
    - fwd_ret_5d
    - fwd_ret_21d
  variant_buffers:
    fwd_ret_5d: 5D
    fwd_ret_21d: 21D
  # Vectorized-backtest thinning step per label: number of schedule slots
  # to advance per trade so holding periods don't overlap.
  rebalance_step:
    fwd_ret_1d: 1
    fwd_ret_5d: 5
    fwd_ret_21d: 21

modeling:
  gbm:
    libraries: [lightgbm]
    preset: default
    device: cpu
    # LightGBM's own CPU default. 63 is the GPU default and was carried over with the
    # device when these runs moved off the GPU, so every CPU fit was quantizing the design
    # matrix into a quarter of the bins the library would have used. Coarser bins are
    # faster and lose split points; the reader running this on a CPU gets what the
    # documentation describes.
    max_bin: 255

causal:
  treatment: mom_skip_recent
  # Bars the treatment's own construction window spans, which is what the placebo block has
  # to cover: permuting mom_skip_recent in blocks shorter than this destroys the serial
  # dependence the refutation exists to preserve, and the resulting p-value reads like a
  # refutation without being one. Declared here rather than inferred, because guessing which
  # element of a window list a column was built from puts a wrong number behind a right-looking
  # one. Derived from the construction, not chosen:
  #
  # `close.shift(skip_recent) / close.shift(momentum[-1]) - 1` in 03_financial_features,
  # which is 21 and 252. Three different numbers live in that one expression and conflating them
  # is easy, so all three are stated: the LOOKBACK is 252 sessions, the oldest price it reads;
  # the RETURN INTERVAL is 231 sessions, from t-252 to t-21, which is what the value measures;
  # and consecutive values OVERLAP in 230 of those 231. The declaration takes the lookback
  # because it is the larger, so a block of this length spans the 231-session dependence
  # whichever way the arithmetic is read.
  treatment_window: 252
  confounders: [vol_gk_21d, vol_gk_63d, zscore_21d]
  method: walk_forward_dml
  # Two things the register above cannot supply, and this is why the number is written here
  # rather than derived. `features.windows.momentum` is a list of bar counts, not a
  # suffix-keyed mapping, and `mom_skip_recent` is not built from any single element of it.
  #
  # And the block has to span this window, which is the reason it is declared at all. A block
  # permutation keeps the treatment's serial dependence while breaking its alignment with the
  # outcome; a block shorter than the construction window breaks the dependence too, which
  # narrows the placebo distribution and makes the p-value read stronger than the evidence is.
  # Sized from the label buffer alone this case study permuted a 252-session momentum column
  # in blocks of 1, 5 and 21.

```

Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT

Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.