본문으로 건너뛰기
라이브러리 문서 전체

펀딩 시점에 맞춘 암호화폐 무기한 선물 연구 설정

코드 Machine Learning for Trading

요약

이 설정은 암호화폐 무기한 선물을 위한 롱·숏 전략 연구 워크플로를 정의합니다. 거래량으로 선별한 19개 자산 투자 대상군, 8시간 펀딩 정산에 맞춘 의사결정, 펀딩 시점 체결을 설정합니다. 주요 목표는 8시간 선행 수익률이며 수익률 및 방향 변형도 포함합니다. 펀딩, 프리미엄, 가격 변동성, 횡단면 특성과 캐리, 평균회귀, 모멘텀, 변동성 가설도 설명합니다.

백테스트 계획은 상위 순위 신호와 포트폴리오 배분기를 비교하고, 거래 비용 민감도를 평가하며, 손절, 추적 손절, 시간 기반 청산 설정을 다양하게 시험합니다. 워크포워드 평가에서는 학습 및 검증 기간 뒤에 날짜가 정해진 홀드아웃을 둡니다. 재현 가능한 분석을 위해 특성과 라벨의 시점, 모델 학습 주기, 적격성, 체결 가정을 명시합니다. 이는 연구 설계 선택이지 전략의 수익성에 대한 증거가 아닙니다. 문서는 펀딩 비율 상한, 극단처럼 보이는 상태가 계속되는 추세, 불안정한 변동성 비율, 겹치는 모멘텀·평균회귀 신호 같은 잠재적 실패 요인을 언급합니다.

핵심 아이디어

  • 펀딩에 맞춘 의사결정은 8시간 주기를 사용하고 정산 시점에 체결합니다.
  • 거래량으로 선별한 무기한 선물 투자 대상군에서 순위 기반 롱·숏 신호를 비교합니다.
  • 특성군은 펀딩 캐리, 프리미엄 양상, 변동성, 횡단면 맥락을 나타냅니다.
  • 포트폴리오 실험에서 배분 방식, 거래 비용, 포지션별 위험 통제를 다양하게 설정합니다.
  • 펀딩 상한과 특성 구간보다 더 오래 지속되는 추세 등 가설에는 유의할 점이 있습니다.

태그

전문
# setup.yaml


```yaml
strategy_id: crypto_perps_funding
setup_version: v1

universe:
  symbols:
    - AAVEUSDT
    - ADAUSDT
    - APTUSDT
    - ATOMUSDT
    - AVAXUSDT
    - BNBUSDT
    - BTCUSDT
    - COMPUSDT
    - DOGEUSDT
    - DOTUSDT
    - ETHUSDT
    - INJUSDT
    - LINKUSDT
    - MKRUSDT
    - NEARUSDT
    - SOLUSDT
    - SUIUSDT
    - UNIUSDT
    - XRPUSDT
  n_assets: 19
  eligibility_rule: top_perps_by_volume
  panel_note: Unbalanced panel; assets enter at listing date (no backfill).

decision:
  cadence: 8_hour_funding_aligned
  snapshot: pre_funding_timestamp
  execution_delay: at_funding_timestamp

# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study (cash + share_type are spec-hash inputs).
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator (inverse_vol, risk_parity, hrp,
# mvo_ledoit_wolf). Crypto perps trade 8-hourly (3 bars/day); 240 bars
# ≈ 80 days of underlying coverage. CS-level ``periods_per_year=365``
# annualizes Sharpe at daily-equivalent grain; allocator windows are
# measured in raw 8h bars regardless.
execution:
  initial_cash: 100_000          # IBKR retail-equivalent cohort default
  share_type: fractional         # Crypto perps trade in fractional contracts
  allocator_lookback: 240        # ~80 days of 8-hourly bars

mapping:
  class: long_short_funding_aligned
  position_state_space: long_short
  entry_logic: threshold_or_rank_based
  sizing: equal_weight_or_risk_parity

costs:
  class: material
  components: [taker_fee, maker_fee]
  # Headline tier used by Ch18 spread-estimation and the cost-comparison
  # analytics in 12_model_analysis. Majors (BTC, ETH, BNB, SOL, XRP) clear
  # with maker fees at the tight spread; alts pay taker.
  fee_schedule:
    taker_bps: 4
    maker_bps: 2

backtest:
  rebalance:
    # A rebalance is skipped when the per-asset weight change is below
    # min_weight_change AND the resulting trade notional is below
    # min_trade_value. The benchmark profile disables thresholds so that
    # full-universe equal-weight (1/N per asset) rebalances at all.
    default:
      min_weight_change: 0.005
      min_trade_value: 100.0
    benchmark:
      min_weight_change: 0.0
      min_trade_value: 0.0
  sweep:
    # Iteration controls per stage. ``signal: 0`` means "all predictions";
    # downstream stages take the top-N from the upstream stage's rank-1.
    # Notebooks read these via get_top_n_predictions(case_study, stage).
    top_n_predictions:
      signal: 0                 # all signal predictions (eq-weight baseline)
      allocation: 10            # top-10 model configs by equal-weight baseline Sharpe
      cost_sensitivity: 1       # top-1 of {signal+allocation} per label
      risk_overlay: 1           # top-1 of {signal+allocation} per label
    # Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
    expensive_allocators_skip: false
    # All others (equal_weight, score_weighted, inverse_vol) take the cheap path.
    # Ch16 signal-stage selection. Long-short by construction. With only
    # ~20 perps, every label exercises both top-k and quintile axes.
    #
    # k is a concentration choice and only means something against the tradeable
    # cross-section. Measured from the label artifacts on 2026-08-23, that is 9 names
    # per decision date at p10, 18 at the median and 19 at p90 - the smallest panel in
    # the fleet. So k=5 is already 28% of the book and k=10 is 56%: both are the
    # diversified end, and a grid of [5, 10] never shows a reader the concentrated
    # side of the tradeoff. k=3 is 17% and supplies it.
    top_k_grid:
      fwd_ret_8h:    [3, 5, 10]
      fwd_ret_24h:   [3, 5, 10]
      fwd_dir_8h:    [3, 5, 10]
      fwd_dir_8h_3c: [3, 5, 10]
    quantile_grid:
      fwd_ret_8h:    [5]
      fwd_ret_24h:   [5]
      fwd_dir_8h:    [5]
      fwd_dir_8h_3c: [5]
    # Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
    # Moment-based allocators (IV/RP/HRP/MVO_LW) all use the CS-level
    # ``execution.allocator_lookback`` (240 8-hourly bars). No max-weight cap.
    # Equal weight is the baseline above, so this list contains alternatives only.
    allocators:
      - {name: score_weighted,  method: score_weighted}
      - {name: inverse_vol,     method: inverse_vol}
      - {name: risk_parity,     method: risk_parity}
      - {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf}
      - {name: hrp,             method: hrp}
      - {name: conformal_weighted, method: conformal_weighted}
    # Ch18 cost sensitivity (bps).
    cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
    # Ch19 risk overlays.
    risk_controls:
      position:
        - {name: stop_loss_3pct,  type: stop_loss,     threshold: 0.03}
        - {name: stop_loss_5pct,  type: stop_loss,     threshold: 0.05}
        - {name: stop_loss_10pct, type: stop_loss,     threshold: 0.10}
        - {name: stop_loss_15pct, type: stop_loss,     threshold: 0.15}
        - {name: trailing_1pct,   type: trailing_stop, threshold: 0.01}
        - {name: trailing_2pct,   type: trailing_stop, threshold: 0.02}
        - {name: trailing_3pct,   type: trailing_stop, threshold: 0.03}
        - {name: trailing_5pct,   type: trailing_stop, threshold: 0.05}
        - {name: trailing_10pct,  type: trailing_stop, threshold: 0.10}
        - {name: trailing_15pct,  type: trailing_stop, threshold: 0.15}
        - {name: trailing_20pct,  type: trailing_stop, threshold: 0.20}
        - {name: time_exit_10,    type: time_exit,     bars: 10}
        - {name: time_exit_20,    type: time_exit,     bars: 20}
        - {name: time_exit_40,    type: time_exit,     bars: 40}

evaluation:
  n_splits: 2
  train_size: 2Y
  val_size: 1Y
  holdout_start: '2024-01-01'
  holdout_end: '2025-12-31'
  calendar: crypto
  periods_per_year: 365  # crypto 7d/wk

labels:
  primary: fwd_ret_8h
  buffer: 8H
  variants:
    - fwd_ret_24h
    - fwd_dir_8h
    - fwd_dir_8h_3c
  variant_buffers:
    fwd_ret_24h: 24H
    fwd_dir_8h: 8H
    fwd_dir_8h_3c: 8H
  # Vectorized-backtest thinning step per label: number of schedule slots
  # to advance per trade so holding periods don't overlap.
  rebalance_step:
    fwd_ret_8h: 1
    fwd_ret_24h: 3   # 8h schedule, 24h horizon -> ceil(24/8) = 3
    fwd_dir_8h: 1
    fwd_dir_8h_3c: 1
  # Continuous return that each classification label is derived from.
  # IC for classification predictions is computed against this column;
  # AUC/accuracy/log_loss are computed against the binary label itself.
  classification_eval_label:
    fwd_dir_8h: fwd_ret_8h
    fwd_dir_8h_3c: fwd_ret_8h

# The feature-specification register and every window `03_financial_features`
# reads. A window typed into the notebook is a second copy of a number the
# warmup audit and the timing figure both have to agree with, so all of them
# are declared once here and bound. Windows are counted in 8-hour settlement
# bars; the map key is the suffix the emitted column carries.
features:
  bar_hours: 8
  # Fee tier, not a liquidity screen: these five clear at the maker spread and
  # the rest pay taker. Same five the `costs.fee_schedule` note above names.
  majors: [BNBUSDT, BTCUSDT, ETHUSDT, SOLUSDT, XRPUSDT]
  ranked: premium_index_close
  # Two features above this absolute rank correlation carry one ordering, so a linear
  # model cannot separate their contributions. F5 cuts the redundancy tree here.
  redundancy_cut: 0.7
  windows:
    premium_momentum: {8h: 1, 24h: 3, 72h: 9, 168h: 21, 336h: 42, 720h: 90}
    premium_volatility: {24h: 3, 72h: 9, 168h: 21, 336h: 42}
    premium_zscore: {7d: 21, 14d: 42}
    premium_dev_mean: {7d: 21, 14d: 42}
    premium_quantile: {7d: 21, 14d: 42, 30d: 90}
    premium_rsi: {24h: 3, 72h: 9}
    price_volatility: {7d: 21, 14d: 42}
    premium_persistence: {7d: 21}
    premium_regime: {72h: 9}
    funding_zscore: {14d: 42}
    funding_half_life: {14d: 42}
    funding_change: {24h: 3}
    funding_cashflow: {7d: 7}
  # Bounds that shape an emitted value rather than guard a denominator. The
  # z-score clip holds a settlement-day outlier off the scale a model reads;
  # the AR(1) clip keeps the half-life finite at a unit root.
  clip:
    zscore: 10.0
    vol_ratio: 10.0
    ar1: 0.999
    half_life: [0.5, 100.0]
  families:
    - name: carry
      pattern: funding_rate|funding_rate_*|cum_positive_funding_7d|funding_half_life_14d|premium_level|premium_rank|premium_zscore_*
      role: signal
      hypothesis: A perpetual whose holders are paying to stay long is crowded, and the crowding unwinds.
      inputs: official funding settlements, premium index close
      lookback: 43
      lag: 0
      frame: per symbol, except the premium percentile which is within the decision timestamp
      representation: level, trailing z-score, cross-sectional percentile, mean-reversion speed
      failure_mode: Funding is clamped by the exchange, so the level saturates in the regimes that matter most.
    - name: mean_reversion
      pattern: premium_dev_mean_*|premium_quantile_pos_*|premium_persistence_*
      role: signal
      hypothesis: A premium far from its own recent range reverts faster than one near the middle of it.
      inputs: premium index close
      lookback: 90
      lag: 0
      frame: per symbol
      representation: deviation from a trailing mean, rolling percentile, sign persistence
      failure_mode: A trending premium looks extreme against its own window for as long as the trend lasts.
    - name: momentum
      pattern: premium_change_*|premium_accel_*
      role: signal
      hypothesis: A premium that has been widening keeps widening over the next settlement or two.
      inputs: premium index close
      lookback: 90
      lag: 0
      frame: per symbol
      representation: differences at six horizons, plus differences between horizons
      failure_mode: Momentum and mean reversion read the same series with opposite signs and cancel.
    - name: volatility
      pattern: premium_vol_*|price_vol_*|vol_ratio_*
      role: state
      hypothesis: Liquidation cascades widen both the premium and the price, and the two carry different information.
      inputs: premium index close, perpetual close
      lookback: 43
      lag: 0
      frame: per symbol
      representation: trailing dispersion at four horizons, plus short-over-long ratios
      failure_mode: A ratio of two dispersions is unstable when the denominator window is quiet.
    - name: cross_sectional
      pattern: premium_vs_median|premium_xs_zscore|xs_funding_dispersion
      role: state
      hypothesis: Whether a premium is high depends on what the rest of the universe is paying that settlement.
      inputs: premium index close, official funding settlements
      lookback: 1
      lag: 0
      frame: within the decision timestamp
      representation: distance from the cross-sectional median, cross-sectional z-score, dispersion
      failure_mode: The panel is unbalanced, so early dates rank against a handful of symbols.
    - name: regime
      pattern: premium_regime_*|premium_rsi_*|funding_session|cost_tier_alt
      role: state
      hypothesis: A sustained premium and the settlement slot condition how any signal should be read.
      inputs: premium index close, symbol, decision timestamp
      lookback: 10
      lag: 0
      frame: per symbol, except the session which is a property of the timestamp
      representation: signed regime average, bounded oscillator, categorical slot and fee tier
      failure_mode: The fee tier is a fixed list, so it does not follow a symbol across a tier change.

# What `04_model_based_features` decides, declared here for the same reason the
# feature windows above are. Every count is in 8-hour settlement bars, the unit
# `features.bar_hours` sets and `features.windows` already uses.
model_based:
  # Trailing settlements a series needs before either model is fitted on it. 500
  # settlements is about five and a half months; below that the leverage term of a
  # GJR recursion and the transition matrix of a two-state chain are estimated off
  # too few regime switches to mean anything.
  #
  # This is a burn-in, not a fold-entry condition. Both models below are fitted on a
  # schedule that runs over the whole history: the first 500 settlements carry no
  # value, and from there the parameters are re-estimated on the cadence each model
  # declares, always on settlements strictly earlier than the ones they then speak
  # for. Nothing about a cross-validation fold enters the fit, so the artifact
  # carries no fold column and a settlement has one value whichever fold selects it.
  min_train_bars: 500
  garch:
    # How often the variance model is re-estimated. 21 settlements is a week. A
    # variance model tracks a level that moves, which is the property the feature
    # exists to report, so it is refreshed faster than the regime model below.
    refit_every: 21
    # The conditional-volatility z-score compares a symbol's current forecast
    # against its own recent level: 90 bars is 30 days.
    vol_zscore_window: 90
    # Bound on the emitted z-score, so one liquidation cascade does not set the
    # scale a model reads. Same role as `features.clip.zscore`.
    zscore_clip: 10.0
  hmm:
    # Calm funding and stressed funding.
    n_states: 2
    # How often the chain is re-estimated. 63 settlements is three weeks. Regime
    # parameters are the slowest-moving thing this notebook fits, and each estimate
    # costs `n_restarts` expectation-maximization searches over the whole history.
    refit_every: 63
    # Expectation-maximization reaches a local optimum, so the fit is repeated
    # from this many starting points and the highest training likelihood is kept.
    # Measured on etfs' panel through the same estimator, ten restarts give ten
    # distinct log-likelihoods, so the search explores rather than repeating itself.
    n_restarts: 10
    # Convergence threshold on the per-settlement log-likelihood gain. Tighter than
    # hmmlearn's 1e-2 default, which is what this notebook has always fitted at and
    # what fx_pairs chose independently. Declared rather than baked into the shared
    # estimator: measured on etfs, the two tolerances move the emitted probability by
    # about 1e-3 on average and flip none of 756 state assignments, so the looser
    # default is not wrong - it is just not what this notebook fits at.
    tol: 1.0e-4
  incremental_ic:
    # A settlement needs this many symbols quoting before its cross-sectional rank
    # correlation is used, and a feature needs this many usable settlements before
    # its mean is reported.
    min_cross_section: 10
    min_decision_times: 20

modeling:
  gbm:
    libraries: [lightgbm]
    preset: default
    device: cpu
    # LightGBM's own CPU default. 63 is the GPU default and was carried over with the
    # device when these runs moved off the GPU, so every CPU fit was quantizing the design
    # matrix into a quarter of the bins the library would have used. Coarser bins are
    # faster and lose split points; the reader running this on a CPU gets what the
    # documentation describes.
    max_bin: 255

causal:
  treatment: premium_zscore_14d
  confounders: [price_vol_14d, funding_rate, premium_dev_mean_14d]
  method: walk_forward_dml

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.