跳至正文
返回文库全部文档

设计与资金费率结算对齐的加密永续合约研究方案

代码 《交易机器学习》

总结

该配置规定了加密货币永续期货的多空策略研究流程。它定义了一个按成交量筛选的19种资产范围,决策与每八小时一次的资金费率结算对齐,并在资金费率结算时点执行。主要目标是八小时前瞻收益,并包含收益和方向变体;该方案还描述了资金费率、溢价、价格波动和横截面特征,以及套息、均值回归、动量和波动率假设。

回测计划比较排名靠前的信号和投资组合配置方法,评估交易成本敏感性,并扫描止损、追踪止损和时间退出控制。滚动评估使用训练期和验证期,之后接一个日期确定的留出期。方案规定特征和标签的时点、模型拟合频率、资格条件和执行假设,以支持可复现分析。这些是研究设计选择,并非任何策略盈利的证据。本文指出的潜在失效情形包括资金费率上限、持续趋势看似极端、波动率比率不稳定,以及动量和均值回归信号重叠。

核心观点

  • 与资金费率对齐的决策每八小时进行一次,并在结算时点执行。
  • 该策略在按成交量筛选的永续期货资产范围内,比较按排名构建的多空信号。
  • 特征类别涵盖资金费率持有收益、溢价表现、波动率和横截面背景。
  • 投资组合实验改变配置方法、交易成本和仓位级风险控制。
  • 所列假设存在局限,包括资金费率有上限,以及趋势持续时间可能超过特征窗口。

标签

全文
# setup.yaml


```yaml
strategy_id: crypto_perps_funding
setup_version: v1

universe:
  symbols:
    - AAVEUSDT
    - ADAUSDT
    - APTUSDT
    - ATOMUSDT
    - AVAXUSDT
    - BNBUSDT
    - BTCUSDT
    - COMPUSDT
    - DOGEUSDT
    - DOTUSDT
    - ETHUSDT
    - INJUSDT
    - LINKUSDT
    - MKRUSDT
    - NEARUSDT
    - SOLUSDT
    - SUIUSDT
    - UNIUSDT
    - XRPUSDT
  n_assets: 19
  eligibility_rule: top_perps_by_volume
  panel_note: Unbalanced panel; assets enter at listing date (no backfill).

decision:
  cadence: 8_hour_funding_aligned
  snapshot: pre_funding_timestamp
  execution_delay: at_funding_timestamp

# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study (cash + share_type are spec-hash inputs).
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator (inverse_vol, risk_parity, hrp,
# mvo_ledoit_wolf). Crypto perps trade 8-hourly (3 bars/day); 240 bars
# ≈ 80 days of underlying coverage. CS-level ``periods_per_year=365``
# annualizes Sharpe at daily-equivalent grain; allocator windows are
# measured in raw 8h bars regardless.
execution:
  initial_cash: 100_000          # IBKR retail-equivalent cohort default
  share_type: fractional         # Crypto perps trade in fractional contracts
  allocator_lookback: 240        # ~80 days of 8-hourly bars

mapping:
  class: long_short_funding_aligned
  position_state_space: long_short
  entry_logic: threshold_or_rank_based
  sizing: equal_weight_or_risk_parity

costs:
  class: material
  components: [taker_fee, maker_fee]
  # Headline tier used by Ch18 spread-estimation and the cost-comparison
  # analytics in 12_model_analysis. Majors (BTC, ETH, BNB, SOL, XRP) clear
  # with maker fees at the tight spread; alts pay taker.
  fee_schedule:
    taker_bps: 4
    maker_bps: 2

backtest:
  rebalance:
    # A rebalance is skipped when the per-asset weight change is below
    # min_weight_change AND the resulting trade notional is below
    # min_trade_value. The benchmark profile disables thresholds so that
    # full-universe equal-weight (1/N per asset) rebalances at all.
    default:
      min_weight_change: 0.005
      min_trade_value: 100.0
    benchmark:
      min_weight_change: 0.0
      min_trade_value: 0.0
  sweep:
    # Iteration controls per stage. ``signal: 0`` means "all predictions";
    # downstream stages take the top-N from the upstream stage's rank-1.
    # Notebooks read these via get_top_n_predictions(case_study, stage).
    top_n_predictions:
      signal: 0                 # all signal predictions (eq-weight baseline)
      allocation: 10            # top-10 model configs by equal-weight baseline Sharpe
      cost_sensitivity: 1       # top-1 of {signal+allocation} per label
      risk_overlay: 1           # top-1 of {signal+allocation} per label
    # Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
    expensive_allocators_skip: false
    # All others (equal_weight, score_weighted, inverse_vol) take the cheap path.
    # Ch16 signal-stage selection. Long-short by construction. With only
    # ~20 perps, every label exercises both top-k and quintile axes.
    #
    # k is a concentration choice and only means something against the tradeable
    # cross-section. Measured from the label artifacts on 2026-08-23, that is 9 names
    # per decision date at p10, 18 at the median and 19 at p90 - the smallest panel in
    # the fleet. So k=5 is already 28% of the book and k=10 is 56%: both are the
    # diversified end, and a grid of [5, 10] never shows a reader the concentrated
    # side of the tradeoff. k=3 is 17% and supplies it.
    top_k_grid:
      fwd_ret_8h:    [3, 5, 10]
      fwd_ret_24h:   [3, 5, 10]
      fwd_dir_8h:    [3, 5, 10]
      fwd_dir_8h_3c: [3, 5, 10]
    quantile_grid:
      fwd_ret_8h:    [5]
      fwd_ret_24h:   [5]
      fwd_dir_8h:    [5]
      fwd_dir_8h_3c: [5]
    # Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
    # Moment-based allocators (IV/RP/HRP/MVO_LW) all use the CS-level
    # ``execution.allocator_lookback`` (240 8-hourly bars). No max-weight cap.
    # Equal weight is the baseline above, so this list contains alternatives only.
    allocators:
      - {name: score_weighted,  method: score_weighted}
      - {name: inverse_vol,     method: inverse_vol}
      - {name: risk_parity,     method: risk_parity}
      - {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf}
      - {name: hrp,             method: hrp}
      - {name: conformal_weighted, method: conformal_weighted}
    # Ch18 cost sensitivity (bps).
    cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
    # Ch19 risk overlays.
    risk_controls:
      position:
        - {name: stop_loss_3pct,  type: stop_loss,     threshold: 0.03}
        - {name: stop_loss_5pct,  type: stop_loss,     threshold: 0.05}
        - {name: stop_loss_10pct, type: stop_loss,     threshold: 0.10}
        - {name: stop_loss_15pct, type: stop_loss,     threshold: 0.15}
        - {name: trailing_1pct,   type: trailing_stop, threshold: 0.01}
        - {name: trailing_2pct,   type: trailing_stop, threshold: 0.02}
        - {name: trailing_3pct,   type: trailing_stop, threshold: 0.03}
        - {name: trailing_5pct,   type: trailing_stop, threshold: 0.05}
        - {name: trailing_10pct,  type: trailing_stop, threshold: 0.10}
        - {name: trailing_15pct,  type: trailing_stop, threshold: 0.15}
        - {name: trailing_20pct,  type: trailing_stop, threshold: 0.20}
        - {name: time_exit_10,    type: time_exit,     bars: 10}
        - {name: time_exit_20,    type: time_exit,     bars: 20}
        - {name: time_exit_40,    type: time_exit,     bars: 40}

evaluation:
  n_splits: 2
  train_size: 2Y
  val_size: 1Y
  holdout_start: '2024-01-01'
  holdout_end: '2025-12-31'
  calendar: crypto
  periods_per_year: 365  # crypto 7d/wk

labels:
  primary: fwd_ret_8h
  buffer: 8H
  variants:
    - fwd_ret_24h
    - fwd_dir_8h
    - fwd_dir_8h_3c
  variant_buffers:
    fwd_ret_24h: 24H
    fwd_dir_8h: 8H
    fwd_dir_8h_3c: 8H
  # Vectorized-backtest thinning step per label: number of schedule slots
  # to advance per trade so holding periods don't overlap.
  rebalance_step:
    fwd_ret_8h: 1
    fwd_ret_24h: 3   # 8h schedule, 24h horizon -> ceil(24/8) = 3
    fwd_dir_8h: 1
    fwd_dir_8h_3c: 1
  # Continuous return that each classification label is derived from.
  # IC for classification predictions is computed against this column;
  # AUC/accuracy/log_loss are computed against the binary label itself.
  classification_eval_label:
    fwd_dir_8h: fwd_ret_8h
    fwd_dir_8h_3c: fwd_ret_8h

# The feature-specification register and every window `03_financial_features`
# reads. A window typed into the notebook is a second copy of a number the
# warmup audit and the timing figure both have to agree with, so all of them
# are declared once here and bound. Windows are counted in 8-hour settlement
# bars; the map key is the suffix the emitted column carries.
features:
  bar_hours: 8
  # Fee tier, not a liquidity screen: these five clear at the maker spread and
  # the rest pay taker. Same five the `costs.fee_schedule` note above names.
  majors: [BNBUSDT, BTCUSDT, ETHUSDT, SOLUSDT, XRPUSDT]
  ranked: premium_index_close
  # Two features above this absolute rank correlation carry one ordering, so a linear
  # model cannot separate their contributions. F5 cuts the redundancy tree here.
  redundancy_cut: 0.7
  windows:
    premium_momentum: {8h: 1, 24h: 3, 72h: 9, 168h: 21, 336h: 42, 720h: 90}
    premium_volatility: {24h: 3, 72h: 9, 168h: 21, 336h: 42}
    premium_zscore: {7d: 21, 14d: 42}
    premium_dev_mean: {7d: 21, 14d: 42}
    premium_quantile: {7d: 21, 14d: 42, 30d: 90}
    premium_rsi: {24h: 3, 72h: 9}
    price_volatility: {7d: 21, 14d: 42}
    premium_persistence: {7d: 21}
    premium_regime: {72h: 9}
    funding_zscore: {14d: 42}
    funding_half_life: {14d: 42}
    funding_change: {24h: 3}
    funding_cashflow: {7d: 7}
  # Bounds that shape an emitted value rather than guard a denominator. The
  # z-score clip holds a settlement-day outlier off the scale a model reads;
  # the AR(1) clip keeps the half-life finite at a unit root.
  clip:
    zscore: 10.0
    vol_ratio: 10.0
    ar1: 0.999
    half_life: [0.5, 100.0]
  families:
    - name: carry
      pattern: funding_rate|funding_rate_*|cum_positive_funding_7d|funding_half_life_14d|premium_level|premium_rank|premium_zscore_*
      role: signal
      hypothesis: A perpetual whose holders are paying to stay long is crowded, and the crowding unwinds.
      inputs: official funding settlements, premium index close
      lookback: 43
      lag: 0
      frame: per symbol, except the premium percentile which is within the decision timestamp
      representation: level, trailing z-score, cross-sectional percentile, mean-reversion speed
      failure_mode: Funding is clamped by the exchange, so the level saturates in the regimes that matter most.
    - name: mean_reversion
      pattern: premium_dev_mean_*|premium_quantile_pos_*|premium_persistence_*
      role: signal
      hypothesis: A premium far from its own recent range reverts faster than one near the middle of it.
      inputs: premium index close
      lookback: 90
      lag: 0
      frame: per symbol
      representation: deviation from a trailing mean, rolling percentile, sign persistence
      failure_mode: A trending premium looks extreme against its own window for as long as the trend lasts.
    - name: momentum
      pattern: premium_change_*|premium_accel_*
      role: signal
      hypothesis: A premium that has been widening keeps widening over the next settlement or two.
      inputs: premium index close
      lookback: 90
      lag: 0
      frame: per symbol
      representation: differences at six horizons, plus differences between horizons
      failure_mode: Momentum and mean reversion read the same series with opposite signs and cancel.
    - name: volatility
      pattern: premium_vol_*|price_vol_*|vol_ratio_*
      role: state
      hypothesis: Liquidation cascades widen both the premium and the price, and the two carry different information.
      inputs: premium index close, perpetual close
      lookback: 43
      lag: 0
      frame: per symbol
      representation: trailing dispersion at four horizons, plus short-over-long ratios
      failure_mode: A ratio of two dispersions is unstable when the denominator window is quiet.
    - name: cross_sectional
      pattern: premium_vs_median|premium_xs_zscore|xs_funding_dispersion
      role: state
      hypothesis: Whether a premium is high depends on what the rest of the universe is paying that settlement.
      inputs: premium index close, official funding settlements
      lookback: 1
      lag: 0
      frame: within the decision timestamp
      representation: distance from the cross-sectional median, cross-sectional z-score, dispersion
      failure_mode: The panel is unbalanced, so early dates rank against a handful of symbols.
    - name: regime
      pattern: premium_regime_*|premium_rsi_*|funding_session|cost_tier_alt
      role: state
      hypothesis: A sustained premium and the settlement slot condition how any signal should be read.
      inputs: premium index close, symbol, decision timestamp
      lookback: 10
      lag: 0
      frame: per symbol, except the session which is a property of the timestamp
      representation: signed regime average, bounded oscillator, categorical slot and fee tier
      failure_mode: The fee tier is a fixed list, so it does not follow a symbol across a tier change.

# What `04_model_based_features` decides, declared here for the same reason the
# feature windows above are. Every count is in 8-hour settlement bars, the unit
# `features.bar_hours` sets and `features.windows` already uses.
model_based:
  # Trailing settlements a series needs before either model is fitted on it. 500
  # settlements is about five and a half months; below that the leverage term of a
  # GJR recursion and the transition matrix of a two-state chain are estimated off
  # too few regime switches to mean anything.
  #
  # This is a burn-in, not a fold-entry condition. Both models below are fitted on a
  # schedule that runs over the whole history: the first 500 settlements carry no
  # value, and from there the parameters are re-estimated on the cadence each model
  # declares, always on settlements strictly earlier than the ones they then speak
  # for. Nothing about a cross-validation fold enters the fit, so the artifact
  # carries no fold column and a settlement has one value whichever fold selects it.
  min_train_bars: 500
  garch:
    # How often the variance model is re-estimated. 21 settlements is a week. A
    # variance model tracks a level that moves, which is the property the feature
    # exists to report, so it is refreshed faster than the regime model below.
    refit_every: 21
    # The conditional-volatility z-score compares a symbol's current forecast
    # against its own recent level: 90 bars is 30 days.
    vol_zscore_window: 90
    # Bound on the emitted z-score, so one liquidation cascade does not set the
    # scale a model reads. Same role as `features.clip.zscore`.
    zscore_clip: 10.0
  hmm:
    # Calm funding and stressed funding.
    n_states: 2
    # How often the chain is re-estimated. 63 settlements is three weeks. Regime
    # parameters are the slowest-moving thing this notebook fits, and each estimate
    # costs `n_restarts` expectation-maximization searches over the whole history.
    refit_every: 63
    # Expectation-maximization reaches a local optimum, so the fit is repeated
    # from this many starting points and the highest training likelihood is kept.
    # Measured on etfs' panel through the same estimator, ten restarts give ten
    # distinct log-likelihoods, so the search explores rather than repeating itself.
    n_restarts: 10
    # Convergence threshold on the per-settlement log-likelihood gain. Tighter than
    # hmmlearn's 1e-2 default, which is what this notebook has always fitted at and
    # what fx_pairs chose independently. Declared rather than baked into the shared
    # estimator: measured on etfs, the two tolerances move the emitted probability by
    # about 1e-3 on average and flip none of 756 state assignments, so the looser
    # default is not wrong - it is just not what this notebook fits at.
    tol: 1.0e-4
  incremental_ic:
    # A settlement needs this many symbols quoting before its cross-sectional rank
    # correlation is used, and a feature needs this many usable settlements before
    # its mean is reported.
    min_cross_section: 10
    min_decision_times: 20

modeling:
  gbm:
    libraries: [lightgbm]
    preset: default
    device: cpu
    # LightGBM's own CPU default. 63 is the GPU default and was carried over with the
    # device when these runs moved off the GPU, so every CPU fit was quantizing the design
    # matrix into a quarter of the bins the library would have used. Coarser bins are
    # faster and lose split points; the reader running this on a CPU gets what the
    # documentation describes.
    max_bin: 255

causal:
  treatment: premium_zscore_14d
  confounders: [price_vol_14d, funding_rate, premium_dev_mean_14d]
  method: walk_forward_dml

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。