Pular para o conteúdo
Todos os documentos da biblioteca

Configuração de pesquisa para perpétuos cripto alinhada ao funding

Código Machine Learning for Trading

Resumo

Esta configuração especifica um fluxo de pesquisa de estratégia long-short para contratos futuros perpétuos de cripto. Define um universo de 19 ativos selecionados por volume, decisões alinhadas aos pagamentos de funding a cada oito horas e execução no horário do funding. O alvo principal é o retorno futuro em oito horas, com variantes de retorno e direção; a configuração também descreve variáveis de funding, prêmio, volatilidade de preço e contexto transversal, além de hipóteses sobre carry, reversão à média, momentum e volatilidade.

O plano de backtest compara os sinais mais bem classificados e métodos de alocação de carteira, avalia a sensibilidade aos custos de transação e testa controles de stop loss, stop móvel e saída por tempo. A avaliação walk-forward usa períodos de treinamento e validação, seguidos por um holdout datado. Os tempos das variáveis e dos rótulos, a frequência de ajuste do modelo, os critérios de elegibilidade e as premissas de execução são especificados para permitir uma análise reproduzível. Essas são decisões de desenho de pesquisa, não evidências de que alguma estratégia seja lucrativa. O documento aponta possíveis modos de falha, como limites para taxas de funding, tendências persistentes que parecem extremas, razões de volatilidade instáveis e sinais sobrepostos de momentum e reversão à média.

Ideias principais

  • Decisões alinhadas ao funding seguem uma cadência de oito horas e são executadas no horário da liquidação.
  • A estratégia compara sinais long-short classificados em um universo de futuros perpétuos selecionado por volume.
  • As famílias de variáveis codificam carry de funding, comportamento do prêmio, volatilidade e contexto transversal.
  • Os experimentos de carteira variam os métodos de alocação, os custos de transação e os controles de risco por posição.
  • As hipóteses declaradas têm ressalvas, como funding limitado e tendências que podem persistir além das janelas das variáveis.

Tags

Texto completo
# setup.yaml


```yaml
strategy_id: crypto_perps_funding
setup_version: v1

universe:
  symbols:
    - AAVEUSDT
    - ADAUSDT
    - APTUSDT
    - ATOMUSDT
    - AVAXUSDT
    - BNBUSDT
    - BTCUSDT
    - COMPUSDT
    - DOGEUSDT
    - DOTUSDT
    - ETHUSDT
    - INJUSDT
    - LINKUSDT
    - MKRUSDT
    - NEARUSDT
    - SOLUSDT
    - SUIUSDT
    - UNIUSDT
    - XRPUSDT
  n_assets: 19
  eligibility_rule: top_perps_by_volume
  panel_note: Unbalanced panel; assets enter at listing date (no backfill).

decision:
  cadence: 8_hour_funding_aligned
  snapshot: pre_funding_timestamp
  execution_delay: at_funding_timestamp

# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study (cash + share_type are spec-hash inputs).
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator (inverse_vol, risk_parity, hrp,
# mvo_ledoit_wolf). Crypto perps trade 8-hourly (3 bars/day); 240 bars
# ≈ 80 days of underlying coverage. CS-level ``periods_per_year=365``
# annualizes Sharpe at daily-equivalent grain; allocator windows are
# measured in raw 8h bars regardless.
execution:
  initial_cash: 100_000          # IBKR retail-equivalent cohort default
  share_type: fractional         # Crypto perps trade in fractional contracts
  allocator_lookback: 240        # ~80 days of 8-hourly bars

mapping:
  class: long_short_funding_aligned
  position_state_space: long_short
  entry_logic: threshold_or_rank_based
  sizing: equal_weight_or_risk_parity

costs:
  class: material
  components: [taker_fee, maker_fee]
  # Headline tier used by Ch18 spread-estimation and the cost-comparison
  # analytics in 12_model_analysis. Majors (BTC, ETH, BNB, SOL, XRP) clear
  # with maker fees at the tight spread; alts pay taker.
  fee_schedule:
    taker_bps: 4
    maker_bps: 2

backtest:
  rebalance:
    # A rebalance is skipped when the per-asset weight change is below
    # min_weight_change AND the resulting trade notional is below
    # min_trade_value. The benchmark profile disables thresholds so that
    # full-universe equal-weight (1/N per asset) rebalances at all.
    default:
      min_weight_change: 0.005
      min_trade_value: 100.0
    benchmark:
      min_weight_change: 0.0
      min_trade_value: 0.0
  sweep:
    # Iteration controls per stage. ``signal: 0`` means "all predictions";
    # downstream stages take the top-N from the upstream stage's rank-1.
    # Notebooks read these via get_top_n_predictions(case_study, stage).
    top_n_predictions:
      signal: 0                 # all signal predictions (eq-weight baseline)
      allocation: 10            # top-10 model configs by equal-weight baseline Sharpe
      cost_sensitivity: 1       # top-1 of {signal+allocation} per label
      risk_overlay: 1           # top-1 of {signal+allocation} per label
    # Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
    expensive_allocators_skip: false
    # All others (equal_weight, score_weighted, inverse_vol) take the cheap path.
    # Ch16 signal-stage selection. Long-short by construction. With only
    # ~20 perps, every label exercises both top-k and quintile axes.
    #
    # k is a concentration choice and only means something against the tradeable
    # cross-section. Measured from the label artifacts on 2026-08-23, that is 9 names
    # per decision date at p10, 18 at the median and 19 at p90 - the smallest panel in
    # the fleet. So k=5 is already 28% of the book and k=10 is 56%: both are the
    # diversified end, and a grid of [5, 10] never shows a reader the concentrated
    # side of the tradeoff. k=3 is 17% and supplies it.
    top_k_grid:
      fwd_ret_8h:    [3, 5, 10]
      fwd_ret_24h:   [3, 5, 10]
      fwd_dir_8h:    [3, 5, 10]
      fwd_dir_8h_3c: [3, 5, 10]
    quantile_grid:
      fwd_ret_8h:    [5]
      fwd_ret_24h:   [5]
      fwd_dir_8h:    [5]
      fwd_dir_8h_3c: [5]
    # Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
    # Moment-based allocators (IV/RP/HRP/MVO_LW) all use the CS-level
    # ``execution.allocator_lookback`` (240 8-hourly bars). No max-weight cap.
    # Equal weight is the baseline above, so this list contains alternatives only.
    allocators:
      - {name: score_weighted,  method: score_weighted}
      - {name: inverse_vol,     method: inverse_vol}
      - {name: risk_parity,     method: risk_parity}
      - {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf}
      - {name: hrp,             method: hrp}
      - {name: conformal_weighted, method: conformal_weighted}
    # Ch18 cost sensitivity (bps).
    cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
    # Ch19 risk overlays.
    risk_controls:
      position:
        - {name: stop_loss_3pct,  type: stop_loss,     threshold: 0.03}
        - {name: stop_loss_5pct,  type: stop_loss,     threshold: 0.05}
        - {name: stop_loss_10pct, type: stop_loss,     threshold: 0.10}
        - {name: stop_loss_15pct, type: stop_loss,     threshold: 0.15}
        - {name: trailing_1pct,   type: trailing_stop, threshold: 0.01}
        - {name: trailing_2pct,   type: trailing_stop, threshold: 0.02}
        - {name: trailing_3pct,   type: trailing_stop, threshold: 0.03}
        - {name: trailing_5pct,   type: trailing_stop, threshold: 0.05}
        - {name: trailing_10pct,  type: trailing_stop, threshold: 0.10}
        - {name: trailing_15pct,  type: trailing_stop, threshold: 0.15}
        - {name: trailing_20pct,  type: trailing_stop, threshold: 0.20}
        - {name: time_exit_10,    type: time_exit,     bars: 10}
        - {name: time_exit_20,    type: time_exit,     bars: 20}
        - {name: time_exit_40,    type: time_exit,     bars: 40}

evaluation:
  n_splits: 2
  train_size: 2Y
  val_size: 1Y
  holdout_start: '2024-01-01'
  holdout_end: '2025-12-31'
  calendar: crypto
  periods_per_year: 365  # crypto 7d/wk

labels:
  primary: fwd_ret_8h
  buffer: 8H
  variants:
    - fwd_ret_24h
    - fwd_dir_8h
    - fwd_dir_8h_3c
  variant_buffers:
    fwd_ret_24h: 24H
    fwd_dir_8h: 8H
    fwd_dir_8h_3c: 8H
  # Vectorized-backtest thinning step per label: number of schedule slots
  # to advance per trade so holding periods don't overlap.
  rebalance_step:
    fwd_ret_8h: 1
    fwd_ret_24h: 3   # 8h schedule, 24h horizon -> ceil(24/8) = 3
    fwd_dir_8h: 1
    fwd_dir_8h_3c: 1
  # Continuous return that each classification label is derived from.
  # IC for classification predictions is computed against this column;
  # AUC/accuracy/log_loss are computed against the binary label itself.
  classification_eval_label:
    fwd_dir_8h: fwd_ret_8h
    fwd_dir_8h_3c: fwd_ret_8h

# The feature-specification register and every window `03_financial_features`
# reads. A window typed into the notebook is a second copy of a number the
# warmup audit and the timing figure both have to agree with, so all of them
# are declared once here and bound. Windows are counted in 8-hour settlement
# bars; the map key is the suffix the emitted column carries.
features:
  bar_hours: 8
  # Fee tier, not a liquidity screen: these five clear at the maker spread and
  # the rest pay taker. Same five the `costs.fee_schedule` note above names.
  majors: [BNBUSDT, BTCUSDT, ETHUSDT, SOLUSDT, XRPUSDT]
  ranked: premium_index_close
  # Two features above this absolute rank correlation carry one ordering, so a linear
  # model cannot separate their contributions. F5 cuts the redundancy tree here.
  redundancy_cut: 0.7
  windows:
    premium_momentum: {8h: 1, 24h: 3, 72h: 9, 168h: 21, 336h: 42, 720h: 90}
    premium_volatility: {24h: 3, 72h: 9, 168h: 21, 336h: 42}
    premium_zscore: {7d: 21, 14d: 42}
    premium_dev_mean: {7d: 21, 14d: 42}
    premium_quantile: {7d: 21, 14d: 42, 30d: 90}
    premium_rsi: {24h: 3, 72h: 9}
    price_volatility: {7d: 21, 14d: 42}
    premium_persistence: {7d: 21}
    premium_regime: {72h: 9}
    funding_zscore: {14d: 42}
    funding_half_life: {14d: 42}
    funding_change: {24h: 3}
    funding_cashflow: {7d: 7}
  # Bounds that shape an emitted value rather than guard a denominator. The
  # z-score clip holds a settlement-day outlier off the scale a model reads;
  # the AR(1) clip keeps the half-life finite at a unit root.
  clip:
    zscore: 10.0
    vol_ratio: 10.0
    ar1: 0.999
    half_life: [0.5, 100.0]
  families:
    - name: carry
      pattern: funding_rate|funding_rate_*|cum_positive_funding_7d|funding_half_life_14d|premium_level|premium_rank|premium_zscore_*
      role: signal
      hypothesis: A perpetual whose holders are paying to stay long is crowded, and the crowding unwinds.
      inputs: official funding settlements, premium index close
      lookback: 43
      lag: 0
      frame: per symbol, except the premium percentile which is within the decision timestamp
      representation: level, trailing z-score, cross-sectional percentile, mean-reversion speed
      failure_mode: Funding is clamped by the exchange, so the level saturates in the regimes that matter most.
    - name: mean_reversion
      pattern: premium_dev_mean_*|premium_quantile_pos_*|premium_persistence_*
      role: signal
      hypothesis: A premium far from its own recent range reverts faster than one near the middle of it.
      inputs: premium index close
      lookback: 90
      lag: 0
      frame: per symbol
      representation: deviation from a trailing mean, rolling percentile, sign persistence
      failure_mode: A trending premium looks extreme against its own window for as long as the trend lasts.
    - name: momentum
      pattern: premium_change_*|premium_accel_*
      role: signal
      hypothesis: A premium that has been widening keeps widening over the next settlement or two.
      inputs: premium index close
      lookback: 90
      lag: 0
      frame: per symbol
      representation: differences at six horizons, plus differences between horizons
      failure_mode: Momentum and mean reversion read the same series with opposite signs and cancel.
    - name: volatility
      pattern: premium_vol_*|price_vol_*|vol_ratio_*
      role: state
      hypothesis: Liquidation cascades widen both the premium and the price, and the two carry different information.
      inputs: premium index close, perpetual close
      lookback: 43
      lag: 0
      frame: per symbol
      representation: trailing dispersion at four horizons, plus short-over-long ratios
      failure_mode: A ratio of two dispersions is unstable when the denominator window is quiet.
    - name: cross_sectional
      pattern: premium_vs_median|premium_xs_zscore|xs_funding_dispersion
      role: state
      hypothesis: Whether a premium is high depends on what the rest of the universe is paying that settlement.
      inputs: premium index close, official funding settlements
      lookback: 1
      lag: 0
      frame: within the decision timestamp
      representation: distance from the cross-sectional median, cross-sectional z-score, dispersion
      failure_mode: The panel is unbalanced, so early dates rank against a handful of symbols.
    - name: regime
      pattern: premium_regime_*|premium_rsi_*|funding_session|cost_tier_alt
      role: state
      hypothesis: A sustained premium and the settlement slot condition how any signal should be read.
      inputs: premium index close, symbol, decision timestamp
      lookback: 10
      lag: 0
      frame: per symbol, except the session which is a property of the timestamp
      representation: signed regime average, bounded oscillator, categorical slot and fee tier
      failure_mode: The fee tier is a fixed list, so it does not follow a symbol across a tier change.

# What `04_model_based_features` decides, declared here for the same reason the
# feature windows above are. Every count is in 8-hour settlement bars, the unit
# `features.bar_hours` sets and `features.windows` already uses.
model_based:
  # Trailing settlements a series needs before either model is fitted on it. 500
  # settlements is about five and a half months; below that the leverage term of a
  # GJR recursion and the transition matrix of a two-state chain are estimated off
  # too few regime switches to mean anything.
  #
  # This is a burn-in, not a fold-entry condition. Both models below are fitted on a
  # schedule that runs over the whole history: the first 500 settlements carry no
  # value, and from there the parameters are re-estimated on the cadence each model
  # declares, always on settlements strictly earlier than the ones they then speak
  # for. Nothing about a cross-validation fold enters the fit, so the artifact
  # carries no fold column and a settlement has one value whichever fold selects it.
  min_train_bars: 500
  garch:
    # How often the variance model is re-estimated. 21 settlements is a week. A
    # variance model tracks a level that moves, which is the property the feature
    # exists to report, so it is refreshed faster than the regime model below.
    refit_every: 21
    # The conditional-volatility z-score compares a symbol's current forecast
    # against its own recent level: 90 bars is 30 days.
    vol_zscore_window: 90
    # Bound on the emitted z-score, so one liquidation cascade does not set the
    # scale a model reads. Same role as `features.clip.zscore`.
    zscore_clip: 10.0
  hmm:
    # Calm funding and stressed funding.
    n_states: 2
    # How often the chain is re-estimated. 63 settlements is three weeks. Regime
    # parameters are the slowest-moving thing this notebook fits, and each estimate
    # costs `n_restarts` expectation-maximization searches over the whole history.
    refit_every: 63
    # Expectation-maximization reaches a local optimum, so the fit is repeated
    # from this many starting points and the highest training likelihood is kept.
    # Measured on etfs' panel through the same estimator, ten restarts give ten
    # distinct log-likelihoods, so the search explores rather than repeating itself.
    n_restarts: 10
    # Convergence threshold on the per-settlement log-likelihood gain. Tighter than
    # hmmlearn's 1e-2 default, which is what this notebook has always fitted at and
    # what fx_pairs chose independently. Declared rather than baked into the shared
    # estimator: measured on etfs, the two tolerances move the emitted probability by
    # about 1e-3 on average and flip none of 756 state assignments, so the looser
    # default is not wrong - it is just not what this notebook fits at.
    tol: 1.0e-4
  incremental_ic:
    # A settlement needs this many symbols quoting before its cross-sectional rank
    # correlation is used, and a feature needs this many usable settlements before
    # its mean is reported.
    min_cross_section: 10
    min_decision_times: 20

modeling:
  gbm:
    libraries: [lightgbm]
    preset: default
    device: cpu
    # LightGBM's own CPU default. 63 is the GPU default and was carried over with the
    # device when these runs moved off the GPU, so every CPU fit was quantizing the design
    # matrix into a quarter of the bins the library would have used. Coarser bins are
    # faster and lose split points; the reader running this on a CPU gets what the
    # documentation describes.
    max_bin: 255

causal:
  treatment: premium_zscore_14d
  confounders: [price_vol_14d, funding_rate, premium_dev_mean_14d]
  method: walk_forward_dml

```

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.