본문으로 건너뛰기
라이브러리 문서 전체

CME 선물 캐리 효과를 평가하는 이중 머신러닝

노트북 Machine Learning for Trading

요약

이 노트북은 변동성, 모멘텀, 횡단면 캐리 순위를 조정한 뒤 선물 캐리가 이후 수익률에 영향을 미치는지 이중 머신러닝으로 추정합니다. 인과적 설명과 예측 성과를 구분합니다. 수익을 낸 백테스트는 캐리가 수익률을 예측한다는 점을 보여줄 수 있지만, 캐리 자체가 수익률의 원인이라고 입증하지는 못합니다. 유연한 방해변수 모델이 교란변수로 처치와 결과를 예측하고, 표본 내 편향을 줄이기 위해 교차 적합을 사용합니다. 변동성 군집과 겹치는 수익률 라벨 때문에 잔차가 종속되므로 이분산성과 자기상관에 강건한 불확실성 추정치를 사용합니다.

반증을 위해 결과 라벨과 처치 구성에서 생기는 종속성을 보존할 수 있도록 연속 블록 단위로 처치를 섞습니다. 캐리는 같은 시점의 가격으로 계산하므로 명시된 구성 기간은 짧지만, 실제 지속성은 여전히 남아 있습니다. 문서는 더 짧은 플라시보 블록 길이에서 상당한 자기상관이 나타나 반증 검정이 완벽하지 않다고 보고합니다. 추정치는 지정된 결과와 조정 조건에 관한 것이지 전략 선택에 관한 것이 아닙니다. 트레이딩 성과를 입증하지도 반증하지도 않습니다. 결과는 선언된 교란변수와 시점에도 좌우되며, 누락된 교란변수로 추정 대상이 암묵적으로 바뀌어서는 안 됩니다.

핵심 아이디어

  • 캐리와 수익률 간 예측 관계만으로 인과 효과가 입증되지는 않습니다.
  • 이중 머신러닝은 교차 적합을 통해 지정된 교란변수로 처치와 결과를 각각 회귀한 잔차를 구합니다.
  • 변동성 군집과 중첩된 결과로 관측치가 종속될 때는 HAC 불확실성 추정치를 사용합니다.
  • 플라시보 블록은 라벨 중첩과 처치 구성 기간 모두에서 발생하는 종속성을 반영해야 합니다.
  • 인과 추정은 표본 외 전략 성과와 다른 질문에 답합니다.

태그

전문
# CME Futures: Double Machine Learning


# CME Futures: Double Machine Learning

This notebook estimates the configured treatment effect for each return horizon. The request
resolves the outcome, treatment, confounders, timing, embargo, nuisance models, and placebo design
before fitting. Missing confounders are an error rather than an invitation to change the estimand.

Nuisance models train on complete timestamp panels before the holdout. HAC uncertainty uses the
declared temporal ordering, and contiguous within-product blocks define the placebo refutation.
`12_model_analysis` reports the causal diagnostics. Trading configuration selection uses
validation backtest Sharpe.

Nothing estimated here competes for that selection. A causal result is not a candidate
configuration: it produces no prediction set, is not eligible for a backtest, and never enters
the funnel. A reader looking for it on the leaderboard will not find it, and that absence is
the design rather than an omission - the two stages answer questions that do not compare.

One consequence for how the result is read: a small or insignificant effect here does not
invalidate the trading strategy, and a large one does not endorse it. The strategy stands or
falls on out-of-sample Sharpe either way. What the estimate changes is what the chapter may
claim about *why* it works.

Prerequisites: `03_financial_features`, `04_model_based_features`, and `05_evaluation`.

## The question, and why this case study in particular has to ask it

The treatment here is `carry_pct` - the term-structure spread between the front contract and
the next one. That is not an arbitrary choice of variable to interrogate. **Carry is the
signal this entire case study trades.** Every model in stages 06 through 10 is, in one form or
another, learning to predict returns from carry and things built out of carry.

So the question this notebook asks is whether the premise underneath all of that holds: does
carry itself move subsequent returns, or does carry merely travel alongside something else
that does? A predictive model does not care - it profits either way, as long as the
association persists. The chapter's explanation of *why* the strategy works cares a great
deal, because an explanation built on the wrong mechanism is wrong even when the strategy is
profitable.

This is also why the answer cannot come from the backtest. A backtest that earns a positive
Sharpe from carry demonstrates that carry predicted returns over that history. It says nothing
about whether carry was the cause, and no amount of out-of-sample performance converts one
into the other.

### What is being adjusted for, and why each one is a candidate explanation

Double machine learning estimates the effect of the treatment after flexibly removing what the
confounders explain of both the treatment and the outcome. It does this by predicting each
from the confounders alone, using models that need not be linear in anything, and then
estimating the effect from what is left over in each. The predictions are cross-fitted, so the
model that residualizes an observation was never fitted on it.

The three confounders are the three rival explanations worth taking seriously here:

- **`vol_21d`.** Contracts in backwardation are frequently contracts under supply stress, and
  stress raises volatility. If volatile contracts earn more simply for being volatile, carry
  would look rewarded without being the reason.
- **`momentum_composite`.** Carry and momentum are correlated in futures - a contract in
  sustained backwardation has often been rising. An unadjusted carry effect could be a
  momentum effect wearing carry's label.
- **`carry_rank`.** This is the subtle one, and it is the reason the confounder list is not
  just "the other features". `carry_rank` is a product's position in the cross-section, while
  `carry_pct` is its level. Those come apart: a contract can have a high carry level in a
  period when every contract does, and rank in the middle of its peers. Adjusting for the rank
  asks whether the carry *level* carries information beyond where the product sits relative to
  the others - which is precisely what a cross-sectional strategy is already exploiting.

### Why the standard error is HAC

The uncertainty around the effect is computed with a heteroskedasticity- and
autocorrelation-consistent estimator, because the ordinary one assumes the residuals are
independent and identically scattered and neither holds on a futures panel. Volatility
clusters, so the scatter is larger in some periods than others; and overlapping outcomes tie
neighbouring residuals together, so they carry information about each other. Both inflate the
effective information in the sample relative to what an ordinary standard error assumes, and
the result is a confidence interval that is too narrow and a t-statistic that is too large.
The HAC estimator does not fix the estimate - it corrects what may be claimed about it.

### Why the label buffer sizes the placebo block here, and not the treatment

The refutation permutes the treatment in contiguous blocks and re-runs the whole estimation,
building a distribution of effects under the hypothesis that the treatment does nothing. The
block has to be long enough to preserve whatever serial dependence the data has, or the
placebo is a weaker opponent than the truth and every p-value looks significant.

Two separate scales create that dependence, and the block spans the longer of them:
`block_size = max(label_buffer, treatment_window)`. The overlapping labels span the label
horizon, and the treatment's own construction window spans itself.

`causal.treatment_window` is 1 for this case study, and the reason is a property of the
construction rather than a judgement. `carry_pct` is
`(c0_price - c1_price) / c0_price * 12`, computed from two prices at the same timestamp.
Nothing rolls, nothing averages, no window is spanned. The label buffer is therefore the
binding scale here: the registered blocks are 5 periods for `fwd_ret_5d` and 21 for
`fwd_ret_21d`, and each row records its own `block_size` beside a `block_size_basis` of
`label_buffer` saying which of the two set it.

**A one-bar construction window is not a claim that the column is serially independent**, and
the two are worth keeping apart. The declared window says how many bars the formula reads;
the column's empirical persistence is a separate quantity and is much larger here.
`12_model_analysis` measures it from the feature panel rather than asserting it, on one row per
product-session because that is the unit the block counts, pooling the autocorrelation within
product with each product demeaned first. Compare each block against the autocorrelation at
that lag: 0.52 at the 5 sessions used for `fwd_ret_5d`, 0.14 at the 21 used for
`fwd_ret_21d`, and indistinguishable from zero by lag 63. The 5-session block leaves real
dependence unpreserved and narrows that label's placebo distribution by some amount neither
notebook quantifies; the 21-session block spans most of it. The narrowing bears mainly on one
of the two refutations here, not equally on both.

The contrast with a rolling treatment is worth holding onto, because it is where this is
usually got wrong. A treatment built from a 252-session window overlaps its neighbours in 251
of them, and a block sized by a 21-day label buffer would shred that overlap; the placebo
distribution narrows, and the p-value collapses toward zero whether or not the effect is
real. That is the case the `max` exists for. The number is declared in `setup.yaml` and
derived from how the column is built, because inferring it from a window list would put a
wrong number behind a right-looking one.

```python
"""Fit the declared CME futures double-machine-learning requests."""

import polars as pl

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    open_study,
    product_universe_table,
)
from case_studies.research import causal_supersedes
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_REDUCTIONS: dict = {}
# Retired by this run: the block-permutation refutation now compares the HAC t-statistic
# rather than the raw effect, so CAUSAL_RUNNER_VERSION moved and every causal identity with
# it. The rows named here hold a p-value computed on the shrunken placebo effects; this run
# supersedes them rather than correcting them, because the statistic is different, not the
# arithmetic. Read out of each registry's current canonical identity per label, 2026-09-10.
SUPERSEDES_CAUSAL: str = '{"fwd_ret_5d": "4abb82b8141c", "fwd_ret_21d": "ac5c1a480d24"}'
```

## Resolve the estimands

Preview row and fold limits must be named in `PREVIEW_REDUCTIONS`. They enter the computation
identity and cannot be mistaken for canonical causal estimates.

**Missing confounders are an error rather than an invitation to change the estimand**, which
is worth stating as a rule because the alternative is so tempting. Dropping a confounder that
failed to resolve leaves a request that still runs and still returns a number - a number
answering a different question, with no adjustment for the thing that went missing, and
nothing in the result marking it. An estimand is the whole specification: outcome, treatment,
confounder set, timing and embargo together. Changing any part of it produces a different
quantity, not a degraded version of the same one.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "preview" and (WORKSPACE is None or not PREVIEW_REDUCTIONS):
    raise ValueError("preview execution requires WORKSPACE and PREVIEW_REDUCTIONS")
universe = product_universe_table()
universe
```

```python
requests = tuple(
    study.causal(
        method="dml",
        label=label,
        execution_tier=EXECUTION_TIER,
        preview_reductions=PREVIEW_REDUCTIONS,
        supersedes=causal_supersedes(
            study,
            SUPERSEDES_CAUSAL,
            label,
            labels=list(ALL_LABELS),
            execution_tier=EXECUTION_TIER,
        ),
    ).resolve()
    for label in ALL_LABELS
)
request_rows = []
for request in requests:
    computation = request.spec["computation"]
    estimand = computation["estimand"]
    population = computation["analysis_population"]
    request_rows.append(
        {
            "label": request.spec["label"],
            "treatment": estimand["treatment"],
            "outcome": estimand["outcome"],
            "outcome_horizon": estimand["outcome_horizon"],
            "confounders": ", ".join(estimand["confounders"]),
            "analysis_rows": population["n_rows"],
            "analysis_timestamps": population["n_timestamps"],
            "request_hash": request.identity,
        }
    )
request_table = pl.DataFrame(request_rows)
request_table
```

## Execute and verify restart

The shared causal runner uses deterministic seeds and persists one immutable result per resolved
request. Reopening the same request must return the same identity and a complete result.

The restart check runs the request twice and requires the same hash. That is cheap here
because the second call is served from the registry, and it is worth doing because a causal
estimate has no external reference to be checked against. A predictive model can be caught by
out-of-sample performance; an effect estimate cannot, so reproducibility is most of what is
available. An estimate that moved between two runs of the same specification would mean the
seed does not control everything the fit depends on, and the number would be one draw from a
distribution nobody characterised.

`SUPERSEDES_CAUSAL` above is how a re-run names the identity it retires, per label. A refit
under a changed specification produces a second current identity for the same label, and
`CausalResult.one` resolves a label to exactly one - so the run has to say which it replaces,
and the retired estimate stays in the registry rather than being deleted.

```python
results = []
for request in requests:
    label = request.spec["label"]
    result = request.run()
    restarted = request.run()
    if not result.complete or restarted.hash != result.hash:
        raise RuntimeError(f"causal request for {label} did not persist completely")
    results.append(
        {
            "label": label,
            "execution_tier": result.execution_tier,
            "complete": result.complete,
            "causal_hash": result.hash,
        }
    )
```

```python
pl.DataFrame(results).sort("label")
```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.