Aprendizado de máquina duplo para testar efeitos de carry em futuros
Resumo
Este notebook usa aprendizado de máquina duplo para estimar se o carry de futuros prevê retornos subsequentes independentemente da volatilidade, do momentum e da classificação transversal de carry. Ele calcula os resíduos de carry e retornos em relação a esses fatores de confusão com modelos flexíveis e cross-fitting, e então estima o efeito a partir dos resíduos. O resultado é avaliado em dois horizontes de retorno. A incerteza HAC considera a volatilidade variável e os resultados sobrepostos, enquanto um teste de refutação por permutação em blocos compara a estatística HAC observada com estimativas após reorganizar os blocos de tratamento.
O desenho distingue o valor preditivo do carry das evidências de que o carry causa retornos: um backtest lucrativo não demonstra o mecanismo, e a estimativa causal não seleciona nem valida uma configuração de trading. O notebook informa que a fórmula do tratamento usa preços em um único horário, enquanto a persistência do tratamento pode se estender além da janela usada para construí-lo. Portanto, o bloco placebo mais curto deixa de preservar parte da dependência serial; o mais longo abrange a maior parte dela. Os resultados também dependem dos fatores de confusão declarados e do momento, portanto remover um fator de confusão ausente mudaria o estimando, em vez de apenas enfraquecer a mesma análise.
Ideias principais
- O aprendizado de máquina duplo estima um efeito de tratamento após ajustar de forma flexível o tratamento e o resultado pelos fatores de confusão especificados.
- O cross-fitting garante que cada observação tenha seus resíduos calculados por modelos auxiliares que não foram treinados com ela.
- A incerteza HAC é usada porque a volatilidade varia ao longo do tempo e rótulos sobrepostos geram resíduos dependentes.
- Os blocos placebo devem refletir tanto a sobreposição de rótulos quanto a janela de construção do tratamento, embora a persistência observada possa se estender além dela.
- Estimativas causais respondem a uma questão sobre mecanismos; elas não escolhem configurações de trading nem substituem a avaliação da estratégia fora da amostra.
Tags
Texto completo
# 11_causal_dml.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # CME Futures: Double Machine Learning
#
# This notebook estimates the configured treatment effect for each return horizon. The request
# resolves the outcome, treatment, confounders, timing, embargo, nuisance models, and placebo design
# before fitting. Missing confounders are an error rather than an invitation to change the estimand.
#
# Nuisance models train on complete timestamp panels before the holdout. HAC uncertainty uses the
# declared temporal ordering, and contiguous within-product blocks define the placebo refutation.
# `12_model_analysis` reports the causal diagnostics. Trading configuration selection uses
# validation backtest Sharpe.
#
# Nothing estimated here competes for that selection. A causal result is not a candidate
# configuration: it produces no prediction set, is not eligible for a backtest, and never enters
# the funnel. A reader looking for it on the leaderboard will not find it, and that absence is
# the design rather than an omission - the two stages answer questions that do not compare.
#
# One consequence for how the result is read: a small or insignificant effect here does not
# invalidate the trading strategy, and a large one does not endorse it. The strategy stands or
# falls on out-of-sample Sharpe either way. What the estimate changes is what the chapter may
# claim about *why* it works.
#
# Prerequisites: `03_financial_features`, `04_model_based_features`, and `05_evaluation`.
# %% [markdown]
# ## The question, and why this case study in particular has to ask it
#
# The treatment here is `carry_pct` - the term-structure spread between the front contract and
# the next one. That is not an arbitrary choice of variable to interrogate. **Carry is the
# signal this entire case study trades.** Every model in stages 06 through 10 is, in one form or
# another, learning to predict returns from carry and things built out of carry.
#
# So the question this notebook asks is whether the premise underneath all of that holds: does
# carry itself move subsequent returns, or does carry merely travel alongside something else
# that does? A predictive model does not care - it profits either way, as long as the
# association persists. The chapter's explanation of *why* the strategy works cares a great
# deal, because an explanation built on the wrong mechanism is wrong even when the strategy is
# profitable.
#
# This is also why the answer cannot come from the backtest. A backtest that earns a positive
# Sharpe from carry demonstrates that carry predicted returns over that history. It says nothing
# about whether carry was the cause, and no amount of out-of-sample performance converts one
# into the other.
#
# ### What is being adjusted for, and why each one is a candidate explanation
#
# Double machine learning estimates the effect of the treatment after flexibly removing what the
# confounders explain of both the treatment and the outcome. It does this by predicting each
# from the confounders alone, using models that need not be linear in anything, and then
# estimating the effect from what is left over in each. The predictions are cross-fitted, so the
# model that residualizes an observation was never fitted on it.
#
# The three confounders are the three rival explanations worth taking seriously here:
#
# - **`vol_21d`.** Contracts in backwardation are frequently contracts under supply stress, and
# stress raises volatility. If volatile contracts earn more simply for being volatile, carry
# would look rewarded without being the reason.
# - **`momentum_composite`.** Carry and momentum are correlated in futures - a contract in
# sustained backwardation has often been rising. An unadjusted carry effect could be a
# momentum effect wearing carry's label.
# - **`carry_rank`.** This is the subtle one, and it is the reason the confounder list is not
# just "the other features". `carry_rank` is a product's position in the cross-section, while
# `carry_pct` is its level. Those come apart: a contract can have a high carry level in a
# period when every contract does, and rank in the middle of its peers. Adjusting for the rank
# asks whether the carry *level* carries information beyond where the product sits relative to
# the others - which is precisely what a cross-sectional strategy is already exploiting.
#
# ### Why the standard error is HAC
#
# The uncertainty around the effect is computed with a heteroskedasticity- and
# autocorrelation-consistent estimator, because the ordinary one assumes the residuals are
# independent and identically scattered and neither holds on a futures panel. Volatility
# clusters, so the scatter is larger in some periods than others; and overlapping outcomes tie
# neighbouring residuals together, so they carry information about each other. Both inflate the
# effective information in the sample relative to what an ordinary standard error assumes, and
# the result is a confidence interval that is too narrow and a t-statistic that is too large.
# The HAC estimator does not fix the estimate - it corrects what may be claimed about it.
#
# ### Why the label buffer sizes the placebo block here, and not the treatment
#
# The refutation permutes the treatment in contiguous blocks and re-runs the whole estimation,
# building a distribution of effects under the hypothesis that the treatment does nothing. The
# block has to be long enough to preserve whatever serial dependence the data has, or the
# placebo is a weaker opponent than the truth and every p-value looks significant.
#
# Two separate scales create that dependence, and the block spans the longer of them:
# `block_size = max(label_buffer, treatment_window)`. The overlapping labels span the label
# horizon, and the treatment's own construction window spans itself.
#
# `causal.treatment_window` is 1 for this case study, and the reason is a property of the
# construction rather than a judgement. `carry_pct` is
# `(c0_price - c1_price) / c0_price * 12`, computed from two prices at the same timestamp.
# Nothing rolls, nothing averages, no window is spanned. The label buffer is therefore the
# binding scale here: the registered blocks are 5 periods for `fwd_ret_5d` and 21 for
# `fwd_ret_21d`, and each row records its own `block_size` beside a `block_size_basis` of
# `label_buffer` saying which of the two set it.
#
# **A one-bar construction window is not a claim that the column is serially independent**, and
# the two are worth keeping apart. The declared window says how many bars the formula reads;
# the column's empirical persistence is a separate quantity and is much larger here.
# `12_model_analysis` measures it from the feature panel rather than asserting it, on one row per
# product-session because that is the unit the block counts, pooling the autocorrelation within
# product with each product demeaned first. Compare each block against the autocorrelation at
# that lag: 0.52 at the 5 sessions used for `fwd_ret_5d`, 0.14 at the 21 used for
# `fwd_ret_21d`, and indistinguishable from zero by lag 63. The 5-session block leaves real
# dependence unpreserved and narrows that label's placebo distribution by some amount neither
# notebook quantifies; the 21-session block spans most of it. The narrowing bears mainly on one
# of the two refutations here, not equally on both.
#
# The contrast with a rolling treatment is worth holding onto, because it is where this is
# usually got wrong. A treatment built from a 252-session window overlaps its neighbours in 251
# of them, and a block sized by a 21-day label buffer would shred that overlap; the placebo
# distribution narrows, and the p-value collapses toward zero whether or not the effect is
# real. That is the case the `max` exists for. The number is declared in `setup.yaml` and
# derived from how the column is built, because inferring it from a window list would put a
# wrong number behind a right-looking one.
# %%
"""Fit the declared CME futures double-machine-learning requests."""
import polars as pl
from case_studies.cme_futures.research_workflow import (
ALL_LABELS,
open_study,
product_universe_table,
)
from case_studies.research import causal_supersedes
# %% tags=["parameters"]
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_REDUCTIONS: dict = {}
# Retired by this run: the block-permutation refutation now compares the HAC t-statistic
# rather than the raw effect, so CAUSAL_RUNNER_VERSION moved and every causal identity with
# it. The rows named here hold a p-value computed on the shrunken placebo effects; this run
# supersedes them rather than correcting them, because the statistic is different, not the
# arithmetic. Read out of each registry's current canonical identity per label, 2026-09-10.
SUPERSEDES_CAUSAL: str = '{"fwd_ret_5d": "4abb82b8141c", "fwd_ret_21d": "ac5c1a480d24"}'
# %% [markdown]
# ## Resolve the estimands
#
# Preview row and fold limits must be named in `PREVIEW_REDUCTIONS`. They enter the computation
# identity and cannot be mistaken for canonical causal estimates.
#
# **Missing confounders are an error rather than an invitation to change the estimand**, which
# is worth stating as a rule because the alternative is so tempting. Dropping a confounder that
# failed to resolve leaves a request that still runs and still returns a number - a number
# answering a different question, with no adjustment for the thing that went missing, and
# nothing in the result marking it. An estimand is the whole specification: outcome, treatment,
# confounder set, timing and embargo together. Changing any part of it produces a different
# quantity, not a degraded version of the same one.
# %%
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "preview" and (WORKSPACE is None or not PREVIEW_REDUCTIONS):
raise ValueError("preview execution requires WORKSPACE and PREVIEW_REDUCTIONS")
universe = product_universe_table()
universe
# %%
requests = tuple(
study.causal(
method="dml",
label=label,
execution_tier=EXECUTION_TIER,
preview_reductions=PREVIEW_REDUCTIONS,
supersedes=causal_supersedes(
study,
SUPERSEDES_CAUSAL,
label,
labels=list(ALL_LABELS),
execution_tier=EXECUTION_TIER,
),
).resolve()
for label in ALL_LABELS
)
request_rows = []
for request in requests:
computation = request.spec["computation"]
estimand = computation["estimand"]
population = computation["analysis_population"]
request_rows.append(
{
"label": request.spec["label"],
"treatment": estimand["treatment"],
"outcome": estimand["outcome"],
"outcome_horizon": estimand["outcome_horizon"],
"confounders": ", ".join(estimand["confounders"]),
"analysis_rows": population["n_rows"],
"analysis_timestamps": population["n_timestamps"],
"request_hash": request.identity,
}
)
request_table = pl.DataFrame(request_rows)
request_table
# %% [markdown]
# ## Execute and verify restart
#
# The shared causal runner uses deterministic seeds and persists one immutable result per resolved
# request. Reopening the same request must return the same identity and a complete result.
#
# The restart check runs the request twice and requires the same hash. That is cheap here
# because the second call is served from the registry, and it is worth doing because a causal
# estimate has no external reference to be checked against. A predictive model can be caught by
# out-of-sample performance; an effect estimate cannot, so reproducibility is most of what is
# available. An estimate that moved between two runs of the same specification would mean the
# seed does not control everything the fit depends on, and the number would be one draw from a
# distribution nobody characterised.
#
# `SUPERSEDES_CAUSAL` above is how a re-run names the identity it retires, per label. A refit
# under a changed specification produces a second current identity for the same label, and
# `CausalResult.one` resolves a label to exactly one - so the run has to say which it replaces,
# and the retired estimate stays in the registry rather than being deleted.
# %%
results = []
for request in requests:
label = request.spec["label"]
result = request.run()
restarted = request.run()
if not result.complete or restarted.hash != result.hash:
raise RuntimeError(f"causal request for {label} did not persist completely")
results.append(
{
"label": label,
"execution_tier": result.execution_tier,
"complete": result.complete,
"causal_hash": result.hash,
}
)
# %% tags=["results"]
pl.DataFrame(results).sort("label")
```Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT
Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.