跳至正文
返回文库全部文档

用双重机器学习估计加密货币永续合约溢价效应

笔记本 《交易机器学习》

总结

本笔记本估计,在干预条件下,加密货币永续合约溢价z分数是否与随后八小时收益相关,同时调整价格波动率、资金费率及溢价偏离近期均值的程度。双重机器学习分别根据这些混杂因素预测结果和处理变量,再分析二者的残差。交叉拟合会将每个观测排除在生成其残差的模型之外,而按时间顺序划分的折和隔离期则防止干扰模型使用未来数据。

本笔记本将识别假设视为核心:估计结果仅针对列出的三个混杂因素进行了调整,无法处理遗漏原因、反向因果或测量误差。文中规定了一种旨在保留序列依赖的分块置换安慰剂检验,块大小依据处理变量构造或收益期限中较长的依赖窗口确定。由于干预效应估计不是交易信号,因果结果与预测模型结果分别登记。本文介绍了方法和证伪设计,但所提供的文本没有给出估计值或安慰剂检验结果;因此,任何因果解释都取决于这里尚未展示的诊断结果和假设。

核心观点

  • 在拟合之前先定义目标估计量中的处理变量、结果、期限和混杂因素。
  • 双重机器学习先根据已声明的混杂因素对处理变量和结果进行残差化,再估计二者之间的关系。
  • 交叉拟合、按时间划分的折和隔离期可减少干扰预测中的过拟合和时间泄漏。
  • 分块置换安慰剂检验应保留序列依赖,而非将单个观测独立打乱。
  • 因果效应估计不能证明预测交易价值,且调整无法处理未列出的混杂因素。

标签

全文
# A different question from the one every other model here asks


# A different question from the one every other model here asks

Every notebook from [`06_linear`](06_linear.ipynb) through [`10_dl_tcn`](10_dl_tcn.ipynb) asks
the same question in different ways: **given what I can see now, what is the best guess for the
next 8-hour return?** A model that answers it well is useful whether or not any of its features
cause anything. Collinear features, proxies, coincidences - all fine, as long as the association
holds out of sample.

This notebook asks something else: **if the premium's z-score were higher, holding the other
declared drivers fixed, would the subsequent return be different?** That is a question about an
intervention, and no amount of predictive accuracy answers it. The two live in the same case
study and must not be reported as though they were the same finding, which is why the causal
result is registered under its own identity and never enters the population that
[`13_backtest`](13_backtest.ipynb) selects from.

## The estimand, stated before anything is fitted

- **Treatment:** `premium_zscore_14d`, the premium's standing relative to its own recent range.
  Continuous, not binary - so the estimate is a slope, the change in expected return per unit of
  treatment, not a difference between two groups.
- **Outcome:** `fwd_ret_8h`, the return over the settlement interval after the decision time.
- **Confounders:** `price_vol_14d`, `funding_rate`, `premium_dev_mean_14d`. These are the
  variables declared to drive both the treatment and the outcome, and adjusting for them is the
  entire identification claim.

That list is short, and its shortness is the honest part of the exercise. **Double machine
learning removes confounding by variables you name.** It does nothing about one you did not, and
nothing about the possibility that the relationship runs the other way. Naming three confounders
is a claim that those three are the relevant ones; the estimate is only as good as that claim.

## What double machine learning actually does

The naive approach - regress the outcome on the treatment and the confounders together - biases
the treatment coefficient whenever the confounders enter nonlinearly, because whatever the linear
term fails to absorb leaks into the treatment. DML avoids that by splitting the problem in two:

1. Predict the **outcome** from the confounders alone, and take the residual.
2. Predict the **treatment** from the confounders alone, and take the residual.
3. Regress residual on residual. What remains is the part of the treatment the confounders do not
  explain, against the part of the outcome they do not explain.

The two nuisance predictions can be any flexible learner, because the residual-on-residual step
is what carries the causal interpretation. **Cross-fitting** is what keeps that step honest: each
observation's residual is computed by a model that did not see it, so an overfitted nuisance
model cannot manufacture a residual correlation. Here the folds are chronological with an
embargo, so the nuisance models are also never fitted on data that comes after what they predict.

## Why the placebo is the part to read carefully

A causal estimate on financial data will produce a number whether or not there is anything there,
so the refutation matters more than the point estimate. The placebo permutes the treatment and
re-estimates, and the estimate should collapse. What makes it a real test rather than a formality
is the **block size**, discussed at the table below: permuting one bar at a time destroys exactly
the serial dependence that makes the original estimate hard to get right, and a placebo that
destroys the difficulty is a test the estimate passes for free.

**Learning objectives.** By the end of this notebook you will be able to:

- State an estimand - treatment, outcome, confounders, horizon - precisely enough that someone
  else could disagree with it.
- Explain what cross-fitting protects against, and why chronological folds with an embargo are
  the right shape for it on overlapping financial labels.
- Say why a block-permutation placebo needs a block long enough to preserve serial dependence,
  and how the block size is derived here rather than chosen.
- Keep a causal estimate and a predictive result in separate reports, and say in one sentence
  why combining them would misdescribe both.

**Book reference:** Chapter 15, causal inference for trading research.

**Prerequisites:** [`03_financial_features`](03_financial_features.ipynb),
[`02_labels`](02_labels.ipynb) and [`05_evaluation`](05_evaluation.ipynb) - the features, the
return labels and the purged walk-forward folds.

**What it writes:** one causal result under its own identity in `run_log/registry.db`. It is read
by [`12_model_analysis`](12_model_analysis.ipynb), which reports it **beside** the predictive
results rather than among them.

```python
import os

import polars as pl

from case_studies.crypto_perps_funding.research_workflow import open_study
from case_studies.research import causal_supersedes
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE = os.environ.get("ML4T_OUTPUT_DIR", "")
LABEL = "fwd_ret_8h"
CONFIG_NAME = "dml"
PREVIEW_REDUCTIONS = {}
OVERRIDES = {}
# The causal identity this run retires. A reader's clone has nothing to retire - `run_log/` is
# not shipped, so a first run meets an empty registry - and `causal_supersedes` withholds the
# declaration against a registry that does not hold it, which is why one committed value is
# right for both. A predecessor named against a registry that *does* hold rows but not this one
# is rejected at the registering write, after the DML fit and every placebo refit are paid for.
#
# The source component of a causal identity is `CAUSAL_RUNNER_VERSION`, a declared integer in
# `case_studies/utils/causal.py`; nothing hashes the file. So an edit to the estimator moves no
# identity unless that constant is raised by hand, and an edit that changes a registered value
# without raising it leaves the next run to hit the cache and serve the old number under new
# code. Once it is raised, `CausalResult.one` resolves a label to exactly one canonical
# identity, so a re-run that does not name what it replaces leaves two live and the next
# notebook fails with "resolved to 2 identities".
# Retired by this run: the block-permutation refutation now compares the HAC t-statistic
# rather than the raw effect, so CAUSAL_RUNNER_VERSION moved and every causal identity with
# it. The rows named here hold a p-value computed on the shrunken placebo effects; this run
# supersedes them rather than correcting them, because the statistic is different, not the
# arithmetic. Read out of each registry's current canonical identity per label, 2026-09-10.
SUPERSEDES_CAUSAL: str = "025f6c2f4e3b"
```

## 1. Resolve the estimand and the refutation contract

Nothing is fitted below. The request resolves to a specification and an identity, and the table
prints the fields that decide what the estimate means: what is being intervened on, what is being
measured, over what horizon, how the nuisance folds are cut, and how the placebo will be built.
Reading them before the fit is the only point at which disagreeing with the estimand is cheap.

`eligible_rows` is the analysis population - the rows that survive having a treatment, an
outcome and every confounder present. Quote that, not the panel height, when describing what the
estimate rests on.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
request = study.causal(
    method="dml",
    label=LABEL,
    config_name=CONFIG_NAME,
    execution_tier=EXECUTION_TIER,
    preview_reductions=PREVIEW_REDUCTIONS,
    overrides=OVERRIDES,
    supersedes=causal_supersedes(
        study, SUPERSEDES_CAUSAL, LABEL, labels=[LABEL], execution_tier=EXECUTION_TIER
    ),
)
resolved = request.resolve()
computation = resolved.spec["computation"]
pl.DataFrame(
    {
        "causal_hash": [resolved.identity],
        "outcome": [computation["estimand"]["outcome"]],
        "treatment": [computation["estimand"]["treatment"]],
        "outcome_horizon": [computation["estimand"]["outcome_horizon"]],
        "n_folds": [computation["cv"]["n_folds"]],
        "embargo_periods": [computation["cv"]["embargo_periods"]],
        "block_size": [computation["refutation"]["block_size"]],
        "block_size_basis": [computation["refutation"]["block_size_basis"]],
        "gap_policy": [computation["refutation"]["temporal_gap_policy"]],
        "eligible_rows": [computation["analysis_population"]["n_rows"]],
    }
)
```

`block_size` is the parameter the refutation lives or dies on, so it is on the
table rather than buried in the spec. The placebo permutes contiguous blocks
within each symbol; a block of one bar is an iid shuffle, which destroys the
serial dependence the placebo is meant to keep and makes the test trivially
easy to pass. Two things create that dependence and the block spans the longer
of them: the overlapping labels span the outcome horizon, and the treatment
spans its own construction window. Here the horizon is a single 8-hour bar
while `premium_zscore_14d` is a 42-bar rolling statistic, so the treatment
window sets the block and `block_size_basis` says so.

## 2. Execute, and register the result separately

The fit runs the nuisance models across the chronological folds, forms the residual-on-residual
estimate, and then pays for the placebo refits - which is where most of the cost is, since the
whole procedure is repeated once per placebo draw.

The check below refuses a result whose *computation* is not the one that was resolved. That is not
defensive coding: the computation carries the estimand, the fold geometry and the refutation
contract, and a result carrying a different one describes a contract that no longer exists. The
notebook downstream would then resolve the label to two live identities and stop.

It compares `spec["computation"]` and not the whole spec, because the spec also carries
`provenance` - the git commit, the platform and the package versions of the run that fitted. Those
record that run's own circumstances and are meant to differ from any later run that reads the row
back from cache. Comparing whole specs asserts that nothing has been committed since, which is not
a property of the estimate, and it makes the notebook raise on every re-run from a different
commit - the opposite of what this check is for. The second run of this notebook fits nothing and
has to say so rather than raise.

Comparing `result.hash` against `resolved.identity` would not do it either: both are
`training_hash_from_spec` of the same resolved spec, so that comparison cannot fail and would
assert nothing. Same check, same reasoning, as `etfs/12_causal_dml`.

```python
result = resolved.run()
if not result.complete:
    raise RuntimeError("causal execution is incomplete")
if result.spec["computation"] != resolved.spec["computation"]:
    raise RuntimeError("the registered causal computation differs from the resolved one")
pl.DataFrame(
    {
        "causal_hash": [result.hash],
        "n_obs": [result.metrics["n_obs"]],
        "complete": [result.complete],
        "execution_tier": [result.execution_tier],
    }
)
```

## Key takeaways and limitations

- **This is not a trading signal and must not be reported as one.** The estimate answers what
  would happen under an intervention on the premium z-score. The backtest in
  [`13_backtest`](13_backtest.ipynb) answers what a strategy would have earned. A number from
  here quoted as evidence for a strategy is a category error, and the separate registration
  exists to make that mistake require deliberate effort.
- **Adjustment covers three named confounders and nothing else.** DML removes bias from variables
  it is given. An omitted common cause, reverse causation from returns to the premium, or a
  confounder measured with error all survive it untouched, and no diagnostic in this notebook
  would reveal them. The identifying assumption is an argument, not an output.
- **The refutation is the load-bearing part.** A point estimate arrives whether or not there is
  anything to estimate. What distinguishes the two is whether the placebo collapses, and whether
  the placebo was built to be hard - which is what the block size decides.
- **The identity hashes the estimator's source.** Any edit to `case_studies/utils/causal.py`
  produces a different causal identity, whether or not it changes a number. That is deliberate: a
  causal claim is a claim about a procedure, and a changed procedure is a new claim until someone
  shows the estimate did not move.
- **A preview run is not a canonical result.** The reduced sample and fold counts a preview
  applies are recorded in its identity and it is barred from the canonical population, so a cheap
  run can never be mistaken for the published estimate.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。