Double Machine Learning for a Crypto Perpetuals Premium Effect
Summary
This notebook estimates whether a crypto perpetuals premium z-score is associated with subsequent eight-hour returns under an intervention, adjusting for price volatility, funding rate, and the premium’s deviation from its recent mean. Double machine learning predicts the outcome and treatment separately from those confounders, then relates the residuals. Cross-fitting keeps each observation out of the models that produce its residual, while chronological folds and an embargo prevent nuisance models from using later data.
The notebook treats identification assumptions as central: the estimate only adjusts for the three named confounders and cannot address omitted causes, reverse causation, or measurement error. It specifies a block-permutation placebo designed to preserve serial dependence, with block size based on the longer dependence window from the treatment construction or return horizon. The causal result is registered separately from predictive model results because an intervention estimate is not a trading signal. The document describes the method and refutation design, but provides no estimate or placebo outcome in the supplied text; any causal interpretation therefore depends on diagnostics and assumptions not demonstrated here.
Key ideas
- The estimand defines the treatment, outcome, horizon, and confounders before fitting.
- Double machine learning residualizes both treatment and outcome against declared confounders before estimating their relationship.
- Cross-fitting and chronological folds with an embargo reduce overfitting and temporal leakage in nuisance predictions.
- A block-permutation placebo should preserve serial dependence rather than shuffle individual observations independently.
- A causal estimate does not establish predictive trading value, and adjustment cannot address unnamed confounders.
Tags
Full text
# A different question from the one every other model here asks
# A different question from the one every other model here asks
Every notebook from [`06_linear`](06_linear.ipynb) through [`10_dl_tcn`](10_dl_tcn.ipynb) asks
the same question in different ways: **given what I can see now, what is the best guess for the
next 8-hour return?** A model that answers it well is useful whether or not any of its features
cause anything. Collinear features, proxies, coincidences - all fine, as long as the association
holds out of sample.
This notebook asks something else: **if the premium's z-score were higher, holding the other
declared drivers fixed, would the subsequent return be different?** That is a question about an
intervention, and no amount of predictive accuracy answers it. The two live in the same case
study and must not be reported as though they were the same finding, which is why the causal
result is registered under its own identity and never enters the population that
[`13_backtest`](13_backtest.ipynb) selects from.
## The estimand, stated before anything is fitted
- **Treatment:** `premium_zscore_14d`, the premium's standing relative to its own recent range.
Continuous, not binary - so the estimate is a slope, the change in expected return per unit of
treatment, not a difference between two groups.
- **Outcome:** `fwd_ret_8h`, the return over the settlement interval after the decision time.
- **Confounders:** `price_vol_14d`, `funding_rate`, `premium_dev_mean_14d`. These are the
variables declared to drive both the treatment and the outcome, and adjusting for them is the
entire identification claim.
That list is short, and its shortness is the honest part of the exercise. **Double machine
learning removes confounding by variables you name.** It does nothing about one you did not, and
nothing about the possibility that the relationship runs the other way. Naming three confounders
is a claim that those three are the relevant ones; the estimate is only as good as that claim.
## What double machine learning actually does
The naive approach - regress the outcome on the treatment and the confounders together - biases
the treatment coefficient whenever the confounders enter nonlinearly, because whatever the linear
term fails to absorb leaks into the treatment. DML avoids that by splitting the problem in two:
1. Predict the **outcome** from the confounders alone, and take the residual.
2. Predict the **treatment** from the confounders alone, and take the residual.
3. Regress residual on residual. What remains is the part of the treatment the confounders do not
explain, against the part of the outcome they do not explain.
The two nuisance predictions can be any flexible learner, because the residual-on-residual step
is what carries the causal interpretation. **Cross-fitting** is what keeps that step honest: each
observation's residual is computed by a model that did not see it, so an overfitted nuisance
model cannot manufacture a residual correlation. Here the folds are chronological with an
embargo, so the nuisance models are also never fitted on data that comes after what they predict.
## Why the placebo is the part to read carefully
A causal estimate on financial data will produce a number whether or not there is anything there,
so the refutation matters more than the point estimate. The placebo permutes the treatment and
re-estimates, and the estimate should collapse. What makes it a real test rather than a formality
is the **block size**, discussed at the table below: permuting one bar at a time destroys exactly
the serial dependence that makes the original estimate hard to get right, and a placebo that
destroys the difficulty is a test the estimate passes for free.
**Learning objectives.** By the end of this notebook you will be able to:
- State an estimand - treatment, outcome, confounders, horizon - precisely enough that someone
else could disagree with it.
- Explain what cross-fitting protects against, and why chronological folds with an embargo are
the right shape for it on overlapping financial labels.
- Say why a block-permutation placebo needs a block long enough to preserve serial dependence,
and how the block size is derived here rather than chosen.
- Keep a causal estimate and a predictive result in separate reports, and say in one sentence
why combining them would misdescribe both.
**Book reference:** Chapter 15, causal inference for trading research.
**Prerequisites:** [`03_financial_features`](03_financial_features.ipynb),
[`02_labels`](02_labels.ipynb) and [`05_evaluation`](05_evaluation.ipynb) - the features, the
return labels and the purged walk-forward folds.
**What it writes:** one causal result under its own identity in `run_log/registry.db`. It is read
by [`12_model_analysis`](12_model_analysis.ipynb), which reports it **beside** the predictive
results rather than among them.
```python
import os
import polars as pl
from case_studies.crypto_perps_funding.research_workflow import open_study
from case_studies.research import causal_supersedes
```
```python
EXECUTION_TIER = "canonical"
WORKSPACE = os.environ.get("ML4T_OUTPUT_DIR", "")
LABEL = "fwd_ret_8h"
CONFIG_NAME = "dml"
PREVIEW_REDUCTIONS = {}
OVERRIDES = {}
# The causal identity this run retires. A reader's clone has nothing to retire - `run_log/` is
# not shipped, so a first run meets an empty registry - and `causal_supersedes` withholds the
# declaration against a registry that does not hold it, which is why one committed value is
# right for both. A predecessor named against a registry that *does* hold rows but not this one
# is rejected at the registering write, after the DML fit and every placebo refit are paid for.
#
# The source component of a causal identity is `CAUSAL_RUNNER_VERSION`, a declared integer in
# `case_studies/utils/causal.py`; nothing hashes the file. So an edit to the estimator moves no
# identity unless that constant is raised by hand, and an edit that changes a registered value
# without raising it leaves the next run to hit the cache and serve the old number under new
# code. Once it is raised, `CausalResult.one` resolves a label to exactly one canonical
# identity, so a re-run that does not name what it replaces leaves two live and the next
# notebook fails with "resolved to 2 identities".
# Retired by this run: the block-permutation refutation now compares the HAC t-statistic
# rather than the raw effect, so CAUSAL_RUNNER_VERSION moved and every causal identity with
# it. The rows named here hold a p-value computed on the shrunken placebo effects; this run
# supersedes them rather than correcting them, because the statistic is different, not the
# arithmetic. Read out of each registry's current canonical identity per label, 2026-09-10.
SUPERSEDES_CAUSAL: str = "025f6c2f4e3b"
```
## 1. Resolve the estimand and the refutation contract
Nothing is fitted below. The request resolves to a specification and an identity, and the table
prints the fields that decide what the estimate means: what is being intervened on, what is being
measured, over what horizon, how the nuisance folds are cut, and how the placebo will be built.
Reading them before the fit is the only point at which disagreeing with the estimand is cheap.
`eligible_rows` is the analysis population - the rows that survive having a treatment, an
outcome and every confounder present. Quote that, not the panel height, when describing what the
estimate rests on.
```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
request = study.causal(
method="dml",
label=LABEL,
config_name=CONFIG_NAME,
execution_tier=EXECUTION_TIER,
preview_reductions=PREVIEW_REDUCTIONS,
overrides=OVERRIDES,
supersedes=causal_supersedes(
study, SUPERSEDES_CAUSAL, LABEL, labels=[LABEL], execution_tier=EXECUTION_TIER
),
)
resolved = request.resolve()
computation = resolved.spec["computation"]
pl.DataFrame(
{
"causal_hash": [resolved.identity],
"outcome": [computation["estimand"]["outcome"]],
"treatment": [computation["estimand"]["treatment"]],
"outcome_horizon": [computation["estimand"]["outcome_horizon"]],
"n_folds": [computation["cv"]["n_folds"]],
"embargo_periods": [computation["cv"]["embargo_periods"]],
"block_size": [computation["refutation"]["block_size"]],
"block_size_basis": [computation["refutation"]["block_size_basis"]],
"gap_policy": [computation["refutation"]["temporal_gap_policy"]],
"eligible_rows": [computation["analysis_population"]["n_rows"]],
}
)
```
`block_size` is the parameter the refutation lives or dies on, so it is on the
table rather than buried in the spec. The placebo permutes contiguous blocks
within each symbol; a block of one bar is an iid shuffle, which destroys the
serial dependence the placebo is meant to keep and makes the test trivially
easy to pass. Two things create that dependence and the block spans the longer
of them: the overlapping labels span the outcome horizon, and the treatment
spans its own construction window. Here the horizon is a single 8-hour bar
while `premium_zscore_14d` is a 42-bar rolling statistic, so the treatment
window sets the block and `block_size_basis` says so.
## 2. Execute, and register the result separately
The fit runs the nuisance models across the chronological folds, forms the residual-on-residual
estimate, and then pays for the placebo refits - which is where most of the cost is, since the
whole procedure is repeated once per placebo draw.
The check below refuses a result whose *computation* is not the one that was resolved. That is not
defensive coding: the computation carries the estimand, the fold geometry and the refutation
contract, and a result carrying a different one describes a contract that no longer exists. The
notebook downstream would then resolve the label to two live identities and stop.
It compares `spec["computation"]` and not the whole spec, because the spec also carries
`provenance` - the git commit, the platform and the package versions of the run that fitted. Those
record that run's own circumstances and are meant to differ from any later run that reads the row
back from cache. Comparing whole specs asserts that nothing has been committed since, which is not
a property of the estimate, and it makes the notebook raise on every re-run from a different
commit - the opposite of what this check is for. The second run of this notebook fits nothing and
has to say so rather than raise.
Comparing `result.hash` against `resolved.identity` would not do it either: both are
`training_hash_from_spec` of the same resolved spec, so that comparison cannot fail and would
assert nothing. Same check, same reasoning, as `etfs/12_causal_dml`.
```python
result = resolved.run()
if not result.complete:
raise RuntimeError("causal execution is incomplete")
if result.spec["computation"] != resolved.spec["computation"]:
raise RuntimeError("the registered causal computation differs from the resolved one")
pl.DataFrame(
{
"causal_hash": [result.hash],
"n_obs": [result.metrics["n_obs"]],
"complete": [result.complete],
"execution_tier": [result.execution_tier],
}
)
```
## Key takeaways and limitations
- **This is not a trading signal and must not be reported as one.** The estimate answers what
would happen under an intervention on the premium z-score. The backtest in
[`13_backtest`](13_backtest.ipynb) answers what a strategy would have earned. A number from
here quoted as evidence for a strategy is a category error, and the separate registration
exists to make that mistake require deliberate effort.
- **Adjustment covers three named confounders and nothing else.** DML removes bias from variables
it is given. An omitted common cause, reverse causation from returns to the premium, or a
confounder measured with error all survive it untouched, and no diagnostic in this notebook
would reveal them. The identifying assumption is an argument, not an output.
- **The refutation is the load-bearing part.** A point estimate arrives whether or not there is
anything to estimate. What distinguishes the two is whether the placebo collapses, and whether
the placebo was built to be hard - which is what the block size decides.
- **The identity hashes the estimator's source.** Any edit to `case_studies/utils/causal.py`
produces a different causal identity, whether or not it changes a number. That is deliberate: a
causal claim is a claim about a procedure, and a changed procedure is a new claim until someone
shows the estimate did not move.
- **A preview run is not a canonical result.** The reduced sample and fold counts a preview
applies are recorded in its identity and it is barred from the canonical population, so a cheap
run can never be mistaken for the published estimate.Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.