Zum Inhalt springen
Alle Bibliotheksdokumente

Kausale DML zur Varianzrisikoprämie in S&P-500-Optionen validieren

Notebook Machine Learning for Trading

Zusammenfassung

Dieses Notebook spezifiziert und führt eine Double-Machine-Learning-Analyse zum Effekt der Varianzrisikoprämie auf Renditen von Short-Optionen bis zum Verfall durch. Vor der Ausführung legt es Behandlung, Ergebnisgröße, Störfaktoren, Zeitablauf, Störmodell, zeitliche Kreuzvalidierung, Kovarianzschätzer und Placebo-Widerlegungsverfahren fest. Die Berechnung übernimmt ein gemeinsam genutzter Ausführer. Dieser prüft erforderliche Störfaktoren, zeitliche Folds, Anpassungen der Störmodelle und den HAC-Standardfehler. Anschließend veröffentlicht er ein Artefakt, das an die aufgelöste Identität der Anfrage gebunden ist.

Eine zentrale methodische Änderung betrifft den Placebo-Test. Das Dokument erklärt, dass eine blockweise Permutation der Behandlung die Residualvarianz erhöhen und Placebo-Effektschätzungen künstlich verengen kann. Das Ergebnis der Widerlegung würde dadurch zu günstig ausfallen. Deshalb vergleicht es HAC-t-Statistiken, die Unsicherheit berücksichtigen, statt rohe Effektschätzungen. Das Notebook führt ein zurückgezogenes Ergebnis als Anlass für die Änderung an, dient jedoch der Ausführung und Validierung; die Interpretation der neuen kausalen Schätzung bleibt einer separaten Analyse vorbehalten. Die Evidenz gilt für diese Behandlung, Ergebnisgröße und deklarierte Analysepopulation. Eine erfolgreiche Ausführung allein belegt keine kausale Schlussfolgerung.

Kernaussagen

  • Die DML-Anfrage legt vor der Anpassung Schätzgröße, Zeitablauf, Störfaktoren, Störmodell und Inferenzdesign fest.
  • Zeitliche Folds und HAC-Kovarianz werden festgelegt, um den Ergebnishorizont abzubilden.
  • Eine Blockpermutation kann Placebo-Effektschätzungen verzerren, wenn sich die Residualvarianz der Behandlung ändert.
  • Der Vergleich von Placebo- und beobachteten HAC-t-Statistiken bezieht Unsicherheit in die Widerlegungsprüfung ein.
  • Ein validiertes Artefakt dokumentiert den erfolgreichen Abschluss der angeforderten Berechnung; die inhaltliche Interpretation bleibt einer späteren Analyse vorbehalten.

Schlagwörter

Volltext
# S&P 500 Options: Causal DML Execution


# S&P 500 Options: Causal DML Execution

This notebook estimates the effect of the variance-risk-premium treatment on the
return-to-expiry outcome. It declares the request through the shared causal boundary and exposes
the resolved estimand, timing, confounders, nuisance model, covariance design, and refutation
protocol before execution.

`11_model_analysis` interprets the causal estimates. This notebook validates the computation
and publishes its artifact only.

Prerequisites: `03_financial_features`, `04_model_based_features`, and `05_evaluation`.

```python
"""Execute the declared S&P 500 options causal DML request."""

import polars as pl

from case_studies.research import causal_supersedes
from case_studies.sp500_options.research_workflow import open_study
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
# Retired by this run: the block-permutation refutation now compares the HAC t-statistic
# rather than the raw effect, so CAUSAL_RUNNER_VERSION moved and every causal identity with
# it. The rows named here hold a p-value computed on the shrunken placebo effects; this run
# supersedes them rather than correcting them, because the statistic is different, not the
# arithmetic. Read out of each registry's current canonical identity per label, 2026-09-10.
SUPERSEDES_CAUSAL: str = "d034b82943c5"
```

## Declared and resolved request

A preview must declare all sample, symbol, fold, or placebo reductions. Canonical execution uses
the complete pre-holdout analysis population.

### What `SUPERSEDES_CAUSAL` retires here

`CausalResult.one` resolves a label to exactly one canonical identity, so a refit has to name the
identity it replaces or the registry is left with two and refuses. The retired identity is
`d034b82943c5`.

What changed is the refutation statistic, not the fit. The placebo loop used to compare each
permuted run's *effect estimate* against the observed effect. Block-permuting the treatment frees
it from the controls, so the first stage can no longer predict it and its residual keeps nearly
all its variance. That residual variance is the whole denominator of the second-stage effect, so
every placebo effect is divided by a larger number than the observed one and the placebo
distribution comes out narrower than the null it stands for. The bias runs one way, toward a
refutation that reads as passed. The comparison is now on the HAC t-statistic, which carries the
denominator in it and cancels the inflation.

The retired identity fitted the same 166,105 observations and reported the same effect of 0.4098
with a HAC standard error of 0.3509, so p = 0.243 under either statistic. Its refutation p was
0.0099, the smallest value 100 draws can report, for an effect whose own t-statistic is 1.17.
Sitting at that floor is the signature. This notebook registers and hands off; the new row is
read and interpreted in `11_model_analysis`.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
request_table = pl.DataFrame(
    {
        "method": ["dml"],
        "label": ["ret_to_expiry"],
        "config_name": ["dml"],
        "execution_tier": [EXECUTION_TIER],
    }
)
request_table
```

```python
request = study.causal(
    **request_table.row(0, named=True),
    preview_reductions=PREVIEW_REDUCTIONS,
    supersedes=causal_supersedes(
        study,
        SUPERSEDES_CAUSAL,
        "ret_to_expiry",
        labels=["ret_to_expiry"],
        execution_tier=EXECUTION_TIER,
    ),
)
resolved = request.resolve()
computation = resolved.spec["computation"]
estimand = computation["estimand"]
causal_plan = pl.DataFrame(
    {
        "treatment": [estimand["treatment"]],
        "outcome": [estimand["outcome"]],
        "confounders": [", ".join(estimand["confounders"])],
        "treatment_observed_at": [estimand["treatment_observed_at"]],
        "outcome_horizon": [estimand["outcome_horizon"]],
        "folds": [computation["cv"]["n_folds"]],
        "embargo_periods": [computation["cv"]["embargo_periods"]],
        "nuisance_model": [computation["model"]["class"]],
        "covariance": ["HAC with the outcome horizon"],
        "placebo_method": [computation["refutation"]["method"]],
        "placebo_block": [computation["refutation"]["block_size"]],
        "placebo_block_basis": [computation["refutation"]["block_size_basis"]],
        "analysis_rows": [computation["analysis_population"]["n_rows"]],
        "training_hash": [resolved.identity],
    }
)
causal_plan
```

## Execute and validate

The shared DML runner fails on missing confounders, invalid temporal folds, incomplete nuisance
fits, or a non-finite HAC standard error. A cached result must match the complete resolved
identity before it can be reused.

```python
if EXECUTION_TIER == "preview" and (not WORKSPACE or not PREVIEW_REDUCTIONS):
    raise ValueError("preview execution requires WORKSPACE and PREVIEW_REDUCTIONS")
result = resolved.run()
if not result.complete or result.hash != resolved.identity:
    raise RuntimeError("causal execution did not publish the complete resolved request")
```

```python
artifact = pl.DataFrame(
    {
        "causal_hash": [result.hash],
        "label": [resolved.spec["label"]],
        "execution_tier": [result.execution_tier],
        "complete": [result.complete],
    }
)
artifact
```

The registered causal artifact is the handoff to `11_model_analysis`. No estimate or empirical
conclusion is interpreted here.

Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT

Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.