رفتن به محتوا
همه اسناد کتابخانه

تحلیل علّی DML برای برآورد اثر صرف ریسک واریانس بر بازده اختیارها

کد یادگیری ماشین برای معامله‌گری

خلاصه

این دفترچه تحلیل یادگیری ماشین دوگانه‌ای را درباره اثر صرف ریسک واریانس بر بازده اختیارهای S&P 500 تا سررسید تعریف و اجرا می‌کند. پیش از اجرا، کمیت هدف، زمان‌بندی مشاهدات، عوامل مخدوش‌کننده، طرح اعتبارسنجی متقاطع زمانی، مدل مزاحم، طرح کوواریانس HAC و روش ابطال با دارونما را آشکار می‌کند. اگر از حالت پیش‌نمایش استفاده شود، درخواست نهایی فقط پس از مشخص شدن کاهش‌های پیش‌نمایش اجرا می‌شود؛ دفترچه نیز کامل بودن نتیجه و تطابق آن با شناسه برنامه‌ریزی‌شده را بررسی می‌کند.

دفترچه همچنین تغییری در مقایسه دارونما مستند می‌کند: اکنون آماره‌های تی HAC را با هم مقایسه می‌کند، نه برآوردهای خام اثر را، زیرا جایگشت بلوکی می‌تواند واریانس باقیمانده درمان را تغییر دهد و اثرهای دارونما را به‌طور گمراه‌کننده‌ای کم‌دامنه نشان دهد. برای انگیزه به نتیجه‌ای پیشین استناد می‌کند، اما نقش اعلام‌شده‌اش اعتبارسنجی و انتشار محاسبه است، نه تفسیر برآورد علّی جدید. بنابراین، نتیجه‌گیری‌های علّی به مرحله تحلیلی جداگانه نیاز دارند و همچنان به کنترل‌ها، فرض‌های زمان‌بندی و طرح ابطال مشخص‌شده وابسته‌اند.

ایده‌های کلیدی

  • تحلیل، اثر علّی صرف ریسک واریانس بر بازده اختیارها تا سررسید را برآورد می‌کند.
  • دفترچه پیش از اجرا، کمیت هدف، زمان‌بندی، عوامل مخدوش‌کننده، فولدها، روش کوواریانس و طرح دارونما را آشکار می‌کند.
  • مقایسه آماره‌های تی HAC در ابطال دارونما، تفاوت‌های واریانس ناشی از جایگشت درمان را در نظر می‌گیرد.
  • پیش از انتشار، نتیجه باید کامل باشد و با شناسه درخواست نهایی مطابقت داشته باشد.
  • این دفترچه اجرایی، برآوردها را برای تفسیر جداگانه تحویل می‌دهد و خود ادعای نتیجه تجربی جدیدی نمی‌کند.

برچسب‌ها

متن کامل
# 10_causal_dml.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # S&P 500 Options: Causal DML Execution
#
# This notebook estimates the effect of the variance-risk-premium treatment on the
# return-to-expiry outcome. It declares the request through the shared causal boundary and exposes
# the resolved estimand, timing, confounders, nuisance model, covariance design, and refutation
# protocol before execution.
#
# `11_model_analysis` interprets the causal estimates. This notebook validates the computation
# and publishes its artifact only.
#
# Prerequisites: `03_financial_features`, `04_model_based_features`, and `05_evaluation`.

# %%
"""Execute the declared S&P 500 options causal DML request."""

import polars as pl

from case_studies.research import causal_supersedes
from case_studies.sp500_options.research_workflow import open_study

# %% tags=["parameters"]
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
# Retired by this run: the block-permutation refutation now compares the HAC t-statistic
# rather than the raw effect, so CAUSAL_RUNNER_VERSION moved and every causal identity with
# it. The rows named here hold a p-value computed on the shrunken placebo effects; this run
# supersedes them rather than correcting them, because the statistic is different, not the
# arithmetic. Read out of each registry's current canonical identity per label, 2026-09-10.
SUPERSEDES_CAUSAL: str = "d034b82943c5"

# %% [markdown]
# ## Declared and resolved request
#
# A preview must declare all sample, symbol, fold, or placebo reductions. Canonical execution uses
# the complete pre-holdout analysis population.
#
# ### What `SUPERSEDES_CAUSAL` retires here
#
# `CausalResult.one` resolves a label to exactly one canonical identity, so a refit has to name the
# identity it replaces or the registry is left with two and refuses. The retired identity is
# `d034b82943c5`.
#
# What changed is the refutation statistic, not the fit. The placebo loop used to compare each
# permuted run's *effect estimate* against the observed effect. Block-permuting the treatment frees
# it from the controls, so the first stage can no longer predict it and its residual keeps nearly
# all its variance. That residual variance is the whole denominator of the second-stage effect, so
# every placebo effect is divided by a larger number than the observed one and the placebo
# distribution comes out narrower than the null it stands for. The bias runs one way, toward a
# refutation that reads as passed. The comparison is now on the HAC t-statistic, which carries the
# denominator in it and cancels the inflation.
#
# The retired identity fitted the same 166,105 observations and reported the same effect of 0.4098
# with a HAC standard error of 0.3509, so p = 0.243 under either statistic. Its refutation p was
# 0.0099, the smallest value 100 draws can report, for an effect whose own t-statistic is 1.17.
# Sitting at that floor is the signature. This notebook registers and hands off; the new row is
# read and interpreted in `11_model_analysis`.

# %%
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
request_table = pl.DataFrame(
    {
        "method": ["dml"],
        "label": ["ret_to_expiry"],
        "config_name": ["dml"],
        "execution_tier": [EXECUTION_TIER],
    }
)
request_table

# %%
request = study.causal(
    **request_table.row(0, named=True),
    preview_reductions=PREVIEW_REDUCTIONS,
    supersedes=causal_supersedes(
        study,
        SUPERSEDES_CAUSAL,
        "ret_to_expiry",
        labels=["ret_to_expiry"],
        execution_tier=EXECUTION_TIER,
    ),
)
resolved = request.resolve()
computation = resolved.spec["computation"]
estimand = computation["estimand"]
causal_plan = pl.DataFrame(
    {
        "treatment": [estimand["treatment"]],
        "outcome": [estimand["outcome"]],
        "confounders": [", ".join(estimand["confounders"])],
        "treatment_observed_at": [estimand["treatment_observed_at"]],
        "outcome_horizon": [estimand["outcome_horizon"]],
        "folds": [computation["cv"]["n_folds"]],
        "embargo_periods": [computation["cv"]["embargo_periods"]],
        "nuisance_model": [computation["model"]["class"]],
        "covariance": ["HAC with the outcome horizon"],
        "placebo_method": [computation["refutation"]["method"]],
        "placebo_block": [computation["refutation"]["block_size"]],
        "placebo_block_basis": [computation["refutation"]["block_size_basis"]],
        "analysis_rows": [computation["analysis_population"]["n_rows"]],
        "training_hash": [resolved.identity],
    }
)
causal_plan

# %% [markdown]
# ## Execute and validate
#
# The shared DML runner fails on missing confounders, invalid temporal folds, incomplete nuisance
# fits, or a non-finite HAC standard error. A cached result must match the complete resolved
# identity before it can be reused.

# %%
if EXECUTION_TIER == "preview" and (not WORKSPACE or not PREVIEW_REDUCTIONS):
    raise ValueError("preview execution requires WORKSPACE and PREVIEW_REDUCTIONS")
result = resolved.run()
if not result.complete or result.hash != resolved.identity:
    raise RuntimeError("causal execution did not publish the complete resolved request")

# %% tags=["results"]
artifact = pl.DataFrame(
    {
        "causal_hash": [result.hash],
        "label": [resolved.spec["label"]],
        "execution_tier": [result.execution_tier],
        "complete": [result.complete],
    }
)
artifact

# %% [markdown]
# The registered causal artifact is the handoff to `11_model_analysis`. No estimate or empirical
# conclusion is interpreted here.

```

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: MIT

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.