עבור לתוכן
כל מסמכי הספרייה

ניתוח אסטרטגיית סטרדל באופציות S&P 500 תוך התחשבות באי־ודאות

קוד Machine Learning for Trading

סיכום

מחברת זו מעריכה אסטרטגיית סטרדל שבועית באופציות S&P 500 באמצעות בקטסטים רשומים ותהליך בחירה קבוע. היא שולפת מהמרשם את התצורה שנבחרה לאחר הגבלת היקום לנכסים נזילים, ואז מעריכה תוצאות אימות ותקופה שמורה בלי להשתמש בתקופה השמורה לבחירת האסטרטגיה. הבקטסט משתמש בגידור דלתא יומי, בסליקה בפקיעה ובעלויות מרווח לכל רגל באופציה ובגידור. המדדים המדווחים כוללים רווחי סמך באמצעות bootstrap לבלוקים, ומשתמשים בהשוואות מזווגות לדעיכה בתקופה השמורה ולהשוואה למדד ייחוס במשקל שווה.

המסמך מתאר הן את מקדם המידע החזוי והן את Sharpe של האסטרטגיה כמתיישבים סטטיסטית עם אפס עבור תווית התשואה הנבחרת עד לפקיעה, ולכן אינו מוצא יתרון ברור בתקופת האימות. תרומתו העיקרית היא מסגרת הערכה ממושמעת ששומרת על אי־ודאות ועל חשבונאות עלויות, ולא ראיה לאסטרטגיה רווחית. התוצאות תלויות בקבוצת המועמדים שהוגדרה, בהגבלות היקום, בהנחות הבקטסט ובמדגם הרשום; המחברת עצמה אינה מאמנת מודלים או מריצה מחדש בקטסטים.

רעיונות מרכזיים

  • האסטרטגיה נבחרת מתוך יקום נזיל שהוגדר מראש, לפני הערכת התקופה השמורה.
  • הבקטסטים מביאים בחשבון מרווחי כניסה לאופציות, מרווחי גידור בנכס הבסיס, גידור דלתא וסליקה בפקיעה.
  • רווחי סמך באמצעות bootstrap לבלוקים והשוואות מזווגות מבטאים אי־ודאות בביצועים ובדעיכת התקופה השמורה.
  • מדדי החיזוי והאסטרטגיה המדווחים מתיישבים עם אפס, ולכן הניתוח אינו מזהה יתרון.
  • המסקנות תלויות בקבוצת המועמדים, בהנחות העלויות ובבקטסטים הרשומים הזמינים.

תגיות

הטקסט המלא
# 18_strategy_analysis.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # S&P 500 Options - Strategy Analysis
#
# This notebook converts the case-study backtest registry for the S&P
# 500 options HTM straddle strategy into a per-case-study strategy
# assessment. Which configuration carries it is resolved from the registry in
# §1 and printed there, not named here: it is a property of the run and it moves
# whenever the sweep is rebuilt, and a name written into prose agrees with the
# registry only until the next rebuild. What is fixed is the selection contract
# - the liquid-universe pin is applied first, then the surface is ranked on
# validation, and the holdout is not consulted.
# All `ret_to_expiry` backtests dispatch
# through the HTM daily-MTM cohort engine (`_run_htm_daily_mtm`): entry on
# the final available session of each Friday week, the underlying delta
# hedged whenever it breaches its threshold, settle at intrinsic value at
# expiration, full per-leg costs (entry-side option spread, underlying hedge
# spread on each session the hedge trades, and an exit-leg option trade only
# where a contract's chain ends before its expiration date). Every metric is reported with its
# block-bootstrap 95% CI; paired holdout comparisons via
# `backtest_paired_metrics` are required before the notebook reports them.
# Cross-case-study comparison is reserved for Chapter 20.
#
# **Learning objectives**
#
# - Read uncertainty-aware backtest metrics for an options strategy whose
#   signal is statistically null (Sharpe and IC both straddle zero) under
#   HTM cost accounting on full per-leg friction.
# - Surface a holdout decay reading without invoking champion/winner/
#   verdict language, and a holdout-vs-EW comparison for the same window.
#
# **Book reference**: Chapter 20, §20.1 - sp500_options anchors the
# "cost model validity" theme in the cross-case-study synthesis.
#
# **Prerequisites**: the case-study pipeline through
# [`17_holdout_backtest`](17_holdout_backtest.ipynb), which registers the holdout result this
# notebook closes on.
#
# **Scope**: no training and no re-backtesting. It does write two derived tables it also reads -
# `backtest_paired_metrics` and `cohort_metrics` - because a case study has to be able to produce
# everything its own notebooks report, and only Chapter 20's synthesis ever wrote them. Both are
# computed from the registered backtests over the population this notebook nominates, and both are
# replaced rather than added to, so a selected configuration that moved does not leave its
# predecessor's rows behind looking current.

# %%
"""S&P 500 Options - Strategy Analysis."""

# ruff: noqa: E402, I001

import json
import sqlite3

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl

# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch  # noqa: F401
import yaml


from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
    BacktestReportMetadata,
    generate_tearsheet_from_run_artifacts,
)

# %% [markdown]
# Load the registry readers and reporting helpers the notebook draws on.

# %%
from case_studies.research import CandidateSet, OfficialPopulation, Study
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.cohort_reporting import cohort_metric_attribution, reportable_pbo
from case_studies.utils.paired_metrics import populate_paired_metrics, rung_for
from case_studies.utils.uncertainty import ENTIRE_REGISTRY, cohort_member_digest
from case_studies.utils.factor_attribution import (
    compute_bootstrap_ci,
    load_factor_data,
    plot_attribution_waterfall,
    run_factor_regression,
)
from case_studies.utils.registry import (
    load_backtest_fold_metrics,
    load_backtest_metrics,
    load_paired_metrics,
)
from case_studies.utils.strategy_analysis import (
    ci_status,
    compute_operating_profile,
    fmt_gate,
    gate1_validation_sharpe_geq_zero,
    gate2_holdout_diff_not_excludes_zero_negatively,
    gate_passes,
    plot_equity_drawdown,
    resolve_solvent_carrier,
    write_strategy_assessment,
)
from utils.paths import get_case_study_dir, get_output_dir
from utils.style import show_with_alt

# %% tags=["parameters"]
MAX_SYMBOLS = 0

# %%
CASE_STUDY = "sp500_options"
# The immutable grid `12_backtest` publishes; the DSR below deflates for its members.
BASELINE_POPULATION = "sp500-options-baseline-validation-v1"
# The nominees `13_portfolio_management` publishes; the selected configuration has to be one of
# them.
STRATEGY_CANDIDATES = "sp500-options-strategy-candidates-v1"
PRIMARY_LABEL = "ret_to_expiry"  # registered HTM strategy label (Appendix A)
PERIODS_PER_YEAR = 252
CASE_DIR = get_case_study_dir(CASE_STUDY)
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

with open(CASE_DIR / "config" / "setup.yaml") as f:
    setup = yaml.safe_load(f)

explorer = BacktestExplorer(CASE_STUDY)
_selection_study = Study.open(CASE_STUDY)
print(explorer)

# %% [markdown]
# Format point estimates and confidence intervals consistently throughout
# the notebook.


# %%
def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
    if point is None:
        return "-"
    p = format(point, fmt)
    if lo is None or hi is None:
        return f"{p} [-, -]"
    return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"


# %% [markdown]
# Format scalar metrics while preserving missing values.


# %%
def _fmt(val: float | None, fmt: str = ".4f") -> str:
    return "-" if val is None else format(val, fmt)


# %% [markdown]
# ## §1 Handoff from model analysis
#
# The strategy phase inherits an HTM short-straddle pipeline: weekly
# entry on a top-k cross-section of `ret_to_expiry`, daily delta hedge
# through the underlying, hold to expiry, full per-leg costs on entry-side
# option spread plus daily underlying hedge spread, and an exit-leg option
# trade only where a contract's chain ends before its expiration date. The
# family, configuration and allocator that carry it are resolved from the
# registry below and printed alongside their identifiers rather than named
# here. The
# liquid universe pin (`UNIVERSE_RESTRICTIONS`) excludes the higher-Sharpe
# full-universe rows from rank-1 selection, so the deployed configuration stays on
# the liquid subset and remains poolable with its holdout replay. The selected configuration
# is fixed before the one holdout evaluation. The upstream
# prediction-side IC and the downstream strategy Sharpe are both
# statistically consistent with zero - there is no IC-vs-Sharpe disconnect
# on this strategy; both metrics agree that no edge has been resolved on
# `ret_to_expiry` over the validation window. `LABEL_RESTRICTIONS` excludes
# the four diagnostic `fwd_ret_*` variants from rank-1 selection (and those
# rows have been dropped from the registry entirely as of the 2026-05-17
# sweep cleanup); the only label that anchors a registered rank-1 here is
# `ret_to_expiry` under HTM cost accounting.

# %%
# Rank-1 = the registered HTM strategy on ret_to_expiry. Resolved
# dynamically from the registry with LABEL_RESTRICTIONS applied so any
# residual diagnostic-variant rows (fwd_ret_5d/_10d/_dh_*) - which inflated
# Sharpes under bps-of-notional costs and cannot anchor a registered HTM
# strategy - are excluded from the cross-stage rank-1 lookup. The val/
# holdout pair shares a training_hash; the notebook reads from those rows
# rather than recomputing.
# The nominated field is resolved BEFORE the ranking and handed to it, not checked against
# the winner afterwards. Those are different tests wherever the conformal branch is taken:
# the common-support re-ranking restricts every series to the timestamps they all share, so a
# row that is never going to win still decides how far that intersection reaches, and
# therefore which admitted candidate does. A membership check on the winner passes while the
# answer has already been moved by a row that was never eligible.
_candidate_hashes = frozenset(CandidateSet.one(_selection_study, name=STRATEGY_CANDIDATES).members)
# `resolve_solvent_carrier` rather than the bare lineage resolver: it applies the same selection and
# additionally refuses a selected configuration whose equity reached zero, whose Sharpe is
# arithmetic on a balance that no longer exists.
_lineage = resolve_solvent_carrier(CASE_STUDY, admitted=_candidate_hashes)
TOP_HASH = _lineage["val_backtest_hash"]
TOP_PHASH = _lineage["val_prediction_hash"]
HO_HASH = _lineage["holdout_backtest_hash"]
HO_PHASH = _lineage["holdout_prediction_hash"]
RANK1_FAMILY = _lineage["family"]
RANK1_CONFIG = _lineage["config_name"]
assert _lineage["label"] == PRIMARY_LABEL, (
    f"Canonical rank-1 label {_lineage['label']!r} != PRIMARY_LABEL "
    f"{PRIMARY_LABEL!r} - LABEL_RESTRICTIONS misconfigured?"
)
assert HO_HASH is not None, (
    "No holdout backtest for the canonical val rank-1. Run "
    "16_holdout_predictions and 17_holdout_backtest, which refit the selected configuration over the "
    "holdout interval and register the result under a training identity of its own. Do not "
    "reach for 20_strategy_synthesis/holdout.py::generate_holdout: it registers its "
    "predictions under the VALIDATION training identity, so what it produces is a "
    "validation-fitted model scored on the holdout window, which is the one thing the "
    "holdout exists to rule out."
)
# The field was applied before the ranking, so this cannot fail on a row from outside it. It stays
# because it is cheap and because it is the assertion a reader needs: the selected configuration
# published below is one of the nominees `13_portfolio_management` froze, not whatever happened to
# rank first in the registry.
if TOP_HASH not in _candidate_hashes:
    raise RuntimeError(
        f"Canonical rank-1 {TOP_HASH} is not a member of {STRATEGY_CANDIDATES} "
        f"({len(_candidate_hashes)} nominees), so the selected configuration was ranked from outside the "
        "set this case study nominated"
    )

# The prediction sets behind the nominated backtests, resolved once and used wherever a
# cohort or a leader is taken below. Both of those questions are about the search this case
# study declared, and the registry holds more than that search: retired generations, the
# full-universe rows the Chapter 18 cost cascade keeps, and any label this notebook does not
# report. Restricting on the universe alone leaves all three in, and a cohort that admits
# them has a different leader and a different deflation than the search being reported.
_db = CASE_DIR / "run_log" / "registry.db"
with sqlite3.connect(str(_db)) as _con:
    _CANDIDATE_PREDICTIONS = tuple(
        sorted(
            row[0]
            for row in _con.execute(
                "SELECT DISTINCT prediction_hash FROM backtest_runs WHERE backtest_hash IN "
                f"({','.join('?' * len(_candidate_hashes))})",
                sorted(_candidate_hashes),
            )
        )
    )
if not _CANDIDATE_PREDICTIONS:
    raise RuntimeError(
        f"{STRATEGY_CANDIDATES} names {len(_candidate_hashes)} backtests and none of them is "
        "in backtest_runs, so the nominated search cannot be resolved to a prediction "
        "population"
    )
print(
    f"Nominated search: {len(_candidate_hashes)} backtests over {len(_CANDIDATE_PREDICTIONS)} prediction sets"
)

# %% [markdown]
# ### The paired rows this notebook reads, built here rather than assumed
#
# §6 reads two `backtest_paired_metrics` kinds keyed on the holdout backtest, and it used to
# load rows that only `20_strategy_synthesis/01_aggregate_synthesis.py` ever wrote. On a
# registry that has been reset, or one whose configuration moved, those rows are absent or belong to
# a holdout hash that no longer exists, and the notebook stopped on a message telling the
# reader to populate the case elsewhere. A case study's own pipeline has to be able to produce
# everything its own notebooks report.
#
# `populate_paired_metrics` is the same producer Chapter 20 calls, and it is given the same
# selection this notebook made: the configuration from the lineage resolver, and the rung
# this case study is pinned to. The pin matters here - rung-1 (mid-to-mid bps) and rung-2
# (full-universe HTM) both carry `universe_filter="full"`, so ranking on the universe alone would
# let either stand in for the liquid configuration the case study actually deploys.
#
# `replace_all` makes the call a snapshot. Registration is an upsert keyed on `(challenger_hash,
# benchmark_hash)`, so it can add this configuration's rows and cannot remove the rows of a selected
# configuration it replaced; those rows name a holdout hash that is no longer the holdout and would
# sit in the table looking current. This case study restricts to one label and one rung, so what
# the call writes is the whole of what belongs here, which is the condition under which pruning to
# it is right.

# %%
# `cohort_metrics` is the same story one table over: this notebook reads the cross-family
# baseline cohort and the deflated Sharpe computed over it, and only Chapter 20 ever wrote
# them. `compute_and_register` is that producer, and it refreshes the whole table rather than
# one row, so it also removes a cohort led by a backtest the registry no longer holds. The
# universe pin goes with it - the full-universe rows are kept for the Chapter 18 cost cascade
# and must not enter a cohort the selected configuration is deflated against.
_cohort_counts = compute_and_register(
    CASE_STUDY,
    universe_filter=rung_for(CASE_STUDY)["universe_filter"],
    prediction_hashes=_CANDIDATE_PREDICTIONS,
    verbose=False,
)
print(
    f"cohort_metrics: {sum(v for k, v in _cohort_counts.items() if k not in ('errors', 'dangling_pruned'))} "
    f"rows across {sorted(k for k in _cohort_counts if k not in ('errors', 'dangling_pruned'))}"
    f", {_cohort_counts['dangling_pruned']} dangling pruned, {_cohort_counts['errors']} errors"
)

# %%
_paired_rows = populate_paired_metrics(
    CASE_STUDY,
    explorer,
    label_restriction=frozenset({PRIMARY_LABEL}),
    rung=rung_for(CASE_STUDY),
    carrier=_lineage,
    # The cohort call above is scoped to `_CANDIDATE_PREDICTIONS` and this one is not:
    # the pairs are selected from every registered prediction set. That disagreement
    # was there before the scope had to be stated; narrowing it changes published
    # numbers, so it is a separate decision from this line.
    prediction_hashes=ENTIRE_REGISTRY,
    periods_per_year=PERIODS_PER_YEAR,
    replace_all=True,
    verbose=False,
)
_written = [row for row in _paired_rows if "skip" not in row]
print(f"backtest_paired_metrics: {len(_written)} pairs written, {len(_paired_rows)} attempted")
for _row in _paired_rows:
    if "skip" in _row:
        print(f"  skipped {_row.get('benchmark_kind', '?')}: {_row['skip']}")

# %% [markdown]
# Resolve the selected configuration's prediction metrics and registered strategy
# specification directly from the registry.

# %%
with sqlite3.connect(str(_db)) as _con:
    _row = _con.execute(
        "SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
        "ic_hac_lag, ic_pct_positive "
        "FROM prediction_metrics WHERE prediction_hash = ?",
        (TOP_PHASH,),
    ).fetchone()
    # Resolve universe_filter and allocator from the rank-1 backtest spec.
    # universe_filter encodes the O'Donovan & Yu 2024 cost-mitigation cascade
    # rung: "full" = rung-2 (whole surface), "liquid" = rung-3 (liquid subset).
    # A missing / null filter means the rank-1 ran on the full surface without
    # the liquid subset selection, which is rung-2 - so it is normalized to
    # "full" here. Ch20 consumers filter on {"full","liquid"}; never emit a
    # null/sentinel that they would silently skip.
    _spec_row = _con.execute(
        "SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
        (TOP_HASH,),
    ).fetchone()
    if _spec_row is None or _spec_row[0] is None:
        raise RuntimeError(
            f"spec_json missing in backtest_runs for rank-1 {TOP_HASH}; "
            "cannot resolve allocator / universe_filter."
        )
    _spec = json.loads(_spec_row[0])
    UNIVERSE_FILTER = _spec.get("strategy", {}).get("signal", {}).get("universe_filter") or "full"
    CASCADE_RUNG = {"full": 2, "liquid": 3}[UNIVERSE_FILTER]
    ALLOCATOR_METHOD = _spec.get("strategy", {}).get("allocation", {}).get("method")
    # Read rather than restated, for the same reason as the allocator beside it: the
    # entry scheme is part of what the selected configuration IS, and three call sites below wrote it
    # as a literal that was only ever checked against the selected configuration of the day.
    SIGNAL_METHOD = _spec.get("strategy", {}).get("signal", {}).get("method")
    TOP_K = _spec.get("strategy", {}).get("signal", {}).get("top_k")
    if not SIGNAL_METHOD or TOP_K is None:
        raise RuntimeError(
            "rank-1 spec declares no strategy.signal method or top_k; "
            "the selected configuration's entry scheme cannot be reported"
        )

# %% [markdown]
# Query the cross-family selection cohort and the narrower family cohort
# separately so DSR and PBO retain their correct attribution.

# %%
_cohort_select = (
    "SELECT k_variants, dsr_raw, dsr_raw_pvalue, dsr_mp, dsr_mp_pvalue, "
    "dsr_er, dsr_er_pvalue, expected_max_sharpe_raw, "
    "expected_max_sharpe_mp, expected_max_sharpe_er, min_trl_periods_raw, "
    "min_trl_periods_mp, min_trl_periods_er, cm.pbo, cm.leader_hash, "
    "cm.pbo_n_combinations, cm.pbo_n_folds, t.config_name, cm.member_digest "
    "FROM cohort_metrics cm "
    "JOIN backtest_runs b ON b.backtest_hash = cm.leader_hash "
    "JOIN prediction_sets p ON p.prediction_hash = b.prediction_hash "
    "JOIN training_runs t ON t.training_hash = p.training_hash "
)
with sqlite3.connect(str(_db)) as _con:
    _search_row = _con.execute(
        _cohort_select + "WHERE cm.cohort_type='stagelabel' AND cm.stage=? AND cm.label=? "
        "AND cm.family IS NULL",
        (_lineage["val_stage"], PRIMARY_LABEL),
    ).fetchone()
    _family_row = _con.execute(
        _cohort_select
        + "WHERE cm.cohort_type='family' AND cm.stage=? AND cm.label=? AND cm.family=?",
        (_lineage["val_stage"], PRIMARY_LABEL, RANK1_FAMILY),
    ).fetchone()

# %% [markdown]
# Normalize a cohort row into the named fields used in the reporting
# cells below.


# %%
def _cohort_payload(row: tuple | None) -> dict | None:
    return (
        {
            "k_variants": row[0],
            "dsr_raw": row[1],
            "dsr_raw_pvalue": row[2],
            "dsr_mp": row[3],
            "dsr_mp_pvalue": row[4],
            "dsr_er": row[5],
            "dsr_er_pvalue": row[6],
            "expected_max_sharpe_raw": row[7],
            "expected_max_sharpe_mp": row[8],
            "expected_max_sharpe_er": row[9],
            "min_trl_periods_raw": row[10],
            "min_trl_periods_mp": row[11],
            "min_trl_periods_er": row[12],
            "pbo": row[13],
            "leader_hash": row[14],
            "pbo_n_combinations": row[15],
            "pbo_n_folds": row[16],
            "leader_config_name": row[17],
            "member_digest": row[18],
        }
        if row is not None
        else None
    )


# %% [markdown]
# Enforce exact cohort ownership before displaying search-wide uncertainty.

# %%
SEARCH_COHORT = _cohort_payload(_search_row)
FAMILY_COHORT = _cohort_payload(_family_row)
if SEARCH_COHORT is None:
    raise RuntimeError("Missing cross-family baseline cohort metrics")
SEARCH_ATTRIBUTION = cohort_metric_attribution(SEARCH_COHORT, TOP_HASH)
# The cohort is fetched at the selected configuration's own stage, so the leader it has to name is
# that stage's leader. This read `stage="signal"` while fetching the cohort at
# `_lineage["val_stage"]`, which agreed only while the selected configuration happened to be a
# baseline row. The current configuration is an allocation row, and the two sides of the comparison
# were then two different stages: the check reported a mismatch that was its own. Same population as
# the cohort above, and for the same reason: `best` with a stage alone ranges over every label and
# every universe in the registry, so the leader it returns can be a row the cohort never saw, and
# the check would then fail on the difference between two populations rather than on a real
# disagreement.
_stage_leader = explorer.best(
    stage=_lineage["val_stage"],
    top_n=1,
    label=PRIMARY_LABEL,
    prediction_hashes=_CANDIDATE_PREDICTIONS,
).row(0, named=True)
if SEARCH_COHORT["leader_hash"] != _stage_leader["backtest_hash"]:
    raise RuntimeError(
        f"Cross-family DSR leader does not match the {_lineage['val_stage']} leader: "
        f"{SEARCH_COHORT['leader_hash']} != {_stage_leader['backtest_hash']}"
    )
# K is the size of the search the DSR deflates for, so it has to be the size of the grid the
# selection ranged over. Two earlier versions of this check named it wrongly. It was first
# pinned at the literal 342 - the count from the pre-rebuild registry, when `12_backtest` swept
# the full and liquid universes together - which describes a search that no longer happens. It
# was then measured off the registry with `explorer.best`, which is worse than a stale literal:
# if part of the grid is missing, the recorded K and the measured count shrink together and the
# check passes on an under-deflated DSR, which is the one thing it exists to catch.
#
# The grid is declared, not measured, and it is the grid of the selected configuration's own stage.
# The nominees `13_portfolio_management` publishes span both stages the selection ranges over - the
# baseline grid `12_backtest` declared and the allocation grid built on top of it - and the DSR that
# deflates the selected configuration is the one computed over the stage it came from. Restricting
# the declared set to that stage is what keeps the three checks describing one search: the cohort,
# its leader, and the members it was computed over.
#
# This read `sp500-options-baseline-validation-v1` unconditionally, which is the same set whenever
# the selected configuration is a baseline row and a different search whenever it is not. Both
# populations are immutable and `require_complete` refuses a missing or partial member, so the count
# cannot shrink to meet a K that already has.
with sqlite3.connect(str(_db)) as _con:
    _stage_of = dict(
        _con.execute(
            "SELECT backtest_hash, stage FROM backtest_runs WHERE backtest_hash IN "
            f"({','.join('?' * len(_candidate_hashes))})",
            sorted(_candidate_hashes),
        ).fetchall()
    )
_baseline_members = tuple(
    sorted(h for h in _candidate_hashes if _stage_of.get(h) == _lineage["val_stage"])
)
if not _baseline_members:
    raise RuntimeError(
        f"{STRATEGY_CANDIDATES} declares no member at the selected configuration's stage "
        f"{_lineage['val_stage']!r}, so there is no declared grid to deflate against"
    )
OfficialPopulation.one(_selection_study, name=BASELINE_POPULATION).require_complete()
# The count alone is not enough. `cohort_metrics` records `member_digest`, an
# order-independent digest of the hashes the correction was actually computed over, precisely
# so a reader can establish that a stored DSR belongs to the cohort it is about to be reported
# against. A cohort that swapped a declared member for a retired or unrelated backtest has the
# same K and a different digest, and only the digest catches it.
_declared_digest = cohort_member_digest(_baseline_members)
if SEARCH_COHORT["member_digest"] is None:
    raise RuntimeError(
        "Cross-family DSR row predates member_digest, so the cohort it was computed over "
        f"cannot be established; recompute the cohort metrics for {STRATEGY_CANDIDATES}"
    )
if SEARCH_COHORT["member_digest"] != _declared_digest:
    raise RuntimeError(
        f"Cross-family DSR was computed over a cohort of {SEARCH_COHORT['k_variants']} "
        f"whose members are not the {len(_baseline_members)} that {STRATEGY_CANDIDATES} "
        f"declares at the {_lineage['val_stage']!r} stage "
        f"(digest {SEARCH_COHORT['member_digest']} != {_declared_digest}); the DSR must "
        "deflate for the complete grid the selection ranged over"
    )
PBO_REPORT = (
    reportable_pbo(FAMILY_COHORT["pbo"], FAMILY_COHORT["pbo_n_combinations"])
    if FAMILY_COHORT
    else {"value": None, "status": "unavailable", "n_combinations": None}
)
if FAMILY_COHORT and FAMILY_COHORT["leader_hash"] != SEARCH_COHORT["leader_hash"]:
    raise RuntimeError("Linear-family PBO and cross-family DSR name different leaders")

ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row

print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={PRIMARY_LABEL}")
print(f"        prediction_hash={TOP_PHASH}, val_backtest_hash={TOP_HASH}")
print(f"        holdout_backtest_hash={HO_HASH} (pred {HO_PHASH})")
print(
    f"        {_lineage['val_stage']} DSR leader={SEARCH_COHORT['leader_hash']} "
    f"(applies_to_carrier={SEARCH_ATTRIBUTION['applies_to_carrier']})"
)
print()
print("Daily-pooled IC (validation, prediction-side):")
print(f"  IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')}  (HAC, lag={int(ic_lag)})")
print(f"  t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3f}")
print(f"  n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f"  CI status: {ci_status(ic_lo, ic_hi)}")

# %% [markdown]
# Daily-pooled IC at the rank-1 prediction set is essentially zero
# and its HAC interval straddles zero. The p-value and positive-day share
# likewise provide no rank-correlation evidence. Section 3 reports a strategy
# Sharpe whose CI also straddles zero on the same prediction set, so
# the IC and Sharpe pictures agree. Both can be summarized
# straightforwardly: on `ret_to_expiry`, with full HTM costs, the
# pipeline does not resolve an edge.
#
# **Kill conditions** are encoded in `setup.yaml`:
#
# 1. `vrp_compression` - VRP < 2% annualized for > 6 months. Not
#    evaluated by this notebook (requires a parametric VRP series, not a
#    backtest registry read).
# 2. `gamma_loss_dominance` - rolling 2-year gamma P&L falls below
#    rolling 2-year VRP collection. Not evaluated for the same reason.
# 3. `cost_erosion` - round-trip option spread consumes > 50% of gross
#    VRP edge. The HTM construction already eliminates the exit-leg
#    spread (settle at intrinsic), so this gate is materially weakened
#    relative to the bps-cost variants - but the entry-leg spread alone
#    remains the dominant friction. The specification block reports the
#    current entry and hedge cost totals; the `15_costs` sweep quantifies
#    the slope against the share of the quoted spread paid on entry.
#
# In addition, the strategy-analysis notebook evaluates two universal gates in §9:
# (i) the validation Sharpe CI lower bound ≥ 0, and
# (ii) the holdout strategy-vs-EW paired CI does not exclude zero on
# the negative side. Both are reported as pass / partial / fail without
# verdict labels.

# %% [markdown]
# ## §2 Search context, family comparison, and lineage
#
# The equal-weight baseline sweep on `ret_to_expiry` covers four model families
# under HTM accounting. The rank-1 lineage anchors a multi-stage path
# (signal → allocation → optional risk overlay) at the training-hash
# level. § 5 says why the standard bps cost-grid is the wrong unit for
# option premium returns and points at the `15_costs` population, which
# runs the sweep in fractions of the quoted half-spread instead.

# %%
ctx = explorer.search_context("signal")
search_table = pl.DataFrame(
    [
        {"metric": "Total signal backtests (all labels)", "value": f"{ctx['total']:,}"},
        {"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
        {"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
        {"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
        {"metric": "% positive Sharpe (all labels)", "value": f"{ctx['pct_positive']:.1f}%"},
        {"metric": "Top-by-Sharpe across all labels", "value": f"{ctx['champion_sharpe']:.3f}"},
        {"metric": "Top-by-Sharpe label", "value": ctx["champion_source"]},
    ]
)
print("Equal-weight baseline search context:")
print(search_table)

# %% [markdown]
# The baseline search context above summarizes the `ret_to_expiry`
# cross-family distribution; the registered strategy uses HTM-specific
# cost handling (no bps-of-notional cost framework applies to option
# premium returns). The four legacy diagnostic labels (`fwd_ret_5d`,
# `fwd_ret_10d`, `fwd_ret_dh_5d`, `fwd_ret_dh_10d`) were dropped from
# the sweep 2026-05-17 - they ran through the vectorized backtest path
# which treats 5d/10d forward returns as daily returns, inflating
# Sharpes to non-credible levels.

# %%
# Family-level baseline Sharpe summary, restricted to ret_to_expiry -
# the only label where the rank-1 HTM strategy lives.
with sqlite3.connect(str(_db)) as _con:
    _famdf = pl.DataFrame(
        _con.execute(
            """
            SELECT
                t.family,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi
            FROM backtest_metrics bm
            JOIN backtest_runs b ON bm.backtest_hash = b.backtest_hash
            JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t ON p.training_hash = t.training_hash
            WHERE b.stage = 'signal'
              AND p.split = 'validation'
              AND t.label = 'ret_to_expiry'
              AND bm.sharpe IS NOT NULL
            """
        ).fetchall(),
        schema=["family", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"],
        orient="row",
    )

# %% [markdown]
# Summarize the complete baseline distribution within each model family.

# %%
family_summary = (
    _famdf.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_q25=pl.col("sharpe").quantile(0.25),
        sharpe_q75=pl.col("sharpe").quantile(0.75),
        sharpe_max=pl.col("sharpe").max(),
        pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
    )
    .sort("sharpe_median", descending=True)
)
print("Family-level baseline Sharpe summary (ret_to_expiry only):")
print(family_summary)

# %%
fig, ax = plt.subplots(figsize=(9, 4), layout="tight")
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()

ax.errorbar(
    medians,
    y,
    xerr=[medians - q25, q75 - medians],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    label="median ±IQR",
)
ax.scatter(maxima, y, marker="x", color="#C62828", s=60, label="max", zorder=5)
ax.axvline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
ax.set_title("Family-level Baseline (equal-weight) Sharpe (ret_to_expiry) - IQR + max")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
show_with_alt(
    fig,
    "Horizontal forest plot with one row per model family. Each row draws the family's median "
    "validation backtest Sharpe as a filled circle, a horizontal bar spanning its interquartile "
    "range, and a red cross at its single best configuration. A dashed vertical line marks zero. "
    "The plot is built to show two things at once: how wide each family's spread of "
    "configurations is, and how far its best run sits above its own median.",
)

# %% [markdown]
# Three families produce zero positive-Sharpe baseline rows on `ret_to_expiry`.
# Linear has two near-zero nonnegative rows among the complete baseline surface.
# The full-universe leader is comparison evidence only; the registered configuration
# is the liquid-universe rank-1. Neither resolves a statistically reliable edge.

# %%
# Lineage stages for the registered HTM rank-1
print(f"Lineage stages registered against prediction {TOP_PHASH} (HTM rank-1):")
with sqlite3.connect(str(_db)) as _con:
    rows = _con.execute(
        """
        SELECT b.stage, COUNT(*) AS n
        FROM backtest_runs b
        WHERE b.prediction_hash = ?
        GROUP BY b.stage
        """,
        (TOP_PHASH,),
    ).fetchall()
    for stage, n in rows:
        print(f"  {stage}: {n} runs")

print()
print(
    "All ret_to_expiry backtests route through the HTM daily-MTM cohort "
    "engine (`_run_htm_daily_mtm`). The equal-weight baseline selects the top-K "
    "cross-section (eq_w_topk); the allocation stage overlays a within-"
    "cross-section weighting on the HTM cohort accounting; the risk_overlay "
    "stage layers position-level controls on top of the allocator weights. "
    "The cost sweep lives in `15_costs`, in fractions of the quoted "
    "half-spread (the standard bps cost-grid is the wrong unit for option "
    "premium returns)."
)

# %% [markdown]
# ## §3 Headline performance with uncertainty
#
# The pinned configuration's performance metrics use 95% block-bootstrap CIs.
# Selection-bias adjustment (DSR raw / MP / ER, k_variants,
# expected_max_sharpe, min_trl_periods) lives in `cohort_metrics` per
# `memory/UNCERTAINTY_ARCHITECTURE.md` - `backtest_metrics` no longer
# carries those columns. The DSR uses the complete cross-family baseline
# cohort and names its full-universe leader. That leader differs from the pinned
# liquid configuration, so these statistics are not presented as configuration metrics.
# Linear-family PBO is displayed separately and suppressed because only two
# CSCV combinations are available. ER is the maintainer-recommended default;
# raw and MP are recorded for context.

# %%
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)

spec_block = {
    "case_study": CASE_STUDY,
    "family": RANK1_FAMILY,
    "config_name": RANK1_CONFIG,
    "label": PRIMARY_LABEL,
    "stage": _lineage["val_stage"],
    "signal_method": SIGNAL_METHOD,
    "allocator": ALLOCATOR_METHOD,
    "top_k": TOP_K,
    "universe_filter": UNIVERSE_FILTER,
    "cascade_rung": CASCADE_RUNG,
    "rebalance_cadence": setup["decision"]["entry_cadence"],
    "rebalance_step_weeks": setup["labels"]["rebalance_step"][PRIMARY_LABEL],
    "hedge_cadence": setup["decision"]["hedge_cadence"],
    "exit_rule": setup["decision"]["exit_time"],
    "cost_model": "HTM daily-MTM cohort: option entry spread + daily underlying hedge spread + commissions",
    "cumulative_entry_cost_premium_units": full["cumulative_entry_cost"],
    "cumulative_hedge_cost_premium_units": full["cumulative_hedge_cost"],
    "n_rebalance_dates": int(full["n_rebalance_dates"])
    if full["n_rebalance_dates"] is not None
    else None,
    "avg_cohorts_open": full["avg_cohorts_open"],
    "validation_window_periods": int(full["n_periods"]),
    "bootstrap_block_length": int(full["bootstrap_block_length"]),
    "bootstrap_n": int(full["bootstrap_n"]),
}
# The stage comes from the selected configuration's own lineage. Naming it "equal-weight baseline"
# was correct only while the cross-stage rank-1 happened to be a signal-stage row; the selected
# configuration in force is an allocation row, and a reader told otherwise has the wrong provenance
# for every number in the block.
print(
    f"Pinned-configuration specification ({_lineage['val_stage']} stage, "
    f"{ALLOCATOR_METHOD or 'equal-weight'} allocation, validation window):"
)
for k, v in spec_block.items():
    print(f"  {k}: {v}")


# %% [markdown]
# Format one uncertainty row consistently across the headline table.


# %%
def _row(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
    return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}


# %% [markdown]
# Collect the selected configuration's bootstrap intervals before adding the
# selection-adjusted diagnostics.

# %%
sharpe_status = ci_status(full["sharpe_ci95_lo"], full["sharpe_ci95_hi"])
sortino_status = ci_status(full["sortino_ci95_lo"], full["sortino_ci95_hi"])
ann_status = ci_status(full["ann_return_ci95_lo"], full["ann_return_ci95_hi"])
mdd_status = ci_status(full["max_dd_ci95_lo"], full["max_dd_ci95_hi"])
calmar_status = ci_status(full["calmar_ci95_lo"], full["calmar_ci95_hi"])
metric_specs = [
    ("Sharpe", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi", sharpe_status),
    ("Sortino", "sortino", "sortino_ci95_lo", "sortino_ci95_hi", sortino_status),
    ("Annualized return", "cagr", "ann_return_ci95_lo", "ann_return_ci95_hi", ann_status),
    ("Max drawdown", "max_drawdown", "max_dd_ci95_lo", "max_dd_ci95_hi", mdd_status),
    ("Calmar", "calmar", "calmar_ci95_lo", "calmar_ci95_hi", calmar_status),
]
performance_rows = [
    _row(label, _fmt(full[point]), _fmt(full[lo]), _fmt(full[hi]), status)
    for label, point, lo, hi, status in metric_specs
]
performance_rows.append(_row("PSR p-value (H0: SR≤0)", _fmt(full["psr_pvalue"]), "-", "-", "n/a"))

# %% [markdown]
# Attribute each DSR variant to the exact cross-family cohort and keep
# the insufficient family PBO visibly separate.

# %%
dsr_rows = []
for suffix, label in [("raw", "DSR_raw"), ("er", "DSR_ER"), ("mp", "DSR_MP")]:
    pvalue = SEARCH_COHORT[f"dsr_{suffix}_pvalue"]
    status = f"{suffix}_p={pvalue:.3f}" if pvalue is not None else "cohort_unavailable"
    dsr_rows.append(
        _row(
            f"{label} ({_lineage['val_stage']} leader {SEARCH_COHORT['leader_hash']})",
            _fmt(SEARCH_COHORT[f"dsr_{suffix}"]),
            "-",
            "-",
            status,
        )
    )

diagnostic_rows = [
    _row(
        "PBO (linear-family cohort)",
        _fmt(PBO_REPORT["value"]) if PBO_REPORT["value"] is not None else "insufficient",
        "-",
        "-",
        PBO_REPORT["status"],
    ),
    _row(
        "k_variants (cross-family search cohort)",
        str(int(SEARCH_COHORT["k_variants"])),
        "-",
        "-",
        "n/a",
    ),
]

# %% [markdown]
# Display the performance and selection diagnostics in one compact table.

# %%
headline = pl.DataFrame(performance_rows + dsr_rows + diagnostic_rows)
print("Selected-configuration performance and exactly attributed search diagnostics:")
print(headline)

# %% [markdown]
# The cross-family DSR row has K equal to the registered baseline count and names
# the complete-grid baseline
# leader, not the pinned liquid configuration. PBO is a separate linear-family
# diagnostic. With only two CSCV combinations, the notebook reports
# "insufficient combinations" instead of interpreting 0.50.

# %%
# Prepare the rank-1 metric stack and reference benchmark for a forest plot.
ew_val = load_benchmark_metrics(CASE_STUDY, PRIMARY_LABEL, period="validation")
ew_ho = load_benchmark_metrics(CASE_STUDY, PRIMARY_LABEL, period="holdout")
forest_metrics = [
    ("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
    ("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
    ("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
    ("Ann. return", full["cagr"], full["ann_return_ci95_lo"], full["ann_return_ci95_hi"]),
]

# %% [markdown]
# Plot the selected configuration intervals against zero and the validation equal-weight
# benchmark.

# %%
fig, ax = plt.subplots(figsize=(8, 4), layout="tight")
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.errorbar(
    points,
    y,
    xerr=[points - los, his - points],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    markersize=7,
)
ax.axvline(0, color="#9E9E9E", linestyle="--", linewidth=0.8)
ax.axvline(
    ew_val["sharpe"],
    color="#43A047",
    linestyle=":",
    linewidth=1.0,
    label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Value")
ax.set_title("Rank-1 (ret_to_expiry, HTM dispatch) Headline Metrics with 95% CIs")
ax.legend(loc="lower right", fontsize=8, frameon=False)
show_with_alt(
    fig,
    "Horizontal forest plot of four headline metrics for the rank-1 configuration, one row each: "
    "Sharpe, Sortino, Calmar and annualized return. Each row draws the point estimate as a circle "
    "with a bar spanning its 95 percent confidence interval. A dashed vertical line marks zero "
    "and a dotted green line marks the equal-weight benchmark's validation Sharpe. All four rows "
    "share one numeric axis labelled Value, so the benchmark line is comparable only to the "
    "Sharpe row.",
)

# %%
# Equity-curve overlay vs validation EW benchmark
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
    pl.read_parquet(strat_returns_path)
    .sort("timestamp")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)

bench_val = (
    load_benchmark_returns(CASE_STUDY, PRIMARY_LABEL, period="validation")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("ew_return").alias("benchmark"))
)

aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
    f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)

cum_strat = np.cumprod(1 + aligned["strategy"].to_numpy()) - 1
cum_bench = np.cumprod(1 + aligned["benchmark"].to_numpy()) - 1
fig, ax = plt.subplots(figsize=(10, 4.2), layout="tight")
ax.plot(aligned["ts"], cum_strat, color="#1565C0", linewidth=1.2, label="Rank-1 strategy (HTM)")
ax.plot(aligned["ts"], cum_bench, color="#43A047", linewidth=1.2, label="EW universe")
ax.axhline(0, color="#9E9E9E", linewidth=0.6, linestyle="--")
ax.set_ylabel("Cumulative return")
ax.set_title("Validation-window cumulative return: rank-1 HTM vs EW universe")
ax.legend(loc="best", frameon=False)
show_with_alt(
    fig,
    "Line chart of cumulative return across the validation overlay window, with two series: the "
    "rank-1 hold-to-maturity strategy and the equal-weight universe. Both are compounded from "
    "per-period returns and start at zero, and a dashed horizontal line marks zero, so the "
    "vertical gap between the lines at any date is the strategy's cumulative excess over the "
    "benchmark.",
)

# %% [markdown]
# The validation Sharpe straddles zero with a CI that spans roughly ±1
# Sharpe in either direction - the validation window does not resolve a
# directional edge for this strategy. Sortino, Calmar, and Ann. return
# all sit in the same straddles_zero status; PSR confirms the read.
# Whichever side of zero the point estimate falls on, the validation
# evidence is null-consistent. § 6's paired holdout vs EW test resolves
# whether the validation/EW gap holds out of sample. Selection accounting
# is supplied by the exact cross-family baseline cohort above (DSR raw /
# MP / ER and K from `cohort_metrics`), whose named leader differs from
# the selected configuration. Linear-family PBO is separately suppressed as underidentified.
# The two near-zero nonnegative baseline rows do not alter the unresolved
# selection-adjusted DSR_ER reading.

# %% [markdown]
# ## §3a Label scope and prior diagnostic exhibit
#
# The equal-weight baseline sweep on this case study now runs on a single label
# - `ret_to_expiry` - under HTM cost accounting (entry-side option
# spread + daily underlying hedge spread, plus an exit-leg option trade only
# where a contract's chain ends before its expiration date).
# Four legacy diagnostic labels (`fwd_ret_5d`, `fwd_ret_10d`,
# `fwd_ret_dh_5d`, `fwd_ret_dh_10d`) were dropped from the sweep on
# 2026-05-17: they routed through the vectorized backtest path that
# treats 5d/10d forward returns as daily returns, inflating Sharpes
# to non-credible levels.
# The structural diagnostic remains valid as prose: equity-style
# bps-of-notional cost accounting understates option spread cost by
# 1-2 orders of magnitude versus the actual % of premium, so any
# equity-style backtester on option premium returns produces
# pathological CAGR / drawdown numerics. The registered strategy
# therefore uses HTM cost accounting on `ret_to_expiry`, and under
# that accounting the cross-section does not resolve a statistically
# significant edge on the validation window.

# %% [markdown]
# ## §4 Risk and drawdown analysis
#
# The validation-window strategy returns paired against the validation
# EW benchmark. Tail-risk metrics come from the registry directly; the
# fold breakdown surfaces the across-fold dispersion that drives the
# bootstrap CI's width.

# %%
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()

pa = PortfolioAnalysis(
    returns=strat_arr,
    benchmark=bench_arr,
    dates=ts_arr,
    periods_per_year=PERIODS_PER_YEAR,
)

dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)

# %%
fig = plot_equity_drawdown(strat_returns_path)
show_with_alt(
    fig,
    "Two stacked panels sharing a date axis. The upper, taller panel plots the strategy's "
    "cumulative return over the backtest window. The lower panel plots drawdown, the fractional "
    "decline from the running peak of that same curve, as a red line over a shaded area, with the "
    "deepest point annotated. The pairing lets the length and depth of each decline be read "
    "against the part of the equity curve that produced it.",
)

# %%
roll = pa.compute_rolling_metrics(windows=[126], metrics=["sharpe", "beta"])
if isinstance(roll, dict):
    print("Rolling-window keys:")
    print({k: type(v).__name__ for k, v in roll.items()})

tail_table = pl.DataFrame(
    [
        {"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
        {"metric": "VaR 95% (daily)", "value": f"{full['var_95']:.4f}"},
        {"metric": "CVaR 95% (daily)", "value": f"{full['cvar_95']:.4f}"},
        {"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
        {"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
        {"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
    ]
)
print()
print("Tail risk profile:")
print(tail_table)

# %%
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
print(f"Per-fold breakdown ({fold_df.height} folds):")
print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
print()
if fold_df.height > 1:
    print(f"Fold Sharpe range: [{fold_df['sharpe'].min():.3f}, {fold_df['sharpe'].max():.3f}]")
    print(f"Fold Sharpe std:   {fold_df['sharpe'].std():.3f}")

# %% [markdown]
# The two fold Sharpes straddle zero and both sit inside the wide
# bootstrap Sharpe CI in Section 3. Tail risk is heavy: negative skew
# and high kurtosis reflect the asymmetric P&L profile
# of a short straddle book - long stretches of small premium decay
# punctuated by large losses when realized vol exceeds implied.
# CVaR and max drawdown are severe, the defining signature
# of the strategy: even with hedging and HTM settlement, a single
# bad-vol regime can wipe out cumulative premium.

# %% [markdown]
# ## §5 Friction budget
#
# The standard `BacktestExplorer.cost_sensitivity()` curve is denominated in bps per leg, which is
# the right convention for equities and futures but the wrong unit for options: a 10% spread on a
# 4% premium is 40 bps of notional, well inside the "positive Sharpe" zone of any equity sweep and
# ruinous for the actual P&L. `15_costs` runs the cost sweep in fractions of the quoted half-spread
# instead, executes it through the same engine as every other backtest, and publishes it as the
# named population `sp500-options-cost-sensitivity-validation-v1`. It reports the resulting Sharpe
# surface across families, universes and spread fractions there, each point with its
# block-bootstrap Sharpe interval, so this notebook does not repeat it.

# %% [markdown]
# ## §6 Holdout closure with paired bootstrap
#
# Two paired tests anchor the holdout read. (i) The holdout rank-1
# versus the validation rank-1 ("did Sharpe hold?"); (ii) the holdout
# rank-1 versus the holdout-window equal-weight benchmark ("did the
# strategy beat random in the holdout?"). Numbers come from
# `backtest_paired_metrics` rows populated for this fixed case lineage;
# never from val_sharpe minus holdout_sharpe arithmetic.

# %%
ho_full = load_backtest_metrics(CASE_STUDY, backtest_hash=HO_HASH).row(0, named=True)
val_full = full

print(f"Validation rank-1 hash: {TOP_HASH} (Sharpe {val_full['sharpe']:+.4f})")
print(f"Holdout rank-1 hash:    {HO_HASH} (Sharpe {ho_full['sharpe']:+.4f})")
print()
print(
    f"Holdout window: {setup['evaluation']['holdout_start']} → {setup['evaluation']['holdout_end']}"
)
print(f"Holdout n_periods: {int(ho_full['n_periods'])}")

# %%
val_ho_pair = load_paired_metrics(
    CASE_STUDY,
    challenger_hash=HO_HASH,
    benchmark_kind="val_rank1_self",
)
if val_ho_pair.is_empty():
    raise RuntimeError(
        "Missing required val_rank1_self paired metrics for current holdout "
        f"{HO_HASH}; populate this case before executing notebook 16"
    )
vh = val_ho_pair.row(0, named=True)

# %% [markdown]
# Format one paired validation-to-holdout metric, preserving its
# bootstrap interval and p-value when available.


# %%
def _diff_row(
    label: str,
    v: float,
    h: float,
    diff: float | None,
    lo: float | None,
    hi: float | None,
    p: float | None,
) -> dict:
    return {
        "metric": label,
        "validation": _fmt(v, ".4f"),
        "holdout": _fmt(h, ".4f"),
        "diff (h−v)": _fmt(diff, ".4f") if diff is not None else "-",
        "diff CI95": f"[{_fmt(lo, '.4f')}, {_fmt(hi, '.4f')}]" if lo is not None else "-",
        "p-value": _fmt(p, ".4f") if p is not None else "-",
    }


# %% [markdown]
# Build the disjoint-window comparison from the stored paired-bootstrap
# row rather than subtracting headline metrics.

# %%
pair_specs = [
    ("Sharpe", "sharpe", "sharpe_diff", "sharpe_diff_ci95", "p_value"),
    ("Annualized return", "cagr", "ret_diff", "ret_diff_ci95", None),
    ("Max drawdown", "max_drawdown", "max_dd_diff", "max_dd_diff_ci95", None),
]
pair_rows = [
    _diff_row(
        label,
        val_full[metric],
        ho_full[metric],
        vh[diff],
        vh[f"{ci_prefix}_lo"],
        vh[f"{ci_prefix}_hi"],
        vh[pvalue] if pvalue else None,
    )
    for label, metric, diff, ci_prefix, pvalue in pair_specs
]
pair_rows.append(
    _diff_row(
        "Information ratio",
        None,
        None,
        vh["info_ratio"],
        vh["info_ratio_ci95_lo"],
        vh["info_ratio_ci95_hi"],
        None,
    )
)
val_ho_table = pl.DataFrame(pair_rows)

# %% [markdown]
# Display the paired decay statistics and state the disjoint-window
# bootstrap construction.

# %%
print("val → holdout paired-bootstrap decay (rank-1 self):")
print(val_ho_table)
print(
    "prob_challenger_wins: "
    + (f"{vh['prob_challenger_wins']:.3f}" if vh["prob_challenger_wins"] is not None else "-")
)
print(f"CI status (Sharpe diff): {ci_status(vh['sharpe_diff_ci95_lo'], vh['sharpe_diff_ci95_hi'])}")
print()
print(
    "Note: validation and holdout windows are disjoint. The paired-metrics "
    "producer therefore uses independent stationary block-bootstrap draws "
    "for the two windows rather than calendar alignment."
)

# %%
ho_vs_ew = load_paired_metrics(
    CASE_STUDY,
    challenger_hash=HO_HASH,
    benchmark_kind="equal_weight_holdout_side_artifact",
)
if ho_vs_ew.is_empty():
    raise RuntimeError(
        "Missing required equal_weight_holdout_side_artifact paired metrics for "
        f"current holdout {HO_HASH}; populate this case before executing notebook 16"
    )
he = ho_vs_ew.row(0, named=True)

print("Holdout strategy vs holdout-window EW universe:")
print(f"  strategy Sharpe:  {ho_full['sharpe']:+.4f}")
print(f"  EW Sharpe:        {ew_ho['sharpe']:+.4f}")
print(
    "  diff Sharpe: "
    f"{_fmt_ci(he['sharpe_diff'], he['sharpe_diff_ci95_lo'], he['sharpe_diff_ci95_hi'], '.4f')}"
)
print("  p_value:                " + (f"{he['p_value']:.4f}" if he["p_value"] is not None else "-"))
print(
    "  prob_challenger_wins:   "
    + (f"{he['prob_challenger_wins']:.3f}" if he["prob_challenger_wins"] is not None else "-")
)
print(
    f"  info_ratio (strategy vs EW): "
    f"{_fmt_ci(he['info_ratio'], he['info_ratio_ci95_lo'], he['info_ratio_ci95_hi'])}"
)
print(f"  CI status: {ci_status(he['sharpe_diff_ci95_lo'], he['sharpe_diff_ci95_hi'])}")

# %% [markdown]
# **Reading.** The two paired tables above are the authoritative holdout
# evidence for the fixed current configuration. Point-estimate ordering alone
# does not establish persistence or benchmark superiority; the paired
# confidence intervals and p-values determine whether either difference
# is resolved. The 2021 window is reported once and is not reused for
# model or strategy selection.

# %% [markdown]
# ## §7 Benchmark-aware diagnostics
#
# Layer 1 reports the universal alpha/beta/IR profile via
# `PortfolioAnalysis` against the equal-weight S&P 500 options universe.
# Layer 2 - equity factor attribution (FF5+MOM) - is in scope per the
# case-study scope deviation table for equity-class case studies, but the model
# fit is structurally compromised on options: FF5+MOM captures linear
# equity factor exposure, while option premium returns are driven by
# vega, gamma, and theta - non-linear payoffs that are orthogonal to
# linear factor loadings. We include the regression for completeness
# but flag the model mismatch in the reading.

# %%
metrics = pa.compute_summary_stats()
attr_df = pl.DataFrame(
    [
        {"metric": "alpha (annualized)", "value": _fmt(getattr(metrics, "alpha", None))},
        {"metric": "beta", "value": _fmt(getattr(metrics, "beta", None), ".3f")},
        {"metric": "information ratio", "value": _fmt(getattr(metrics, "information_ratio", None))},
        {"metric": "tracking error", "value": _fmt(getattr(metrics, "tracking_error", None))},
        {"metric": "up capture", "value": _fmt(getattr(metrics, "up_capture", None))},
        {"metric": "down capture", "value": _fmt(getattr(metrics, "down_capture", None))},
    ]
)
print("Layer 1: rank-1 vs validation EW universe (PortfolioAnalysis):")
print(attr_df)

# %%
# Placebo regression: residual α and HAC t-stat on EW alone.
import statsmodels.api as sm
from case_studies.utils.warning_policy import apply_notebook_warning_policy

apply_notebook_warning_policy()

X = sm.add_constant(bench_arr)
ols = sm.OLS(strat_arr, X).fit(cov_type="HAC", cov_kwds={"maxlags": 5})
alpha_daily = ols.params[0]
alpha_t = ols.tvalues[0]
alpha_p = ols.pvalues[0]
beta_ew = ols.params[1]
beta_t = ols.tvalues[1]
print()
print("Placebo regression vs EW universe (HAC, maxlags=5):")
print(f"  α (daily) = {alpha_daily:.6f}, α (annualized) = {alpha_daily * PERIODS_PER_YEAR:.4f}")
print(f"  α t-stat = {alpha_t:.3f}, p = {alpha_p:.3f}")
print(f"  β        = {beta_ew:.3f}, β t-stat = {beta_t:.3f}")
print(f"  CI status (α): {'excludes_zero_strong' if alpha_p < 0.05 else 'straddles_zero'}")

# %%
# Layer 2: FF5+MOM factor regression with HAC standard errors.
strategy_rets_pd = pd.Series(
    strat_arr,
    index=pd.to_datetime(aligned["ts"].to_list()),
    name="strategy",
)
_period_start = str(strategy_rets_pd.index.min().date())
_period_end = str(strategy_rets_pd.index.max().date())

_factors = load_factor_data(start=_period_start, end=_period_end)
_reg = run_factor_regression(
    strategy_rets_pd,
    _factors,
    model="ff5_mom",
    hac_lags=5,
    dollar_neutral=False,
)

print()
print("Layer 2: FF5+MOM factor attribution (HAC, daily):")
print(f"  n_obs:           {_reg['n_obs']}")
print(f"  R²:              {_reg['r_squared']:.3f}")
print(
    f"  α (annualized):  {_reg['alpha_annualized']:+.4f}  "
    f"(t={_reg['alpha_t_stat']:.2f}, p={_reg['alpha_p_value']:.3f})"
)
print(f"  Strategy Sharpe: {_reg['strategy_sharpe']:+.3f}")
print(f"  Residual Sharpe: {_reg['residual_sharpe']:+.3f}")
print()
print("Factor betas (HAC):")
for factor, beta in _reg["betas"].items():
    t = _reg["t_stats"][factor]
    sig = "*" if _reg["p_values"][factor] < 0.05 else ""
    print(f"  {factor:8s}: {beta:+.4f}  (t={t:+.2f}){sig}")

# %%
# Block bootstrap CIs on alpha and factor betas
_boot = compute_bootstrap_ci(
    strategy_rets_pd,
    _factors,
    model="ff5_mom",
    n_boot=500,
    block_size=20,
    dollar_neutral=False,
    seed=42,
)

if _boot.get("n_boot", 0) > 0:
    print(f"\nBootstrap CIs (n={_boot['n_boot']}, block=20 days):")
    print(
        f"  α (annualized): [{_boot['alpha_ann_lo']:+.4f}, {_boot['alpha_ann_hi']:+.4f}] (95% CI)"
    )
    for factor in _reg["factor_columns"]:
        key_lo = f"{factor}_lo"
        key_hi = f"{factor}_hi"
        if key_lo in _boot:
            print(f"  {factor:8s}: [{_boot[key_lo]:+.4f}, {_boot[key_hi]:+.4f}]")

# %%
fig_attr = plot_attribution_waterfall(
    _reg, title="S&P 500 Options HTM: FF5+MOM Attribution (diagnostic)"
)
show_with_alt(
    fig_attr,
    "Bar chart with one bar per FF5 factor plus momentum and a final gray Residual bar, against a "
    "y axis labelled Sharpe Contribution, each bar annotated with its value and a dashed "
    "horizontal line marking the strategy's total Sharpe. The factor bars are each coefficient's "
    "absolute value as a share of the total, rescaled to the factor-explained part of Sharpe, so "
    "they show relative exposure and not an exact return decomposition.",
)

# %% [markdown]
# Layer-1 placebo regression vs EW: α (annualized) and HAC t-stat read
# directly off the print block above. Layer-2 FF5+MOM produces a low R²
# by design - the linear factor model captures whatever residual equity
# delta leaks through the daily delta hedge, but the dominant return
# drivers (gamma, vega, theta) sit outside the FF5+MOM linear span.
# The α point estimate is therefore not interpretable as
# "factor-adjusted edge" for an options strategy in the same sense as
# for an equity portfolio; it's a residual after a structurally
# incomplete factor model. The diagnostic value is in the *betas*: a
# significant Mkt-RF loading would indicate incomplete delta hedging
# (the daily rehedge with a 0.10 net-delta band leaks some directional
# exposure); a significant MOM loading would suggest the model selects
# straddles on names whose option markets lag equity momentum (a
# microstructure artefact).
#
# A volatility-native attribution (VIX, VVIX, term-structure slope,
# variance risk premium factors) is the right model for this strategy.
# That construction is a Ch20 extension (the dispersion-trading
# literature provides it); per strategy-analysis convention this notebook reports
# what's available without synthesizing a new factor m

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: MIT

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.