Pular para o conteúdo
Todos os documentos da biblioteca

Análise de incerteza em estratégia de straddle de opções do S&P 500

Código Machine Learning for Trading

Resumo

Este notebook avalia uma estratégia semanal de straddle de opções do S&P 500 usando backtests registrados e um processo de seleção fixo. Ele recupera a configuração selecionada no registro após aplicar uma restrição de universo líquido e, então, avalia os resultados de validação e holdout sem usar o holdout para escolher a estratégia. O backtest usa hedge delta diário, liquidação no vencimento e custos de spread por perna de opção e hedge. As métricas relatadas incluem intervalos de confiança por bootstrap em blocos, e comparações pareadas são usadas para a deterioração no holdout e a comparação com um benchmark de pesos iguais.

O documento caracteriza o coeficiente de informação das previsões e o Sharpe da estratégia como estatisticamente compatíveis com zero para o rótulo selecionado de retorno até o vencimento; portanto, não identifica uma vantagem definida no período de validação. Sua principal contribuição é um método disciplinado de avaliação que preserva a incerteza e contabiliza os custos, e não evidências de uma estratégia rentável. Os resultados continuam dependendo do conjunto declarado de candidatas, das restrições de universo, das premissas do backtest e da amostra registrada; o próprio notebook não treina modelos nem refaz backtests.

Ideias principais

  • A estratégia é selecionada de um universo líquido indicado antes da avaliação do holdout.
  • Os backtests contabilizam spreads de entrada das opções, spreads do hedge no ativo subjacente, hedge delta e liquidação no vencimento.
  • Intervalos de confiança por bootstrap em blocos e comparações pareadas expressam a incerteza do desempenho e da deterioração no holdout.
  • As métricas relatadas para previsões e estratégia são compatíveis com zero, portanto a análise não identifica uma vantagem.
  • As conclusões dependem do conjunto de candidatas, das premissas de custos e dos backtests registrados disponíveis.

Tags

Texto completo
# 18_strategy_analysis.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # S&P 500 Options - Strategy Analysis
#
# This notebook converts the case-study backtest registry for the S&P
# 500 options HTM straddle strategy into a per-case-study strategy
# assessment. Which configuration carries it is resolved from the registry in
# §1 and printed there, not named here: it is a property of the run and it moves
# whenever the sweep is rebuilt, and a name written into prose agrees with the
# registry only until the next rebuild. What is fixed is the selection contract
# - the liquid-universe pin is applied first, then the surface is ranked on
# validation, and the holdout is not consulted.
# All `ret_to_expiry` backtests dispatch
# through the HTM daily-MTM cohort engine (`_run_htm_daily_mtm`): entry on
# the final available session of each Friday week, the underlying delta
# hedged whenever it breaches its threshold, settle at intrinsic value at
# expiration, full per-leg costs (entry-side option spread, underlying hedge
# spread on each session the hedge trades, and an exit-leg option trade only
# where a contract's chain ends before its expiration date). Every metric is reported with its
# block-bootstrap 95% CI; paired holdout comparisons via
# `backtest_paired_metrics` are required before the notebook reports them.
# Cross-case-study comparison is reserved for Chapter 20.
#
# **Learning objectives**
#
# - Read uncertainty-aware backtest metrics for an options strategy whose
#   signal is statistically null (Sharpe and IC both straddle zero) under
#   HTM cost accounting on full per-leg friction.
# - Surface a holdout decay reading without invoking champion/winner/
#   verdict language, and a holdout-vs-EW comparison for the same window.
#
# **Book reference**: Chapter 20, §20.1 - sp500_options anchors the
# "cost model validity" theme in the cross-case-study synthesis.
#
# **Prerequisites**: the case-study pipeline through
# [`17_holdout_backtest`](17_holdout_backtest.ipynb), which registers the holdout result this
# notebook closes on.
#
# **Scope**: no training and no re-backtesting. It does write two derived tables it also reads -
# `backtest_paired_metrics` and `cohort_metrics` - because a case study has to be able to produce
# everything its own notebooks report, and only Chapter 20's synthesis ever wrote them. Both are
# computed from the registered backtests over the population this notebook nominates, and both are
# replaced rather than added to, so a selected configuration that moved does not leave its
# predecessor's rows behind looking current.

# %%
"""S&P 500 Options - Strategy Analysis."""

# ruff: noqa: E402, I001

import json
import sqlite3

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl

# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch  # noqa: F401
import yaml


from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
    BacktestReportMetadata,
    generate_tearsheet_from_run_artifacts,
)

# %% [markdown]
# Load the registry readers and reporting helpers the notebook draws on.

# %%
from case_studies.research import CandidateSet, OfficialPopulation, Study
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.cohort_reporting import cohort_metric_attribution, reportable_pbo
from case_studies.utils.paired_metrics import populate_paired_metrics, rung_for
from case_studies.utils.uncertainty import ENTIRE_REGISTRY, cohort_member_digest
from case_studies.utils.factor_attribution import (
    compute_bootstrap_ci,
    load_factor_data,
    plot_attribution_waterfall,
    run_factor_regression,
)
from case_studies.utils.registry import (
    load_backtest_fold_metrics,
    load_backtest_metrics,
    load_paired_metrics,
)
from case_studies.utils.strategy_analysis import (
    ci_status,
    compute_operating_profile,
    fmt_gate,
    gate1_validation_sharpe_geq_zero,
    gate2_holdout_diff_not_excludes_zero_negatively,
    gate_passes,
    plot_equity_drawdown,
    resolve_solvent_carrier,
    write_strategy_assessment,
)
from utils.paths import get_case_study_dir, get_output_dir
from utils.style import show_with_alt

# %% tags=["parameters"]
MAX_SYMBOLS = 0

# %%
CASE_STUDY = "sp500_options"
# The immutable grid `12_backtest` publishes; the DSR below deflates for its members.
BASELINE_POPULATION = "sp500-options-baseline-validation-v1"
# The nominees `13_portfolio_management` publishes; the selected configuration has to be one of
# them.
STRATEGY_CANDIDATES = "sp500-options-strategy-candidates-v1"
PRIMARY_LABEL = "ret_to_expiry"  # registered HTM strategy label (Appendix A)
PERIODS_PER_YEAR = 252
CASE_DIR = get_case_study_dir(CASE_STUDY)
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

with open(CASE_DIR / "config" / "setup.yaml") as f:
    setup = yaml.safe_load(f)

explorer = BacktestExplorer(CASE_STUDY)
_selection_study = Study.open(CASE_STUDY)
print(explorer)

# %% [markdown]
# Format point estimates and confidence intervals consistently throughout
# the notebook.


# %%
def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
    if point is None:
        return "-"
    p = format(point, fmt)
    if lo is None or hi is None:
        return f"{p} [-, -]"
    return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"


# %% [markdown]
# Format scalar metrics while preserving missing values.


# %%
def _fmt(val: float | None, fmt: str = ".4f") -> str:
    return "-" if val is None else format(val, fmt)


# %% [markdown]
# ## §1 Handoff from model analysis
#
# The strategy phase inherits an HTM short-straddle pipeline: weekly
# entry on a top-k cross-section of `ret_to_expiry`, daily delta hedge
# through the underlying, hold to expiry, full per-leg costs on entry-side
# option spread plus daily underlying hedge spread, and an exit-leg option
# trade only where a contract's chain ends before its expiration date. The
# family, configuration and allocator that carry it are resolved from the
# registry below and printed alongside their identifiers rather than named
# here. The
# liquid universe pin (`UNIVERSE_RESTRICTIONS`) excludes the higher-Sharpe
# full-universe rows from rank-1 selection, so the deployed configuration stays on
# the liquid subset and remains poolable with its holdout replay. The selected configuration
# is fixed before the one holdout evaluation. The upstream
# prediction-side IC and the downstream strategy Sharpe are both
# statistically consistent with zero - there is no IC-vs-Sharpe disconnect
# on this strategy; both metrics agree that no edge has been resolved on
# `ret_to_expiry` over the validation window. `LABEL_RESTRICTIONS` excludes
# the four diagnostic `fwd_ret_*` variants from rank-1 selection (and those
# rows have been dropped from the registry entirely as of the 2026-05-17
# sweep cleanup); the only label that anchors a registered rank-1 here is
# `ret_to_expiry` under HTM cost accounting.

# %%
# Rank-1 = the registered HTM strategy on ret_to_expiry. Resolved
# dynamically from the registry with LABEL_RESTRICTIONS applied so any
# residual diagnostic-variant rows (fwd_ret_5d/_10d/_dh_*) - which inflated
# Sharpes under bps-of-notional costs and cannot anchor a registered HTM
# strategy - are excluded from the cross-stage rank-1 lookup. The val/
# holdout pair shares a training_hash; the notebook reads from those rows
# rather than recomputing.
# The nominated field is resolved BEFORE the ranking and handed to it, not checked against
# the winner afterwards. Those are different tests wherever the conformal branch is taken:
# the common-support re-ranking restricts every series to the timestamps they all share, so a
# row that is never going to win still decides how far that intersection reaches, and
# therefore which admitted candidate does. A membership check on the winner passes while the
# answer has already been moved by a row that was never eligible.
_candidate_hashes = frozenset(CandidateSet.one(_selection_study, name=STRATEGY_CANDIDATES).members)
# `resolve_solvent_carrier` rather than the bare lineage resolver: it applies the same selection and
# additionally refuses a selected configuration whose equity reached zero, whose Sharpe is
# arithmetic on a balance that no longer exists.
_lineage = resolve_solvent_carrier(CASE_STUDY, admitted=_candidate_hashes)
TOP_HASH = _lineage["val_backtest_hash"]
TOP_PHASH = _lineage["val_prediction_hash"]
HO_HASH = _lineage["holdout_backtest_hash"]
HO_PHASH = _lineage["holdout_prediction_hash"]
RANK1_FAMILY = _lineage["family"]
RANK1_CONFIG = _lineage["config_name"]
assert _lineage["label"] == PRIMARY_LABEL, (
    f"Canonical rank-1 label {_lineage['label']!r} != PRIMARY_LABEL "
    f"{PRIMARY_LABEL!r} - LABEL_RESTRICTIONS misconfigured?"
)
assert HO_HASH is not None, (
    "No holdout backtest for the canonical val rank-1. Run "
    "16_holdout_predictions and 17_holdout_backtest, which refit the selected configuration over the "
    "holdout interval and register the result under a training identity of its own. Do not "
    "reach for 20_strategy_synthesis/holdout.py::generate_holdout: it registers its "
    "predictions under the VALIDATION training identity, so what it produces is a "
    "validation-fitted model scored on the holdout window, which is the one thing the "
    "holdout exists to rule out."
)
# The field was applied before the ranking, so this cannot fail on a row from outside it. It stays
# because it is cheap and because it is the assertion a reader needs: the selected configuration
# published below is one of the nominees `13_portfolio_management` froze, not whatever happened to
# rank first in the registry.
if TOP_HASH not in _candidate_hashes:
    raise RuntimeError(
        f"Canonical rank-1 {TOP_HASH} is not a member of {STRATEGY_CANDIDATES} "
        f"({len(_candidate_hashes)} nominees), so the selected configuration was ranked from outside the "
        "set this case study nominated"
    )

# The prediction sets behind the nominated backtests, resolved once and used wherever a
# cohort or a leader is taken below. Both of those questions are about the search this case
# study declared, and the registry holds more than that search: retired generations, the
# full-universe rows the Chapter 18 cost cascade keeps, and any label this notebook does not
# report. Restricting on the universe alone leaves all three in, and a cohort that admits
# them has a different leader and a different deflation than the search being reported.
_db = CASE_DIR / "run_log" / "registry.db"
with sqlite3.connect(str(_db)) as _con:
    _CANDIDATE_PREDICTIONS = tuple(
        sorted(
            row[0]
            for row in _con.execute(
                "SELECT DISTINCT prediction_hash FROM backtest_runs WHERE backtest_hash IN "
                f"({','.join('?' * len(_candidate_hashes))})",
                sorted(_candidate_hashes),
            )
        )
    )
if not _CANDIDATE_PREDICTIONS:
    raise RuntimeError(
        f"{STRATEGY_CANDIDATES} names {len(_candidate_hashes)} backtests and none of them is "
        "in backtest_runs, so the nominated search cannot be resolved to a prediction "
        "population"
    )
print(
    f"Nominated search: {len(_candidate_hashes)} backtests over {len(_CANDIDATE_PREDICTIONS)} prediction sets"
)

# %% [markdown]
# ### The paired rows this notebook reads, built here rather than assumed
#
# §6 reads two `backtest_paired_metrics` kinds keyed on the holdout backtest, and it used to
# load rows that only `20_strategy_synthesis/01_aggregate_synthesis.py` ever wrote. On a
# registry that has been reset, or one whose configuration moved, those rows are absent or belong to
# a holdout hash that no longer exists, and the notebook stopped on a message telling the
# reader to populate the case elsewhere. A case study's own pipeline has to be able to produce
# everything its own notebooks report.
#
# `populate_paired_metrics` is the same producer Chapter 20 calls, and it is given the same
# selection this notebook made: the configuration from the lineage resolver, and the rung
# this case study is pinned to. The pin matters here - rung-1 (mid-to-mid bps) and rung-2
# (full-universe HTM) both carry `universe_filter="full"`, so ranking on the universe alone would
# let either stand in for the liquid configuration the case study actually deploys.
#
# `replace_all` makes the call a snapshot. Registration is an upsert keyed on `(challenger_hash,
# benchmark_hash)`, so it can add this configuration's rows and cannot remove the rows of a selected
# configuration it replaced; those rows name a holdout hash that is no longer the holdout and would
# sit in the table looking current. This case study restricts to one label and one rung, so what
# the call writes is the whole of what belongs here, which is the condition under which pruning to
# it is right.

# %%
# `cohort_metrics` is the same story one table over: this notebook reads the cross-family
# baseline cohort and the deflated Sharpe computed over it, and only Chapter 20 ever wrote
# them. `compute_and_register` is that producer, and it refreshes the whole table rather than
# one row, so it also removes a cohort led by a backtest the registry no longer holds. The
# universe pin goes with it - the full-universe rows are kept for the Chapter 18 cost cascade
# and must not enter a cohort the selected configuration is deflated against.
_cohort_counts = compute_and_register(
    CASE_STUDY,
    universe_filter=rung_for(CASE_STUDY)["universe_filter"],
    prediction_hashes=_CANDIDATE_PREDICTIONS,
    verbose=False,
)
print(
    f"cohort_metrics: {sum(v for k, v in _cohort_counts.items() if k not in ('errors', 'dangling_pruned'))} "
    f"rows across {sorted(k for k in _cohort_counts if k not in ('errors', 'dangling_pruned'))}"
    f", {_cohort_counts['dangling_pruned']} dangling pruned, {_cohort_counts['errors']} errors"
)

# %%
_paired_rows = populate_paired_metrics(
    CASE_STUDY,
    explorer,
    label_restriction=frozenset({PRIMARY_LABEL}),
    rung=rung_for(CASE_STUDY),
    carrier=_lineage,
    # The cohort call above is scoped to `_CANDIDATE_PREDICTIONS` and this one is not:
    # the pairs are selected from every registered prediction set. That disagreement
    # was there before the scope had to be stated; narrowing it changes published
    # numbers, so it is a separate decision from this line.
    prediction_hashes=ENTIRE_REGISTRY,
    periods_per_year=PERIODS_PER_YEAR,
    replace_all=True,
    verbose=False,
)
_written = [row for row in _paired_rows if "skip" not in row]
print(f"backtest_paired_metrics: {len(_written)} pairs written, {len(_paired_rows)} attempted")
for _row in _paired_rows:
    if "skip" in _row:
        print(f"  skipped {_row.get('benchmark_kind', '?')}: {_row['skip']}")

# %% [markdown]
# Resolve the selected configuration's prediction metrics and registered strategy
# specification directly from the registry.

# %%
with sqlite3.connect(str(_db)) as _con:
    _row = _con.execute(
        "SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
        "ic_hac_lag, ic_pct_positive "
        "FROM prediction_metrics WHERE prediction_hash = ?",
        (TOP_PHASH,),
    ).fetchone()
    # Resolve universe_filter and allocator from the rank-1 backtest spec.
    # universe_filter encodes the O'Donovan & Yu 2024 cost-mitigation cascade
    # rung: "full" = rung-2 (whole surface), "liquid" = rung-3 (liquid subset).
    # A missing / null filter means the rank-1 ran on the full surface without
    # the liquid subset selection, which is rung-2 - so it is normalized to
    # "full" here. Ch20 consumers filter on {"full","liquid"}; never emit a
    # null/sentinel that they would silently skip.
    _spec_row = _con.execute(
        "SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
        (TOP_HASH,),
    ).fetchone()
    if _spec_row is None or _spec_row[0] is None:
        raise RuntimeError(
            f"spec_json missing in backtest_runs for rank-1 {TOP_HASH}; "
            "cannot resolve allocator / universe_filter."
        )
    _spec = json.loads(_spec_row[0])
    UNIVERSE_FILTER = _spec.get("strategy", {}).get("signal", {}).get("universe_filter") or "full"
    CASCADE_RUNG = {"full": 2, "liquid": 3}[UNIVERSE_FILTER]
    ALLOCATOR_METHOD = _spec.get("strategy", {}).get("allocation", {}).get("method")
    # Read rather than restated, for the same reason as the allocator beside it: the
    # entry scheme is part of what the selected configuration IS, and three call sites below wrote it
    # as a literal that was only ever checked against the selected configuration of the day.
    SIGNAL_METHOD = _spec.get("strategy", {}).get("signal", {}).get("method")
    TOP_K = _spec.get("strategy", {}).get("signal", {}).get("top_k")
    if not SIGNAL_METHOD or TOP_K is None:
        raise RuntimeError(
            "rank-1 spec declares no strategy.signal method or top_k; "
            "the selected configuration's entry scheme cannot be reported"
        )

# %% [markdown]
# Query the cross-family selection cohort and the narrower family cohort
# separately so DSR and PBO retain their correct attribution.

# %%
_cohort_select = (
    "SELECT k_variants, dsr_raw, dsr_raw_pvalue, dsr_mp, dsr_mp_pvalue, "
    "dsr_er, dsr_er_pvalue, expected_max_sharpe_raw, "
    "expected_max_sharpe_mp, expected_max_sharpe_er, min_trl_periods_raw, "
    "min_trl_periods_mp, min_trl_periods_er, cm.pbo, cm.leader_hash, "
    "cm.pbo_n_combinations, cm.pbo_n_folds, t.config_name, cm.member_digest "
    "FROM cohort_metrics cm "
    "JOIN backtest_runs b ON b.backtest_hash = cm.leader_hash "
    "JOIN prediction_sets p ON p.prediction_hash = b.prediction_hash "
    "JOIN training_runs t ON t.training_hash = p.training_hash "
)
with sqlite3.connect(str(_db)) as _con:
    _search_row = _con.execute(
        _cohort_select + "WHERE cm.cohort_type='stagelabel' AND cm.stage=? AND cm.label=? "
        "AND cm.family IS NULL",
        (_lineage["val_stage"], PRIMARY_LABEL),
    ).fetchone()
    _family_row = _con.execute(
        _cohort_select
        + "WHERE cm.cohort_type='family' AND cm.stage=? AND cm.label=? AND cm.family=?",
        (_lineage["val_stage"], PRIMARY_LABEL, RANK1_FAMILY),
    ).fetchone()

# %% [markdown]
# Normalize a cohort row into the named fields used in the reporting
# cells below.


# %%
def _cohort_payload(row: tuple | None) -> dict | None:
    return (
        {
            "k_variants": row[0],
            "dsr_raw": row[1],
            "dsr_raw_pvalue": row[2],
            "dsr_mp": row[3],
            "dsr_mp_pvalue": row[4],
            "dsr_er": row[5],
            "dsr_er_pvalue": row[6],
            "expected_max_sharpe_raw": row[7],
            "expected_max_sharpe_mp": row[8],
            "expected_max_sharpe_er": row[9],
            "min_trl_periods_raw": row[10],
            "min_trl_periods_mp": row[11],
            "min_trl_periods_er": row[12],
            "pbo": row[13],
            "leader_hash": row[14],
            "pbo_n_combinations": row[15],
            "pbo_n_folds": row[16],
            "leader_config_name": row[17],
            "member_digest": row[18],
        }
        if row is not None
        else None
    )


# %% [markdown]
# Enforce exact cohort ownership before displaying search-wide uncertainty.

# %%
SEARCH_COHORT = _cohort_payload(_search_row)
FAMILY_COHORT = _cohort_payload(_family_row)
if SEARCH_COHORT is None:
    raise RuntimeError("Missing cross-family baseline cohort metrics")
SEARCH_ATTRIBUTION = cohort_metric_attribution(SEARCH_COHORT, TOP_HASH)
# The cohort is fetched at the selected configuration's own stage, so the leader it has to name is
# that stage's leader. This read `stage="signal"` while fetching the cohort at
# `_lineage["val_stage"]`, which agreed only while the selected configuration happened to be a
# baseline row. The current configuration is an allocation row, and the two sides of the comparison
# were then two different stages: the check reported a mismatch that was its own. Same population as
# the cohort above, and for the same reason: `best` with a stage alone ranges over every label and
# every universe in the registry, so the leader it returns can be a row the cohort never saw, and
# the check would then fail on the difference between two populations rather than on a real
# disagreement.
_stage_leader = explorer.best(
    stage=_lineage["val_stage"],
    top_n=1,
    label=PRIMARY_LABEL,
    prediction_hashes=_CANDIDATE_PREDICTIONS,
).row(0, named=True)
if SEARCH_COHORT["leader_hash"] != _stage_leader["backtest_hash"]:
    raise RuntimeError(
        f"Cross-family DSR leader does not match the {_lineage['val_stage']} leader: "
        f"{SEARCH_COHORT['leader_hash']} != {_stage_leader['backtest_hash']}"
    )
# K is the size of the search the DSR deflates for, so it has to be the size of the grid the
# selection ranged over. Two earlier versions of this check named it wrongly. It was first
# pinned at the literal 342 - the count from the pre-rebuild registry, when `12_backtest` swept
# the full and liquid universes together - which describes a search that no longer happens. It
# was then measured off the registry with `explorer.best`, which is worse than a stale literal:
# if part of the grid is missing, the recorded K and the measured count shrink together and the
# check passes on an under-deflated DSR, which is the one thing it exists to catch.
#
# The grid is declared, not measured, and it is the grid of the selected configuration's own stage.
# The nominees `13_portfolio_management` publishes span both stages the selection ranges over - the
# baseline grid `12_backtest` declared and the allocation grid built on top of it - and the DSR that
# deflates the selected configuration is the one computed over the stage it came from. Restricting
# the declared set to that stage is what keeps the three checks describing one search: the cohort,
# its leader, and the members it was computed over.
#
# This read `sp500-options-baseline-validation-v1` unconditionally, which is the same set whenever
# the selected configuration is a baseline row and a different search whenever it is not. Both
# populations are immutable and `require_complete` refuses a missing or partial member, so the count
# cannot shrink to meet a K that already has.
with sqlite3.connect(str(_db)) as _con:
    _stage_of = dict(
        _con.execute(
            "SELECT backtest_hash, stage FROM backtest_runs WHERE backtest_hash IN "
            f"({','.join('?' * len(_candidate_hashes))})",
            sorted(_candidate_hashes),
        ).fetchall()
    )
_baseline_members = tuple(
    sorted(h for h in _candidate_hashes if _stage_of.get(h) == _lineage["val_stage"])
)
if not _baseline_members:
    raise RuntimeError(
        f"{STRATEGY_CANDIDATES} declares no member at the selected configuration's stage "
        f"{_lineage['val_stage']!r}, so there is no declared grid to deflate against"
    )
OfficialPopulation.one(_selection_study, name=BASELINE_POPULATION).require_complete()
# The count alone is not enough. `cohort_metrics` records `member_digest`, an
# order-independent digest of the hashes the correction was actually computed over, precisely
# so a reader can establish that a stored DSR belongs to the cohort it is about to be reported
# against. A cohort that swapped a declared member for a retired or unrelated backtest has the
# same K and a different digest, and only the digest catches it.
_declared_digest = cohort_member_digest(_baseline_members)
if SEARCH_COHORT["member_digest"] is None:
    raise RuntimeError(
        "Cross-family DSR row predates member_digest, so the cohort it was computed over "
        f"cannot be established; recompute the cohort metrics for {STRATEGY_CANDIDATES}"
    )
if SEARCH_COHORT["member_digest"] != _declared_digest:
    raise RuntimeError(
        f"Cross-family DSR was computed over a cohort of {SEARCH_COHORT['k_variants']} "
        f"whose members are not the {len(_baseline_members)} that {STRATEGY_CANDIDATES} "
        f"declares at the {_lineage['val_stage']!r} stage "
        f"(digest {SEARCH_COHORT['member_digest']} != {_declared_digest}); the DSR must "
        "deflate for the complete grid the selection ranged over"
    )
PBO_REPORT = (
    reportable_pbo(FAMILY_COHORT["pbo"], FAMILY_COHORT["pbo_n_combinations"])
    if FAMILY_COHORT
    else {"value": None, "status": "unavailable", "n_combinations": None}
)
if FAMILY_COHORT and FAMILY_COHORT["leader_hash"] != SEARCH_COHORT["leader_hash"]:
    raise RuntimeError("Linear-family PBO and cross-family DSR name different leaders")

ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row

print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={PRIMARY_LABEL}")
print(f"        prediction_hash={TOP_PHASH}, val_backtest_hash={TOP_HASH}")
print(f"        holdout_backtest_hash={HO_HASH} (pred {HO_PHASH})")
print(
    f"        {_lineage['val_stage']} DSR leader={SEARCH_COHORT['leader_hash']} "
    f"(applies_to_carrier={SEARCH_ATTRIBUTION['applies_to_carrier']})"
)
print()
print("Daily-pooled IC (validation, prediction-side):")
print(f"  IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')}  (HAC, lag={int(ic_lag)})")
print(f"  t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3f}")
print(f"  n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f"  CI status: {ci_status(ic_lo, ic_hi)}")

# %% [markdown]
# Daily-pooled IC at the rank-1 prediction set is essentially zero
# and its HAC interval straddles zero. The p-value and positive-day share
# likewise provide no rank-correlation evidence. Section 3 reports a strategy
# Sharpe whose CI also straddles zero on the same prediction set, so
# the IC and Sharpe pictures agree. Both can be summarized
# straightforwardly: on `ret_to_expiry`, with full HTM costs, the
# pipeline does not resolve an edge.
#
# **Kill conditions** are encoded in `setup.yaml`:
#
# 1. `vrp_compression` - VRP < 2% annualized for > 6 months. Not
#    evaluated by this notebook (requires a parametric VRP series, not a
#    backtest registry read).
# 2. `gamma_loss_dominance` - rolling 2-year gamma P&L falls below
#    rolling 2-year VRP collection. Not evaluated for the same reason.
# 3. `cost_erosion` - round-trip option spread consumes > 50% of gross
#    VRP edge. The HTM construction already eliminates the exit-leg
#    spread (settle at intrinsic), so this gate is materially weakened
#    relative to the bps-cost variants - but the entry-leg spread alone
#    remains the dominant friction. The specification block reports the
#    current entry and hedge cost totals; the `15_costs` sweep quantifies
#    the slope against the share of the quoted spread paid on entry.
#
# In addition, the strategy-analysis notebook evaluates two universal gates in §9:
# (i) the validation Sharpe CI lower bound ≥ 0, and
# (ii) the holdout strategy-vs-EW paired CI does not exclude zero on
# the negative side. Both are reported as pass / partial / fail without
# verdict labels.

# %% [markdown]
# ## §2 Search context, family comparison, and lineage
#
# The equal-weight baseline sweep on `ret_to_expiry` covers four model families
# under HTM accounting. The rank-1 lineage anchors a multi-stage path
# (signal → allocation → optional risk overlay) at the training-hash
# level. § 5 says why the standard bps cost-grid is the wrong unit for
# option premium returns and points at the `15_costs` population, which
# runs the sweep in fractions of the quoted half-spread instead.

# %%
ctx = explorer.search_context("signal")
search_table = pl.DataFrame(
    [
        {"metric": "Total signal backtests (all labels)", "value": f"{ctx['total']:,}"},
        {"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
        {"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
        {"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
        {"metric": "% positive Sharpe (all labels)", "value": f"{ctx['pct_positive']:.1f}%"},
        {"metric": "Top-by-Sharpe across all labels", "value": f"{ctx['champion_sharpe']:.3f}"},
        {"metric": "Top-by-Sharpe label", "value": ctx["champion_source"]},
    ]
)
print("Equal-weight baseline search context:")
print(search_table)

# %% [markdown]
# The baseline search context above summarizes the `ret_to_expiry`
# cross-family distribution; the registered strategy uses HTM-specific
# cost handling (no bps-of-notional cost framework applies to option
# premium returns). The four legacy diagnostic labels (`fwd_ret_5d`,
# `fwd_ret_10d`, `fwd_ret_dh_5d`, `fwd_ret_dh_10d`) were dropped from
# the sweep 2026-05-17 - they ran through the vectorized backtest path
# which treats 5d/10d forward returns as daily returns, inflating
# Sharpes to non-credible levels.

# %%
# Family-level baseline Sharpe summary, restricted to ret_to_expiry -
# the only label where the rank-1 HTM strategy lives.
with sqlite3.connect(str(_db)) as _con:
    _famdf = pl.DataFrame(
        _con.execute(
            """
            SELECT
                t.family,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi
            FROM backtest_metrics bm
            JOIN backtest_runs b ON bm.backtest_hash = b.backtest_hash
            JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t ON p.training_hash = t.training_hash
            WHERE b.stage = 'signal'
              AND p.split = 'validation'
              AND t.label = 'ret_to_expiry'
              AND bm.sharpe IS NOT NULL
            """
        ).fetchall(),
        schema=["family", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"],
        orient="row",
    )

# %% [markdown]
# Summarize the complete baseline distribution within each model family.

# %%
family_summary = (
    _famdf.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_q25=pl.col("sharpe").quantile(0.25),
        sharpe_q75=pl.col("sharpe").quantile(0.75),
        sharpe_max=pl.col("sharpe").max(),
        pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
    )
    .sort("sharpe_median", descending=True)
)
print("Family-level baseline Sharpe summary (ret_to_expiry only):")
print(family_summary)

# %%
fig, ax = plt.subplots(figsize=(9, 4), layout="tight")
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()

ax.errorbar(
    medians,
    y,
    xerr=[medians - q25, q75 - medians],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    label="median ±IQR",
)
ax.scatter(maxima, y, marker="x", color="#C62828", s=60, label="max", zorder=5)
ax.axvline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
ax.set_title("Family-level Baseline (equal-weight) Sharpe (ret_to_expiry) - IQR + max")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
show_with_alt(
    fig,
    "Horizontal forest plot with one row per model family. Each row draws the family's median "
    "validation backtest Sharpe as a filled circle, a horizontal bar spanning its interquartile "
    "range, and a red cross at its single best configuration. A dashed vertical line marks zero. "
    "The plot is built to show two things at once: how wide each family's spread of "
    "configurations is, and how far its best run sits above its own median.",
)

# %% [markdown]
# Three families produce zero positive-Sharpe baseline rows on `ret_to_expiry`.
# Linear has two near-zero nonnegative rows among the complete baseline surface.
# The full-universe leader is comparison evidence only; the registered configuration
# is the liquid-universe rank-1. Neither resolves a statistically reliable edge.

# %%
# Lineage stages for the registered HTM rank-1
print(f"Lineage stages registered against prediction {TOP_PHASH} (HTM rank-1):")
with sqlite3.connect(str(_db)) as _con:
    rows = _con.execute(
        """
        SELECT b.stage, COUNT(*) AS n
        FROM backtest_runs b
        WHERE b.prediction_hash = ?
        GROUP BY b.stage
        """,
        (TOP_PHASH,),
    ).fetchall()
    for stage, n in rows:
        print(f"  {stage}: {n} runs")

print()
print(
    "All ret_to_expiry backtests route through the HTM daily-MTM cohort "
    "engine (`_run_htm_daily_mtm`). The equal-weight baseline selects the top-K "
    "cross-section (eq_w_topk); the allocation stage overlays a within-"
    "cross-section weighting on the HTM cohort accounting; the risk_overlay "
    "stage layers position-level controls on top of the allocator weights. "
    "The cost sweep lives in `15_costs`, in fractions of the quoted "
    "half-spread (the standard bps cost-grid is the wrong unit for option "
    "premium returns)."
)

# %% [markdown]
# ## §3 Headline performance with uncertainty
#
# The pinned configuration's performance metrics use 95% block-bootstrap CIs.
# Selection-bias adjustment (DSR raw / MP / ER, k_variants,
# expected_max_sharpe, min_trl_periods) lives in `cohort_metrics` per
# `memory/UNCERTAINTY_ARCHITECTURE.md` - `backtest_metrics` no longer
# carries those columns. The DSR uses the complete cross-family baseline
# cohort and names its full-universe leader. That leader differs from the pinned
# liquid configuration, so these statistics are not presented as configuration metrics.
# Linear-family PBO is displayed separately and suppressed because only two
# CSCV combinations are available. ER is the maintainer-recommended default;
# raw and MP are recorded for context.

# %%
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)

spec_block = {
    "case_study": CASE_STUDY,
    "family": RANK1_FAMILY,
    "config_name": RANK1_CONFIG,
    "label": PRIMARY_LABEL,
    "stage": _lineage["val_stage"],
    "signal_method": SIGNAL_METHOD,
    "allocator": ALLOCATOR_METHOD,
    "top_k": TOP_K,
    "universe_filter": UNIVERSE_FILTER,
    "cascade_rung": CASCADE_RUNG,
    "rebalance_cadence": setup["decision"]["entry_cadence"],
    "rebalance_step_weeks": setup["labels"]["rebalance_step"][PRIMARY_LABEL],
    "hedge_cadence": setup["decision"]["hedge_cadence"],
    "exit_rule": setup["decision"]["exit_time"],
    "cost_model": "HTM daily-MTM cohort: option entry spread + daily underlying hedge spread + commissions",
    "cumulative_entry_cost_premium_units": full["cumulative_entry_cost"],
    "cumulative_hedge_cost_premium_units": full["cumulative_hedge_cost"],
    "n_rebalance_dates": int(full["n_rebalance_dates"])
    if full["n_rebalance_dates"] is not None
    else None,
    "avg_cohorts_open": full["avg_cohorts_open"],
    "validation_window_periods": int(full["n_periods"]),
    "bootstrap_block_length": int(full["bootstrap_block_length"]),
    "bootstrap_n": int(full["bootstrap_n"]),
}
# The stage comes from the selected configuration's own lineage. Naming it "equal-weight baseline"
# was correct only while the cross-stage rank-1 happened to be a signal-stage row; the selected
# configuration in force is an allocation row, and a reader told otherwise has the wrong provenance
# for every number in the block.
print(
    f"Pinned-configuration specification ({_lineage['val_stage']} stage, "
    f"{ALLOCATOR_METHOD or 'equal-weight'} allocation, validation window):"
)
for k, v in spec_block.items():
    print(f"  {k}: {v}")


# %% [markdown]
# Format one uncertainty row consistently across the headline table.


# %%
def _row(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
    return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}


# %% [markdown]
# Collect the selected configuration's bootstrap intervals before adding the
# selection-adjusted diagnostics.

# %%
sharpe_status = ci_status(full["sharpe_ci95_lo"], full["sharpe_ci95_hi"])
sortino_status = ci_status(full["sortino_ci95_lo"], full["sortino_ci95_hi"])
ann_status = ci_status(full["ann_return_ci95_lo"], full["ann_return_ci95_hi"])
mdd_status = ci_status(full["max_dd_ci95_lo"], full["max_dd_ci95_hi"])
calmar_status = ci_status(full["calmar_ci95_lo"], full["calmar_ci95_hi"])
metric_specs = [
    ("Sharpe", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi", sharpe_status),
    ("Sortino", "sortino", "sortino_ci95_lo", "sortino_ci95_hi", sortino_status),
    ("Annualized return", "cagr", "ann_return_ci95_lo", "ann_return_ci95_hi", ann_status),
    ("Max drawdown", "max_drawdown", "max_dd_ci95_lo", "max_dd_ci95_hi", mdd_status),
    ("Calmar", "calmar", "calmar_ci95_lo", "calmar_ci95_hi", calmar_status),
]
performance_rows = [
    _row(label, _fmt(full[point]), _fmt(full[lo]), _fmt(full[hi]), status)
    for label, point, lo, hi, status in metric_specs
]
performance_rows.append(_row("PSR p-value (H0: SR≤0)", _fmt(full["psr_pvalue"]), "-", "-", "n/a"))

# %% [markdown]
# Attribute each DSR variant to the exact cross-family cohort and keep
# the insufficient family PBO visibly separate.

# %%
dsr_rows = []
for suffix, label in [("raw", "DSR_raw"), ("er", "DSR_ER"), ("mp", "DSR_MP")]:
    pvalue = SEARCH_COHORT[f"dsr_{suffix}_pvalue"]
    status = f"{suffix}_p={pvalue:.3f}" if pvalue is not None else "cohort_unavailable"
    dsr_rows.append(
        _row(
            f"{label} ({_lineage['val_stage']} leader {SEARCH_COHORT['leader_hash']})",
            _fmt(SEARCH_COHORT[f"dsr_{suffix}"]),
            "-",
            "-",
            status,
        )
    )

diagnostic_rows = [
    _row(
        "PBO (linear-family cohort)",
        _fmt(PBO_REPORT["value"]) if PBO_REPORT["value"] is not None else "insufficient",
        "-",
        "-",
        PBO_REPORT["status"],
    ),
    _row(
        "k_variants (cross-family search cohort)",
        str(int(SEARCH_COHORT["k_variants"])),
        "-",
        "-",
        "n/a",
    ),
]

# %% [markdown]
# Display the performance and selection diagnostics in one compact table.

# %%
headline = pl.DataFrame(performance_rows + dsr_rows + diagnostic_rows)
print("Selected-configuration performance and exactly attributed search diagnostics:")
print(headline)

# %% [markdown]
# The cross-family DSR row has K equal to the registered baseline count and names
# the complete-grid baseline
# leader, not the pinned liquid configuration. PBO is a separate linear-family
# diagnostic. With only two CSCV combinations, the notebook reports
# "insufficient combinations" instead of interpreting 0.50.

# %%
# Prepare the rank-1 metric stack and reference benchmark for a forest plot.
ew_val = load_benchmark_metrics(CASE_STUDY, PRIMARY_LABEL, period="validation")
ew_ho = load_benchmark_metrics(CASE_STUDY, PRIMARY_LABEL, period="holdout")
forest_metrics = [
    ("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
    ("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
    ("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
    ("Ann. return", full["cagr"], full["ann_return_ci95_lo"], full["ann_return_ci95_hi"]),
]

# %% [markdown]
# Plot the selected configuration intervals against zero and the validation equal-weight
# benchmark.

# %%
fig, ax = plt.subplots(figsize=(8, 4), layout="tight")
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.errorbar(
    points,
    y,
    xerr=[points - los, his - points],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    markersize=7,
)
ax.axvline(0, color="#9E9E9E", linestyle="--", linewidth=0.8)
ax.axvline(
    ew_val["sharpe"],
    color="#43A047",
    linestyle=":",
    linewidth=1.0,
    label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Value")
ax.set_title("Rank-1 (ret_to_expiry, HTM dispatch) Headline Metrics with 95% CIs")
ax.legend(loc="lower right", fontsize=8, frameon=False)
show_with_alt(
    fig,
    "Horizontal forest plot of four headline metrics for the rank-1 configuration, one row each: "
    "Sharpe, Sortino, Calmar and annualized return. Each row draws the point estimate as a circle "
    "with a bar spanning its 95 percent confidence interval. A dashed vertical line marks zero "
    "and a dotted green line marks the equal-weight benchmark's validation Sharpe. All four rows "
    "share one numeric axis labelled Value, so the benchmark line is comparable only to the "
    "Sharpe row.",
)

# %%
# Equity-curve overlay vs validation EW benchmark
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
    pl.read_parquet(strat_returns_path)
    .sort("timestamp")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)

bench_val = (
    load_benchmark_returns(CASE_STUDY, PRIMARY_LABEL, period="validation")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("ew_return").alias("benchmark"))
)

aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
    f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)

cum_strat = np.cumprod(1 + aligned["strategy"].to_numpy()) - 1
cum_bench = np.cumprod(1 + aligned["benchmark"].to_numpy()) - 1
fig, ax = plt.subplots(figsize=(10, 4.2), layout="tight")
ax.plot(aligned["ts"], cum_strat, color="#1565C0", linewidth=1.2, label="Rank-1 strategy (HTM)")
ax.plot(aligned["ts"], cum_bench, color="#43A047", linewidth=1.2, label="EW universe")
ax.axhline(0, color="#9E9E9E", linewidth=0.6, linestyle="--")
ax.set_ylabel("Cumulative return")
ax.set_title("Validation-window cumulative return: rank-1 HTM vs EW universe")
ax.legend(loc="best", frameon=False)
show_with_alt(
    fig,
    "Line chart of cumulative return across the validation overlay window, with two series: the "
    "rank-1 hold-to-maturity strategy and the equal-weight universe. Both are compounded from "
    "per-period returns and start at zero, and a dashed horizontal line marks zero, so the "
    "vertical gap between the lines at any date is the strategy's cumulative excess over the "
    "benchmark.",
)

# %% [markdown]
# The validation Sharpe straddles zero with a CI that spans roughly ±1
# Sharpe in either direction - the validation window does not resolve a
# directional edge for this strategy. Sortino, Calmar, and Ann. return
# all sit in the same straddles_zero status; PSR confirms the read.
# Whichever side of zero the point estimate falls on, the validation
# evidence is null-consistent. § 6's paired holdout vs EW test resolves
# whether the validation/EW gap holds out of sample. Selection accounting
# is supplied by the exact cross-family baseline cohort above (DSR raw /
# MP / ER and K from `cohort_metrics`), whose named leader differs from
# the selected configuration. Linear-family PBO is separately suppressed as underidentified.
# The two near-zero nonnegative baseline rows do not alter the unresolved
# selection-adjusted DSR_ER reading.

# %% [markdown]
# ## §3a Label scope and prior diagnostic exhibit
#
# The equal-weight baseline sweep on this case study now runs on a single label
# - `ret_to_expiry` - under HTM cost accounting (entry-side option
# spread + daily underlying hedge spread, plus an exit-leg option trade only
# where a contract's chain ends before its expiration date).
# Four legacy diagnostic labels (`fwd_ret_5d`, `fwd_ret_10d`,
# `fwd_ret_dh_5d`, `fwd_ret_dh_10d`) were dropped from the sweep on
# 2026-05-17: they routed through the vectorized backtest path that
# treats 5d/10d forward returns as daily returns, inflating Sharpes
# to non-credible levels.
# The structural diagnostic remains valid as prose: equity-style
# bps-of-notional cost accounting understates option spread cost by
# 1-2 orders of magnitude versus the actual % of premium, so any
# equity-style backtester on option premium returns produces
# pathological CAGR / drawdown numerics. The registered strategy
# therefore uses HTM cost accounting on `ret_to_expiry`, and under
# that accounting the cross-section does not resolve a statistically
# significant edge on the validation window.

# %% [markdown]
# ## §4 Risk and drawdown analysis
#
# The validation-window strategy returns paired against the validation
# EW benchmark. Tail-risk metrics come from the registry directly; the
# fold breakdown surfaces the across-fold dispersion that drives the
# bootstrap CI's width.

# %%
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()

pa = PortfolioAnalysis(
    returns=strat_arr,
    benchmark=bench_arr,
    dates=ts_arr,
    periods_per_year=PERIODS_PER_YEAR,
)

dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)

# %%
fig = plot_equity_drawdown(strat_returns_path)
show_with_alt(
    fig,
    "Two stacked panels sharing a date axis. The upper, taller panel plots the strategy's "
    "cumulative return over the backtest window. The lower panel plots drawdown, the fractional "
    "decline from the running peak of that same curve, as a red line over a shaded area, with the "
    "deepest point annotated. The pairing lets the length and depth of each decline be read "
    "against the part of the equity curve that produced it.",
)

# %%
roll = pa.compute_rolling_metrics(windows=[126], metrics=["sharpe", "beta"])
if isinstance(roll, dict):
    print("Rolling-window keys:")
    print({k: type(v).__name__ for k, v in roll.items()})

tail_table = pl.DataFrame(
    [
        {"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
        {"metric": "VaR 95% (daily)", "value": f"{full['var_95']:.4f}"},
        {"metric": "CVaR 95% (daily)", "value": f"{full['cvar_95']:.4f}"},
        {"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
        {"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
        {"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
    ]
)
print()
print("Tail risk profile:")
print(tail_table)

# %%
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
print(f"Per-fold breakdown ({fold_df.height} folds):")
print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
print()
if fold_df.height > 1:
    print(f"Fold Sharpe range: [{fold_df['sharpe'].min():.3f}, {fold_df['sharpe'].max():.3f}]")
    print(f"Fold Sharpe std:   {fold_df['sharpe'].std():.3f}")

# %% [markdown]
# The two fold Sharpes straddle zero and both sit inside the wide
# bootstrap Sharpe CI in Section 3. Tail risk is heavy: negative skew
# and high kurtosis reflect the asymmetric P&L profile
# of a short straddle book - long stretches of small premium decay
# punctuated by large losses when realized vol exceeds implied.
# CVaR and max drawdown are severe, the defining signature
# of the strategy: even with hedging and HTM settlement, a single
# bad-vol regime can wipe out cumulative premium.

# %% [markdown]
# ## §5 Friction budget
#
# The standard `BacktestExplorer.cost_sensitivity()` curve is denominated in bps per leg, which is
# the right convention for equities and futures but the wrong unit for options: a 10% spread on a
# 4% premium is 40 bps of notional, well inside the "positive Sharpe" zone of any equity sweep and
# ruinous for the actual P&L. `15_costs` runs the cost sweep in fractions of the quoted half-spread
# instead, executes it through the same engine as every other backtest, and publishes it as the
# named population `sp500-options-cost-sensitivity-validation-v1`. It reports the resulting Sharpe
# surface across families, universes and spread fractions there, each point with its
# block-bootstrap Sharpe interval, so this notebook does not repeat it.

# %% [markdown]
# ## §6 Holdout closure with paired bootstrap
#
# Two paired tests anchor the holdout read. (i) The holdout rank-1
# versus the validation rank-1 ("did Sharpe hold?"); (ii) the holdout
# rank-1 versus the holdout-window equal-weight benchmark ("did the
# strategy beat random in the holdout?"). Numbers come from
# `backtest_paired_metrics` rows populated for this fixed case lineage;
# never from val_sharpe minus holdout_sharpe arithmetic.

# %%
ho_full = load_backtest_metrics(CASE_STUDY, backtest_hash=HO_HASH).row(0, named=True)
val_full = full

print(f"Validation rank-1 hash: {TOP_HASH} (Sharpe {val_full['sharpe']:+.4f})")
print(f"Holdout rank-1 hash:    {HO_HASH} (Sharpe {ho_full['sharpe']:+.4f})")
print()
print(
    f"Holdout window: {setup['evaluation']['holdout_start']} → {setup['evaluation']['holdout_end']}"
)
print(f"Holdout n_periods: {int(ho_full['n_periods'])}")

# %%
val_ho_pair = load_paired_metrics(
    CASE_STUDY,
    challenger_hash=HO_HASH,
    benchmark_kind="val_rank1_self",
)
if val_ho_pair.is_empty():
    raise RuntimeError(
        "Missing required val_rank1_self paired metrics for current holdout "
        f"{HO_HASH}; populate this case before executing notebook 16"
    )
vh = val_ho_pair.row(0, named=True)

# %% [markdown]
# Format one paired validation-to-holdout metric, preserving its
# bootstrap interval and p-value when available.


# %%
def _diff_row(
    label: str,
    v: float,
    h: float,
    diff: float | None,
    lo: float | None,
    hi: float | None,
    p: float | None,
) -> dict:
    return {
        "metric": label,
        "validation": _fmt(v, ".4f"),
        "holdout": _fmt(h, ".4f"),
        "diff (h−v)": _fmt(diff, ".4f") if diff is not None else "-",
        "diff CI95": f"[{_fmt(lo, '.4f')}, {_fmt(hi, '.4f')}]" if lo is not None else "-",
        "p-value": _fmt(p, ".4f") if p is not None else "-",
    }


# %% [markdown]
# Build the disjoint-window comparison from the stored paired-bootstrap
# row rather than subtracting headline metrics.

# %%
pair_specs = [
    ("Sharpe", "sharpe", "sharpe_diff", "sharpe_diff_ci95", "p_value"),
    ("Annualized return", "cagr", "ret_diff", "ret_diff_ci95", None),
    ("Max drawdown", "max_drawdown", "max_dd_diff", "max_dd_diff_ci95", None),
]
pair_rows = [
    _diff_row(
        label,
        val_full[metric],
        ho_full[metric],
        vh[diff],
        vh[f"{ci_prefix}_lo"],
        vh[f"{ci_prefix}_hi"],
        vh[pvalue] if pvalue else None,
    )
    for label, metric, diff, ci_prefix, pvalue in pair_specs
]
pair_rows.append(
    _diff_row(
        "Information ratio",
        None,
        None,
        vh["info_ratio"],
        vh["info_ratio_ci95_lo"],
        vh["info_ratio_ci95_hi"],
        None,
    )
)
val_ho_table = pl.DataFrame(pair_rows)

# %% [markdown]
# Display the paired decay statistics and state the disjoint-window
# bootstrap construction.

# %%
print("val → holdout paired-bootstrap decay (rank-1 self):")
print(val_ho_table)
print(
    "prob_challenger_wins: "
    + (f"{vh['prob_challenger_wins']:.3f}" if vh["prob_challenger_wins"] is not None else "-")
)
print(f"CI status (Sharpe diff): {ci_status(vh['sharpe_diff_ci95_lo'], vh['sharpe_diff_ci95_hi'])}")
print()
print(
    "Note: validation and holdout windows are disjoint. The paired-metrics "
    "producer therefore uses independent stationary block-bootstrap draws "
    "for the two windows rather than calendar alignment."
)

# %%
ho_vs_ew = load_paired_metrics(
    CASE_STUDY,
    challenger_hash=HO_HASH,
    benchmark_kind="equal_weight_holdout_side_artifact",
)
if ho_vs_ew.is_empty():
    raise RuntimeError(
        "Missing required equal_weight_holdout_side_artifact paired metrics for "
        f"current holdout {HO_HASH}; populate this case before executing notebook 16"
    )
he = ho_vs_ew.row(0, named=True)

print("Holdout strategy vs holdout-window EW universe:")
print(f"  strategy Sharpe:  {ho_full['sharpe']:+.4f}")
print(f"  EW Sharpe:        {ew_ho['sharpe']:+.4f}")
print(
    "  diff Sharpe: "
    f"{_fmt_ci(he['sharpe_diff'], he['sharpe_diff_ci95_lo'], he['sharpe_diff_ci95_hi'], '.4f')}"
)
print("  p_value:                " + (f"{he['p_value']:.4f}" if he["p_value"] is not None else "-"))
print(
    "  prob_challenger_wins:   "
    + (f"{he['prob_challenger_wins']:.3f}" if he["prob_challenger_wins"] is not None else "-")
)
print(
    f"  info_ratio (strategy vs EW): "
    f"{_fmt_ci(he['info_ratio'], he['info_ratio_ci95_lo'], he['info_ratio_ci95_hi'])}"
)
print(f"  CI status: {ci_status(he['sharpe_diff_ci95_lo'], he['sharpe_diff_ci95_hi'])}")

# %% [markdown]
# **Reading.** The two paired tables above are the authoritative holdout
# evidence for the fixed current configuration. Point-estimate ordering alone
# does not establish persistence or benchmark superiority; the paired
# confidence intervals and p-values determine whether either difference
# is resolved. The 2021 window is reported once and is not reused for
# model or strategy selection.

# %% [markdown]
# ## §7 Benchmark-aware diagnostics
#
# Layer 1 reports the universal alpha/beta/IR profile via
# `PortfolioAnalysis` against the equal-weight S&P 500 options universe.
# Layer 2 - equity factor attribution (FF5+MOM) - is in scope per the
# case-study scope deviation table for equity-class case studies, but the model
# fit is structurally compromised on options: FF5+MOM captures linear
# equity factor exposure, while option premium returns are driven by
# vega, gamma, and theta - non-linear payoffs that are orthogonal to
# linear factor loadings. We include the regression for completeness
# but flag the model mismatch in the reading.

# %%
metrics = pa.compute_summary_stats()
attr_df = pl.DataFrame(
    [
        {"metric": "alpha (annualized)", "value": _fmt(getattr(metrics, "alpha", None))},
        {"metric": "beta", "value": _fmt(getattr(metrics, "beta", None), ".3f")},
        {"metric": "information ratio", "value": _fmt(getattr(metrics, "information_ratio", None))},
        {"metric": "tracking error", "value": _fmt(getattr(metrics, "tracking_error", None))},
        {"metric": "up capture", "value": _fmt(getattr(metrics, "up_capture", None))},
        {"metric": "down capture", "value": _fmt(getattr(metrics, "down_capture", None))},
    ]
)
print("Layer 1: rank-1 vs validation EW universe (PortfolioAnalysis):")
print(attr_df)

# %%
# Placebo regression: residual α and HAC t-stat on EW alone.
import statsmodels.api as sm
from case_studies.utils.warning_policy import apply_notebook_warning_policy

apply_notebook_warning_policy()

X = sm.add_constant(bench_arr)
ols = sm.OLS(strat_arr, X).fit(cov_type="HAC", cov_kwds={"maxlags": 5})
alpha_daily = ols.params[0]
alpha_t = ols.tvalues[0]
alpha_p = ols.pvalues[0]
beta_ew = ols.params[1]
beta_t = ols.tvalues[1]
print()
print("Placebo regression vs EW universe (HAC, maxlags=5):")
print(f"  α (daily) = {alpha_daily:.6f}, α (annualized) = {alpha_daily * PERIODS_PER_YEAR:.4f}")
print(f"  α t-stat = {alpha_t:.3f}, p = {alpha_p:.3f}")
print(f"  β        = {beta_ew:.3f}, β t-stat = {beta_t:.3f}")
print(f"  CI status (α): {'excludes_zero_strong' if alpha_p < 0.05 else 'straddles_zero'}")

# %%
# Layer 2: FF5+MOM factor regression with HAC standard errors.
strategy_rets_pd = pd.Series(
    strat_arr,
    index=pd.to_datetime(aligned["ts"].to_list()),
    name="strategy",
)
_period_start = str(strategy_rets_pd.index.min().date())
_period_end = str(strategy_rets_pd.index.max().date())

_factors = load_factor_data(start=_period_start, end=_period_end)
_reg = run_factor_regression(
    strategy_rets_pd,
    _factors,
    model="ff5_mom",
    hac_lags=5,
    dollar_neutral=False,
)

print()
print("Layer 2: FF5+MOM factor attribution (HAC, daily):")
print(f"  n_obs:           {_reg['n_obs']}")
print(f"  R²:              {_reg['r_squared']:.3f}")
print(
    f"  α (annualized):  {_reg['alpha_annualized']:+.4f}  "
    f"(t={_reg['alpha_t_stat']:.2f}, p={_reg['alpha_p_value']:.3f})"
)
print(f"  Strategy Sharpe: {_reg['strategy_sharpe']:+.3f}")
print(f"  Residual Sharpe: {_reg['residual_sharpe']:+.3f}")
print()
print("Factor betas (HAC):")
for factor, beta in _reg["betas"].items():
    t = _reg["t_stats"][factor]
    sig = "*" if _reg["p_values"][factor] < 0.05 else ""
    print(f"  {factor:8s}: {beta:+.4f}  (t={t:+.2f}){sig}")

# %%
# Block bootstrap CIs on alpha and factor betas
_boot = compute_bootstrap_ci(
    strategy_rets_pd,
    _factors,
    model="ff5_mom",
    n_boot=500,
    block_size=20,
    dollar_neutral=False,
    seed=42,
)

if _boot.get("n_boot", 0) > 0:
    print(f"\nBootstrap CIs (n={_boot['n_boot']}, block=20 days):")
    print(
        f"  α (annualized): [{_boot['alpha_ann_lo']:+.4f}, {_boot['alpha_ann_hi']:+.4f}] (95% CI)"
    )
    for factor in _reg["factor_columns"]:
        key_lo = f"{factor}_lo"
        key_hi = f"{factor}_hi"
        if key_lo in _boot:
            print(f"  {factor:8s}: [{_boot[key_lo]:+.4f}, {_boot[key_hi]:+.4f}]")

# %%
fig_attr = plot_attribution_waterfall(
    _reg, title="S&P 500 Options HTM: FF5+MOM Attribution (diagnostic)"
)
show_with_alt(
    fig_attr,
    "Bar chart with one bar per FF5 factor plus momentum and a final gray Residual bar, against a "
    "y axis labelled Sharpe Contribution, each bar annotated with its value and a dashed "
    "horizontal line marking the strategy's total Sharpe. The factor bars are each coefficient's "
    "absolute value as a share of the total, rescaled to the factor-explained part of Sharpe, so "
    "they show relative exposure and not an exact return decomposition.",
)

# %% [markdown]
# Layer-1 placebo regression vs EW: α (annualized) and HAC t-stat read
# directly off the print block above. Layer-2 FF5+MOM produces a low R²
# by design - the linear factor model captures whatever residual equity
# delta leaks through the daily delta hedge, but the dominant return
# drivers (gamma, vega, theta) sit outside the FF5+MOM linear span.
# The α point estimate is therefore not interpretable as
# "factor-adjusted edge" for an options strategy in the same sense as
# for an equity portfolio; it's a residual after a structurally
# incomplete factor model. The diagnostic value is in the *betas*: a
# significant Mkt-RF loading would indicate incomplete delta hedging
# (the daily rehedge with a 0.10 net-delta band leaks some directional
# exposure); a significant MOM loading would suggest the model selects
# straddles on names whose option markets lag equity momentum (a
# microstructure artefact).
#
# A volatility-native attribution (VIX, VVIX, term-structure slope,
# variance risk premium factors) is the right model for this strategy.
# That construction is a Ch20 extension (the dispersion-trading
# literature provides it); per strategy-analysis convention this notebook reports
# what's available without synthesizing a new factor m

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.