Перейти к содержимому
Все документы библиотеки

Оценка стратегий NASDAQ-100 парными сравнениями и бутстрэп-интервалами

Код Machine Learning for Trading

Сводка

В документе описан процесс оценки зарегистрированных бэктестов микроструктуры NASDAQ-100. Он считывает существующие запуски, не обучая модели и не выполняя бэктесты повторно, отслеживает цепочки выбранных стратегий, сравнивает кандидатов с эталоном с равными весами и оценивает результаты на валидации и отложенной выборке. Основной метод сравнения сопоставляет доходности стратегий за одни и те же дни, напрямую учитывая общие движения рынка вместо вычитания независимых сводных показателей. Для оценки неопределённости используются интервалы блочного бутстрэпа, сохраняющие зависимость между последовательными дневными доходностями.

В ноутбуке также заданы критерии прохождения валидации и проверки на отложенной выборке, рассматриваются затраты и дополнительные меры контроля риска, а структурированная оценка сохраняется для последующего объединения результатов нескольких исследований. При отсутствии производных таблиц реестра их можно создать по основному отбору, заданному тематическим исследованием; уже заполненные реестры считываются без изменений. Эти процедуры делают анализ воспроизводимым и позволяют точнее описывать неопределённость, однако прохождение критерия не доказывает, что стратегия сохранит эффективность или будет применима в других условиях. Документ описывает аналитическую основу; предоставленный фрагмент не содержит фактических результатов стратегии или значений интервалов.

Ключевые идеи

  • Сравнивай стратегии за одни и те же даты, чтобы учитывать общую рыночную доходность через парные разницы.
  • При оценке интервалов неопределённости используй блочную выборку последовательных дней, чтобы учитывать зависимость дневных доходностей.
  • Считывай зарегистрированные бэктесты и сохраняй их цепочку происхождения, не запуская их повторно и не переписывая результаты вручную.
  • Используй критерии валидации и отложенной выборки как ограниченные проверки данных, а не доказательство будущей эффективности.
  • Заполняй производные таблицы реестра, только если они отсутствуют или устарели относительно заданного основного набора.

Теги

Полный текст
# 20_strategy_analysis.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # NASDAQ-100 Microstructure — Strategy Analysis
#
# This notebook turns the backtest registry for the NASDAQ-100 microstructure
# case study into one strategy assessment. It reports every metric with an
# interval rather than as a point, compares strategies against each other in
# pairs rather than by subtracting summary numbers, and reads the holdout once.
# Comparison against other case studies is reserved for the synthesis chapter.
#
# Two ideas run through it. **Paired comparison**: two strategies evaluated on
# the same days are compared day by day, so the shared market movement cancels
# and what is left is the difference between them. Subtracting one summary
# statistic from another throws that pairing away and produces an interval far
# wider than the evidence warrants. **Block bootstrap**: intervals are built by
# resampling contiguous blocks of days rather than individual days, because
# returns on consecutive days are related and resampling them independently would
# understate the uncertainty.
#
# **Learning objectives**
#
# - Read backtest metrics with their intervals from the registry rather than
#   transcribing point estimates
# - Trace one selected lineage through the pipeline stages and attribute the
#   change at each stage to the field that stage varied
# - Compare a strategy against an equal-weight benchmark on the same days, on
#   both the validation and the holdout window
# - State a gate so that it can fail, and say what passing it does and does not
#   establish
#
# **Book reference**: Chapter 20, §20.1 (the §9 handoff feeds Ch20's
# cross-case-study aggregation).
#
# **Prerequisites**: case-study pipeline through `14_backtest`; the
# locked registry (`case_studies/nasdaq100_microstructure/run_log/registry.db`).
#
# **Scope**: no training and no re-backtesting. The notebook reads what
# `14_backtest` through `17_costs` registered, and derives
# `cohort_metrics` and `backtest_paired_metrics` from those runs itself when
# they are absent, so the case study no longer depends on a chapter-20
# notebook having been run first. On an already-populated registry it writes
# nothing.

# %%
"""NASDAQ-100 Microstructure — Strategy Analysis."""

import json
import sqlite3

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl

# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch  # noqa: F401
import yaml
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
    BacktestReportMetadata,
    generate_tearsheet_from_run_artifacts,
)

from case_studies.research import read_only_study
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.factor_attribution import (
    compute_bootstrap_ci,
    compute_rolling_exposures,
    format_attribution_summary,
    load_factor_data,
    plot_attribution_waterfall,
    plot_rolling_exposures,
    run_factor_regression,
)
from case_studies.utils.notebook_contracts import (
    degenerate_prediction_sql,
    derived_tables_off_canonical_universe,
    strategy_input_counts,
)
from case_studies.utils.paired_metrics import populate_paired_metrics, rung_for
from case_studies.utils.registry import (
    load_backtest_fold_metrics,
    load_backtest_metrics,
    load_paired_metrics,
)
from case_studies.utils.strategy_analysis import (
    ci_status,
    compute_operating_profile,
    fmt_gate,
    gate1_validation_sharpe_geq_zero,
    gate2_holdout_diff_not_excludes_zero_negatively,
    gate_passes,
    plot_concentration_curve,
    plot_equity_drawdown,
    plot_sharpe_waterfall,
    resolve_solvent_carrier,
    write_strategy_assessment,
)
from case_studies.utils.sweep_config import get_universe_filters_for
from case_studies.utils.uncertainty import ENTIRE_REGISTRY, NO_CARRIER
from utils.paths import get_output_dir
from utils.style import show_with_alt

# Figures go through `show_with_alt`, and none of them calls `tight_layout()`. Both are
# measured rather than stylistic: `utils/style` and `matplotlibrc` set
# `figure.constrained_layout.use`, so `tight_layout()` warns and fights the layout engine
# already running, and `fig.show()` on a non-interactive backend warns that the canvas
# cannot be shown - two UserWarnings per figure, written into the rendered cell. The six
# notebooks of this case study already at `done` use this form and neither of the others.
#
# The alt strings name structure, axes and reference lines, never which series wins. An
# ordering is a registry result a rebuild can reverse, and alt text is prose no rebuild
# revisits, so a ranking written here would go stale silently on the one surface a reader
# who cannot see the chart depends on.

# %% tags=["parameters"]
CASE_STUDY = "nasdaq100_microstructure"
MAX_SYMBOLS = 0
# Where this notebook reads. Empty means the canonical registry; a smoke run passes the
# workspace the stages before it wrote to.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""

# %% [markdown]
# This notebook reads; it registers nothing, and that decides how it opens the registry. Every
# route through `open_study` ends in `Study.activate()`, which rewrites `ML4T_OUTPUT_DIR` for
# the rest of the process and clears the caches keyed on it, so every later
# `get_case_study_dir` answers for a different directory than the one resolved here. On the
# canonical tier with no workspace that route is `Study.regenerate`, which refuses outright
# unless `features`, `labels` and `run_log` are symlinks - true in a maintainer worktree, false
# in every clean clone.
#
# `read_only_study` is the form that does not activate: it resolves the root for the tier and
# workspace it is given and hands back a `Study.at` over it, which points the path helpers at
# that root rather than clearing the variable. `CASE_DIR` is that root, and every question this
# notebook asks - the catalog, the lineage, the populations, the artifacts - is answered from
# it.

# %%
study = read_only_study(
    CASE_STUDY,
    workspace=WORKSPACE or None,
    execution_tier=EXECUTION_TIER,
    entry_point="20_strategy_analysis",
)
CASE_DIR = study.root
PERIODS_PER_YEAR = 252  # strategy returns are aggregated to daily before Sharpe
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

with open(CASE_DIR / "config" / "setup.yaml") as f:
    setup = yaml.safe_load(f)

explorer = BacktestExplorer(CASE_STUDY)
print(explorer)

# %% [markdown]
# ### Produce the derived tables this notebook reads
#
# `cohort_metrics` and `backtest_paired_metrics` are read below - the first for the deflated
# Sharpe and PBO columns, the second to find the validation counterpart of a holdout run. Until
# now nothing in this case study wrote either. The header of this notebook said so outright:
# they "were populated by `20_strategy_synthesis/01_aggregate_synthesis.py`", a chapter-20
# notebook. A case study that depends upward into the chapter that aggregates it cannot be run
# from its own pipeline, and its numbers exist only after something outside it happens to run.
#
# Both producers already exist as case-study-scoped functions, extracted from those very Ch20
# loops - `populate_paired_metrics` says so in its own docstring - and `cme_futures/17` is the
# notebook that uses them. This adopts the same pattern.
#
# Population is idempotent (keyed upserts, reproducing stored values exactly), but it still
# rewrites the registry, so it runs only when a table is absent or empty. On an already-populated
# registry this notebook stays a pure reader and touches nothing.

# %%
_db = CASE_DIR / "run_log" / "registry.db"
_counts_before = strategy_input_counts(CASE_DIR)
_n_backtests = _counts_before["backtest_runs"]
_n_cohorts = _counts_before["cohort_metrics"]
_n_pairs = _counts_before["backtest_paired_metrics"]

if _n_backtests == 0:
    # Refuse rather than report. Every figure and gate below is computed from backtest runs, so
    # with none registered this notebook does not produce a weaker answer - it produces an empty
    # one that reads like a finished analysis. `14_backtest` is what fills this.
    raise RuntimeError(
        f"{CASE_STUDY} has no registered backtest runs, so there is nothing to analyse. "
        "Run 14_backtest through 17_costs first; this notebook reads what they "
        "register and computes no backtests of its own."
    )

# Both producers need this case study's canonical selection, and passing neither is not a
# smaller version of passing both - it is a different selection. `setup.yaml` says the
# full-universe variant "is NOT a canonical rank-1 / cohort / DSR candidate" and lives only in
# the 17_costs comparison, so a cohort computed without the filter mixes rows the case study
# excludes by declaration into the leaders, trial counts, DSR and PBO. The filter is read from
# `setup.yaml` rather than written here, so it cannot disagree with the sweep that produced the
# runs. The rung pin is the matching statement for paired metrics: nasdaq's carrier is the
# cost-feasible ensemble, fixed before the holdout was opened.
_UNIVERSE_FILTER = get_universe_filters_for(CASE_STUDY)[0]
_RUNG = rung_for(CASE_STUDY)

# The two tables are filled independently. A single `if either is empty` guard recomputed both,
# and `compute_and_register` writes with `replace_all=True` - so a registry with correct cohorts
# and an empty paired table would have had its cohorts replaced as a side effect of populating
# the pairs.
# A row count answers "has this been populated", which is not what a rerun needs to know. Both
# tables are written from a selection, and one populated by an earlier run that selected
# differently - before this notebook passed the universe filter, say - is fully populated and
# wrong. Neither table records the selection that produced it, so it is recovered from what the
# rows point at: a canonical row cannot reference a backtest run outside the canonical universe.
_stale = derived_tables_off_canonical_universe(CASE_DIR, _UNIVERSE_FILTER)
for _table in sorted(_stale):
    print(f"{_table} references runs outside {_UNIVERSE_FILTER}, so it is rebuilt")

if _n_cohorts == 0 or "cohort_metrics" in _stale:
    # `compute_and_register` writes with `replace_all=True`, so a rebuild that computes no
    # cohort empties the table rather than leaving the previous one - the opposite of what
    # `populate_paired_metrics` does on the same failure. An emptied table is also clean by
    # the universe check below, which reads rows and finds none, so the wipe would pass
    # unnoticed and every cohort figure would render blank. Catch it here, where the before
    # and after counts are both in hand.
    _n_cohorts_before = _n_cohorts
    # Cohorts are scoped by the rung's universe filter and by nothing else: every
    # prediction set on that rung counts toward K, retired generations included.
    _counts = compute_and_register(
        CASE_STUDY, universe_filter=_UNIVERSE_FILTER, prediction_hashes=ENTIRE_REGISTRY
    )
    _n_cohorts = sum(_counts[k] for k in ("family", "stagelabel", "label"))
    if _n_cohorts == 0:
        # Backtests are registered - the refusal above guarantees it - so a run that computes
        # no cohort means the canonical selection matched nothing, or matched too few variants
        # per group for a cohort to exist. Either way the selection-bias figures have no input,
        # and rendering them blank reads like an analysis that found nothing rather than one
        # that was never computed. That is the same reason this notebook refuses an empty
        # `backtest_runs`, so it refuses here too.
        _lost = (
            f" and replaced {_n_cohorts_before} existing rows with none"
            if _n_cohorts_before > 0
            else ""
        )
        raise RuntimeError(
            f"rebuilding cohort_metrics for {_UNIVERSE_FILTER} produced no cohort{_lost}. "
            f"Check that {_UNIVERSE_FILTER} matches the backtest runs this case study "
            f"registered, and that the sweep left more than one variant per cohort."
        )
    print(f"populated cohort_metrics: {_n_cohorts} rows")
else:
    print(f"already populated: cohort_metrics {_n_cohorts} rows")

# Rebuilt every run, where cohorts above are rebuilt only when they look wrong. The universe
# recovery cannot answer this table's question. It reads what the rows point at, and a pair
# written under an earlier pin points at a backtest that is still inside the current one: when
# `family == "ensemble"` left this case study's pin on 2026-09-14 the rank-1 rung moved from the
# ensemble at +0.566 to `deep_learning/nlinear` at +2.300, and the ensemble row it had already
# written stayed canonical, stayed cost-feasible and stayed `fwd_ret_15m`. Nothing about it
# reads as stale, because being outside the pin and no longer being what the pin selects are
# different things, and only the first is recoverable from the row. A row count answers neither.
# So it is recomputed rather than validated. `replace_all=True` makes the call a complete
# snapshot, and it costs about a second against this notebook's 25 to 40.
# `NO_CARRIER` is the raw-Sharpe ranking inside the producer rather than a resolved
# lineage: this case study pins a rung the canonical resolver does not know about.
_pairs = populate_paired_metrics(
    CASE_STUDY,
    explorer,
    rung=_RUNG,
    replace_all=True,
    carrier=NO_CARRIER,
    prediction_hashes=ENTIRE_REGISTRY,
)
_n_pairs = sum(1 for row in _pairs if "skip" not in row)
print(f"populated backtest_paired_metrics: {_n_pairs} pairs")

# A rebuild that could not produce a canonical table leaves the previous one in place, which
# is the right call for the data - a stale table is recoverable and deleted rows are not - but
# the wrong one for this notebook, which would go on to read those rows and present selection
# statistics computed under a universe the chapter does not claim. Refuse instead of rendering
# them. An unpopulated registry has nothing off-universe to find and never reaches this.
_still_stale = derived_tables_off_canonical_universe(CASE_DIR, _UNIVERSE_FILTER)
if _still_stale:
    raise RuntimeError(
        f"{', '.join(sorted(_still_stale))} still reference runs outside "
        f"{_UNIVERSE_FILTER} after rebuilding. The rebuild produced nothing to replace them "
        f"with, so every figure below would describe a selection this case study does not "
        f"make. Re-run the backtests for {_UNIVERSE_FILTER} before this notebook."
    )


def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
    if point is None:
        return "—"
    p = format(point, fmt)
    if lo is None or hi is None:
        return f"{p} [—, —]"
    return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"


def _fmt(val: float | None, fmt: str = ".4f") -> str:
    return "—" if val is None else format(val, fmt)


# %% [markdown]
# ## §1 Handoff from model analysis
#
# The strategy phase inherits one configuration, chosen once on validation results
# only, across the stages that are eligible for the holdout - signal, allocation
# and risk overlay.
#
# The selection is made by `resolve_solvent_carrier`, the shared resolver, and not by
# ranking a Sharpe column here. The same resolver answers `18_holdout_predictions` and
# `19_holdout_backtest`, so this page reports the configuration those notebooks refitted
# and priced. It refuses rather than selecting past a problem, on three conditions that
# are all permanent: a rank-1 with no recorded `max_drawdown`, one whose drawdown reached
# -100% so its Sharpe is arithmetic on an account that no longer exists, and one fitted
# under a conformal calibration version `run_backtest` will no longer execute.
#
# **On this case study the eligible stages come to one.** `UNIVERSE_RESTRICTIONS` pins
# rank-1 to the `cost_feasible` universe that `backtest.sweep.universe_filter` declares
# canonical, and `15_portfolio_management` and `16_risk_management` both register
# specs carrying no `universe_filter`. So the allocation and risk-overlay stages
# contribute no eligible rows here and the carrier is a baseline-stage configuration.
#
# The label is taken from the selection rather than assumed. The pool spans every
# declared label, so the chosen strategy may rest on a variant label rather than the
# primary one, and every loader downstream follows the selection rather than the default.
#
# **Here it does.** `config/setup.yaml` declares `fwd_ret_15m` primary, and the
# selection lands on `fwd_dir_15m`, the direction variant, because that is where the
# best validation Sharpe is. The consequence is worth stating rather than leaving a
# reader to infer it from two labels appearing in different tables: the holdout was
# spent on the variant. The primary label is where the stage sweeps ran, carrying all
# 60 allocation and all 20 risk-overlay backtests against the variant's none, while
# the variant carries the single holdout row. So the allocation and overlay evidence
# in the sections below describes one label and the out-of-sample evidence in §6
# describes another. Each is read on the configuration that produced it and neither
# stands in for the other.

# %%
carrier = resolve_solvent_carrier(CASE_STUDY)
TOP_HASH = carrier["val_backtest_hash"]
TOP_PHASH = carrier["val_prediction_hash"]
RANK1_FAMILY = carrier["family"]
RANK1_CONFIG = carrier["config_name"]
RANK1_LABEL = carrier["label"]
PRIMARY_LABEL = RANK1_LABEL

_db = CASE_DIR / "run_log" / "registry.db"
with sqlite3.connect(str(_db)) as _con:
    _row = _con.execute(
        "SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
        "ic_hac_lag, ic_pct_positive "
        "FROM prediction_metrics WHERE prediction_hash = ?",
        (TOP_PHASH,),
    ).fetchone()

ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row

print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={PRIMARY_LABEL}")
print(f"        prediction_hash={TOP_PHASH}, backtest_hash={TOP_HASH}")
print()
print("Daily-pooled IC (validation):")
print(f"  IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')}  (HAC, lag={int(ic_lag)})")
print(f"  t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3f}")
print(f"  n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f"  CI status: {ci_status(ic_lo, ic_hi)}")

# %% [markdown]
# The cells above report the selected prediction set's own score and interval
# from `prediction_metrics`. Read the interval rather than the point: it says
# whether the ordering the strategy trades on is distinguishable from no
# ordering at all, which is the precondition for anything downstream.
#
# The probabilistic Sharpe ratio reported later asks a related question about
# the strategy rather than the signal - whether its Sharpe is distinguishable
# from zero once the shape of its return distribution is accounted for. A
# strategy with occasional large gains and many small losses can show a
# flattering Sharpe that this correction discounts.
#
# This case study declares no strategy-specific stopping conditions in
# `setup.yaml`, so §9 evaluates two conditions that apply to any strategy:
# whether the validation interval's lower bound reaches zero, and whether the
# holdout comparison against the benchmark excludes zero on the negative side.
# Each is reported as pass, partial or fail.

# %% [markdown]
# ## §2 Search context, family comparison, and lineage waterfall
#
# The baseline stage produced many validation backtests across the model families
# and label horizons. The selected row is one outcome of that search, and a
# search over many configurations produces a highest value even when none of the
# configurations has an edge.
#
# Two things put it in context. Where it sits within its own family's
# distribution says whether it stands apart from its neighbours or is the top of
# a tightly packed band. How it changes through the pipeline stages says which
# stage moved it, and by how much relative to the uncertainty in that move.

# %%
ctx = explorer.search_context("signal")
search_table = pl.DataFrame(
    [
        {"metric": "Total signal backtests", "value": f"{ctx['total']:,}"},
        {"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
        {"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
        {"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
        {"metric": "% positive Sharpe", "value": f"{ctx['pct_positive']:.1f}%"},
        {"metric": "Top-by-Sharpe in this sweep", "value": f"{ctx['champion_sharpe']:.3f}"},
        {"metric": "Top-by-Sharpe percentile", "value": f"{ctx['champion_percentile']:.1f}%"},
    ]
)
print("Baseline-stage search context:")
print(search_table)

# %% [markdown]
# The family comparison reads each backtest's stored interval rather than
# recomputing one, so the plot can show error bars instead of points alone.
#
# It is scoped to the canonical universe, because the baseline stage holds two and they are
# not distributed evenly across families. `backtest.sweep.signal_passes` re-runs only
# `mechanism_top_n` predictions on the full universe, and those are whichever families
# pass 1 ranked highest, so pooling both would move one family's median by rows the others
# have no counterpart for. Labels are pooled deliberately: the question here is what a
# family's spread of outcomes looks like, and this case study trains most families on
# every label.

# %%
with sqlite3.connect(str(_db)) as _con:
    _famdf = pl.DataFrame(
        _con.execute(
            """
            SELECT
                t.family,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi
            FROM backtest_metrics bm
            JOIN backtest_runs b      ON bm.backtest_hash = b.backtest_hash
            JOIN prediction_sets p    ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t      ON p.training_hash   = t.training_hash
            WHERE b.stage = 'signal'
              AND p.split = 'validation'
              AND bm.sharpe IS NOT NULL
              AND (bm.num_trades IS NULL OR bm.num_trades > 0)
              AND COALESCE(
                  json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full'
              ) = ?
            """,
            (_UNIVERSE_FILTER or "full",),
        ).fetchall(),
        schema=["family", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"],
        orient="row",
    )
if _famdf.is_empty():
    msg = (
        f"No baseline-stage validation backtests on the {_UNIVERSE_FILTER or 'full'} universe "
        f"for {CASE_STUDY}, so there is no family distribution to plot. The rest of this "
        "notebook reads the same universe, so this is the first place a registry missing it "
        "shows up."
    )
    raise RuntimeError(msg)

family_summary = (
    _famdf.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_q25=pl.col("sharpe").quantile(0.25),
        sharpe_q75=pl.col("sharpe").quantile(0.75),
        sharpe_max=pl.col("sharpe").max(),
        pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
    )
    .sort("sharpe_median", descending=True)
)
print("Family-level baseline Sharpe summary:")
print(family_summary)

# %%
fig, ax = plt.subplots(figsize=(9, 4))
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()

ax.errorbar(
    medians,
    y,
    xerr=[medians - q25, q75 - medians],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    label="median ±IQR",
)
ax.scatter(maxima, y, marker="x", color="#C62828", s=60, label="max", zorder=5)
ax.axvline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
ax.set_title("Family-level Signal Sharpe — IQR + max")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
show_with_alt(
    fig,
    "Horizontal chart of validation Sharpe by model family. One row per family, named on "
    "the vertical axis which runs top to bottom. A filled circle marks that family's "
    "median Sharpe with a horizontal bar spanning its interquartile range, and a red "
    "cross marks its single best configuration. A dashed vertical line sits at zero "
    "Sharpe. The gap between a family's median and its cross is the spread the best draw "
    "was taken from.",
)

# %% [markdown]
# **How to read the family distribution.** The bars show each family's spread of
# validation outcomes, not one number per family. That is deliberate: the sweep
# ran many configurations per family, so a family's best result is the maximum of
# a sample, and maxima grow with sample size whether or not the family has an
# edge.
#
# Compare the medians to compare families, and read the width of each band as
# what separates a configuration from its neighbours. Where the bands overlap
# heavily, the families are producing the same range of outcomes and the ordering
# of their maxima is not evidence that one is better.
#
# The cost model charged here is the engine's, applied identically to every
# configuration, so differences between families are not differences in what they
# were charged. Every row is on one universe for the same reason.

# %% [markdown]
# The lineage gives one backtest per pipeline stage for this prediction. Each
# stage's stored interval is pulled alongside it so the change from one stage to
# the next can be read against the uncertainty in each.

# %%
lineage = explorer.champion_lineage(TOP_PHASH)
ci_lo: dict[str, float] = {}
ci_hi: dict[str, float] = {}
for stage_name, info in lineage.items():
    bm = load_backtest_metrics(CASE_STUDY, backtest_hash=info["backtest_hash"])
    if not bm.is_empty():
        row = bm.row(0, named=True)
        ci_lo[stage_name] = row.get("sharpe_ci95_lo")
        ci_hi[stage_name] = row.get("sharpe_ci95_hi")

print("Lineage stages present for selected prediction:")
for s, info in lineage.items():
    lo, hi = ci_lo.get(s), ci_hi.get(s)
    print(f"  {s}: hash={info['backtest_hash']}, Sharpe={_fmt_ci(info['sharpe'], lo, hi)}")

missing_stages = [s for s in ("cost_sensitivity", "risk_overlay") if s not in lineage]
if missing_stages:
    print()
    print(f"Stage transitions not run for this prediction: {missing_stages}")

# %%
fig = plot_sharpe_waterfall(lineage, ci_lo=ci_lo, ci_hi=ci_hi)
show_with_alt(
    fig,
    "Waterfall of Sharpe across the selected configuration's locked lineage: signal, then "
    "allocation, then cost, then risk overlay, one bar per stage in that order. Each bar carries "
    "asymmetric error bars for its block-bootstrap 95 percent confidence interval. Those are "
    "marginal intervals on each stage and not a test between stages: whether one stage differs "
    "from the next is the paired difference printed below the figure, which can exclude zero "
    "while two bars overlap.",
)

# %% [markdown]
# Stage transitions are read from the stored paired comparisons rather than
# recomputed here. A stage registered against a different prediction set has no
# transition for this lineage, and is reported as not run rather than
# reconstructed from unpaired numbers.

# %%
ALLOC_HASH = lineage.get("allocation", {}).get("backtest_hash")
if ALLOC_HASH is not None:
    pair = load_paired_metrics(
        CASE_STUDY,
        challenger_hash=ALLOC_HASH,
        benchmark_kind="signal_leader",
    )
    if not pair.is_empty():
        r = pair.row(0, named=True)
        print("Stage transition: allocation challenger vs signal leader")
        print(
            f"  sharpe_diff = "
            f"{_fmt_ci(r['sharpe_diff'], r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}"
        )
        print(f"  p_value = {r['p_value']:.3f}")
        print(f"  prob_challenger_wins = {r['prob_challenger_wins']:.3f}")
        print(f"  CI status: {ci_status(r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}")

# %% [markdown]
# Live values for the stage transitions printed above are read from
# `backtest_paired_metrics`. On this CS, every stage delta is small in
# magnitude — the position-level overlays do not produce a stage-delta
# CI that excludes zero on the positive side. The signal economics
# remain the binding constraint; allocator choice and overlay choice
# modulate the magnitude of loss rather than producing a credibility-
# resolved positive transition.

# %%
# top_k is an entry-scheme parameter, so the sweep that varies it registers at
# stage=signal on this case study: the carrier's rows are 5, 10 and 20 positions
# there and none at the allocation stage, which `concentration_curve` defaults to.
# Asking for the default raised rather than returning nothing, which is the point
# of that refusal: this cell printed "no concentration data" and dropped its
# figure for as long as the registry held no rows for the carrier at all.
conc_df = explorer.concentration_curve(TOP_PHASH, stage="signal")
if not conc_df.is_empty():
    fig = plot_concentration_curve(conc_df)
    show_with_alt(
        fig,
        "Line chart of Sharpe against portfolio concentration. The horizontal axis is the "
        "number of positions held, top_k, and the vertical axis is the Sharpe the selected "
        "prediction achieves at each. The shape rather than any single point is the content: "
        "it shows whether concentrating the book helped or hurt.",
    )
    best_per_k = conc_df.sort("sharpe", descending=True).group_by("top_k").first().sort("top_k")
    print("Concentration sweep: best Sharpe by top_k:")
    print(best_per_k.select("top_k", "allocator", "sharpe", "max_drawdown"))
else:
    print("No concentration data: no top_k sweep registered for this prediction.")

# %% [markdown]
# **How to read the concentration curve.** Position count trades two things off
# against each other. Holding few names concentrates capital into the ordering's
# strongest views, so a correct ordering is rewarded more and an incorrect one
# punished more, and each name turns over more often as the ranking shifts.
# Holding many names dilutes both, and as the count approaches the size of the
# universe the portfolio converges on the universe itself, leaving nothing for
# the ordering to express.
#
# The useful reading is the shape rather than the maximum. A curve with a clear
# interior peak says breadth is a real decision for this strategy. A flat curve
# says the ordering is not strong enough for concentration to matter, and the
# apparent best count is picking out noise.

# %% [markdown]
# ## §3 Headline performance with uncertainty
#
# The selected specification is the validation-window backtest associated
# with the highest baseline Sharpe. Every metric is reported with
# its block-bootstrap interval from `backtest_metrics`; the equity overlay
# shows the cumulative trajectory against the equal-weight NASDAQ-100
# universe benchmark.

# %%
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)
full = dict(full)  # mutate-friendly

_db = CASE_DIR / "run_log" / "registry.db"
_cohort_leader_hash = None
_cohort_leader_sharpe = None
_cohort_stage = None
with sqlite3.connect(str(_db)) as _con:
    _stage_for_rank1 = _con.execute(
        "SELECT stage FROM backtest_runs WHERE backtest_hash = ?",
        (TOP_HASH,),
    ).fetchone()[0]
for _try_stage in dict.fromkeys((_stage_for_rank1, "allocation", "signal")):
    with sqlite3.connect(str(_db)) as _con:
        _cohort = _con.execute(
            """
            SELECT dsr_er, dsr_er_pvalue, expected_max_sharpe_er,
                   min_trl_periods_er, dsr_mp, dsr_mp_pvalue,
                   expected_max_sharpe_mp, min_trl_periods_mp,
                   dsr_raw, dsr_raw_pvalue,
                   expected_max_sharpe_raw, min_trl_periods_raw,
                   n_trials_effective_mp, n_trials_effective_er,
                   pbo, pbo_n_combinations, pbo_n_folds,
                   k_variants, leader_hash, leader_sharpe
            FROM cohort_metrics
            WHERE cohort_type='family' AND stage=? AND label=? AND family=?
              AND leader_sharpe IS NOT NULL
            """,
            (_try_stage, RANK1_LABEL, RANK1_FAMILY),
        ).fetchone()
    if _cohort is not None:
        (
            full["dsr_er"],
            full["dsr_er_pvalue"],
            full["expected_max_sharpe_er"],
            full["min_trl_periods_er"],
            full["dsr_mp"],
            full["dsr_mp_pvalue"],
            full["expected_max_sharpe_mp"],
            full["min_trl_periods_mp"],
            full["dsr_raw"],
            full["dsr_raw_pvalue"],
            full["expected_max_sharpe_raw"],
            full["min_trl_periods_raw"],
            full["n_trials_effective_mp"],
            full["n_trials_effective_er"],
            full["pbo"],
            full["pbo_n_combinations"],
            full["pbo_n_folds"],
            full["k_variants"],
            _cohort_leader_hash,
            _cohort_leader_sharpe,
        ) = _cohort
        _cohort_stage = _try_stage
        # Mirror DSR_ER onto the legacy `dsr` / `dsr_pvalue` / `expected_max_sharpe`
        # fields so downstream consumers that still read those keys see the
        # ER-default values.
        full["dsr"] = full["dsr_er"]
        full["dsr_pvalue"] = full["dsr_er_pvalue"]
        full["expected_max_sharpe"] = full["expected_max_sharpe_er"]
        full["min_trl_periods"] = full["min_trl_periods_er"]
        break

spec_block = {
    "case_study": CASE_STUDY,
    "family": RANK1_FAMILY,
    "config_name": RANK1_CONFIG,
    "label": RANK1_LABEL,
    "signal_method": lineage["signal"].get("signal_method"),
    "top_k": lineage["signal"].get("top_k"),
    "allocation": lineage.get("allocation", {}).get("allocator"),
    "risk_overlay": lineage.get("risk_overlay", {}).get("risk_name"),
    "rebalance_step_bars": setup["labels"]["rebalance_step"][RANK1_LABEL],
    "cost_assumption": "engine cost model: 5 bps commission + 2 bps slippage at the baseline stage; per_share_plus_spread sensitivity in §5",
    "validation_window_periods": int(full["n_periods"]) if full["n_periods"] is not None else None,
    "num_trades": int(full["num_trades"]) if full["num_trades"] is not None else None,
    "avg_turnover": full["avg_turnover"],
    "bootstrap_block_length": int(full["bootstrap_block_length"])
    if full["bootstrap_block_length"] is not None
    else None,
    "bootstrap_n": int(full["bootstrap_n"]) if full["bootstrap_n"] is not None else None,
}
print("Rank-1 specification (cross-stage selection, validation window):")
for k, v in spec_block.items():
    print(f"  {k}: {v}")

_block = int(full["bootstrap_block_length"])
_rstep = setup["labels"]["rebalance_step"][RANK1_LABEL]
print(
    f"  audit: bootstrap_block_length={_block} day(s); "
    f"rebalance_step={_rstep} bar(s) ({_rstep * 15} minutes)"
)

# Cohort cascade audit: surface where DSR comes from when the selected lineage
# stage's cohort row is missing.
if _cohort_stage and _cohort_stage != _stage_for_rank1:
    print(
        f"  cohort cascade: selected stage `{_stage_for_rank1}` has no "
        f"cohort_metrics row (cross-stage timestamp-alignment limit); "
        f"DSR/EMS/MinTRL/PBO sourced from family cohort `{_cohort_stage}/"
        f"{RANK1_LABEL}/{RANK1_FAMILY}` (k={int(full['k_variants'])}, "
        f"leader_hash={_cohort_leader_hash}, leader_sharpe="
        f"{_cohort_leader_sharpe:+.3f}; spine sibling differs by "
        f"{full['sharpe'] - _cohort_leader_sharpe:+.3f})."
    )
elif _cohort_stage:
    print(
        f"  cohort: family cohort `{_cohort_stage}/{RANK1_LABEL}/"
        f"{RANK1_FAMILY}` (k={int(full['k_variants'])}, leader_hash="
        f"{_cohort_leader_hash}); spine sibling differs by "
        f"{full['sharpe'] - _cohort_leader_sharpe:+.3f}."
    )


# %%
def _row(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
    return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}


sharpe_status = ci_status(full["sharpe_ci95_lo"], full["sharpe_ci95_hi"])
sortino_status = ci_status(full["sortino_ci95_lo"], full["sortino_ci95_hi"])
ann_status = ci_status(full["ann_return_ci95_lo"], full["ann_return_ci95_hi"])
mdd_status = ci_status(full["max_dd_ci95_lo"], full["max_dd_ci95_hi"])
calmar_status = ci_status(full["calmar_ci95_lo"], full["calmar_ci95_hi"])

headline = pl.DataFrame(
    [
        _row(
            "Sharpe",
            _fmt(full["sharpe"]),
            _fmt(full["sharpe_ci95_lo"]),
            _fmt(full["sharpe_ci95_hi"]),
            sharpe_status,
        ),
        _row(
            "Sortino",
            _fmt(full["sortino"]),
            _fmt(full["sortino_ci95_lo"]),
            _fmt(full["sortino_ci95_hi"]),
            sortino_status,
        ),
        _row(
            "Annualized return",
            _fmt(full["cagr"]),
            _fmt(full["ann_return_ci95_lo"]),
            _fmt(full["ann_return_ci95_hi"]),
            ann_status,
        ),
        _row(
            "Max drawdown",
            _fmt(full["max_drawdown"]),
            _fmt(full["max_dd_ci95_lo"]),
            _fmt(full["max_dd_ci95_hi"]),
            mdd_status,
        ),
        _row(
            "Calmar",
            _fmt(full["calmar"]),
            _fmt(full["calmar_ci95_lo"]),
            _fmt(full["calmar_ci95_hi"]),
            calmar_status,
        ),
        {
            "metric": "PSR p-value (H0: SR≤0)",
            "point": _fmt(full["psr_pvalue"]),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_ER (selection-adjusted)",
            "point": _fmt(full.get("dsr_er")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_ER p-value",
            "point": _fmt(full.get("dsr_er_pvalue")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_MP",
            "point": _fmt(full.get("dsr_mp")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_MP p-value",
            "point": _fmt(full.get("dsr_mp_pvalue")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_RAW (naive raw-K)",
            "point": _fmt(full.get("dsr_raw")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_RAW p-value",
            "point": _fmt(full.get("dsr_raw_pvalue")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": (
                f"Expected max Sharpe (k={int(full['k_variants'])})"
                if full.get("k_variants") is not None
                else "Expected max Sharpe"
            ),
            "point": _fmt(full.get("expected_max_sharpe_er")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "PBO (CSCV)",
            "point": _fmt(full.get("pbo"), ".3f"),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
    ]
)
print("Rank-1 headline metrics with 95% CIs:")
print(headline)

# %%
# Forest plot: selected-lineage metrics with CI bars + reference lines
ew_val = load_benchmark_metrics(CASE_STUDY, PRIMARY_LABEL, period="validation")
forest_metrics = [
    ("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
    ("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
    ("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
    ("Ann. return", full["cagr"], full["ann_return_ci95_lo"], full["ann_return_ci95_hi"]),
]

fig, ax = plt.subplots(figsize=(8, 4))
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.errorbar(
    points,
    y,
    xerr=[points - los, his - points],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    markersize=7,
)
ax.axvline(0, color="#9E9E9E", linestyle="--", linewidth=0.8)
ax.axvline(
    ew_val["sharpe"],
    color="#43A047",
    linestyle=":",
    linewidth=1.0,
    label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
_alloc_sharpe = lineage.get("allocation", {}).get("sharpe")
if _alloc_sharpe is not None:
    ax.axvline(
        _alloc_sharpe,
        color="#E53935",
        linestyle=":",
        linewidth=1.0,
        label=f"Allocation selection Sharpe ({_alloc_sharpe:.2f})",
    )
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Value")
ax.set_title("Rank-1 Headline Metrics with 95% CIs")
ax.legend(loc="lower right", fontsize=8, frameon=False)
show_with_alt(
    fig,
    "Forest plot of the selected configuration's headline metrics. One row per metric, named on "
    "the vertical axis, with a point estimate and a horizontal bar for its 95 percent confidence "
    "interval, read on a shared horizontal value axis. A dashed vertical line marks zero, and "
    "further vertical reference lines mark the equal-weight validation Sharpe and the "
    "allocation-stage Sharpe. Those two lines are Sharpe values, so only the Sharpe row is "
    "comparable with them; the Sortino, Calmar and annualized-return rows share the axis but "
    "measure different quantities.",
)

# %% [markdown]
# The strategy's returns are stored daily while the benchmark is stored at the
# decision cadence, so the benchmark is aggregated to daily before the two are
# aligned. Comparing series at different frequencies would otherwise pair a day
# against a fraction of one.

# %%
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
    pl.read_parquet(strat_returns_path)
    .sort("timestamp")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)

bench_val = (
    load_benchmark_returns(CASE_STUDY, PRIMARY_LABEL, period="validation")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .group_by("ts")
    .agg(((1 + pl.col("ew_return")).product() - 1).alias("benchmark"))
    .sort("ts")
)

aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
    f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)

cum_strat = np.cumprod(1 + aligned["strategy"].to_numpy()) - 1
cum_bench = np.cumprod(1 + aligned["benchmark"].to_numpy()) - 1
fig, ax = plt.subplots(figsize=(10, 4.2))
ax.plot(aligned["ts"], cum_strat, color="#1565C0", linewidth=1.2, label="Selected strategy")
ax.plot(aligned["ts"], cum_bench, color="#43A047", linewidth=1.2, label="EW universe")
ax.axhline(0, color="#9E9E9E", linewidth=0.6, linestyle="--")
ax.set_ylabel("Cumulative return")
ax.set_title("Validation-window cumulative return: selected strategy vs EW universe")
ax.legend(loc="best", frameon=False)
show_with_alt(
    fig,
    "Two cumulative return paths over the validation window, plotted against time. One "
    "line is the selected strategy and the other the equal-weight universe, distinguished "
    "in the legend; the vertical axis is cumulative return and a dashed horizontal line "
    "marks zero. The comparison is the point: the strategy's path is only informative "
    "beside the universe it traded in.",
)

# %% [markdown]
# **How to read the selection-bias adjustment.** The deflated Sharpe ratio asks
# what the highest result from a search of this size would look like if none of
# the configurations had an edge, and discounts the observed result by that
# amount. The more configurations searched, the higher the bar it sets.
#
# Its inputs are the number of variants in the cohort searched, the shape of the
# return distribution, and the length of the record. All three are printed above
# rather than assumed, because the count of variants in particular decides how
# large the adjustment is, and a count that quietly includes duplicates of an
# already-chosen configuration inflates the correction.
#
# The companion quantity is the track record length required for a result of this
# size to become distinguishable from zero. When it exceeds the history
# available, the honest statement is that the window cannot settle the question,
# whatever the point estimate.

# %% [markdown]
# ## §4 Risk and drawdown analysis
#
# Risk metrics use the validation-window strategy returns paired against the validation
# EW benchmark. The drawdown panel surfaces the worst episode and whether it recovered;
# rolling Sharpe and rolling beta say whether either was steady or driven by part of the
# window.
#
# The lower panel is worth reading for its shape and not only its depth: whether the
# losses arrived steadily across the window or inside one episode, and whether anything
# was recovered afterwards. That is a description of when the strategy lost, and it is as
# far as the panel goes. It does not identify a cause - a strategy whose costs exceed its
# edge still has winning trades and partial recoveries, and a steady decline can happen
# with no cost at all - so §5 varies the cost directly rather than inferring it from this
# curve.

# %%
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()

pa = PortfolioAnalysis(
    returns=strat_arr,
    benchmark=bench_arr,
    dates=ts_arr,
    periods_per_year=PERIODS_PER_YEAR,
)

dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)

# %%
fig = plot_equity_drawdown(strat_returns_path)
show_with_alt(
    fig,
    "Two stacked panels sharing a time axis. The upper panel is the selected strategy's "
    "cumulative return over the window and the lower panel is its drawdown, the distance "
    "below the running maximum at each moment. Reading a date down both panels pairs a "
    "level of return with the loss being carried to hold it.",
)

# %%
# Rolling Sharpe + rolling beta. `aligned` is daily - the strategy's returns are stored
# that way and the benchmark is compounded to match - so the window is 126 trading days,
# about six months of a validation window that runs roughly 253.
ROLL_WINDOW = 126
roll = pa.compute_rolling_metrics(windows=[ROLL_WINDOW], metrics=["sharpe", "beta"])

# Summarised rather than printed as an object. This used to print the types of the
# result's members, which says nothing about the strategy, and the section's own prose
# promises these two series - a computed result that reaches no reader is the same as one
# never computed.
_roll_rows = []
for _label, _store in (("Rolling Sharpe", roll.sharpe), ("Rolling beta", roll.beta)):
    _series = _store.get(ROLL_WINDOW)
    _valid = _series.drop_nulls() if _series is not None else None
    if _valid is None or _valid.len() == 0:
        _roll_rows.append({"series": _label, "n": 0, "min": "n/a", "median": "n/a", "max": "n/a"})
        continue
    _roll_rows.append(
        {
            "series": _label,
            "n": _valid.len(),
            "min": f"{_valid.min():+.3f}",
            "median": f"{_valid.median():+.3f}",
            "max": f"{_valid.max():+.3f}",
        }
    )
print(f"Rolling metrics over {ROLL_WINDOW} trading days ({aligned.height} days in the window):")
print(pl.DataFrame(_roll_rows))

# Tail-risk read straight from registry (already computed and stored)
tail_table = pl.DataFrame(
    [
        {"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
        {"metric": "VaR 95% (daily)", "value": f"{full['var_95']:.4f}"},
        {"metric": "CVaR 95% (daily)", "value": f"{full['cvar_95']:.4f}"},
        {"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
        {"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
        {"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
    ]
)
print()
print("Tail risk profile:")
print(tail_table)

# %%
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
print(f"Per-fold breakdown ({fold_df.height} folds):")
print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
print()
print(f"Fold Sharpe range: [{fold_df['sharpe'].min():.3f}, {fold_df['sharpe'].max():.3f}]")
print(f"Fold Sharpe std:   {fold_df['sharpe'].std():.3f}")

# %% [markdown]
# **How to read the four printouts above.** Each answers a question the others cannot,
# and none of their values is transcribed here: this notebook re-resolves its carrier from
# the registry on every run, so a number written into this cell would describe whichever
# configuration was rank-1 the day it was written.
#
# *Per-fold Sharpes* say whether the headline belongs to the strategy or to one block.
# Two folds agreeing in sign is weak evidence; two disagreeing is strong evidence against.
# With `evaluation.n_splits: 2` there is no third block to break the tie, so a split
# verdict is a reason to distrust the headline rather than to average it.
#
# *Skewness and kurtosis* say what the distribution behind that Sharpe looks like.
# Negative skew with high kurtosis is the shape a Sharpe flatters, because Sharpe reads
# only the first two moments: many small gains against occasional large losses. Positive
# skew is the opposite trade, and it is what a strategy paying a per-trade cost for
# occasional large wins looks like.
#
# *Tail ratio* is the 95th percentile over the absolute 5th, so it compares where the two
# tails begin and not how much sits beyond those points. Near 1 the two boundaries are the
# same size. It answers a different question from skewness, which the extremes past those
# percentiles drive, so the two can point opposite ways on the same series - rare large
# losses pull the skew negative while leaving the 5th percentile where it was. Disagreement
# between them is a fact about the distribution rather than a sign that one is wrong.
#
# *The rolling window* locates the result in time: each point is the Sharpe of the 126
# trading days ending there, so the summary reports how far the estimate moved over the
# window rather than a single number for the whole period. Read the spread between the
# minimum and the maximum, not the sign. Consecutive windows share 125 of their 126 days,
# so the series is autocorrelated by construction, and a validation window of roughly 253
# days holds about two non-overlapping windows. Sampling variation alone produces sign
# changes on a stationary series at this length, so a crossing is not evidence of a
# regime, and an estimate that stays positive throughout does not establish that
# performance was stable. Rolling beta is read the same way and carries the same
# qualification: each point is an estimate over 126 overlapping days, and a book whose
# population beta is zero still returns nonzero sample betas.
#
# Dollar neutrality and benchmark beta are two different properties and only one of them is
# estimated. Dollar neutrality is a fact about the weights - whether the long and short
# notionals sum to zero - and is read off the allocation directly. Beta is a fact about how
# those positions co-move with the benchmark, so a book can be exactly dollar-neutral and
# still carry beta whenever its longs and shorts differ in benchmark sensitivity. Rolling
# beta is the estimate of that second quantity, and a wandering series is a reason to
# measure the exposure with its uncertainty rather than a reading of it.

# %% [markdown]
# ## §5 Friction budget & cost sensitivity
#
# Cost sensitivity is where a short-rebalancing strategy is decided, because
# every cost is charged per trade and this one trades often. Two grids are read
# here: a **cost grid** that walks commission and slippage from zero upward, and
# a **risk-overlay grid** of rules that close a position on a condition other
# than the signal.
#
# They answer different questions. The cost grid asks how much friction the
# strategy can absorb before its edge runs out, which states the execution
# quality it requires rather than the profit it made under one assumption. The
# overlay grid asks whether trading *less* on positions that are going wrong
# preserves more than the extra trading costs. The two were swept independently,
# with overlays applied at zero engine cost, so each result is attributable to
# the field its sweep varied.
#
# **Neither grid is restricted to the selected lineage, and they do not sit on
# the same universe.** Both read their whole stage across labels, which is what
# the per-horizon stratification below is for. The cost rows come from
# `17_costs`, whose pool is pinned to the canonical `cost_feasible` universe; the
# overlay rows come from `16_risk_management`, which overlays allocation-stage
# specs that carry no `universe_filter` and are therefore full-universe. So a
# cost curve and an overlay row in this section are not two readings of one
# portfolio, and the gradient across thresholds is what §5.2 supports rather than
# a level comparable with §5.1.

# %% [markdown]
# ### §5.1 Cost sensitivity stratified by label horizon

# %%
with sqlite3.connect(str(_db)) as _con:
    cost_df = pl.DataFrame(
        _con.execute(
            """
            SELECT
                b.spec_json,
                t.family,
                t.config_name,
                t.label,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi,
                bm.max_drawdown,
                bm.cagr,
                bm.num_trades,
                bm.avg_turnover
            FROM backtest_runs b
            JOIN backtest_metrics bm ON bm.backtest_hash  = b.backtest_hash
            JOIN prediction_sets p   ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t     ON p.training_hash   = t.training_hash
            WHERE b.stage = 'cost_sensitivity'
              AND p.split = 'validation'
              AND bm.sharpe IS NOT NULL
              AND (bm.num_trades IS NULL OR bm.num_trades > 0)
            """
        ).fetchall(),
        schema=[
            "spec_json",
            "family",
            "config_name",
            "label",
            "sharpe",
            "sharpe_ci95_lo",
            "sharpe_ci95_hi",
            "max_drawdown",
            "cagr",
            "num_trades",
            "avg_turnover",
        ],
        orient="row",
    )


def _cost_bps(spec_str: str) -> float:
    """Cost per leg in bps from the locked spec.

    `backtest_config.commission.rate` and `backtest_config.slippage.rate`
    are decimal fractions; sum × 10000 gives total per-leg cost in bps.
    """
    spec = json.loads(spec_str)
    bc = spec.get("backtest_config", {})
    comm = bc.get("commission", {}) or {}
    slip = bc.get("slippage", {}) or {}
    return float((comm.get("rate", 0) + slip.get("rate", 0)) * 10000)


cost_df = cost_df.with_columns(
    pl.col("spec_json").map_elements(_cost_bps, return_dtype=pl.Float64).alias("cost_bps")
)

# Aggregate per (label, cost_bps): max-config Sharpe + envelope
cost_curve_per_label = (
    cost_df.group_by(["label", "cost_bps"])
    .agg(
        n=pl.len(),
        sharpe_max=pl.col("sharpe").max(),
        sharpe_min=pl.col("sharpe").min(),
        sharpe_median=pl.col("sharpe").median(),
        # CI lower bound *of the best-Sharpe row* per (label, cost):
        # this is the credibility band on the headline cell.
        best_ci_lo=pl.col("sharpe_ci95_lo").top_k_by("sharpe", 1).first(),
        best_ci_hi=pl.col("sharpe_ci95_hi").top_k_by("sharpe", 1).first(),
        best_n_trades=pl.col("num_trades").top_k_by("sharpe", 1).first(),
        best_avg_turnover=pl.col("avg_turnover").top_k_by("sharpe", 1).first(),
        best_family=pl.col("family").top_k_by("sharpe", 1).first(),
        best_config=pl.col("config_name").top_k_by("sharpe", 1).first(),
    )
    .sort(["label", "cost_bps"])
)
print("Per-label cost grid (validation, best Sharpe per (label, cost) cell):")
with pl.Config(tbl_rows=33):
    print(
        cost_curve_per_label.select(
            [
                "label",
                "cost_bps",
                "sharpe_max",
                "best_ci_lo",
                "best_ci_hi",
                "sharpe_median",
                "best_n_trades",
                "best_family",
                "best_config",
            ]
        )
    )

# %%
# Per-label cost curve: one line per label horizon
fig, ax = plt.subplots(figsize=(9.5, 4.5))
label_order = ["fwd_ret_5m", "fwd_ret_15m", "fwd_ret_60m"]
label_colors = {
    "fwd_ret_5m": "#C62828",
    "fwd_ret_15m": "#FB8C00",
    "fwd_ret_60m": "#1565C0",
}
for lbl in label_order:
    sub = cost_curve_per_label.filter(pl.col("label") == lbl).sort("cost_bps")
    if sub.is_empty():
        continue
    xs = sub["cost_bps"].to_numpy()
    ax.plot(
        xs,
        sub["sharpe_max"].to_numpy(),
        color=label_colors[lbl],
        linewidth=1.6,
        marker="o",
        markersize=4,
        label=f"{lbl} (best config)",
    )
    ax.fill_between(
        xs,
        sub["best_ci_lo"].to_numpy(),
        sub["best_ci_hi"].to_numpy(),
        color=label_colors[lbl],
        alpha=0.10,
    )

ax.axhline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.axhline(
    ew_val["sharpe"],
    color="#43A047",
    linewidth=0.9,
    linestyle=":",
    label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
ax.axvspan(1.0, 3.0, color="#43A047", alpha=0.08, label="large-cap spread (1–3 bps)")
ax.axvspan(3.0, 8.0, color="#FB8C00", alpha=0.08, label="mid-cap spread (3–8 bps)")
ax.axvline(
    5.0, color="#C62828", linewidth=0.8, linestyle=":", label="protocol friction floor (5 bps)"
)
ax.set_xlabel("Per-leg cost (bps)")
ax.set_ylabel("Sharpe (validation, best config per cost level)")
ax.set_title("Cost sensitivity by label horizon — best config per (label, cost)")
ax.legend(loc="best", fontsize=8, frameon=False)
show_with_alt(
    fig,
    "Line chart of validation Sharpe against per-leg cost in basis points, one line per label "
    "horizon, each showing the best configuration at every cost level with a shaded band around "
    "it. A dashed horizontal line marks zero Sharpe. Two shaded vertical spans mark realistic "
    "spreads, roughly 1 to 3 basis points for large caps and 3 to 8 for mid caps, and a dotted "
    "vertical line marks the protocol's 5 basis point friction floor. That floor is a reference "
    "and not what this case study charges: config/setup.yaml bills a per-share commission plus "
    "symbol-specific half-spreads, not a flat rate.",
)

# %% [markdown]
# Trade counts at zero cost say how exposed each label is to friction before

Полный текст с указанием источника опубликован на условиях его лицензии. Лицензия: MIT

Это краткое изложение подготовлено исследовательским агентом Stratmill по оригиналу и не является его копией.