مواد پر جائیں
لائبریری کی تمام دستاویزات

غیریقینی کو ملحوظ رکھنے والا ETF اسٹریٹیجی اور بینچ مارک تجزیہ

نوٹ بک Machine Learning for Trading

خلاصہ

یہ نوٹ بک اسٹریٹیجی کے پورے عمل میں رجسٹرڈ ETF بیک ٹیسٹس کا جائزہ لیتی ہے، سگنلز اور اثاثوں کی تخصیص سے لے کر لاگت اور رسک اوورلیز تک۔ یہ ماڈلز فٹ کرنے یا نئے بیک ٹیسٹ شامل کرنے کے بجائے موجودہ پیش گوئیاں اور بیک ٹیسٹس پڑھتی ہے، اور رجسٹرڈ نتائج سے گروہوں اور جوڑی وار تقابل کے پیمانے اخذ کرتی ہے۔ رپورٹ کردہ تجزیے میں اسٹریٹیجی کے پیمانوں کے لیے بوٹسٹریپ اعتماد کے وقفے، انتخابی تعصب کے پیمانے، مرحلہ وار تقابل، اور توثیق و ہولڈ آؤٹ کے ادوار میں مساوی وزن والے ETF کائنات کے بینچ مارک سے تقابل شامل ہے۔

یہ بینچ مارک کے مقابلے میں الفا، بیٹا اور انفارمیشن ریشو بھی دیکھتی ہے، اور فیکٹر ایٹریبیوشن کو پلیسیبو پورٹ فولیو کے ساتھ استعمال کرتی ہے تاکہ انتخاب سے متعلق ایکسپوژرز کو مختلف اثاثوں کی کائنات سے پیدا ہونے والے ایکسپوژرز سے الگ کرنے میں مدد ملے۔ تجزیہ صرف نقطہ تخمینوں کی بنیاد پر اسٹریٹیجیز کی درجہ بندی کرنے کے بجائے غیریقینی برقرار رکھنے کے لیے بنایا گیا ہے۔ اس کے نتائج رجسٹرڈ آبادی، بینچ مارک اور دستیاب ہولڈ آؤٹ ڈیٹا پر منحصر ہیں؛ نوٹ بک کی اپنی وضاحت خبردار کرتی ہے کہ ایکویٹیز، بانڈز، کموڈیٹیز اور کرنسیوں پر مشتمل کائنات میں فیکٹرز کی وضاحتی قوت محدود ہے۔ فراہم کردہ اقتباس طریقۂ کار بیان کرتا ہے، مگر اسٹریٹیجی کے نتائج کے پیمانے نہیں دیتا۔

اہم خیالات

  • بیک ٹیسٹ کی کارکردگی کو غیریقینی کے ساتھ سمجھنے کے لیے بوٹسٹریپ وقفے اور احتمال پر مبنی پیمانے استعمال کریں۔
  • نقطہ تخمینوں کے فرق کے بجائے ریٹرن سیریز کے جوڑی وار تقابل سے مراحل اور اسٹریٹیجیز کا جائزہ لیں۔
  • توثیق اور ہولڈ آؤٹ ادوار میں ETF اسٹریٹیجیز کا مساوی وزن والے کائنات کے بینچ مارک سے جائزہ لیں۔
  • فیکٹر ایٹریبیوشن اور پلیسیبو پورٹ فولیو کائنات سے چلنے والے ایکسپوژرز کو اسٹریٹیجی کے انتخاب کے اثرات سے الگ کرنے میں مدد دیتے ہیں۔
  • تجزیہ رجسٹرڈ بیک ٹیسٹس اور دستیاب ہولڈ آؤٹ ڈیٹا پر منحصر ہے، اور مختلف اثاثوں کی ساخت فیکٹرز کی وضاحتی قوت محدود کرتی ہے۔

ٹیگز

مکمل متن
# ETFs - Strategy Analysis


# ETFs - Strategy Analysis

This notebook reads the ETF case study's whole registered pipeline - every backtest from the
signal sweep through the risk overlay - and turns it into one strategy assessment. Every metric
carries a block-bootstrap confidence interval, every comparison between two strategies goes
through a paired bootstrap rather than a difference of point estimates, and the holdout closure
is read the same way. Comparison across case studies is Chapter 20's subject.

**Learning objectives**

- Read uncertainty-aware backtest metrics - Sharpe with its interval, PSR, DSR - from the
  registry rather than transcribing point estimates.
- Trace one strategy configuration through the four pipeline stages
  (signal to allocation → cost → risk) with paired-bootstrap stage
  transitions.
- Use the equal-weight ETF universe benchmark, sliced into validation and holdout periods, for
  both the equity-curve overlay and the holdout strategy-versus-benchmark paired test.
- Layer 1 + Layer 2 benchmark-aware diagnostics: PortfolioAnalysis
  alpha/beta/IR against the equal-weight universe, plus an FF5+MOM
  factor attribution with placebo-portfolio control. The cross-asset
  ETF universe (equities / bonds / commodities / currencies) means
  factor R² is structurally limited; the placebo benchmark separates
  universe-driven from selection-driven factor exposure.

**Book reference**: Chapter 20, §20.1 (the §9 handoff feeds Ch20's
cross-case-study aggregation).

**Prerequisites**: case-study pipeline through `19_holdout_backtest`;
the locked registry (`case_studies/etfs/run_log/registry.db`).

**Scope**: no training and no re-backtesting. It does write two derived tables,
`cohort_metrics` and `backtest_paired_metrics`, and that is a deliberate change from the
read-only scope this notebook used to declare.

Both tables are derived from backtests that already exist - selection-bias statistics over the
cohorts, and paired-bootstrap comparisons between registered return series. Nothing is refitted
and no backtest is added. They were previously produced by a chapter-20 notebook looping over
every case study, which made a case study's own strategy analysis unreadable until a later
chapter had been run, and left both tables empty for any reader working the case study in order.
A stage that cannot be read without running a chapter that comes after it is not a stage. So the
notebook that has every stage in front of it produces them, and re-running it recomputes only
what is missing.

```python
"""ETFs - Strategy Analysis."""

import json
import sqlite3

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl

# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch  # noqa: F401
import yaml
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
    BacktestReportMetadata,
    generate_tearsheet_from_run_artifacts,
)

from case_studies.research import open_study, split_unpublished_members
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.factor_attribution import (
    compute_bootstrap_ci,
    compute_rolling_exposures,
    format_attribution_summary,
    load_factor_data,
    plot_attribution_waterfall,
    plot_rolling_exposures,
    run_factor_regression,
    run_placebo_benchmark,
)
from case_studies.utils.paired_metrics import populate_paired_metrics
from case_studies.utils.registry import (
    load_backtest_fold_metrics,
    load_backtest_metrics,
    load_paired_metrics,
    load_prediction_index,
)
from case_studies.utils.strategy_analysis import (
    ci_status,
    compute_operating_profile,
    fmt_gate,
    gate1_validation_sharpe_geq_zero,
    gate2_holdout_diff_not_excludes_zero_negatively,
    gate_passes,
    plot_concentration_curve,
    plot_equity_drawdown,
    plot_sharpe_waterfall,
    resolve_holdout_self_backtest,
    resolve_solvent_carrier,
    write_strategy_assessment,
)
from case_studies.utils.uncertainty import STAGE_SEQUENCE, descends_from
from utils.paths import get_case_study_dir, get_output_dir
from utils.style import show_with_alt
```

```python
# MAX_SYMBOLS is gone. Nothing below read it, and a declared cap the run does not apply is
# worse than no cap: a reader who sets it gets a full run and no warning. The two names that
# remain are read - `_declares_tier_and_workspace` (tests/pm_helpers.py) looks for exactly
# this pair - which is why they stay bound although nothing below references them. Without
# them the canonical branch regenerates in place, which needs symlinks a CI checkout has not
# got.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
```

The study is opened before any path or registry read. Under the preview tier, opening it
activates a workspace and rewrites `ML4T_OUTPUT_DIR` process-wide; a `CASE_DIR` or a
`BacktestExplorer` built first would address the released registry while everything after it
reads the preview one.

```python
CASE_STUDY = "etfs"
PRIMARY_LABEL = "fwd_ret_21d"
PERIODS_PER_YEAR = 252  # NYSE calendar, daily bars
study = open_study(CASE_STUDY, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
CASE_DIR = get_case_study_dir(CASE_STUDY)
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

with open(CASE_DIR / "config" / "setup.yaml") as f:
    setup = yaml.safe_load(f)

explorer = BacktestExplorer(CASE_STUDY)
print(explorer)
```

**Which prediction sets their publishers still stand behind.** A refit publishes a second
generation under the same population name, and the generation it replaced stays in the registry:
complete, current under a schema version that has not moved, and carrying every backtest the
previous sweep registered for it. The configuration this notebook describes is chosen by
backtest Sharpe over that pool, so without the lineage the analysis can be about a strategy the
case study no longer publishes.

```python
LIVE_PREDICTIONS = (
    split_unpublished_members(
        study,
        load_prediction_index(CASE_STUDY, split="validation"),
    )
    .live["prediction_hash"]
    .to_list()
)
if not LIVE_PREDICTIONS:
    raise RuntimeError(
        f"no live prediction sets for {CASE_STUDY}/validation; run the model stage and "
        "14_backtest first"
    )
print(f"Live prediction sets: {len(LIVE_PREDICTIONS):,}")

# Both derived tables fill in two waves - stage transitions once allocation, cost and risk have
# run, holdout kinds only once the holdout has been evaluated - so the predicate is whether every
# kind this notebook reads is present, not whether anything is.
PAIRED_KINDS = (
    "signal_leader",
    "allocation_leader",
    "cost_sensitivity_leader",
    "val_rank1_self",
    "equal_weight_holdout_side_artifact",
)
COHORT_TYPES = ("family", "stagelabel", "label")


def _derived_table_state() -> tuple[set[str], set[str]]:
    """Which paired kinds and cohort granularities the registry already holds."""
    with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as db:
        tables = {r[0] for r in db.execute("SELECT name FROM sqlite_master WHERE type='table'")}
        kinds = (
            {
                r[0]
                for r in db.execute("SELECT DISTINCT benchmark_kind FROM backtest_paired_metrics")
            }
            if "backtest_paired_metrics" in tables
            else set()
        )
        cohorts = (
            {r[0] for r in db.execute("SELECT DISTINCT cohort_type FROM cohort_metrics")}
            if "cohort_metrics" in tables
            else set()
        )
    return kinds, cohorts


def _stale_derived_rows(live: list[str]) -> tuple[int, int]:
    """Derived rows built over anything this notebook does not report as live.

    Presence of every enum value says the tables were built; it does not say they were built
    over the population this notebook reports. A refit retires the generation a previous run
    led with, and its cohort and paired rows survive intact under those same enum values.

    A leader and a challenger are the visible halves. A cohort is also stale when a retired
    prediction is one of the variants the correction was computed over - the leader can be
    live while the trial count and the deflated Sharpe are not - and a pair is also stale
    when its benchmark side is retired, which is the side the difference is measured
    against. `member_digest` records the cohort's members, so a cohort whose digest is
    absent cannot be shown to be live either and is recomputed rather than trusted.

    Cohort membership is not queryable here - `backtest_runs` carries neither label nor
    family - so the membership test is that the cohort's stage still holds a retired
    backtest at all. `compute_and_register` refreshes the whole table rather than one row,
    so triggering it too readily costs a recompute and can never report a stale number.
    """
    with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as db:
        tables = {r[0] for r in db.execute("SELECT name FROM sqlite_master WHERE type='table'")}
        payload = json.dumps(live)
        stale_cohort = (
            db.execute(
                """
                SELECT COUNT(*) FROM cohort_metrics cm
                WHERE cm.member_digest IS NULL
                   OR cm.leader_hash IN (
                          SELECT backtest_hash FROM backtest_runs
                          WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
                      )
                   OR EXISTS (
                          SELECT 1 FROM backtest_runs r
                          WHERE r.stage IS cm.stage
                            AND r.prediction_hash NOT IN (SELECT value FROM json_each(?1))
                      )
                """,
                (payload,),
            ).fetchone()[0]
            if "cohort_metrics" in tables
            else 0
        )
        stale_paired = (
            db.execute(
                """
                SELECT COUNT(*) FROM backtest_paired_metrics pm
                WHERE pm.challenger_hash IN (
                          SELECT backtest_hash FROM backtest_runs
                          WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
                      )
                   OR pm.benchmark_hash IN (
                          SELECT backtest_hash FROM backtest_runs
                          WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
                      )
                """,
                (payload,),
            ).fetchone()[0]
            if "backtest_paired_metrics" in tables
            else 0
        )
    return stale_cohort, stale_paired


# §1 reports the configuration the validation stages ranked first, and the paired-metrics
# producer below has to be told what it is. Left to itself the producer ranks the registry on
# raw Sharpe, which is a second selector beside this one - it applies neither the common-support
# re-ranking a conformal candidate forces nor the restrictions the resolver holds. The two agree
# on this registry today, which is why the published rows are right; etfs now has a second fully
# backtested label, so agreement is a fact about this registry rather than a property of either
# ranking. So the selection is resolved here, before anything is written, and the two rankings
# are required to agree rather than assumed to.
```

```python
_HOLDOUT_STAGES = ("signal", "allocation", "risk_overlay")
_FIELD = (
    pl.concat(
        [
            explorer.best(stage=s, top_n=2000, prediction_hashes=LIVE_PREDICTIONS)
            for s in _HOLDOUT_STAGES
        ],
        how="diagonal_relaxed",
    )
    .filter(pl.col("family") != "benchmark")
    .sort("sharpe", descending=True)
    .unique(subset=["prediction_hash"], keep="first", maintain_order=True)
)
ADMITTED = frozenset(_FIELD["backtest_hash"].to_list())
top_signal = _FIELD.head(1)
TOP_HASH = top_signal.row(0, named=True)["backtest_hash"]
TOP_PHASH = top_signal.row(0, named=True)["prediction_hash"]
RANK1_FAMILY = top_signal.row(0, named=True)["family"]
RANK1_CONFIG = top_signal.row(0, named=True)["config_name"]
# Every label the case study declares is backtested equal-weight and competes here; what wins
# decides the label everything below is keyed to. Reading `labels.primary` instead was only ever
# right by coincidence, and etfs has a second fully backtested label, so the coincidence is not
# one to rely on. The primary is kept as the declared value, to say when the two differ.
SELECTED_LABEL = top_signal.row(0, named=True)["label"]

# The resolver is given the same field, not the whole registry: it re-ranks on common timestamp
# support wherever a conformal candidate is present, so a row this notebook never admitted would
# otherwise decide how far the intersection reaches and therefore which admitted row wins.
#
# `resolve_solvent_carrier` rather than the bare lineage resolver, so a selected configuration whose
# equity reached zero is refused rather than reported. A long-short book with no margin call keeps
# compounding through zero, so every metric it reports after that point - including a Sharpe high
# enough to top a ranking - is arithmetic on a balance that no longer exists.
CARRIER = resolve_solvent_carrier(CASE_STUDY, admitted=ADMITTED)
if CARRIER["val_backtest_hash"] != TOP_HASH:
    raise RuntimeError(
        "this notebook's ranking and the canonical resolver disagree on the selected configuration: "
        f"{TOP_HASH} against {CARRIER['val_backtest_hash']}. Everything below reports the "
        "first and the paired rows would be written against the second, so the decay would "
        "compare the right holdout against a different strategy."
    )
```

```python
have_kinds, have_cohorts = _derived_table_state()
stale_cohort, stale_paired = _stale_derived_rows(LIVE_PREDICTIONS)
missing_kinds = sorted(set(PAIRED_KINDS) - have_kinds)
missing_cohorts = sorted(set(COHORT_TYPES) - have_cohorts)
if missing_cohorts or stale_cohort:
    reason = (
        f"{len(missing_cohorts)} granularity(ies) missing"
        if missing_cohorts
        else f"{stale_cohort} row(s) led by a retired prediction"
    )
    counts = compute_and_register(CASE_STUDY, prediction_hashes=LIVE_PREDICTIONS)
    print(
        f"cohort_metrics: recomputed {sum(counts.values())} rows across {sorted(counts)} ({reason})"
    )
else:
    print(
        f"cohort_metrics: {len(have_cohorts)} granularities present and every leader is live, "
        "nothing recomputed"
    )
if missing_kinds or stale_paired:
    reason = (
        f"missing {', '.join(missing_kinds)}"
        if missing_kinds
        else f"{stale_paired} pair(s) challenged by a retired prediction"
    )
    # `replace_all=False` is additive: the pairs this call does not produce stay. That is
    # what this notebook has always done, and it is stated now because the argument
    # decides what the table a reader loads below contains.
    rows = populate_paired_metrics(
        CASE_STUDY,
        prediction_hashes=LIVE_PREDICTIONS,
        carrier=CARRIER,
        replace_all=False,
    )
    written = sum(1 for r in rows if "skip" not in r)
    print(f"backtest_paired_metrics: wrote {written} pairs ({reason})")
else:
    print(
        "backtest_paired_metrics: every kind this notebook reads is present and every "
        "challenger is live, nothing recomputed"
    )

# Both producers are additive: they write the rows this run's population produces and leave
# every other row where it is. So the rebuild above cannot remove a cohort or a pair built over
# a generation that has since been retired, and re-checking only which *kinds* are present would
# report a table that still holds them as clean. The staleness check is run again and its
# residual stated.
#
# Neither producer is asked to prune. `_prune_paired_metrics` deletes the complement of what a
# run wrote, and a run scoped to one label's live predictions has not written the pairs Chapter
# 20 registered for the other labels - pruning to that partial set would delete them. The rows
# are inert for this notebook either way: every reader below resolves a pair by an explicit
# challenger hash taken from the live lineage, never by scanning the table, so a stale row has
# nothing here that would read it. It is stated rather than ignored because that is a property
# of this notebook's readers and not of the table.
residual_cohort, residual_paired = _stale_derived_rows(LIVE_PREDICTIONS)
if residual_cohort or residual_paired:
    print(
        f"derived tables still hold {residual_cohort} cohort row(s) and {residual_paired} "
        "pair(s) from retired generations; no reader below resolves either"
    )
else:
    print("derived tables hold no rows from retired generations")

have_kinds, have_cohorts = _derived_table_state()
STILL_MISSING_KINDS = sorted(set(PAIRED_KINDS) - have_kinds)
if STILL_MISSING_KINDS:
    # Named rather than left to surface as an empty frame eight cells later. The holdout kinds
    # are absent until the holdout has been evaluated, which is a stage that has not run rather
    # than a failure of this one.
    print(f"  still unavailable: {', '.join(STILL_MISSING_KINDS)}")


def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
    """Compact `point [lo, hi]` formatter with NULL-safety."""
    if point is None:
        return "-"
    p = format(point, fmt)
    if lo is None or hi is None:
        return f"{p} [-, -]"
    return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"


def _fmt(val: float | None, fmt: str = ".4f") -> str:
    return "-" if val is None else format(val, fmt)


# The stage transitions below are read off `champion_lineage`, which takes the highest-Sharpe
# backtest at each stage independently. Two consecutive entries therefore share a prediction and
# nothing else: `populate_paired_metrics` says so where it builds them - "the pair is a stage
# comparison, not a demonstrated parent and child". Nothing in `backtest_paired_metrics` records
# which axes moved, so a difference produced by three simultaneous changes and one produced by a
# single change are stored identically and read alike.
#
# The cost stage makes that concrete rather than theoretical. `cost_sensitivity` is a monotone
# grid - the same strategy priced at seventeen cost levels - so its Sharpe maximum is the
# zero-cost point by construction, in this case study and in every other. Taking it as "the cost
# stage's leader" makes the allocation-to-cost transition a comparison of the same returns with
# friction switched off, and reports the saving as a gain the cost model contributed. Its interval
# is tight and its p-value is zero because the two series are nearly identical, which is a
# property of the comparison rather than evidence for it.
#
# So each transition prints the axes that actually differ, and the reading is qualified when more
# than one of them moved.
_COST_KEYS = ("commission", "slippage")


def _axes(backtest_hash: str) -> dict:
    """The comparable axes of one registered backtest, read from its stored specification."""
    with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as db:
        row = db.execute(
            "SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?", (backtest_hash,)
        ).fetchone()
    if row is None or not row[0]:
        return {}
    spec = json.loads(row[0])
    config = spec.get("backtest_config", {})
    strategy = spec.get("strategy", {})
    axes = {
        "allocator": strategy.get("allocation", {}).get("method"),
        "concentration": strategy.get("signal", {}).get("top_k"),
        "risk overlay": (strategy.get("risk") or {}).get("name"),
    }
    for key in _COST_KEYS:
        axes[key] = json.dumps(config.get(key, {}), sort_keys=True)
    return axes


def _priced(axes: dict) -> bool:
    """Whether this backtest charges anything at all to trade."""
    for key in _COST_KEYS:
        model = json.loads(axes.get(key) or "{}")
        if any(type(v) in (int, float) and v > 0 for v in model.values()):
            return True
    return False


def _changed(challenger_hash: str, benchmark_hash: str) -> list[str]:
    """Which axes differ between a transition's two sides, in reading order."""
    chal, bench = _axes(challenger_hash), _axes(benchmark_hash)
    if not chal or not bench:
        return []
    moved = [name for name in chal if chal[name] != bench[name]]
    # The two cost keys always move together here and name one decision, so they read as one.
    if set(_COST_KEYS) <= set(moved):
        moved = [m for m in moved if m not in _COST_KEYS] + ["cost model"]
    return moved
```

## §1 What the strategy phase inherits

The strategy phase does not choose a model. It receives one: **the configuration with the
highest validation backtest Sharpe** across the holdout-eligible stages, where a configuration
is the whole package - model, label, feature set, backtest settings and any risk overlay - and
the checkpoint is part of it. Everything below describes that one configuration rather than
comparing candidates. Note which metric selected it: not the information coefficient, which
orders nothing here, and not a chapter's headline figure. [`13_model_analysis`](13_model_analysis.ipynb)
is where the population was described; the choosing happened in the backtest stages.

The prediction-side information coefficient printed below is the upstream prior on everything
that follows. If the ranking the strategy trades on is not credibly different from zero, a
strategy Sharpe that looks good is a fact about the portfolio construction and the window rather
than about the signal - and the interval on the IC is what says which case this is.

```python
_db = CASE_DIR / "run_log" / "registry.db"
with sqlite3.connect(str(_db)) as _con:
    _row = _con.execute(
        "SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
        "ic_hac_lag, ic_pct_positive "
        "FROM prediction_metrics WHERE prediction_hash = ?",
        (TOP_PHASH,),
    ).fetchone()
ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row

print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={SELECTED_LABEL}")
print(f"        prediction_hash={TOP_PHASH}, backtest_hash={TOP_HASH}")
print()
print("Daily-pooled IC (validation):")
print(f"  IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')}  (HAC, lag={int(ic_lag)})")
print(f"  t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3f}")
print(f"  n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f"  CI status: {ci_status(ic_lo, ic_hi)}")
```

**Read the interval, not the point.** A daily-pooled IC is an average of many daily rank
correlations on an overlapping label, so consecutive days are dependent and the HAC correction is
what makes its interval mean anything. Whether that interval clears zero is the question; the
magnitude is small on this panel either way, because the instruments are themselves diversified
and there is little idiosyncratic variation left to rank.

**What follows from each answer.** An interval clearing zero sets the expectation that the
strategy Sharpe in §3 should also be credible, and makes it worth asking what went wrong if it is
not. An interval spanning zero says the opposite: whatever §3 reports has no upstream support,
and a good Sharpe there needs explaining rather than celebrating.

**Kill conditions are not declared in `setup.yaml`.** §9 evaluates two universal gates: the
validation Sharpe interval's lower bound against zero, and the holdout strategy-versus-equal-
weight paired interval against zero on the negative side. Both are reported as pass, partial or
fail, and neither is a judgement on the strategy.

## §2 Where the leader sits in the search that produced it

A leader is the maximum of a search, and the maximum of a search is not the same quantity as the
performance of a strategy chosen in advance. The table below puts it back in context: how many
backtests the signal stage produced, where their Sharpe ratios fell, and how far above the middle
of that distribution the leader sits.

The wider that distribution and the more configurations in it, the more of the leader's margin is
attributable to having looked. §9's selection-adjusted statistics price that directly; this
section is where the shape it prices becomes visible.

```python
ctx = explorer.search_context("signal", prediction_hashes=LIVE_PREDICTIONS)
search_table = pl.DataFrame(
    [
        {"metric": "Total signal backtests", "value": f"{ctx['total']:,}"},
        {"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
        {"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
        {"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
        {"metric": "% positive Sharpe", "value": f"{ctx['pct_positive']:.1f}%"},
        {"metric": "Top-by-Sharpe in this sweep", "value": f"{ctx['champion_sharpe']:.3f}"},
        {"metric": "Top-by-Sharpe percentile", "value": f"{ctx['champion_percentile']:.1f}%"},
    ]
)
print("Signal-stage search context:")
print(search_table)
```

```python
with sqlite3.connect(str(_db)) as _con:
    _famdf = pl.DataFrame(
        _con.execute(
            """
            SELECT
                t.family,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi
            FROM backtest_metrics bm
            JOIN backtest_runs b ON bm.backtest_hash = b.backtest_hash
            JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t  ON p.training_hash = t.training_hash
            WHERE b.stage = 'signal'
              AND p.split = 'validation'
              AND bm.sharpe IS NOT NULL
              AND (bm.num_trades IS NULL OR bm.num_trades > 0)
              AND b.prediction_hash IN (SELECT value FROM json_each(?))
            """,
            (json.dumps(LIVE_PREDICTIONS),),
        ).fetchall(),
        schema=["family", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"],
        orient="row",
    )

family_summary = (
    _famdf.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_q25=pl.col("sharpe").quantile(0.25),
        sharpe_q75=pl.col("sharpe").quantile(0.75),
        sharpe_max=pl.col("sharpe").max(),
        pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
    )
    .sort("sharpe_median", descending=True)
)
print("Family-level signal-stage Sharpe summary:")
print(family_summary)
```

```python
fig, ax = plt.subplots(figsize=(9, 4))
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()

ax.errorbar(
    medians,
    y,
    xerr=[medians - q25, q75 - medians],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    label="median ±IQR",
)
ax.scatter(maxima, y, marker="x", color="#C62828", s=60, label="max", zorder=5)
ax.axvline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
ax.set_title("Baseline Sharpe by family: interquartile range and maximum")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
fig.tight_layout()
show_with_alt(
    fig,
    "Validation Sharpe by model family, one row per family: a marker at the median with a bar "
    "spanning the interquartile range, a cross at the family maximum, and a dashed line at zero.",
)
```

**The median and the maximum answer different questions, which is why both are drawn.** A
family's median says what a configuration drawn from it typically does; its maximum says what its
strongest one did, and the strongest is what a search returns. A family whose maximum stands far above
its own median is a family whose leader is mostly a draw from a wide distribution.

**Where the families sit relative to each other matters less than how much they overlap.** Read
the interquartile bars: where they cover each other, the ordering between those families is not
something the signal stage decided, and the selected configuration's family may hold that
position because of the stages that came after rather than because of the signal.

**The lineage waterfall below is where that is settled.** It tracks one prediction through
signal, allocation, cost and risk, so a leader that arrives in the lead late - lifted by portfolio
construction rather than by its ranking - is visible as a rising line rather than as a high
starting point. Neither shape is better; they are different claims about where the performance
came from.

```python
lineage = explorer.champion_lineage(TOP_PHASH)
ci_lo: dict[str, float] = {}
ci_hi: dict[str, float] = {}
for stage_name, info in lineage.items():
    bm = load_backtest_metrics(CASE_STUDY, backtest_hash=info["backtest_hash"])
    if not bm.is_empty():
        row = bm.row(0, named=True)
        ci_lo[stage_name] = row.get("sharpe_ci95_lo")
        ci_hi[stage_name] = row.get("sharpe_ci95_hi")

print("Lineage stages present for rank-1 prediction:")
for s, info in lineage.items():
    lo, hi = ci_lo.get(s), ci_hi.get(s)
    print(f"  {s}: hash={info['backtest_hash']}, Sharpe={_fmt_ci(info['sharpe'], lo, hi)}")
```

```python
fig = plot_sharpe_waterfall(lineage, ci_lo=ci_lo, ci_hi=ci_hi)
show_with_alt(
    fig,
    "One bar per stage of the locked lineage, from the baseline backtest through allocation, cost "
    "and risk overlay, each carrying its block-bootstrap interval as an error bar.",
)
```

```python
# Stage-transition deltas via load_paired_metrics — never recompute paired
# metrics inline. ETFs is one of the few CSs whose rank-1 prediction has
# all four pipeline stages registered against it, so each transition has
# a populated paired row.
# The pairs are the consecutive stages this prediction actually has, taken in
# STAGE_SEQUENCE order, which is the order the backtests run: size positions,
# apply risk controls, then measure what realistic costs take off the winner.
present = [s for s in STAGE_SEQUENCE if lineage.get(s)]
for stage_name in STAGE_SEQUENCE:
    if lineage.get(stage_name) is None:
        print(f"Stage {stage_name}: not run for this prediction")
for prev_stage, stage_name in zip(present, present[1:]):
    kind = f"{prev_stage}_leader"
    stage_info = lineage[stage_name]
    # champion_lineage selects each stage independently, so two entries can be
    # siblings rather than parent and child. The producer writes a paired row
    # only where the later stage carries the earlier one's whole strategy
    # prefix; asking for one it declined to write is normal, and the reason is
    # worth naming rather than reporting as a missing row.
    if not descends_from(
        stage_info.get("_strategy", {}),
        lineage[prev_stage].get("_strategy", {}),
        prev_stage,
    ):
        print(
            f"Stage {stage_name}: not paired against {kind} - this run does not "
            f"descend from the {prev_stage} leader, so the two are siblings and "
            f"the difference between them is not a stage effect."
        )
        continue
    pair = load_paired_metrics(
        CASE_STUDY,
        challenger_hash=stage_info["backtest_hash"],
        benchmark_kind=kind,
    )
    if pair.is_empty():
        print(f"Stage {stage_name}: paired row missing for kind={kind}")
        continue
    r = pair.row(0, named=True)
    print(f"{stage_name} challenger vs {kind}:")
    print(
        f"  sharpe_diff = "
        f"{_fmt_ci(r['sharpe_diff'], r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}"
    )
    print(f"  p_value = {r['p_value']:.3f}")
    print(f"  prob_challenger_wins = {r['prob_challenger_wins']:.3f}")
    print(f"  CI status: {ci_status(r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}")
    moved = _changed(stage_info["backtest_hash"], r["benchmark_hash"])
    print(f"  axes that differ: {', '.join(moved) if moved else 'none'}")
    chal_axes, bench_axes = _axes(stage_info["backtest_hash"]), _axes(r["benchmark_hash"])
    if "cost model" in moved and _priced(bench_axes) and not _priced(chal_axes):
        # Read the other way round, which is the direction the comparison supports.
        print(
            f"  the challenger is the frictionless member of the cost grid and the benchmark is "
            f"priced, so what this measures is the price of friction: {r['sharpe_diff']:.3f} "
            f"Sharpe, not a gain the cost stage produced"
        )
    if len(moved) > 1:
        print(
            f"  {len(moved)} axes moved at once, so none of the difference is attributable to "
            f"{stage_name} specifically"
        )
    print()
```

Read the transitions printed above against the order the backtests run:
positions are sized, risk controls are applied to the sized strategy, and
costs are charged against the strategy that survives both. Each printed row
compares a stage against the leader of the stage before it, so a positive
`sharpe_diff` is what that one step added and nothing else.

Two readings need care. A confidence interval that straddles zero means the
step's contribution is not resolved at this sample - the point estimate still
says which way it leaned, and `prob_challenger_wins` says how often it led
across bootstrap draws, but neither is a rejection. And a transition reported
as siblings rather than a pair is not a gap in the evidence: the two stages
were selected independently and the later one does not carry the earlier
one's configuration, so their difference mixes the stage with everything else
that differs between them, and no paired row is written for it.

```python
# The stage is named rather than defaulted. This notebook plots one line per
# allocator, so it wants the allocation stage, and it happens to be what the old
# default gave it - but most predictions in this registry hold signal rows and no
# allocation rows, so a carrier that had not been through the allocator menu used
# to make this cell raise with the stage nowhere in the call.
conc_df = explorer.concentration_curve(TOP_PHASH, stage="allocation")
if not conc_df.is_empty():
    fig = plot_concentration_curve(conc_df)
    show_with_alt(
        fig,
        "Validation Sharpe against the number of positions held, one marker per top-k with the "
        "best allocator at that k annotated beside it and the best k highlighted.",
    )
    best_per_k = conc_df.sort("sharpe", descending=True).group_by("top_k").first().sort("top_k")
    print("Allocation: best Sharpe by top_k:")
    print(best_per_k.select("top_k", "allocator", "sharpe", "max_drawdown"))
else:
    print("No concentration data - allocation stage absent for this prediction.")
```

**Concentration is the portfolio decision the ranking does not make.** Holding fewer funds uses
more of the signal's confidence and less of the universe's diversification; holding more does the
reverse. On a cross-asset universe that trade has a second edge, because the funds are drawn from
equities, bonds, commodities and currencies, and a tight selection can end up inside one of those
rather than across them.

The curve is read for where it flattens rather than for its maximum. A maximum picked off this
scan is a selection over the same validation window everything else was selected on, and the
concentration the lineage actually uses is the one the allocation stage registered.

## §3 What the strategy earned, with its uncertainty

One strategy, one validation window, every figure with the interval the block bootstrap gives it.
The block bootstrap rather than an independent one because daily strategy returns are serially
dependent, and resampling them independently would produce an interval far tighter than the data
support.

The equity overlay puts the strategy against the equal-weight ETF universe over the same dates.
That benchmark is the honest comparator for this case study: it holds the same instruments with
no model at all, so the gap between the two curves is what the ranking bought.

```python
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)

spec_block = {
    "case_study": CASE_STUDY,
    "family": RANK1_FAMILY,
    "config_name": RANK1_CONFIG,
    "label": SELECTED_LABEL,
    "signal_method": lineage["signal"].get("signal_method"),
    "top_k": lineage["signal"].get("top_k"),
    "allocation": lineage.get("allocation", {}).get("allocator"),
    "cost_assumption": "no costs at signal stage; sensitivity in §5",
    "risk_overlay": lineage.get("risk_overlay", {}).get("risk_name"),
    "validation_window_periods": int(full["n_periods"]),
    "num_trades": int(full["num_trades"]) if full["num_trades"] is not None else None,
    "avg_turnover": full.get("avg_turnover"),
    "bootstrap_block_length": int(full["bootstrap_block_length"]),
    "bootstrap_n": int(full["bootstrap_n"]),
}
print("The selected configuration specification (signal stage, validation window):")
for k, v in spec_block.items():
    print(f"  {k}: {v}")

# Audit: bootstrap_block_length resolves from the label horizon (21 trading
# days for fwd_ret_21d). Rebalance cadence in setup.yaml is monthly
# month-end (~21 bars) - the two encode the same autocorrelation scale.
_block = int(full["bootstrap_block_length"])
print(
    f"  audit: bootstrap_block_length={_block} days "
    f"(rebalance cadence={setup['decision']['cadence']} ≈ 21 bars)"
)
```

```python
sharpe_status = ci_status(full["sharpe_ci95_lo"], full["sharpe_ci95_hi"])
sortino_status = ci_status(full["sortino_ci95_lo"], full["sortino_ci95_hi"])
ann_status = ci_status(full["ann_return_ci95_lo"], full["ann_return_ci95_hi"])
mdd_status = ci_status(full["max_dd_ci95_lo"], full["max_dd_ci95_hi"])
calmar_status = ci_status(full["calmar_ci95_lo"], full["calmar_ci95_hi"])


def _hrow(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
    return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}


# The selection-adjusted columns join in from cohort_metrics and carry no per-row interval, so
# their CI cells read as unavailable rather than as a computed bound.
headline = pl.DataFrame(
    [
        _hrow(
            "Sharpe",
            _fmt(full["sharpe"]),
            _fmt(full["sharpe_ci95_lo"]),
            _fmt(full["sharpe_ci95_hi"]),
            sharpe_status,
        ),
        _hrow(
            "Sortino",
            _fmt(full["sortino"]),
            _fmt(full["sortino_ci95_lo"]),
            _fmt(full["sortino_ci95_hi"]),
            sortino_status,
        ),
        _hrow(
            "Annualized return",
            _fmt(full["cagr"]),
            _fmt(full["ann_return_ci95_lo"]),
            _fmt(full["ann_return_ci95_hi"]),
            ann_status,
        ),
        _hrow(
            "Max drawdown",
            _fmt(full["max_drawdown"]),
            _fmt(full["max_dd_ci95_lo"]),
            _fmt(full["max_dd_ci95_hi"]),
            mdd_status,
        ),
        _hrow(
            "Calmar",
            _fmt(full["calmar"]),
            _fmt(full["calmar_ci95_lo"]),
            _fmt(full["calmar_ci95_hi"]),
            calmar_status,
        ),
        {
            "metric": "PSR p-value (H0: SR≤0)",
            "point": _fmt(full["psr_pvalue"]),
            "ci95_lo": "-",
            "ci95_hi": "-",
            "status": "n/a",
        },
        {
            "metric": "DSR (selection-adjusted)",
            "point": _fmt(full["dsr"]),
            "ci95_lo": "-",
            "ci95_hi": "-",
            "status": "n/a",
        },
        {
            "metric": "Expected max Sharpe",
            "point": _fmt(full["expected_max_sharpe"]),
            "ci95_lo": "-",
            "ci95_hi": "-",
            "status": "n/a",
        },
        {
            "metric": "PBO",
            "point": _fmt(full["pbo"]),
            "ci95_lo": "-",
            "ci95_hi": "-",
            "status": "n/a",
        },
    ]
)
print("The selected configuration headline metrics with 95% CIs:")
print(headline)
```

```python
# Forest plot: rank-1 metrics with CI bars + reference lines
ew_val = load_benchmark_metrics(CASE_STUDY, SELECTED_LABEL, period="validation")
forest_metrics = [
    ("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
    ("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
    ("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
    ("Ann. return", full["cagr"], full["ann_return_ci95_lo"], full["ann_return_ci95_hi"]),
]

fig, ax = plt.subplots(figsize=(8, 4))
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.errorbar(
    points,
    y,
    xerr=[points - los, his - points],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    markersize=7,
)
ax.axvline(0, color="#9E9E9E", linestyle="--", linewidth=0.8)
# `load_benchmark_metrics` returns None when the benchmark JSON is absent, which its docstring
# states and which is the ordinary case in a workspace that holds only what this run wrote. The
# reference line is a comparison against the equal-weight baseline, so without it there is
# nothing to draw - and drawing the rest of the forest is still worth doing. Same shape as the
# risk-overlay line below, which has always been conditional.
if ew_val is not None:
    ax.axvline(
        ew_val["sharpe"],
        color="#43A047",
        linestyle=":",
        linewidth=1.0,
        label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
    )
risk_sharpe = lineage.get("risk_overlay", {}).get("sharpe")
if risk_sharpe is not None:
    ax.axvline(
        risk_sharpe,
        color="#E53935",
        linestyle=":",
        linewidth=1.0,
        label=f"Risk-overlay leading Sharpe ({risk_sharpe:.2f})",
    )
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Value")
ax.set_title("The selected configuration Headline Metrics with 95% CIs")
ax.legend(loc="lower right", fontsize=8, frameon=False)
fig.tight_layout()
show_with_alt(
    fig,
    "One row per headline metric, each a point estimate with a bar spanning its 95% interval, "
    "against a dashed line at zero and dotted reference lines for the equal-weight and risk- "
    "overlay comparisons where those are available.",
)
```

```python
# Equity-curve overlay vs validation EW benchmark
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
    pl.read_parquet(strat_returns_path)
    .sort("timestamp")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)

bench_val = (
    load_benchmark_returns(CASE_STUDY, SELECTED_LABEL, period="validation")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("ew_return").alias("benchmark"))
)

aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
    f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)

cum_strat = np.cumprod(1 + aligned["strategy"].to_numpy()) - 1
cum_bench = np.cumprod(1 + aligned["benchmark"].to_numpy()) - 1
fig, ax = plt.subplots(figsize=(10, 4.2))
ax.plot(
    aligned["ts"],
    cum_strat,
    color="#1565C0",
    linewidth=1.2,
    label="The selected configuration strategy",
)
ax.plot(aligned["ts"], cum_bench, color="#43A047", linewidth=1.2, label="EW universe")
ax.axhline(0, color="#9E9E9E", linewidth=0.6, linestyle="--")
ax.set_ylabel("Cumulative return")
ax.set_title("Validation-window cumulative return: rank-1 strategy vs EW universe")
ax.legend(loc="best", frameon=False)
fig.tight_layout()
show_with_alt(
    fig,
    "Cumulative return over the validation window, one line for the selected strategy and one for "
    "the equal-weight universe, against a dashed line at zero.",
)
```

**The lower bound of the Sharpe interval is what the first gate reads**, and it answers a
narrower question than the point estimate: not how well the strategy did, but whether this window
is consistent with it having no edge at all. A point estimate cannot answer that; an interval can.

**The probabilistic Sharpe ratio is a second, parametric answer to the same question.** It
corrects for the skew and kurtosis a Sharpe ratio assumes away. Where the two disagree, the
bootstrap interval is the one that made fewer assumptions.

**The drawdown interval is not a test.** A drawdown is negative by construction, so an interval
excluding zero says nothing. What it bounds is the magnitude, and the width is what to size
against rather than the point.

**The selection-adjusted statistics are what price the search**, and they read as unavailable
rather than as favourable when `cohort_metrics` holds no row for this lineage. A deflated Sharpe
absent from a table is not a deflation of zero. Where they are absent the in-sample guard is the
probabilistic Sharpe alone and the decisive evidence moves to §6 - a weaker position, and worth
saying so.

## §4 Risk and drawdown analysis

Risk metrics use the validation-window strategy returns paired
against the validation EW benchmark. The drawdown panel surfaces the
worst episode and recovery; rolling Sharpe and rolling beta locate
when the strategy decoupled from the universe. ETFs are among the
most liquid instruments in this book, so drawdowns reflect signal
decay or cross-asset rotation reversals rather than execution
slippage.

```python
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()

pa = PortfolioAnalysis(
    returns=strat_arr,
    benchmark=bench_arr,
    dates=ts_arr,
    periods_per_year=PERIODS_PER_YEAR,
)

dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)
```

```python
fig = plot_equity_drawdown(strat_returns_path)
show_with_alt(
    fig,
    "Two panels sharing a date axis: cumulative return of the strategy above, and its peak-to- "
    "trough drawdown below.",
)
```

```python
# Rolling Sharpe + rolling beta (window 126 ~ 6 months)
roll = pa.compute_rolling_metrics(windows=[126], metrics=["sharpe", "beta"])
print("Rolling-window keys:")
print({k: type(v).__name__ for k, v in roll.items()} if isinstance(roll, dict) else roll)

# Tail risk straight from registry
tail_table = pl.DataFrame(
    [
        {"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
        {"metric": "VaR 95% (daily)", "value": f"{full['var_95']:.4f}"},
        {"metric": "CVaR 95% (daily)", "value": f"{full['cvar_95']:.4f}"},
        {"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
        {"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
        {"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
    ]
)
print()
print("Tail risk profile:")
print(tail_table)
```

```python
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
if fold_df.height > 0:
    print(f"Per-fold breakdown ({fold_df.height} folds):")
    print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
    print()
    print(f"Fold Sharpe range: [{fold_df['sharpe'].min():.3f}, {fold_df['sharpe'].max():.3f}]")
    print(f"Fold Sharpe std:   {fold_df['sharpe'].std():.3f}")
else:
    print(
        "Per-fold metrics not populated for this backtest_hash. "
        "ETFs rank-1 was bootstrapped against the consolidated validation "
        "window rather than expanding-window folds; the strategy headline "
        "uses the consolidated CI, and §6's val→ho paired test substitutes "
        "for an explicit per-fold stability check."
    )
```

**Depth and duration are different risks and only one is in the Sharpe.** A strategy can recover
quickly from a deep fall or sit underwater for years after a shallow one, and the second is what
ends a mandate. The panel reports both.

**Kurtosis decides whether the interval above can be trusted.** Heavy tails mean the observed
Sharpe rests on fewer effective observations than the sample size suggests, and a risk overlay
clipping the left tail harder than the right shows up here as skew moving with the overlay rather
than with the signal.

**Rolling Sharpe and rolling beta locate when the strategy stopped tracking its universe.** A
cross-asset strategy that is long equities most of the time has a beta near one and no
diversifying behaviour to show for itself. Where beta falls is where the model rotated into bonds
or commodities, and whether that helped is in the rolling Sharpe over the same dates.

## §5 How much friction the strategy tolerates

ETFs are among the most liquid instruments in this book: the spread on the largest funds is a
fraction of a basis point, while sector, international and commodity funds cost several. The cost
stage walked a per-leg grid over the selected configuration; the curve below is how its Sharpe
responds, with bands marking the most-liquid end of the universe and the typical one.

The reading that matters is the distance between the declared cost and where the curve crosses
zero, not the Sharpe at any single level. [`17_costs`](17_costs.ipynb) computes that crossing
directly and reports it as a bound when the grid does not reach it.

```python
with sqlite3.connect(str(_db)) as _con:
    cost_df = pl.DataFrame(
        _con.execute(
            """
            SELECT
                b.spec_json,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi,
                bm.max_drawdown
            FROM backtest_runs b
            JOIN backtest_metrics bm ON bm.backtest_hash = b.backtest_hash
            JOIN prediction_sets p   ON b.prediction_hash = p.prediction_hash
            WHERE b.stage = 'cost_sensitivity'
              AND p.split = 'validation'
              AND bm.sharpe IS NOT NULL
              AND (bm.num_trades IS NULL OR bm.num_trades > 0)
            """
        ).fetchall(),
        schema=[
            "spec_json",
            "sharpe",
            "sharpe_ci95_lo",
            "sharpe_ci95_hi",
            "max_drawdown",
        ],
        orient="row",
    )


def _cost_bps(spec_str: str) -> float:
    """Per-leg cost (bps) extracted from locked spec.

    Cost rates live under backtest_config.commission.rate +
    backtest_config.slippage.rate as decimal fractions; sum × 10_000
    is total per-leg cost in bps.
    """
    spec = json.loads(spec_str)
    bc = spec.get("backtest_config", {})
    comm = bc.get("commission", {}) or {}
    slip = bc.get("slippage", {}) or {}
    return float((comm.get("rate", 0) + slip.get("rate", 0)) * 10000)


cost_df = cost_df.with_columns(
    pl.col("spec_json").map_elements(_cost_bps, return_dtype=pl.Float64).alias("cost_bps")
)
cost_curve = (
    cost_df.group_by("cost_bps")
    .agg(
        sharpe_max=pl.col("sharpe").max(),
        sharpe_min=pl.col("sharpe").min(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_ci_lo=pl.col("sharpe_ci95_lo").min(),
        sharpe_ci_hi=pl.col("sharpe_ci95_hi").max(),
        n=pl.len(),
    )
    .sort("cost_bps")
)
print("Cost sensitivity curve (validation, all configs):")
print(cost_curve)
```

```python
fig, ax = plt.subplots(figsize=(9, 4))
xs = cost_curve["cost_bps"].to_numpy()
ax.fill_between(
    xs,
    cost_curve["sharpe_ci_lo"].to_numpy(),
    cost_curve["sharpe_ci_hi"].to_numpy(),
    alpha=0.18,
    color="#5B9BD5",
    label="best–worst CI envelope across configs",
)
ax.plot(
    xs,
    cost_curve["sharpe_median"].to_numpy(),
    color="#1565C0",
    linewidth=1.4,
    label="median Sharpe",
)
ax.plot(
    xs,
    cost_curve["sharpe_max"].to_numpy(),
    color="#43A047",
    linewidth=1.0,
    linestyle="--",
    label="best-config Sharpe",
)
ax.axhline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
# Realistic ETF friction:
# - Most-liquid ETF spread (SPY/QQQ/IWM): 0.2-1 bps round-trip
# - Typical ETF spread: 2-5 bps round-trip
ax.axvspan(0.2, 1.0, color="#43A047", alpha=0.10, label="most-liquid ETF (0.2–1 bps)")
ax.axvspan(2.0, 5.0, color="#FB8C00", alpha=0.10, label="typical ETF (2 to 5 bps)")
ax.set_xlabel("Per-leg cost (bps)")
ax.set_ylabel("Sharpe (validation)")
ax.set_title("Cost sensitivity - etfs (validation, baseline+allocation+cost stages)")
ax.legend(loc="best", fontsize=8, frameon=False)
fig.tight_layout()
show_with_alt(
    fig,
    "Sharpe against per-leg cost in basis points: the median across configurations as a line, the "
    "best configuration dashed, the best-to-worst envelope shaded, a dashed line at zero, and "
    "shaded bands for the most-liquid and typical ETF spread ranges.",
)
```

```python
# Breakeven cost: where the best-config Sharpe lower bound crosses zero.
crossing_rows = cost_curve.filter(pl.col("sharpe_ci_lo") > 0)
if not crossing_rows.is_empty():
    breakeven = crossing_rows["cost_bps"].max()
    print(f"Sharpe CI lower bound stays > 0 up to: {breakeven:.0f} bps")
else:
    print("Sharpe CI lower bound never exceeds 0 across the cost grid.")

# Setup-encoded cost configuration
cost_config = setup.get("costs", {})
print()
print("Realistic ETF friction (per setup.yaml):")
for k, v in cost_config.items():
    print(f"  {k}: {v}")
print("See Chapter 18 for the transaction-cost framework.")
```

**Read the dispersion across configurations against the slope of the cost curve.** Where the
spread between configurations at one cost level is wider than the effect of moving several
levels, friction is not what decides this strategy's fate and the choice of configuration is.
That is the ordinary situation for a monthly strategy, and it is a statement about the cadence
rather than about the model. [`17_costs`](17_costs.ipynb) computes the crossing directly.

## §6 The holdout, opened once

Everything above is measured on the window the strategy was selected on. This section is the only
out-of-sample evidence in the case study, and the holdout may be spent once: it is read here, and
nothing that follows may be used to choose anything.

Two paired tests. The first asks whether the edge held between the window that selected the
strategy and the window that did not. The second asks whether it beat the equal-weight universe
*inside* the holdout - the same instruments, the same dates, no model. Both come from
`backtest_paired_metrics`, never from subtracting one Sharpe from another.

The anchor is the holdout backtest that replays the selected configuration's **strategy**, not
the highest-Sharpe holdout backtest sharing its training hash. Matching on the strategy keeps the
anchor on the lineage that was actually selected, even where an experimental side-channel
allocator shares the holdout prediction set and posts a higher holdout Sharpe. The
`val_rank1_self` pair is written against the canonical lineage's holdout hash, so it is findable
only under that match.

When no such backtest exists the section reports which state the registry is in and computes
nothing further. The ordinary case is that the holdout has not been evaluated yet, which a
reader working the case study in order meets before it has, and that is a stage that has not
run rather than a failure of this one.

```python
holdout_replay = resolve_holdout_self_backtest(CASE_STUDY, TOP_HASH)
HOLDOUT_AVAILABLE = holdout_replay.found
HO_HASH = holdout_replay.backtest_hash

print(f"Validation rank-1 hash: {TOP_HASH}")
if HOLDOUT_AVAILABLE:
    print(f"Holdout rank-1 hash:    {HO_HASH}")
else:
    print(f"Holdout closure unavailable: {holdout_replay.reason}")

val_full = full
ho_full = (
    load_backtest_metrics(CASE_STUDY, backtest_hash=HO_HASH).row(0, named=True)
    if HOLDOUT_AVAILABLE
    else None
)
```

A missing paired row is reported as missing rather than filled with NaN. A NaN decay propagates
into the table below, into the section-9 gate, and into the assessment artifact, where it prints
as a dash that a reader cannot distinguish from a computed zero - and the gate would then be
evaluated on a comparison that was never made. Nothing here is computed from a substitute.

```python
vh = None
if HOLDOUT_AVAILABLE:
    val_ho_pair = load_paired_metrics(
        CASE_STUDY, challenger_hash=HO_HASH, benchmark_kind="val_rank1_self"
    )
    if val_ho_pair.is_empty():
        print(
            "The holdout replay is registered but carries no val_rank1_self pair, so the "
            "validation-to-holdout decay cannot be computed. The populator writes that pair "
            "only when both series overlap enough to bootstrap."
        )
    else:
        vh = val_ho_pair.row(0, named=True)
VAL_HO_DECAY_AVAILABLE = vh is not None


def _diff_row(
    label: str,
    v: float,
    h: float,
    diff: float | None,
    lo: float | None,
    hi: float | None,
    p: float | None,
) -> dict:
    return {
        "metric": label,
        "validation": _fmt(v, ".4f") if v is not None else "-",
        "holdout": _fmt(h, ".4f") if h is not None else "-",
        "diff (h-v)": _fmt(diff, ".4f") if diff is not None else "-",
        "diff CI95": (f"[{_fmt(lo, '.4f')}, {_fmt(hi, '.4f')}]" if lo is not None else "-"),
        "p-value": _fmt(p, ".4f") if p is not None else "-",
    }


if not VAL_HO_DECAY_AVAILABLE:
    print("No validation-to-holdout decay to report.")
else:
    val_ho_table = pl.DataFrame(
        [
            _diff_row(
                "Sharpe",
                val_full["sharpe"],
                ho_full["sharpe"],
                vh["sharpe_diff"],
                vh["sharpe_diff_ci95_lo"],
                vh["sharpe_diff_ci95_hi"],
                vh["p_value"],
            ),
            _diff_row(
                "Annualized return",
                val_full["cagr"],
                ho_full["cagr"],
                vh["ret_diff"],
                vh["ret_diff_ci95_lo"],
                vh["ret_diff_ci95_hi"],
                None,
            ),
            _diff_row(
                "Max drawdown",
                val_full["max_drawdown"],
                ho_full["max_drawdown"],
                vh["max_dd_diff"],
                vh["max_dd_diff_ci95_lo"],
                vh["max_dd_diff_ci95_hi"],
                None,
            ),
            _diff_row(
                "Information ratio",
                None,
                None,
                vh["info_ratio"],
                vh["info_ratio_ci95_lo"],
                vh["info_ratio_ci95_hi"],
                None,
            ),
        ]
    )
    print("validation to holdout paired-bootstrap decay (rank-1 self):")
    print(val_ho_table)
    print(f"prob_challenger_wins: {vh['prob_challenger_wins']:.3f}")
    print(
        "CI status (Sharpe diff): "
        f"{ci_status(vh['sharpe_diff_ci95_lo'], vh['sharpe_diff_ci95_hi'])}"
    )
    print()
    print(
        "The validation and holdout windows share no observations, so there is no difference "
        "series to pair on and the populator bootstraps each window separately over its whole "
        "length. Nothing is truncated and no two draws are paired. That is the absence of a "
        "pairing, not independence - the two Sharpes are the same strategy in adjacent periods "
        "and stay dependent. Read the interval as

ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT

یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔