Pular para o conteúdo
Todos os documentos da biblioteca

Avaliação de estratégias de microestrutura NASDAQ-100 com incerteza pareada

Notebook Machine Learning for Trading

Resumo

Este notebook transforma backtests de microestrutura NASDAQ-100 registrados em uma avaliação de estratégia, sem treinar modelos nem executar novamente os backtests. Ele lê métricas e linhagem do registro, cria tabelas derivadas de grupos e comparações pareadas quando necessário e compara estratégias com um benchmark de pesos iguais nas mesmas datas. Seu método central é analisar as diferenças diárias de retorno, em vez de subtrair estatísticas-resumo calculadas separadamente, preservando as informações das observações correspondentes. Ele usa intervalos de bootstrap em blocos para considerar a dependência entre retornos diários consecutivos.

A análise também acompanha uma estratégia selecionada pelas etapas do fluxo de processamento, avalia evidências de validação e holdout e define critérios que podem falhar. Inclui outros materiais de diagnóstico, como sensibilidade a custos, camadas de risco por posição e análise de exposição a fatores, com resultados reunidos em um artefato de avaliação da estratégia. O documento ressalta que passar por um critério é uma evidência limitada, não uma prova de vantagem duradoura. Os resultados dependem do universo de candidatos registrado e do histórico disponível; a incerteza continua relevante, e o notebook reserva deliberadamente as comparações entre estudos de caso para uma síntese separada.

Ideias principais

  • Compare estratégias nas mesmas datas usando diferenças pareadas de retorno para preservar informações compartilhadas do mercado.
  • Intervalos de bootstrap em blocos preservam parte da dependência serial que uma reamostragem diária independente ignoraria.
  • A linhagem do registro pode mostrar qual etapa do fluxo e qual configuração produziram o resultado avaliado.
  • Os critérios devem ser definidos de modo que possam falhar e interpretados como evidências limitadas.
  • Análises de custos, camadas de risco e fatores complementam métricas de desempenho sem eliminar a incerteza.

Tags

Texto completo
# NASDAQ-100 Microstructure — Strategy Analysis


# NASDAQ-100 Microstructure — Strategy Analysis

This notebook turns the backtest registry for the NASDAQ-100 microstructure
case study into one strategy assessment. It reports every metric with an
interval rather than as a point, compares strategies against each other in
pairs rather than by subtracting summary numbers, and reads the holdout once.
Comparison against other case studies is reserved for the synthesis chapter.

Two ideas run through it. **Paired comparison**: two strategies evaluated on
the same days are compared day by day, so the shared market movement cancels
and what is left is the difference between them. Subtracting one summary
statistic from another throws that pairing away and produces an interval far
wider than the evidence warrants. **Block bootstrap**: intervals are built by
resampling contiguous blocks of days rather than individual days, because
returns on consecutive days are related and resampling them independently would
understate the uncertainty.

**Learning objectives**

- Read backtest metrics with their intervals from the registry rather than
  transcribing point estimates
- Trace one selected lineage through the pipeline stages and attribute the
  change at each stage to the field that stage varied
- Compare a strategy against an equal-weight benchmark on the same days, on
  both the validation and the holdout window
- State a gate so that it can fail, and say what passing it does and does not
  establish

**Book reference**: Chapter 20, §20.1 (the §9 handoff feeds Ch20's
cross-case-study aggregation).

**Prerequisites**: case-study pipeline through `14_backtest`; the
locked registry (`case_studies/nasdaq100_microstructure/run_log/registry.db`).

**Scope**: no training and no re-backtesting. The notebook reads what
`14_backtest` through `17_costs` registered, and derives
`cohort_metrics` and `backtest_paired_metrics` from those runs itself when
they are absent, so the case study no longer depends on a chapter-20
notebook having been run first. On an already-populated registry it writes
nothing.

```python
"""NASDAQ-100 Microstructure — Strategy Analysis."""

import json
import sqlite3

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl

# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch  # noqa: F401
import yaml
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
    BacktestReportMetadata,
    generate_tearsheet_from_run_artifacts,
)

from case_studies.research import read_only_study
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.factor_attribution import (
    compute_bootstrap_ci,
    compute_rolling_exposures,
    format_attribution_summary,
    load_factor_data,
    plot_attribution_waterfall,
    plot_rolling_exposures,
    run_factor_regression,
)
from case_studies.utils.notebook_contracts import (
    degenerate_prediction_sql,
    derived_tables_off_canonical_universe,
    strategy_input_counts,
)
from case_studies.utils.paired_metrics import populate_paired_metrics, rung_for
from case_studies.utils.registry import (
    load_backtest_fold_metrics,
    load_backtest_metrics,
    load_paired_metrics,
)
from case_studies.utils.strategy_analysis import (
    ci_status,
    compute_operating_profile,
    fmt_gate,
    gate1_validation_sharpe_geq_zero,
    gate2_holdout_diff_not_excludes_zero_negatively,
    gate_passes,
    plot_concentration_curve,
    plot_equity_drawdown,
    plot_sharpe_waterfall,
    resolve_solvent_carrier,
    write_strategy_assessment,
)
from case_studies.utils.sweep_config import get_universe_filters_for
from case_studies.utils.uncertainty import ENTIRE_REGISTRY, NO_CARRIER
from utils.paths import get_output_dir
from utils.style import show_with_alt

# Figures go through `show_with_alt`, and none of them calls `tight_layout()`. Both are
# measured rather than stylistic: `utils/style` and `matplotlibrc` set
# `figure.constrained_layout.use`, so `tight_layout()` warns and fights the layout engine
# already running, and `fig.show()` on a non-interactive backend warns that the canvas
# cannot be shown - two UserWarnings per figure, written into the rendered cell. The six
# notebooks of this case study already at `done` use this form and neither of the others.
#
# The alt strings name structure, axes and reference lines, never which series wins. An
# ordering is a registry result a rebuild can reverse, and alt text is prose no rebuild
# revisits, so a ranking written here would go stale silently on the one surface a reader
# who cannot see the chart depends on.
```

```python
CASE_STUDY = "nasdaq100_microstructure"
MAX_SYMBOLS = 0
# Where this notebook reads. Empty means the canonical registry; a smoke run passes the
# workspace the stages before it wrote to.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
```

This notebook reads; it registers nothing, and that decides how it opens the registry. Every
route through `open_study` ends in `Study.activate()`, which rewrites `ML4T_OUTPUT_DIR` for
the rest of the process and clears the caches keyed on it, so every later
`get_case_study_dir` answers for a different directory than the one resolved here. On the
canonical tier with no workspace that route is `Study.regenerate`, which refuses outright
unless `features`, `labels` and `run_log` are symlinks - true in a maintainer worktree, false
in every clean clone.

`read_only_study` is the form that does not activate: it resolves the root for the tier and
workspace it is given and hands back a `Study.at` over it, which points the path helpers at
that root rather than clearing the variable. `CASE_DIR` is that root, and every question this
notebook asks - the catalog, the lineage, the populations, the artifacts - is answered from
it.

```python
study = read_only_study(
    CASE_STUDY,
    workspace=WORKSPACE or None,
    execution_tier=EXECUTION_TIER,
    entry_point="20_strategy_analysis",
)
CASE_DIR = study.root
PERIODS_PER_YEAR = 252  # strategy returns are aggregated to daily before Sharpe
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

with open(CASE_DIR / "config" / "setup.yaml") as f:
    setup = yaml.safe_load(f)

explorer = BacktestExplorer(CASE_STUDY)
print(explorer)
```

### Produce the derived tables this notebook reads

`cohort_metrics` and `backtest_paired_metrics` are read below - the first for the deflated
Sharpe and PBO columns, the second to find the validation counterpart of a holdout run. Until
now nothing in this case study wrote either. The header of this notebook said so outright:
they "were populated by `20_strategy_synthesis/01_aggregate_synthesis.py`", a chapter-20
notebook. A case study that depends upward into the chapter that aggregates it cannot be run
from its own pipeline, and its numbers exist only after something outside it happens to run.

Both producers already exist as case-study-scoped functions, extracted from those very Ch20
loops - `populate_paired_metrics` says so in its own docstring - and `cme_futures/17` is the
notebook that uses them. This adopts the same pattern.

Population is idempotent (keyed upserts, reproducing stored values exactly), but it still
rewrites the registry, so it runs only when a table is absent or empty. On an already-populated
registry this notebook stays a pure reader and touches nothing.

```python
_db = CASE_DIR / "run_log" / "registry.db"
_counts_before = strategy_input_counts(CASE_DIR)
_n_backtests = _counts_before["backtest_runs"]
_n_cohorts = _counts_before["cohort_metrics"]
_n_pairs = _counts_before["backtest_paired_metrics"]

if _n_backtests == 0:
    # Refuse rather than report. Every figure and gate below is computed from backtest runs, so
    # with none registered this notebook does not produce a weaker answer - it produces an empty
    # one that reads like a finished analysis. `14_backtest` is what fills this.
    raise RuntimeError(
        f"{CASE_STUDY} has no registered backtest runs, so there is nothing to analyse. "
        "Run 14_backtest through 17_costs first; this notebook reads what they "
        "register and computes no backtests of its own."
    )

# Both producers need this case study's canonical selection, and passing neither is not a
# smaller version of passing both - it is a different selection. `setup.yaml` says the
# full-universe variant "is NOT a canonical rank-1 / cohort / DSR candidate" and lives only in
# the 17_costs comparison, so a cohort computed without the filter mixes rows the case study
# excludes by declaration into the leaders, trial counts, DSR and PBO. The filter is read from
# `setup.yaml` rather than written here, so it cannot disagree with the sweep that produced the
# runs. The rung pin is the matching statement for paired metrics: nasdaq's carrier is the
# cost-feasible ensemble, fixed before the holdout was opened.
_UNIVERSE_FILTER = get_universe_filters_for(CASE_STUDY)[0]
_RUNG = rung_for(CASE_STUDY)

# The two tables are filled independently. A single `if either is empty` guard recomputed both,
# and `compute_and_register` writes with `replace_all=True` - so a registry with correct cohorts
# and an empty paired table would have had its cohorts replaced as a side effect of populating
# the pairs.
# A row count answers "has this been populated", which is not what a rerun needs to know. Both
# tables are written from a selection, and one populated by an earlier run that selected
# differently - before this notebook passed the universe filter, say - is fully populated and
# wrong. Neither table records the selection that produced it, so it is recovered from what the
# rows point at: a canonical row cannot reference a backtest run outside the canonical universe.
_stale = derived_tables_off_canonical_universe(CASE_DIR, _UNIVERSE_FILTER)
for _table in sorted(_stale):
    print(f"{_table} references runs outside {_UNIVERSE_FILTER}, so it is rebuilt")

if _n_cohorts == 0 or "cohort_metrics" in _stale:
    # `compute_and_register` writes with `replace_all=True`, so a rebuild that computes no
    # cohort empties the table rather than leaving the previous one - the opposite of what
    # `populate_paired_metrics` does on the same failure. An emptied table is also clean by
    # the universe check below, which reads rows and finds none, so the wipe would pass
    # unnoticed and every cohort figure would render blank. Catch it here, where the before
    # and after counts are both in hand.
    _n_cohorts_before = _n_cohorts
    # Cohorts are scoped by the rung's universe filter and by nothing else: every
    # prediction set on that rung counts toward K, retired generations included.
    _counts = compute_and_register(
        CASE_STUDY, universe_filter=_UNIVERSE_FILTER, prediction_hashes=ENTIRE_REGISTRY
    )
    _n_cohorts = sum(_counts[k] for k in ("family", "stagelabel", "label"))
    if _n_cohorts == 0:
        # Backtests are registered - the refusal above guarantees it - so a run that computes
        # no cohort means the canonical selection matched nothing, or matched too few variants
        # per group for a cohort to exist. Either way the selection-bias figures have no input,
        # and rendering them blank reads like an analysis that found nothing rather than one
        # that was never computed. That is the same reason this notebook refuses an empty
        # `backtest_runs`, so it refuses here too.
        _lost = (
            f" and replaced {_n_cohorts_before} existing rows with none"
            if _n_cohorts_before > 0
            else ""
        )
        raise RuntimeError(
            f"rebuilding cohort_metrics for {_UNIVERSE_FILTER} produced no cohort{_lost}. "
            f"Check that {_UNIVERSE_FILTER} matches the backtest runs this case study "
            f"registered, and that the sweep left more than one variant per cohort."
        )
    print(f"populated cohort_metrics: {_n_cohorts} rows")
else:
    print(f"already populated: cohort_metrics {_n_cohorts} rows")

# Rebuilt every run, where cohorts above are rebuilt only when they look wrong. The universe
# recovery cannot answer this table's question. It reads what the rows point at, and a pair
# written under an earlier pin points at a backtest that is still inside the current one: when
# `family == "ensemble"` left this case study's pin on 2026-09-14 the rank-1 rung moved from the
# ensemble at +0.566 to `deep_learning/nlinear` at +2.300, and the ensemble row it had already
# written stayed canonical, stayed cost-feasible and stayed `fwd_ret_15m`. Nothing about it
# reads as stale, because being outside the pin and no longer being what the pin selects are
# different things, and only the first is recoverable from the row. A row count answers neither.
# So it is recomputed rather than validated. `replace_all=True` makes the call a complete
# snapshot, and it costs about a second against this notebook's 25 to 40.
# `NO_CARRIER` is the raw-Sharpe ranking inside the producer rather than a resolved
# lineage: this case study pins a rung the canonical resolver does not know about.
_pairs = populate_paired_metrics(
    CASE_STUDY,
    explorer,
    rung=_RUNG,
    replace_all=True,
    carrier=NO_CARRIER,
    prediction_hashes=ENTIRE_REGISTRY,
)
_n_pairs = sum(1 for row in _pairs if "skip" not in row)
print(f"populated backtest_paired_metrics: {_n_pairs} pairs")

# A rebuild that could not produce a canonical table leaves the previous one in place, which
# is the right call for the data - a stale table is recoverable and deleted rows are not - but
# the wrong one for this notebook, which would go on to read those rows and present selection
# statistics computed under a universe the chapter does not claim. Refuse instead of rendering
# them. An unpopulated registry has nothing off-universe to find and never reaches this.
_still_stale = derived_tables_off_canonical_universe(CASE_DIR, _UNIVERSE_FILTER)
if _still_stale:
    raise RuntimeError(
        f"{', '.join(sorted(_still_stale))} still reference runs outside "
        f"{_UNIVERSE_FILTER} after rebuilding. The rebuild produced nothing to replace them "
        f"with, so every figure below would describe a selection this case study does not "
        f"make. Re-run the backtests for {_UNIVERSE_FILTER} before this notebook."
    )


def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
    if point is None:
        return "—"
    p = format(point, fmt)
    if lo is None or hi is None:
        return f"{p} [—, —]"
    return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"


def _fmt(val: float | None, fmt: str = ".4f") -> str:
    return "—" if val is None else format(val, fmt)
```

## §1 Handoff from model analysis

The strategy phase inherits one configuration, chosen once on validation results
only, across the stages that are eligible for the holdout - signal, allocation
and risk overlay.

The selection is made by `resolve_solvent_carrier`, the shared resolver, and not by
ranking a Sharpe column here. The same resolver answers `18_holdout_predictions` and
`19_holdout_backtest`, so this page reports the configuration those notebooks refitted
and priced. It refuses rather than selecting past a problem, on three conditions that
are all permanent: a rank-1 with no recorded `max_drawdown`, one whose drawdown reached
-100% so its Sharpe is arithmetic on an account that no longer exists, and one fitted
under a conformal calibration version `run_backtest` will no longer execute.

**On this case study the eligible stages come to one.** `UNIVERSE_RESTRICTIONS` pins
rank-1 to the `cost_feasible` universe that `backtest.sweep.universe_filter` declares
canonical, and `15_portfolio_management` and `16_risk_management` both register
specs carrying no `universe_filter`. So the allocation and risk-overlay stages
contribute no eligible rows here and the carrier is a baseline-stage configuration.

The label is taken from the selection rather than assumed. The pool spans every
declared label, so the chosen strategy may rest on a variant label rather than the
primary one, and every loader downstream follows the selection rather than the default.

**Here it does.** `config/setup.yaml` declares `fwd_ret_15m` primary, and the
selection lands on `fwd_dir_15m`, the direction variant, because that is where the
best validation Sharpe is. The consequence is worth stating rather than leaving a
reader to infer it from two labels appearing in different tables: the holdout was
spent on the variant. The primary label is where the stage sweeps ran, carrying all
60 allocation and all 20 risk-overlay backtests against the variant's none, while
the variant carries the single holdout row. So the allocation and overlay evidence
in the sections below describes one label and the out-of-sample evidence in §6
describes another. Each is read on the configuration that produced it and neither
stands in for the other.

```python
carrier = resolve_solvent_carrier(CASE_STUDY)
TOP_HASH = carrier["val_backtest_hash"]
TOP_PHASH = carrier["val_prediction_hash"]
RANK1_FAMILY = carrier["family"]
RANK1_CONFIG = carrier["config_name"]
RANK1_LABEL = carrier["label"]
PRIMARY_LABEL = RANK1_LABEL

_db = CASE_DIR / "run_log" / "registry.db"
with sqlite3.connect(str(_db)) as _con:
    _row = _con.execute(
        "SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
        "ic_hac_lag, ic_pct_positive "
        "FROM prediction_metrics WHERE prediction_hash = ?",
        (TOP_PHASH,),
    ).fetchone()

ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row

print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={PRIMARY_LABEL}")
print(f"        prediction_hash={TOP_PHASH}, backtest_hash={TOP_HASH}")
print()
print("Daily-pooled IC (validation):")
print(f"  IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')}  (HAC, lag={int(ic_lag)})")
print(f"  t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3f}")
print(f"  n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f"  CI status: {ci_status(ic_lo, ic_hi)}")
```

The cells above report the selected prediction set's own score and interval
from `prediction_metrics`. Read the interval rather than the point: it says
whether the ordering the strategy trades on is distinguishable from no
ordering at all, which is the precondition for anything downstream.

The probabilistic Sharpe ratio reported later asks a related question about
the strategy rather than the signal - whether its Sharpe is distinguishable
from zero once the shape of its return distribution is accounted for. A
strategy with occasional large gains and many small losses can show a
flattering Sharpe that this correction discounts.

This case study declares no strategy-specific stopping conditions in
`setup.yaml`, so §9 evaluates two conditions that apply to any strategy:
whether the validation interval's lower bound reaches zero, and whether the
holdout comparison against the benchmark excludes zero on the negative side.
Each is reported as pass, partial or fail.

## §2 Search context, family comparison, and lineage waterfall

The baseline stage produced many validation backtests across the model families
and label horizons. The selected row is one outcome of that search, and a
search over many configurations produces a highest value even when none of the
configurations has an edge.

Two things put it in context. Where it sits within its own family's
distribution says whether it stands apart from its neighbours or is the top of
a tightly packed band. How it changes through the pipeline stages says which
stage moved it, and by how much relative to the uncertainty in that move.

```python
ctx = explorer.search_context("signal")
search_table = pl.DataFrame(
    [
        {"metric": "Total signal backtests", "value": f"{ctx['total']:,}"},
        {"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
        {"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
        {"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
        {"metric": "% positive Sharpe", "value": f"{ctx['pct_positive']:.1f}%"},
        {"metric": "Top-by-Sharpe in this sweep", "value": f"{ctx['champion_sharpe']:.3f}"},
        {"metric": "Top-by-Sharpe percentile", "value": f"{ctx['champion_percentile']:.1f}%"},
    ]
)
print("Baseline-stage search context:")
print(search_table)
```

The family comparison reads each backtest's stored interval rather than
recomputing one, so the plot can show error bars instead of points alone.

It is scoped to the canonical universe, because the baseline stage holds two and they are
not distributed evenly across families. `backtest.sweep.signal_passes` re-runs only
`mechanism_top_n` predictions on the full universe, and those are whichever families
pass 1 ranked highest, so pooling both would move one family's median by rows the others
have no counterpart for. Labels are pooled deliberately: the question here is what a
family's spread of outcomes looks like, and this case study trains most families on
every label.

```python
with sqlite3.connect(str(_db)) as _con:
    _famdf = pl.DataFrame(
        _con.execute(
            """
            SELECT
                t.family,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi
            FROM backtest_metrics bm
            JOIN backtest_runs b      ON bm.backtest_hash = b.backtest_hash
            JOIN prediction_sets p    ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t      ON p.training_hash   = t.training_hash
            WHERE b.stage = 'signal'
              AND p.split = 'validation'
              AND bm.sharpe IS NOT NULL
              AND (bm.num_trades IS NULL OR bm.num_trades > 0)
              AND COALESCE(
                  json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full'
              ) = ?
            """,
            (_UNIVERSE_FILTER or "full",),
        ).fetchall(),
        schema=["family", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"],
        orient="row",
    )
if _famdf.is_empty():
    msg = (
        f"No baseline-stage validation backtests on the {_UNIVERSE_FILTER or 'full'} universe "
        f"for {CASE_STUDY}, so there is no family distribution to plot. The rest of this "
        "notebook reads the same universe, so this is the first place a registry missing it "
        "shows up."
    )
    raise RuntimeError(msg)

family_summary = (
    _famdf.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_q25=pl.col("sharpe").quantile(0.25),
        sharpe_q75=pl.col("sharpe").quantile(0.75),
        sharpe_max=pl.col("sharpe").max(),
        pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
    )
    .sort("sharpe_median", descending=True)
)
print("Family-level baseline Sharpe summary:")
print(family_summary)
```

```python
fig, ax = plt.subplots(figsize=(9, 4))
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()

ax.errorbar(
    medians,
    y,
    xerr=[medians - q25, q75 - medians],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    label="median ±IQR",
)
ax.scatter(maxima, y, marker="x", color="#C62828", s=60, label="max", zorder=5)
ax.axvline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
ax.set_title("Family-level Signal Sharpe — IQR + max")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
show_with_alt(
    fig,
    "Horizontal chart of validation Sharpe by model family. One row per family, named on "
    "the vertical axis which runs top to bottom. A filled circle marks that family's "
    "median Sharpe with a horizontal bar spanning its interquartile range, and a red "
    "cross marks its single best configuration. A dashed vertical line sits at zero "
    "Sharpe. The gap between a family's median and its cross is the spread the best draw "
    "was taken from.",
)
```

**How to read the family distribution.** The bars show each family's spread of
validation outcomes, not one number per family. That is deliberate: the sweep
ran many configurations per family, so a family's best result is the maximum of
a sample, and maxima grow with sample size whether or not the family has an
edge.

Compare the medians to compare families, and read the width of each band as
what separates a configuration from its neighbours. Where the bands overlap
heavily, the families are producing the same range of outcomes and the ordering
of their maxima is not evidence that one is better.

The cost model charged here is the engine's, applied identically to every
configuration, so differences between families are not differences in what they
were charged. Every row is on one universe for the same reason.

The lineage gives one backtest per pipeline stage for this prediction. Each
stage's stored interval is pulled alongside it so the change from one stage to
the next can be read against the uncertainty in each.

```python
lineage = explorer.champion_lineage(TOP_PHASH)
ci_lo: dict[str, float] = {}
ci_hi: dict[str, float] = {}
for stage_name, info in lineage.items():
    bm = load_backtest_metrics(CASE_STUDY, backtest_hash=info["backtest_hash"])
    if not bm.is_empty():
        row = bm.row(0, named=True)
        ci_lo[stage_name] = row.get("sharpe_ci95_lo")
        ci_hi[stage_name] = row.get("sharpe_ci95_hi")

print("Lineage stages present for selected prediction:")
for s, info in lineage.items():
    lo, hi = ci_lo.get(s), ci_hi.get(s)
    print(f"  {s}: hash={info['backtest_hash']}, Sharpe={_fmt_ci(info['sharpe'], lo, hi)}")

missing_stages = [s for s in ("cost_sensitivity", "risk_overlay") if s not in lineage]
if missing_stages:
    print()
    print(f"Stage transitions not run for this prediction: {missing_stages}")
```

```python
fig = plot_sharpe_waterfall(lineage, ci_lo=ci_lo, ci_hi=ci_hi)
show_with_alt(
    fig,
    "Waterfall of Sharpe across the selected configuration's locked lineage: signal, then "
    "allocation, then cost, then risk overlay, one bar per stage in that order. Each bar carries "
    "asymmetric error bars for its block-bootstrap 95 percent confidence interval. Those are "
    "marginal intervals on each stage and not a test between stages: whether one stage differs "
    "from the next is the paired difference printed below the figure, which can exclude zero "
    "while two bars overlap.",
)
```

Stage transitions are read from the stored paired comparisons rather than
recomputed here. A stage registered against a different prediction set has no
transition for this lineage, and is reported as not run rather than
reconstructed from unpaired numbers.

```python
ALLOC_HASH = lineage.get("allocation", {}).get("backtest_hash")
if ALLOC_HASH is not None:
    pair = load_paired_metrics(
        CASE_STUDY,
        challenger_hash=ALLOC_HASH,
        benchmark_kind="signal_leader",
    )
    if not pair.is_empty():
        r = pair.row(0, named=True)
        print("Stage transition: allocation challenger vs signal leader")
        print(
            f"  sharpe_diff = "
            f"{_fmt_ci(r['sharpe_diff'], r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}"
        )
        print(f"  p_value = {r['p_value']:.3f}")
        print(f"  prob_challenger_wins = {r['prob_challenger_wins']:.3f}")
        print(f"  CI status: {ci_status(r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}")
```

Live values for the stage transitions printed above are read from
`backtest_paired_metrics`. On this CS, every stage delta is small in
magnitude — the position-level overlays do not produce a stage-delta
CI that excludes zero on the positive side. The signal economics
remain the binding constraint; allocator choice and overlay choice
modulate the magnitude of loss rather than producing a credibility-
resolved positive transition.

```python
# top_k is an entry-scheme parameter, so the sweep that varies it registers at
# stage=signal on this case study: the carrier's rows are 5, 10 and 20 positions
# there and none at the allocation stage, which `concentration_curve` defaults to.
# Asking for the default raised rather than returning nothing, which is the point
# of that refusal: this cell printed "no concentration data" and dropped its
# figure for as long as the registry held no rows for the carrier at all.
conc_df = explorer.concentration_curve(TOP_PHASH, stage="signal")
if not conc_df.is_empty():
    fig = plot_concentration_curve(conc_df)
    show_with_alt(
        fig,
        "Line chart of Sharpe against portfolio concentration. The horizontal axis is the "
        "number of positions held, top_k, and the vertical axis is the Sharpe the selected "
        "prediction achieves at each. The shape rather than any single point is the content: "
        "it shows whether concentrating the book helped or hurt.",
    )
    best_per_k = conc_df.sort("sharpe", descending=True).group_by("top_k").first().sort("top_k")
    print("Concentration sweep: best Sharpe by top_k:")
    print(best_per_k.select("top_k", "allocator", "sharpe", "max_drawdown"))
else:
    print("No concentration data: no top_k sweep registered for this prediction.")
```

**How to read the concentration curve.** Position count trades two things off
against each other. Holding few names concentrates capital into the ordering's
strongest views, so a correct ordering is rewarded more and an incorrect one
punished more, and each name turns over more often as the ranking shifts.
Holding many names dilutes both, and as the count approaches the size of the
universe the portfolio converges on the universe itself, leaving nothing for
the ordering to express.

The useful reading is the shape rather than the maximum. A curve with a clear
interior peak says breadth is a real decision for this strategy. A flat curve
says the ordering is not strong enough for concentration to matter, and the
apparent best count is picking out noise.

## §3 Headline performance with uncertainty

The selected specification is the validation-window backtest associated
with the highest baseline Sharpe. Every metric is reported with
its block-bootstrap interval from `backtest_metrics`; the equity overlay
shows the cumulative trajectory against the equal-weight NASDAQ-100
universe benchmark.

```python
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)
full = dict(full)  # mutate-friendly

_db = CASE_DIR / "run_log" / "registry.db"
_cohort_leader_hash = None
_cohort_leader_sharpe = None
_cohort_stage = None
with sqlite3.connect(str(_db)) as _con:
    _stage_for_rank1 = _con.execute(
        "SELECT stage FROM backtest_runs WHERE backtest_hash = ?",
        (TOP_HASH,),
    ).fetchone()[0]
for _try_stage in dict.fromkeys((_stage_for_rank1, "allocation", "signal")):
    with sqlite3.connect(str(_db)) as _con:
        _cohort = _con.execute(
            """
            SELECT dsr_er, dsr_er_pvalue, expected_max_sharpe_er,
                   min_trl_periods_er, dsr_mp, dsr_mp_pvalue,
                   expected_max_sharpe_mp, min_trl_periods_mp,
                   dsr_raw, dsr_raw_pvalue,
                   expected_max_sharpe_raw, min_trl_periods_raw,
                   n_trials_effective_mp, n_trials_effective_er,
                   pbo, pbo_n_combinations, pbo_n_folds,
                   k_variants, leader_hash, leader_sharpe
            FROM cohort_metrics
            WHERE cohort_type='family' AND stage=? AND label=? AND family=?
              AND leader_sharpe IS NOT NULL
            """,
            (_try_stage, RANK1_LABEL, RANK1_FAMILY),
        ).fetchone()
    if _cohort is not None:
        (
            full["dsr_er"],
            full["dsr_er_pvalue"],
            full["expected_max_sharpe_er"],
            full["min_trl_periods_er"],
            full["dsr_mp"],
            full["dsr_mp_pvalue"],
            full["expected_max_sharpe_mp"],
            full["min_trl_periods_mp"],
            full["dsr_raw"],
            full["dsr_raw_pvalue"],
            full["expected_max_sharpe_raw"],
            full["min_trl_periods_raw"],
            full["n_trials_effective_mp"],
            full["n_trials_effective_er"],
            full["pbo"],
            full["pbo_n_combinations"],
            full["pbo_n_folds"],
            full["k_variants"],
            _cohort_leader_hash,
            _cohort_leader_sharpe,
        ) = _cohort
        _cohort_stage = _try_stage
        # Mirror DSR_ER onto the legacy `dsr` / `dsr_pvalue` / `expected_max_sharpe`
        # fields so downstream consumers that still read those keys see the
        # ER-default values.
        full["dsr"] = full["dsr_er"]
        full["dsr_pvalue"] = full["dsr_er_pvalue"]
        full["expected_max_sharpe"] = full["expected_max_sharpe_er"]
        full["min_trl_periods"] = full["min_trl_periods_er"]
        break

spec_block = {
    "case_study": CASE_STUDY,
    "family": RANK1_FAMILY,
    "config_name": RANK1_CONFIG,
    "label": RANK1_LABEL,
    "signal_method": lineage["signal"].get("signal_method"),
    "top_k": lineage["signal"].get("top_k"),
    "allocation": lineage.get("allocation", {}).get("allocator"),
    "risk_overlay": lineage.get("risk_overlay", {}).get("risk_name"),
    "rebalance_step_bars": setup["labels"]["rebalance_step"][RANK1_LABEL],
    "cost_assumption": "engine cost model: 5 bps commission + 2 bps slippage at the baseline stage; per_share_plus_spread sensitivity in §5",
    "validation_window_periods": int(full["n_periods"]) if full["n_periods"] is not None else None,
    "num_trades": int(full["num_trades"]) if full["num_trades"] is not None else None,
    "avg_turnover": full["avg_turnover"],
    "bootstrap_block_length": int(full["bootstrap_block_length"])
    if full["bootstrap_block_length"] is not None
    else None,
    "bootstrap_n": int(full["bootstrap_n"]) if full["bootstrap_n"] is not None else None,
}
print("Rank-1 specification (cross-stage selection, validation window):")
for k, v in spec_block.items():
    print(f"  {k}: {v}")

_block = int(full["bootstrap_block_length"])
_rstep = setup["labels"]["rebalance_step"][RANK1_LABEL]
print(
    f"  audit: bootstrap_block_length={_block} day(s); "
    f"rebalance_step={_rstep} bar(s) ({_rstep * 15} minutes)"
)

# Cohort cascade audit: surface where DSR comes from when the selected lineage
# stage's cohort row is missing.
if _cohort_stage and _cohort_stage != _stage_for_rank1:
    print(
        f"  cohort cascade: selected stage `{_stage_for_rank1}` has no "
        f"cohort_metrics row (cross-stage timestamp-alignment limit); "
        f"DSR/EMS/MinTRL/PBO sourced from family cohort `{_cohort_stage}/"
        f"{RANK1_LABEL}/{RANK1_FAMILY}` (k={int(full['k_variants'])}, "
        f"leader_hash={_cohort_leader_hash}, leader_sharpe="
        f"{_cohort_leader_sharpe:+.3f}; spine sibling differs by "
        f"{full['sharpe'] - _cohort_leader_sharpe:+.3f})."
    )
elif _cohort_stage:
    print(
        f"  cohort: family cohort `{_cohort_stage}/{RANK1_LABEL}/"
        f"{RANK1_FAMILY}` (k={int(full['k_variants'])}, leader_hash="
        f"{_cohort_leader_hash}); spine sibling differs by "
        f"{full['sharpe'] - _cohort_leader_sharpe:+.3f}."
    )
```

```python
def _row(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
    return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}


sharpe_status = ci_status(full["sharpe_ci95_lo"], full["sharpe_ci95_hi"])
sortino_status = ci_status(full["sortino_ci95_lo"], full["sortino_ci95_hi"])
ann_status = ci_status(full["ann_return_ci95_lo"], full["ann_return_ci95_hi"])
mdd_status = ci_status(full["max_dd_ci95_lo"], full["max_dd_ci95_hi"])
calmar_status = ci_status(full["calmar_ci95_lo"], full["calmar_ci95_hi"])

headline = pl.DataFrame(
    [
        _row(
            "Sharpe",
            _fmt(full["sharpe"]),
            _fmt(full["sharpe_ci95_lo"]),
            _fmt(full["sharpe_ci95_hi"]),
            sharpe_status,
        ),
        _row(
            "Sortino",
            _fmt(full["sortino"]),
            _fmt(full["sortino_ci95_lo"]),
            _fmt(full["sortino_ci95_hi"]),
            sortino_status,
        ),
        _row(
            "Annualized return",
            _fmt(full["cagr"]),
            _fmt(full["ann_return_ci95_lo"]),
            _fmt(full["ann_return_ci95_hi"]),
            ann_status,
        ),
        _row(
            "Max drawdown",
            _fmt(full["max_drawdown"]),
            _fmt(full["max_dd_ci95_lo"]),
            _fmt(full["max_dd_ci95_hi"]),
            mdd_status,
        ),
        _row(
            "Calmar",
            _fmt(full["calmar"]),
            _fmt(full["calmar_ci95_lo"]),
            _fmt(full["calmar_ci95_hi"]),
            calmar_status,
        ),
        {
            "metric": "PSR p-value (H0: SR≤0)",
            "point": _fmt(full["psr_pvalue"]),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_ER (selection-adjusted)",
            "point": _fmt(full.get("dsr_er")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_ER p-value",
            "point": _fmt(full.get("dsr_er_pvalue")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_MP",
            "point": _fmt(full.get("dsr_mp")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_MP p-value",
            "point": _fmt(full.get("dsr_mp_pvalue")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_RAW (naive raw-K)",
            "point": _fmt(full.get("dsr_raw")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "DSR_RAW p-value",
            "point": _fmt(full.get("dsr_raw_pvalue")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": (
                f"Expected max Sharpe (k={int(full['k_variants'])})"
                if full.get("k_variants") is not None
                else "Expected max Sharpe"
            ),
            "point": _fmt(full.get("expected_max_sharpe_er")),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
        {
            "metric": "PBO (CSCV)",
            "point": _fmt(full.get("pbo"), ".3f"),
            "ci95_lo": "—",
            "ci95_hi": "—",
            "status": "n/a",
        },
    ]
)
print("Rank-1 headline metrics with 95% CIs:")
print(headline)
```

```python
# Forest plot: selected-lineage metrics with CI bars + reference lines
ew_val = load_benchmark_metrics(CASE_STUDY, PRIMARY_LABEL, period="validation")
forest_metrics = [
    ("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
    ("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
    ("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
    ("Ann. return", full["cagr"], full["ann_return_ci95_lo"], full["ann_return_ci95_hi"]),
]

fig, ax = plt.subplots(figsize=(8, 4))
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.errorbar(
    points,
    y,
    xerr=[points - los, his - points],
    fmt="o",
    color="#1565C0",
    ecolor="#5B9BD5",
    elinewidth=2.0,
    capsize=4,
    markersize=7,
)
ax.axvline(0, color="#9E9E9E", linestyle="--", linewidth=0.8)
ax.axvline(
    ew_val["sharpe"],
    color="#43A047",
    linestyle=":",
    linewidth=1.0,
    label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
_alloc_sharpe = lineage.get("allocation", {}).get("sharpe")
if _alloc_sharpe is not None:
    ax.axvline(
        _alloc_sharpe,
        color="#E53935",
        linestyle=":",
        linewidth=1.0,
        label=f"Allocation selection Sharpe ({_alloc_sharpe:.2f})",
    )
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Value")
ax.set_title("Rank-1 Headline Metrics with 95% CIs")
ax.legend(loc="lower right", fontsize=8, frameon=False)
show_with_alt(
    fig,
    "Forest plot of the selected configuration's headline metrics. One row per metric, named on "
    "the vertical axis, with a point estimate and a horizontal bar for its 95 percent confidence "
    "interval, read on a shared horizontal value axis. A dashed vertical line marks zero, and "
    "further vertical reference lines mark the equal-weight validation Sharpe and the "
    "allocation-stage Sharpe. Those two lines are Sharpe values, so only the Sharpe row is "
    "comparable with them; the Sortino, Calmar and annualized-return rows share the axis but "
    "measure different quantities.",
)
```

The strategy's returns are stored daily while the benchmark is stored at the
decision cadence, so the benchmark is aggregated to daily before the two are
aligned. Comparing series at different frequencies would otherwise pair a day
against a fraction of one.

```python
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
    pl.read_parquet(strat_returns_path)
    .sort("timestamp")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)

bench_val = (
    load_benchmark_returns(CASE_STUDY, PRIMARY_LABEL, period="validation")
    .with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
    .group_by("ts")
    .agg(((1 + pl.col("ew_return")).product() - 1).alias("benchmark"))
    .sort("ts")
)

aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
    f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)

cum_strat = np.cumprod(1 + aligned["strategy"].to_numpy()) - 1
cum_bench = np.cumprod(1 + aligned["benchmark"].to_numpy()) - 1
fig, ax = plt.subplots(figsize=(10, 4.2))
ax.plot(aligned["ts"], cum_strat, color="#1565C0", linewidth=1.2, label="Selected strategy")
ax.plot(aligned["ts"], cum_bench, color="#43A047", linewidth=1.2, label="EW universe")
ax.axhline(0, color="#9E9E9E", linewidth=0.6, linestyle="--")
ax.set_ylabel("Cumulative return")
ax.set_title("Validation-window cumulative return: selected strategy vs EW universe")
ax.legend(loc="best", frameon=False)
show_with_alt(
    fig,
    "Two cumulative return paths over the validation window, plotted against time. One "
    "line is the selected strategy and the other the equal-weight universe, distinguished "
    "in the legend; the vertical axis is cumulative return and a dashed horizontal line "
    "marks zero. The comparison is the point: the strategy's path is only informative "
    "beside the universe it traded in.",
)
```

**How to read the selection-bias adjustment.** The deflated Sharpe ratio asks
what the highest result from a search of this size would look like if none of
the configurations had an edge, and discounts the observed result by that
amount. The more configurations searched, the higher the bar it sets.

Its inputs are the number of variants in the cohort searched, the shape of the
return distribution, and the length of the record. All three are printed above
rather than assumed, because the count of variants in particular decides how
large the adjustment is, and a count that quietly includes duplicates of an
already-chosen configuration inflates the correction.

The companion quantity is the track record length required for a result of this
size to become distinguishable from zero. When it exceeds the history
available, the honest statement is that the window cannot settle the question,
whatever the point estimate.

## §4 Risk and drawdown analysis

Risk metrics use the validation-window strategy returns paired against the validation
EW benchmark. The drawdown panel surfaces the worst episode and whether it recovered;
rolling Sharpe and rolling beta say whether either was steady or driven by part of the
window.

The lower panel is worth reading for its shape and not only its depth: whether the
losses arrived steadily across the window or inside one episode, and whether anything
was recovered afterwards. That is a description of when the strategy lost, and it is as
far as the panel goes. It does not identify a cause - a strategy whose costs exceed its
edge still has winning trades and partial recoveries, and a steady decline can happen
with no cost at all - so §5 varies the cost directly rather than inferring it from this
curve.

```python
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()

pa = PortfolioAnalysis(
    returns=strat_arr,
    benchmark=bench_arr,
    dates=ts_arr,
    periods_per_year=PERIODS_PER_YEAR,
)

dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)
```

```python
fig = plot_equity_drawdown(strat_returns_path)
show_with_alt(
    fig,
    "Two stacked panels sharing a time axis. The upper panel is the selected strategy's "
    "cumulative return over the window and the lower panel is its drawdown, the distance "
    "below the running maximum at each moment. Reading a date down both panels pairs a "
    "level of return with the loss being carried to hold it.",
)
```

```python
# Rolling Sharpe + rolling beta. `aligned` is daily - the strategy's returns are stored
# that way and the benchmark is compounded to match - so the window is 126 trading days,
# about six months of a validation window that runs roughly 253.
ROLL_WINDOW = 126
roll = pa.compute_rolling_metrics(windows=[ROLL_WINDOW], metrics=["sharpe", "beta"])

# Summarised rather than printed as an object. This used to print the types of the
# result's members, which says nothing about the strategy, and the section's own prose
# promises these two series - a computed result that reaches no reader is the same as one
# never computed.
_roll_rows = []
for _label, _store in (("Rolling Sharpe", roll.sharpe), ("Rolling beta", roll.beta)):
    _series = _store.get(ROLL_WINDOW)
    _valid = _series.drop_nulls() if _series is not None else None
    if _valid is None or _valid.len() == 0:
        _roll_rows.append({"series": _label, "n": 0, "min": "n/a", "median": "n/a", "max": "n/a"})
        continue
    _roll_rows.append(
        {
            "series": _label,
            "n": _valid.len(),
            "min": f"{_valid.min():+.3f}",
            "median": f"{_valid.median():+.3f}",
            "max": f"{_valid.max():+.3f}",
        }
    )
print(f"Rolling metrics over {ROLL_WINDOW} trading days ({aligned.height} days in the window):")
print(pl.DataFrame(_roll_rows))

# Tail-risk read straight from registry (already computed and stored)
tail_table = pl.DataFrame(
    [
        {"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
        {"metric": "VaR 95% (daily)", "value": f"{full['var_95']:.4f}"},
        {"metric": "CVaR 95% (daily)", "value": f"{full['cvar_95']:.4f}"},
        {"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
        {"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
        {"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
    ]
)
print()
print("Tail risk profile:")
print(tail_table)
```

```python
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
print(f"Per-fold breakdown ({fold_df.height} folds):")
print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
print()
print(f"Fold Sharpe range: [{fold_df['sharpe'].min():.3f}, {fold_df['sharpe'].max():.3f}]")
print(f"Fold Sharpe std:   {fold_df['sharpe'].std():.3f}")
```

**How to read the four printouts above.** Each answers a question the others cannot,
and none of their values is transcribed here: this notebook re-resolves its carrier from
the registry on every run, so a number written into this cell would describe whichever
configuration was rank-1 the day it was written.

*Per-fold Sharpes* say whether the headline belongs to the strategy or to one block.
Two folds agreeing in sign is weak evidence; two disagreeing is strong evidence against.
With `evaluation.n_splits: 2` there is no third block to break the tie, so a split
verdict is a reason to distrust the headline rather than to average it.

*Skewness and kurtosis* say what the distribution behind that Sharpe looks like.
Negative skew with high kurtosis is the shape a Sharpe flatters, because Sharpe reads
only the first two moments: many small gains against occasional large losses. Positive
skew is the opposite trade, and it is what a strategy paying a per-trade cost for
occasional large wins looks like.

*Tail ratio* is the 95th percentile over the absolute 5th, so it compares where the two
tails begin and not how much sits beyond those points. Near 1 the two boundaries are the
same size. It answers a different question from skewness, which the extremes past those
percentiles drive, so the two can point opposite ways on the same series - rare large
losses pull the skew negative while leaving the 5th percentile where it was. Disagreement
between them is a fact about the distribution rather than a sign that one is wrong.

*The rolling window* locates the result in time: each point is the Sharpe of the 126
trading days ending there, so the summary reports how far the estimate moved over the
window rather than a single number for the whole period. Read the spread between the
minimum and the maximum, not the sign. Consecutive windows share 125 of their 126 days,
so the series is autocorrelated by construction, and a validation window of roughly 253
days holds about two non-overlapping windows. Sampling variation alone produces sign
changes on a stationary series at this length, so a crossing is not evidence of a
regime, and an estimate that stays positive throughout does not establish that
performance was stable. Rolling beta is read the same way and carries the same
qualification: each point is an estimate over 126 overlapping days, and a book whose
population beta is zero still returns nonzero sample betas.

Dollar neutrality and benchmark beta are two different properties and only one of them is
estimated. Dollar neutrality is a fact about the weights - whether the long and short
notionals sum to zero - and is read off the allocation directly. Beta is a fact about how
those positions co-move with the benchmark, so a book can be exactly dollar-neutral and
still carry beta whenever its longs and shorts differ in benchmark sensitivity. Rolling
beta is the estimate of that second quantity, and a wandering series is a reason to
measure the exposure with its uncertainty rather than a reading of it.

## §5 Friction budget & cost sensitivity

Cost sensitivity is where a short-rebalancing strategy is decided, because
every cost is charged per trade and this one trades often. Two grids are read
here: a **cost grid** that walks commission and slippage from zero upward, and
a **risk-overlay grid** of rules that close a position on a condition other
than the signal.

They answer different questions. The cost grid asks how much friction the
strategy can absorb before its edge runs out, which states the execution
quality it requires rather than the profit it made under one assumption. The
overlay grid asks whether trading *less* on positions that are going wrong
preserves more than the extra trading costs. The two were swept independently,
with overlays applied at zero engine cost, so each result is attributable to
the field its sweep varied.

**Neither grid is restricted to the selected lineage, and they do not sit on
the same universe.** Both read their whole stage across labels, which is what
the per-horizon stratification below is for. The cost rows come from
`17_costs`, whose pool is pinned to the canonical `cost_feasible` universe; the
overlay rows come from `16_risk_management`, which overlays allocation-stage
specs that carry no `universe_filter` and are therefore full-universe. So a
cost curve and an overlay row in this section are not two readings of one
portfolio, and the gradient across thresholds is what §5.2 supports rather than
a level comparable with §5.1.

### §5.1 Cost sensitivity stratified by label horizon

```python
with sqlite3.connect(str(_db)) as _con:
    cost_df = pl.DataFrame(
        _con.execute(
            """
            SELECT
                b.spec_json,
                t.family,
                t.config_name,
                t.label,
                bm.sharpe,
                bm.sharpe_ci95_lo,
                bm.sharpe_ci95_hi,
                bm.max_drawdown,
                bm.cagr,
                bm.num_trades,
                bm.avg_turnover
            FROM backtest_runs b
            JOIN backtest_metrics bm ON bm.backtest_hash  = b.backtest_hash
            JOIN prediction_sets p   ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t     ON p.training_hash   = t.training_hash
            WHERE b.stage = 'cost_sensitivity'
              AND p.split = 'validation'
              AND bm.sharpe IS NOT NULL
              AND (bm.num_trades IS NULL OR bm.num_trades > 0)
            """
        ).fetchall(),
        schema=[
            "spec_json",
            "family",
            "config_name",
            "label",
            "sharpe",
            "sharpe_ci95_lo",
            "sharpe_ci95_hi",
            "max_drawdown",
            "cagr",
            "num_trades",
            "avg_turnover",
        ],
        orient="row",
    )


def _cost_bps(spec_str: str) -> float:
    """Cost per leg in bps from the locked spec.

    `backtest_config.commission.rate` and `backtest_config.slippage.rate`
    are decimal fractions; sum × 10000 gives total per-leg cost in bps.
    """
    spec = json.loads(spec_str)
    bc = spec.get("backtest_config", {})
    comm = bc.get("commission", {}) or {}
    slip = bc.get("slippage", {}) or {}
    return float((comm.get("rate", 0) + slip.get("rate", 0)) * 10000)


cost_df = cost_df.with_columns(
    pl.col("spec_json").map_elements(_cost_bps, return_dtype=pl.Float64).alias("cost_bps")
)

# Aggregate per (label, cost_bps): max-config Sharpe + envelope
cost_curve_per_label = (
    cost_df.group_by(["label", "cost_bps"])
    .agg(
        n=pl.len(),
        sharpe_max=pl.col("sharpe").max(),
        sharpe_min=pl.col("sharpe").min(),
        sharpe_median=pl.col("sharpe").median(),
        # CI lower bound *of the best-Sharpe row* per (label, cost):
        # this is the credibility band on the headline cell.
        best_ci_lo=pl.col("sharpe_ci95_lo").top_k_by("sharpe", 1).first(),
        best_ci_hi=pl.col("sharpe_ci95_hi").top_k_by("sharpe", 1).first(),
        best_n_trades=pl.col("num_trades").top_k_by("sharpe", 1).first(),
        best_avg_turnover=pl.col("avg_turnover").top_k_by("sharpe", 1).first(),
        best_family=pl.col("family").top_k_by("sharpe", 1).first(),
        best_config=pl.col("config_name").top_k_by("sharpe", 1).first(),
    )
    .sort(["label", "cost_bps"])
)
print("Per-label cost grid (validation, best Sharpe per (label, cost) cell):")
with pl.Config(tbl_rows=33):
    print(
        cost_curve_per_label.select(
            [
                "label",
                "cost_bps",
                "sharpe_max",
                "best_ci_lo",
                "best_ci_hi",
                "sharpe_median",
                "best_n_trades",
                "best_family",
                "best_config",
            ]
        )
    )
```

```python
# Per-label cost curve: one line per label horizon
fig, ax = plt.subplots(figsize=(9.5, 4.5))
label_order = ["fwd_ret_5m", "fwd_ret_15m", "fwd_ret_60m"]
label_colors = {
    "fwd_ret_5m": "#C62828",
    "fwd_ret_15m": "#FB8C00",
    "fwd_ret_60m": "#1565C0",
}
for lbl in label_order:
    sub = cost_curve_per_label.filter(pl.col("label") == lbl).sort("cost_bps")
    if sub.is_empty():
        continue
    xs = sub["cost_bps"].to_numpy()
    ax.plot(
        xs,
        sub["sharpe_max"].to_numpy(),
        color=label_colors[lbl],
        linewidth=1.6,
        marker="o",
        markersize=4,
        label=f"{lbl} (best config)",
    )
    ax.fill_between(
        xs,
        sub["best_ci_lo"].to_numpy(),
        sub["best_ci_hi"].to_numpy(),
        color=label_colors[lbl],
        alpha=0.10,
    )

ax.axhline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.axhline(
    ew_val["sharpe"],
    color="#43A047",
    linewidth=0.9,
    linestyle=":",
    label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
ax.axvspan(1.0, 3.0, color="#43A047", alpha=0.08, label="large-cap spread (1–3 bps)")
ax.axvspan(3.0, 8.0, color="#FB8C00", alpha=0.08, label="mid-cap spread (3–8 bps)")
ax.axvline(
    5.0, color="#C62828", linewidth=0.8, linestyle=":", label="protocol friction floor (5 bps)"
)
ax.set_xlabel("Per-leg cost (bps)")
ax.set_ylabel("Sharpe (validation, best config per cost level)")
ax.set_title("Cost sensitivity by label horizon — best config per (label, cost)")
ax.legend(loc="best", fontsize=8, frameon=False)
show_with_alt(
    fig,
    "Line chart of validation Sharpe against per-leg cost in basis points, one line per label "
    "horizon, each showing the best configuration at every cost level with a shaded band around "
    "it. A dashed horizontal line marks zero Sharpe. Two shaded vertical spans mark realistic "
    "spreads, roughly 1 to 3 basis points for large caps and 3 to 8 for mid caps, and a dotted "
    "vertical line marks the protocol's 5 basis point friction floor. That floor is a reference "
    "and not what this case study charges: config/setup.yaml bills a per-share commission plus "
    "symbol-specific half-spreads, not a flat rate.",
)
```

Trade counts at zero cost say how exposed each label is to friction before any
cost is charged. A configuration that trades twice as often loses twice as much
to the same per-trade cost, so this count predicts the slope of each label's
cost curve rather than its level.

```python
turnover_table = (
    cost_df.filter(pl.col("cost_bps") == 0)
    .group_by("label")
    .agg(
        n_configs=pl.len(),
        max_sharpe=pl.col("sharpe").max(),
        median_sharpe=pl.col("sharpe").median(),
        max_num_trades=pl.col("num_trades").max(),
        min_num_trades=pl.col("num_trades").min(),
        median_avg_turnover=pl.col("avg_turnover").median(),
    )
    .sort("median_avg_turnover", descending=True)
)
print("Trade count and turnover at zero cost by label horizon:")
print(turnover_table)
```

Two breakeven costs are reported per label. The first is the highest cost at
which the point estimate stays positive. The second is the highest cost at
which the interval's lower bound stays positive, which is the stricter and more
useful of the two: 

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.