Saltar al contenido
Todos los documentos de la biblioteca

Agregación y comparación de backtests en estudios de trading

Código Machine Learning for Trading

Resumen

Este cuaderno consulta registros de backtests en nueve estudios de caso y prepara tablas comparables para análisis posteriores. Organiza los resultados por clase de activo y frecuencia de los datos, y registra el rendimiento en las etapas de señal, asignación, costes y riesgo, junto con información sobre familias de modelos y resultados de validación frente al periodo de reserva. El proceso de selección compara especificaciones completas de estrategia —la señal, el método de asignación y el componente de riesgo— y sigue la configuración seleccionada hasta una ejecución de reserva con reentrenamiento. Si esta no produce un backtest de reserva utilizable, recurre a otro candidato elegible según la clasificación de validación.

El cuaderno también mide cuán concentrados o estables parecen los mejores resultados de validación. Presenta la diferencia entre las configuraciones mejor clasificadas y las de menor rango, la variación entre segmentos y la frecuencia con que la configuración líder obtiene resultados positivos en los distintos segmentos. Una agregación completa se niega a escribir resultados para todo el capítulo si falta algún registro previsto o está vacío, para evitar que un conjunto parcial de datos parezca una comparación completa. Estos diagnósticos ayudan a matizar la evidencia entre casos, pero los resúmenes heredan los datos, costes, etiquetas y supuestos de backtest de cada estudio; las clasificaciones de validación y los resultados del periodo de reserva no deben interpretarse como rendimiento universal de una estrategia.

Ideas clave

  • Compara backtests entre estudios de caso mediante registros y tablas resumen coherentes por etapa.
  • Define una especificación de estrategia con su señal, método de asignación y componente de riesgo considerados en conjunto.
  • Sigue la especificación de validación seleccionada hasta la evaluación del periodo de reserva y recurre a otra si su ejecución no es utilizable.
  • Usa la amplitud del grupo superior y los resultados por segmento para evaluar cuánto puede depender una selección del ruido de clasificación.
  • Exige todos los registros antes de publicar agregados de alcance completo, para que una cobertura incompleta no parezca completa.

Etiquetas

Texto completo
# 01_aggregate_synthesis.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Aggregate Synthesis
#
# **Docker image**: `ml4t`
#
# This notebook queries all 9 case study registries via `BacktestExplorer`
# and builds cross-dataset comparison DataFrames for the remaining Ch20 notebooks.
#
# **Data source**: `registry.db` per case study (no JSON files needed).
#
# **Learning Objectives**:
# - Query per-case-study backtest registries for signal, allocation, cost, and risk metrics
# - Build cross-dataset comparison tables
# - Export summary DataFrames for downstream notebooks (02–06)
#
# **Book Reference**: Chapter 20, Section 20.1 (First-pass results across nine case studies)
#
# **Prerequisites**: Case studies must have run Ch16–19 backtests.

# %%
"""Ch20 Aggregate Synthesis — query registries and compare all 9 case studies."""

import json
import sqlite3
from functools import cache
from pathlib import Path

import polars as pl
import yaml
from IPython.display import Markdown, display

from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_returns
from case_studies.utils.strategy_analysis import (
    allocation_method_of,
    compute_cost_bps,
    rank_one,
    training_run_fitted_for_the_holdout,
)
from utils.paths import REPO_ROOT, get_case_study_dir, get_chapter_dir

# %% tags=["parameters"]
MAX_SYMBOLS = 0
# When non-empty, restricts the cross-CS iteration to the given subset.
# Used by the per-CS pipeline driver to populate `backtest_paired_metrics`
# for a single CS after its holdout has landed, without re-running the
# full 9-CS aggregation.
CASE_STUDIES: list[str] = []
# Test-only: in an isolated test registry, nasdaq's out-of-band cost-feasible
# selected configuration is absent, so its spine cannot resolve. Production leaves this False
# (a missing selection fails loudly); the test harness sets it True so cost/risk
# for such a case study are reported not-applicable instead of raising.
ALLOW_MISSING_SPINE = False
# A full-mode run overwrites the nine-case-study artifacts every downstream notebook and the
# chapter figures read. Refuse to start one unless all nine registries are present and carry
# backtests, so a partial aggregation cannot be published as a complete one. The test harness
# sets this False because its isolated registry is not the production store; a subset run
# (`CASE_STUDIES` non-empty) is the per-case-study driver path and is never full mode.
REQUIRE_ALL_REGISTRIES = True

# %%
OUTPUT_DIR = get_chapter_dir(20) / "output"
OUTPUT_DIR.mkdir(exist_ok=True)

ALL_CASE_STUDIES = [
    "etfs",
    "crypto_perps_funding",
    "nasdaq100_microstructure",
    "sp500_equity_option_analytics",
    "us_firm_characteristics",
    # FX rank-1 is linear/ridge_a100.0 on fwd_ret_21d (val Sharpe +0.048,
    # holdout +0.194), resolved after the 2026-06-01 DL-lookback fix. The
    # earlier deep_learning/tcn selection (val +0.108 / holdout -1.59) was an
    # artifact of gappy validation folds (lookback=60 warmup consumed each
    # fold's head); those sets were purged and the clean lineage re-resolved.
    # See backtest_audit.md and project_registry_hash_collisions.
    "fx_pairs",
    "cme_futures",
    "sp500_options",
    "us_equities_panel",
]
if CASE_STUDIES:
    ALL_CASE_STUDIES = [cs for cs in ALL_CASE_STUDIES if cs in set(CASE_STUDIES)]

DISPLAY_NAMES = {
    "etfs": "ETFs",
    "crypto_perps_funding": "Crypto",
    "nasdaq100_microstructure": "NASDAQ-100",
    "sp500_equity_option_analytics": "S&P 500 Eq+Opt",
    "us_firm_characteristics": "US Firms",
    "fx_pairs": "FX Pairs",
    "cme_futures": "CME Futures",
    "sp500_options": "S&P 500 Options",
    "us_equities_panel": "US Equities",
}

# %%
ASSET_CLASS_MAP = {
    "etfs": "equity_etf",
    "crypto_perps_funding": "crypto",
    "nasdaq100_microstructure": "equity_micro",
    "sp500_equity_option_analytics": "equity_options",
    "us_firm_characteristics": "equity_firm",
    "fx_pairs": "fx",
    "cme_futures": "futures",
    "sp500_options": "options",
    "us_equities_panel": "equity_panel",
}

FREQ_MAP = {
    "etfs": "daily",
    "crypto_perps_funding": "8h",
    "nasdaq100_microstructure": "15min",
    "sp500_equity_option_analytics": "daily",
    "us_firm_characteristics": "monthly",
    "fx_pairs": "daily",
    "cme_futures": "daily",
    "sp500_options": "daily",
    "us_equities_panel": "daily",
}

# %% [markdown]
# ## Load Registries
#
# Create a `BacktestExplorer` for each case study that has a registry.

# %%
explorers: dict[str, BacktestExplorer] = {}
configs: dict[str, dict] = {}
unreadable_registries: list[str] = []
empty_registries: list[str] = []

for cs in ALL_CASE_STUDIES:
    try:
        explorers[cs] = BacktestExplorer(cs)
        setup_path = get_case_study_dir(cs) / "config" / "setup.yaml"
        if setup_path.exists():
            configs[cs] = yaml.safe_load(setup_path.read_text())
        else:
            configs[cs] = {}
        summary = explorers[cs].summary()
        total = sum(summary.values())
        if total == 0:
            empty_registries.append(cs)
            print(f"  [EMPTY] {cs}: registry.db present, zero backtest runs")
        else:
            print(f"  [OK] {cs}: {total} backtests ({summary})")
    except FileNotFoundError:
        unreadable_registries.append(cs)
        print(f"  [MISSING] {cs}: no registry.db")

print(f"\nLoaded: {len(explorers)}/{len(ALL_CASE_STUDIES)} case studies")

# %% [markdown]
# ### The full-mode precondition
#
# This notebook overwrites the nine-case-study artifacts that notebooks 02 through 08 and the
# chapter figures read. A run that finds only some of the nine registries produces an output
# indistinguishable from a complete one, and on 2026-08-28 that is exactly what happened: a
# synthesis run from a worktree carrying three registries stamped itself production and
# published a holdout Sharpe under prose describing nine case studies. The push gate caught it;
# the notebook did not.
#
# So a full-mode run refuses to continue unless every case study named above has a readable
# registry holding backtests. A subset run - `CASE_STUDIES` non-empty, the per-case-study
# driver path that repopulates one case study's paired metrics after its holdout lands - is not
# full mode and is not covered by the check.


# %%
def refuse_partial_full_mode(
    *,
    expected: list[str],
    subset: list[str],
    unreadable: list[str],
    empty: list[str],
    enforce: bool = True,
) -> None:
    """Raise unless a full-mode run can see every registry it claims to aggregate.

    A subset run - ``subset`` non-empty - is the per-case-study driver path and is not
    full mode, so it is never refused. ``enforce`` is the seeded-test-registry escape and
    is False in exactly one place, ``tests/overrides.yaml``.
    """
    if not enforce or subset or not (unreadable or empty):
        return
    raise RuntimeError(
        f"Full-mode synthesis needs all {len(expected)} registries present and holding "
        "backtests, and this checkout does not have them. Refusing before anything is "
        "written, because the artifacts this notebook overwrites are read as a complete "
        "set covering every case study.\n"
        f"  no registry.db:      {unreadable or 'none'}\n"
        f"  zero backtest runs:  {empty or 'none'}\n"
        "Run this once every case study has registered its backtests and the fleet has "
        "stopped writing to them. Passing CASE_STUDIES is not a way round this: a subset "
        "run repopulates one case study's paired metrics in its own registry and writes "
        "none of the chapter-wide artifacts."
    )


refuse_partial_full_mode(
    expected=ALL_CASE_STUDIES,
    subset=CASE_STUDIES,
    unreadable=unreadable_registries,
    empty=empty_registries,
    enforce=REQUIRE_ALL_REGISTRIES,
)

# %% [markdown]
# ## Top-Cluster Diagnostics
#
# Rather than pre-committing to a single selected configuration per case study, we inspect the
# *cluster* of top configurations on the validation split. A signal with genuine predictive
# structure shows a thick top of the distribution: many configurations sit within a
# fold-standard-error of the top-ranked Sharpe, and the implied pick is insensitive to small
# perturbations in the selection rule. A thin cluster - a large gap between the top-ranked
# configuration and the tenth - suggests the top result is closer to a tail draw than to a
# stable optimum.
#
# For each case study we report the top-ranked Sharpe, the tenth-ranked Sharpe where ten
# configurations exist, the spread between them, the mean per-fold Sharpe, and the number of
# folds in which the top-ranked configuration has positive Sharpe. These are measurements that
# feed the downstream narrative.
#
# ### The selection rule
#
# **A backtest's full strategy specification is the signal method, the allocation method and the
# risk overlay taken together.** Naming all three is what makes a validation result and a
# holdout result comparable, because it pins every stage rather than the signal alone.
#
# Each case study's selected configuration is the highest-Sharpe validation backtest across
# those three pipeline stages. The deployed holdout configuration is that same specification,
# retrained on holdout data. When the holdout retrain produces no usable backtest at it -
# degenerate predictions, a vol window that does not match the history available, a universe
# filter that rejects the sample, or another generation failure - the rule falls back to the
# next-highest validation Sharpe that does have a usable holdout, and so on until one succeeds.
# The helper implementing that walk feeds the holdout query and the lineage resolver, which pin
# each validation and holdout pair to one specification.

# %% [markdown]
# ### Two restrictions, and why the selection needs both
#
# **A label restriction**, so the cluster diagnostics and the Chapter 20 holdout retrain rank the
# same thing. sp500_options trains a hold-to-maturity label with coherent option costs alongside
# four fixed-horizon straddle labels priced through the vectorized path with a generic
# basis-point cost. The Chapter 20 narrative uses the hold-to-maturity label as its
# option-strategy reference, so restricting the cluster diagnostics to that same label keeps the
# §20.1 top-cluster numbers aligned with the §20.5 and §20.6 narrative.
#
# Those two section numbers are right, and they are recorded here because they will not look it.
# Eleven references in this chapter pointed at sections that exist and do not carry what was
# claimed, and a sweep for §20.5 in a chapter-20 notebook now finds this one and sees the same
# shape. It is not the same: §20.5's Table 20.6 carries an sp500_options allocator row, and which
# row that is depends on the label pinned here; §20.6 carries the option cost model, which is the
# hold-to-maturity accounting rather than the basis-point sweep the other four labels get. Both
# targets hold material this restriction decides. Do not retarget it.
#
# **An execution-regime restriction**, because sp500_options is evaluated under the
# O'Donovan-Yu (2025) cost-mitigation cascade, whose three rungs are a naive round trip, full
# hold-to-maturity, and hold-to-maturity restricted to the liquid bottom-spread quintile. The
# registered strategy is the third rung; the second is the demoted variant §18.8 discusses. The
# first two rungs both carry the same universe filter, so filtering on that column
# alone leaves `ORDER BY sharpe DESC LIMIT 1` free to pick whichever of the two happens to score
# higher in the current data. Pinning the universe filter *and* the exit rule together is what
# makes the selected row deterministic and coherent with hold-to-maturity. Case studies with no
# entry here skip the filter altogether.

# %%
# The rung pins are imported, not restated. This file used to carry its own copy of both
# predicates and of the dict around them, verbatim, and `paired_metrics.populate_paired_metrics`
# carries the other - both write `backtest_paired_metrics`, so a pin corrected on one side only
# would let one of them overwrite the other's rows with a differently-selected lineage. The
# duplication is how that drift happens, and the mirror keys beside each predicate
# (`universe_filter`, `exit_at_max_days`, `label`) exist for the SQL paths and `progression(...)`
# calls that cannot take a polars expression.
from case_studies.utils.paired_metrics import RUNG_PINS as _CLUSTER_RUNG_RESTRICTIONS  # noqa: E402
from case_studies.utils.strategy_analysis import (  # noqa: E402
    LABEL_RESTRICTIONS as _CLUSTER_LABEL_RESTRICTIONS,
)
from case_studies.utils.strategy_analysis import (  # noqa: E402
    NoSelectableCandidates,
    resolve_solvent_carrier,
    selectable_validation_candidates,
)


@cache
def _canonical_carrier(cs: str) -> dict | None:
    """The configuration this case study reports, from the resolver that decides it.

    This notebook built the same cross-stage rank-1 by hand in four places - concatenating
    `explorer.best` over signal, allocation and risk_overlay, dropping benchmark families,
    applying `LABEL_RESTRICTIONS` and `RUNG_PINS`, and taking the highest Sharpe. Three of
    them wanted the winner and read it from here; the fourth, `_val_rank1_carrier`, walks the
    whole field and takes it from `selectable_validation_candidates`, which is the same
    ranking one step earlier. That is not the ranking the case studies report. `resolve_canonical_rank1_lineage` re-ranks the field
    on exact common timestamp support whenever a conformal candidate is in it, because a
    conformal allocator abstains until it is calibrated and books zeros over the abstention,
    and it applies `UNIVERSE_RESTRICTIONS` and `CARRIER_PINS` besides. Measured 2026-09-18
    against the nine canonical registries, the two rankings named different configurations on
    two case studies: `fx_pairs` (`linear/ridge_a1000000.0` at Sharpe 0.4121
    against `deep_learning/lstm_h64` at comparison Sharpe 0.3091) and
    `nasdaq100_microstructure` (`deep_learning/nlinear` on fwd_ret_15m at 2.3001 against
    `gbm/default_multiclass` on fwd_dir_15m at 2.4159). `spine_prediction_hash` is what
    `05_portfolio_allocation` and Figure 20.7 pin their allocator comparison to, so on
    `fx_pairs` the chapter compared allocators on a configuration the case study does not
    report.

    Returns None where the resolver finds nothing selectable, which is the state the
    hand-built rankings reported as an empty frame. An insolvent or mis-calibrated carrier
    still raises: it is a sweep to fix, not a case study to skip.
    """
    try:
        return resolve_solvent_carrier(cs)
    except NoSelectableCandidates:
        return None


def _retired(cs: str) -> frozenset[str]:
    """Identities a later generation retired, from the same helper `populate_paired_metrics`
    uses. Both write `backtest_paired_metrics`, so a disagreement here would let a Chapter 20
    run overwrite the corrected pairs with a retired lineage."""
    from case_studies.utils.paired_metrics import _retired_prediction_hashes

    return _retired_prediction_hashes(cs)


@cache
def _live_predictions(cs: str) -> list[str] | None:
    """What this case study currently publishes, or None when it declares no populations.

    Membership, not the complement of retirement. A prediction no population ever listed has
    not been retired by anyone, so ranking over "everything not retired" admits experimental
    results the case study never published; ranking over the members in force does not.

    Applied inside the query rather than to its result, because `best()` applies its SQL
    `LIMIT top_n` first - a row filtered afterwards has already consumed a slot and can hide a
    live candidate below the cut.
    """
    from case_studies.research.population import published_members_at

    published = published_members_at(get_case_study_dir(cs), member_kind="prediction")
    if published is None:
        return None
    if not published:
        # `best()` tests this argument for truthiness, so an empty list would read as "no
        # filter" and rank everything. A study that declares populations and publishes
        # nothing has nothing to report, which is a refusal rather than a wide-open ranking.
        raise RuntimeError(f"{cs} declares populations but publishes no prediction identities")
    return sorted(published)


def _best_live(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
    """`explorer.best` narrowed to what the case study still publishes."""
    return explorer.best(stage=stage, top_n=top_n, prediction_hashes=_live_predictions(cs))


def _best_pinned(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
    """`explorer.best` for a stage, fetching enough rows that a rung-restricted
    cohort survives the post-hoc predicate filter.

    `best()` extracts `universe_filter` from `spec_json` in Python, after the
    SQL `LIMIT top_n`. For nasdaq the pinned cost-feasible carrier sits below
    the full-universe in-sample maxima, so a small `top_n` truncates it before
    `_apply_rung_restriction` runs. Pull all rows for restricted case studies."""
    live = _live_predictions(cs)
    if cs in _CLUSTER_RUNG_RESTRICTIONS:
        return explorer.best(stage=stage, top_n=1_000_000, prediction_hashes=live)
    return explorer.best(stage=stage, top_n=top_n, prediction_hashes=live)


def _apply_rung_restriction(df: pl.DataFrame, cs: str) -> pl.DataFrame:
    """Filter `df` to the case study's pinned rung, if one is configured.

    Returns the input untouched if no restriction applies. The helper
    relies on `BacktestExplorer.best()` always emitting both
    `universe_filter` and `exit_at_max_days` columns; if a future
    schema regression drops them, the polars `filter` will raise a
    column-not-found error rather than silently allowing the rank-1
    selection to drift back to the cross-rung max."""
    rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
    if rung is None or df.is_empty():
        return df
    return df.filter(rung["predicate"])


# There is no selected configuration pin here, and there is no mechanism for one.
# `_CARRIER_PIN_PREDICATES` held `"us_firm_characteristics": pl.col("config_name") ==
# "default_huber"` until 2026-08-25, copied from `case_studies.utils.strategy_analysis.CARRIER_PINS`
# and translated into a config-name predicate, under a "keep in sync" comment doing the job a
# mechanism should.
#
# It had not been in sync for a rebuild. Against the current registry `default_huber` is the
# WEAKEST of the ten configs that reached the allocation stage (48 validation backtests, best
# Sharpe 2.128, 2.075 average - tenth of ten), while the documented rule selects `leaves_63_mse`
# (59 backtests, 3.116). So this restricted one case study to its worst advanced configuration
# while every notebook inside that case study reported its best.
#
# Worse than the hash pin removed from `CARRIER_PINS` the same day, because a hash pin dies
# loudly: every hash changes when a sweep is rebuilt, so it resolves to nothing and stops. A
# config-name predicate survives the rebuild and keeps selecting, silently and wrongly.
#
# The mapping stayed empty behind an `_apply_carrier_pin` that could no longer fire, which is a
# second implementation of a rule nothing applied. A selected configuration restriction needed here
# again is `carrier_pins.carrier_config_name(cs)`, which resolves an owner's pin to its config
# through the registry - the thing the copy existed to avoid, and the thing that would have failed
# loudly rather than filtering to the wrong config.


def _progression_for(
    explorer: "BacktestExplorer",
    pred_hash: str,
    cs: str,
) -> pl.DataFrame:
    """Call `progression()` with the case study's rung scope, if any."""
    rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
    if rung is None:
        return explorer.progression(pred_hash)
    return explorer.progression(
        pred_hash,
        universe_filter=rung["universe_filter"],
        exit_at_max_days=rung["exit_at_max_days"],
    )


# Stages whose registry numbers should never be reported for a case study,
# either because the strategy makes the stage structurally meaningless (HTM
# short-straddle has no allocator choice, no bps cost sweep) or because the
# legacy registry contains deprecated entries that pre-date the strategy
# redesign. Consumed by both `build_backtest_rows` and the `synthesis_dict`
# sanitizer below so the in-notebook attrition funnel and the JSON artifact
# cannot drift on the same case study.
_STAGES_NOT_APPLICABLE: dict[str, set[str]] = {
    "sp500_options": {"costs", "risk"},
    "us_firm_characteristics": {"risk"},
    "nasdaq100_microstructure": {"allocation", "costs", "risk"},
}

# Per-CS rationale strings for `not_applicable_reason` fields written into
# `synthesis_dict`. Keyed by (cs, stage) so two case studies that skip the
# same stage for different structural reasons render different prose.
_STAGE_NA_REASONS: dict[tuple[str, str], str] = {
    ("sp500_options", "allocation"): ("HTM short-straddle has fixed 1/n_roll cohort weighting"),
    ("sp500_options", "costs"): ("option costs use §18.8 bid-ask accounting, not bps sweep"),
    ("sp500_options", "risk"): ("HTM expiration structure sets risk profile"),
    ("us_firm_characteristics", "risk"): (
        "vectorized-engine path; portfolio overlays purged 2026-05-17"
    ),
    ("nasdaq100_microstructure", "allocation"): (
        "carrier is a signal-stage slot strategy; the slot mechanism is the sizing rule"
    ),
    ("nasdaq100_microstructure", "costs"): (
        "timing-corrected broad carrier cost grid deferred to v3.1"
    ),
    ("nasdaq100_microstructure", "risk"): (
        "timing-corrected broad carrier risk grid deferred to v3.1"
    ),
}


def _stage_applicable(cs: str, stage: str) -> bool:
    """Return False if `cs` has stage `stage` declared not applicable.

    `stage` must be one of the canonical labels used by
    ``_STAGES_NOT_APPLICABLE`` itself (`allocation`, `costs`, `risk`).
    Callers in this notebook always pass canonical literals."""
    return stage not in _STAGES_NOT_APPLICABLE.get(cs, set())


# %%
cluster_rows = []
for cs, explorer in explorers.items():
    top = _best_pinned(explorer, cs, "signal", 200)
    if top.is_empty() or top["sharpe"][0] is None:
        continue
    if "family" in top.columns:
        top = top.filter(pl.col("family") != "benchmark")
    label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
    if label_restriction and "label" in top.columns:
        top = top.filter(pl.col("label").is_in(list(label_restriction)))
    top = _apply_rung_restriction(top, cs)
    if top.is_empty() or top["sharpe"][0] is None:
        continue
    rank1 = top["sharpe"][0]
    rank10 = top["sharpe"][9] if top.height >= 10 else None
    spread = (rank1 - rank10) if rank10 is not None else None
    # Fold-level stability for the rank-1 backtest
    try:
        bt_hash = top["backtest_hash"][0]
        fold_df = explorer.fold_performance(bt_hash)
        if not fold_df.is_empty():
            mean_fold_sh = float(fold_df["sharpe"].mean())
            se_fold_sh = (
                float(fold_df["sharpe"].std(ddof=1) / (fold_df.height**0.5))
                if fold_df.height > 1
                else None
            )
            n_folds_pos = int(fold_df.filter(pl.col("sharpe") > 0).height)
            n_folds = fold_df.height
        else:
            mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0
    except Exception:
        mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0

    cluster_rows.append(
        {
            "case_study": DISPLAY_NAMES.get(cs, cs),
            "cs_id": cs,
            "rank1_sharpe": rank1,
            "rank10_sharpe": rank10,
            "rank1_rank10_spread": spread,
            "mean_per_fold_sharpe": mean_fold_sh,
            "fold_sharpe_se": se_fold_sh,
            "n_folds_pos": n_folds_pos,
            "n_folds": n_folds,
            "n_configs": int(top.height),
        }
    )

cluster_df = pl.DataFrame(cluster_rows)
if not cluster_df.is_empty():
    print("\n=== Rank-1 Cluster Diagnostics (validation) ===")
    print(
        cluster_df.select(
            "case_study",
            "rank1_sharpe",
            "rank10_sharpe",
            "rank1_rank10_spread",
            "mean_per_fold_sharpe",
            "fold_sharpe_se",
            "n_folds_pos",
            "n_folds",
            "n_configs",
        )
    )

# %% [markdown]
# Read the table as a tuple: (rank1 Sharpe, rank10 Sharpe, spread, fold-SE,
# folds-positive). A small rank1→rank10 spread relative to the fold-SE signals
# a thick top-of-distribution. Folds-positive close to the total fold count
# signals temporal stability. Both can be read off per case study without
# collapsing the evidence into a single label.

# %% [markdown]
# ## Overview Table
#
# The nine case studies differ in asset class, rebalancing frequency, universe
# size, and the cost assumption each one carries. The table records those four
# properties so that a later result can be attributed to the setting it was
# measured in rather than to the model that produced it.

# %%
overview_rows = []
for cs, explorer in explorers.items():
    setup = configs.get(cs, {})
    cost_bps = compute_cost_bps(setup)
    universe = setup.get("universe", {})
    # A futures universe is sized in products rather than in assets, so
    # cme_futures declares `n_products` where the others declare `n_assets`.
    # Reading only the latter reported that case study as an empty universe.
    n_assets = (
        universe.get("n_assets")
        or universe.get("n_products")
        or len(universe.get("symbols", []))
        or 0
    )
    primary_label = setup.get("labels", {}).get("primary", "")

    families = explorer.compare_families(stage="signal")

    overview_rows.append(
        {
            "case_study": DISPLAY_NAMES.get(cs, cs),
            "cs_id": cs,
            "asset_class": ASSET_CLASS_MAP.get(cs, "unknown"),
            "frequency": FREQ_MAP.get(cs, "daily"),
            "universe": n_assets,
            "primary_label": primary_label,
            "cost_bps": cost_bps,
            "n_model_families": len(families) if not families.is_empty() else 0,
        }
    )

overview_df = pl.DataFrame(overview_rows)
overview_df.select("case_study", "asset_class", "frequency", "universe", "cost_bps")

# %% [markdown]
# The test bed covers equity ETFs, crypto perpetuals, intraday microstructure, equity plus
# options, firm characteristics, FX, futures, pure options, and a broad equity panel. The
# `cost_bps` column of the table above records the transaction-cost assumption each one
# carries, so a later result can be read against the cost regime it was measured under.

# %% [markdown]
# ## Model IC Comparison
#
# Mean IC by model family across case studies, queried from `prediction_metrics`.
# Each cell shows the average IC across all configurations within a family,
# filtered to each case study's primary label.

# %%
ic_rows = []
for cs, explorer in explorers.items():
    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    if not db_path.exists():
        continue

    # Filter by primary label so IC values match book prose
    primary_label = configs.get(cs, {}).get("labels", {}).get("primary", "")

    db = sqlite3.connect(str(db_path))
    # Best IC per family on primary label only
    # Exclude causal_dml: it estimates treatment effects, not predictive IC.
    # NOTE: best_ic and best_ic_daily are independent per-family MAXes — they may
    # come from *different* predictions. This is intentional ("best daily IC in
    # the family"), not "the daily IC of the best-by-fold model".
    query = """
        SELECT t.family, MAX(pm.ic_mean) AS best_ic,
               MAX(pm.ic_mean_daily) AS best_ic_daily,
               AVG(pm.ic_mean) AS mean_ic, COUNT(*) AS n_preds
        FROM training_runs t
        JOIN prediction_sets p ON t.training_hash = p.training_hash
        JOIN prediction_metrics pm ON p.prediction_hash = pm.prediction_hash
        WHERE p.split != 'holdout'
          AND pm.ic_mean IS NOT NULL
          AND t.family != 'causal_dml'
    """
    params: tuple = ()
    if primary_label:
        query += "      AND t.label = ?\n"
        params = (primary_label,)
    query += "    GROUP BY t.family"

    rows = db.execute(query, params).fetchall()
    db.close()

    for family, best_ic, best_ic_daily, mean_ic, n_preds in rows:
        ic_rows.append(
            {
                "case_study": DISPLAY_NAMES.get(cs, cs),
                "family": family,
                "ic_mean": mean_ic,
                "ic_best": best_ic,
                "ic_best_daily": best_ic_daily,
                "n_predictions": n_preds,
            }
        )

ic_df = pl.DataFrame(ic_rows)

# %%
if not ic_df.is_empty():
    ic_pivot = ic_df.pivot(on="family", index="case_study", values="ic_mean").sort("case_study")
else:
    ic_pivot = pl.DataFrame()

# %% [markdown]
# ### Mean IC by Model Family

# %%
ic_pivot

# %% [markdown]
# Each cell is the mean of `ic_mean` over that family's non-holdout prediction sets, taken at
# the primary label the case study declares in `setup.yaml`, with `causal_dml` excluded because
# its runs are not fit to predict. Families are comparable within a row, since the label is
# fixed across the row; they are not comparable across rows, because each case study declares a
# different primary label, from a fifteen-minute forward return to a twenty-one-day one, and
# prices a different instrument. A negative mean IC marks a case study where prediction is
# difficult under that label rather than a defect in the family. §20.3 carries this table as
# Table 20.4, and §20.6 works through how an option case study's raw IC translates into Sharpe
# once single-name execution costs are charged.

# %% [markdown]
# ## Backtest Comparison
#
# Cross-dataset comparison of pipeline outcomes. Each row takes the highest-Sharpe result at
# each stage **independently**, so the signal that tops one column may come from a different
# model than the allocation that tops the next.


# %%
def build_backtest_rows():
    """Build backtest comparison rows from all case study explorers."""
    bt_rows = []
    for cs, explorer in explorers.items():
        summary = explorer.summary()

        # Best signal-stage result — exclude benchmark families (equal_weight,
        # etc.) since §20.4's model comparison is about trained models, not
        # passive baselines. Also apply case-study label and universe-filter
        # restrictions so the Ch20 rank-1 is HTM-coherent for sp500_options
        # and pinned to the Rung-3 liquid subset, which is what
        # `strategy_analysis.UNIVERSE_RESTRICTIONS` holds ({"sp500_options":
        # "liquid"}). The Rung-2 full universe is retained only for the §18.8
        # cascade comparison and never anchors the deployed carrier.
        label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
        signal_candidates = _best_pinned(explorer, cs, "signal", 200)
        if not signal_candidates.is_empty() and "family" in signal_candidates.columns:
            signal_candidates = signal_candidates.filter(pl.col("family") != "benchmark")
        if (
            label_restriction
            and "label" in signal_candidates.columns
            and not signal_candidates.is_empty()
        ):
            signal_candidates = signal_candidates.filter(
                pl.col("label").is_in(list(label_restriction))
            )
        signal_candidates = _apply_rung_restriction(signal_candidates, cs)
        best_signal = signal_candidates.head(1)
        signal_sharpe = best_signal["sharpe"][0] if not best_signal.is_empty() else None
        best_source = best_signal["source"][0] if not best_signal.is_empty() else ""

        # Carrier-pred pin for cost/risk. Case studies with a rung restriction
        # (nasdaq cost-feasible ensemble) carry their headline cost/risk on the
        # selected prediction only; the full-universe sweep rows are the
        # Ch18/Ch19 cost-defeat demonstration and must not pool into the
        # cross-case comparison. Other case studies pass None (no pin) and keep
        # the registry-wide aggregation unchanged.
        carrier_pred = (
            best_signal["prediction_hash"][0]
            if cs in _CLUSTER_RUNG_RESTRICTIONS and not best_signal.is_empty()
            else None
        )

        # Best allocation-stage result (same filters). For case studies that
        # declare the allocation stage not applicable (e.g. sp500_options HTM),
        # the registry numbers come from deprecated runs, so report None to
        # match the synthesis_dict sanitizer below.
        if _stage_applicable(cs, "allocation"):
            alloc_candidates = _best_live(explorer, cs, "allocation", 200)
            if not alloc_candidates.is_empty() and "family" in alloc_candidates.columns:
                alloc_candidates = alloc_candidates.filter(pl.col("family") != "benchmark")
            if (
                label_restriction
                and "label" in alloc_candidates.columns
                and not alloc_candidates.is_empty()
            ):
                alloc_candidates = alloc_candidates.filter(
                    pl.col("label").is_in(list(label_restriction))
                )
            alloc_candidates = _apply_rung_restriction(alloc_candidates, cs)
            best_alloc = alloc_candidates.head(1)
            alloc_sharpe = best_alloc["sharpe"][0] if not best_alloc.is_empty() else None
            # The allocator that produced `alloc_sharpe`, read from that row's own spec, so
            # the name and the number describe one configuration. See
            # `strategy_analysis.allocation_method_of`.
            best_allocator = allocation_method_of(
                cs, best_alloc["backtest_hash"][0] if not best_alloc.is_empty() else None
            )
        else:
            alloc_sharpe = None
            best_allocator = ""

        # Cost sensitivity (gated by stage policy)
        survives_costs = None
        if _stage_applicable(cs, "costs"):
            cost_df = explorer.cost_sensitivity(prediction_hash=carrier_pred)
            if not cost_df.is_empty():
                zero_cost = cost_df.filter(pl.col("cost_bps") == 0)
                survives_costs = not zero_cost.is_empty() and zero_cost["sharpe"].max() > 0

        # Risk overlay (gated by stage policy)
        best_overlay = ""
        managed_sharpe = None
        if _stage_applicable(cs, "risk"):
            risk_df = explorer.risk_impact(prediction_hash=carrier_pred)
            if not risk_df.is_empty():
                # rank_one, not a one-key sort: overlays that never trigger book the
                # baseline Sharpe exactly, so ties at the top are ordinary here and a
                # one-key sort would let frame order decide the name reported below.
                best_risk_row = rank_one(risk_df, by="sharpe", name="risk_name")
                best_overlay = best_risk_row["risk_name"][0]
                managed_sharpe = best_risk_row["sharpe"][0]

        # Spine rank-1 prediction_hash - the configuration the case study reports, taken
        # from `_canonical_carrier` rather than ranked a second time here. Figure 20.7 and
        # `05_portfolio_allocation` both read this value, and Ch20 prose Tables 20.5-20.7
        # quote the resolver, so the two have to be one answer.
        _carrier = _canonical_carrier(cs)
        spine_pred_hash = _carrier["val_prediction_hash"] if _carrier else None

        bt_rows.append(
            {
                "case_study": DISPLAY_NAMES.get(cs, cs),
                "case_study_id": cs,
                "spine_prediction_hash": spine_pred_hash,
                "n_signal": summary.get("signal", 0),
                "n_allocation": summary.get("allocation", 0),
                "n_cost": summary.get("cost_sensitivity", 0),
                "n_risk": summary.get("risk_overlay", 0),
                "best_source": best_source,
                "signal_sharpe": signal_sharpe,
                "best_allocator": best_allocator,
                "alloc_sharpe": alloc_sharpe,
                "survives_costs": survives_costs,
                "best_overlay": best_overlay,
                "managed_sharpe": managed_sharpe,
            }
        )
    return bt_rows


# %%
bt_rows = build_backtest_rows()

# %%
bt_df = pl.DataFrame(bt_rows)
print("\nPipeline Comparison:")
print(
    bt_df.select(
        "case_study",
        "signal_sharpe",
        "alloc_sharpe",
        "survives_costs",
        "managed_sharpe",
    )
)

# %% [markdown]
# Read the baseline column of the table above for how many case studies enter the pipeline with
# a positive baseline-stage Sharpe and which do not. Those counts move whenever a registry is
# rebuilt, which is why they are in the table rather than in this sentence.
#
# The lineage table below traces each selected prediction across the stages in the order the
# backtests run: baseline, allocation, risk overlay, then the cost sweep charged against
# whatever survived. A Sharpe that rises from one column to the next is what that stage added,
# and only where the later stage carries the earlier one's configuration - the paired rows above
# say which transitions meet that test. NASDAQ-100 is excluded from that comparison in v3.0
# because its timing-corrected broad cost and risk grids are deferred to v3.1.

# %% [markdown]
# ## Paired-Bootstrap Comparison vs Equal-Weight Benchmark
#
# Each case study's selected baseline-stage backtest, under the same label, universe-filter and
# rung restrictions used for the cluster diagnostics, is compared to its equal-weight benchmark
# using a **paired stationary block bootstrap on daily strategy returns**. Block length is derived from
# ``setup.yaml.labels.{label}.rebalance_step`` (falling back to the optimal
# block size, never below the label horizon). Reported quantities:
#
# - ``sharpe_diff`` with a bootstrap confidence interval
# - ``ret_diff``, the annualized return difference, with its confidence interval
# - ``info_ratio`` of the daily-return difference
# - ``prob_challenger_wins`` — bootstrap fraction in which challenger Sharpe
#   exceeds the benchmark
# - ``p_value`` — two-sided bootstrap p-value for ``sharpe_diff = 0``
#
# Results land in ``backtest_paired_metrics`` (per case study) and roll up
# into the cross-dataset table below. Intervals are at the conventional confidence level the
# bootstrap call sets. This is the right unit of uncertainty for the headline claim about the
# selected configuration: the Sharpe **difference against the passive baseline that experienced
# the same market conditions**, rather than the Sharpe alone.


# %%
def _benchmark_returns_from_artifact(
    cs: str, label: str, period: str = "overall"
) -> tuple[str, pl.DataFrame, str] | None:
    """Resolve the side-artifact equal-weight benchmark for ``(cs, label)``.

    The benchmark is the daily-MTM EW reference series persisted by
    ``scripts/compute_vectorized_ew_benchmark.py`` at
    ``case_studies/{cs}/benchmark/{label}.parquet``. Single, well-defined
    methodology per (cs, label) — no universe/rung/cadence ambiguity that
    a registry-side ``family='benchmark'`` lookup would have to disambiguate.

    ``period`` selects the time window slice (``"overall"`` or ``"holdout"``)
    applied by ``load_benchmark_returns``. Classification-label fallback to
    the matching ``fwd_ret_*`` artifact applies in both periods.

    Returns ``(synthetic_hash, returns_df, resolved_label)`` where
    ``synthetic_hash`` is a deterministic identifier safe to use as the PK
    column in ``backtest_paired_metrics`` (which has no FK on
    ``benchmark_hash``) and ``resolved_label`` is the actual label whose
    artifact was loaded — equal to ``label`` unless the classification
    fallback fired, in which case it's the matching ``fwd_ret_*`` label.
    Returns ``None`` if the artifact is missing.
    """
    df = load_benchmark_returns(cs, label, period=period)
    bench_label = label
    if df.is_empty() or "ew_return" not in df.columns:
        # Fallback: classification labels (``fwd_class_*``, ``fwd_dir_*``,
        # ``fwd_tb_*``, ``fwd_carry_*``) share the same forecast window as
        # their continuous counterpart (``fwd_ret_*``). The EW universe over
        # the same window is identical regardless of the label being
        # predicted, so map e.g. ``fwd_class_1m`` to ``fwd_ret_1m``.
        fallback = None
        for prefix in ("fwd_class_", "fwd_dir_", "fwd_tb_", "fwd_carry_"):
            if label.startswith(prefix):
                fallback = "fwd_ret_" + label[len(prefix) :]
                break
        if fallback is None:
            return None
        df = load_benchmark_returns(cs, fallback, period=period)
        if df.is_empty() or "ew_return" not in df.columns:
            return None
        bench_label = fallback
    suffix = "" if period == "overall" else f":{period}"
    bench_hash = f"side_ew:{cs}:{bench_label}{suffix}"
    return (
        bench_hash,
        df.select(
            pl.col("timestamp").cast(pl.Date).alias("timestamp"),
            pl.col("ew_return").cast(pl.Float64).alias("ret"),
        ),
        bench_label,
    )


def _aligned_returns(cs: str, h: str) -> pl.DataFrame | None:
    """Load and normalize a backtest's daily returns; columns ``[timestamp, ret]``."""
    parquet = get_case_study_dir(cs) / "run_log" / "backtest" / h / "daily_returns.parquet"
    if not parquet.exists():
        return None
    df = pl.read_parquet(parquet)
    ret_col = next(
        (c for c in ("daily_return", "ret", "return", "value") if c in df.columns),
        df.columns[-1],
    )
    ts_col = next(
        (c for c in ("timestamp", "date", "datetime") if c in df.columns),
        df.columns[0],
    )
    return df.select(
        pl.col(ts_col).cast(pl.Date).alias("timestamp"),
        pl.col(ret_col).cast(pl.Float64).alias("ret"),
    )


# %%
import numpy as np

from case_studies.utils.uncertainty import (
    SIGNAL_BASELINE_BY_CASE_STUDY,
    STAGE_SEQUENCE,
    compute_independent_diff_uncertainty,
    compute_paired_uncertainty,
    descends_from,
    joint_returns,
)


def _min_paired_n(ppy: int) -> int:
    """Minimum series length for paired-bootstrap stability, frequency-aware.

    The ~21 floor was written for daily cadences (about a month of obs).
    Monthly case studies (e.g. ``us_firm_characteristics``) have ~12 holdout
    observations by design, and ``compute_paired_uncertainty`` runs cleanly
    on n=12. Scale the floor with ``ppy`` so monthly/weekly CSs aren't
    blocked by a daily-tuned guard.
    """
    if ppy <= 12:  # monthly
        return 6
    if ppy <= 52:  # weekly
        return 12
    return 21  # daily / 8h / intraday


# Distinguish skipped CSs from real failures so empty cross-dataset rollups
# aren't indistinguishable from a silent crash.
paired_rows: list[dict] = []
paired_skips: list[dict] = []

for cs, explorer in explorers.items():
    # The leader is the configuration the case study reports, not a ranking rebuilt here.
    # `_canonical_carrier` documents why the two are not the same ordering.
    carrier = _canonical_carrier(cs)
    if carrier is None:
        paired_skips.append({"case_study": cs, "reason": "no_selectable_candidates"})
        continue
    leader_hash = carrier["val_backtest_hash"]
    leader_label = carrier["label"]
    if not leader_label:
        paired_skips.append({"case_study": cs, "reason": "no_label_on_leader"})
        continue
    bench_resolution = _benchmark_returns_from_artifact(cs, leader_label)
    if not bench_resolution:
        paired_skips.append(
            {"case_study": cs, "reason": f"no_benchmark_artifact_for_label:{leader_label}"}
        )
        continue
    benchmark_hash, base, resolved_bench_label = bench_resolution

    chal = _aligned_returns(cs, leader_hash)
    if chal is None:
        paired_skips.append({"case_study": cs, "reason": "no_challenger_returns_parquet"})
        continue

    ppy = {"daily": 252, "weekly": 52, "monthly": 12, "8h": 1095}.get(
        FREQ_MAP.get(cs, "daily"), 252
    )
    min_n = _min_paired_n(ppy)
    aligned = chal.join(base, on="timestamp", how="inner", suffix="_b")
    if aligned.height < min_n:
        paired_skips.append(
            {"case_study": cs, "reason": f"insufficient_overlap:n={aligned.height}"}
        )
        continue
    # A strategy against a benchmark: the leader's leading flat run is warmup before its
    # first signal, not a position it held, so the sample starts where both are trading.
    c_arr, b_arr = joint_returns(aligned["ret"].to_numpy(), aligned["ret_b"].to_numpy())
    if c_arr.size < min_n:
        paired_skips.append(
            {"case_study": cs, "reason": f"insufficient_after_coerce:n={c_arr.size}"}
        )
        continue
    paired = compute_paired_uncertainty(
        c_arr,
        b_arr,
        periods_per_year=ppy,
        case_study=cs,
        label=leader_label,
        n_boot=2000,
        seed=42,
    )
    if not paired:
        paired_skips.append({"case_study": cs, "reason": "paired_uncertainty_empty"})
        continue

    # Side-artifact benchmark — deterministic across (cs, label), no
    # universe/rung ambiguity, no fallback-by-recency.
    benchmark_kind = f"{SIGNAL_BASELINE_BY_CASE_STUDY.get(cs, 'equal_weight')}_side_artifact"
    paired_rows.append(
        {
            "case_study": DISPLAY_NAMES.get(cs, cs),
            "label": leader_label,
            "benchmark_label": resolved_bench_label,  # may differ from leader_label when classification fallback fired
            "sharpe_diff": paired.get("sharpe_diff"),
            "sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
            "sharpe_diff_ci_hi": paired.get("sharpe_diff_ci95_hi"),
            "ret_diff": paired.get("ret_diff"),
            "info_ratio": paired.get("info_ratio"),
            "p_value": paired.get("p_value"),
            "prob_wins": paired.get("prob_challenger_wins"),
            "block": paired.get("bootstrap_block_length"),
            "n_boot": paired.get("bootstrap_n"),
        }
    )

paired_df = pl.DataFrame(paired_rows)
if not paired_df.is_empty():
    print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
    print(
        paired_df.select(
            "case_study",
            "label",
            "benchmark_label",
            "sharpe_diff",
            "sharpe_diff_ci_lo",
            "sharpe_diff_ci_hi",
            "info_ratio",
            "prob_wins",
            "p_value",
        )
    )
else:
    print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
    print("No paired-bootstrap rows produced — see skip table below for reasons.")

if paired_skips:
    print("\nSkipped case studies:")
    for s in paired_skips:
        print(f"  - {s['case_study']:<32}  {s['reason']}")

# Loud invariant — a 0/N or all-skipped outcome is now obvious in the
# notebook output instead of buried under "no paired-bootstrap rows."
print(f"\npaired={len(paired_rows)}/{len(explorers)}, skipped={len(paired_skips)}/{len(explorers)}")

# %% [markdown]
# Read each row as the selected challenger's annualized Sharpe **minus** the equal-weight
# benchmark's, with a confidence interval from the paired stationary block bootstrap on the
# daily-return difference; the information ratio summarizes the
# excess-return-to-tracking-error ratio; ``prob_wins`` is the fraction of
# bootstrap resamples in which the challenger beat the benchmark; ``p_value``
# tests ``H0: sharpe_diff = 0``. A confident "the model adds skill over the
# passive baseline" claim requires (i) the CI excludes zero, (ii) ``prob_wins``
# close to 1, and (iii) a small ``p_value``. Cases where the CI straddles
# zero are not failures — they signal that the apparent Sharpe gap is within
# block-bootstrap sampling error and should be reported as such.

# %% [markdown]
# ## Paired metrics — full coverage for strategy-analysis notebook
#
# The block above populates the first pair type, the selected baseline signal against
# equal-weight over the whole window. The strategy-analysis notebook (per-CS strategy notebooks)
# requires five additional pair types per case study to render §2 (stage-
# transition waterfall), §6 (holdout decay + holdout-vs-benchmark) and §7
# (benchmark-aware diagnostics) without inline bootstrap recomputation.
#
# The pair set:
#
# 1. selected signal (overall) ↔ equal-weight (overall) — populated above
# 2. selected signal (holdout) ↔ equal-weight (holdout window)
# 3. the selected configuration on holdout ↔ the same configuration on
#    validation (same lineage decay; min-length truncation since the
#    windows are disjoint)
# 4-6. one pair per consecutive stage transition the prediction actually has,
#    in ``STAGE_SEQUENCE`` order: allocation ↔ signal, risk-overlay ↔
#    allocation, cost-sensitivity ↔ risk-overlay. A case study that did not
#    run a stage yields fewer pairs, and a stage that does not carry the
#    previous stage's configuration yields none for that transition - the two
#    were selected independently and their difference is not a stage effect.
#
# Pair #3 truncates both series to ``min(len(val), len(ho))`` to satisfy
# ``compute_paired_uncertainty``'s equal-length precondition. The CI is
# interpreted as bootstrap resampling Sharpe in each window independently
# and taking the difference; the truncation is preserved in the
# ``benchmark_kind`` value (``val_rank1_self`` always carries the truncation
# caveat). All pairs use the same paired stationary block bootstrap helper.


# %%
def _full_strategy_spec_from_backtest(db: sqlite3.Connection, bt_hash: str) -> dict | None:
    """Pull the full strategy spec dict (signal + allocation + risk) from
    `bt_hash`'s spec_json. Returns None if the row is missing or signal has
    no `method` field.

    A backtest's full specification is the tuple (signal, allocation, risk). Pinning
    the val→holdout pair on this full spec keeps the comparison apples-to-
    apples; pinning on signal alone allows MAX(sharpe) to surface a holdout
    row with a different allocation (e.g. conformal_weighted) or risk overlay
    than the validation rank-1 carrier.
    """
    row = db.execute(
        "SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
        (bt_hash,),
    ).fetchone()
    if not row:
        return None
    strat = json.loads(row[0]).get("strategy", {})
    sig = strat.get("signal", {})
    if not sig.get("method"):
        return None
    alloc = strat.get("allocation") or {}
    risk = strat.get("risk") or {}
    return {
        "signal": {
            "method": sig.get("method"),
            "top_k": sig.get("top_k"),
            "percentile": sig.get("percentile"),
        },
        "allocation": {
            "method": alloc.get("method"),
            "top_k": alloc.get("top_k"),
            "long_short": alloc.get("long_short"),
        },
        "risk": {
            "name": risk.get("name"),
        },
    }


def _val_rank1_carrier(cs: str) -> dict | None:
    """Return ``{'spec', 'prediction_hash'}`` for ``cs``'s validation rank-1 carrier.

    The prediction hash is carried out alongside the spec because the holdout resolver
    needs it: naming all three stages pins the configuration AND the checkpoint, and without
    it a case study that registered several checkpoints against one strategy is ambiguous
    and the resolver refuses. It was determinable all along - this walk had it in hand and
    threw it away - so refusing there would have dropped a case study out of the
    reader-facing holdout table for want of a value one line above.

    The val rank-1 *full strategy* spec for ``cs`` — the
    highest-Sharpe validation backtest across (signal, allocation,
    risk_overlay) stages — walking candidates by val Sharpe descending until
    one with a matching holdout backtest at the SAME full spec is found.

    Implements the selection rule documented in §20.1: the deployed
    holdout for each case study is the val rank-1 across all three pipeline
    stages, retrained on holdout data; when retrain produces no usable
    holdout at that full spec (degenerate predictions, vol-window mismatch,
    universe filter rejection, etc.) the walk falls through to the next
    candidate by val Sharpe. The first val candidate with a registered
    holdout backtest at the same (signal, allocation, risk) tuple defines
    the apples-to-apples carrier pair.

    Returns None when no val candidate up to rank ~200 has a matching
    holdout under the case study's label / rung restrictions.
    """
    explorer = explorers.get(cs)
    if explorer is None:
        return None
    # The field the resolver ranks, in the resolver's order, rather than a concat of
    # `explorer.best` re-filtered here: the walk starts at the carrier `_canonical_carrier`
    # names and falls through in the same order the resolver would. Every row is kept - no
    # dedup by prediction_hash - because when the rank-1 configuration has no matching holdout
    # retrain but a same-prediction lower-Sharpe variant (a different allocator or risk
    # overlay) does, a dedup would jump to a different prediction instead of accepting the
    # same-prediction variant as the apples-to-apples match.
    try:
        candidates = selectable_validation_candidates(cs)
    except NoSelectableCandidates:
        # The helper raises on an empty pool rather than returning one, and this walk's
        # callers read `None` as "no holdout pair for this case study" - the state the
        # hand-built ranking reported as an empty frame. A pool with nothing eligible in it
        # is that state, not a reason to stop aggregating the other eight.
        return None
    label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)

    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
    db = sqlite3.connect(str(db_path))
    try:
        for candidate in candidates[:200]:
            bt_hash = candidate["backtest_hash"]
            spec = _full_strategy_spec_from_backtest(db, bt_hash)
            if spec is None:
                continue
            # The probe below asks whether THIS candidate has a holdout, so it matches the
            # candidate's own configuration and checkpoint and not only its strategy spec.
            #
            # Matching the spec alone made the walk stop at a candidate whose own checkpoint
            # had no holdout whenever a sibling checkpoint had one at the same spec. The
            # resolver, handed that configuration, then finds nothing for it - and the walk has
            # already stopped, so the case study reports no holdout while one exists for a
            # later candidate. Advancing instead is what makes the fall-through the resolver
            # no longer performs unnecessary rather than merely forbidden.
            carrier_row = db.execute(
                """
                SELECT t.family, t.config_name, t.label,
                       p.checkpoint_value, p.checkpoint_kind
                FROM prediction_sets p
                JOIN training_runs t ON t.training_hash = p.training_hash
                WHERE p.prediction_hash = ?
                """,
                (candidate["prediction_hash"],),
            ).fetchone()
            if carrier_row is None:
                continue
            spec_clauses, spec_params = _full_strategy_clauses(spec)
            ho_clauses = ["p.split = 'holdout'"] + spec_clauses
            if _retired(cs):
                ho_clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
            ho_params: list[object] = list(spec_params)
            if _retired(cs):
                ho_params.append(json.dumps(sorted(_retired(cs))))
            if label_restriction:
                placeholders = ",".join("?" for _ in label_restriction)
                ho_clauses.append(f"t.label IN ({placeholders})")
                ho_params.extend(sorted(label_restriction))
            if rung is not None:
                ho_clauses.append(
                    "COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
                )
                ho_params.append(rung["universe_filter"])
                if rung["exit_at_max_days"] is None:
                    ho_clauses.append(
                        "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
                    )
                else:
                    ho_clauses.append(
                        "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?"
                    )
                    ho_params.append(rung["exit_at_max_days"])
            # `t.spec_json` rather than `1`, and no LIMIT: the probe has to apply the same
            # eligibility test the resolver applies, and that test is not expressible in SQL.
            #
            # A model fitted on the validation folds can publish predictions over the holdout
            # window, so `p.split = 'holdout'` with a non-null Sharpe is not enough to make a
            # row a holdout result. The resolver drops those through
            # `training_run_fitted_for_the_holdout`; a probe that admitted them would stop the
            # walk at a candidate whose only holdout is validation-fitted, the resolver would
            # then find nothing eligible for it, and the case study would report no holdout
            # while a later candidate had a real one.
            probe_rows = db.execute(
                f"""
                SELECT t.spec_json FROM prediction_sets p
                JOIN training_runs t ON p.training_hash = t.training_hash
                JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
                                     AND b.stage IN ('signal','allocation','risk_overlay','holdout')
                JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
                WHERE {" AND ".join(ho_clauses)}
                  AND t.family = ?
                  AND t.config_name = ?
                  AND t.label = ?
                  AND p.checkpoint_value IS ?
                  AND p.checkpoint_kind IS ?
                  AND bm.sharpe IS NOT NULL
                """,
                ho_params + list(carrier_row),
            ).fetchall()
            row = any(training_run_fitted_for_the_holdout(probe[0]) for probe in probe_rows)
            if row:
                return {"spec": spec, "prediction_hash": candidate["prediction_hash"]}
    finally:
        db.close()
    return None


def _full_strategy_clauses(spec: dict | None) -> tuple[list[str], list[object]]:
    """Build SQL WHERE clauses + params that pin a backtest row to the full
    strategy spec (signal + allocation + risk). Empty list returned when
    spec is None (no constraint).

    Pinning on the full spec ensures `MAX(sharpe)` over candidate holdout
    backtests cannot surface a different allocator (e.g. conformal_weighted
    when val rank-1 was score_weighted) or a different risk overlay than
    the validation carrier — the val→holdout pair stays apples-to-apples
    on the full pipelin

Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT

Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.