Zum Inhalt springen
Alle Bibliotheksdokumente

Backtests aus Fallstudien zusammenführen und vergleichen

Code Machine Learning for Trading

Zusammenfassung

Dieses Notebook fragt Backtest-Register aus neun Fallstudien ab und erstellt vergleichbare Tabellen für nachgelagerte Analysen. Es ordnet Ergebnisse nach Anlageklasse und Datenfrequenz und erfasst die Performance über die Phasen Signal, Allokation, Kosten und Risiko hinweg, einschließlich Angaben zur Modellfamilie und Vergleichen zwischen Validierung und Holdout. Im Auswahlprozess werden vollständige Strategiespezifikationen – Signal, Allokationsmethode und Risikoüberlagerung – verglichen. Die ausgewählte Konfiguration wird dann bis zu einem neu trainierten Holdout-Lauf weiterverfolgt. Falls dieser keinen nutzbaren Holdout-Backtest erzeugen kann, wird auf einen anderen geeigneten Kandidaten aus der nach Validierungsergebnissen geordneten Auswahl zurückgegriffen.

Das Notebook misst außerdem, wie konzentriert oder stabil die besten Validierungsergebnisse erscheinen. Es berichtet den Abstand zwischen hoch und niedriger eingestuften Konfigurationen, die Variation auf Fold-Ebene und wie häufig die führende Konfiguration über die Folds hinweg positive Ergebnisse erzielt. Eine vollständige Zusammenführung verweigert kapitelweite Ausgaben, falls ein erwartetes Register fehlt oder leer ist. So kann ein unvollständiger Datensatz nicht mit einem vollständigen Vergleich verwechselt werden. Diese Diagnosen helfen, Evidenz aus mehreren Fällen einzuordnen. Die Zusammenfassungen übernehmen jedoch die Daten, Kosten, Labels und Backtest-Annahmen der jeweiligen Fallstudie. Validierungsranglisten und Holdout-Ergebnisse sind daher nicht als universelle Strategieperformance zu betrachten.

Kernaussagen

  • Vergleichen Sie Backtests aus Fallstudien anhand von Registern und einheitlichen Zusammenfassungstabellen zu den einzelnen Phasen.
  • Definieren Sie eine Strategiespezifikation gemeinsam anhand ihres Signals, ihrer Allokationsmethode und ihrer Risikoüberlagerung.
  • Verfolgen Sie die in der Validierung ausgewählte Spezifikation bis zur Holdout-Bewertung und nutzen Sie einen Ersatz, falls ihr Holdout-Lauf unbrauchbar ist.
  • Nutzen Sie die Breite der Spitzengruppe und Ergebnisse auf Fold-Ebene, um die Empfindlichkeit einer Auswahl gegenüber Rangordnungsrauschen zu beurteilen.
  • Verlangen Sie vor der Veröffentlichung vollständiger Zusammenfassungen alle Register, damit eine unvollständige Abdeckung nicht vollständig erscheint.

Schlagwörter

Volltext
# 01_aggregate_synthesis.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Aggregate Synthesis
#
# **Docker image**: `ml4t`
#
# This notebook queries all 9 case study registries via `BacktestExplorer`
# and builds cross-dataset comparison DataFrames for the remaining Ch20 notebooks.
#
# **Data source**: `registry.db` per case study (no JSON files needed).
#
# **Learning Objectives**:
# - Query per-case-study backtest registries for signal, allocation, cost, and risk metrics
# - Build cross-dataset comparison tables
# - Export summary DataFrames for downstream notebooks (02–06)
#
# **Book Reference**: Chapter 20, Section 20.1 (First-pass results across nine case studies)
#
# **Prerequisites**: Case studies must have run Ch16–19 backtests.

# %%
"""Ch20 Aggregate Synthesis — query registries and compare all 9 case studies."""

import json
import sqlite3
from functools import cache
from pathlib import Path

import polars as pl
import yaml
from IPython.display import Markdown, display

from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_returns
from case_studies.utils.strategy_analysis import (
    allocation_method_of,
    compute_cost_bps,
    rank_one,
    training_run_fitted_for_the_holdout,
)
from utils.paths import REPO_ROOT, get_case_study_dir, get_chapter_dir

# %% tags=["parameters"]
MAX_SYMBOLS = 0
# When non-empty, restricts the cross-CS iteration to the given subset.
# Used by the per-CS pipeline driver to populate `backtest_paired_metrics`
# for a single CS after its holdout has landed, without re-running the
# full 9-CS aggregation.
CASE_STUDIES: list[str] = []
# Test-only: in an isolated test registry, nasdaq's out-of-band cost-feasible
# selected configuration is absent, so its spine cannot resolve. Production leaves this False
# (a missing selection fails loudly); the test harness sets it True so cost/risk
# for such a case study are reported not-applicable instead of raising.
ALLOW_MISSING_SPINE = False
# A full-mode run overwrites the nine-case-study artifacts every downstream notebook and the
# chapter figures read. Refuse to start one unless all nine registries are present and carry
# backtests, so a partial aggregation cannot be published as a complete one. The test harness
# sets this False because its isolated registry is not the production store; a subset run
# (`CASE_STUDIES` non-empty) is the per-case-study driver path and is never full mode.
REQUIRE_ALL_REGISTRIES = True

# %%
OUTPUT_DIR = get_chapter_dir(20) / "output"
OUTPUT_DIR.mkdir(exist_ok=True)

ALL_CASE_STUDIES = [
    "etfs",
    "crypto_perps_funding",
    "nasdaq100_microstructure",
    "sp500_equity_option_analytics",
    "us_firm_characteristics",
    # FX rank-1 is linear/ridge_a100.0 on fwd_ret_21d (val Sharpe +0.048,
    # holdout +0.194), resolved after the 2026-06-01 DL-lookback fix. The
    # earlier deep_learning/tcn selection (val +0.108 / holdout -1.59) was an
    # artifact of gappy validation folds (lookback=60 warmup consumed each
    # fold's head); those sets were purged and the clean lineage re-resolved.
    # See backtest_audit.md and project_registry_hash_collisions.
    "fx_pairs",
    "cme_futures",
    "sp500_options",
    "us_equities_panel",
]
if CASE_STUDIES:
    ALL_CASE_STUDIES = [cs for cs in ALL_CASE_STUDIES if cs in set(CASE_STUDIES)]

DISPLAY_NAMES = {
    "etfs": "ETFs",
    "crypto_perps_funding": "Crypto",
    "nasdaq100_microstructure": "NASDAQ-100",
    "sp500_equity_option_analytics": "S&P 500 Eq+Opt",
    "us_firm_characteristics": "US Firms",
    "fx_pairs": "FX Pairs",
    "cme_futures": "CME Futures",
    "sp500_options": "S&P 500 Options",
    "us_equities_panel": "US Equities",
}

# %%
ASSET_CLASS_MAP = {
    "etfs": "equity_etf",
    "crypto_perps_funding": "crypto",
    "nasdaq100_microstructure": "equity_micro",
    "sp500_equity_option_analytics": "equity_options",
    "us_firm_characteristics": "equity_firm",
    "fx_pairs": "fx",
    "cme_futures": "futures",
    "sp500_options": "options",
    "us_equities_panel": "equity_panel",
}

FREQ_MAP = {
    "etfs": "daily",
    "crypto_perps_funding": "8h",
    "nasdaq100_microstructure": "15min",
    "sp500_equity_option_analytics": "daily",
    "us_firm_characteristics": "monthly",
    "fx_pairs": "daily",
    "cme_futures": "daily",
    "sp500_options": "daily",
    "us_equities_panel": "daily",
}

# %% [markdown]
# ## Load Registries
#
# Create a `BacktestExplorer` for each case study that has a registry.

# %%
explorers: dict[str, BacktestExplorer] = {}
configs: dict[str, dict] = {}
unreadable_registries: list[str] = []
empty_registries: list[str] = []

for cs in ALL_CASE_STUDIES:
    try:
        explorers[cs] = BacktestExplorer(cs)
        setup_path = get_case_study_dir(cs) / "config" / "setup.yaml"
        if setup_path.exists():
            configs[cs] = yaml.safe_load(setup_path.read_text())
        else:
            configs[cs] = {}
        summary = explorers[cs].summary()
        total = sum(summary.values())
        if total == 0:
            empty_registries.append(cs)
            print(f"  [EMPTY] {cs}: registry.db present, zero backtest runs")
        else:
            print(f"  [OK] {cs}: {total} backtests ({summary})")
    except FileNotFoundError:
        unreadable_registries.append(cs)
        print(f"  [MISSING] {cs}: no registry.db")

print(f"\nLoaded: {len(explorers)}/{len(ALL_CASE_STUDIES)} case studies")

# %% [markdown]
# ### The full-mode precondition
#
# This notebook overwrites the nine-case-study artifacts that notebooks 02 through 08 and the
# chapter figures read. A run that finds only some of the nine registries produces an output
# indistinguishable from a complete one, and on 2026-08-28 that is exactly what happened: a
# synthesis run from a worktree carrying three registries stamped itself production and
# published a holdout Sharpe under prose describing nine case studies. The push gate caught it;
# the notebook did not.
#
# So a full-mode run refuses to continue unless every case study named above has a readable
# registry holding backtests. A subset run - `CASE_STUDIES` non-empty, the per-case-study
# driver path that repopulates one case study's paired metrics after its holdout lands - is not
# full mode and is not covered by the check.


# %%
def refuse_partial_full_mode(
    *,
    expected: list[str],
    subset: list[str],
    unreadable: list[str],
    empty: list[str],
    enforce: bool = True,
) -> None:
    """Raise unless a full-mode run can see every registry it claims to aggregate.

    A subset run - ``subset`` non-empty - is the per-case-study driver path and is not
    full mode, so it is never refused. ``enforce`` is the seeded-test-registry escape and
    is False in exactly one place, ``tests/overrides.yaml``.
    """
    if not enforce or subset or not (unreadable or empty):
        return
    raise RuntimeError(
        f"Full-mode synthesis needs all {len(expected)} registries present and holding "
        "backtests, and this checkout does not have them. Refusing before anything is "
        "written, because the artifacts this notebook overwrites are read as a complete "
        "set covering every case study.\n"
        f"  no registry.db:      {unreadable or 'none'}\n"
        f"  zero backtest runs:  {empty or 'none'}\n"
        "Run this once every case study has registered its backtests and the fleet has "
        "stopped writing to them. Passing CASE_STUDIES is not a way round this: a subset "
        "run repopulates one case study's paired metrics in its own registry and writes "
        "none of the chapter-wide artifacts."
    )


refuse_partial_full_mode(
    expected=ALL_CASE_STUDIES,
    subset=CASE_STUDIES,
    unreadable=unreadable_registries,
    empty=empty_registries,
    enforce=REQUIRE_ALL_REGISTRIES,
)

# %% [markdown]
# ## Top-Cluster Diagnostics
#
# Rather than pre-committing to a single selected configuration per case study, we inspect the
# *cluster* of top configurations on the validation split. A signal with genuine predictive
# structure shows a thick top of the distribution: many configurations sit within a
# fold-standard-error of the top-ranked Sharpe, and the implied pick is insensitive to small
# perturbations in the selection rule. A thin cluster - a large gap between the top-ranked
# configuration and the tenth - suggests the top result is closer to a tail draw than to a
# stable optimum.
#
# For each case study we report the top-ranked Sharpe, the tenth-ranked Sharpe where ten
# configurations exist, the spread between them, the mean per-fold Sharpe, and the number of
# folds in which the top-ranked configuration has positive Sharpe. These are measurements that
# feed the downstream narrative.
#
# ### The selection rule
#
# **A backtest's full strategy specification is the signal method, the allocation method and the
# risk overlay taken together.** Naming all three is what makes a validation result and a
# holdout result comparable, because it pins every stage rather than the signal alone.
#
# Each case study's selected configuration is the highest-Sharpe validation backtest across
# those three pipeline stages. The deployed holdout configuration is that same specification,
# retrained on holdout data. When the holdout retrain produces no usable backtest at it -
# degenerate predictions, a vol window that does not match the history available, a universe
# filter that rejects the sample, or another generation failure - the rule falls back to the
# next-highest validation Sharpe that does have a usable holdout, and so on until one succeeds.
# The helper implementing that walk feeds the holdout query and the lineage resolver, which pin
# each validation and holdout pair to one specification.

# %% [markdown]
# ### Two restrictions, and why the selection needs both
#
# **A label restriction**, so the cluster diagnostics and the Chapter 20 holdout retrain rank the
# same thing. sp500_options trains a hold-to-maturity label with coherent option costs alongside
# four fixed-horizon straddle labels priced through the vectorized path with a generic
# basis-point cost. The Chapter 20 narrative uses the hold-to-maturity label as its
# option-strategy reference, so restricting the cluster diagnostics to that same label keeps the
# §20.1 top-cluster numbers aligned with the §20.5 and §20.6 narrative.
#
# Those two section numbers are right, and they are recorded here because they will not look it.
# Eleven references in this chapter pointed at sections that exist and do not carry what was
# claimed, and a sweep for §20.5 in a chapter-20 notebook now finds this one and sees the same
# shape. It is not the same: §20.5's Table 20.6 carries an sp500_options allocator row, and which
# row that is depends on the label pinned here; §20.6 carries the option cost model, which is the
# hold-to-maturity accounting rather than the basis-point sweep the other four labels get. Both
# targets hold material this restriction decides. Do not retarget it.
#
# **An execution-regime restriction**, because sp500_options is evaluated under the
# O'Donovan-Yu (2025) cost-mitigation cascade, whose three rungs are a naive round trip, full
# hold-to-maturity, and hold-to-maturity restricted to the liquid bottom-spread quintile. The
# registered strategy is the third rung; the second is the demoted variant §18.8 discusses. The
# first two rungs both carry the same universe filter, so filtering on that column
# alone leaves `ORDER BY sharpe DESC LIMIT 1` free to pick whichever of the two happens to score
# higher in the current data. Pinning the universe filter *and* the exit rule together is what
# makes the selected row deterministic and coherent with hold-to-maturity. Case studies with no
# entry here skip the filter altogether.

# %%
# The rung pins are imported, not restated. This file used to carry its own copy of both
# predicates and of the dict around them, verbatim, and `paired_metrics.populate_paired_metrics`
# carries the other - both write `backtest_paired_metrics`, so a pin corrected on one side only
# would let one of them overwrite the other's rows with a differently-selected lineage. The
# duplication is how that drift happens, and the mirror keys beside each predicate
# (`universe_filter`, `exit_at_max_days`, `label`) exist for the SQL paths and `progression(...)`
# calls that cannot take a polars expression.
from case_studies.utils.paired_metrics import RUNG_PINS as _CLUSTER_RUNG_RESTRICTIONS  # noqa: E402
from case_studies.utils.strategy_analysis import (  # noqa: E402
    LABEL_RESTRICTIONS as _CLUSTER_LABEL_RESTRICTIONS,
)
from case_studies.utils.strategy_analysis import (  # noqa: E402
    NoSelectableCandidates,
    resolve_solvent_carrier,
    selectable_validation_candidates,
)


@cache
def _canonical_carrier(cs: str) -> dict | None:
    """The configuration this case study reports, from the resolver that decides it.

    This notebook built the same cross-stage rank-1 by hand in four places - concatenating
    `explorer.best` over signal, allocation and risk_overlay, dropping benchmark families,
    applying `LABEL_RESTRICTIONS` and `RUNG_PINS`, and taking the highest Sharpe. Three of
    them wanted the winner and read it from here; the fourth, `_val_rank1_carrier`, walks the
    whole field and takes it from `selectable_validation_candidates`, which is the same
    ranking one step earlier. That is not the ranking the case studies report. `resolve_canonical_rank1_lineage` re-ranks the field
    on exact common timestamp support whenever a conformal candidate is in it, because a
    conformal allocator abstains until it is calibrated and books zeros over the abstention,
    and it applies `UNIVERSE_RESTRICTIONS` and `CARRIER_PINS` besides. Measured 2026-09-18
    against the nine canonical registries, the two rankings named different configurations on
    two case studies: `fx_pairs` (`linear/ridge_a1000000.0` at Sharpe 0.4121
    against `deep_learning/lstm_h64` at comparison Sharpe 0.3091) and
    `nasdaq100_microstructure` (`deep_learning/nlinear` on fwd_ret_15m at 2.3001 against
    `gbm/default_multiclass` on fwd_dir_15m at 2.4159). `spine_prediction_hash` is what
    `05_portfolio_allocation` and Figure 20.7 pin their allocator comparison to, so on
    `fx_pairs` the chapter compared allocators on a configuration the case study does not
    report.

    Returns None where the resolver finds nothing selectable, which is the state the
    hand-built rankings reported as an empty frame. An insolvent or mis-calibrated carrier
    still raises: it is a sweep to fix, not a case study to skip.
    """
    try:
        return resolve_solvent_carrier(cs)
    except NoSelectableCandidates:
        return None


def _retired(cs: str) -> frozenset[str]:
    """Identities a later generation retired, from the same helper `populate_paired_metrics`
    uses. Both write `backtest_paired_metrics`, so a disagreement here would let a Chapter 20
    run overwrite the corrected pairs with a retired lineage."""
    from case_studies.utils.paired_metrics import _retired_prediction_hashes

    return _retired_prediction_hashes(cs)


@cache
def _live_predictions(cs: str) -> list[str] | None:
    """What this case study currently publishes, or None when it declares no populations.

    Membership, not the complement of retirement. A prediction no population ever listed has
    not been retired by anyone, so ranking over "everything not retired" admits experimental
    results the case study never published; ranking over the members in force does not.

    Applied inside the query rather than to its result, because `best()` applies its SQL
    `LIMIT top_n` first - a row filtered afterwards has already consumed a slot and can hide a
    live candidate below the cut.
    """
    from case_studies.research.population import published_members_at

    published = published_members_at(get_case_study_dir(cs), member_kind="prediction")
    if published is None:
        return None
    if not published:
        # `best()` tests this argument for truthiness, so an empty list would read as "no
        # filter" and rank everything. A study that declares populations and publishes
        # nothing has nothing to report, which is a refusal rather than a wide-open ranking.
        raise RuntimeError(f"{cs} declares populations but publishes no prediction identities")
    return sorted(published)


def _best_live(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
    """`explorer.best` narrowed to what the case study still publishes."""
    return explorer.best(stage=stage, top_n=top_n, prediction_hashes=_live_predictions(cs))


def _best_pinned(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
    """`explorer.best` for a stage, fetching enough rows that a rung-restricted
    cohort survives the post-hoc predicate filter.

    `best()` extracts `universe_filter` from `spec_json` in Python, after the
    SQL `LIMIT top_n`. For nasdaq the pinned cost-feasible carrier sits below
    the full-universe in-sample maxima, so a small `top_n` truncates it before
    `_apply_rung_restriction` runs. Pull all rows for restricted case studies."""
    live = _live_predictions(cs)
    if cs in _CLUSTER_RUNG_RESTRICTIONS:
        return explorer.best(stage=stage, top_n=1_000_000, prediction_hashes=live)
    return explorer.best(stage=stage, top_n=top_n, prediction_hashes=live)


def _apply_rung_restriction(df: pl.DataFrame, cs: str) -> pl.DataFrame:
    """Filter `df` to the case study's pinned rung, if one is configured.

    Returns the input untouched if no restriction applies. The helper
    relies on `BacktestExplorer.best()` always emitting both
    `universe_filter` and `exit_at_max_days` columns; if a future
    schema regression drops them, the polars `filter` will raise a
    column-not-found error rather than silently allowing the rank-1
    selection to drift back to the cross-rung max."""
    rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
    if rung is None or df.is_empty():
        return df
    return df.filter(rung["predicate"])


# There is no selected configuration pin here, and there is no mechanism for one.
# `_CARRIER_PIN_PREDICATES` held `"us_firm_characteristics": pl.col("config_name") ==
# "default_huber"` until 2026-08-25, copied from `case_studies.utils.strategy_analysis.CARRIER_PINS`
# and translated into a config-name predicate, under a "keep in sync" comment doing the job a
# mechanism should.
#
# It had not been in sync for a rebuild. Against the current registry `default_huber` is the
# WEAKEST of the ten configs that reached the allocation stage (48 validation backtests, best
# Sharpe 2.128, 2.075 average - tenth of ten), while the documented rule selects `leaves_63_mse`
# (59 backtests, 3.116). So this restricted one case study to its worst advanced configuration
# while every notebook inside that case study reported its best.
#
# Worse than the hash pin removed from `CARRIER_PINS` the same day, because a hash pin dies
# loudly: every hash changes when a sweep is rebuilt, so it resolves to nothing and stops. A
# config-name predicate survives the rebuild and keeps selecting, silently and wrongly.
#
# The mapping stayed empty behind an `_apply_carrier_pin` that could no longer fire, which is a
# second implementation of a rule nothing applied. A selected configuration restriction needed here
# again is `carrier_pins.carrier_config_name(cs)`, which resolves an owner's pin to its config
# through the registry - the thing the copy existed to avoid, and the thing that would have failed
# loudly rather than filtering to the wrong config.


def _progression_for(
    explorer: "BacktestExplorer",
    pred_hash: str,
    cs: str,
) -> pl.DataFrame:
    """Call `progression()` with the case study's rung scope, if any."""
    rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
    if rung is None:
        return explorer.progression(pred_hash)
    return explorer.progression(
        pred_hash,
        universe_filter=rung["universe_filter"],
        exit_at_max_days=rung["exit_at_max_days"],
    )


# Stages whose registry numbers should never be reported for a case study,
# either because the strategy makes the stage structurally meaningless (HTM
# short-straddle has no allocator choice, no bps cost sweep) or because the
# legacy registry contains deprecated entries that pre-date the strategy
# redesign. Consumed by both `build_backtest_rows` and the `synthesis_dict`
# sanitizer below so the in-notebook attrition funnel and the JSON artifact
# cannot drift on the same case study.
_STAGES_NOT_APPLICABLE: dict[str, set[str]] = {
    "sp500_options": {"costs", "risk"},
    "us_firm_characteristics": {"risk"},
    "nasdaq100_microstructure": {"allocation", "costs", "risk"},
}

# Per-CS rationale strings for `not_applicable_reason` fields written into
# `synthesis_dict`. Keyed by (cs, stage) so two case studies that skip the
# same stage for different structural reasons render different prose.
_STAGE_NA_REASONS: dict[tuple[str, str], str] = {
    ("sp500_options", "allocation"): ("HTM short-straddle has fixed 1/n_roll cohort weighting"),
    ("sp500_options", "costs"): ("option costs use §18.8 bid-ask accounting, not bps sweep"),
    ("sp500_options", "risk"): ("HTM expiration structure sets risk profile"),
    ("us_firm_characteristics", "risk"): (
        "vectorized-engine path; portfolio overlays purged 2026-05-17"
    ),
    ("nasdaq100_microstructure", "allocation"): (
        "carrier is a signal-stage slot strategy; the slot mechanism is the sizing rule"
    ),
    ("nasdaq100_microstructure", "costs"): (
        "timing-corrected broad carrier cost grid deferred to v3.1"
    ),
    ("nasdaq100_microstructure", "risk"): (
        "timing-corrected broad carrier risk grid deferred to v3.1"
    ),
}


def _stage_applicable(cs: str, stage: str) -> bool:
    """Return False if `cs` has stage `stage` declared not applicable.

    `stage` must be one of the canonical labels used by
    ``_STAGES_NOT_APPLICABLE`` itself (`allocation`, `costs`, `risk`).
    Callers in this notebook always pass canonical literals."""
    return stage not in _STAGES_NOT_APPLICABLE.get(cs, set())


# %%
cluster_rows = []
for cs, explorer in explorers.items():
    top = _best_pinned(explorer, cs, "signal", 200)
    if top.is_empty() or top["sharpe"][0] is None:
        continue
    if "family" in top.columns:
        top = top.filter(pl.col("family") != "benchmark")
    label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
    if label_restriction and "label" in top.columns:
        top = top.filter(pl.col("label").is_in(list(label_restriction)))
    top = _apply_rung_restriction(top, cs)
    if top.is_empty() or top["sharpe"][0] is None:
        continue
    rank1 = top["sharpe"][0]
    rank10 = top["sharpe"][9] if top.height >= 10 else None
    spread = (rank1 - rank10) if rank10 is not None else None
    # Fold-level stability for the rank-1 backtest
    try:
        bt_hash = top["backtest_hash"][0]
        fold_df = explorer.fold_performance(bt_hash)
        if not fold_df.is_empty():
            mean_fold_sh = float(fold_df["sharpe"].mean())
            se_fold_sh = (
                float(fold_df["sharpe"].std(ddof=1) / (fold_df.height**0.5))
                if fold_df.height > 1
                else None
            )
            n_folds_pos = int(fold_df.filter(pl.col("sharpe") > 0).height)
            n_folds = fold_df.height
        else:
            mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0
    except Exception:
        mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0

    cluster_rows.append(
        {
            "case_study": DISPLAY_NAMES.get(cs, cs),
            "cs_id": cs,
            "rank1_sharpe": rank1,
            "rank10_sharpe": rank10,
            "rank1_rank10_spread": spread,
            "mean_per_fold_sharpe": mean_fold_sh,
            "fold_sharpe_se": se_fold_sh,
            "n_folds_pos": n_folds_pos,
            "n_folds": n_folds,
            "n_configs": int(top.height),
        }
    )

cluster_df = pl.DataFrame(cluster_rows)
if not cluster_df.is_empty():
    print("\n=== Rank-1 Cluster Diagnostics (validation) ===")
    print(
        cluster_df.select(
            "case_study",
            "rank1_sharpe",
            "rank10_sharpe",
            "rank1_rank10_spread",
            "mean_per_fold_sharpe",
            "fold_sharpe_se",
            "n_folds_pos",
            "n_folds",
            "n_configs",
        )
    )

# %% [markdown]
# Read the table as a tuple: (rank1 Sharpe, rank10 Sharpe, spread, fold-SE,
# folds-positive). A small rank1→rank10 spread relative to the fold-SE signals
# a thick top-of-distribution. Folds-positive close to the total fold count
# signals temporal stability. Both can be read off per case study without
# collapsing the evidence into a single label.

# %% [markdown]
# ## Overview Table
#
# The nine case studies differ in asset class, rebalancing frequency, universe
# size, and the cost assumption each one carries. The table records those four
# properties so that a later result can be attributed to the setting it was
# measured in rather than to the model that produced it.

# %%
overview_rows = []
for cs, explorer in explorers.items():
    setup = configs.get(cs, {})
    cost_bps = compute_cost_bps(setup)
    universe = setup.get("universe", {})
    # A futures universe is sized in products rather than in assets, so
    # cme_futures declares `n_products` where the others declare `n_assets`.
    # Reading only the latter reported that case study as an empty universe.
    n_assets = (
        universe.get("n_assets")
        or universe.get("n_products")
        or len(universe.get("symbols", []))
        or 0
    )
    primary_label = setup.get("labels", {}).get("primary", "")

    families = explorer.compare_families(stage="signal")

    overview_rows.append(
        {
            "case_study": DISPLAY_NAMES.get(cs, cs),
            "cs_id": cs,
            "asset_class": ASSET_CLASS_MAP.get(cs, "unknown"),
            "frequency": FREQ_MAP.get(cs, "daily"),
            "universe": n_assets,
            "primary_label": primary_label,
            "cost_bps": cost_bps,
            "n_model_families": len(families) if not families.is_empty() else 0,
        }
    )

overview_df = pl.DataFrame(overview_rows)
overview_df.select("case_study", "asset_class", "frequency", "universe", "cost_bps")

# %% [markdown]
# The test bed covers equity ETFs, crypto perpetuals, intraday microstructure, equity plus
# options, firm characteristics, FX, futures, pure options, and a broad equity panel. The
# `cost_bps` column of the table above records the transaction-cost assumption each one
# carries, so a later result can be read against the cost regime it was measured under.

# %% [markdown]
# ## Model IC Comparison
#
# Mean IC by model family across case studies, queried from `prediction_metrics`.
# Each cell shows the average IC across all configurations within a family,
# filtered to each case study's primary label.

# %%
ic_rows = []
for cs, explorer in explorers.items():
    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    if not db_path.exists():
        continue

    # Filter by primary label so IC values match book prose
    primary_label = configs.get(cs, {}).get("labels", {}).get("primary", "")

    db = sqlite3.connect(str(db_path))
    # Best IC per family on primary label only
    # Exclude causal_dml: it estimates treatment effects, not predictive IC.
    # NOTE: best_ic and best_ic_daily are independent per-family MAXes — they may
    # come from *different* predictions. This is intentional ("best daily IC in
    # the family"), not "the daily IC of the best-by-fold model".
    query = """
        SELECT t.family, MAX(pm.ic_mean) AS best_ic,
               MAX(pm.ic_mean_daily) AS best_ic_daily,
               AVG(pm.ic_mean) AS mean_ic, COUNT(*) AS n_preds
        FROM training_runs t
        JOIN prediction_sets p ON t.training_hash = p.training_hash
        JOIN prediction_metrics pm ON p.prediction_hash = pm.prediction_hash
        WHERE p.split != 'holdout'
          AND pm.ic_mean IS NOT NULL
          AND t.family != 'causal_dml'
    """
    params: tuple = ()
    if primary_label:
        query += "      AND t.label = ?\n"
        params = (primary_label,)
    query += "    GROUP BY t.family"

    rows = db.execute(query, params).fetchall()
    db.close()

    for family, best_ic, best_ic_daily, mean_ic, n_preds in rows:
        ic_rows.append(
            {
                "case_study": DISPLAY_NAMES.get(cs, cs),
                "family": family,
                "ic_mean": mean_ic,
                "ic_best": best_ic,
                "ic_best_daily": best_ic_daily,
                "n_predictions": n_preds,
            }
        )

ic_df = pl.DataFrame(ic_rows)

# %%
if not ic_df.is_empty():
    ic_pivot = ic_df.pivot(on="family", index="case_study", values="ic_mean").sort("case_study")
else:
    ic_pivot = pl.DataFrame()

# %% [markdown]
# ### Mean IC by Model Family

# %%
ic_pivot

# %% [markdown]
# Each cell is the mean of `ic_mean` over that family's non-holdout prediction sets, taken at
# the primary label the case study declares in `setup.yaml`, with `causal_dml` excluded because
# its runs are not fit to predict. Families are comparable within a row, since the label is
# fixed across the row; they are not comparable across rows, because each case study declares a
# different primary label, from a fifteen-minute forward return to a twenty-one-day one, and
# prices a different instrument. A negative mean IC marks a case study where prediction is
# difficult under that label rather than a defect in the family. §20.3 carries this table as
# Table 20.4, and §20.6 works through how an option case study's raw IC translates into Sharpe
# once single-name execution costs are charged.

# %% [markdown]
# ## Backtest Comparison
#
# Cross-dataset comparison of pipeline outcomes. Each row takes the highest-Sharpe result at
# each stage **independently**, so the signal that tops one column may come from a different
# model than the allocation that tops the next.


# %%
def build_backtest_rows():
    """Build backtest comparison rows from all case study explorers."""
    bt_rows = []
    for cs, explorer in explorers.items():
        summary = explorer.summary()

        # Best signal-stage result — exclude benchmark families (equal_weight,
        # etc.) since §20.4's model comparison is about trained models, not
        # passive baselines. Also apply case-study label and universe-filter
        # restrictions so the Ch20 rank-1 is HTM-coherent for sp500_options
        # and pinned to the Rung-3 liquid subset, which is what
        # `strategy_analysis.UNIVERSE_RESTRICTIONS` holds ({"sp500_options":
        # "liquid"}). The Rung-2 full universe is retained only for the §18.8
        # cascade comparison and never anchors the deployed carrier.
        label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
        signal_candidates = _best_pinned(explorer, cs, "signal", 200)
        if not signal_candidates.is_empty() and "family" in signal_candidates.columns:
            signal_candidates = signal_candidates.filter(pl.col("family") != "benchmark")
        if (
            label_restriction
            and "label" in signal_candidates.columns
            and not signal_candidates.is_empty()
        ):
            signal_candidates = signal_candidates.filter(
                pl.col("label").is_in(list(label_restriction))
            )
        signal_candidates = _apply_rung_restriction(signal_candidates, cs)
        best_signal = signal_candidates.head(1)
        signal_sharpe = best_signal["sharpe"][0] if not best_signal.is_empty() else None
        best_source = best_signal["source"][0] if not best_signal.is_empty() else ""

        # Carrier-pred pin for cost/risk. Case studies with a rung restriction
        # (nasdaq cost-feasible ensemble) carry their headline cost/risk on the
        # selected prediction only; the full-universe sweep rows are the
        # Ch18/Ch19 cost-defeat demonstration and must not pool into the
        # cross-case comparison. Other case studies pass None (no pin) and keep
        # the registry-wide aggregation unchanged.
        carrier_pred = (
            best_signal["prediction_hash"][0]
            if cs in _CLUSTER_RUNG_RESTRICTIONS and not best_signal.is_empty()
            else None
        )

        # Best allocation-stage result (same filters). For case studies that
        # declare the allocation stage not applicable (e.g. sp500_options HTM),
        # the registry numbers come from deprecated runs, so report None to
        # match the synthesis_dict sanitizer below.
        if _stage_applicable(cs, "allocation"):
            alloc_candidates = _best_live(explorer, cs, "allocation", 200)
            if not alloc_candidates.is_empty() and "family" in alloc_candidates.columns:
                alloc_candidates = alloc_candidates.filter(pl.col("family") != "benchmark")
            if (
                label_restriction
                and "label" in alloc_candidates.columns
                and not alloc_candidates.is_empty()
            ):
                alloc_candidates = alloc_candidates.filter(
                    pl.col("label").is_in(list(label_restriction))
                )
            alloc_candidates = _apply_rung_restriction(alloc_candidates, cs)
            best_alloc = alloc_candidates.head(1)
            alloc_sharpe = best_alloc["sharpe"][0] if not best_alloc.is_empty() else None
            # The allocator that produced `alloc_sharpe`, read from that row's own spec, so
            # the name and the number describe one configuration. See
            # `strategy_analysis.allocation_method_of`.
            best_allocator = allocation_method_of(
                cs, best_alloc["backtest_hash"][0] if not best_alloc.is_empty() else None
            )
        else:
            alloc_sharpe = None
            best_allocator = ""

        # Cost sensitivity (gated by stage policy)
        survives_costs = None
        if _stage_applicable(cs, "costs"):
            cost_df = explorer.cost_sensitivity(prediction_hash=carrier_pred)
            if not cost_df.is_empty():
                zero_cost = cost_df.filter(pl.col("cost_bps") == 0)
                survives_costs = not zero_cost.is_empty() and zero_cost["sharpe"].max() > 0

        # Risk overlay (gated by stage policy)
        best_overlay = ""
        managed_sharpe = None
        if _stage_applicable(cs, "risk"):
            risk_df = explorer.risk_impact(prediction_hash=carrier_pred)
            if not risk_df.is_empty():
                # rank_one, not a one-key sort: overlays that never trigger book the
                # baseline Sharpe exactly, so ties at the top are ordinary here and a
                # one-key sort would let frame order decide the name reported below.
                best_risk_row = rank_one(risk_df, by="sharpe", name="risk_name")
                best_overlay = best_risk_row["risk_name"][0]
                managed_sharpe = best_risk_row["sharpe"][0]

        # Spine rank-1 prediction_hash - the configuration the case study reports, taken
        # from `_canonical_carrier` rather than ranked a second time here. Figure 20.7 and
        # `05_portfolio_allocation` both read this value, and Ch20 prose Tables 20.5-20.7
        # quote the resolver, so the two have to be one answer.
        _carrier = _canonical_carrier(cs)
        spine_pred_hash = _carrier["val_prediction_hash"] if _carrier else None

        bt_rows.append(
            {
                "case_study": DISPLAY_NAMES.get(cs, cs),
                "case_study_id": cs,
                "spine_prediction_hash": spine_pred_hash,
                "n_signal": summary.get("signal", 0),
                "n_allocation": summary.get("allocation", 0),
                "n_cost": summary.get("cost_sensitivity", 0),
                "n_risk": summary.get("risk_overlay", 0),
                "best_source": best_source,
                "signal_sharpe": signal_sharpe,
                "best_allocator": best_allocator,
                "alloc_sharpe": alloc_sharpe,
                "survives_costs": survives_costs,
                "best_overlay": best_overlay,
                "managed_sharpe": managed_sharpe,
            }
        )
    return bt_rows


# %%
bt_rows = build_backtest_rows()

# %%
bt_df = pl.DataFrame(bt_rows)
print("\nPipeline Comparison:")
print(
    bt_df.select(
        "case_study",
        "signal_sharpe",
        "alloc_sharpe",
        "survives_costs",
        "managed_sharpe",
    )
)

# %% [markdown]
# Read the baseline column of the table above for how many case studies enter the pipeline with
# a positive baseline-stage Sharpe and which do not. Those counts move whenever a registry is
# rebuilt, which is why they are in the table rather than in this sentence.
#
# The lineage table below traces each selected prediction across the stages in the order the
# backtests run: baseline, allocation, risk overlay, then the cost sweep charged against
# whatever survived. A Sharpe that rises from one column to the next is what that stage added,
# and only where the later stage carries the earlier one's configuration - the paired rows above
# say which transitions meet that test. NASDAQ-100 is excluded from that comparison in v3.0
# because its timing-corrected broad cost and risk grids are deferred to v3.1.

# %% [markdown]
# ## Paired-Bootstrap Comparison vs Equal-Weight Benchmark
#
# Each case study's selected baseline-stage backtest, under the same label, universe-filter and
# rung restrictions used for the cluster diagnostics, is compared to its equal-weight benchmark
# using a **paired stationary block bootstrap on daily strategy returns**. Block length is derived from
# ``setup.yaml.labels.{label}.rebalance_step`` (falling back to the optimal
# block size, never below the label horizon). Reported quantities:
#
# - ``sharpe_diff`` with a bootstrap confidence interval
# - ``ret_diff``, the annualized return difference, with its confidence interval
# - ``info_ratio`` of the daily-return difference
# - ``prob_challenger_wins`` — bootstrap fraction in which challenger Sharpe
#   exceeds the benchmark
# - ``p_value`` — two-sided bootstrap p-value for ``sharpe_diff = 0``
#
# Results land in ``backtest_paired_metrics`` (per case study) and roll up
# into the cross-dataset table below. Intervals are at the conventional confidence level the
# bootstrap call sets. This is the right unit of uncertainty for the headline claim about the
# selected configuration: the Sharpe **difference against the passive baseline that experienced
# the same market conditions**, rather than the Sharpe alone.


# %%
def _benchmark_returns_from_artifact(
    cs: str, label: str, period: str = "overall"
) -> tuple[str, pl.DataFrame, str] | None:
    """Resolve the side-artifact equal-weight benchmark for ``(cs, label)``.

    The benchmark is the daily-MTM EW reference series persisted by
    ``scripts/compute_vectorized_ew_benchmark.py`` at
    ``case_studies/{cs}/benchmark/{label}.parquet``. Single, well-defined
    methodology per (cs, label) — no universe/rung/cadence ambiguity that
    a registry-side ``family='benchmark'`` lookup would have to disambiguate.

    ``period`` selects the time window slice (``"overall"`` or ``"holdout"``)
    applied by ``load_benchmark_returns``. Classification-label fallback to
    the matching ``fwd_ret_*`` artifact applies in both periods.

    Returns ``(synthetic_hash, returns_df, resolved_label)`` where
    ``synthetic_hash`` is a deterministic identifier safe to use as the PK
    column in ``backtest_paired_metrics`` (which has no FK on
    ``benchmark_hash``) and ``resolved_label`` is the actual label whose
    artifact was loaded — equal to ``label`` unless the classification
    fallback fired, in which case it's the matching ``fwd_ret_*`` label.
    Returns ``None`` if the artifact is missing.
    """
    df = load_benchmark_returns(cs, label, period=period)
    bench_label = label
    if df.is_empty() or "ew_return" not in df.columns:
        # Fallback: classification labels (``fwd_class_*``, ``fwd_dir_*``,
        # ``fwd_tb_*``, ``fwd_carry_*``) share the same forecast window as
        # their continuous counterpart (``fwd_ret_*``). The EW universe over
        # the same window is identical regardless of the label being
        # predicted, so map e.g. ``fwd_class_1m`` to ``fwd_ret_1m``.
        fallback = None
        for prefix in ("fwd_class_", "fwd_dir_", "fwd_tb_", "fwd_carry_"):
            if label.startswith(prefix):
                fallback = "fwd_ret_" + label[len(prefix) :]
                break
        if fallback is None:
            return None
        df = load_benchmark_returns(cs, fallback, period=period)
        if df.is_empty() or "ew_return" not in df.columns:
            return None
        bench_label = fallback
    suffix = "" if period == "overall" else f":{period}"
    bench_hash = f"side_ew:{cs}:{bench_label}{suffix}"
    return (
        bench_hash,
        df.select(
            pl.col("timestamp").cast(pl.Date).alias("timestamp"),
            pl.col("ew_return").cast(pl.Float64).alias("ret"),
        ),
        bench_label,
    )


def _aligned_returns(cs: str, h: str) -> pl.DataFrame | None:
    """Load and normalize a backtest's daily returns; columns ``[timestamp, ret]``."""
    parquet = get_case_study_dir(cs) / "run_log" / "backtest" / h / "daily_returns.parquet"
    if not parquet.exists():
        return None
    df = pl.read_parquet(parquet)
    ret_col = next(
        (c for c in ("daily_return", "ret", "return", "value") if c in df.columns),
        df.columns[-1],
    )
    ts_col = next(
        (c for c in ("timestamp", "date", "datetime") if c in df.columns),
        df.columns[0],
    )
    return df.select(
        pl.col(ts_col).cast(pl.Date).alias("timestamp"),
        pl.col(ret_col).cast(pl.Float64).alias("ret"),
    )


# %%
import numpy as np

from case_studies.utils.uncertainty import (
    SIGNAL_BASELINE_BY_CASE_STUDY,
    STAGE_SEQUENCE,
    compute_independent_diff_uncertainty,
    compute_paired_uncertainty,
    descends_from,
    joint_returns,
)


def _min_paired_n(ppy: int) -> int:
    """Minimum series length for paired-bootstrap stability, frequency-aware.

    The ~21 floor was written for daily cadences (about a month of obs).
    Monthly case studies (e.g. ``us_firm_characteristics``) have ~12 holdout
    observations by design, and ``compute_paired_uncertainty`` runs cleanly
    on n=12. Scale the floor with ``ppy`` so monthly/weekly CSs aren't
    blocked by a daily-tuned guard.
    """
    if ppy <= 12:  # monthly
        return 6
    if ppy <= 52:  # weekly
        return 12
    return 21  # daily / 8h / intraday


# Distinguish skipped CSs from real failures so empty cross-dataset rollups
# aren't indistinguishable from a silent crash.
paired_rows: list[dict] = []
paired_skips: list[dict] = []

for cs, explorer in explorers.items():
    # The leader is the configuration the case study reports, not a ranking rebuilt here.
    # `_canonical_carrier` documents why the two are not the same ordering.
    carrier = _canonical_carrier(cs)
    if carrier is None:
        paired_skips.append({"case_study": cs, "reason": "no_selectable_candidates"})
        continue
    leader_hash = carrier["val_backtest_hash"]
    leader_label = carrier["label"]
    if not leader_label:
        paired_skips.append({"case_study": cs, "reason": "no_label_on_leader"})
        continue
    bench_resolution = _benchmark_returns_from_artifact(cs, leader_label)
    if not bench_resolution:
        paired_skips.append(
            {"case_study": cs, "reason": f"no_benchmark_artifact_for_label:{leader_label}"}
        )
        continue
    benchmark_hash, base, resolved_bench_label = bench_resolution

    chal = _aligned_returns(cs, leader_hash)
    if chal is None:
        paired_skips.append({"case_study": cs, "reason": "no_challenger_returns_parquet"})
        continue

    ppy = {"daily": 252, "weekly": 52, "monthly": 12, "8h": 1095}.get(
        FREQ_MAP.get(cs, "daily"), 252
    )
    min_n = _min_paired_n(ppy)
    aligned = chal.join(base, on="timestamp", how="inner", suffix="_b")
    if aligned.height < min_n:
        paired_skips.append(
            {"case_study": cs, "reason": f"insufficient_overlap:n={aligned.height}"}
        )
        continue
    # A strategy against a benchmark: the leader's leading flat run is warmup before its
    # first signal, not a position it held, so the sample starts where both are trading.
    c_arr, b_arr = joint_returns(aligned["ret"].to_numpy(), aligned["ret_b"].to_numpy())
    if c_arr.size < min_n:
        paired_skips.append(
            {"case_study": cs, "reason": f"insufficient_after_coerce:n={c_arr.size}"}
        )
        continue
    paired = compute_paired_uncertainty(
        c_arr,
        b_arr,
        periods_per_year=ppy,
        case_study=cs,
        label=leader_label,
        n_boot=2000,
        seed=42,
    )
    if not paired:
        paired_skips.append({"case_study": cs, "reason": "paired_uncertainty_empty"})
        continue

    # Side-artifact benchmark — deterministic across (cs, label), no
    # universe/rung ambiguity, no fallback-by-recency.
    benchmark_kind = f"{SIGNAL_BASELINE_BY_CASE_STUDY.get(cs, 'equal_weight')}_side_artifact"
    paired_rows.append(
        {
            "case_study": DISPLAY_NAMES.get(cs, cs),
            "label": leader_label,
            "benchmark_label": resolved_bench_label,  # may differ from leader_label when classification fallback fired
            "sharpe_diff": paired.get("sharpe_diff"),
            "sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
            "sharpe_diff_ci_hi": paired.get("sharpe_diff_ci95_hi"),
            "ret_diff": paired.get("ret_diff"),
            "info_ratio": paired.get("info_ratio"),
            "p_value": paired.get("p_value"),
            "prob_wins": paired.get("prob_challenger_wins"),
            "block": paired.get("bootstrap_block_length"),
            "n_boot": paired.get("bootstrap_n"),
        }
    )

paired_df = pl.DataFrame(paired_rows)
if not paired_df.is_empty():
    print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
    print(
        paired_df.select(
            "case_study",
            "label",
            "benchmark_label",
            "sharpe_diff",
            "sharpe_diff_ci_lo",
            "sharpe_diff_ci_hi",
            "info_ratio",
            "prob_wins",
            "p_value",
        )
    )
else:
    print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
    print("No paired-bootstrap rows produced — see skip table below for reasons.")

if paired_skips:
    print("\nSkipped case studies:")
    for s in paired_skips:
        print(f"  - {s['case_study']:<32}  {s['reason']}")

# Loud invariant — a 0/N or all-skipped outcome is now obvious in the
# notebook output instead of buried under "no paired-bootstrap rows."
print(f"\npaired={len(paired_rows)}/{len(explorers)}, skipped={len(paired_skips)}/{len(explorers)}")

# %% [markdown]
# Read each row as the selected challenger's annualized Sharpe **minus** the equal-weight
# benchmark's, with a confidence interval from the paired stationary block bootstrap on the
# daily-return difference; the information ratio summarizes the
# excess-return-to-tracking-error ratio; ``prob_wins`` is the fraction of
# bootstrap resamples in which the challenger beat the benchmark; ``p_value``
# tests ``H0: sharpe_diff = 0``. A confident "the model adds skill over the
# passive baseline" claim requires (i) the CI excludes zero, (ii) ``prob_wins``
# close to 1, and (iii) a small ``p_value``. Cases where the CI straddles
# zero are not failures — they signal that the apparent Sharpe gap is within
# block-bootstrap sampling error and should be reported as such.

# %% [markdown]
# ## Paired metrics — full coverage for strategy-analysis notebook
#
# The block above populates the first pair type, the selected baseline signal against
# equal-weight over the whole window. The strategy-analysis notebook (per-CS strategy notebooks)
# requires five additional pair types per case study to render §2 (stage-
# transition waterfall), §6 (holdout decay + holdout-vs-benchmark) and §7
# (benchmark-aware diagnostics) without inline bootstrap recomputation.
#
# The pair set:
#
# 1. selected signal (overall) ↔ equal-weight (overall) — populated above
# 2. selected signal (holdout) ↔ equal-weight (holdout window)
# 3. the selected configuration on holdout ↔ the same configuration on
#    validation (same lineage decay; min-length truncation since the
#    windows are disjoint)
# 4-6. one pair per consecutive stage transition the prediction actually has,
#    in ``STAGE_SEQUENCE`` order: allocation ↔ signal, risk-overlay ↔
#    allocation, cost-sensitivity ↔ risk-overlay. A case study that did not
#    run a stage yields fewer pairs, and a stage that does not carry the
#    previous stage's configuration yields none for that transition - the two
#    were selected independently and their difference is not a stage effect.
#
# Pair #3 truncates both series to ``min(len(val), len(ho))`` to satisfy
# ``compute_paired_uncertainty``'s equal-length precondition. The CI is
# interpreted as bootstrap resampling Sharpe in each window independently
# and taking the difference; the truncation is preserved in the
# ``benchmark_kind`` value (``val_rank1_self`` always carries the truncation
# caveat). All pairs use the same paired stationary block bootstrap helper.


# %%
def _full_strategy_spec_from_backtest(db: sqlite3.Connection, bt_hash: str) -> dict | None:
    """Pull the full strategy spec dict (signal + allocation + risk) from
    `bt_hash`'s spec_json. Returns None if the row is missing or signal has
    no `method` field.

    A backtest's full specification is the tuple (signal, allocation, risk). Pinning
    the val→holdout pair on this full spec keeps the comparison apples-to-
    apples; pinning on signal alone allows MAX(sharpe) to surface a holdout
    row with a different allocation (e.g. conformal_weighted) or risk overlay
    than the validation rank-1 carrier.
    """
    row = db.execute(
        "SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
        (bt_hash,),
    ).fetchone()
    if not row:
        return None
    strat = json.loads(row[0]).get("strategy", {})
    sig = strat.get("signal", {})
    if not sig.get("method"):
        return None
    alloc = strat.get("allocation") or {}
    risk = strat.get("risk") or {}
    return {
        "signal": {
            "method": sig.get("method"),
            "top_k": sig.get("top_k"),
            "percentile": sig.get("percentile"),
        },
        "allocation": {
            "method": alloc.get("method"),
            "top_k": alloc.get("top_k"),
            "long_short": alloc.get("long_short"),
        },
        "risk": {
            "name": risk.get("name"),
        },
    }


def _val_rank1_carrier(cs: str) -> dict | None:
    """Return ``{'spec', 'prediction_hash'}`` for ``cs``'s validation rank-1 carrier.

    The prediction hash is carried out alongside the spec because the holdout resolver
    needs it: naming all three stages pins the configuration AND the checkpoint, and without
    it a case study that registered several checkpoints against one strategy is ambiguous
    and the resolver refuses. It was determinable all along - this walk had it in hand and
    threw it away - so refusing there would have dropped a case study out of the
    reader-facing holdout table for want of a value one line above.

    The val rank-1 *full strategy* spec for ``cs`` — the
    highest-Sharpe validation backtest across (signal, allocation,
    risk_overlay) stages — walking candidates by val Sharpe descending until
    one with a matching holdout backtest at the SAME full spec is found.

    Implements the selection rule documented in §20.1: the deployed
    holdout for each case study is the val rank-1 across all three pipeline
    stages, retrained on holdout data; when retrain produces no usable
    holdout at that full spec (degenerate predictions, vol-window mismatch,
    universe filter rejection, etc.) the walk falls through to the next
    candidate by val Sharpe. The first val candidate with a registered
    holdout backtest at the same (signal, allocation, risk) tuple defines
    the apples-to-apples carrier pair.

    Returns None when no val candidate up to rank ~200 has a matching
    holdout under the case study's label / rung restrictions.
    """
    explorer = explorers.get(cs)
    if explorer is None:
        return None
    # The field the resolver ranks, in the resolver's order, rather than a concat of
    # `explorer.best` re-filtered here: the walk starts at the carrier `_canonical_carrier`
    # names and falls through in the same order the resolver would. Every row is kept - no
    # dedup by prediction_hash - because when the rank-1 configuration has no matching holdout
    # retrain but a same-prediction lower-Sharpe variant (a different allocator or risk
    # overlay) does, a dedup would jump to a different prediction instead of accepting the
    # same-prediction variant as the apples-to-apples match.
    try:
        candidates = selectable_validation_candidates(cs)
    except NoSelectableCandidates:
        # The helper raises on an empty pool rather than returning one, and this walk's
        # callers read `None` as "no holdout pair for this case study" - the state the
        # hand-built ranking reported as an empty frame. A pool with nothing eligible in it
        # is that state, not a reason to stop aggregating the other eight.
        return None
    label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)

    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
    db = sqlite3.connect(str(db_path))
    try:
        for candidate in candidates[:200]:
            bt_hash = candidate["backtest_hash"]
            spec = _full_strategy_spec_from_backtest(db, bt_hash)
            if spec is None:
                continue
            # The probe below asks whether THIS candidate has a holdout, so it matches the
            # candidate's own configuration and checkpoint and not only its strategy spec.
            #
            # Matching the spec alone made the walk stop at a candidate whose own checkpoint
            # had no holdout whenever a sibling checkpoint had one at the same spec. The
            # resolver, handed that configuration, then finds nothing for it - and the walk has
            # already stopped, so the case study reports no holdout while one exists for a
            # later candidate. Advancing instead is what makes the fall-through the resolver
            # no longer performs unnecessary rather than merely forbidden.
            carrier_row = db.execute(
                """
                SELECT t.family, t.config_name, t.label,
                       p.checkpoint_value, p.checkpoint_kind
                FROM prediction_sets p
                JOIN training_runs t ON t.training_hash = p.training_hash
                WHERE p.prediction_hash = ?
                """,
                (candidate["prediction_hash"],),
            ).fetchone()
            if carrier_row is None:
                continue
            spec_clauses, spec_params = _full_strategy_clauses(spec)
            ho_clauses = ["p.split = 'holdout'"] + spec_clauses
            if _retired(cs):
                ho_clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
            ho_params: list[object] = list(spec_params)
            if _retired(cs):
                ho_params.append(json.dumps(sorted(_retired(cs))))
            if label_restriction:
                placeholders = ",".join("?" for _ in label_restriction)
                ho_clauses.append(f"t.label IN ({placeholders})")
                ho_params.extend(sorted(label_restriction))
            if rung is not None:
                ho_clauses.append(
                    "COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
                )
                ho_params.append(rung["universe_filter"])
                if rung["exit_at_max_days"] is None:
                    ho_clauses.append(
                        "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
                    )
                else:
                    ho_clauses.append(
                        "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?"
                    )
                    ho_params.append(rung["exit_at_max_days"])
            # `t.spec_json` rather than `1`, and no LIMIT: the probe has to apply the same
            # eligibility test the resolver applies, and that test is not expressible in SQL.
            #
            # A model fitted on the validation folds can publish predictions over the holdout
            # window, so `p.split = 'holdout'` with a non-null Sharpe is not enough to make a
            # row a holdout result. The resolver drops those through
            # `training_run_fitted_for_the_holdout`; a probe that admitted them would stop the
            # walk at a candidate whose only holdout is validation-fitted, the resolver would
            # then find nothing eligible for it, and the case study would report no holdout
            # while a later candidate had a real one.
            probe_rows = db.execute(
                f"""
                SELECT t.spec_json FROM prediction_sets p
                JOIN training_runs t ON p.training_hash = t.training_hash
                JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
                                     AND b.stage IN ('signal','allocation','risk_overlay','holdout')
                JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
                WHERE {" AND ".join(ho_clauses)}
                  AND t.family = ?
                  AND t.config_name = ?
                  AND t.label = ?
                  AND p.checkpoint_value IS ?
                  AND p.checkpoint_kind IS ?
                  AND bm.sharpe IS NOT NULL
                """,
                ho_params + list(carrier_row),
            ).fetchall()
            row = any(training_run_fitted_for_the_holdout(probe[0]) for probe in probe_rows)
            if row:
                return {"spec": spec, "prediction_hash": candidate["prediction_hash"]}
    finally:
        db.close()
    return None


def _full_strategy_clauses(spec: dict | None) -> tuple[list[str], list[object]]:
    """Build SQL WHERE clauses + params that pin a backtest row to the full
    strategy spec (signal + allocation + risk). Empty list returned when
    spec is None (no constraint).

    Pinning on the full spec ensures `MAX(sharpe)` over candidate holdout
    backtests cannot surface a different allocator (e.g. conformal_weighted
    when val rank-1 was score_weighted) or a different risk overlay than
    the validation carrier — the val→holdout pair stays apples-to-apples
    on the full pipelin

Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT

Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.