Перейти к содержимому
Все документы библиотеки

Парный бутстрэп для сравнения этапов отбора в бэктестах

Код Machine Learning for Trading

Сводка

В документе описан процессор парных сравнений результатов на разных этапах отбора стратегий. Сравниваются отобранные сигналы с эталонами с равными весами, выборки отложенного периода с валидационными, а также переходы между этапами распределения капитала, анализа чувствительности к издержкам и применения защитных правил. В большинстве случаев используется парный стационарный блочный бутстрэп, сохраняющий взаимосвязь согласованных рядов доходности. Для сравнения валидации и отложенного периода, где наблюдения не перекрываются, выполняется независимая выборка по окнам; полученный интервал измеряет разницу между этими окнами, но не свидетельствует об устойчивом преимуществе стратегии.

Процессор также учитывает практические ограничения отбора: специфичные для случая ограничения меток или ступеней стратегии, минимальную длину выборки с учётом частоты, отсутствие артефактов доходности и удаление устаревших строк реестра после полного запуска. Эти детали обеспечивают единообразную отчётность в разных исследованиях, но сам модуль служит инфраструктурой для оценки неопределённости, а не торговой стратегией или эмпирическим доказательством превосходства какого-либо метода отбора. Доверительные интервалы бутстрэпа зависят от доступной истории доходности и схемы сравнения; непересекающиеся периоды не делают оценки результатов статистически независимыми.

Ключевые идеи

  • Парный стационарный блочный бутстрэп использует согласованные наблюдения доходности для оценки неопределённости разницы результатов.
  • Для непересекающихся окон валидации и отложенного периода нужны независимые выборки; интервал измеряет разницу между окнами.
  • Сравнения отслеживают изменение результатов на этапах отбора сигнала, распределения капитала, учёта издержек и применения защитных правил.
  • Минимальные пороги размера выборки с учётом частоты наблюдений позволяют учесть различное число данных в исследованиях.
  • Результаты бутстрэпа описывают неопределённость при заданных выборке и схеме анализа, но не доказывают устойчивое торговое преимущество.

Теги

Полный текст
# paired_metrics.py


```py
"""Compute-and-register the per-case-study ``backtest_paired_metrics`` table.

In-repo home of the paired-bootstrap producer that the strategy-analysis
notebook (``NN_strategy_analysis.py``) runs so a reader who never touches
Chapter 20 still lands a populated ``backtest_paired_metrics`` table. It renders
§2 (stage-transition waterfall), §6 (holdout decay + holdout-vs-benchmark) and
§7 (benchmark-aware diagnostics) without inline bootstrap recomputation.

The logic is extracted verbatim from the two producer loops in
``20_strategy_synthesis/01_aggregate_synthesis.py``; the only change is that the
per-case-study selection config (label restriction, rung pin, carrier pin,
frequency) is passed in as arguments rather than read from Ch20 module globals,
so a single case study can run standalone. The Chapter-20 aggregate can later
call this function per case study and become a pure reader.

Six pair types are produced per case study:

1. signal rank-1 (overall) ↔ equal-weight (overall)
2. cross-stage rank-1 holdout ↔ equal-weight (holdout window)
3. holdout rank-1 ↔ validation rank-1 of the same lineage (disjoint windows)
4. allocation rank-1 ↔ signal rank-1 (same window, stage transition)
5. cost-sensitivity rank-1 ↔ allocation rank-1 (same window)
6. risk-overlay rank-1 ↔ cost-sensitivity rank-1 (same window)

All pairs use the paired stationary block bootstrap
(``compute_paired_uncertainty``); pair #3 uses independent per-window draws
(``compute_independent_diff_uncertainty``) because its two windows share no
observations, so there is no difference series to pair on. Disjointness removes the
pairing; it does not make the two Sharpes independent, and the interval that comes
back is calibrated for the gap between those two windows rather than for the
strategy having one edge across both. See that function for the measurement.
"""

from __future__ import annotations

import json
import sqlite3
from functools import cache
from pathlib import Path

import numpy as np
import polars as pl

from case_studies.utils.analytics import DISPLAY_NAMES
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_returns
from case_studies.utils.notebook_contracts import degenerate_prediction_hashes
from case_studies.utils.registry.registration import register_paired_metrics
from case_studies.utils.strategy_analysis import (
    is_refit_of,
    training_run_fitted_for_the_holdout,
)
from case_studies.utils.uncertainty import (
    SIGNAL_BASELINE_BY_CASE_STUDY,
    STAGE_SEQUENCE,
    CarrierScope,
    EntireRegistry,
    NoCarrier,
    PredictionScope,
    compute_independent_diff_uncertainty,
    compute_paired_uncertainty,
    descends_from,
    joint_returns,
)
from utils.paths import get_case_study_dir

# Cross-stage rank-1 pooling stages - mirrors strategy_analysis.SELECTION_STAGES.
_PAIRED_STAGES = ("signal", "allocation", "risk_overlay")


def _min_paired_n(ppy: int) -> int:
    """Minimum series length for paired-bootstrap stability, frequency-aware.

    The ~21 floor was written for daily cadences (about a month of obs).
    Monthly case studies (e.g. ``us_firm_characteristics``) have ~12 holdout
    observations by design, and ``compute_paired_uncertainty`` runs cleanly
    on n=12. Scale the floor with ``ppy`` so monthly/weekly CSs aren't
    blocked by a daily-tuned guard.
    """
    if ppy <= 12:  # monthly
        return 6
    if ppy <= 52:  # weekly
        return 12
    return 21  # daily / 8h / intraday


# The two case studies whose canonical strategy is pinned to one rung of a cascade, mirroring
# `20_strategy_synthesis/01_aggregate_synthesis.py::_CLUSTER_RUNG_RESTRICTIONS`. Held here as
# well because a case study's own strategy-analysis notebook has to make the same selection
# without importing a chapter, and `tests/test_rung_pins_match_chapter_20.py` fails if the two
# definitions drift.
#
# sp500_options: rung-1 (mid-to-mid bps) and rung-2 (full-universe HTM) both carry
# `universe_filter="full"`, so filtering on the universe alone leaves `ORDER BY sharpe DESC
# LIMIT 1` free to pick whichever rung is higher in current data. The pin combines the universe
# with `exit_at_max_days` so the rank-1 row is deterministic and HTM-coherent.
#
# nasdaq100_microstructure: the cost-feasible sweep on the primary label, matched on design
# attributes any registry can satisfy and chosen before the holdout was opened.
#
# The pin used to name `family == "ensemble"` as well. The mean-forecast ensemble existed
# because the per-model baseline on this case study was not worth reporting, and that is no
# longer the case: measured 2026-09-14 on the cost-feasible pool, `deep_learning/nlinear` on
# `fwd_ret_15m` reaches +2.300 against the ensemble's +0.566, so pinning to the ensemble
# anchored every paired comparison to the weakest thing in the pool. The family clause is
# gone; the ensemble rows stay in the registry and stay selectable, they are simply no
# longer the only thing the pin can choose.
#
# The label is part of the pin and was not always, and it does more work now that family is
# not. The pool spans four declared labels, so universe alone leaves `ORDER BY sharpe DESC
# LIMIT 1` free to choose among them - `fwd_dir_15m` reaches +2.416, above the primary
# label's best - and the rank-1 rung would move onto a label this case study is not featured
# on with nothing announcing it. A pin that omits a dimension selects along it silently.
#
# `fwd_ret_15m` is what the book prints for this case study, in Table 11.6
# (`NASDAQ-100 15m | fwd_ret_15m`), in Chapter 12's case-study table (`15 minutes | forward
# return`) and in Chapter 13's (`15 minutes | NLinear`). `config/setup.yaml::labels.primary`
# says the same, and `tests/test_rung_pin_label.py` asserts the two do not drift apart.
RUNG_PINS: dict[str, dict] = {
    "sp500_options": {
        "predicate": (pl.col("universe_filter") == "liquid") & pl.col("exit_at_max_days").is_null(),
        "universe_filter": "liquid",
        "exit_at_max_days": None,
    },
    "nasdaq100_microstructure": {
        "predicate": (pl.col("universe_filter") == "cost_feasible")
        & (pl.col("label") == "fwd_ret_15m"),
        "universe_filter": "cost_feasible",
        "exit_at_max_days": None,
        # Mirrors the predicate for the SQL paths that cannot take a polars expression.
        "label": "fwd_ret_15m",
    },
}


def rung_for(cs: str) -> dict | None:
    """The rung this case study is pinned to, or None where it is not pinned."""
    return RUNG_PINS.get(cs)


def _best_for_rung(
    explorer: BacktestExplorer,
    stage: str,
    rung: dict | None,
    top_n: int = 2000,
    prediction_hashes: list[str] | None = None,
) -> pl.DataFrame:
    """``explorer.best`` for a stage, fetching enough rows that the pin survives.

    ``best()`` reads ``universe_filter`` out of ``spec_json`` in Python, and truncates to
    ``top_n`` after that; ``_apply_rung_restriction`` runs later still. For nasdaq the pinned
    cost-feasible carrier sits below the full-universe in-sample maxima, so a small ``top_n``
    truncates it before the predicate is ever applied and the pin silently selects nothing -
    or, worse, the best surviving row that was never the carrier.

    A pinned case study therefore asks for every row, which ``best`` now spells ``top_n=0``.
    It was a literal million until ``sweep_config.top_n_cap`` gave 0 that meaning, and a cohort
    past a million would have been truncated rather than refused.

    Ch20 solves this with the same widening (`_best_pinned`); the extraction into this module
    dropped it, which is why every pinned selection here has to go through this helper rather
    than call ``explorer.best`` directly.
    """
    return explorer.best(
        stage=stage,
        top_n=0 if rung is not None else top_n,
        prediction_hashes=prediction_hashes,
    )


def _apply_rung_restriction(df: pl.DataFrame, rung: dict | None) -> pl.DataFrame:
    """Filter ``df`` to the case study's pinned rung, if one is configured.

    Returns the input untouched if ``rung`` is None. The helper relies on
    ``BacktestExplorer.best()`` always emitting both ``universe_filter`` and
    ``exit_at_max_days`` columns; if a future schema regression drops them, the
    polars ``filter`` raises column-not-found rather than silently letting the
    rank-1 selection drift back to the cross-rung max.
    """
    if rung is None or df.is_empty():
        return df
    return df.filter(rung["predicate"])


def _benchmark_returns_from_artifact(
    cs: str, label: str, period: str = "overall"
) -> tuple[str, pl.DataFrame, str] | None:
    """Resolve the side-artifact equal-weight benchmark for ``(cs, label)``.

    The benchmark is the daily-MTM EW reference series persisted by
    ``scripts/compute_vectorized_ew_benchmark.py`` at
    ``case_studies/{cs}/benchmark/{label}.parquet``. Single, well-defined
    methodology per (cs, label). ``period`` selects the window slice
    (``"overall"`` or ``"holdout"``). Classification-label fallback to the
    matching ``fwd_ret_*`` artifact applies in both periods.

    Returns ``(synthetic_hash, returns_df, resolved_label)`` or ``None`` when
    the artifact is missing. ``synthetic_hash`` is a deterministic identifier
    safe as the PK column in ``backtest_paired_metrics`` (no FK on
    ``benchmark_hash``).
    """
    df = load_benchmark_returns(cs, label, period=period)
    bench_label = label
    if df.is_empty() or "ew_return" not in df.columns:
        fallback = None
        for prefix in ("fwd_class_", "fwd_dir_", "fwd_tb_", "fwd_carry_"):
            if label.startswith(prefix):
                fallback = "fwd_ret_" + label[len(prefix) :]
                break
        if fallback is None:
            return None
        df = load_benchmark_returns(cs, fallback, period=period)
        if df.is_empty() or "ew_return" not in df.columns:
            return None
        bench_label = fallback
    suffix = "" if period == "overall" else f":{period}"
    bench_hash = f"side_ew:{cs}:{bench_label}{suffix}"
    return (
        bench_hash,
        df.select(
            pl.col("timestamp").cast(pl.Date).alias("timestamp"),
            pl.col("ew_return").cast(pl.Float64).alias("ret"),
        ),
        bench_label,
    )


def _aligned_returns(cs: str, h: str) -> pl.DataFrame | None:
    """Load and normalize a backtest's daily returns; columns ``[timestamp, ret]``."""
    parquet = get_case_study_dir(cs) / "run_log" / "backtest" / h / "daily_returns.parquet"
    if not parquet.exists():
        return None
    df = pl.read_parquet(parquet)
    ret_col = next(
        (c for c in ("daily_return", "ret", "return", "value") if c in df.columns),
        df.columns[-1],
    )
    ts_col = next(
        (c for c in ("timestamp", "date", "datetime") if c in df.columns),
        df.columns[0],
    )
    return df.select(
        pl.col(ts_col).cast(pl.Date).alias("timestamp"),
        pl.col(ret_col).cast(pl.Float64).alias("ret"),
    )


def _full_strategy_spec_from_backtest(db: sqlite3.Connection, bt_hash: str) -> dict | None:
    """Pull the full strategy spec dict (signal + allocation + risk) from
    ``bt_hash``'s spec_json. Returns None if the row is missing or signal has
    no ``method`` field.

    The carrier of a backtest is the tuple (signal, allocation, risk). Pinning
    the val→holdout pair on this full spec keeps the comparison apples-to-
    apples; pinning on signal alone allows MAX(sharpe) to surface a holdout
    row with a different allocation or risk overlay than the validation rank-1.
    """
    row = db.execute(
        "SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
        (bt_hash,),
    ).fetchone()
    if not row:
        return None
    strat = json.loads(row[0]).get("strategy", {})
    sig = strat.get("signal", {})
    if not sig.get("method"):
        return None
    alloc = strat.get("allocation") or {}
    risk = strat.get("risk") or {}
    return {
        "signal": {
            "method": sig.get("method"),
            "top_k": sig.get("top_k"),
            "percentile": sig.get("percentile"),
        },
        "allocation": {
            "method": alloc.get("method"),
            "top_k": alloc.get("top_k"),
            "long_short": alloc.get("long_short"),
        },
        "risk": {
            "name": risk.get("name"),
        },
    }


def _full_strategy_clauses(spec: dict | None) -> tuple[list[str], list[object]]:
    """Build SQL WHERE clauses + params pinning a backtest row to the full
    strategy spec (signal + allocation + risk). Empty list when spec is None.

    Pinning on the full spec ensures ``MAX(sharpe)`` over candidate holdout
    backtests cannot surface a different allocator or risk overlay than the
    validation carrier — the val→holdout pair stays apples-to-apples on the
    full pipeline configuration, not just the signal.
    """
    if not spec:
        return [], []
    clauses: list[str] = []
    params: list[object] = []

    sig = spec.get("signal") or {}
    method = sig.get("method")
    if method is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.signal.method') IS NULL")
    else:
        clauses.append("json_extract(b.spec_json, '$.strategy.signal.method') = ?")
        params.append(method)
    top_k = sig.get("top_k")
    if top_k is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.signal.top_k') IS NULL")
    else:
        clauses.append("CAST(json_extract(b.spec_json, '$.strategy.signal.top_k') AS INTEGER) = ?")
        params.append(int(top_k))
    pct = sig.get("percentile")
    if pct is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.signal.percentile') IS NULL")
    else:
        clauses.append(
            "CAST(json_extract(b.spec_json, '$.strategy.signal.percentile') AS REAL) = ?"
        )
        params.append(float(pct))

    alloc = spec.get("allocation") or {}
    am = alloc.get("method")
    if am is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.allocation.method') IS NULL")
    else:
        clauses.append("json_extract(b.spec_json, '$.strategy.allocation.method') = ?")
        params.append(am)
    ak = alloc.get("top_k")
    if ak is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.allocation.top_k') IS NULL")
    else:
        clauses.append(
            "CAST(json_extract(b.spec_json, '$.strategy.allocation.top_k') AS INTEGER) = ?"
        )
        params.append(int(ak))
    als = alloc.get("long_short")
    if als is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.allocation.long_short') IS NULL")
    else:
        clauses.append(
            "CAST(json_extract(b.spec_json, '$.strategy.allocation.long_short') AS INTEGER) = ?"
        )
        params.append(int(bool(als)))

    risk = spec.get("risk") or {}
    risk_name = risk.get("name")
    if risk_name is None:
        clauses.append("json_extract(b.spec_json, '$.strategy.risk.name') IS NULL")
    else:
        clauses.append("json_extract(b.spec_json, '$.strategy.risk.name') = ?")
        params.append(risk_name)

    return clauses, params


@cache
def _retired_prediction_hashes(cs: str) -> frozenset[str]:
    """Every prediction identity a later generation retired, on either split.

    Retirement is recorded member-wise on the *validation* population, and a prediction
    identity includes its split, so a retired validation hash never equals its holdout
    sibling and cannot filter it directly. What identifies the same model state across the
    two is the training run together with the checkpoint it was scored at, so this expands
    the recorded set along that key.

    Member-wise, not run-wise: a population that moved one checkpoint of a run and kept
    another has retired one checkpoint, and the sibling that did not move is still current.

    ``superseded_members_at`` rather than ``superseded_members``, because the latter takes a
    ``Study`` and every ``Study.open`` branch ends in ``activate()``, which clears
    ``ML4T_OUTPUT_DIR`` and would re-point a preview or isolated workspace at the released
    registry - answering for a different registry than these queries read.
    """
    from case_studies.research.population import superseded_members_at

    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    if not db_path.exists():
        return frozenset()
    recorded = superseded_members_at(case_dir, member_kind="prediction")
    if not recorded:
        return frozenset()

    db = sqlite3.connect(str(db_path))
    try:
        rows = db.execute(
            "SELECT prediction_hash, training_hash, checkpoint_kind, checkpoint_value "
            "FROM prediction_sets"
        ).fetchall()
    finally:
        db.close()
    retired_states = {
        (training_hash, checkpoint_kind, checkpoint_value)
        for prediction_hash, training_hash, checkpoint_kind, checkpoint_value in rows
        if prediction_hash in recorded
    }
    return frozenset(
        prediction_hash
        for prediction_hash, training_hash, checkpoint_kind, checkpoint_value in rows
        if (training_hash, checkpoint_kind, checkpoint_value) in retired_states
    )


def _val_rank1_carrier(
    cs: str,
    explorer: BacktestExplorer,
    *,
    label_restriction: frozenset[str] | None,
    rung: dict | None,
    prediction_hashes: list[str] | None = None,
    retired_hashes: frozenset[str] | None = None,
) -> dict | None:
    """Return the val rank-1 *full strategy* spec for ``cs`` — the
    highest-Sharpe validation backtest across (signal, allocation,
    risk_overlay) stages — walking candidates by val Sharpe descending until
    one with a matching holdout backtest at the SAME full spec is found.

    Returns None when no val candidate up to rank ~200 has a matching holdout
    under the case study's label / rung restrictions.
    """
    cand = pl.concat(
        [
            _best_for_rung(explorer, s, rung, prediction_hashes=prediction_hashes)
            for s in ("signal", "allocation", "risk_overlay")
        ],
        how="diagonal_relaxed",
    )
    if cand.is_empty() or "backtest_hash" not in cand.columns:
        return None
    cand = _eligible_candidates(cs, cand, label_restriction=label_restriction, rung=rung)
    if cand.is_empty():
        return None
    # Do NOT dedup by prediction_hash here — the walk needs every registered
    # (signal, allocation, risk_overlay) tuple so a same-prediction lower-sharpe
    # variant can serve as the apples-to-apples carrier.
    cand = cand.sort("sharpe", descending=True)

    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    db = sqlite3.connect(str(db_path))
    try:
        for i in range(min(cand.height, 200)):
            bt_hash = cand["backtest_hash"][i]
            spec = _full_strategy_spec_from_backtest(db, bt_hash)
            if spec is None:
                continue
            spec_clauses, spec_params = _full_strategy_clauses(spec)
            ho_clauses = ["p.split = 'holdout'"] + spec_clauses
            ho_params: list[object] = list(spec_params)
            # The same exclusion `_holdout_lineage_for` applies. Without it a retired holdout
            # makes this candidate look eligible, the walk stops here, and that call then
            # filters the row out and returns nothing - instead of advancing to the next live
            # candidate that does have a holdout.
            if retired_hashes:
                ho_clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
                ho_params.append(json.dumps(sorted(retired_hashes)))
            if label_restriction:
                placeholders = ",".join("?" for _ in label_restriction)
                ho_clauses.append(f"t.label IN ({placeholders})")
                ho_params.extend(sorted(label_restriction))
            if rung is not None:
                ho_clauses.append(
                    "COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
                )
                ho_params.append(rung["universe_filter"])
                if rung["exit_at_max_days"] is None:
                    ho_clauses.append(
                        "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
                    )
                else:
                    ho_clauses.append(
                        "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?"
                    )
                    ho_params.append(rung["exit_at_max_days"])
            # The probe asks whether THIS candidate has a holdout, so it matches the
            # candidate's own configuration and checkpoint, not only its strategy spec, and
            # applies the same eligibility test `_holdout_lineage_for` applies.
            #
            # Spec alone stopped the walk wherever a SIBLING checkpoint had a holdout at the
            # same spec. The caller then pinned this candidate, the resolver found nothing of
            # its own, and the case study reported no holdout although a later candidate had
            # one. `training_run_fitted_for_the_holdout` is the other half: a model fitted on
            # the validation folds publishes over the holdout window, so split plus a non-null
            # Sharpe does not make a row a holdout result, and it is not expressible in SQL.
            carrier_row = db.execute(
                """
                SELECT t.family, t.config_name, t.label,
                       p.checkpoint_value, p.checkpoint_kind
                FROM prediction_sets p
                JOIN training_runs t ON t.training_hash = p.training_hash
                WHERE p.prediction_hash = ?
                """,
                (cand["prediction_hash"][i],),
            ).fetchone()
            if carrier_row is None:
                continue
            probe_rows = db.execute(
                f"""
                SELECT t.spec_json FROM prediction_sets p
                JOIN training_runs t ON p.training_hash = t.training_hash
                JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
                                     AND b.stage IN ('signal','allocation','risk_overlay','holdout')
                JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
                WHERE {" AND ".join(ho_clauses)}
                  AND t.family = ?
                  AND t.config_name = ?
                  AND t.label = ?
                  AND p.checkpoint_value IS ?
                  AND p.checkpoint_kind IS ?
                  AND bm.sharpe IS NOT NULL
                """,
                ho_params + list(carrier_row),
            ).fetchall()
            if any(training_run_fitted_for_the_holdout(probe[0]) for probe in probe_rows):
                return {"spec": spec, "prediction_hash": cand["prediction_hash"][i]}
    finally:
        db.close()
    return None


def _holdout_lineage_for(
    cs: str,
    leader_label: str,
    strategy_spec: dict | None = None,
    *,
    label_restriction: frozenset[str] | None,
    rung: dict | None,
    prefer_prediction_hash: str | None = None,
    retired_hashes: frozenset[str] | None = None,
) -> dict | None:
    """Return ``{backtest_hash, prediction_hash, family, config_name, label}``
    for the highest-Sharpe holdout backtest registered in this case study,
    honoring per-CS cluster restrictions but **not** the leader's label.

    ``leader_label`` is intentionally unused in the SQL — kept for call-site
    symmetry with ``_val_backtest_for_lineage``. The holdout's *own* label is
    returned so callers can pair it against matching benchmarks (the label may
    differ from the validation rank-1 when ``generate_holdout``'s degeneracy
    fallback accepts a candidate on a different label).

    When ``strategy_spec`` is provided, the holdout pick is restricted to
    backtests with the same full (signal, allocation, risk) tuple as val's
    rank-1 carrier, so the val→holdout comparison stays apples-to-apples.
    """
    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    if not db_path.exists():
        return None

    clauses = ["p.split = 'holdout'"]
    params: list[object] = []
    if label_restriction:
        placeholders = ",".join("?" for _ in label_restriction)
        clauses.append(f"t.label IN ({placeholders})")
        params.extend(label_restriction)
    if rung is not None:
        clauses.append(
            "COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
        )
        params.append(rung["universe_filter"])
        if rung["exit_at_max_days"] is None:
            clauses.append(
                "json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
            )
        else:
            clauses.append("json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?")
            params.append(rung["exit_at_max_days"])
    if retired_hashes:
        # ``_retired_prediction_hashes`` carries the recorded validation retirements across
        # to the holdout rows that share a training run and checkpoint, so this filters on the
        # identity the row actually has.
        #
        # It reaches a holdout scored from the same trained model. One produced by the
        # canonical *retrain* registers its own training hash and shares no key with the
        # validation identity that was retired - measured: no holdout training spec in any
        # registry here references a validation identity, so there is nothing to join on.
        #
        # What prevents a retired carrier from reaching a holdout is upstream instead:
        # `strategy_analysis.resolve_canonical_rank1_lineage` ranks over published members
        # only, so a retrain created from here on descends from a live carrier by
        # construction. Holdout rows written
        # before that are stale artifacts of an earlier selection, and they are regenerated.
        clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
        params.append(json.dumps(sorted(retired_hashes)))
    spec_clauses, spec_params = _full_strategy_clauses(strategy_spec)
    clauses.extend(spec_clauses)
    params.extend(spec_params)
    where_sql = " AND ".join(clauses)

    db = sqlite3.connect(str(db_path))
    db.row_factory = sqlite3.Row
    try:
        # Same-lineage preference: given the validation rank-1's own prediction
        # set, prefer a holdout sharing its trained model AND its checkpoint.
        # Both are read from that one row here rather than accepted as separate
        # arguments, because a caller that passes the training hash and forgets
        # the checkpoint reintroduces the defect while looking correct.
        #
        # The checkpoint has to be pinned because a trained model registers one
        # prediction set per declared checkpoint and they share a strategy spec,
        # so the training hash alone leaves one indistinguishable candidate per
        # checkpoint. This branch writes the hash the ``val_rank1_self`` pair is
        # stored under and ``select_holdout_self_backtest`` reads it back, so the
        # two must agree. Ordering by ``backtest_hash`` rather than by Sharpe
        # keeps the choice off holdout performance either way.
        if prefer_prediction_hash is not None:
            carrier = db.execute(
                """
                SELECT t.family, t.config_name, t.label,
                       p.checkpoint_value, p.checkpoint_kind, t.spec_json
                FROM prediction_sets p
                JOIN training_runs t ON t.training_hash = p.training_hash
                WHERE p.prediction_hash = ?
                """,
                (prefer_prediction_hash,),
            ).fetchone()
            if carrier is not None:
                carrier_spec_json = carrier["spec_json"]
                carrier_key = list(carrier)[:5]
                # Matched on the declared configuration, not on the carrier's training
                # hash. A holdout prediction produced correctly carries a NEW training
                # identity - the same configuration refitted on the holdout fold - so
                # matching on the validation training hash finds only a holdout scored
                # from the validation-fitted model. ``select_holdout_self_backtest``
                # reads back the hash this branch writes the ``val_rank1_self`` pair
                # under, so the two apply the same rule.
                rows_ = db.execute(
                    f"""
                    SELECT t.family, t.config_name, t.label,
                           p.prediction_hash, b.backtest_hash, t.spec_json
                    FROM prediction_sets p
                    JOIN training_runs t ON p.training_hash = t.training_hash
                    JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
                                         AND b.stage IN
                                             ('signal','allocation','risk_overlay','holdout')
                    JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
                    WHERE {where_sql}
                      AND t.family = ?
                      AND t.config_name = ?
                      AND t.label = ?
                      AND p.checkpoint_value IS ?
                      AND p.checkpoint_kind IS ?
                    ORDER BY b.backtest_hash
                    """,
                    params + carrier_key,
                ).fetchall()
                # Naming the carrier pins the configuration and the checkpoint, and that is
                # normally one candidate. It is not guaranteed to be: one prediction set can
                # carry several backtests - a replay under a different strategy spec, an
                # experimental allocator sharing the holdout prediction - and they survive
                # this filter together. Returning the first in `backtest_hash` order would
                # decide on nothing the carrier determines, which is the same defect the
                # unpinned branch below refuses, so it refuses here too rather than only
                # where the caller happened not to pin.
                # Fitted for the holdout AND a refit of this specification. The four
                # columns the query filters on are a configuration's NAME, and a name is
                # reused across generations - refit a study after its features change and the
                # new runs carry the same family, config_name, label and checkpoint as the
                # old ones. On the current registries fx_pairs has 144 configuration groups
                # spanning more than one feature-artifact generation and etfs has 10, so
                # without the second condition a holdout fitted on features the study no
                # longer publishes can be the sole coarse match and get reported as the
                # carrier's own holdout.
                eligible = [
                    candidate
                    for candidate in rows_
                    if training_run_fitted_for_the_holdout(candidate["spec_json"])
                    and is_refit_of(candidate["spec_json"], carrier_spec_json)
                ]
                if len({candidate["backtest_hash"] for candidate in eligible}) > 1:
                    raise ValueError(
                        f"{len(eligible)} holdout backtests match the pinned carrier "
                        f"{prefer_prediction_hash} for {cs} - same configuration, same "
                        "checkpoint, same strategy. Choosing between them would rank the "
                        "holdout on its own result. Retire the replays that are not this "
                        "study's holdout, so one candidate remains."
                    )
                if eligible:
                    resolved = dict(eligible[0])
                    resolved.pop("spec_json")
                    return resolved
            # A named carrier with no eligible holdout is an ANSWER, not a reason to look
            # elsewhere. Falling through to the unpinned query below made the pin a mere
            # preference: where the selected checkpoint had no holdout and a sibling
            # checkpoint did, the fallback returned the sibling's, and the reader-facing
            # table then reported a checkpoint that validation never selected. There is no
            # weaker sense in which that is the carrier's holdout.
            #
            # `carrier is None` lands here too, and for the same reason: the caller named a
            # prediction set the registry does not have, and "some other holdout" is not a
            # better answer to that than none.
            return None

        # The caller named no carrier, so this is the unpinned fallback. Two filters, and
        # neither is optional.
        #
        # Only runs actually refitted for the holdout are eligible: a model fitted on the
        # validation folds can publish predictions over the holdout window, and it is not a
        # holdout result whatever its Sharpe.
        #
        # And what survives has to be ONE candidate. Anything that reaches here cannot be
        # separated on what the registry records, so picking among them means picking by
        # holdout Sharpe - choosing the evaluation by its own result, which is the one thing
        # this module must never do. The refusal is the answer; the caller resolves it by
        # naming the validation carrier, not by this function guessing.
        rows = db.execute(
            f"""
            SELECT DISTINCT t.family, t.config_name, t.label,
                   p.prediction_hash, b.backtest_hash, p.training_hash, t.spec_json
            FROM prediction_sets p
            JOIN training_runs t ON p.training_hash = t.training_hash
            JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
                                 AND b.stage IN ('signal','allocation','risk_overlay','holdout')
            JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
            WHERE {where_sql}
            ORDER BY b.backtest_hash
            """,
            params,
        ).fetchall()
    finally:
        db.close()
    rows = [row for row in rows if training_run_fitted_for_the_holdout(row["spec_json"])]
    if not rows:
        return None
    # Ambiguity is ambiguity however it arises, and it arises three ways: several trained
    # models, several checkpoints of one trained model (one prediction set each, sharing a
    # strategy spec), and several backtests hanging off one prediction set. Only the first
    # used to refuse; the other two fell through to `rows[0]` under `ORDER BY b.backtest_hash`,
    # which is arbitrary with respect to the configuration and therefore picks on the holdout's
    # own result as surely as ordering by Sharpe would.
    if len({row["backtest_hash"] for row in rows}) > 1:
        lineages = {row["training_hash"] for row in rows}
        raise ValueError(
            f"{len(rows)} holdout backtests across {len(lineages)} trained model(s) were "
            f"refitted for the holdout and match the carrier spec for {cs}. Choosing between "
            "them would rank the holdout on its own result. Pass the validation rank-1's "
            "prediction hash as `prefer_prediction_hash`, which pins the configuration and the "
            "checkpoint, so the holdout is resolved from the carrier that was selected rather "
            "than from the holdout scores."
        )
    row = rows[0]
    return {
        k: row[k] for k in ("family", "config_name", "label", "prediction_hash", "backtest_hash")
    }


def _val_backtest_for_lineage(
    cs: str,
    family: str,
    config_name: str,
    label: str,
    *,
    prediction_hashes: list[str] | None = None,
) -> str | None:
    """Return the highest-Sharpe validation signal-stage backtest_hash for the
    given (family, config_name, label) lineage, or None if absent.

    Used by ``val_rank1_self`` pair construction so the comparison stays
    *within* a lineage when the holdout retrain came from a fallback candidate.
    """
    case_dir = get_case_study_dir(cs)
    db_path = case_dir / "run_log" / "registry.db"
    if not db_path.exists():
        return None
    db = sqlite3.connect(str(db_path))
    try:
        row = db.execute(
            """
            SELECT b.backtest_hash
            FROM prediction_sets p
            JOIN training_runs t ON p.training_hash = t.training_hash
            JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
                                 AND b.stage = 'signal'
            JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
            WHERE p.split = 'validation'
              AND t.family = ?
              AND t.config_name = ?
              AND t.label = ?
              {scope}
            ORDER BY bm.sharpe DESC NULLS LAST
            LIMIT 1
            """.format(
                scope=(
                    " AND p.prediction_hash IN (SELECT value FROM json_each(?))"
                    if prediction_hashes is not None
                    else ""
                )
            ),
            (family, config_name, label)
            + ((json.dumps(list(prediction_hashes)),) if prediction_hashes is not None else ()),
        ).fetchone()
    finally:
        db.close()
    return row[0] if row else None


def _populate_pair(
    cs,
    challenger_hash,
    benchmark_hash,
    benchmark_kind,
    challenger_returns,
    benchmark_returns,
    ppy,
    label,
    *,
    disjoint_windows: bool = False,
    challenger_overlays_baseline: bool = False,
    benchmark_label: str | None = None,
    write_case_dir: Path | None = None,
):
    """Compute and register one paired-metric row. Idempotent UPSERT.

    With ``disjoint_windows=True`` (val→holdout decay), each side is bootstrapped
    over its full window and the difference distribution is built from those draws,
    because two windows that share no timestamps leave no difference series to
    resample. That is the absence of a pairing, not independence: see
    :func:`case_studies.utils.uncertainty.compute_independent_diff_uncertainty` for
    what the resulting interval does and does not cover. Otherwise, the streams are
    inner-joined on timestamp and a paired stationary bootstrap runs on the aligned
    diff series.

    ``challenger_overlays_baseline`` says what a leading flat run on the challenger
    means, and the two pair shapes here answer differently. Against the equal-weight
    benchmark the challenger is an independent strategy whose returns begin at the
    first bar rather than at its first signal, so those rows are warmup and the
    default drops them. Across a stage transition the challenger is built on top of
    the baseline and both are live from the same session, so a flat challenger there
    is a position it chose to hold and the comparison keeps it. See
    :func:`case_studies.utils.uncertainty.joint_returns`.
    """
    min_n = _min_paired_n(ppy)
    if disjoint_windows:
        c_arr = challenger_returns.sort("timestamp")["ret"].to_numpy()
        b_arr = benchmark_returns.sort("timestamp")["ret"].to_numpy()
        finite_c = np.isfinite(c_arr)
        finite_b = np.isfinite(b_arr)
        c_arr, b_arr = c_arr[finite_c], b_arr[finite_b]
        if c_arr.size < min_n or b_arr.size < min_n:
            return {
                "cs": cs,
                "kind": benchmark_kind,
                "label": label,
                "benchmark_label": benchmark_label if benchmark_label is not None else label,
                "skip": f"insufficient_disjoint:n_c={c_arr.size},n_b={b_arr.size}",
            }
        n_overlap = min(c_arr.size, b_arr.size)
        paired = compute_independent_diff_uncertainty(
            c_arr,
            b_arr,
            periods_per_year=ppy,
            case_study=cs,
            label=label,
            n_boot=2000,
            seed=42,
        )
    else:
        aligned = challenger_returns.join(
            benchmark_returns, on="timestamp", how="inner", suffix="_b"
        )
        if aligned.height < min_n:
            return {
                "cs": cs,
                "kind": benchmark_kind,
                "label": label,
                "benchmark_label": benchmark_label if benchmark_label is not None else label,
                "skip": f"insufficient_overlap:n={aligned.height}",
            }
        c_arr = aligned["ret"].to_numpy()
        b_arr = aligned["ret_b"].to_numpy()
        # Measured here, applied once inside `compute_paired_uncertainty`, which trims
        # whatever it is handed. Both rules happen to survive being applied twice, so passing
        # the trimmed pair on would work today; it would stop working silently the moment a
        # rule stops being idempotent, and the figure registered as `n_overlap` would then
        # name a sample the bootstrap never ran on.
        n_overlap = joint_returns(
            c_arr, b_arr, challenger_overlays_baseline=challenger_overlays_baseline
        )[0].size
        if n_overlap < min_n:
            return {
                "cs": cs,
                "kind": benchmark_kind,
                "label": label,
                "benchmark_label": benchmark_label if benchmark_label is not None else label,
                "skip": f"insufficient_after_coerce:n={n_overlap}",
            }
        paired = compute_paired_uncertainty(
            c_arr,
            b_arr,
            periods_per_year=ppy,
            case_study=cs,
            label=label,
            n_boot=2000,
            seed=42,
            challenger_overlays_baseline=challenger_overlays_baseline,
        )

    if not paired:
        return {
            "cs": cs,
            "kind": benchmark_kind,
            "label": label,
            "benchmark_label": benchmark_label if benchmark_label is not None else label,
            "skip": "uncertainty_empty",
        }
    register_paired_metrics(
        cs,
        challenger_hash,
        benchmark_hash,
        paired,
        benchmark_kind=benchmark_kind,
        periods_per_year=ppy,
        case_dir=write_case_dir,
    )
    # disjoint path: paired carries n_c/n_b (post-coerce per-side sizes); use
    # min so n_overlap reflects what the bootstrap actually used. paired path:
    # no n_c/n_b, n_overlap already the post-`joint_returns` length.
    n_actual = n_overlap
    n_c = paired.get("n_c")
    n_b = paired.get("n_b")
    if n_c is not None and n_b is not None:
        n_actual = int(min(float(n_c), float(n_b)))
    return {
        "cs": cs,
        "kind": benchmark_kind,
        "label": label,
        "benchmark_label": benchmark_label if benchmark_label is not None else label,
        "n_overlap": n_actual,
        "sharpe_diff": paired.get("sharpe_diff"),
        "sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
        "sharpe_diff_ci_hi": paired.get("sharpe_diff_ci95_hi"),
        "info_ratio": paired.get("info_ratio"),
        "p_value": paired.get("p_value"),
    }


def _drop_retired_generations(cs: str, cand):
    """Candidates whose own publisher still publishes them.

    Every ranking in this module sorts on `sharpe` over whatever the registry holds, and a
    superseded generation is still complete, still `current` under its schema version, and
    still ranks. Passing a resolved `carrier` fixes only the pairs that consult it; Pair #1
    ranks for itself, so the retired row won there and the validation-side pair was written
    against a backtest the case study no longer publishes - measured on fx_pairs, where the
    live carrier then had no challenger row at all and the strategy-analysis notebook refused
    for want of evidence that had been written under the retired hash.

    Both sides are filtered, because a retired generation reaches a ranking through either.
    The prediction side is the one that hides: a refit that changes no numbers publishes
    identical predictions under a new identity, so old and new carry the same Sharpe to the
    last digit and the sort returns whichever it likes. It goes through
    ``_retired_prediction_hashes`` rather than the recorded set, so a holdout prediction is
    dropped along with the validation generation it was retrained from.
    """
    from case_studies.research.population import superseded_members_at

    if cand is None or cand.is_empty():
        return cand
    if "backtest_hash" in cand.columns:
        retired = superseded_members_at(get_case_study_dir(cs), member_kind="backtest")
        if retired:
            cand = cand.filter(~pl.col("backtest_hash").is_in(list(retired)))
    if "prediction_hash" in cand.columns:
        retired_predictions = _retired_prediction_hashes(cs)
        if retired_predictions:
            cand = cand.filter(~pl.col("prediction_hash").is_in(list(retired_predictions)))
    return cand


def _drop_degenerate_predictions(cs: str, cand):
    """Candidates whose prediction set selection refuses to consider.

    A LASSO or ElasticNet fit that shrinks every coefficient to zero on a fold predicts a
    constant there, so that fold ranks nothing and the pooled IC computed over it is biased
    rather than a model result. ``degenerate_prediction_sql`` states the rule - both limbs of
    it, the NULL IC of an all-tied fold and the denormal one of a fold constant only to display
    precision - and
    ``selectable_validation_candidates`` applies it, which is why the published carrier cannot
    be one of these.

    A pair is the other publication path and had no such filter. The sweep backtests the whole
    declared population rather than a shortlist, so the registry does hold backtests on
    degenerate sets: measured on us_equities_panel 2026-09-14, 15 of its prediction sets are
    degenerate, four signal-stage backtests stand on two of them, and both sat at Sharpe 0.6062
    against a 0.8977 leader - third and fourth in the stage, so ranking alone did not catch it
    and gives no reason to expect it to as the sweep continues.

    Applied wherever ``_drop_retired_generations`` is, for the same reason: every ranking in
    this module sorts on ``sharpe`` over whatever the registry holds.
    """
    if cand is None or cand.is_empty() or "prediction_hash" not in cand.columns:
        return cand
    degenerate = degenerate_prediction_hashes(get_case_study_dir(cs))
    if not degenerate:
        return cand
    return cand.filter(~pl.col("prediction_hash").is_in(list(degenerate)))


def _eligible_candidates(
    cs: str, cand, *, label_restriction: frozenset[str] | None, rung: dict | None
):
    """Every filter a ranking in this module owes its candidate pool, in one place.

    The three rankings here - the carrier walk, pair #1, and the no-carrier leader - had this
    chain written out three times, which is how the degeneracy filter came to be missing from
    all of them while selection had it: a filter added to one copy is not added to the others,
    and nothing reads as wrong at any single site.

    Retirement and degeneracy are both facts about whether the row may be published at all.
    The benchmark exclusion is about what a challenger is. The label and rung restrictions are
    the case study's own scope. What is NOT here is the carrier pin, which pair #1 deliberately
    does not apply.
    """
    cand = _drop_retired_generations(cs, cand)
    cand = _drop_degenerate_predictions(cs, cand)
    if cand is None or cand.is_empty():
        return cand
    if "family" in cand.columns:
        cand = cand.filter(pl.col("family") != "benchmark")
    if label_restriction and "label" in cand.columns:
        cand = cand.filter(pl.col("label").is_in(list(label_restriction)))
    return _apply_rung_restriction(cand, rung)


def populate_paired_metrics(
    cs: str,
    explorer: BacktestExplorer | None = None,
    *,
    label_restriction: frozenset[str] | None = None,
    rung: dict | None = None,
    carrier: CarrierScope,
    periods_per_year: int | None = None,
    verbose: bool = True,
    replace_all: bool,
    write_case_dir: Path | None = None,
    prediction_hashes: PredictionScope,
) -> list[dict]:
    """Compute all six paired-bootstrap pair types for ``cs`` and register them.

    Extracted from the two producer loops in
    ``20_strategy_synthesis/01_aggregate_synthesis.py``, specialized to a single
    case study. The per-CS selection config that Ch20 reads from module globals
    is passed in:

    * ``label_restriction`` - ``strategy_analysis.LABEL_RESTRICTIONS.get(cs)`` (e.g.
      sp500_options → ``frozenset({'ret_to_expiry'})``); None for most CSs.
    * ``rung`` - ``{"predicate", "universe_filter", "exit_at_max_days"}``, plus
      ``label`` where the pin names one, for the rung-pinned CSs (sp500_options,
      nasdaq100_microstructure); None else. A line naming
      ``us_firm_characteristics -> config_name == 'default_huber'`` stood here
      until 2026-09-14, dangling under this bullet and describing a pin deleted on
      2026-08-25 for selecting that case study's weakest advanced configuration
      (`20_strategy_synthesis/01_aggregate_synthesis.py:365`).
    * ``periods_per_year`` — the annualization factor. Defaults to the case
      study's own ``evaluation.periods_per_year`` declaration rather than to a
      cadence, so a caller that omits it gets its own scale instead of someone
      else's.
    * ``carrier`` — a ``resolve_canonical_rank1_lineage`` result, or ``NO_CARRIER``.
      With a lineage, pairs #2-6 use its validation and holdout backtests instead of
      re-ranking the registry here, and pair #1 is pinned to its validation backtest
      too - the code below refuses rather than ranking when a carrier is supplied,
      because a pair #1 registered under a backtest the case study does not report
      leaves its carrier with no validation-to-benchmark evidence. This bullet said
      pair #1 was unaffected until 2026-09-14, which had not been true since that
      refusal landed. ``NO_CARRIER`` keeps the legacy ranking, which is not
      the canonical selection - it orders on raw Sharpe and applies neither the
      common-support re-ranking nor the restrictions the resolver holds - so a caller
      that can resolve the lineage should pass it. The rung-pinned case studies
      (sp500_options, nasdaq100_microstructure) restrict on a dimension the resolver
      does not know, which is why the legacy ranking still exists at all.

    ``replace_all`` makes the call a complete snapshot: pairs it did not write are
    deleted, so a rebuild under a different selection does not leave the previous
    selection's rows behind. Registration alone is an UPSERT keyed on
    ``(challenger_hash, benchmark_hash)``, which cannot remove a row it no longer
    produces, so False is additive: the previous selection's rows survive alongside
    the corrected ones. A call that writes nothing prunes nothing - that is a failed
    rebuild, not an empty snapshot.

    ``periods_per_year`` used to be ``freq: str = "daily"``, resolved through a
    name-to-count map. That default is silently right for the six case studies that
    annualize at 252 and silently wrong for the rest: ``us_firm_characteristics`` is
    monthly, so every Sharpe difference and interval it wrote was scaled by sqrt(252)
    rather than sqrt(12), a factor of 4.58 on numbers a notebook prints as its holdout
    closure. Worse, ``_min_paired_n(252)`` returns 21, so the twelve observation
    holdout pairs were skipped and the table was missing rows with nothing recording
    the omission.

    An integer rather than a cadence name because the name was only ever converted
    back to a number, and the caller holds the number. Reading the declaration by
    default is what stops the next non-252 case study inheriting the wrong scale by
    saying nothing.

    Returns the list of per-pair summary dicts (mirrors the ``paired_rows`` +
    ``extra_paired_rows`` the Ch20 producer builds); each pair is also written to
    ``backtest_paired_metrics`` via ``register_paired_metrics``.

    ``prediction_hashes`` restricts every candidate read to that population, or
    ``ENTIRE_REGISTRY`` for the whole-registry read. A pair is a comparison between two
    strategies the caller reports; selecting either side from the whole registry lets a
    retired generation be the challenger or the benchmark, and the difference is then
    measured against a strategy its own publisher replaced. Detecting those rows and
    rebuilding without this would write them back unchanged.

    ``carrier``, ``replace_all`` and ``prediction_hashes`` carry no default. Each decides
    what the numbers this writes are computed over, and each used to default to the widest
    reading, so omitting one type-checked, ran, and produced plausible rows that were wrong
    exactly when the registry held something the caller does not report - invisible in
    review, in CI and in the output. Requiring them costs one line per call site and makes
    the wide readers greppable. See ``uncertainty.ENTIRE_REGISTRY``.

    ``write_case_dir`` redirects the registry *write* to an alternate case dir
    (reads still come from the live tree) — used by the verification harness to
    write into a temp registry copy non-destructively.
    """
    if explorer is None:
        explorer = BacktestExplorer(cs)
    # Both scopes are stated by the caller and carry no default; the sentinels are
    # normalized here so the rest of the body reads the same as it did when they were
    # `None`. `ENTIRE_REGISTRY` is the whole-registry read, `NO_CARRIER` the raw-Sharpe
    # re-rank. See `uncertainty.ENTIRE_REGISTRY` for why neither is a default.
    live = None if isinstance(prediction_hashes, EntireRegistry) else list(prediction_hashes)
    lineage = None if isinstance(carrier, NoCarrier) else carrier
    if periods_per_year is None:
        from case_studies.utils.uncertainty import periods_per_year_from_setup

        periods_per_year = int(periods_per_year_from_setup(cs))
    ppy = int(periods_per_year)
    rows: list[dict] = []
    written_keys: set[tuple[str, str]] = set()

    def _pair(cs_, challenger_hash, benchmark_hash, *args, **kwargs):
        result = _populate_pair(cs_, challenger_hash, benchmark_hash, *args, **kwargs)
        if "skip" not in result:
            written_keys.add((challenger_hash, benchmark_hash))
        return result

    # -- Pair #1: signal rank-1 (overall) ↔ equal-weight (overall) -----------
    cand = pl.concat(
        [_best_for_rung(explorer, s, rung, prediction_hashes=live) for s in _PAIRED_STAGES],
        how="diagonal_relaxed",
    )
    skip_pair1 = False
    if cand.is_empty() or "backtest_hash" not in cand.columns:
        skip_pair1 = True
    if not skip_pair1:
        # NB: pair #1 (Ch20 Loop A) applies ONLY the rung restriction — no
        # carrier pin — unlike pairs #2-6 (Loop B), which apply both. Preserve
        # that asymmetry so carrier-pinned CSs (us_firm_characteristics) match.
        cand = _eligible_candidates(cs, cand, label_restriction=label_restriction, rung=rung)
        if cand.is_empty():
            skip_pair1 = True
    if not skip_pair1:
        # This ranking is on raw Sharpe, and the canonical resolver is not: when a conformal
        # candidate is in the field it re-ranks everything on exact common timestamp support,
        # so it can return a lower raw-Sharpe row. It also happens that the two tie exactly -
        # a risk overlay that never binds produces the same returns as the allocation stage
        # under it, to the last digit - and the dedupe then keeps whichever the sort emitted.
        #
        # Either way the pair ends up registered under a backtest the case study does not
        # report, and a notebook asking for its carrier's validation-to-benchmark evidence
        # finds none. So a supplied carrier is used rather than ranked against: the caller
        # resolved it through the canonical selection, which is the answer this ranking is a
        # cheaper approximation of. With no carrier the sort stands, with `backtest_hash` as
        # a final key so the choice is at least deterministic.
        carrier_backtest = str(lineage["val_backtest_hash"]) if lineage else None
        cand1 = cand.sort(["sharpe", "backtest_hash"], descending=[True, False]).unique(
            subset=["prediction_hash"], keep="first", maintain_order=True
        )
        if carrier_backtest is not None:
            pinned = cand.filter(pl.col("backtest_hash") == carrier_backtest)
            if pinned.is_empty():
                raise RuntimeError(
                    f"the carrier {carrier_backtest} passed for {cs} is not among the "
                    f"{cand.height} candidates this ranking sees. Pair #1 would be registered "
                    "under a different backtest than the one the case study reports."
                )
            cand1 = pinned
        leader_hash = cand1["backtest_hash"][0]
        leader_label = cand1["label"][0] if "label" in cand1.columns else None
        if leader_label:
            bench_resolution = _benchmark_returns_from_artifact(cs, leader_label)
            chal = _aligned_returns(cs, leader_hash)
            if bench_resolution and chal is not None:
                benchmark_hash, base, resolved_bench_label = bench_resolution
                min_n = _min_paired_n(ppy)
                aligned = chal.join(base, on="timestamp", how="inner", suffix="_b")
                if aligned.height >= min_n:
                    c_arr = aligned["ret"].to_numpy()
                    b_arr = aligned["ret_b"].to_numpy()
                    # Sized here, trimmed once inside `compute_paired_uncertainty`; see
                    # `_populate_pair` for why the pair is not trimmed on the way in.
                    if joint_returns(c_arr, b_arr)[0].size >= min_n:
                        paired = compute_paired_uncertainty(
                            c_arr,
                            b_arr,
                            periods_per_year=ppy,
                            case_study=cs,
                            label=leader_label,
                            n_boot=2000,
                            seed=42,
                        )
                        if paired:
                            benchmark_kind = (
                                f"{SIGNAL_BASELINE_BY_CASE_STUDY.get(cs, 'equal_weight')}"
                                "_side_artifact"
                            )
                            register_paired_metrics(
                                cs,
                                leader_hash,
                                benchmark_hash,
                                paired,
                                benchmark_kind=benchmark_kind,
                                periods_per_year=ppy,
                                case_dir=write_case_dir,
                            )
                            written_keys.add((leader_hash, benchmark_hash))
                            rows.append(
                                {
                                    "case_study": DISPLAY_NAMES.get(cs, cs),
                                    "kind": benchmark_kind,
                                    "label": leader_label,
                                    "benchmark_label": resolved_bench_label,
                                    "sharpe_diff": paired.get("sharpe_diff"),
                                    "sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
                                    "sharp

Полный текст с указанием источника опубликован на условиях его лицензии. Лицензия: MIT

Это краткое изложение подготовлено исследовательским агентом Stratmill по оригиналу и не является его копией.