رفتن به محتوا
همه اسناد کتابخانه

رتبه‌بندی مقاوم بک‌تست و ارزیابی استراتژی

کد یادگیری ماشین برای معامله‌گری

خلاصه

این سند ابزارهایی برای مقایسه استراتژی‌های معاملاتی ثبت‌شده و گردآوری ارزیابی‌های ساختاریافته مطالعات موردی شرح می‌دهد. یک روش رتبه‌بندی اصلی، سری‌های بازده را پیش از محاسبه سنجه‌های سبد به زمان‌مهرهای دقیقاً مشترک محدود می‌کند تا هر نامزد بر مبنای داده‌های یکسان سنجیده شود. زمان‌مهرهای تکراری و هم‌پوشانی ناکافی را رد می‌کند و استراتژی‌ای را که نابودی سرمایه را تجربه کرده، غیرقابل‌رتبه‌بندی می‌داند؛ حتی اگر پنجره مقایسه مشترک پیش از آن زیان پایان یابد. کد همچنین نشان می‌دهد محدودیت‌های برچسب و مجموعه نمادها چگونه می‌توانند تعیین کنند کدام نامزدها برای انتخاب استراتژی مرجع مطالعه موردی واجد شرایط‌اند.

ابزارهای ارزیابی، انتخاب‌های اعتبارسنجی را با بررسی پیکربندی و شرایط پس از اجرا به بازاجرای مجموعه نگه‌داشته‌شده پیوند می‌دهند و خلاصه‌های مدل، بک‌تست، تخصیص، هزینه و ریسک را برای گزارش‌گیری گرد می‌آورند. سند توضیح می‌دهد چرا نبود خروجی‌های مجموعه نگه‌داشته‌شده می‌تواند مرحله‌ای عادی از روند کار باشد، نه جست‌وجویی ناموفق. این‌ها تمهیدات حفاظتی و سازوکارهای گزارش‌دهی‌اند، نه گواه سودآوری استراتژی. محتوای قابل‌مشاهده بر پیاده‌سازی متمرکز است و نمونه‌هایی از اصلاح قواعد انتخاب دارد که در غیر این صورت ممکن بود بک‌تست‌های تشخیصی را بالاتر از نامزدهای مناسب‌تر قرار دهند.

ایده‌های کلیدی

  • سری‌های بازده نامزدها را در زمان‌مهرهای دقیقاً مشترک رتبه‌بندی کنید تا مقایسه منصفانه باشد.
  • استراتژی‌ای که نابودی سرمایه را تجربه می‌کند نباید بر اساس بخشی از تاریخچه پیش از نابودی رتبه‌بندی شود.
  • محدودیت‌های واجدشرایط‌بودن می‌توانند مشخص کنند کدام برچسب‌ها و مجموعه نمادها برای انتخاب مرجع مناسب‌اند.
  • تطبیق مجموعه نگه‌داشته‌شده باید تأیید کند که بازاجرا با پیکربندی درخواستی مطابقت دارد.
  • خلاصه‌های ارزیابی، شواهد اعتبارسنجی و ریسک را سامان می‌دهند، اما سودآوری را ثابت نمی‌کنند.

برچسب‌ها

متن کامل
# strategy_analysis.py


```py
"""Strategy analysis figure helpers and assessment writer.

Companion to ``BacktestExplorer`` - produces the figures and structured
artifacts for each case study's ``strategy_analysis.py`` notebook.

Usage::

    from case_studies.utils.strategy_analysis import (
        plot_sharpe_waterfall,
        plot_concentration_curve,
        plot_equity_drawdown,
        load_holdout_metrics,
        write_strategy_assessment,
        load_strategy_assessment,
    )
"""

from __future__ import annotations

import json
import math
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import UTC, datetime
from pathlib import Path
from typing import Any, Literal

import matplotlib.pyplot as plt
import numpy as np
import polars as pl

from case_studies.utils.carrier_pins import CARRIER_PINS
from case_studies.utils.notebook_contracts import (
    degenerate_prediction_sql,
    full_coverage_prediction_sql,
)
from case_studies.utils.uncertainty import STAGE_SEQUENCE
from case_studies.utils.warning_policy import warn_the_reader

# ---------------------------------------------------------------------------
# Canonical rank-1 resolution (LABEL_RESTRICTIONS-aware)
# ---------------------------------------------------------------------------
#
# Per-CS whitelist of labels eligible to anchor the registered strategy. The
# only entry today is sp500_options, restricted to ret_to_expiry because the
# four legacy diagnostic variants (fwd_ret_5d, fwd_ret_10d, fwd_ret_dh_5d,
# fwd_ret_dh_10d) were dropped from the sweep + registry 2026-05-17 - they
# went through the vectorized backtest path which treats their 5d/10d
# forward returns as daily returns, inflating Sharpes (e.g. fwd_ret_10d
# allocation Sharpe ~6.5) to non-credible levels. ret_to_expiry runs through
# the HTM daily-MTM cohort path and is the only label with an honest cost
# model for this CS. This is the only definition in the tree, which
# ``tests/test_carrier_routing_contract.py`` checks by reading every module: the retired
# ``20_strategy_synthesis/holdout.py`` kept a second copy in sync by comment, and it drifted.
LABEL_RESTRICTIONS: dict[str, frozenset[str]] = {
    "sp500_options": frozenset({"ret_to_expiry"}),
}


# Per-CS canonical universe pin: case_study -> strategy.signal.universe_filter
# value eligible to anchor the registered rank-1. sp500_options trades only the
# liquid (bottom-quintile half-spread) subset - the full-universe round-trip
# option spread consumes the variance-risk-premium edge, so full-universe rows
# are excluded from rank-1 selection (the full universe is retained only for the
# Ch18 htm_cost_cascade comparison, never as the deployed carrier). Without this
# pin, full-universe allocation backtests registered by the standard sweep
# (e.g. the 2026-05-31 L1-grid rollout) leak into rank-1 by raw Sharpe and
# orphan the liquid-lineage holdout. This is the only definition; the holdout selection
# applies it by going through :func:`selectable_validation_candidates` rather than by
# repeating the filter.
#
# nasdaq100_microstructure is the same arrangement one universe over. Its
# ``setup.yaml`` declares ``backtest.sweep.universe_filter: cost_feasible`` and says
# beside it that "the full-universe variant is NOT a canonical rank-1 / cohort / DSR
# candidate; it lives only in the 17_costs.py full-vs-screened comparison". Nothing
# enforced that: the sweep's pass 2 registers full-universe reference arms so
# ``17_costs`` can price the screen, and this resolver ranked them beside the screened
# rows by raw Sharpe. The entry is what makes the declaration true rather than stated.
#
# Measured 2026-09-13 before adding it: the registry held 96 full-universe signal rows
# and the rank-1 was the same row either way - `gbm/default_multiclass` on
# `fwd_dir_15m`, a cost-feasible slot configuration at Sharpe 2.416 - so this changes
# no published value today. That is the point at which to close a hole, not after a
# full-universe row has won and moved a chapter.
#
# ``test_declared_canonical_universe_is_pinned`` asserts the two stay together: a case
# study that declares the key and is missing from here has an unenforced declaration.
UNIVERSE_RESTRICTIONS: dict[str, str] = {
    "nasdaq100_microstructure": "cost_feasible",
    "sp500_options": "liquid",
}


# Carrier choices are owner-controlled in ``carrier_pins`` and use validation
# information only. The corrected S&P 500 options carrier is the liquid-universe
# cross-stage rank-1. Two alternative allocator rows tie its Sharpe exactly, so
# the deterministic tie-break preserves the simpler equal-weight baseline spec.


def _ruined(returns) -> bool:
    """Whether this return path took its account through zero equity anywhere."""
    from case_studies.utils.backtest_runner import first_ruin_index

    return first_ruin_index(returns) is not None


def rank_returns_on_common_support(
    returns_by_hash: dict[str, pl.DataFrame], *, periods_per_year: int
) -> pl.DataFrame:
    """Rank backtests after restricting every return series to exact common support."""
    if not returns_by_hash:
        raise ValueError("No return series supplied for common-support ranking")

    normalized: dict[str, pl.DataFrame] = {}
    common_timestamps: set[Any] | None = None
    for backtest_hash, frame in returns_by_hash.items():
        return_col = next(
            (name for name in ("daily_return", "return", "returns") if name in frame.columns),
            None,
        )
        if "timestamp" not in frame.columns or return_col is None:
            raise ValueError(
                f"{backtest_hash}: expected timestamp plus a return column; got {frame.columns}"
            )
        selected = (
            frame.select("timestamp", pl.col(return_col).alias("daily_return"))
            .with_columns(pl.col("timestamp").cast(pl.Datetime("ns")))
            .sort("timestamp")
        )
        if selected["timestamp"].n_unique() != selected.height:
            raise ValueError(f"{backtest_hash}: duplicate timestamps in daily returns")
        normalized[backtest_hash] = selected
        timestamps = set(selected["timestamp"].to_list())
        common_timestamps = (
            timestamps if common_timestamps is None else common_timestamps & timestamps
        )

    if common_timestamps is None or len(common_timestamps) < 2:
        raise ValueError("Backtest candidates have fewer than two common timestamps")

    from case_studies.utils.backtest_runner import compute_portfolio_metrics

    common = sorted(common_timestamps)
    common_frame = pl.DataFrame({"timestamp": common}, schema={"timestamp": pl.Datetime("ns")})
    common_ns = common_frame["timestamp"].cast(pl.Int64).to_list()
    rows: list[dict[str, Any]] = []
    for backtest_hash, frame in normalized.items():
        aligned = frame.join(common_frame, on="timestamp", how="inner").sort("timestamp")
        if aligned["timestamp"].cast(pl.Int64).to_list() != common_ns:
            raise ValueError(f"{backtest_hash}: failed exact common-support alignment")
        metrics = compute_portfolio_metrics(
            aligned["daily_return"].to_numpy(),
            periods_per_year=periods_per_year,
            uncertainty=False,
            trim_leading_zeros=False,
        )
        # A path the engine stopped at ruin carries no Sharpe, by design: a ratio
        # of a mean to a dispersion describes a process that continues
        # . The candidate stays on the frame so the
        # caller can see it was compared, and sorts below every solvent one.
        #
        # Ruin is read off the *whole* series, not off the common-support slice.
        # The intersection can end before the period that wiped the account out,
        # and a bankrupt book restricted to the months before it went bankrupt is
        # not a rankable strategy - it would outrank a solvent candidate with a
        # negative Sharpe on the strength of the window the comparison happened
        # to choose.
        #
        # None rather than the NaN the metrics carry: polars sorts a NaN first on
        # a descending sort, which is the one place that matters here.
        sharpe = metrics["sharpe"]
        if _ruined(frame["daily_return"].to_numpy()) or (sharpe is not None and math.isnan(sharpe)):
            sharpe = None
        rows.append(
            {
                "backtest_hash": backtest_hash,
                "sharpe": None if sharpe is None else float(sharpe),
                "n_periods": aligned.height,
                "start": common[0],
                "end": common[-1],
            }
        )
    return pl.DataFrame(rows).sort("sharpe", descending=True, nulls_last=True)


def rank_backtests_on_common_support(
    case_study: str, backtest_hashes: list[str], *, periods_per_year: int
) -> pl.DataFrame:
    """Load registered returns and rank them on their exact timestamp intersection."""
    from utils.paths import get_case_study_dir

    backtest_root = get_case_study_dir(case_study) / "run_log" / "backtest"
    returns_by_hash: dict[str, pl.DataFrame] = {}
    for backtest_hash in backtest_hashes:
        path = backtest_root / backtest_hash / "daily_returns.parquet"
        if not path.exists():
            raise FileNotFoundError(f"Missing registered return artifact: {path}")
        returns_by_hash[backtest_hash] = pl.read_parquet(path)
    return rank_returns_on_common_support(returns_by_hash, periods_per_year=periods_per_year)


@dataclass(frozen=True)
class HoldoutSelfBacktest:
    """The outcome of looking for a validation run's holdout replay.

    ``select_holdout_self_backtest`` answers with a hash or with ``None``, and ``None``
    covers four different states of the registry: the validation backtest is not
    registered, its prediction set is not, no holdout prediction set exists for the
    configuration at all, or holdout backtests exist and none replays the validation
    strategy. A strategy-analysis notebook that raises on ``None`` therefore tells its
    reader nothing about which, and the most common of the four - the holdout stage has
    simply not been run yet - is a normal state for anyone working the notebooks in
    order rather than a defect.

    ``reason`` is a sentence for the rendered page. It names the validation run that was
    searched for, so a reader can see the search was well formed and is not being told
    that something went wrong.

    ``training_hash`` is the identity that produced the holdout being returned, and it is
    here because a hash alone cannot answer the question that matters. The lookup matches
    on the declared configuration plus three post-conditions, and none of them reads a
    clock: a returned hash says "a registered holdout is a valid refit of the configuration
    you asked about", never "the holdout was taken while that configuration was rank-1".
    Those separate whenever the field kept growing after the window was spent.

    The gap that makes concrete: neither this
    lookup nor `18_holdout_predictions` could express *a holdout exists for this
    configuration, produced by a generation that is no longer reproducible*. The notebook
    derives the refit it would perform and compares; with the identity returned here, a
    caller can make that comparison too instead of taking the hash on trust.

    It is the registered identity and not a verdict on purpose. Deciding whether a
    generation is superseded needs the ranking, and this lookup is called from inside
    `resolve_canonical_rank1_lineage` - resolving the ranking here would re-enter it.
    """

    backtest_hash: str | None
    reason: str | None = None
    training_hash: str | None = None
    """The training identity behind the returned holdout, or None when nothing was found."""

    @property
    def found(self) -> bool:
        return self.backtest_hash is not None


HoldoutRefitStatus = Literal["refit", "not_out_of_sample", "unattributable"]


def holdout_refit_status(training_spec_json: str | None) -> HoldoutRefitStatus:
    """What a training run's own specification says about how it was fitted.

    Three answers, not two, and the third is why this exists. Read from the training
    specification rather than from the prediction set's split, because the split says where
    the predictions land and says nothing about what the model saw while fitting - a model
    fitted on the validation folds can publish predictions over the holdout window, and that
    is exactly the mistake the holdout exists to rule out.

    ``refit``
        The run's CV declares the holdout fold. This is a holdout evaluation.
    ``not_out_of_sample``
        The run records a CV split and it is not the holdout. **This is a statement about
        the record, not a finding about the fit.** The retired
        `20_strategy_synthesis/holdout.py::generate_holdout`, deleted on 2026-09-12,
        genuinely refit on a holdout
        fold and then registered the predictions under the *validation* training identity,
        whose CV says ``validation`` - and the rows it wrote are still in the registries -
        so a row answering this way is either a validation-fitted model published over the
        holdout window or a real refit filed under the wrong identity, and the registry
        cannot tell them apart. Either way it may not be reported as a holdout result, and
        either way only an operator can say which it is.
    ``unattributable``
        The run records no CV split, so nothing can be concluded either way.

    Neither of the last two authorizes a deletion on its own, which is the correction to
    make about this function: a non-holdout CV declaration does not establish that no
    holdout evaluation occurred. What they authorize is a refusal that names the row, which
    is what nothing did before - the notebooks filtered on the two-valued predicate and
    these rows were invisible to the refusal and to everything after it.
    """
    if not training_spec_json:
        return "unattributable"
    try:
        cv = (json.loads(training_spec_json).get("computation") or {}).get("cv") or {}
    except (TypeError, ValueError):
        return "unattributable"
    split = cv.get("split")
    if split is None:
        return "unattributable"
    return "refit" if split == "holdout" else "not_out_of_sample"


def training_run_fitted_for_the_holdout(training_spec_json: str | None) -> bool:
    """True when a training run's own CV declares the holdout fold.

    This is what separates a refit from a validation-fitted model scored on a later
    window. See :func:`holdout_refit_status`, of which this is the two-valued reading: a
    run that cannot be shown to have been refitted answers False, which is the right
    default for a lineage lookup and the wrong one for deciding what to delete.
    """
    return holdout_refit_status(training_spec_json) == "refit"


def registered_holdout_generations(case_dir: str | Path) -> list[dict[str, Any]]:
    """Every holdout prediction set in a registry, with what its specification says.

    One implementation. This was copied into `etfs/18_holdout_predictions`,
    `sp500_options/16_holdout_predictions` and
    `us_firm_characteristics/15_holdout_predictions`, and the copies had already drifted:
    two of them delete through the schema-derived cascade in `registry/maintenance.py` and
    the third listed the child tables by hand, which is the version that aborts on a
    registry holding a `cohort_metrics.leader_hash` row.

    ``status`` is :func:`holdout_refit_status` on the training run behind each set.
    ``checkpoint`` is part of the identity rather than a detail of it: one training run
    publishes one prediction set per declared checkpoint, and moving the selection from one
    checkpoint to another is a different configuration evaluated on the same window.
    Identity on the training hash alone would read that as the same generation and let both
    stand.
    """
    import sqlite3

    db_path = Path(case_dir) / "run_log" / "registry.db"
    with sqlite3.connect(f"file:{db_path}?mode=ro", uri=True) as conn:
        rows = conn.execute(
            """
            SELECT p.prediction_hash, p.training_hash, p.checkpoint_kind, p.checkpoint_value,
                   t.config_name, t.spec_json
            FROM prediction_sets p
            JOIN training_runs t ON t.training_hash = p.training_hash
            WHERE p.split = 'holdout'
            ORDER BY p.prediction_hash
            """
        ).fetchall()
    return [
        {
            "prediction_hash": prediction_hash,
            "training_hash": training_hash,
            "checkpoint": (checkpoint_kind, checkpoint_value),
            "config_name": config_name,
            "status": holdout_refit_status(training_spec_json),
            "refitted": holdout_refit_status(training_spec_json) == "refit",
        }
        for (
            prediction_hash,
            training_hash,
            checkpoint_kind,
            checkpoint_value,
            config_name,
            training_spec_json,
        ) in rows
    ]


@dataclass(frozen=True)
class HoldoutGenerationsToRetire:
    """What a holdout run finds already registered against its window, in three buckets.

    Splitting them is the point. The rule that governs each is different, and the code this
    replaces had only the first bucket - it filtered the registered generations on
    ``refitted``, so a row that was NOT a refit was invisible to the refusal and invisible
    to the deletion that follows it. Nothing owned removing one, and the note written when a
    stale `us_firm_characteristics` row was removed by hand on 2026-08-30 said exactly that:
    the next case study to reach the stage would accumulate the same second generation and
    the same silence.
    """

    superseded: tuple[dict[str, Any], ...]
    """Refits of a *different* configuration on the same window.

    A second holdout evaluation, and replacing one is a research decision rather than
    maintenance: deleting the rows does not undo having observed them. Whoever owns the
    sweep decides, which is what the notebooks' ``REPLACE_HOLDOUT`` switch is for.
    """

    not_out_of_sample: tuple[dict[str, Any], ...]
    """Rows whose training run records a CV split that is not the holdout.

    Two different things produce this record and it cannot separate them. One is a
    validation-fitted model publishing over the holdout window, the defect `29f13165`
    fixed, which is not a holdout evaluation at all. The other is a genuine refit filed
    under the validation training identity, which is what the retired
    `20_strategy_synthesis/holdout.py::generate_holdout`, deleted on 2026-09-12, did -
    it built a holdout fold,
    trained on it, then registered the predictions against `candidate["training_hash"]`.
    Removing it removed the producer; the rows it already wrote are why this bucket stays.

    So this bucket is refused rather than deleted. A row in it may not be reported as a
    holdout result, because on the first reading nothing out of sample was measured and on
    the second the identity is wrong; but deleting it automatically would, on the second
    reading, destroy a real evaluation on the strength of a record that is known to be
    unreliable for exactly this distinction.
    """

    unattributable: tuple[dict[str, Any], ...]
    """Rows whose training run records no CV split, so neither can be concluded.

    Refused, and with no way past: a missing specification is not evidence that a run was
    not refitted, and there is nothing recorded to adjudicate from.
    """


def holdout_generations_to_retire(
    case_dir: str | Path,
    *,
    this_generation: tuple[str, tuple[Any, Any]],
) -> HoldoutGenerationsToRetire:
    """Divide what is already registered against the holdout window by what governs it.

    ``this_generation`` is ``(training_hash, (checkpoint_kind, checkpoint_value))`` for the
    run about to be registered; a generation equal to it is this run and is in no bucket.
    """
    superseded, not_out_of_sample, unattributable = [], [], []
    for row in registered_holdout_generations(case_dir):
        if (row["training_hash"], row["checkpoint"]) == this_generation:
            continue
        if row["status"] == "refit":
            superseded.append(row)
        elif row["status"] == "not_out_of_sample":
            not_out_of_sample.append(row)
        else:
            unattributable.append(row)
    return HoldoutGenerationsToRetire(
        superseded=tuple(superseded),
        not_out_of_sample=tuple(not_out_of_sample),
        unattributable=tuple(unattributable),
    )


class HoldoutWindowSpent(RuntimeError):
    """A second evaluation of a window this case study reports as unseen was refused."""


def refuse_a_second_look(
    retire: HoldoutGenerationsToRetire,
    *,
    this_configuration: str,
    this_training_hash: str,
    checkpoint: tuple[Any, Any],
    retiring: Sequence[str] = (),
) -> tuple[dict[str, Any], ...]:
    """Refuse a second evaluation of a spent holdout window, unless it is named.

    The default is refusal, and the override is per generation rather than per run. A
    boolean would be set once and left set, and the guard would then be decorative; naming
    the prediction set means each override is a statement about one window that somebody had
    to look up. ``retiring`` is a sequence of prediction hashes the operator accepts
    retiring, and the run proceeds only when it names exactly what is registered.

    Returns the rows being retired, so the caller can put them in the render. A second look
    that proceeds deliberately has to say so where a reader sees it, not only in the launch
    line - the registry would otherwise show one evaluation of the window and a reader would
    have no way to learn there had been two.

    Only ``superseded`` is overridable. The other two buckets are not a decision anybody can
    make by naming a hash:

    * ``unattributable`` - the training runs record no CV split, so whether they were
      refitted for the holdout cannot be established either way. Naming one asserts a fact
      the registry does not hold.
    * ``not_out_of_sample`` - the row is either a validation-fitted model published over the
      window or a refit filed under its validation identity, and the registry cannot tell
      those apart. An override here would not authorize a second look; it would authorize
      reporting something that may never have been out of sample.
    """
    if retire.unattributable:
        raise HoldoutWindowSpent(
            "the holdout window carries prediction sets whose training runs record no CV "
            "split, so whether they were refitted for the holdout cannot be established: "
            + ", ".join(
                f"{row['prediction_hash']} (training {row['training_hash']})"
                for row in retire.unattributable
            )
            + ". Establish what produced them before registering another evaluation on the "
            "same window. This is not what `retiring` is for: naming one of these would "
            "assert something the registry does not record."
        )
    if retire.not_out_of_sample:
        raise HoldoutWindowSpent(
            "the holdout window carries prediction sets whose training runs declare a CV "
            "split other than the holdout: "
            + ", ".join(
                f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
                for row in retire.not_out_of_sample
            )
            + ". Each is either a validation-fitted model published over the window, which "
            "is not an out-of-sample result, or a refit registered under its validation "
            "training identity, and the registry cannot tell those apart. Resolve it "
            "through the registry's own lifecycle, which records that the row was retired."
        )

    registered = {row["prediction_hash"] for row in retire.superseded}
    named = dict.fromkeys(retiring)

    unknown = [hash_ for hash_ in named if hash_ not in registered]
    if unknown:
        raise HoldoutWindowSpent(
            f"the run authorizes retiring {', '.join(unknown)}, which this window does not "
            "carry. An override that names something absent is either aimed at another case "
            "study or has outlived the generation it was written for, and either way it "
            "would sit in the launch line authorizing whatever arrives next. Registered "
            "here: " + (", ".join(sorted(registered)) or "nothing") + "."
        )

    unnamed = [row for row in retire.superseded if row["prediction_hash"] not in named]
    if unnamed:
        kind, value = checkpoint
        raise HoldoutWindowSpent(
            "the holdout window already carries a refit of a different configuration: "
            + ", ".join(
                f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
                for row in unnamed
            )
            + f". This run would evaluate {this_configuration} (training "
            f"{this_training_hash}, checkpoint {kind}={value}) on the same window, which "
            "would be a second configuration measured on a period this case study reports "
            "as unseen. Deleting the earlier rows does not undo having observed their "
            "result - the selection that produced this configuration may have been informed "
            "by the earlier holdout number, and no deletion reaches that. To take the second "
            "look deliberately, name what is being retired: "
            "RETIRE_HOLDOUT_GENERATIONS:json="
            + json.dumps(sorted(row["prediction_hash"] for row in unnamed))
            + ". The run then records it where a reader sees it."
        )
    return retire.superseded


# What a holdout refit is allowed to change, and nothing else. Everything outside this set
# has to agree with the validation run, because the holdout is defined as *that configuration*
# refitted on a later window - not as another run that happens to share its name.
#
#   cv                        the definition of the refit: split, folds, identity, request
#   expected_prediction_keys  derived from the fold geometry, so it moves with cv
#   input_data_spec           carries the fold splits and a fingerprint over them. Its
#                             `artifacts` are still checked, through the top-level
#                             `feature_artifacts` that duplicates them.
#   runtime_identity          the environment the run happened in
#   source_identity           the notebook and commit that ran it
#   model.effective_params_by_fold   keyed by fold number, which the refit renumbers
#
# This is a denylist rather than an allowlist on purpose: a field added to the specification
# later is compared by default, so the check tightens as the spec grows instead of silently
# ignoring the new field.
# Fields a refit may differ in for reasons that are not the fold set: the CV specification it
# was asked for, where its inputs came from, and the two identity stamps. The fold-derived
# fields are NOT listed here - they come from `FOLD_DERIVED_FIELDS`, which is the declaration
# of what a holdout refit is required to recompute.
_REFIT_MAY_CHANGE = frozenset({"cv", "input_data_spec", "runtime_identity", "source_identity"})


def _refit_comparable(training_spec_json: str | None) -> dict | None:
    """A training specification reduced to what a refit must preserve.

    The fold-derived fields are skipped by reading the same tuple the holdout deriver writes
    them from, rather than by a list kept in step with it by hand. The hand-kept list matched
    two of the three and compared `macro_context.resolved_fold_digest`, which
    `_resolved_macro_digest` computes over `case.splits` - so it differs whenever a refit is
    correct, and the one holdout that reached this check was rejected for changing a field it
    was required to change.

    Scoped to the named subfield, never to its container: the other five entries under
    `macro_context` are the macro *configuration* - `series`, `policy`, `alignment`,
    `availability_lag_days`, `version` - and a run that differs in any of them is not a refit
    of this specification at all.
    """
    if not training_spec_json:
        return None
    # Imported in the function because `case_studies.research` pulls in the workspace on import
    # and this module is reached from notebooks that open no study.
    from case_studies.research.holdout import FOLD_DERIVED_FIELDS

    computation = dict(json.loads(training_spec_json).get("computation") or {})
    for key in _REFIT_MAY_CHANGE:
        computation.pop(key, None)
    for container, field in FOLD_DERIVED_FIELDS:
        if container == "computation":
            computation.pop(field, None)
            continue
        nested = dict(computation.get(container) or {})
        if nested:
            nested.pop(field, None)
            computation[container] = nested
    return computation


def is_refit_of(holdout_spec_json: str | None, validation_spec_json: str | None) -> bool:
    """True when a holdout training run is the validation run's own configuration refitted.

    Family, configuration name, label and checkpoint are what the queries can filter on in
    SQL, and they are not enough to identify a configuration. They are a *name*, and a name is
    reused across generations: refit a study after its features change and the new runs carry
    the same four values as the old ones. Measured on the current registries, fx_pairs has 144
    configuration groups spanning more than one feature-artifact generation and etfs has 10 -
    so on those two case studies the coarse filter alone can return a holdout fitted on
    features the study no longer publishes, and report it as the selected carrier's.

    Comparing the specifications closes that. The feature artifact digests are the field that
    catches the stale generation, but the comparison is deliberately not limited to them: model
    hyperparameters, feature names, the task and the sampling all have to agree too, because a
    holdout that differs in any of them is not a refit of what was selected.

    A run with no recorded specification answers False. It cannot be shown to be a refit, and
    the holdout lineage is not a place to assume.
    """
    holdout = _refit_comparable(holdout_spec_json)
    validation = _refit_comparable(validation_spec_json)
    if holdout is None or validation is None:
        return False
    return holdout == validation


def resolve_holdout_self_backtest(
    case_study: str,
    val_backtest_hash: str,
    *,
    labels: Sequence[str] | None = None,
    admitted: frozenset[str] | None = None,
) -> HoldoutSelfBacktest:
    """Find the holdout replay of a validation run's strategy, or say why there is none.

    A thin wrapper over :func:`_resolve_holdout_self_backtest`: every "not found" answer
    goes through :func:`_refuse_a_selection_disagreement` first, so a caller that selected
    a different carrier than the holdout notebooks ran is told that rather than being told
    the holdout has not been produced.

    ``labels`` and ``admitted`` are the caller's own selection scope, and they are carried
    into that diagnosis rather than dropped. A carrier chosen from a preview's label subset
    or from a frozen candidate set is a correct answer within its scope, and comparing it
    against an unrestricted resolution would reject it for disagreeing with a question it
    was never asked.
    """
    found = _resolve_holdout_self_backtest(case_study, val_backtest_hash)
    if found.backtest_hash is None:
        _refuse_a_selection_disagreement(
            case_study, val_backtest_hash, labels=labels, admitted=admitted
        )
    return found


def _resolve_holdout_self_backtest(
    case_study: str,
    val_backtest_hash: str,
) -> HoldoutSelfBacktest:
    """Find the holdout backtest that replays a validation run's strategy, or say why not.

    The lookup itself. Callers go through :func:`resolve_holdout_self_backtest`, which adds
    the selection-disagreement diagnosis to every "not found" answer.

    This is the canonical ``val_rank1_self`` lineage anchor for the section 6 holdout
    closure: the holdout backtest produced by replaying the validation rank-1 strategy on
    the holdout prediction set. Matching by strategy spec, rather than by taking the
    highest holdout Sharpe among candidates sharing the ``training_hash``, keeps the
    lookup robust against experimental side-channel allocators - ``conformal_weighted``
    most of all - that share the holdout prediction set but diverge from the validation
    rank-1's allocator. Without that guard an allocator variant whose holdout Sharpe
    happens to be higher silently displaces the anchor, and the
    ``backtest_paired_metrics`` ``val_rank1_self`` pair, written against the canonical
    lineage's holdout hash, is then never found.

    The checkpoint is part of the model configuration, so the replay is pinned to the
    validation prediction set's own ``checkpoint_value`` and ``checkpoint_kind`` as well
    as its ``training_hash``. One trained model registers one prediction set per declared
    checkpoint and the strategy spec is identical across them, so ``training_hash`` alone
    leaves several indistinguishable holdout candidates. Resolving those by holdout
    Sharpe - which an earlier ``ORDER BY bm.sharpe DESC`` did - reads the holdout to
    choose among configurations, which ``reference/CASE_STUDY_PIPELINE.md`` section 6
    forbids outright.

    Raises when the pinned lineage is still ambiguous, rather than picking one.
    """
    import sqlite3

    from utils.paths import get_case_study_dir

    db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
    with sqlite3.connect(str(db_path)) as db:
        row = db.execute(
            "SELECT prediction_hash, spec_json FROM backtest_runs WHERE backtest_hash = ?",
            (val_backtest_hash,),
        ).fetchone()
        if row is None:
            return HoldoutSelfBacktest(
                None,
                f"the selected validation backtest {val_backtest_hash} is not registered in "
                f"{case_study}'s run log, so there is nothing to look for a holdout replay of",
            )
        val_pred_hash, val_spec_json = row
        val_strategy = json.loads(val_spec_json).get("strategy", {})

        train_row = db.execute(
            """
            SELECT training_hash, checkpoint_value, checkpoint_kind
            FROM prediction_sets WHERE prediction_hash = ?
            """,
            (val_pred_hash,),
        ).fetchone()
        if train_row is None:
            return HoldoutSelfBacktest(
                None,
                f"the selected validation backtest {val_backtest_hash} names prediction set "
                f"{val_pred_hash}, which is not registered, so the configuration to replay "
                "cannot be identified",
            )
        training_hash, checkpoint_value, checkpoint_kind = train_row

        val_train_row = db.execute(
            "SELECT family, config_name, label, spec_json FROM training_runs "
            "WHERE training_hash = ?",
            (training_hash,),
        ).fetchone()
        configuration = val_train_row[:3] if val_train_row is not None else None
        val_training_spec_json = val_train_row[3] if val_train_row is not None else None
        if configuration is None:
            return HoldoutSelfBacktest(
                None,
                f"the training run {training_hash} behind validation backtest "
                f"{val_backtest_hash} is not registered, so the configuration to look for a "
                "holdout refit of cannot be named",
            )

        # ``IS`` is SQLite's null-safe equality: a configuration with no
        # checkpoint dimension stores NULL on both sides and must still match,
        # while ``=`` would drop it.
        #
        # The join is on the declared configuration rather than on the validation
        # training hash. A holdout prediction produced correctly carries a NEW
        # training identity - it is the same configuration refitted on the holdout
        # fold, and the identity covers the CV interval - so matching on the
        # validation training hash can only ever find a holdout scored from the
        # validation-fitted model, which is the thing the holdout exists to avoid.
        candidates = db.execute(
            """
            SELECT b.backtest_hash, b.spec_json, t.training_hash, t.spec_json
            FROM backtest_runs b
            JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
            JOIN training_runs t ON t.training_hash = p.training_hash
            WHERE p.split = 'holdout'
              AND p.checkpoint_value IS ?
              AND p.checkpoint_kind IS ?
              AND t.family = ?
              AND t.config_name = ?
              AND t.label = ?
            ORDER BY b.backtest_hash
            """,
            (checkpoint_value, checkpoint_kind, *configuration),
        ).fetchall()

    checkpoint = f"checkpoint {checkpoint_kind}={checkpoint_value}"
    if not candidates:
        return HoldoutSelfBacktest(
            None,
            f"no holdout backtest is registered for the configuration behind validation run "
            f"{val_backtest_hash} ({configuration[0]}/{configuration[1]} on "
            f"{configuration[2]}, {checkpoint}), so the holdout has not been evaluated for "
            "this case study",
        )

    # Three conditions, and the SQL above can express none of them. The strategy spec has to
    # be the validation run's, so the anchor is a replay of what was selected rather than a
    # neighbouring allocator that shares the holdout prediction. The training run behind it has
    # to have been refitted for the holdout: a model fitted on the validation folds can publish
    # predictions over the holdout window, and accepting one would report the exact thing the
    # holdout exists to rule out. And it has to be a refit of THIS specification rather than of
    # a configuration with the same name - see `is_refit_of`, which is what stops a holdout
    # fitted on a superseded feature generation from being reported as the carrier's.
    matched = sorted(
        {
            (bh, holdout_training_hash)
            for bh, spec_json, holdout_training_hash, training_spec_json in candidates
            if json.loads(spec_json).get("strategy", {}) == val_strategy
            and training_run_fitted_for_the_holdout(training_spec_json)
            and is_refit_of(training_spec_json, val_training_spec_json)
        }
    )
    if not matched:
        return HoldoutSelfBacktest(
            None,
            f"{len(candidates)} holdout backtests are registered for "
            f"{configuration[0]}/{configuration[1]} on {configuration[2]} ({checkpoint}), and "
            "none of them replays that run's strategy from a run refitted for the holdout "
            "under the same specification, so the anchor is not a replay of what was "
            "selected",
        )
    if len(matched) > 1:
        raise ValueError(
            f"holdout replay for {val_backtest_hash} is ambiguous: "
            f"{[bh for bh, _ in matched]} are all "
            f"{configuration[0]}/{configuration[1]} on {configuration[2]}, refitted for the "
            f"holdout at {checkpoint}, with one strategy spec"
        )
    backtest_hash, holdout_training_hash = matched[0]
    return HoldoutSelfBacktest(backtest_hash, training_hash=holdout_training_hash)


def _refuse_a_selection_disagreement(
    case_study: str,
    val_backtest_hash: str,
    *,
    labels: Sequence[str] | None = None,
    admitted: frozenset[str] | None = None,
) -> None:
    """Raise when a *different* carrier has the holdout the caller could not find.

    A strategy-analysis notebook that ranks its pool's Sharpe column directly can select a
    different configuration than `resolve_solvent_carrier` does - a Sharpe computed over a
    configuration's own available history is not comparable across configurations that
    priced different spans, so the raw ranking rewards whichever candidate had the most
    forgiving window. Measured on cme_futures (992 signal / 120 allocation / 28 risk_overlay
    backtests, rebuilt 2026-08-30): the raw column answered latent_factors/sdf on
    fwd_ret_21d at 1.274, the resolver gbm/leaves_31_mse on fwd_ret_5d at 1.236 raw and
    1.294 once compared over the 1,270 sessions they all price.

    The failure that followed was silent and read as the wrong thing. `17_holdout_predictions`
    and `18_holdout_backtest` resolve through `resolve_solvent_carrier`, so the notebook then
    asked for the holdout replay of a configuration those notebooks never ran, got None, and
    printed "not produced yet" while the holdout sat in the registry. A reader concluded the
    holdout had not been run.

    Only that exact shape raises: a *registered* hash other than the canonical one was asked
    about, and the canonical carrier has a replay. A case study whose holdout stage genuinely
    has not run reports "not produced yet" as before, which is a normal state for anyone
    working the notebooks in order. An unregistered hash keeps its own answer, which is more
    useful than this one - a hash the run log has never seen is a stale constant rather than
    a carrier chosen from a pool. And a resolver that cannot answer - an empty registry, a
    case study with no rank-1 - leaves the original answer standing rather than turning a
    missing holdout into a resolver error.
    """
    import sqlite3

    from utils.paths import get_case_study_dir

    # Only a hash that is actually registered can be a *selection*. When the caller names a
    # backtest the run log has never seen, `_resolve_holdout_self_backtest` already says
    # exactly that, and it is the more useful answer - a stale hardcoded hash, not a
    # carrier chosen from a pool.
    try:
        db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
        with sqlite3.connect(f"file:{db_path}?mode=ro", uri=True) as db:
            registered = db.execute(
                "SELECT 1 FROM backtest_runs WHERE backtest_hash = ?", (val_backtest_hash,)
            ).fetchone()
    except sqlite3.Error:
        return
    if registered is None:
        return

    try:
        canonical = resolve_canonical_rank1_lineage(case_study, admitted=admitted, labels=labels)
    except Exception:  # noqa: BLE001 - this diagnoses; it must never replace the real answer
        return
    canonical_hash = canonical.get("val_backtest_hash")
    if not canonical_hash or canonical_hash == val_backtest_hash:
        return
    if canonical.get("holdout_backtest_hash") is None:
        return
    raise RuntimeError(
        f"{case_study}: no holdout replays validation backtest {val_backtest_hash}, but "
        f"{canonical['holdout_backtest_hash']} replays {canonical_hash} - "
        f"{canonical.get('family')}/{canonical.get('config_name')} on "
        f"{canonical.get('label')}, which is what `resolve_solvent_carrier` selects and what "
        "the holdout notebooks ran. This is a selection disagreement, not a missing holdout: "
        "the caller ranked its own pool and chose a different carrier. Resolve the carrier "
        "through `resolve_solvent_carrier`, which re-ranks candidates on exact common "
        "timestamp support and applies LABEL_RESTRICTIONS, UNIVERSE_RESTRICTIONS and "
        "CARRIER_PINS."
    )


def select_holdout_self_backtest(
    case_study: str,
    val_backtest_hash: str,
    *,
    labels: Sequence[str] | None = None,
    admitted: frozenset[str] | None = None,
) -> str | None:
    """The holdout replay's hash, or ``None`` when there is not exactly one.

    Kept for callers that only need the hash. A notebook that has to tell its reader
    what is missing wants ``resolve_holdout_self_backtest``, which carries the reason.
    """
    return resolve_holdout_self_backtest(
        case_study, val_backtest_hash, labels=labels, admitted=admitted
    ).backtest_hash


class NoSelectableCandidates(RuntimeError):
    """The case study currently publishes no configuration that may be selected.

    Every refusal `selectable_validation_candidates` raises for an empty pool is this
    type, so a caller asking *whether* a selection exists can answer "no" without
    swallowing an unrelated failure. It subclasses ``RuntimeError`` because that is what
    the selector raised before the type existed, and every caller that reports the refusal
    rather than branching on it keeps working unchanged.

    `has_holdout_predictions` is the caller that needs the distinction: it reports whether
    a holdout already covers the current top-N, and an initialised registry with no
    eligible validation backtest is a legitimate "not yet", not an error. It used to catch
    `ValueError`, which is what the pool it built itself raised; routing it through the
    canonical selector changed the type under it. The caller that made this matter was
    `20_strategy_synthesis/00_holdout_predictions.py`, which put the call outside its
    generation loop's handler, so one un-run case study stopped every case study after
    it. That notebook is retired with the rest of Chapter 20's holdout generation; the
    distinction stays because a "not yet" reported as a failure is wrong wherever it is
    read.
    """


SELECTION_STAGES: tuple[str, ...] = ("signal", "allocation", "risk_overlay")
"""The pool a configuration is selected from.

``reference/CASE_STUDY_PIPELINE.md`` section 5: the selected configuration is the
highest validation backtest Sharpe across these three stages, exactly.
``cost_sensitivity`` is a perturbation analysis rather than an alternative strategy and
is excluded; the ``benchmark`` family is excluded separately, below.
"""


def selectable_validation_candidates(
    case_study: str,
    *,
    admitted: frozenset[str] | None = None,
    labels: Sequence[str] | None = None,
    families: Sequence[str] | None = None,
) -> list[dict[str, Any]]:
    """Every validation backtest this case study may select from, best Sharpe first.

    One implementation, because holdout selection had two. The retired
    ``20_strategy_synthesis/holdout.py::select_best_models`` built its own pool out of
    ``BacktestExplorer.best`` per stage and this function's caller built one in SQL, and
    the two were held together by a comment. They applied
    different eligibility filters - membership on the prediction side there, a
    retired-set exclusion on both sides here - and different orderings, so they could
    name different configurations for the single holdout use. Measured on ``fx_pairs``
    2026-09-07: ``deep_learning/tcn`` on ``fwd_ret_21d`` has two backtests tied at Sharpe
    0.2639142245820113, and ``select_best_models`` answered ``9402978117e9`` while this
    path answered ``56070f34dff1``. Same model, two strategy specifications, and the
    holdout replays the specification exactly.

    Eligibility, in the order it is applied:

    * the stage pool (:data:`SELECTION_STAGES`) on ``split='validation'``, non-null
      Sharpe, ``family != 'benchmark'``, and no degenerate prediction set;
    * ``LABEL_RESTRICTIONS`` and ``UNIVERSE_RESTRICTIONS``, or a ``CARRIER_PINS`` entry
      where the owner has recorded one, which supersedes the ranking entirely;
    * ``families`` and ``labels``, the caller's own narrowing - ``labels`` replaces the
      declared restriction rather than adding to it, so a preview that narrowed its pool
      resolves inside the pool it narrowed to;
    * **publication**: an identity is selectable only where the population its producer
      publishes still lists it, asked per member kind and on both sides of the join.
    * **what the sweep refused**: a prediction set a sweep measured and dropped for covering
      less than the cross-section its feature panels offered it. Read from the registry
      rather than recomputed, because the check costs what the sweep's startup costs and
      because two implementations of one rule is what let this resolver be the looser of the
      two. Absence is not admission: a member no sweep has measured has no row and is left
      where it is.

    Publication is the membership question, not the exclusion one, and the two are not
    the same set. A prediction that no population ever listed was retired by nobody, so
    an exclusion set admits it and a membership set does not - which is how an
    experimental result its case study never published reaches a ranking.
    :func:`published_members_at` already subtracts what is retired, so membership is
    never the weaker test; where a registry declares no population of a kind it answers
    ``None`` and that kind places no constraint, which is the state of a fixture or of a
    registry written before the mechanism existed.

    ``admitted``, when given, is applied here rather than checked against the winner
    afterwards - see :func:`resolve_canonical_rank1_lineage`, whose common-support
    re-ranking is decided by the whole field and not only by the row that wins.

    Each candidate is a dict with ``backtest_hash``, ``prediction_hash``, ``stage``,
    ``training_hash``, ``family``, ``config_name``, ``label``, ``sharpe`` and
    ``spec_json``. Ordering is Sharpe descending, then the signal-only specification
    ahead of one carrying an allocation block, then ``backtest_hash`` ascending - a total
    order, so two callers ranking the same registry cannot disagree.
    """
    import sqlite3

    from case_studies.research.population import published_members_at
    from case_studies.utils.notebook_contracts import (
        predictions_the_sweep_refused,
    )
    from utils.paths import get_case_study_dir

    case_dir = get_case_study_dir(case_study)
    db_path = case_dir / "run_log" / "registry.db"
    label_filter = tuple(labels) if labels is not None else LABEL_RESTRICTIONS.get(case_study)
    universe_pin = UNIVERSE_RESTRICTIONS.get(case_study)
    carrier_pin = CARRIER_PINS.get(case_study)
    columns = (
        "backtest_hash",
        "prediction_hash",
        "stage",
        "training_hash",
        "family",
        "config_name",
        "label",
        "sharpe",
        "spec_json",
    )

    base_select = """
        SELECT b.backtest_hash, b.prediction_hash, b.stage,
               t.training_hash, t.family, t.config_name, t.label,
               bm.sharpe, b.spec_json
        FROM backtest_runs b
        JOIN backtest_metrics bm ON bm.backtest_hash = b.backtest_hash
        JOIN prediction_sets p ON p.prediction_hash = b.prediction_hash
        JOIN training_runs t ON t.training_hash = p.training_hash
        JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
    """
    # A Sharpe computed over fewer decision dates than its peers describes a different
    # sample, so it is not comparable with theirs and must not be ranked beside them
    # (`reference/CASE_STUDY_PIPELINE.md` section 10). `select_best_models` applied this
    # through `BacktestExplorer.best` and this resolver did not, which is one more way the
    # two pools could differ, and the collapsed path takes the stronger of the two rather
    # than the one that happened to be shorter. Measured across all nine production
    # registries 2026-09-07: it removes 18 of etfs' 2,038 ranked rows, none anywhere else,
    # and moves no rank-1.
    #
    # The explorer's other bar, `num_trades > 0`, is not carried over. It is not part of
    # the selection rule - a flat period is an observation of zero, not an abstention - and
    # it removes nothing from any of the nine.
    #
    # The maximum is taken WITHIN the published population, not across the whole table, and
    # the difference is not cosmetic. The bar keeps rows whose `ic_n_days` equals the
    # maximum for their `(split, family, label)`; a retired prediction scored over a longer
    # window sets that maximum, and every live row for the same family and label then falls
    # below it and is dropped by the query. Filtering publication afterwards cannot put them
    # back - they never came out of the database. `BacktestExplorer.best` passes the
    # population for exactly this reason, and dropping it while collapsing the two selectors
    # would have been a regression on the one case study whose refit changed a fold count.
    published_predictions = published_members_at(case_dir, member_kind="prediction")
    if published_predictions is not None and not published_predictions:
        raise NoSelectableCandidates(
            f"{case_study} declares prediction populations and publishes no prediction "
            "identities, so there is nothing it may select. Re-run the stage that publishes "
            "them rather than ranking over an empty population."
        )
    coverage_params: tuple = ()
    if published_predictions is None:
        coverage_bar = full_coverage_prediction_sql("p", "t", "pm")
    else:
        coverage_bar = full_coverage_prediction_sql(
            "p", "t", "pm", population_subquery="SELECT value FROM json_each(?)"
        )
        coverage_params = (json.dumps(sorted(published_predictions)),)

    if carrier_pin:
        # Documented a-priori carrier pin: resolve directly to the pinned
        # validation backtest rather than the max-Sharpe cross-stage rank-1.
        # The owner-controlled pin is a validation-time choice. Current-lineage
        # carrier decisions are deferred until all model producers finish.
        val_sql = base_select + (
            " WHERE b.backtest_hash LIKE ?"
            " AND p.split = 'validation'"
            " AND bm.sharpe IS NOT NULL"
            + degenerate_prediction_sql("p.prediction_hash")
            + coverage_bar
            + " ORDER BY bm.sharpe DESC, b.backtest_hash ASC"
        )
        params: tuple = (carrier_pin + "%",) + coverage_params
    else:
        stages = ",".join("?" for _ in SELECTION_STAGES)
        val_sql = base_select + (
            f" WHERE b.stage IN ({stages})"
            " AND p.split = 'validation'"
            " AND bm.sharpe IS NOT NULL"
            " AND t.family != 'benchmark'"
            + degenerate_prediction_sql("p.prediction_hash")
            + coverage_bar
        )
        params = tuple(SELECTION_STAGES) + coverage_params
        if label_filter:
            placeholders = ",".join("?" for _ in label_filter)
            val_sql += f" AND t.label IN ({placeholders})"
            params += tuple(label_filter)
        if families:
            placeholders = ",".join("?" for _ in families)
            val_sql += f" AND t.family IN ({placeholders})"
            params += tuple(families)
        if universe_pin:
            val_sql += " AND json_extract(b.spec_json, '$.strategy.signal.universe_filter') = ?"
            params += (universe_pin,)
        # Tie-break: among rows with identical Sharpe (e.g. the equal-weight baseline
        # equal-weight selection and its economically identical equal_weight
        # allocation-stage re-run, which share a prediction), prefer the
        # signal-only spec (no allocation block). That is the spec the holdout
        # is replayed from, so the canonical lineage stays poolable with its
        # holdout. Final ``backtest_hash`` key makes the order deterministic.
        val_sql += (
            " ORDER BY bm.sharpe DESC,"
            " (json_extract(b.spec_json, '$.strategy.allocation') IS NULL) DESC,"
            " b.backtest_hash ASC"
        )

    db = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
    try:
        rows = db.execute(val_sql, params).fetchall()
    finally:
        db.close()
    candidates = [dict(zip(columns, row, strict=True)) for row in rows]

    # A superseded generation is still complete, still `current` under its schema version,
    # and still ranks. `identity_status` says the registry understands the row; it says
    # nothing about whether the row is the one its producer still publishes, which is
    # recorded in the population lineage instead. Without this the carrier can be resolved
    # from a retired generation - measured on fx_pairs, where a rebuilt allocation stage
    # left the retired conformal-v2 backtest ranking first and every downstream notebook
    # refused it as unreproducible rather than selecting the generation in force.
    #
    # Both sides of the join, because a retired generation reaches the ranking through
    # either. The backtest side is the obvious one. The prediction side is the one that
    # survived unnoticed: a refit that changes no numbers - a relabel, a re-key, a rerun
    # that reproduces its inputs - publishes value-for-value identical predictions under a
    # new identity, so the old and new rows carry the SAME Sharpe to the last digit. On an
    # exact tie the ORDER BY returns whichever row it likes, and "whichever it likes" was
    # observed returning the retired one. Measured on sp500_equity_option_analytics: three
    # candidates at sharpe 1.965796084396144, and the resolver took a training run from a
    # superseded generation, against which a full 17-point cost surface was then registered.
    #
    # That is the shape worth remembering: the tie is produced BY CONSTRUCTION whenever a
    # refit changes nothing, so every lane that has ever superseded a population is exposed,
    # and the defect is invisible wherever no tie exists and silently wrong wherever one
    # does. Which is why it survived.
    #
    # The question is asked per NAME rather than globally, for both kinds. The naive
    # "retired by someone and listed by nobody in force" reads as equivalent and is not:
    # one identity is legitimately listed under several names, and a narrowed or preview
    # run keeps its own frozen snapshot in force forever.
    ranked = len(candidates)
    published_backtests = published_members_at(case_dir, member_kind="backtest")
    if published_backtests is not None and not published_backtests:
        raise NoSelectableCandidates(
            f"{case_study} declares backtest populations and publishes no backtest "
            "identities, so there is nothing it may select. Re-run the stage that publishes "
            "them rather than ranking over an empty population."
        )
    for key, published in (
        ("backtest_hash", published_backtests),
        ("prediction_hash", published_predictions),
    ):
        if published is None:
            continue
        candidates = [row for row in candidates if row[key] in published]

    # What the sweep measured and refused. `full_coverage_prediction_sql` above is the bar this
    # query can express, and it counts decision days: a family that scores every day for half
    # the universe ties the day count and ranks beside families that scored all of it. The
    # sweep charges every member against the feature panel it was offered and drops the ones
    # that fall short, and until it recorded that answer the two rules disagreed with the
    # resolver on the looser side - measured on nasdaq100_microstructure 2026-09-13, where
    # pinning the label made the carrier a 67.4%-coverage prediction set with an IC of
    # -0.00125 that the sweep had already stopped backtesting.
    #
    # Only what a sweep POSITIVELY recorded as short is dropped here. A member nothing has
    # measured has no row and is left exactly where it is, so a registry swept before the
    # record existed keeps t

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: MIT

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.