עבור לתוכן
כל מסמכי הספרייה

מדדי בקטסט מהימנים ואנואליזציה של רשת תשואות

קוד Machine Learning for Trading

סיכום

המסמך מתאר עקרונות מימוש למנוע בקטסט משותף ומתמקד בפרשנות נכונה של רשתות תשואה בעת אנואליזציה של ביצועים. הוא אומד את תדירות הדגימה הנצפית לפי הפערים הטיפוסיים בין תצפיות תשואה, משמיט פערים ארוכים באופן חריג ואז משווה את האומדן לתדירות שהוגדרה במחקר המקרה. המוסכמה המוגדרת נשמרת כשהרשת הנצפית תואמת לה; רשת גסה יותר במידה ניכרת יכולה להעיד שהתשואות דוללו למועדי איזון מחדש, ולחייב גורם אנואליזציה נמוך יותר.

הקטע מסביר גם מדוע בקטסטים צריכים להפסיק לצבור תשואות אחרי שההון בחשבון מגיע לאפס: תשואות נמוכות ממינוס אחת עלולות להפוך יתרות מצטברות מאוחרות יותר וליצור מדדי שארפ או צמיחה מטעים. הוא ממליץ להגביל את ההפסד להון הכולל ולהתייחס לתשואות מאוחרות יותר כאפס. אלה אמצעי הגנה למדידה ולסימולציה, ולא אסטרטגיית מסחר. הטקסט שסופק חלקי וכולל פרטי מימוש ספציפיים; התאמת התדירויות שלו תלויה בחותמות זמן ובהגדרות מהימנות, ואילו הטיפול המפושט בפשיטת רגל אינו מדמה דרישות להפקדת ביטחונות, נושים או כללי פירוק אחרים.

רעיונות מרכזיים

  • האנואליזציה צריכה לשקף את המרווחים בין תצפיות התשואה, ולא להניח שהם נקבעים לפי תדירות האיזון מחדש.
  • פערים טיפוסיים בחותמות זמן יכולים להבחין בין רשת תשואות דלילה לבין רשת עם תקופות חסרות בודדות.
  • מספר התקופות המוגדרות בשנה מותאם לתדירות הנצפית כדי להימנע ממדדי ביצועים לא עקביים.
  • צבירה מעבר להפסד מלא עלולה ליצור תשואות לא תקפות לאחר פשיטת רגל וסטטיסטיקות מטעות.
  • הגבלת ההון לאפס היא כלל חשבון מפושט ואינה מייצגת את מנגנוני הביטחונות בפירוט.

תגיות

הטקסט המלא
# backtest_runner.py


```py
"""Core backtest execution — engine-first, used by BOTH demo and sweep notebooks.

This module provides a single ``run_backtest()`` function that:
1. Converts predictions to target weights via strategy_spec["signal"]
2. Dispatches to engine or vectorized path
3. Optionally registers the result in registry.db
4. Returns a unified result object

The key invariant is that **sweep notebooks call the same function as demo
notebooks**. There is no separate vectorized reimplementation for sweeps.

Usage::

    from case_studies.utils.backtest_runner import run_backtest

    result = run_backtest(
        case_study="etfs",
        prediction_hash="abc123",
        strategy_spec=spec,
        prices=prices,
        predictions=predictions,
    )
"""

from __future__ import annotations

import math
import warnings
from copy import deepcopy
from dataclasses import dataclass
from datetime import date, datetime, time
from typing import Any, cast
from zoneinfo import ZoneInfo

import numpy as np
import polars as pl

from case_studies.utils.backtest_loaders import (
    declared_rebalance_step,
    get_backtest_config,
)
from case_studies.utils.backtest_presets import (
    apply_calendar_session_enforcement,
    ensure_backtest_spec,
    is_backtest_spec,
    runtime_backtest_config,
    strategy_view,
)
from case_studies.utils.signals import build_target_weights_from_config
from case_studies.utils.warning_policy import warn_the_reader

# ---------------------------------------------------------------------------
# Periods per year for Sharpe annualization
# ---------------------------------------------------------------------------

# NOTE: a `cadence -> periods_per_year` table used to live here. It was the
# source of a real defect: it assumed each return observation is one rebalance
# period, which is false for every case study that rebalances slowly and marks
# to market daily. Measured grid frequencies are 252/yr for sp500_options and
# nasdaq100_microstructure despite weekly cohorts and intraday signals, and
# 252/yr for monthly-rebalanced etfs. Annualize with
# `resolve_periods_per_year` below, which reads the declared
# `evaluation.periods_per_year` and reconciles it against the grid on hand.


def observed_periods_per_year(daily_returns) -> float:
    """Sampling rate of a return frame, or 0.0 if it cannot be measured.

    Rows divided by elapsed span would answer a different question - how densely
    the frame covers its own window - and would read a daily series with a
    six-month hole in it as a twice-weekly one. What the annualization needs is
    the spacing between consecutive observations, so the rate is measured over
    the intervals that are typical of the series and gaps far longer than the
    typical spacing are dropped from both the count and the span.

    A weekday grid has intervals of one day and three across a weekend, both
    within the cut, and comes out near 252. A grid thinned to every fifth
    session has intervals near seven days throughout and comes out near 52. A
    delisting hole or a suspended market drops out of the span instead of
    dragging the rate down with it.
    """
    n = len(daily_returns)
    if n <= 2:
        return 0.0
    all_ts = daily_returns["timestamp"].unique().sort()
    # `total_seconds()` returns whole seconds, so a gap under one second truncates to 0 and the
    # filter below deletes it. That is safe here because of the grids this is called on, not
    # because of the unit: the finest return grid any case study evaluates on is minute-level,
    # so the smallest real gap is 60. It would go false the day something evaluates on seconds,
    # and it would fail quietly - every gap zeroed, the array emptied, 0.0 returned, and
    # `reconcile_periods_per_year` keeping the declared constant with nothing saying why.
    # Switch to `dt.total_nanoseconds()` at that point rather than widening the filter.
    # Found 2026-09-10 by the sweep behind
    # 03_market_microstructure/04_itch_order_lifecycle_analysis, where the same pair of calls
    # zeroed every sub-second order lifetime and then dropped the whole population.
    gaps = all_ts.diff().drop_nulls().dt.total_seconds().to_numpy().astype(float)
    gaps = gaps[gaps > 0]
    if gaps.size == 0:
        return 0.0

    typical = float(np.median(gaps))
    if typical <= 0:
        return 0.0
    kept = gaps[gaps <= 5.0 * typical]
    if kept.size == 0:
        return 0.0

    span_years = float(kept.sum()) / (365.25 * 86400)
    return kept.size / span_years if span_years > 0.01 else 0.0


def reconcile_periods_per_year(declared: int, daily_returns, *, case_study: str = "") -> int:
    """Keep the declared factor unless the frame on hand contradicts it outright.

    The declared `evaluation.periods_per_year` is the authority: it describes
    the convention of the registered return grid, and using it keeps metrics on
    the round factor the rest of the pipeline shares. But it is a case-study
    constant, and one path writes a grid that does not match it. The vectorized
    path thins predictions to rebalance dates before aggregating, so with
    `sp500_options`'s rebalance_step of 5 a daily prediction grid leaves about
    74 rows a year against a declared 252.

    The override is one-directional, because thinning is the only mechanism by
    which the written grid departs from the declared convention, and it only
    coarsens. A frame that reads *finer* than declared is therefore not evidence
    that the declared value is wrong - an intraday grid whose case study still
    declares 252 is a setup.yaml that has not caught up, and inflating Sharpe by
    the square root of a factor of 139 is not the way to report that. It stays on
    the declared value and says so in the log.

    A factor of two of headroom on the coarse side keeps holidays and partial
    years on the declared value, and still catches thinning, which coarsens by
    the rebalance step and so lands well outside it.
    """
    import logging

    observed = observed_periods_per_year(daily_returns)
    if not observed or not declared:
        return declared or int(observed) or 252

    ratio = observed / declared
    if ratio > 2.0:
        logging.getLogger(__name__).warning(
            "%s: return grid is %.1f obs/yr, finer than the %d declared in setup.yaml. "
            "Annualizing at the declared value; declare the grid's own frequency if "
            "this case study writes intraday returns.",
            case_study or "backtest",
            observed,
            declared,
        )
        return declared
    if ratio >= 0.5:
        return declared

    logging.getLogger(__name__).warning(
        "%s: return grid is %.1f obs/yr but setup.yaml declares %d; annualizing at "
        "the observed rate. Thinning to rebalance dates is the usual cause.",
        case_study or "backtest",
        observed,
        declared,
    )
    return int(observed)


def resolve_periods_per_year(case_study: str, daily_returns) -> int:
    """Annualization factor for a return series, resolved once per backtest.

    Every path that annualizes a registered return grid goes through here, so
    the overall metrics, the per-fold metrics and the backfill cannot disagree
    about what the same frame is.
    """
    from case_studies.utils.uncertainty import periods_per_year_from_setup

    try:
        declared = int(periods_per_year_from_setup(case_study))
    except (FileNotFoundError, KeyError):
        declared = 0

    return reconcile_periods_per_year(declared, daily_returns, case_study=case_study)


def overall_periods_per_year(case_study: str | None, calendar: str, daily_returns) -> int:
    """Annualization for a whole-window backtest, whichever engine produced it.

    The engine path used to read the exchange calendar, which returns ~252 for
    every equity venue and so disagreed with a monthly case study by a factor of
    21, and with its own per-fold metrics. A named case study now resolves the
    same way everywhere; the calendar remains for a call that has no case study
    and therefore no declared value to read.
    """
    if case_study:
        return resolve_periods_per_year(case_study, daily_returns)
    return calendar_periods_per_year(calendar)


# Calendar name → exchange_calendars MIC code
_CALENDAR_TO_XCAL: dict[str, str] = {
    "NYSE": "XNYS",
    "CME": "us_futures",
    "FX": "24/5",
    "crypto": "24/7",
}

# Cache for calendar session counts
_calendar_ppy_cache: dict[str, int] = {}


def calendar_periods_per_year(calendar: str) -> int:
    """Get trading days per year for a calendar using exchange_calendars.

    Computes the average number of sessions over a 10-year window
    (2015-2024) and caches the result.
    """
    if calendar in _calendar_ppy_cache:
        return _calendar_ppy_cache[calendar]

    xcal_name = _CALENDAR_TO_XCAL.get(calendar, calendar)

    try:
        import exchange_calendars as xcals

        cal = xcals.get_calendar(xcal_name)
        total = sum(
            len(cal.sessions_in_range(f"{y}-01-01", f"{y}-12-31")) for y in range(2015, 2025)
        )
        ppy = round(total / 10)
    except Exception:
        # Fallback if exchange_calendars unavailable or calendar unknown
        ppy = 252

    _calendar_ppy_cache[calendar] = ppy
    return ppy


# ---------------------------------------------------------------------------
# Portfolio metrics via ml4t-diagnostic
# ---------------------------------------------------------------------------


# ---------------------------------------------------------------------------
# Ruin: an account that loses its capital stops there
# ---------------------------------------------------------------------------

# What a bankrupt path reports, and why it reports that.
#
# A long-short book can lose more than it holds in a single period: the long leg
# stops at -100%, but the short leg's loss is unbounded, and a squeeze on a
# concentrated short costs more than the account holds. Compounding straight
# through zero is not a larger version of that loss, it is a different quantity.
# Once equity is negative, `(1 + r)` inverts the sign of every later period, so a
# gain on the underlying reduces the balance and the series after that point is
# not a return series. Measured on the `us_firm_characteristics` registry,
# 2026-09-07: 46 registered runs hold a period return below -100%, and the worst
# of them reports sharpe 1.547 and a positive cagr against a total return of
# -202.3.
#
# The engine models no creditor, so the largest loss it can express is the
# capital: the period that would take equity through zero is truncated to
# exactly -100%, and every later period is 0.0, because there is nothing left to
# trade. That is the "floor equity at zero" option of the three the issue offers.
# The other two need something the case studies do not declare - a margin rule
# needs a maintenance requirement and a borrow rate, and halting the path without
# flooring it leaves a shorter series that still compares against full-length
# ones as though the difference were performance.
#
# Descriptive statistics of the stopped path are then true statements:
# `total_return`, `max_drawdown` and `cagr` all read -1.0, and a mean drawdown
# over a sweep is back inside [-1, 0] rather than the -8.15 a sweep of
# twenty-four with five insolvent members reported. The risk-adjusted ratios are
# not statements about the stopped path at all. A ratio of a mean to a dispersion
# describes a process that continues; this one ended. So they are registered as
# NULL - `BacktestExplorer.best()` already reads `sharpe IS NOT NULL` as "not
# rankable", and SQL `AVG` and `ORDER BY ... DESC` both keep NULLs out of the way
# on their own. Nothing downstream has to know about ruin to stop ranking it.
#
# `ruin` and `ruin_period` are registered as metrics, never as identity fields:
# they say what happened in a run, not what was asked for.
#
# The unrankable value is NaN and not None, which is a decision about the readers
# rather than about the metric. SQLite stores a NaN as NULL, so the registry and
# every SQL reader see exactly the intended "no value" - `ORDER BY sharpe DESC`
# puts it last, `AVG` skips it, and the fifteen `sharpe IS NOT NULL` filters drop
# it. A None would reach the notebooks that format a metric as `f"{sharpe:.3f}"`
# as a TypeError several frames from the cause; NaN formats as `nan`, which is
# what those lines should print.
#
# The cost is that polars sorts a NaN *first* on a descending sort, so anything
# ranking these dicts in memory before they reach SQLite has to say what it
# wants. `rank_returns_on_common_support` is the one place that does, and it
# converts to null and passes `nulls_last=True`.
RUIN_UNRANKABLE_METRICS: tuple[str, ...] = (
    "sharpe",
    "sortino",
    "calmar",
    "omega",
    "stability",
    "tail_ratio",
)


def first_ruin_index(returns) -> int | None:
    """Index of the first period whose loss takes cumulative equity to zero.

    ``None`` when the path stays solvent. A period return below -100% is not by
    itself ruin - prior gains can absorb it - so the test is on the equity curve
    rather than on any single return.
    """
    arr = np.asarray(returns, dtype=float)
    if arr.size == 0:
        return None
    equity = np.cumprod(1.0 + arr)
    # Non-finite equity is a broken series rather than a bankrupt one, and
    # silently reporting it as ruin would hide the breakage. It is excluded here
    # and left to `_safe` downstream.
    hit = np.flatnonzero(np.isfinite(equity) & (equity <= 0.0))
    return int(hit[0]) if hit.size else None


def apply_ruin_stop(returns) -> tuple[np.ndarray, int | None]:
    """Stop a return series at the period that wipes the account out.

    Returns ``(series, ruin_index)``. The ruin period is truncated to exactly
    -1.0 and every later period is set to 0.0. A solvent series is returned
    unchanged with ``None``. Idempotent: a stopped series reports the same index
    and is not changed again, so a caller that applies this to both the return
    frame and the array it computes metrics from cannot make the two disagree.
    """
    arr = np.asarray(returns, dtype=float)
    index = first_ruin_index(arr)
    if index is None:
        return arr, None
    stopped = arr.copy()
    stopped[index] = -1.0
    stopped[index + 1 :] = 0.0
    return stopped, index


def stop_returns_at_ruin(
    daily_returns: pl.DataFrame, column: str = "daily_return"
) -> tuple[pl.DataFrame, int | None]:
    """`apply_ruin_stop` over a ``[timestamp, daily_return]`` frame.

    The engine paths also emit an equity curve, a trade log and a fill log. Those
    are the broker's raw record of what it did, including the trade that took the
    account out, and they are deliberately left alone: the stop is a statement
    about what the account had left to compound, not a claim that the fills did
    not happen.
    """
    stopped, index = apply_ruin_stop(daily_returns[column].to_numpy())
    if index is None:
        return daily_returns, None
    return daily_returns.with_columns(pl.Series(column, stopped)), index


def _apply_ruin_semantics(
    out: dict[str, float | None], ruin_period: int | None, *, uncertainty_columns: bool = True
) -> dict[str, float | None]:
    """Stamp `ruin`/`ruin_period` and, for a stopped path, what it actually reports.

    The three descriptive metrics are stated here rather than read back from the
    diagnostic library. A stopped path ends at exactly zero equity by
    construction, so total return, CAGR and maximum drawdown are known to be
    -1.0 - and the library does not always produce them. On a path that ruins in
    its first period the drawdown series divides by a running maximum of zero and
    comes back all-NaN, which `_safe` turns into 0.0; a single-period path takes
    the short branch in the caller, where every value is 0.0. Both reported ruin
    beside a drawdown of nothing.
    """
    out["ruin"] = 0.0 if ruin_period is None else 1.0
    out["ruin_period"] = None if ruin_period is None else float(ruin_period)
    if ruin_period is None:
        return out
    out["total_return"] = -1.0
    out["max_drawdown"] = -1.0
    out["cagr"] = -1.0
    # Every ratio that ranks one path against another, and every confidence band
    # drawn around one, is undefined for a path that ended.
    out.update(dict.fromkeys(RUIN_UNRANKABLE_METRICS, float("nan")))
    if uncertainty_columns:
        from case_studies.utils.registry.store import _BACKTEST_UNCERTAINTY_COLUMNS

        out.update(dict.fromkeys(_BACKTEST_UNCERTAINTY_COLUMNS, float("nan")))
    return out


def compute_portfolio_metrics(
    returns: np.ndarray,
    *,
    periods_per_year: int = 252,
    case_study: str | None = None,
    label: str | None = None,
    uncertainty: bool = True,
    uncertainty_n_boot: int = 1000,
    uncertainty_seed: int = 0,
    trim_leading_zeros: bool = False,
) -> dict[str, float | None]:
    """Compute portfolio metrics using ml4t-diagnostic.

    Replaces hand-rolled Sharpe/drawdown/etc. with the library's
    validated implementation. When ``uncertainty=True`` (default) the returned
    dict is extended with block-bootstrap CIs, Lo/LdP-2025 Sharpe SE,
    Newey-West HAC SE for annualized return, and PSR p-value — driven by
    :func:`case_studies.utils.uncertainty.compute_backtest_uncertainty`.

    Parameters
    ----------
    returns : np.ndarray
        Array of period returns (daily or per-rebalance).
    periods_per_year : int
        Annualization factor (252 for daily, 52 for weekly, etc.).
    case_study, label : optional
        Used by the block-length resolver to pick rebalance_step from setup.yaml.
    uncertainty : bool, default True
        If False, skip the bootstrap (fast path for sweep inner loops).
    uncertainty_n_boot, uncertainty_seed : int
        Bootstrap configuration.
    trim_leading_zeros : bool, default False
        Legacy first-non-zero strip. Kept for callers that pass pre-canonical
        return series (e.g., raw engine output without canonical-window slice).
        Production callers (``_run_engine`` and the retrofit pipeline) pass
        ``False`` because they slice to the canonical (cs, label, split) window
        first, which preserves real "no-trade" days at the start of the window
        as legitimate zero-return periods rather than stripping them.

    Returns
    -------
    dict[str, float | None]
        Metric name → value. Keys match the existing backtest_metrics schema,
        plus uncertainty columns when ``uncertainty=True``, plus ``ruin`` and
        ``ruin_period``. A path stopped at ruin carries ``None`` for every
        ranking metric and every uncertainty column - see
        ``RUIN_UNRANKABLE_METRICS``.
    """
    from ml4t.diagnostic.evaluation import PortfolioAnalysis

    if trim_leading_zeros and len(returns) > 0:
        nonzero = np.flatnonzero(np.asarray(returns) != 0.0)
        if len(nonzero) > 0:
            returns = returns[nonzero[0] :]

    # Stop the path at the period that takes equity through zero, before anything
    # is measured off it. See RUIN_UNRANKABLE_METRICS above for what a stopped
    # path reports and why.
    returns, ruin_period = apply_ruin_stop(returns)

    if len(returns) < 2:
        return _apply_ruin_semantics(
            {
                "sharpe": 0.0,
                "sortino": 0.0,
                "total_return": 0.0,
                "max_drawdown": 0.0,
                "cagr": 0.0,
                "calmar": 0.0,
                "volatility": 0.0,
                "win_rate": 0.0,
                "omega": 0.0,
                "var_95": 0.0,
                "cvar_95": 0.0,
                "stability": 0.0,
                "skewness": 0.0,
                "kurtosis": 0.0,
                "tail_ratio": 0.0,
                "n_periods": int(len(returns)),
            },
            ruin_period,
            uncertainty_columns=False,
        )

    analysis = PortfolioAnalysis(returns=returns, periods_per_year=periods_per_year)
    with warnings.catch_warnings():
        warnings.filterwarnings(
            "ignore",
            message="Precision loss occurred in moment calculation",
            category=RuntimeWarning,
        )
        pm = analysis.compute_summary_stats()

    def _safe(v: float) -> float:
        """Sanitize metric value: handle complex, inf, nan."""
        if isinstance(v, complex):
            v = v.real
        if not np.isfinite(v):
            return 0.0
        return float(v)

    out = {
        "sharpe": _safe(pm.sharpe_ratio),
        "sortino": _safe(pm.sortino_ratio),
        "total_return": _safe(pm.total_return),
        "max_drawdown": _safe(pm.max_drawdown),
        "cagr": _safe(pm.annual_return),
        "calmar": _safe(pm.calmar_ratio),
        "volatility": _safe(pm.annual_volatility),
        "win_rate": _safe(pm.win_rate),
        "omega": _safe(pm.omega_ratio),
        "var_95": _safe(pm.var_95),
        "cvar_95": _safe(pm.cvar_95),
        "stability": _safe(pm.stability),
        "skewness": _safe(pm.skewness),
        "kurtosis": _safe(pm.kurtosis),
        "tail_ratio": _safe(pm.tail_ratio),
        "n_periods": int(len(returns)),
    }

    out = _apply_ruin_semantics(out, ruin_period)
    if ruin_period is not None:
        return out

    if uncertainty and len(returns) >= 4:
        try:
            from case_studies.utils.uncertainty import compute_backtest_uncertainty

            unc = compute_backtest_uncertainty(
                returns,
                periods_per_year=periods_per_year,
                case_study=case_study,
                label=label,
                n_boot=uncertainty_n_boot,
                seed=uncertainty_seed,
            )
            out.update(unc)
        except Exception as exc:  # pragma: no cover - never block point estimates
            warnings.warn(
                f"compute_backtest_uncertainty failed: {exc}; point metrics returned without CIs",
                stacklevel=2,
            )

    return out


# ---------------------------------------------------------------------------
# Result container
# ---------------------------------------------------------------------------


@dataclass
class BacktestRunResult:
    """Unified result from both engine and vectorized paths."""

    daily_returns: pl.DataFrame  # [timestamp, daily_return]
    metrics: dict[str, float]
    strategy_spec: dict
    prediction_hash: str
    backtest_hash: str | None = None
    # Engine-only fields
    engine_result: Any = None  # BacktestResult from ml4t-backtest
    weights: pl.DataFrame | None = None
    execution_mode: str = "engine"


# ---------------------------------------------------------------------------
# Risk-control trigger counts
# ---------------------------------------------------------------------------

# The controls the engine knows how to build, and therefore the ones it can count.
# The names are the `type` values a spec declares under `strategy.risk`.
RISK_POSITION_RULE_TYPES = ("stop_loss", "trailing_stop", "time_exit")
RISK_PORTFOLIO_LIMIT_TYPES = ("max_drawdown", "daily_loss")
RISK_CONTROL_TYPES = RISK_POSITION_RULE_TYPES + RISK_PORTFOLIO_LIMIT_TYPES


class RiskTriggerLog:
    """How many times each declared risk control acted during one backtest.

    A risk overlay that never fired was indistinguishable from one that was never
    installed, and the two have opposite meanings. The
    registry carried `num_trades` and the performance metrics and nothing else, so
    a reader comparing an overlay against the strategy it was laid on could
    conclude "the control acted" from a difference and nothing at all from a
    match: a stop can close a position the next rebalance would have closed
    anyway, leaving the trade count where it was. `rules/notebook-standards.md`
    C17 records the failure that hides behind that - 56 registered
    `crypto_perps_funding/16_risk_management` results whose Sharpe, drawdown and
    trade count matched the unprotected book in every digit, because the
    configuration declared its controls in a shape the engine does not read and
    nothing was installed.

    So the counts distinguish three states, and the third is the one that was
    missing:

    * nothing declared - every count is ``None``, which is what "no overlay" means;
    * declared and never fired - ``0``;
    * declared and fired - the number of times.

    A fourth state, declared but not installable on this execution path, is not
    represented because it is refused instead: registering a row named for a
    control the path cannot apply is the C17 failure itself.

    Counts are per control *type*, not per declared control: two stops in one spec
    share a key and their firings sum. Every sweep in the tree declares one
    control per backtest, so the distinction has no reader today.
    """

    def __init__(self) -> None:
        self._counts: dict[str, int] = {}

    def declare(self, control: str) -> None:
        """Record that this control was installed, before it has fired."""
        self._counts.setdefault(control, 0)

    def record(self, control: str, n: int = 1) -> None:
        self._counts[control] = self._counts.get(control, 0) + n

    def as_metrics(self) -> dict[str, float | None]:
        """One registered column per control type, plus the total.

        Every column is ``None`` when no control was declared, so a NULL reads as
        "no overlay here" rather than as "not measured".
        """
        if not self._counts:
            return {"risk_triggers": None} | {
                f"risk_triggers_{name}": None for name in RISK_CONTROL_TYPES
            }
        return {"risk_triggers": float(sum(self._counts.values()))} | {
            f"risk_triggers_{name}": (float(self._counts[name]) if name in self._counts else None)
            for name in RISK_CONTROL_TYPES
        }


class _CountedPositionRule:
    """Delegate to a position rule and count the exits it produces.

    Structural, not inherited: ``ml4t.backtest.risk.PositionRule`` is a Protocol
    and ``RuleChain`` is a plain dataclass over ``rules``, so a wrapper with
    ``evaluate`` composes wherever the real rule does.

    Only exits are counted - ``EXIT_FULL`` and ``EXIT_PARTIAL``. ``HOLD`` and
    ``ADJUST_STOP`` are not firings: moving a stop level is the control working
    without acting on the book, and counting it would report a rule that closed
    nothing as the most active one in the sweep. No rule in ml4t-backtest 0.1.3
    returns ``ADJUST_STOP``, so the exclusion is for rules added later.
    """

    def __init__(self, rule, control: str, log: RiskTriggerLog) -> None:
        self._rule = rule
        self._control = control
        self._log = log
        log.declare(control)

    def evaluate(self, state):
        from ml4t.backtest.risk import ActionType

        action = self._rule.evaluate(state)
        if action.action in (ActionType.EXIT_FULL, ActionType.EXIT_PARTIAL):
            self._log.record(self._control)
        return action


def _counted_portfolio_limit(limit, control: str, log: RiskTriggerLog):
    """Wrap a portfolio limit so each breach *episode* is counted once.

    ``RiskManager.update`` re-checks every limit on every bar and returns whatever
    is breached, so a halted book breaches again on each of the bars that follow.
    Counting those would report one drawdown halt as several hundred firings. The
    count moves on the rising edge only: not breached, then breached.
    """
    from ml4t.backtest.risk import PortfolioLimit

    class _CountedPortfolioLimit(PortfolioLimit):
        def __init__(self) -> None:
            self._limit = limit
            self._breached = False
            log.declare(control)

        def check(self, state):
            result = self._limit.check(state)
            breached = bool(result.breached)
            if breached and not self._breached:
                log.record(control)
            self._breached = breached
            return result

    return _CountedPortfolioLimit()


def _target_weights_by_timestamp(
    weights: pl.DataFrame,
) -> dict[date | datetime, dict[str, float]]:
    """Build deterministic timestamp and symbol ordered engine targets."""
    duplicate_count = weights.select(pl.struct("timestamp", "symbol").is_duplicated().sum()).item()
    if duplicate_count:
        raise ValueError(
            f"Target weights contain {duplicate_count} duplicate timestamp-symbol rows"
        )
    targets: dict[date | datetime, dict[str, float]] = {}
    for row in weights.sort("timestamp", "symbol").iter_rows(named=True):
        timestamp = row["timestamp"]
        if timestamp not in targets:
            targets[timestamp] = {}
        targets[timestamp][row["symbol"]] = row["weight"]
    return targets


def _engine_timestamp(
    value: object,
    *,
    feed_is_date: bool,
    feed_timezone: str | None,
    configured_timezone: str,
) -> date | datetime:
    if feed_is_date:
        if isinstance(value, datetime):
            return value.date()
        if isinstance(value, date):
            return value
        raise TypeError(f"engine target timestamp must be date-like, got {type(value).__name__}")
    if isinstance(value, datetime):
        timestamp = value
    elif isinstance(value, date):
        timestamp = datetime.combine(value, time.min)
    else:
        raise TypeError(f"engine target timestamp must be date-like, got {type(value).__name__}")
    if feed_timezone is not None:
        zone = ZoneInfo(feed_timezone)
        return (
            timestamp.replace(tzinfo=zone)
            if timestamp.tzinfo is None
            else timestamp.astimezone(zone)
        )
    if timestamp.tzinfo is not None:
        timestamp = timestamp.astimezone(ZoneInfo(configured_timezone)).replace(tzinfo=None)
    return timestamp


# ---------------------------------------------------------------------------
# Weight precomputation (for risk sweep reuse)
# ---------------------------------------------------------------------------


def precompute_weights(
    predictions: pl.DataFrame,
    strategy_spec: dict,
    prices: pl.DataFrame,
    *,
    label: str = "",
    case_study: str = "",
    prediction_hash: str | None = None,
    conformal_widths: pl.DataFrame | None = None,
) -> pl.DataFrame:
    """Compute allocation weights from a strategy spec, without running the engine.

    Use this to avoid redundant MVO/HRP computation in Ch19 risk sweeps
    where the same allocation weights are tested with different risk overlays.

    ``prediction_hash`` is required for the ``conformal_weighted`` allocation
    method (it loads that prediction's conformal widths); other methods ignore it.

    Returns
    -------
    pl.DataFrame
        Weights [timestamp, symbol, weight] ready for ``run_backtest(precomputed_weights=...)``.
    """
    predictions = normalize_prediction_columns(predictions)
    strategy = strategy_view(strategy_spec)
    signal_config = strategy["signal"]
    rebal_spec = strategy.get("rebalance", {})
    # Before ranking, not after allocating. `run_backtest` narrows the predictions it is
    # handed and then ranks them, so a caller that brings its own weights has to narrow at
    # the same point or the two paths build different portfolios under one identity. Doing
    # it to the finished weights is not the same operation and is worse than not doing it:
    # with `top_k=2` picking A and B, dropping B leaves a book half in cash, where narrowing
    # first picks A and C at the intended weight each.
    predictions = apply_traded_universe(predictions, prices, signal_config, case_study=case_study)
    weights = build_target_weights_from_config(predictions, signal_config)
    alloc_spec = strategy.get("allocation")
    if alloc_spec:
        cadence = strategy.get("rebalance", {}).get("cadence", "")
        weights = _apply_allocation(
            weights,
            predictions,
            prices,
            alloc_spec,
            cadence=cadence,
            label=label,
            case_study=case_study,
            prediction_hash=prediction_hash,
            conformal_widths=conformal_widths,
            rebalance_step=rebal_spec.get("step"),
        )
    return weights


# ---------------------------------------------------------------------------
# Strategy spec construction
# ---------------------------------------------------------------------------


# ---------------------------------------------------------------------------
# Prediction normalization
# ---------------------------------------------------------------------------


def normalize_prediction_columns(df: pl.DataFrame) -> pl.DataFrame:
    """Normalize prediction columns to canonical [timestamp, symbol, y_score, ...]."""
    renames = {}

    # Time column: date → timestamp
    if "timestamp" not in df.columns and "date" in df.columns:
        renames["date"] = "timestamp"

    # Entity column: asset/product/stock_id/entity → symbol
    if "symbol" not in df.columns:
        for col in ("asset", "product", "stock_id", "entity"):
            if col in df.columns:
                renames[col] = "symbol"
                break

    # Score column
    if "y_score" not in df.columns:
        if "prediction" in df.columns:
            renames["prediction"] = "y_score"

    if "y_true" not in df.columns and "actual" in df.columns:
        renames["actual"] = "y_true"
    if "fold_id" not in df.columns and "fold" in df.columns:
        renames["fold"] = "fold_id"

    if renames:
        df = df.rename(renames)

    # Cast types
    if "timestamp" in df.columns:
        ts_dtype = df.schema["timestamp"]
        if ts_dtype == pl.Date:
            df = df.with_columns(pl.col("timestamp").cast(pl.Datetime("us")))
        elif ts_dtype in (pl.String, pl.Utf8):
            df = df.with_columns(pl.col("timestamp").str.to_datetime().cast(pl.Datetime("us")))
        elif hasattr(ts_dtype, "time_zone") and ts_dtype.time_zone:
            df = df.with_columns(pl.col("timestamp").dt.replace_time_zone(None))

    if "symbol" in df.columns and df.schema["symbol"] != pl.String:
        df = df.with_columns(pl.col("symbol").cast(pl.String))

    return df


# Tolerant-by-design cap on (timestamp, symbol) join misses between a
# classification label and its continuous-return counterpart. >10% null
# rate indicates a regeneration mismatch between the two label parquets
# (not source-data sparsity), and is escalated to a hard error rather
# than silently dropping rows from the backtest. Callers operating in
# a legitimately high-null regime can override via the ``max_null_rate``
# parameter on ``substitute_continuous_return_for_classification``.
_MAX_NULL_RATE = 0.10

# Polars integer dtypes for symbol id columns (e.g., us_firm ``stock_id``).
# Used by ``_align_symbol_dtype`` to detect numeric-vs-string mismatches.
_INT_SYMBOL_DTYPES = (
    pl.UInt8,
    pl.UInt16,
    pl.UInt32,
    pl.UInt64,
    pl.Int8,
    pl.Int16,
    pl.Int32,
    pl.Int64,
)


def _align_symbol_dtype(
    target: pl.DataFrame,
    other: pl.DataFrame,
    *,
    case_study: str,
    target_side: str,
    other_side: str,
) -> pl.DataFrame:
    """Cast ``other['symbol']`` to ``target['symbol'].dtype``, failing loudly.

    Polars will silently raise ``InvalidOperationError`` when a string
    column with real tickers (e.g. ``"AAPL"``) is cast to integer — the
    error message names neither the case study nor the column origin,
    making diagnostics painful. This helper detects the pl.Utf8 ↔
    integer mismatch and surfaces a context-rich error before the cast,
    keeping the same behavior for compatible cases (same dtype, or
    same-kind cast).
    """
    target_dtype = target["symbol"].dtype
    other_dtype = other["symbol"].dtype
    if other_dtype == target_dtype:
        return other
    target_is_int = target_dtype in _INT_SYMBOL_DTYPES
    other_is_str = other_dtype in (pl.Utf8, pl.String)
    other_is_int = other_dtype in _INT_SYMBOL_DTYPES
    target_is_str = target_dtype in (pl.Utf8, pl.String)
    if target_is_int and other_is_str:
        # Probe: every value must parse as the target integer dtype.
        try:
            return other.with_columns(pl.col("symbol").cast(target_dtype))
        except Exception as exc:  # noqa: BLE001 — surface Polars's opaque error
            raise TypeError(
                f"_align_symbol_dtype: incompatible symbol representations for "
                f"case_study={case_study!r}: {target_side}.symbol is "
                f"{target_dtype} (numeric ids) but {other_side}.symbol is "
                f"{other_dtype} (likely tickers, not parseable as integer). "
                f"Underlying Polars error: {exc}"
            ) from exc
    if other_is_int and target_is_str:
        return other.with_columns(pl.col("symbol").cast(target_dtype))
    # Same-kind cast (e.g., Int32 → Int64, Utf8 → String alias).
    return other.with_columns(pl.col("symbol").cast(target_dtype))


def substitute_continuous_return_for_classification(
    predictions: pl.DataFrame,
    case_study: str,
    label: str,
    *,
    max_null_rate: float = _MAX_NULL_RATE,
) -> pl.DataFrame:
    """Replace binary y_true with the underlying continuous return for classification labels.

    The vectorized backtest computes ``gross_ret = weight * y_true``. For
    regression labels y_true is the forward return; for classification
    labels (fwd_class_*, fwd_dir_*) it is the binary class indicator, so
    the product collapses into a position-weighted accuracy proxy rather
    than economic P&L. We substitute y_true with the continuous return
    declared in setup.yaml::labels.classification_eval_label.

    Returns predictions unchanged when ``label`` is not registered as a
    classification target (i.e., regression labels pass through).
    """
    if not label:
        return predictions
    from pathlib import Path as _Path

    import yaml as _yaml

    from utils import CASE_STUDIES_DIR

    setup_path = _Path(CASE_STUDIES_DIR) / case_study / "config" / "setup.yaml"
    if not setup_path.exists():
        return predictions
    setup = _yaml.safe_load(setup_path.read_text())
    mapping = (setup.get("labels") or {}).get("classification_eval_label") or {}
    if label not in mapping:
        return predictions

    eval_label = str(mapping[label])
    # `get_case_study_dir`, not `CASE_STUDIES_DIR`. The setup file above is configuration and
    # lives in the repository; a label parquet is generated output and lives wherever
    # `ML4T_OUTPUT_DIR` puts it, which is what every other reader of one resolves. Reading it
    # from the checkout meant this path could only work where the artifacts happened to sit
    # beside the source: a run under output isolation raised FileNotFoundError for a label it
    # had just written. Measured in CI on `crypto_perps_funding` 13_backtest, where the first
    # classification label reached - `fwd_dir_8h` - looked for `fwd_ret_8h.parquet` under
    # /app/case_studies/... while the labels were in the isolated output root.
    from utils.paths import get_case_study_dir as _case_dir

    eval_path = _case_dir(case_study) / "labels" / f"{eval_label}.parquet"
    if not eval_path.exists():
        raise FileNotFoundError(
            f"Continuous-return label {eval_label!r} expected at {eval_path} "
            f"for classification label {label!r} but not found. Required so the "
            f"vectorized backtest can compute economic P&L instead of weight × binary."
        )
    eval_df = pl.read_parquet(eval_path).select(["timestamp", "symbol", eval_label])

    # Dedupe-assert eval_df on the join key before the left join. A duplicate
    # (timestamp, symbol) row in the continuous-return parquet would fan out
    # ``predictions`` silently, inflating downstream weight × y_true into a
    # wrong-but-plausible P&L (the very failure mode this function is meant
    # to prevent on the classification path).
    eval_h0 = eval_df.height
    eval_h_uniq = eval_df.unique(subset=["timestamp", "symbol"]).height
    if eval_h_uniq != eval_h0:
        raise ValueError(
            f"substitute_continuous_return_for_classification: continuous-return "
            f"label parquet at {eval_path} has {eval_h0 - eval_h_uniq} duplicate "
            f"(timestamp, symbol) rows ({eval_h_uniq} unique). Re-run the upstream "
            f"label step for case_study={case_study!r} to produce a unique-keyed "
            f"parquet."
        )
    eval_df = eval_df.unique(subset=["timestamp", "symbol"], keep="first")

    # Harmonize join-key dtypes to match the (already-normalized) predictions frame.
    if eval_df["timestamp"].dtype != predictions["timestamp"].dtype:
        if eval_df["timestamp"].dtype == pl.Date:
            eval_df = eval_df.with_columns(pl.col("timestamp").cast(pl.Datetime("us")))
        eval_df = eval_df.cast({"timestamp": predictions["timestamp"].dtype})
    eval_df = _align_symbol_dtype(
        predictions,
        eval_df,
        case_study=case_study,
        target_side="predictions",
        other_side=f"labels/{eval_label}.parquet",
    )

    pred_h0 = predictions.height
    joined = (
        predictions.drop("y_true")
        .join(eval_df, on=["timestamp", "symbol"], how="left")
        .rename({eval_label: "y_true"})
    )
    # Height-assert: ``left`` should never produce more rows than the left frame
    # carried in. Belt-and-suspenders for the dedupe assertion above.
    if joined.height != pred_h0:
        raise RuntimeError(
            f"substitute_continuous_return_for_classification: left join "
            f"changed row count {pred_h0} -> {joined.height} for case_study="
            f"{case_study!r} label={label!r}. eval_df keys are not unique "
            f"on (timestamp, symbol) after dedupe — internal invariant broken."
        )

    n_null = int(joined["y_true"].null_count())
    if n_null > 0:
        n_total = joined.height
        null_rate = n_null / n_total
        # Tolerant-by-design caps at ``max_null_rate`` (default
        # ``_MAX_NULL_RATE`` = 10%); above that, raise. >10% null rate
        # indicates a regeneration mismatch between the classification and
        # continuous-return parquets, not source-data sparsity.
        if null_rate > max_null_rate:
            raise ValueError(
                f"substitute_continuous_return_for_classification: "
                f"{n_null}/{n_total} ({null_rate:.2%}) predictions for "
                f"classification label {label!r} have no matching {eval_label!r} "
                f"value after join on (timestamp, symbol); exceeds "
                f"max_null_rate={max_null_rate:.2%}. Null rate above "
                f"{max_null_rate:.0%} indicates a regeneration mismatch "
                f"between the classification and continuous-return label "
                f"parquets; re-run the upstream label step for {case_study!r}."
            )
        warn_the_reader(
            f"{n_null}/{n_total} ({null_rate:.4%}) predictions for classification label "
            f"{label!r} have no matching {eval_label!r} value after join on "
            f"(timestamp, symbol); dropping those rows.",
            source="substitute_continuous_return_for_classification",
            key=(case_study, label, eval_label, n_null, n_total),
        )
        joined = joined.filter(pl.col("y_true").is_not_null())
    return joined


def _apply_cost_feasible_filter(
    predictions: pl.DataFrame,
    case_study: str,
    prediction_hash: str | None,
) -> pl.DataFrame:
    """Restrict predictions to the frozen, per-split cost-feasible universe.

    The split is resolved from the prediction set's registry entry; the
    symbol list is read from ``setup.yaml::universe.cost_feasible.{split}``.
    Raises if the split cannot be resolved or the list is absent — a silent
    full-universe fallback would change the registered result.
    """
    from pathlib import Path as _Path

    import yaml as _yaml

    from case_studies.utils.cv_window import lookup_split
    from utils import CASE_STUDIES_DIR

    if not prediction_hash:
        raise ValueError(
            "universe_filter='cost_feasible' requires a prediction_hash to "
            f"resolve the split for case_study={case_study!r}; got none."
        )
    split = lookup_split(case_study, prediction_hash)
    if split not in ("validation", "holdout"):
        raise ValueError(
            f"universe_filter='cost_feasible' could not resolve split for "
            f"prediction_hash={prediction_hash!r} (case_study={case_study!r}); "
            f"lookup_split returned {split!r}. The prediction set must be "
            f"registered with a 'validation' or 'holdout' split first."
        )
    setup = _yaml.safe_load(
        (_Path(CASE_STUDIES_DIR) / case_study / "config" / "setup.yaml").read_text()
    )
    symbols = (((setup.get("universe") or {}).get("cost_feasible")) or {}).get(split)
    if not symbols:
        raise KeyError(
            f"setup.yaml::universe.cost_feasible.{split} missing/empty for "
            f"case_study={case_study!r}; required when "
            f"signal.universe_filter='cost_feasible'."
        )
    filtered = predictions.filter(pl.col("symbol").is_in(list(symbols)))
    if filtered.is_empty() and not predictions.is_empty():
        raise ValueError(
            f"universe_filter='cost_feasible' produced an empty frame for "
            f"case_study={case_study!r} split={split!r}: the prediction set's "
            f"symbols do not intersect the frozen cost-feasible list (e.g. a "
            f"point-in-time ticker mismatch like FB/META). Refusing to run a "
            f"zero-row backtest — same 'no silent fallback' intent as above."
        )
    return filtered


def apply_universe_filter(
    predictions: pl.DataFrame,
    prices: pl.DataFrame,
    case_study: str,
    signal_config: dict | None,
    prediction_hash: str | None = None,
) -> pl.DataFrame:
    """Apply spec-declared universe restriction to predictions before backtest.

    When ``signal_config["universe_filter"] == "liquid"`` (sp500_options
    rung-3 in the O'Donovan-Yu / Muravyev-Pearson HTM cost cascade), the
    backtest must restrict each rebalance date to the tightest-quoted
    subset of the universe. The quantile lives in
    ``setup.yaml::backtest.sweep.htm_cost_cascade.liquid_quantile``; the
    spread column is ``instr_rel_spread`` on the prices frame.

    When ``signal_config["universe_filter"] == "cost_feasible"``
    (nasdaq100_microstructure), the backtest restricts predictions to a
    FROZEN, per-split symbol list committed under
    ``setup.yaml::universe.cost_feasible.{validation,holdout}``. The list is
    the cost-feasible universe — the cheapest-to-trade names by round-trip
    cost, profiled strictly before each window (no look-ahead — see
    ``build_cost_feasible_universe.py``); the split is resolved from the
    prediction set's registry entry via ``lookup_split``. Like ``liquid``,
    only the filter *name* enters the backtest hash, not the resolved symbols.

    Returns predictions unchanged when no filter applies. Built into
    ``run_backtest`` so any caller — sweep notebooks, the case studies' holdout
    notebooks, ad-hoc scripts — gets the same filter as the bespoke sp500_options
    pipeline, driven purely by the strategy spec.
    """
    if not signal_config:
        return predictions
    uf = str(signal_config.get("universe_filter", "")).strip().lower()
    if uf in ("", "full", "none"):
        return predictions
    if uf == "cost_feasible":
        return _apply_cost_feasible_filter(predictions, case_study, prediction_hash)
    if uf != "liquid":
        raise ValueError(
            f"universe_filter={uf!r} not supported. Allowed: 'liquid', 'cost_feasible', or 'full'."
        )
    if "instr_rel_spread" not in prices.columns:
        raise ValueError(
            f"universe_filter='liquid' requires 'instr_rel_spread' on the prices "
            f"frame for case_study={case_study!r}; got columns={list(prices.columns)}."
        )
    # Deliberately pinned to the repository copy rather than resolved through
    # get_case_study_dir/get_htm_cost_cascade: liquid_quantile is NOT part of the
    # backtest spec and so never enters backtest_hash. Letting an experiment copy
    # vary it would let two different quantiles collide on one hash against the
    # registry the experiment inherits. Making it experiment-editable therefore
    # requires plumbing it into strategy.signal and the hash first, and dropping
    # the hardcoded LIQUID_QUANTILE prefilters in sp500_options 12_backtest.py
    # and 13_portfolio_management.py - a hash-identity change, not a config change.
    from pathlib import Path as _Path

    import yaml as _yaml

    from utils import CASE_STUDIES_DIR

    setup = _yaml.safe_load(
        (_Path(CASE_STUDIES_DIR) / case_study / "config" / "setup.yaml").read_text()
    )
    cascade = (((setup.get("backtest") or {}).get("sweep") or {}).get("htm_cost_cascade")) or {}
    if "liquid_quantile" not in cascade:
        raise KeyError(
            f"setup.yaml::backtest.sweep.htm_cost_cascade.liquid_quantile missing for "
            f"case_study={case_study!r}; required when signal.universe_filter='liquid'."
        )
    liquid_quantile = float(cascade["liquid_quantile"])

    # Daily quantile of relative half-spread; ties broken with rank('min').
    # Collapse timestamp to the date grain before grouping so any caller
    # supplying sub-daily or unnormalized intraday bars still produces a
    # within-date rank rather than a within-bar rank (mirrors the bespoke
    # sp500_options sweep). Dedupe ``(date, symbol)`` to one row per
    # (date, symbol) — taking the min half-spread when multiple bars share
    # a date — so the rank denominator is symbol-count, not bar-count.
    half = (
        prices.select(
            pl.col("timestamp").cast(pl.Date).alias("_date"),
            pl.col("symbol"),
            (pl.col("instr_rel_spread") / 2).alias("_hs"),
        )
        .group_by(["_date", "symbol"])
        .agg(pl.col("_hs").min())
    )
    liquid_keys = (
        half.with_columns(
            (pl.col("_hs").rank("min").over("_date") / pl.col("_hs").count().over("_date")).alias(
                "_q"
            )
        )
        .filter(pl.col("_q") <= liquid_quantile)
        .select([pl.col("_date").alias("timestamp"), pl.col("symbol")])
    )
    if liquid_keys["timestamp"].dtype != predictions["timestamp"].dtype:
        # Predictions stamps are typically Datetime("us") at midnight; cast
        # back from Date so the semi-join key types match exactly.
        if predictions["timestamp"].dtype == pl.Datetime("us"):
            liquid_keys = liquid_keys.with_columns(pl.col("timestamp").cast(pl.Datetime("us")))
        else:
            liquid_keys = liquid_keys.cast({"timestamp": predictions["timestamp"].dtype})
    liquid_keys = _align_symbol_dtype(
        predictions,
        liquid_keys,
        case_study=case_study,
        target_side="predictions",
        other_side="prices",
    )
    return predictions.join(liquid_keys, on=["timestamp", "symbol"], how="semi")


# ---------------------------------------------------------------------------
# Core backtest function
# ---------------------------------------------------------------------------


def resolved_allow_short_selling(
    strategy_spec: dict,
    precomputed_weights: pl.DataFrame | None = None,
) -> bool:
    """Return the account short-selling flag implied by a resolved strategy spec.

    Identity-defining, so it has exactly one implementation: callers that hash a
    spec before handing it to :func:`run_backtest` must derive the flag through
    this function, or the spec they hashed and the spec that runs will differ.
    """
    signal_config = strategy_view(strategy_spec)["signal"]
    allow_short = bool(
        strategy_spec["backtest_config"]["account"].get("allow_short_selling", False)
    ) or bool(signal_config.get("long_short", False))
    allow_short = allow_short or (
        str(signal_config.get("direction", "long_only")).strip().lower() == "short_only"
    )
    if precomputed_weights is not None:
        allow_short = allow_short or bool(precomputed_weights.filter(pl.col("weight") < 0).height)
    return allow_short


def _restore_ruin_nans(metrics: dict) -> dict:
    """Put the NaNs back that SQLite turned into NULLs on the way in.

    `compute_portfolio_metrics` reports a stopped path's ranking metrics as NaN
    precisely so a notebook formatting `f"{sharpe:.3f}"` prints `nan` rather than
    raising on a None. SQLite stores that NaN as NULL, so the skip-if-complete
    branch - which reads the metrics back out of the registry rather than
    recomputing them - handed the None straight to those same lines. A cached
    bankrupt run therefore failed where a fresh one printed
    the neighbour of the defect above.

    Only where `ruin` says the NULL was written on purpose. Everywhere else a
    NULL means the metric was never computed, and inventing a NaN for it would
    claim a measurement that was not made.
    """
    if metrics.get("ruin") != 1.0:
        return metrics
    from case_studies.utils.registry.store import _BACKTEST_UNCERTAINTY_COLUMNS

    for name in (*RUIN_UNRANKABLE_METRICS, *_BACKTEST_UNCERTAINTY_COLUMNS):
        if name in metrics and metrics[name] is None:
            metrics[name] = float("nan")
    return metrics


def apply_traded_universe(
    predictions: pl.DataFrame,
    prices: pl.DataFrame,
    signal_config: dict | None,
    *,
    case_study: str,
) -> pl.DataFrame:
    """Narrow *predictions* to the universe the spec says this run trades.

     Returns them unchanged when ``signal_config`` declares no ``traded_universe``,
     which is every full run and every spec written before the key existed.

     A declaration is a claim about the panel, so it is checked against the panel
     rather than trusted: the digest here and the one the caller hashed both cover
     the sorted symbol list, so a panel that is not the one the caller declared stops
     the run instead of registering a result under an identity that describes a
     different portfolio. That is the whole point of the key - a reduced run must not
     be able to hash like the full run over the same predictions
    .

     The narrowing itself is what makes ``MAX_SYMBOLS`` reduce on the vectorized path,
     where ``gross_ret = weight * y_true`` is computed from the predictions and never
     consults ``prices``. On the engine path it is close to a no-op, because an
     unpriced name could not fill there anyway; doing it in both places keeps
     ``n_assets`` - which the notebooks read off the panel to decide which ``top_k``
     schemes are feasible - describing the cross-section the sweep actually ranks.
    """
    from case_studies.utils.backtest_presets import traded_universe_declaration

    declared = (signal_config or {}).get("traded_universe")
    if not declared:
        return predictions
    if "symbol" not in prices.columns:
        raise ValueError(
            f"{case_study}: the spec declares a traded universe but the price panel has no "
            f"'symbol' column; got columns={list(prices.columns)}"
        )
    panel = traded_universe_declaration(prices)
    if panel["digest"] != declared.get("digest"):
        raise ValueError(
            f"{case_study}: the price panel is not the universe this spec declares. "
            f"Declared {declared.get('n_symbols')} symbols "
            f"(digest {declared.get('digest')}), panel holds {panel['n_symbols']} "
            f"(digest {panel['digest']}). The declaration is part of backtest_hash, so "
            "running against a different panel would register this portfolio under "
            "another one's identity."
        )
    universe = prices.select("symbol").unique()
    if universe["symbol"].dtype != predictions["symbol"].dtype:
        universe = _align_symbol_dtype(
            predictions,
            universe,
            case_study=case_study,
            target_side="predictions",
            other_side="price panel",
        )
    return predictions.join(universe, on="symbol", how="semi")


def warn_if_the_panel_does_not_bound_the_universe(
    predictions: pl.DataFrame, prices: pl.DataFrame, *, case_study: str, label: str
) -> None:
    """Say when a price panel is not the universe a vectorized run will trade.

    The vectorized path computes ``gross_ret = weight * y_true`` from the
    predictions frame and reads ``prices`` only for the rebalance calendar, which
    is the same set of decision dates whichever symbols are in the panel. So
    ``load_backtest_prices_for(..., max_symbols=N)`` reduces the price frame and
    nothing else, and a preview taken at a reduced ``MAX_SYMBOLS`` is the
    production sweep with the production cost per backtest. Measured on
    us_firm_characteristics/11_backtest, 2026-08-24: 8 predictions x 4 schemes at
    300 symbols and at 3,708 gave bit-identical Sharpe, CAGR and drawdown across
    all 32 backtests, in 21 s against 19 s.

    This warns and does not act, which is a decision the engine cannot make on
    its own. Both ways of acting were tried and both are wrong here:

    * **Narrowing the predictions to the panel** makes the traded universe decide
      the portfolio without entering the backtest identity. The caller hashes its
      specification before this module sees the run -
      ``us_firm_characteristics/11_backtest.py:273`` computes
      ``backtest_hash_from_parts`` and skips a matching run - so a reduced preview
      would be served the full-universe result. Stamping the universe inside
      ``run_backtest`` is too late for that check, and a symbol count is not a
      universe in any case: ``{A, B}`` and ``{A, C}`` are different portfolios.
    * **Refusing the run** stops a preview that is legitimately configured this
      way. Measured on CI 2026-09-07: the `us_firm_characteristics` fixture holds
      a 5-symbol price panel against 20-symbol predictions, and refusing took
      four notebooks down.

    What closes this is the caller declaring the universe it trades, in the spec,
    before it hashes - which is a notebook change. Until then the parameter is
    inert here and this says so at the moment it happens.

    Said twice, on purpose. 22 of the 35 notebooks that call ``run_backtest`` or
    ``run_plumbing_test`` install a blanket ``warnings.filterwarnings("ignore")``
    at import, measured 2026-09-07 over ``case_studies/*/[0-9]*.py``::

        cme_futures                    1 of 1    fx_pairs                0 of 5
        etfs                           5 of 5    us_equities_panel       0 of 4
        nasdaq100_microstructure       5 of 5    crypto_perps_funding    1 of 5
        sp500_equity_option_analytics  5 of 5
        us_firm_characteristics        5 of 5

    Five case studies filter every backtest-calling notebook they have, two filter
    none, and ``crypto_perps_funding`` filters only ``18_holdout_backtest.py``. So
    a diagnostic that only warns is silent for the readers of five of the seven,
    and which of the seven this is being read in is not something this function
    can know. (76 of the 195 numbered case-study notebooks install it in total; a
    recursive grep also finds an archived notebook and the line you are reading,
    which is why that count comes back as 78.)
    The ``warnings`` call is what a library caller and the tests read; the print is
    what survives the filter and lands in the rendered cell. It is emitted once per
    (case study, label

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: MIT

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.