跳至正文
返回文库全部文档

使用计入成本的回测检验横截面信号

笔记本 《交易机器学习》

总结

本笔记本使用统一的回测引擎,将模型预测评估为 NASDAQ-100 交易策略。首先用随机信号进行流程检查:如果计入成本后仍持续盈利,就可能意味着存在前视、数据未对齐或成本未计入等问题。随后比较预测与入场规则的组合,包括等权重的高排名投资组合,以及通过保留仓位直至其被替代来限制换手率的规则。结果会被记录,以便后续进行统计分析,包括计算 Deflated Sharpe 并比较不同模型系列。

笔记本强调,仅凭信号排名质量无法决定交易结果。换手率、流动性、标的筛选和执行成本都会改变实际表现,因此应分别比较经过筛选和未筛选的标的范围,以及不同交易规则。组合模型被视为稳定预测的方法,而非保证带来更高收益的方法。需要注意,报告的 Deflated Sharpe 有效试验次数反映的是固定设计下的模型系列选择,并未计入上游完整的配置搜索。留出集留待后续评估,而不用于此处选择配置。

核心观点

  • 随机信号回测在计入成本后不应持续盈利,因此可用于检查流程错误。
  • 即使底层预测排名不变,控制换手率也可能显著改变结果。
  • 横截面排名良好,并不保证扣除换手率、流动性和执行成本后仍能盈利。
  • 分别比较标的筛选和换手率规则,以识别各自的影响。
  • 解读 Deflated Sharpe 结果时,应考虑有效试验次数估计包含了哪些选择步骤。

标签

全文
# NASDAQ-100 Microstructure: Backtest & Signal Evaluation


# NASDAQ-100 Microstructure: Backtest & Signal Evaluation

**Chapter 16 — Strategy Simulation**

This notebook translates Ch11–15 model outputs into backtested strategies for
the NASDAQ-100 microstructure case study: 15-minute decisions across the 113
stocks of the declared 115 that carry prices.
The backtest runs through the **ml4t-backtest engine** using 15-minute OHLCV
bars constructed from AlgoSeek TAQ trade prices — the same data that supports
position-level risk controls, realistic execution simulation, and proper cost
accounting in downstream chapters (Ch17–19).

1. **Plumbing test** — verify the engine pipeline produces no spurious alpha
2. **Parametric sweep** — test all (prediction × signal method) combinations
3. **Statistical analysis** — DSR, family comparison, IC-to-Sharpe relationship

Sections 1–2 generate new backtest results (write to registry). Section 3
is read-only — it queries the registry via `BacktestExplorer` and can be
re-run independently without re-running the sweep.

**Book Reference:** Chapter 16, Sections 16.4–16.8

**Prerequisites:** Completed model training (Ch11–15) for this case study.

```python
"""Ch16 Backtest & Signal Evaluation — NASDAQ-100 Microstructure case study."""

import json
import sqlite3
import time

import polars as pl

from case_studies.research import (
    SUPERSEDES_LIVE,
    OfficialPopulation,
    open_study,
    population_supersedes,
    prediction_rows_at,
    published_population_names_at,
    reuse_disclosure,
    superseded_members_at,
)
from case_studies.utils.backtest_loaders import get_backtest_config, load_backtest_prices_for
from case_studies.utils.backtest_presets import (
    build_backtest_spec,
    serializable_backtest_spec,
    traded_universe_declaration,
)
from case_studies.utils.backtest_runner import (
    normalize_prediction_columns,
    run_backtest,
    run_plumbing_test,
)
from case_studies.utils.ensemble import (
    ensemble_training_spec,
    load_ensemble_declaration,
    mean_forecast,
    resolve_members,
)
from case_studies.utils.notebook_contracts import (
    degenerate_prediction_hashes,
    excluded_families,
    prediction_members_in_force,
)
from case_studies.utils.registry import (
    backtest_hash_from_parts,
    load_existing_backtest_hashes,
    load_prediction_index,
    read_predictions,
)
from case_studies.utils.registry.registration import (
    register_prediction_set,
    register_training_run,
)
from case_studies.utils.sweep_config import (
    get_entry_schemes_for,
    get_signal_passes_for,
    get_top_k_values_for,
    get_top_n_predictions,
    get_universe_filters_for,
)
from utils.paths import get_case_study_dir
from utils.style import show_with_alt
```

```python
CASE_STUDY_ID = "nasdaq100_microstructure"
LABEL = ""
SPLIT = "validation"
# Zero means the smallest top_k from setup.yaml backtest.sweep.top_k_grid.
TOP_K = 0
MAX_SYMBOLS = 0
FORCE_REBACKTEST = False  # Set True to re-backtest even if a complete backtest_hash exists
TOP_N_PREDICTIONS = None
# Zero means every feasible entry scheme. A positive value keeps the first N, which is what makes
# a reduced run of this notebook finish in minutes: `backtest.sweep.signal_nasdaq100` crosses two
# selection methods, three quantiles, two directions, three slot counts, two hold windows and two
# signal-exit thresholds, and 42 of those combinations are feasible on this panel. A backtest of
# this panel was measured at 31.8 s on 2026-09-11, so the full 45 arms is about 34 minutes for a
# single prediction set. The same lever as `MAX_RISK_VARIANTS` in 16 and `MAX_COST_POINTS` in the
# cost notebooks. Truncating the list keeps the canonical equal-weight arms, which are what
# `signal_passes.baseline_schemes` names, so a reduced run still completes pass 1.
MAX_ENTRY_SCHEMES = 0
# Both names stay bound here although nothing below reads them: that is what makes the harness
# force preview and supply a workspace - `_declares_tier_and_workspace` in `tests/pm_helpers.py`
# looks for exactly this pair. Without them the canonical branch regenerates in place, which
# needs generated-artifact symlinks a CI checkout does not have.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
```

## 1. Setup & Plumbing Test

Before the sweep, check the machinery on a signal that carries no information.
Replacing the predictions with random scores and running the same engine should
not produce a profitable strategy. If it does, the profit came from the
pipeline rather than the signal - a lookahead in the fill timing, a misaligned
join, a cost model that never charges.

The check is one-sided. Random trading pays the spread on every rebalance
without any compensating edge, so a clearly negative result is the expected
outcome and not a failure. What would fail is profit that persists after
costs.

The study is opened before anything resolves a path or reads the registry. Opening it
activates a root and rewrites `ML4T_OUTPUT_DIR` process-wide, and every later
`get_case_study_dir`, prediction index and registry write resolves against that variable. A
`CASE_DIR` bound before this line points at the released registry while the sweep writes to
the workspace, and the two never meet: the sweep finds nothing registered and every reader
scoped to hashes from the other root comes back empty.

```python
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)

CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
bt_config = get_backtest_config(CASE_STUDY_ID)
if TOP_N_PREDICTIONS is None:
    TOP_N_PREDICTIONS = get_top_n_predictions(CASE_STUDY_ID, "signal")

if not LABEL:
    LABEL = bt_config.primary_label

print(
    f"""=== Protocol Term Sheet ===
  Case study:    {CASE_STUDY_ID}
  Label:         {LABEL}
  Calendar:      {bt_config.calendar}
  Cadence:       {bt_config.cadence}
  Commission:    {bt_config.commission_bps:.1f} bps
  Slippage:      {bt_config.slippage_bps:.1f} bps
  Total cost:    {bt_config.commission_bps + bt_config.slippage_bps:.1f} bps/leg
  Long/short:    {bt_config.long_short}
""",
    flush=True,
)
if excluded_families(CASE_STUDY_ID):
    print(
        "Active-model filter: excluding "
        f"{', '.join(sorted(excluded_families(CASE_STUDY_ID)))} pending corrected reruns",
        flush=True,
    )
```

```python
prices = load_backtest_prices_for(CASE_STUDY_ID, LABEL, split="validation", max_symbols=MAX_SYMBOLS)

# `MAX_SYMBOLS` reduces the price panel, and until the run says so in its own specification
# that reduction did not reach `backtest_hash`: a reduced run and the full run over the same
# predictions hashed alike, so the second was served the first's result and the reduction
# bought nothing (ml4t/agent-workspace#911). Declaring it here, before anything is hashed,
# gives a reduced run an identity of its own; `run_backtest` checks the panel against the
# declaration and narrows the predictions to it, so the sweep ranks the cross-section this
# says it ranks and `n_assets` above describes that same set. A full run declares nothing and
# is byte-identical to before.
# A reduced run is a preview run. Refused on the canonical tier so a narrowed result can
# never land in the registry the book's numbers come from, and so the two can never sit in
# one registry to be ranked against each other: `resolve_best_predictions` takes MAX(sharpe)
# over every backtest of a prediction, and a Sharpe earned over a handful of names would
# advance a configuration ahead of one earned over the whole panel. `us_equities_panel` 16
# through 19 already refuse the parameter this way, and `canonically_refused_parameters`
# reads the refusal out of the source, so the canonical fixture path drops the name rather
# than handing the notebook something its first cell raises on.
if EXECUTION_TIER == "canonical" and MAX_SYMBOLS:
    raise ValueError(
        "MAX_SYMBOLS narrows the universe this run trades, which makes it a different "
        "portfolio from the declared one and gives it its own backtest identity "
        "(ml4t/agent-workspace#911). A canonical run trades the declared universe: set "
        "MAX_SYMBOLS=0, or run under EXECUTION_TIER='preview' with a WORKSPACE."
    )
TRADED_UNIVERSE = traded_universe_declaration(prices) if MAX_SYMBOLS else None

n_assets = prices["symbol"].n_unique()
if TOP_K == 0:
    _feasible_top_k = get_top_k_values_for(CASE_STUDY_ID, LABEL, n_assets)
    if not _feasible_top_k:
        raise ValueError(
            f"top_k_grid for {LABEL!r} in {CASE_STUDY_ID} has no value < "
            f"n_assets={n_assets}; declare a feasible k in setup.yaml"
        )
    TOP_K = _feasible_top_k[0]
print(f"Prices: {len(prices):,} rows, {n_assets} assets; plumbing-test TOP_K={TOP_K}", flush=True)
```

```python
strategy_spec = build_backtest_spec(
    CASE_STUDY_ID,
    bt_config,
    prices=prices,
    traded_universe=TRADED_UNIVERSE,
    prediction_hash="plumbing_test",
    initial_cash=bt_config.initial_cash,
    chapter="ch16",
    # The cadence is per-label here (setup.yaml::decision.cadence_by_label), so a spec
    # built without one would be built against a default no run uses.
    label=LABEL,
    signal={
        "method": "score_weighted_top_k",
        "top_k": TOP_K,
        "long_short": bt_config.long_short,
    },
)

try:
    random_sharpe = run_plumbing_test(
        CASE_STUDY_ID,
        prices,
        strategy_spec,
        top_k=TOP_K,
        initial_cash=bt_config.initial_cash,
        calendar=bt_config.calendar,
    )

    status = "FAIL" if random_sharpe > 1.5 else "PASS"
    print(f"Random signal Sharpe: {random_sharpe:.3f}  [{status}]", flush=True)

    if random_sharpe > 1.5:
        print("WARNING: Random signal produces positive Sharpe — investigate pipeline", flush=True)
    elif random_sharpe < -1.5:
        print(
            "NOTE: Strongly negative random Sharpe reflects turnover drag under quote-aware costs",
            flush=True,
        )
except ValueError as e:
    if "zero variance" in str(e).lower():
        print(f"Plumbing test skipped: {e} (too few assets for meaningful test)", flush=True)
        random_sharpe = 0.0
    else:
        raise
```

## 2. The Signal Sweep

Every combination of prediction and entry scheme below runs through the same
`run_backtest()` call as a one-off backtest would, so the sweep and a single
backtest cannot diverge.

The sweep runs in two passes. The first trades every prediction three ways -
equal weight over the top 5, 10 and 20 names - on the cost-feasible universe,
which is the universe the strategy this chapter ends on actually trades. The
second takes the predictions that scored highest in the first and asks what the
entry mechanism does to them, across the slot grid declared in
`backtest.sweep.signal_nasdaq100`. The same three equal-weight arms are also run
without the cost screen, on those same predictions: that unscreened pair is what
Section 3 reads.

Ranking over the whole universe is the naive approach the feasibility analysis
warned against, and it is kept here as the comparison rather than as the main
grid. The ordering will often place its strongest views on the least liquid
names in the panel, and at this rebalancing frequency each of those positions is
entered and exited repeatedly.

The sweep also separates two things that are easily conflated: how well a
prediction orders the cross-section, and how much trading that ordering
provokes. A prediction whose scores change sharply from bar to bar produces
more position changes than one whose scores move smoothly, at the same
ordering quality. The cost of those changes is charged here and not in any
score.

A grid cell is skipped when its identity is already registered or already
queued by this sweep. Two schemes can resolve to one identity - a backtest is
defined by what it computes, not by the label its scheme carries - and without
the queued half of that test the first computes the result while the second
reuses its cache and is still counted as work done, so the sweep reports more
backtests than the registry holds.

```python
pred_index = load_prediction_index(
    CASE_STUDY_ID,
    label=LABEL,
    split=SPLIT,
)
if not pred_index.is_empty():
    # Exclude causal_dml (not a trading signal) and the synthetic ensemble
    # forecast — the ensemble is a cost-feasible-universe construct introduced
    # in Section 4 (Act 2), not part of the full-universe baseline sweep.
    pred_index = pred_index.filter(~pl.col("family").is_in(["causal_dml", "ensemble"]))

if pred_index.is_empty():
    msg = f"No predictions found for {CASE_STUDY_ID}/{LABEL}/{SPLIT}"
    raise RuntimeError(msg)

# `load_prediction_index` answers "what predictions exist in this registry" and leaves
# admissibility to the caller. A backtest is asking a different question - "what should
# I trade" - and the difference is exactly the conditions the catalog computes. Without
# this filter the sweep consumes rows the official population cannot contain: a run that
# reports every backtest completed and zero failed, off a population that resolves empty.
# That is not a late failure, it is a loud success on the wrong rows, and it costs the
# full sweep to discover.
# `complete` is the whole test: the catalog already requires a current identity before a row
# can be complete, and the tier is decided by which registry the rows were read from, not by
# a column comparison. Re-asserting either here would reject a preview run's own rows.
# The catalog is read off `CASE_DIR`, the directory `load_prediction_index` just read, and
# not by opening a `Study`: every `Study.open` branch ends in `activate()`, which would both
# answer for a different registry than the one being filtered and re-point the rest of the
# notebook - including where `run_backtest(register=True)` writes - at the activated root.
_catalog = prediction_rows_at(CASE_DIR)
# `complete` does not answer supersession, and the two look like the same question. A row is
# complete when its artifacts and fold metrics are present, and `identity_status == "current"`
# only says the registry still understands the schema it was written under. Neither moves when
# a model notebook refits: the retired generation's rows stay complete and current, and a sweep
# selecting on the catalog alone runs over both generations at once. It does not fail - it
# succeeds over twice the population and freezes the mixture into every backtest downstream.
#
# Measured against this registry on 2026-09-11: it records one supersedes edge, on
# `nasdaq100_microstructure-linear-validation-v1`. That name's retired generation lists 61
# prediction identities and its generation in force shares none of them, because a refit moves
# every hash. None of the 61 is in the catalog, so this join removes nothing today, and every
# one of the catalog's rows belongs to some declared population.
#
# A no-op is what this filter looks like whenever a refit takes the rows it retired out of the
# catalog with it, which is what the linear refit did. It stops being one the first time a
# generation is retired while its rows stay - and nothing here guarantees which of the two a
# given refit will be, which is the reason to hold the join rather than to decide per refit.
_retired = superseded_members_at(CASE_DIR)
_admissible = _catalog.filter(
    pl.col("complete") & ~pl.col("prediction_hash").is_in(list(_retired))
).select("prediction_hash")
_offered = len(pred_index)
_offered_hashes = set(pred_index["prediction_hash"])
pred_index = pred_index.join(_admissible, on="prediction_hash", how="inner")
if pred_index.is_empty():
    msg = (
        f"{_offered} prediction sets exist for {CASE_STUDY_ID}/{LABEL}/{SPLIT} but none is "
        "admissible. Either every row is missing an artifact or a fold metric, or carries an "
        "identity this schema version no longer recognises; or every row belongs to a "
        f"generation its own name has moved past ({len(_retired)} identities in this registry "
        "are retired that way). Backtesting them would produce a full sweep over a population "
        "that cannot be official. Re-run the model notebooks on the research interface first."
    )
    raise RuntimeError(msg)
if len(pred_index) < _offered:
    # Two conditions were tested and they call for different work, so they are counted apart: an
    # incomplete row needs its own fit finished, a superseded one needs nothing - it is a
    # retired generation and the sweep is right to leave it. Supersession is named first because
    # it decides the row on its own; completing a retired row would not readmit it.
    _dropped = _offered_hashes - set(pred_index["prediction_hash"])
    _dropped_superseded = len(_dropped & _retired)
    print(
        f"  Excluded {len(_dropped)} of {_offered} prediction sets: "
        f"{len(_dropped) - _dropped_superseded} not complete, "
        f"{_dropped_superseded} superseded by a later generation of their own population",
        flush=True,
    )

# Membership, and it runs here rather than beside the exclusion above for two reasons
# that are easy to get wrong. The admissibility refusal has to stay reachable: a pool
# emptied by the catalog filter must raise the error naming incompleteness and
# supersession, not this one, which would report an empty pool as unpublished. And the
# _dropped accounting above splits its exclusions into two counts by subtracting what
# survives, so a membership drop landing before it would be counted as not complete.
#
# Membership is a different set from the exclusion above, and the weaker one is what
# this sweep had (ml4t/agent-workspace#1088). "Not retired" admits a
# prediction no population ever listed - an experiment its case study never published is
# retired by nobody - while "published" does not. `published_members_at` subtracts the
# retired set itself, so this is never the looser test; it is the same shape
# `sp500_equity_option_analytics/14_backtest:292` already applies, and running the nine
# on two different rules is the thing worth removing even where the answer agrees.
#
# The two rules agree on the publication question here - the registry publishes 784
# prediction sets and holds 784 - but do not read that as the filter being inert. Measured
# on the 2026-09-13 canonical run, it narrowed the pool in two steps and both bit:
#
#     220 of 784 members dropped for covering less of the cross-section than their
#         feature panels offered them
#     564 prediction sets in the populations in force
#     80 further sets excluded on this label as complete, not retired, and published
#         by no population in force
#     162 predictions swept
#
# The first named drop was a `deep_learning/lstm_h64` set carrying 81.2% of the (entity,
# session) pairs its input panel offered it. Without this filter the previous sweep's
# pass-2 eight was four `nlinear` checkpoints at 67.4% coverage and four L1-linear sets
# that are constants - two distinct forecast values across 8,006,995 rows - so 180 of its
# 360 mechanism backtests priced a 45-arm entry grid over predictions that rank nothing.
#
# `None` means the registry declares no populations at all - a fixture, or a clean clone -
# and there is then nothing to filter against. Left unscoped in that case rather than
# narrowed to nothing, which is what `prediction_members_in_force` returns `None` to say.
#
# It is slow, and silently so: it opens every published member's artifact to check symbol
# coverage, which on this registry is 784 parquet files totalling 94.3 GB. Measured on the
# 2026-09-13 canonical run at **8,608 s - 2 h 23 m, 11.0 s per member** - before the
# sweep's first backtest, with nothing printed while it runs because every print in this
# cell comes after the call returns. That is the cost of the check rather than a hang, and
# it is said here because the next reader watching a quiet log is the one who needs to
# know. It is also label-agnostic: it coverage-checks every published member whatever the
# caller is sweeping, so two thirds of those reads are for labels this run never touches.
# Scoping it to the candidates the caller can act on is a change to the shared helper and
# not to this notebook, which is why it is described here rather than made here.
_members, _population_notes = prediction_members_in_force(study, CASE_DIR)
for _note in _population_notes:
    print(f"  {_note}", flush=True)
if _members is not None:
    print(f"  {len(_members):,} prediction sets in the populations in force", flush=True)
    _before_membership = set(pred_index["prediction_hash"])
    pred_index = pred_index.filter(pl.col("prediction_hash").is_in(list(_members)))
    _unpublished = _before_membership - set(pred_index["prediction_hash"])
    if _unpublished:
        print(
            f"  Excluded {len(_unpublished)} further prediction sets that are complete and "
            "not retired but that no population in force publishes",
            flush=True,
        )
    if pred_index.is_empty():
        msg = (
            f"{len(_before_membership)} prediction sets for {CASE_STUDY_ID}/{LABEL}/{SPLIT} "
            "are complete and not retired, and no population in force publishes any of "
            f"them. The registry publishes {len(_members):,} prediction identities under "
            f"{len(published_population_names_at(CASE_DIR))} population names; this label's "
            "rows are in none of them, so the model notebook that wrote them registered "
            "them outside the population mechanism. Re-run it before sweeping."
        )
        raise RuntimeError(msg)

# The pool the sweep is entitled to draw on, captured HERE: after the membership and
# coverage filter above, and before the preview truncation below. The ensemble reads it
# rather than the catalog-level `_admissible`, which is only "complete and not retired"
# and so predates both. Captured before the truncation because `TOP_N_PREDICTIONS` narrows
# what a preview run *sweeps*, not what the case study is allowed to average.
SWEPT_POOL = set(pred_index["prediction_hash"])

if TOP_N_PREDICTIONS > 0:
    pred_index = pred_index.head(TOP_N_PREDICTIONS)

n_predictions = len(pred_index)
print(f"Predictions to sweep: {n_predictions}", flush=True)
ic_min, ic_max = pred_index["ic_mean"].min(), pred_index["ic_mean"].max()
print(
    f"  IC range: {ic_min:.4f} — {ic_max:.4f}"
    if ic_min is not None
    else "  IC range: not yet computed",
    flush=True,
)
```

```python
# Passed as the cross-section width when the question is what setup.yaml declares rather than
# what this panel can trade: it is above every concentration any grid here declares, so
# `get_entry_schemes_for` drops nothing for feasibility and returns the declared names.
NO_FEASIBILITY_LIMIT = 1_000_000

entry_schemes = get_entry_schemes_for(
    CASE_STUDY_ID, LABEL, n_assets, long_short=bt_config.long_short
)
if MAX_ENTRY_SCHEMES:
    entry_schemes = entry_schemes[:MAX_ENTRY_SCHEMES]

# The universe axis, and which arms are crossed with which of its values.
#
# `backtest.sweep.universe_filter` declares `cost_feasible` here, and `universe.cost_feasible`
# names the 50 validation and 50 holdout symbols `apply_universe_filter` restricts to. Nothing
# in this notebook read either: every registered spec omitted `signal.universe_filter`, so
# `BacktestExplorer` defaulted all of them to "full". Act 1 worked by that accident, and the
# two readers that ask for the declared value got nothing - Section 4 below printed an empty
# table and exited 0, `17_costs` priced nothing and exited 0, and only `20_strategy_analysis`
# failed. Measured 2026-09-09: 202 signal backtests registered, 0 carrying the key.
_declared_universes = [u for u in get_universe_filters_for(CASE_STUDY_ID) if u]
universe_filters: list[str | None] = [None, *_declared_universes]

# The grid is walked in two passes, declared under `backtest.sweep.signal_passes`.
#
# Crossing every arm with every universe over every prediction is what this cell used to
# build, and on this case study that is 741 prediction sets × 117 arms × 2 universes =
# 173,394 backtests. One backtest of this panel was measured at 31.8 s on 2026-09-11, which
# puts the cross-product at about 1,500 hours. No other case study in the book declares more
# than four entry schemes or more than one universe.
#
# Pass 1 asks which predictions are worth studying, on three equal-weight concentrations and
# on the universe the carrier trades. Pass 2 asks what the entry mechanism does to them, and
# asks it only of the predictions pass 1 ranked highest. The two questions have different
# widths, and one cross-product answered both at the width of the wider one.
_passes = get_signal_passes_for(CASE_STUDY_ID)
_by_name = {s["name"]: s for s in entry_schemes}

if _passes is None:
    # No plan declared: the full cross-product, which is what every other case study runs.
    baseline_arms = [(s, u) for u in universe_filters for s in entry_schemes]
    mechanism_arms: list[tuple[dict, str | None]] = []
    mechanism_top_n = 0
else:
    # A name in `baseline_schemes` goes missing for two reasons and only one of them is a
    # defect. `get_entry_schemes_for` drops a concentration the cross-section cannot realize,
    # so any run narrower than the declared grid loses names by construction. The CI fixture
    # trades twelve symbols long-short, and a long-short selection needs 2k distinct names,
    # so `ew_top5` is realizable there and neither `ew_top10` nor `ew_top20` is - measured by
    # calling this function at n_assets=12, which reproduces that run's arm list exactly. A
    # name no declared grid produces at any width is a typo in setup.yaml.
    #
    # Asking the same function again with the feasibility rule lifted separates the two without
    # restating anywhere how a scheme is named. `MAX_ENTRY_SCHEMES` was the previous test and it
    # answers a different question - it is one of several reductions, and the fixture reaches
    # here with it unset - so a narrowed panel raised the typo error against a correct config.
    _declared_names = {
        s["name"]
        for s in get_entry_schemes_for(
            CASE_STUDY_ID,
            LABEL,
            NO_FEASIBILITY_LIMIT,
            long_short=bt_config.long_short,
            ranked_width=NO_FEASIBILITY_LIMIT,
        )
    }
    _undeclared = [n for n in _passes["baseline_schemes"] if n not in _declared_names]
    if _undeclared:
        msg = (
            f"backtest.sweep.signal_passes.baseline_schemes names {_undeclared}, which "
            f"get_entry_schemes_for does not produce for {LABEL} at any cross-section width. "
            f"Names the declared grids do produce: {sorted(_declared_names)[:8]}"
        )
        raise KeyError(msg)
    _baseline_names = [n for n in _passes["baseline_schemes"] if n in _by_name]
    if not _baseline_names:
        # Every declared baseline concentration is above what this run can rank. Pass 1 would
        # sweep nothing, rank nothing and report a completed run, which is the failure the
        # check above used to catch by accident.
        msg = (
            f"none of backtest.sweep.signal_passes.baseline_schemes "
            f"{_passes['baseline_schemes']} is feasible for {LABEL} on this run: "
            f"{n_assets} assets in the price panel, long_short={bt_config.long_short}, and "
            f"every declared concentration needs more names than that leaves. Declare a "
            f"smaller concentration, or widen the panel and the fitting stages together."
        )
        raise RuntimeError(msg)
    baseline_arms = [(_by_name[n], _passes["baseline_universe"]) for n in _baseline_names]
    mechanism_arms = [
        (s, _passes["baseline_universe"]) for s in entry_schemes if s["name"] not in _baseline_names
    ]
    # The reference universe carries the canonical concentrations only. It is what `17_costs`
    # reads for its full-vs-screened comparison and what Act 1 below reads for the unscreened
    # baseline. It is not a second copy of the sweep.
    mechanism_arms += [
        (_by_name[n], _passes["reference_universe"])
        for n in _passes["reference_schemes"]
        if n in _by_name
    ]
    mechanism_top_n = _passes["mechanism_top_n"]

print(f"\nEntry schemes ({len(entry_schemes)}):", flush=True)
for es in entry_schemes:
    print(f"  {es['name']}: {es['method']} (top_k={es.get('top_k', '-')})", flush=True)
print(
    "Universes: " + ", ".join("full" if u is None else str(u) for u in universe_filters),
    flush=True,
)

n_pass2_predictions = min(mechanism_top_n, n_predictions) if mechanism_arms else 0
n_pass1 = n_predictions * len(baseline_arms)
n_pass2 = n_pass2_predictions * len(mechanism_arms)
total_backtests = n_pass1 + n_pass2
print(
    f"\nPass 1 (baseline): {n_predictions} predictions × {len(baseline_arms)} arms "
    f"= {n_pass1} backtests",
    flush=True,
)
print(
    f"Pass 2 (mechanism): top {n_pass2_predictions} by pass-1 Sharpe × "
    f"{len(mechanism_arms)} arms = {n_pass2} backtests",
    flush=True,
)
print(f"Total grid: {total_backtests} backtests", flush=True)
```

```python
t0 = time.time()
tally = {"completed": 0, "skipped": 0, "failed": 0}
existing_hashes = load_existing_backtest_hashes(CASE_STUDY_ID, stage="signal")
# Identities already registered, plus the ones this sweep has queued. Both mean
# "running this grid cell would add nothing", which is what the skip test needs.
planned = set(existing_hashes)
print(f"Existing signal-stage hashes in registry: {len(existing_hashes):,}", flush=True)


def run_arms(rows, arms, pass_name):
    """Run one pass: every arm in `arms` against every prediction in `rows`.

    Both passes walk the same loop, so what separates them is only which
    predictions and which arms they are handed. A pass that is given no arms
    returns nothing rather than raising: a preview run truncates the scheme list
    and can legitimately leave pass 2 empty.
    """
    out = []
    n_cells = len(rows) * len(arms)
    if not n_cells:
        print(f"\n{pass_name}: no grid cells", flush=True)
        return out
    print(
        f"\n{pass_name}: {len(rows)} predictions × {len(arms)} arms = {n_cells} cells", flush=True
    )

    for i, pred_row in enumerate(rows.iter_rows(named=True)):
        pred_hash = pred_row["prediction_hash"]
        source = pred_row["source"]
        ic_mean = pred_row["ic_mean"]

        pending_schemes = []

        for j, (scheme, universe) in enumerate(arms):
            idx = i * len(arms) + j + 1

            signal = {
                "method": scheme["method"],
                "top_k": scheme.get("top_k", 20),
                "long_short": bt_config.long_short,
            }
            signal.update({k: v for k, v in scheme.items() if k not in ("name", "method")})
            # Only when a filter is declared. A `None` written into the spec is not the same as
            # an absent key: it would move `backtest_hash` for every row registered before this
            # axis existed, and `get_universe_filters_for` normalizes "full" and "none" to None
            # for that reason.
            if universe is not None:
                signal["universe_filter"] = universe
            spec = build_backtest_spec(
                CASE_STUDY_ID,
                bt_config,
                prices=prices,
                traded_universe=TRADED_UNIVERSE,
                prediction_hash=pred_hash,
                initial_cash=bt_config.initial_cash,
                chapter="ch16",
                label=LABEL,
                signal=signal,
            )
            backtest_hash = backtest_hash_from_parts(pred_hash, serializable_backtest_spec(spec))

            if backtest_hash in planned:
                # Recorded, not just counted. A skipped cell is a cell of this pass whose
                # result is already in the registry, and pass 2 ranks on the pass-1 cells
                # rather than on the subset this invocation happened to compute.
                tally["skipped"] += 1
                out.append(
                    {
                        "prediction_hash": pred_hash,
                        "source": source,
                        "family": pred_row["family"],
                        "config_name": pred_row["config_name"],
                        "ic_mean": ic_mean,
                        "signal_method": scheme["name"],
                        "universe": "full" if universe is None else universe,
                        "pass": pass_name,
                        "backtest_hash": backtest_hash,
                        "ran": False,
                        # The metrics of a skipped cell are in the registry, not here, and
                        # the two halves of `out` have to share a schema for polars to
                        # build one frame from them.
                        "sharpe": None,
                        "total_return": None,
                        "max_drawdown": None,
                        "cagr": None,
                        "volatility": None,
                        "num_trades": None,
                        "failure": None,
                    }
                )
                continue
            planned.add(backtest_hash)
            pending_schemes.append((idx, scheme, universe, spec))

        if not pending_schemes:
            continue

        predictions = normalize_prediction_columns(read_predictions(CASE_STUDY_ID, pred_hash))

        for idx, scheme, universe, spec in pending_schemes:
            record = {
                "prediction_hash": pred_hash,
                "source": source,
                "ic_mean": ic_mean,
                "family": pred_row["family"],
                "config_name": pred_row["config_name"],
                "signal_method": scheme["name"],
                "universe": "full" if universe is None else universe,
                "pass": pass_name,
                "ran": True,
            }
            try:
                result = run_backtest(
                    CASE_STUDY_ID,
                    pred_hash,
                    spec,
                    prices=prices,
                    predictions=predictions,
                    label=LABEL,
                    register=True,
                    force_rebacktest=FORCE_REBACKTEST,
                    initial_cash=bt_config.initial_cash,
                    calendar=bt_config.calendar,
                )
                record.update(
                    backtest_hash=result.backtest_hash,
                    sharpe=result.metrics["sharpe"],
                    total_return=result.metrics["total_return"],
                    max_drawdown=result.metrics["max_drawdown"],
                    cagr=result.metrics.get("cagr", 0.0),
                    volatility=result.metrics.get("volatility", 0.0),
                    num_trades=result.metrics["num_trades"],
                    failure=None,
                )
                tally["completed"] += 1
                if result.backtest_hash:
                    existing_hashes.add(result.backtest_hash)
                    planned.add(result.backtest_hash)
            except Exception as exc:  # noqa: BLE001 - one cell must not end the sweep
                # Say what threw and which cell it was. A counter the log never
                # expands is how 144 of 360 pass-2 cells failed on `fwd_ret_5m`
                # while the run exited 0 and label coverage reported ok: every
                # one of them was a constant-forecast prediction meeting a slot
                # arm that cannot rank it, and nothing on the way out said so.
                # `record` already carries the family, config and arm, so the
                # message identifies the cell rather than only the exception.
                tally["failed"] += 1
                print(
                    f"  FAILED [{idx}/{n_cells}] {record['family']}/{record['config_name']} "
                    f"{record['signal_method']} on {record['universe']} "
                    f"(pred {pred_hash[:12]}): {type(exc).__name__}: {exc}",
                    flush=True,
                )
                record.update(
                    backtest_hash=None,
                    sharpe=None,
                    total_return=None,
                    max_drawdown=None,
                    cagr=None,
                    volatility=None,
                    num_trades=None,
                    failure=f"{type(exc).__name__}: {exc}",
                )
            out.append(record)

            if idx % 20 == 0 or idx == n_cells:
                elapsed = time.time() - t0
                rate = idx / elapsed if elapsed > 0 else 0
                print(
                    f"  [{idx}/{n_cells}] {elapsed:.0f}s ({rate:.1f} bt/s) | "
                    f"completed: {tally['completed']} skipped: {tally['skipped']} "
                    f"failed: {tally['failed']}",
                    flush=True,
                )
    return out


baseline_results = run_arms(pred_index, baseline_arms, "Pass 1 (baseline)")
```

### Which predictions pass 2 studies

The mechanism grid runs on the predictions that scored highest in pass 1, and
on nothing else. The ranking is by pass-1 validation Sharpe.

Not by information coefficient. `load_prediction_index` returns its rows
ordered by `ic_mean` descending, so taking the head of it - which is what
`top_n_predictions.signal` would do - would let a rank correlation decide which
models are backtested at all. A rank correlation and a traded result disagree
whenever turnover differs between two predictions of the same ordering quality,
which is the whole subject of this chapter. Every case study in the book
declares `signal: 0` for that reason.

What this does inherit from pass 1 is pass 1's own arbitrariness: the three
equal-weight concentrations are one way to trade a prediction, and a prediction
that suits the slot mechanism but not equal weight will not reach pass 2. That
is a property of a two-pass design and not of this particular grid, and it is
the price of not running the cross-product.

```python
# The Sharpe comes from the registry and not from `baseline_results`, although pass 1 just
# produced both. A cell whose identity was already registered is skipped rather than re-run -
# that is what makes this sweep resumable - and a skipped cell computes no metrics for this
# invocation to hold. Ranking on what this invocation computed would therefore rank on
# whatever pass 1 had left to do: the whole pass on a first run, a fragment of it after an
# interruption, and nothing at all on a re-run of a finished sweep, which would silently move
# the selection or empty it. The registry holds every pass-1 cell either way.
#
# Built from the two hash fields under an explicit schema rather than from the records
# whole. `pl.DataFrame` infers a schema from the first 100 rows, the metric columns of a
# skipped record are all null, and a run that skips its first hundred cells and then
# computes one would hand a float to a column inferred as Null and fail at construction -
# before the `.select` that discards those columns ever runs. Resuming a long sweep is
# exactly the case that orders the records that way.
_cells = pl.DataFrame(
    [
        {"prediction_hash": r["prediction_hash"], "backtest_hash": r["backtest_hash"]}
        for r in baseline_results
    ],
    schema={"prediction_hash": pl.String, "backtest_hash": pl.String},
    orient="row",
).drop_nulls()
pass2_index = pred_index.head(0)
if mechanism_arms and not _cells.is_empty():
    _conn = sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db"))
    _metrics = pl.read_database(
        "SELECT backtest_hash, sharpe FROM backtest_metrics",
        connection=_conn,
        schema_overrides={"sharpe": pl.Float64},
    )
    _conn.close()
    # `drop_nulls` is the complete test here, and only because the Sharpe crosses SQLite.
    # A ruined path writes `float("nan")` for every ranking metric
    # (`backtest_runner.RUIN_UNRANKABLE_METRICS`), and a NaN read back into polars would
    # survive `drop_nulls` and then sort AHEAD of every finite Sharpe under
    # `descending=True`, so a bankrupt strategy would take a pass-2 slot. SQLite has no
    # NaN: it stores one as NULL, which is what makes the null test sufficient. Read the
    # scores from anywhere but the registry and this line needs `is_finite` as well.
    _scored = _cells.join(_metrics, on="backtest_hash", how="inner").drop_nulls("sharpe")
    if not _scored.is_empty():
        _ranked = (
            _scored.group_by("prediction_hash")
            .agg(best_sharpe=pl.col("sharpe").max())
            .sort("best_sharpe", descending=True)
            .head(mechanism_top_n)
        )
        pass2_index = pred_index.join(
            _ranked.select("prediction_hash"), on="prediction_hash", how="inner"
        )
        print(
            _ranked.join(
                pred_index.select("prediction_hash", "source", "family"),
                on="prediction_hash",
                how="left",
            ).select("source", "family", "best_sharpe")
        )

if mechanism_arms and pass2_index.is_empty():
    # Pass 1 registered no metric for any of its cells, so pass 2 has nothing to select and
    # would run over an empty frame while the sweep reported success.
    msg = (
        f"pass 1 ran {len(_cells)} grid cells for {LABEL} and the registry holds a Sharpe "
        "for none of them, so the mechanism grid has nothing to select. Either every "
        "pass-1 backtest failed, or the metrics were written to a different registry than "
        f"{CASE_DIR / 'run_log' / 'registry.db'}."
    )
    raise RuntimeError(msg)
```

```python
mechanism_results = run_arms(pass2_index, mechanism_arms, "Pass 2 (mechanism)")
results = baseline_results + mechanism_results

elapsed = time.time() - t0
completed, skipped, failed = tally["completed"], tally["skipped"], tally["failed"]
print(
    f"\nSweep complete in {elapsed:.0f}s: {reuse_disclosure(completed, skipped, failed)}",
    flush=True,
)
```

## 3. Full-Universe Signal Evaluation (Act 1)

This section is **read-only** — it queries the registry via `BacktestExplorer`
and reads the rows that carry no cost-feasibility screen. The cost-feasible
carrier is Section 4.

What it reads is the unscreened arm of pass 2: equal weight over the top 5, 10
and 20 names, on the predictions pass 1 ranked highest, with every name in the
panel eligible. The slot mechanism is not in these rows - it is run on the
screened universe only - so the comparison this section draws is between
concentrations and between model families at one entry rule, and not between
entry rules. The entry-rule comparison is Section 4's.

Reading the unscreened arm against the screened one is what makes the cost
screen a measured decision rather than a declared one: the same predictions and
the same three ways of trading them, with and without the constraint.

```python
from case_studies.utils.backtest_explorer import BacktestExplorer

explorer = BacktestExplorer(CASE_STUDY_ID)
print(repr(explorer))

# Scope Act 1 to the full universe. The cost-feasible carrier rows
# (universe_filter == "cost_feasible") are analyzed in Section 4.
all_signal = explorer.best(stage="signal", top_n=99999)
full_signal = all_signal.filter(pl.col("universe_filter") == "full")
print(
    f"Full-universe signal backtests: {full_signal.height:,} "
    f"(of {all_signal.height:,} total signal-stage rows)"
)
```

### What turnover does to the same ordering

The two entry schemes differ in how much trading they provoke from the same
ordering. Equal-weight top-k re-ranks and rebalances at every decision time, so
any name that drifts across the cut-off is sold and another bought. The slot
mechanism holds a fixed number of positions and replaces one only when a
candidate scores above the position it would displace, which leaves a position
alone while it stays competitive.

Comparing the two on the same predictions isolates the cost of turnover from
the quality of the ordering, because only the trading rule differs. The
per-method summary below reports the spread of outcomes rather than one figure
per method: on a grid this size the extremes are the configurations most likely
to be reading noise, and the median says more about the method itself.

```python
method_split = (
    full_signal.group_by("signal_method")
    .agg(
        n=pl.len(),
        pos_frac=(pl.col("sharpe") > 0).mean().round(3),
        sharpe_max=pl.col("sharpe").max().round(2),
        sharpe_median=pl.col("sharpe").median().round(2),
    )
    .sort("n", descending=True)
)
print(method_split)
```

### The strongest validation configurations

The highest-scoring configurations from the sweep, kept for the out-of-sample
test later in the pipeline. Two cautions attach to this table.

It is the top of a large grid, so the configurations in it are the ones that
suited this particular validation window best, and part of what put them there
is chance. The larger the grid, the more of the top is chance.

Nothing here is a selection. The configuration carried forward is chosen once,
under the rule in the strategy analysis notebook, and tested on the holdout
exactly once.

```python
top = full_signal.sort("sharpe", descending=True).head(10)
print(top.select("source", "signal_method", "sharpe", "cagr", "max_drawdown"))
```

### Model family comparison across the universe

Grouping the sweep by the model family that produced each prediction asks
whether a family's ranking quality carries through to a traded result.

The two need not agree, and the reason is turnover. A model can order the
cross-section well and still produce a poor strategy if its scores jump between
decision times, because each jump is a trade and each trade pays the spread. A
discrete score, such as one derived from predicted class membership, changes in
steps and provokes more of those than a continuous one at the same ordering
quality.

```python
families = (
    full_signal.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_max=pl.col("sharpe").max(),
        sharpe_median=pl.col("sharpe").median(),
    )
    .sort("sharpe_max", descending=True)
)
print(families)
```

```python
import matplotlib.pyplot as plt

fig, axes = plt.subplots(1, 2, figsize=(14, 5))

# Sharpe distribution histogram (full universe)
if not full_signal.is_empty():
    axes[0].hist(full_signal["sharpe"].to_numpy(), bins=30, edgecolor="white")
    axes[0].axvline(0, color="red", linestyle="--", linewidth=1)
    axes[0].set_xlabel("Sharpe Ratio")
    axes[0].set_ylabel("Count")
    axes[0].set_title("Full-Universe Sweep: Turnover Drives the Sharpe Distribution")

    # IC vs Sharpe
    axes[1].scatter(
        full_signal["ic_mean"].fill_null(0).to_numpy(),
        full_signal["sharpe"].to_numpy(),
        alpha=0.4,
        s=20,
    )
    axes[1].set_xlabel("Prediction IC (mean)")
    axes[1].set_ylabel("Backtest Sharpe")
    axes[1].set_title("IC → Sharpe: Better Prediction = Better Trading?")

# No `tight_layout()`: `matplotlibrc` sets the layout engine for every figure, so a second
# layout pass writes a warning into the render rather than improving it.
show_with_alt(
    fig,
    "Two panels side by side from the full-universe baseline sweep. Left: a 30-bin "
    "histogram of backtest Sharpe ratios, counts on the vertical axis and Sharpe on the "
    "horizontal, with a dashed red vertical line at zero. Right: a scatter of backtest "
    "Sharpe on the vertical axis against mean prediction IC on the horizontal, one "
    "translucent point per backtested configuration. The two panels answer the same "
    "question from opposite directions: how much of the Sharpe spread the histogram shows "
    "is accounted for by prediction quality.",
)
```

### Sharpe vs Trade Count Diagnostic

At 15-minute cadence, the relationship between Sharpe and trade count
reveals the cost dominance problem. Strategies with high trade counts
(active rebalancing) have deeply negative Sharpe because cost drag scales
linearly with the number of trades. The only configurations with positive
Sharpe are those with very few trades — either degenerate strategies
(near-constant predictions from DL models) or extreme concentration
(top_k=5 with infrequent position changes).

This diagnostic motivates the cadence × cost analysis in Ch18: reducing
the rebalance frequency is the structural solution to cost dominance.

The query below reads the full-universe runs only. Rows produced under the
screened universe are excluded so the relationship between trade count and
outcome is read on the unscreened baseline rather than on a mixture of the two.

```python
db_path = CASE_DIR / "run_log" / "registry.db"
conn = sqlite3.connect(str(db_path))

trade_df = pl.read_database(
    """
    SELECT
        br.backtest_hash,
        bm.sharpe,
        bm.num_trades
    FROM backtest_runs br
    JOIN backtest_metrics bm ON br.backtest_hash = bm.backtest_hash
    WHERE br.stage = 'signal'
      AND json_extract(br.spec_json, '$.strategy.signal.universe_filter') IS NULL
    """,
    connection=conn,
    schema_overrides={"num_trades": pl.Float64, "sharpe": pl.Float64},
)
conn.close()

trade_df = trade_df.drop_nulls("num_trades")
print(f"Signal-stage backtests with trade data: {len(trade_df)}")
print(f"Trade count range: {trade_df['num_trades'].min():.0f} — {trade_df['num_trades'].max():.0f}")
print(f"Median trades: {trade_df['num_trades'].median():.0f}")
```

```python
if not trade_df.is_empty():
    fig, axes = plt.subplots(1, 2, figsize=(14, 5))

    # Scatter: Sharpe vs trades
    axes[0].scatter(
        trade_df["num_trades"].to_numpy(),
        trade_df["sharpe"].to_numpy(),
        alpha=0.15,
        s=10,
    )
    axes[0].axhline(0, color="red", linestyle="--", linewidth=1)
    axes[0].set_xlabel("Number of Trades")
    axes[0].set_ylabel("Sharpe Ratio")
    axes[0].set_title("Sharpe vs Trade Count (baseline stage)")

    # Highlight: positive Sharpe only
    positive = trade_df.filter(pl.col("sharpe") > 0)
    if not positive.is_empty():
        axes[0].scatter(
            positive["num_trades"].to_numpy(),
            positive["sharpe"].to_numpy(),
            color="green",
            alpha=0.6,
            s=20,
            label=f"Sharpe > 0 ({len(positive)})",
        )
        axes[0].legend()

    # Histogram: trade count distribution
    axes[1].hist(trade_df["num_trades"].to_numpy(), bins=50, edgecolor="white")
    axes[1].set_xlabel("Number of Trades")
    axes[1].set_ylabel("Count")
    axes[1].set_title("Distribution of Trade Counts")

    # See the note on the figure above: the layout engine is set repository-wide.
    show_with_alt(
        fig,
        "Two panels side by side. Left: a scatter of backtest Sharpe on the vertical axis "
        "against number of trades on the horizontal, one translucent grey point per "
        "backtested configuration, with a dashed red horizontal line at zero Sharpe and "
        "the configurations above zero redrawn in green and counted in the legend. Right: "
        "a 50-bin histogram of the same trade counts, counts on the vertical axis. Read "
        "the horizontal position of the green points against the bulk of the histogram.",
    )

    # Summary: positive-Sharpe strategies are low-trade
    if not positive.is_empty():
        print(
            f"\nPositive-Sharpe backtests: {len(positive)} / {len(trade_df)} "
            f"({len(positive) / len(trade_df):.1%})"
        )
        print(f"  Median trades (positive Sharpe): {positive['num_trades'].median():.0f}")
        print(f"  Median trades (all): {trade_df['num_trades'].median():.0f}")
        print(
            f"  → Positive Sharpe strategies trade "
            f"{trade_df['num_trades'].median() / max(positive['num_trades'].median(), 1):.1f}x less"
        )
```

A prediction that changes little from bar to bar triggers few position changes,
and so pays little in spread, whatever its ordering quality. That is model
smoothness standing in for a decision about how often to trade, and it is not a
decision anyone made. `17_costs` makes it explicitly, by sweeping the rebalance
cadence itself.

## 4. The Cost-Feasible Carrier (Act 2)

Act 1 established that ranking across all 113 names and rebalancing every bar
is cost-defeated. Act 2 applies the **cost-feasibility screen** from the
feasibility analysis (restrict to the cost-feasible universe — the
cheapest-to-trade names, frozen per split) and replaces every-bar rebalancing
with a turnover-controlled **slot mechanism**: a fixed number of concurrent
positions (slots), each entered when its signal clears a rolling per-symbol
percentile gate and held to a maximum horizon, displaced only by fresher
signals. The book then trades on composition changes rather than re-ranking
every bar. See `case_studies/utils/slot_strategy.py` for the mechanism.

The design shown here fixes the number of concurrent slots, the maximum time a
position may be held, the percentile a signal must clear to enter, and the exit
rule. The grid those four were chosen from runs on the screened universe only,
which is also the universe the chosen design is read on: a mechanism whose
purpose is to cut the cost of trading a cost-expensive tail has nothing to
choose between on a panel that still contains the tail. These backtests are
registered under `universe_filter='cost_feasible'` and are queried directly.

### The forecast the chapter carries

The strategy this chapter ends on does not trade one model. It trades the mean
of every regularized LightGBM forecast in the study - twelve configurations,
three losses across four leaf counts - as a single ordering.

That object has to be built before it can be traded, and it is built here. The
members are resolved from what is registered rather than listed, each enters at
its last checkpoint, and the average is registered as a prediction set of its
own under `family='ensemble'`. Everything downstream then reads it exactly the
way it reads a model: the backtest below, the holdout notebooks, and the
cross-case-study comparison in `20_strategy_analysis`.

Averaging is not a way of getting a better forecast, and normally it does not
produce one. What it produces is a forecast nobody chose. The rule that picks
it is fixed before the validation window is read, so the validation result is
a measurement of that rule rather than the best of the twelve results it could
have reported.

Its last checkpoint and not its best one, for the same reason: a checkpoint
chosen on validation is a choice made on the window the result is then read on.

```python
_ens = load_ensemble_declaration(CASE_STUDY_ID)
ensemble_prediction_hash = None
# How many configurations the mean forecast averages. `_members` cannot answer that far
# down the notebook: the name holds the in-force prediction identities from the membership
# filter until the branch below rebinds it to `resolve_members`' frame, so a run that skips
# the ensemble leaves a frozenset where a later `.height` would be read.
ENSEMBLE_MEMBER_COUNT = None
if _ens is None:
    print("No ensemble declared for this case study.", flush=True)
elif EXECUTION_TIER != "canonical":
    # A preview training spec has to identity-cover its reductions, and this one has no
    # reductions of its own to cover - it is an average of whatever the members are. The
    # ensemble is a canonical-tier object; a preview run reads the canonical registry's
    # rows if they are there and otherwise leaves Section 4's ensemble line empty.
    print(f"Ensemble skipped: it is a canonical-tier object and this run is {EXECUTION_TIER}.")
else:
    # The pool the sweep actually ran on, so the two stages agree by construction rather
    # than by comment. This used to pass `_admissible`, which is the catalog filter -
    # complete and not retired - computed before `prediction_members_in_force` narrowed
    # the pool on publication and cross-sectional coverage. On the 2026-09-13 run that gap
    # was 784 catalog rows against 162 swept, so an unpublished or undercovered member
    # could have entered the ensemble while the sweep refused to backtest it, and
    # publishing the ensemble would have carried it downstream anyway.
    #
    # Degenerate sets are excluded on top, because the sweep has no degeneracy filter of
    # its own and a constant forecast averaged into a mean is not a weaker member - it is
    # a fixed offset. Four L1-linear sets in this registry hold two distinct forecast
    # values across 8,006,995 rows.
    _eligible = SWEPT_POOL - degenerate_prediction_hashes(CASE_DIR)
    _members = resolve_members(
        CASE_DIR,
        label=LABEL,
        family=str(_ens["member_family"]),
        split=SPLIT,
        max_num_leaves=int(_ens["max_num_leaves"]),
        admissible=_eligible,
    )
    _member_spec_json = sqlite3.connect(
        f"file:{CASE_DIR / 'run_log' / 'registry.db'}?mode=ro", uri=True
    )
    try:
        _row = _member_spec_json.execute(
            "SELECT spec_json FROM training_runs WHERE training_hash = ?",
            (_members["training_hash"][0],),
        ).fetchone()
    finally:
        _member_spec_json.close()
    _member_spec = json.loads(_row[0])
    _task = (_member_spec.get("computation", {}).get("task") or {}).get("type")

    if _task != "regression":
        # Averaging class scores is a different construction from averaging forecasts - it
        # needs the class values and the continuous column the classes were cut from, and
        # the registry asks for both. The featured carrier is a continuous label, so the
        # construction is not built rather than guessed at.
        print(
            f"Ensemble skipped for {LABEL}: its {_ens['member_family']} members are "
            f"{_task} models, and a mean forecast is defined here for regression only.",
            flush=True,
        )
    else:
        ENSEMBLE_MEMBER_COUNT = _members.height
        print(
            f"Ensemble members ({_members.height}, {_ens['member_family']} at "
            f"<= {_ens['max_num_leaves']} leaves, last checkpoint):",
            flush=True,
        )
        for _m in _members.iter_rows(named=True):
            print(
                f"  {_m['config_name']:22} leaves {_m['num_leaves']:>3}  "
                f"checkpoint {_m['checkpoint_value']}",
                flush=True,
            )

        _spec = ensemble_training_spec(
            _row[0],
            members=_members,
            method=str(_ens["method"]),
            config_name=str(_ens["config_name"]),
            provenance={
                "entry_point": "case_studies.utils.ensemble",
                "notebook_path": "14_backtest",
                "member_family": str(_ens["member_family"]),
            },
        )
        _ens_training_hash = register_training_run(
            CASE_STUDY_ID, _spec, case_dir=CASE_DIR, entry_point="14_backtest.py"
        )
        _averaged = mean_forecast(CASE_STUDY_ID, _members["prediction_hash"].to_list())
        ensemble_prediction_hash = register_prediction_set(
            CASE_STUDY_ID,
            _ens_training_hash,
            checkpoint_kind="final",
            checkpoint_value=None,
            split=SPLIT,
            predictions=_averaged,
            expected_keys=_averaged.select("fold_id", "symbol", "timestamp"),
            label=LABEL,
            case_dir=CASE_DIR,
        )
        # Registration is not publication, and selection asks the second question. An
        # identity is selectable only where the population its producer publishes still
        # lists it - `selectable_validation_candidates` tests membership, not exclusion -
        # so a prediction set no population ever named is invisible to the ranking however
        # complete it is. Every model notebook 

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。