본문으로 건너뛰기
라이브러리 문서 전체

커버리지 고려 백테스트로 주식 신호 평가

코드 Machine Learning for Trading

요약

이 노트북은 검증 예측을 롱온리 주식 포트폴리오로 평가합니다. 각 레이블의 주기에 따라 결정 일정을 정하고, 상위 K개 종목을 동일 비중으로 선택합니다. 시드가 있는 무작위 신호 실행으로 종가부터 다음 시가까지의 사건 순서를 점검한 뒤 같은 백테스트 과정으로 포트폴리오 집중도 선택을 스윕합니다. 옵션에서 파생된 변수는 예측 입력으로 쓰고, 거래 대상은 주식입니다.

이 분석은 검증 날짜 커버리지가 비슷한 예측 세트끼리만 비교하고 비용을 반영합니다. 블록 부트스트랩 구간, 선택 조정 지표, 백테스트 과최적화 확률을 통해 불확실성을 논의하며, 검증 결과로 선택한 변형이 우연히 더 좋아 보일 수 있다고 경고합니다. 또한 유니버스가 현재 구성 종목으로 이루어져 생존 편향이 미래 해석을 제한한다고 설명합니다. 진행 표는 전략 단계마다 등록된 결과 중 최선의 결과를 비교하지만 각 단계의 인과적 기여도를 분리하지는 않습니다. 홀드아웃 평가는 진행하지 않습니다.

핵심 아이디어

  • 백테스트는 결정 시점의 종가에서 신호를 만들고 다음 이용 가능한 시가에 주문을 제출합니다.
  • 동일 비중 상위 K개 포트폴리오는 비중 산정 규칙을 고정한 채 집중도를 바꿉니다.
  • 샤프 비율을 비교하기 전에 예측 세트의 검증 날짜 커버리지가 비슷한지 확인해야 합니다.
  • 부트스트랩 구간과 선택 조정 통계는 여러 변형 중 선택된 결과를 맥락에 맞게 해석하는 데 도움이 됩니다.
  • 현재 구성 종목 유니버스의 검증 성과는 연구 근거이지 미래 성과 추정치가 아닙니다.

태그

전문
# 14_backtest.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # S&P 500 Equity+Options: Backtest and Signal Evaluation
#
# **Chapter 16 - Strategy Simulation**
#
# This notebook translates validation predictions into long-only equity portfolios on the
# label's own decision grid - `decision.cadence_by_label` puts the two 10-day labels on a
# biweekly schedule and leaves the 5-day ones weekly - and evaluates the equal-weight top-$K$
# baseline. The resolved cadence is printed in the term sheet below. Option-derived features are
# predictive inputs only; the strategy trades equities.
#
# **Learning objectives**
#
# - verify decision-at-close and next-open execution with a random-signal smoke test;
# - sweep equal-weight top-5, top-10, and top-20 portfolios for the primary label;
# - rank only prediction sets with comparable validation-date coverage;
# - interpret bootstrap and selection-adjusted uncertainty without opening the holdout.
#
# Sections 1-2 register only missing primary-label baselines. Section 3 is
# read-only and evaluates all five labels already present in the registry.
#
# **Book Reference:** Chapter 16, Sections 16.4–16.8
#
# **Prerequisites:** completed validation predictions from Chapters 11-15 and
# the case-study protocol in `config/setup.yaml`.

# %%
"""Ch16 backtest and signal evaluation for S&P 500 equity and option features."""

import time

import matplotlib.pyplot as plt
import polars as pl

from case_studies.research import (
    OfficialPopulation,
    Study,
    attest_sweep,
    open_study,
    open_sweep_attempt,
    population_supersedes,
    predictions_identity,
    published_population_names_at,
    reuse_disclosure,
    sweep_plan_name,
)
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.backtest_loaders import get_backtest_config, load_backtest_prices_for
from case_studies.utils.backtest_presets import (
    build_backtest_spec,
    serializable_backtest_spec,
    traded_universe_declaration,
)
from case_studies.utils.backtest_runner import (
    normalize_prediction_columns,
    run_backtest,
    run_plumbing_test,
)
from case_studies.utils.notebook_contracts import prediction_members_in_force
from case_studies.utils.notebook_render import selection_adjusted_leader_table
from case_studies.utils.registry import (
    backtest_dir,
    backtest_hash_from_parts,
    load_existing_backtest_hashes,
    load_prediction_index,
    read_predictions,
    resolve_best_predictions,
)
from case_studies.utils.sweep_config import (
    get_entry_schemes_for,
    get_top_k_values_for,
    get_top_n_predictions,
)
from utils.paths import get_case_study_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, FIGSIZE, add_message_title, show_with_alt, zero_line

# %% tags=["parameters"]
CASE_STUDY_ID = "sp500_equity_option_analytics"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
LABEL = ""
SPLIT = "validation"
# Zero means the smallest top_k from setup.yaml backtest.sweep.top_k_grid.
TOP_K = 0
MAX_SYMBOLS = 0
FORCE_REBACKTEST = False  # Set True to re-backtest even if a complete backtest_hash exists
TOP_N_PREDICTIONS = None
SEED = 42

# %%
set_global_seeds(SEED)

# %% [markdown]
# ## 1. Setup & Plumbing Test
#
# The protocol forms signals at Friday's close and submits orders for the next
# available open, normally Monday. This event ordering prevents Friday-close
# features from receiving a same-bar fill. A seeded random-signal run is a
# plumbing smoke test: it can detect gross engine bias, but it cannot prove that
# every research choice is unbiased.

# %% [markdown]
# ### What is asked for, and what it resolves to
#
# The parameters above are the request; the values the backtest runs on are resolved here and
# carry different names. Keeping them apart means a run can print both, and a resolved value can
# never quietly overwrite the request that produced it. Precedence is the same throughout: an
# injected parameter wins, otherwise the case study's own declaration.

# %%
# A run given a workspace reads and registers there rather than in the released case
# directory, and `open_study` is what activates that root. Activation rewrites
# `ML4T_OUTPUT_DIR` for the rest of the process, so it has to happen before the first
# `get_case_study_dir` rather than beside the registry read further down: `CASE_DIR` has to
# already answer for the workspace.
#
# `WORKSPACE` is read at both tiers. It used to be read on the preview branch only, so a
# canonical run that passed one was answered with `Study.regenerate` and registered its
# backtests in the published store while its caller read from the workspace it asked for -
# no exception, no warning, and an exit status that said the run had refused (#1100). A
# canonical run with a workspace is the same full-fidelity sweep writing to that root, which
# is what a rehearsal against a private registry needs. A preview still requires one, because
# a preview with no workspace has nowhere of its own to write.
_workspace_study = None
if EXECUTION_TIER == "preview" and not WORKSPACE:
    raise ValueError("preview execution requires WORKSPACE")
if WORKSPACE or EXECUTION_TIER == "preview":
    _workspace_study = open_study(
        CASE_STUDY_ID,
        execution_tier=EXECUTION_TIER,
        workspace=WORKSPACE or None,
        entry_point="14_backtest",
    )
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
bt_config = get_backtest_config(CASE_STUDY_ID)
TOP_N = (
    TOP_N_PREDICTIONS
    if TOP_N_PREDICTIONS is not None
    else get_top_n_predictions(CASE_STUDY_ID, "signal")
)
BACKTEST_LABEL = LABEL or bt_config.primary_label

print(f"""Protocol term sheet
  Case study:    {CASE_STUDY_ID}
  Label:         {BACKTEST_LABEL}
  Calendar:      {bt_config.calendar}
  Cadence:       {bt_config.cadence_for(BACKTEST_LABEL)}
  Commission:    {bt_config.commission_bps:.1f} bps
  Slippage:      {bt_config.slippage_bps:.1f} bps
  Total cost:    {bt_config.commission_bps + bt_config.slippage_bps:.1f} bps/leg
  Long/short:    {bt_config.long_short}
""")

# %%
prices = load_backtest_prices_for(
    CASE_STUDY_ID, BACKTEST_LABEL, split="validation", max_symbols=MAX_SYMBOLS
)
n_assets = prices["symbol"].n_unique()

# `MAX_SYMBOLS` reduces the price panel, and until the run says so in its own specification
# that reduction did not reach `backtest_hash`: a reduced run and the full run over the same
# predictions hashed alike, so the second was served the first's result and the reduction
# bought nothing (ml4t/agent-workspace#911). Declaring it here, before anything is hashed,
# gives a reduced run an identity of its own; `run_backtest` checks the panel against the
# declaration and narrows the predictions to it, so the sweep ranks the cross-section this
# says it ranks and `n_assets` above describes that same set. A full run declares nothing and
# is byte-identical to before.
# A reduced run is a preview run. Refused on the canonical tier so a narrowed result can
# never land in the registry the book's numbers come from, and so the two can never sit in
# one registry to be ranked against each other: `resolve_best_predictions` takes MAX(sharpe)
# over every backtest of a prediction, and a Sharpe earned over a handful of names would
# advance a configuration ahead of one earned over the whole panel. `us_equities_panel` 16
# through 19 already refuse the parameter this way, and `canonically_refused_parameters`
# reads the refusal out of the source, so the canonical fixture path drops the name rather
# than handing the notebook something its first cell raises on.
if EXECUTION_TIER == "canonical" and MAX_SYMBOLS:
    raise ValueError(
        "MAX_SYMBOLS narrows the universe this run trades, which makes it a different "
        "portfolio from the declared one and gives it its own backtest identity "
        "(ml4t/agent-workspace#911). A canonical run trades the declared universe: set "
        "MAX_SYMBOLS=0, or run under EXECUTION_TIER='preview' with a WORKSPACE."
    )
TRADED_UNIVERSE = traded_universe_declaration(prices) if MAX_SYMBOLS else None

# Called unconditionally, because the call is the feasibility check: it raises when no
# declared k fits `n_assets`. It used to sit in the `else` of `if TOP_K:`, so a papermill
# TOP_K skipped the check as well as the default it was there to supply. At MAX_SYMBOLS: 3
# with TOP_K: 1 that is what let this notebook run a `6 predictions x 0 schemes` sweep,
# register nothing, and report an analysis of rows the fixture already held
# (ml4t/agent-workspace#1084).
_feasible_top_k = get_top_k_values_for(CASE_STUDY_ID, BACKTEST_LABEL, n_assets)
PLUMBING_TOP_K = TOP_K if TOP_K else _feasible_top_k[0]
print(
    f"Price support: {len(prices):,} rows across {n_assets} historical symbols; "
    f"plumbing-test top-K={PLUMBING_TOP_K}"
)

# %%
strategy_spec = build_backtest_spec(
    CASE_STUDY_ID,
    bt_config,
    prices=prices,
    traded_universe=TRADED_UNIVERSE,
    prediction_hash="plumbing_test",
    initial_cash=bt_config.initial_cash,
    chapter="ch16",
    signal={
        "method": "score_weighted_top_k",
        "top_k": PLUMBING_TOP_K,
        "long_short": bt_config.long_short,
    },
    label=BACKTEST_LABEL,
)

try:
    random_sharpe = run_plumbing_test(
        CASE_STUDY_ID,
        prices,
        strategy_spec,
        top_k=PLUMBING_TOP_K,
        seed=SEED,
        initial_cash=bt_config.initial_cash,
        calendar=bt_config.calendar,
    )

    status = "PASS" if abs(random_sharpe) < 1.5 else "FAIL"
    print(f"Random signal Sharpe: {random_sharpe:.3f}  [{status}]")

    if abs(random_sharpe) >= 1.5:
        print("WARNING: random signal produces non-trivial Sharpe; inspect the pipeline")
except ValueError as e:
    if "zero variance" in str(e).lower():
        print(f"Plumbing test skipped: {e} (too few assets for meaningful test)")
        random_sharpe = 0.0
    else:
        raise

# %% [markdown]
# ## 2. Parametric Sweep
#
# Sweep all primary-label prediction and concentration combinations through the
# same `run_backtest()` function used for a single strategy. The sweep stores
# every non-degenerate prediction result; the analysis below applies the stricter
# maximum-coverage eligibility rule before ranking candidates.
#
# The top-$K$ grid isolates concentration while holding sizing constant at equal
# weight. Weekly top-5, top-10, and top-20 portfolios reveal whether a ranking
# signal holds as the selected tail broadens.

# %% [markdown]
# **A population is immutable and the registry keeps every generation, so a candidate set built
# straight from it counts retired members beside current ones.** Refitting a configuration under a
# corrected estimator publishes a new snapshot that supersedes the old one; both stay readable, and
# nothing in the registry read path filters on that - `case_studies/utils/registry/queries.py`
# contains no occurrence of `supersed`. Without the filter both generations of a refitted
# configuration enter the ranking as separate candidates, with near-identical scores, and the
# published leaders are then fewer distinct strategies than they appear to be.
#
# `prediction_members_in_force` is that filter, and it takes two steps because neither is enough
# alone. It unions what each name publishes now - `OfficialPopulation.one` resolves the one
# generation in a name's chain that nothing supersedes, refusing rather than guessing if the chain
# has forked - and then subtracts the members those names have retired. The subtraction is needed
# because a narrowed or preview run freezes its own snapshot of whatever the catalog held that day
# and stays in force under its own name forever, so the union alone hands a retired generation back
# through the frozen name that still lists it.
#
# A registry that publishes no population at all - a fixture, or a reader's clean clone - is not
# the same as one whose populations are empty, and the filter is skipped rather than applied to
# nothing. It says so where it does that; the sweep below then rests on catalog admissibility.

# %%
# `Study.at` is the read-only form: one root, no activation. These notebooks only read the
# populations - their backtests reach the registry by their own paths - and every other way in
# ends in `activate()`, which rewrites `ML4T_OUTPUT_DIR` process-wide. `open_study` at the
# canonical tier with no workspace routes to `Study.regenerate`, which refuses unless
# `features`, `labels` and `run_log` are symlinks: true in a maintainer worktree, false in
# every clean clone and CI run. Given a workspace it routes to `Study.open` instead, which is
# why the branch above opens one whenever `WORKSPACE` is set rather than only for a preview.
# `CASE_DIR` is already the directory this notebook resolved, including under a workspace, so
# asking it directly answers for the registry the rest of the notebook reads.
_study = (
    _workspace_study
    if _workspace_study is not None
    else Study.at(CASE_DIR, case_study=CASE_STUDY_ID, entry_point="14_backtest")
)
_members, _population_notes = prediction_members_in_force(_study, CASE_DIR)
for _note in _population_notes:
    print(_note)
CURRENT_MEMBERS = _members
# Kept for the refusal below, which has to name what scoped the read. `_population_notes`
# reports gaps in coverage, not which populations are in force, and by the time the
# baseline read comes back empty the only thing that explains it is the names.
CURRENT_POPULATIONS = sorted(published_population_names_at(_study.root))
if CURRENT_MEMBERS is not None:
    print(f"{len(CURRENT_MEMBERS):,} prediction sets in the populations in force")

# %%
pred_index = load_prediction_index(
    CASE_STUDY_ID,
    label=BACKTEST_LABEL,
    split=SPLIT,
    # Without this the index resolves the case directory itself, which is the canonical one even
    # when everything else in this notebook is reading a preview workspace: the query took no
    # part in the activation that `open_study` performed. A preview then reports "No predictions
    # found" while its own registry holds them - 247 of them on 2026-09-06.
    case_dir=CASE_DIR,
)
if CURRENT_MEMBERS is not None:
    pred_index = pred_index.filter(pl.col("prediction_hash").is_in(CURRENT_MEMBERS))

if pred_index.is_empty():
    msg = f"No predictions found for {CASE_STUDY_ID}/{BACKTEST_LABEL}/{SPLIT}"
    raise RuntimeError(msg)

if TOP_N > 0:
    pred_index = pred_index.head(TOP_N)

n_predictions = len(pred_index)
print(f"Predictions to sweep: {n_predictions}")
ic_min, ic_max = pred_index["ic_mean"].min(), pred_index["ic_mean"].max()
if ic_min is not None:
    print(f"  Fold-aggregate IC range: {ic_min:.4f} to {ic_max:.4f}")
else:
    print("  IC range: not yet computed")

# %%
entry_schemes = get_entry_schemes_for(
    CASE_STUDY_ID, BACKTEST_LABEL, n_assets, long_short=bt_config.long_short
)
n_schemes = len(entry_schemes)

print(f"\nEntry schemes ({n_schemes}):")
for es in entry_schemes:
    print(f"  {es['name']}: {es['method']} (top_k={es.get('top_k', '-')})")

total_backtests = n_predictions * n_schemes
print(
    f"\nTotal grid: {n_predictions} predictions × {n_schemes} schemes = {total_backtests} backtests"
)

# %% [markdown]
# For each prediction set, content-addressed hashes identify concentration
# variants already present in the registry. Only missing specifications need
# execution.


# %%
def _artifact_matches_prediction_window(backtest_hash, predictions):
    artifact = backtest_dir(CASE_STUDY_ID, backtest_hash) / "daily_returns.parquet"
    if not artifact.exists():
        return False
    expected = predictions.select(pl.col("timestamp").cast(pl.Date).unique().sort())
    observed = pl.read_parquet(artifact).select(pl.col("timestamp").cast(pl.Date).unique().sort())
    return observed.equals(expected)


def _planned_backtests(pred_row):
    """Every backtest this sweep intends for one prediction set, identified before it runs.

    The identity is a function of the prediction hash and the specification, so the whole grid
    is nameable without reading a single prediction parquet. That is what lets the plan be
    published before execution rather than after it.
    """
    planned = []
    pred_hash = pred_row["prediction_hash"]
    for scheme in entry_schemes:
        signal = {
            "method": scheme["method"],
            "top_k": scheme.get("top_k", 20),
            "long_short": bt_config.long_short,
        }
        signal.update({k: v for k, v in scheme.items() if k not in ("name", "method")})
        spec = build_backtest_spec(
            CASE_STUDY_ID,
            bt_config,
            prices=prices,
            traded_universe=TRADED_UNIVERSE,
            prediction_hash=pred_hash,
            initial_cash=bt_config.initial_cash,
            chapter="ch16",
            signal=signal,
            label=BACKTEST_LABEL,
        )
        backtest_hash = backtest_hash_from_parts(pred_hash, serializable_backtest_spec(spec))
        planned.append(
            {
                "prediction_hash": pred_hash,
                "source": pred_row["source"],
                "scheme": scheme,
                "spec": spec,
                "backtest_hash": backtest_hash,
            }
        )
    return planned


# %% [markdown]
# A single execution helper keeps the sweep cell focused on orchestration. It
# returns the content hash on success and an error message on failure.


# %%
def _execute_one(pred_hash, spec, predictions, force_rebacktest):
    try:
        result = run_backtest(
            CASE_STUDY_ID,
            pred_hash,
            spec,
            prices=prices,
            predictions=predictions,
            label=BACKTEST_LABEL,
            register=True,
            force_rebacktest=FORCE_REBACKTEST or force_rebacktest,
            initial_cash=bt_config.initial_cash,
            calendar=bt_config.calendar,
        )
    except Exception as exc:
        return None, str(exc)
    return result.backtest_hash, None


# %% [markdown]
# Progress counts include cached and newly executed combinations. A complete
# registry therefore finishes quickly without reading every prediction parquet.

# %% [markdown]
# The grid is published as an official population *before* it executes, and checked with
# `require_complete` after. Recording it afterwards would say nothing: a sweep that stops
# part-way registers no plan and leaves the previous, smaller, complete one in force, so the
# freeze downstream would read an interruption as a finished run. Published first, an
# interrupted sweep leaves a current plan whose members are not all registered, which is what
# `require_complete` reports.
#
# The name carries which prediction sets the grid was planned against, so "has this baseline
# been run against the predictions in force" is a lookup rather than an inference over its
# members. Coverage - one rankable baseline per declared label - is a floor and answers a
# different question: it catches a label that never started, not a grid that stopped half way.

# %%
planned_by_prediction = [
    (pred_row["prediction_hash"], _planned_backtests(pred_row))
    for pred_row in pred_index.iter_rows(named=True)
]
planned = [row for _, rows in planned_by_prediction for row in rows]
if len(planned) != total_backtests:
    raise RuntimeError(f"planned {len(planned)} backtests, expected {total_backtests}")

BASELINE_POPULATION = sweep_plan_name(
    CASE_STUDY_ID, BACKTEST_LABEL, "signal", predictions_identity(CURRENT_MEMBERS)
)
# The generation this run retires, per population name. A grid that has grown - a new entry
# scheme, a wider `TOP_N` - is a changed population under a live name and has to say which one
# it replaces; the refusal prints the current hash. Absent for a name this registry has never
# held, which is every clean clone and every first run of a label.
SUPERSEDES_BASELINE_POPULATIONS: dict[str, str] = {}

_plan = None
try:
    _writable = (
        _workspace_study
        if _workspace_study is not None
        else open_study(CASE_STUDY_ID, entry_point="14_backtest")
    )
except PermissionError as exc:
    print(f"Not recording the baseline plan here: {exc}")
else:
    # `storage_root` is the registry this run writes, which is not always `root`: a preview's
    # `root` stays the case directory while its writes go to `<workspace>/.preview/<case>`. A
    # canonical run, with or without a workspace, writes its own root. Either way the guard
    # compares it against the directory the sweep above read, so a run that reads one registry
    # and records its plan in another is refused rather than recorded.
    if _writable.storage_root(EXECUTION_TIER) != CASE_DIR:
        raise RuntimeError(
            f"14 ran its sweep against {CASE_DIR} but opened a study writing to {_writable.storage_root(EXECUTION_TIER)}. "
            "Recording the plan there would describe a registry this run did not write."
        )
    if EXECUTION_TIER != "canonical":
        # `OfficialPopulation.create` calls `_refuse_preview_activation`, and rightly: a
        # population is a durable claim about what this case study publishes, and a preview is
        # discarded with its workspace. The sweep below still executes and still registers its
        # backtests - what is skipped is the published name over them. `_plan` stays None, which
        # the attestation below already tests for, because the same state arises when the study
        # is not writable at all.
        print(
            f"{EXECUTION_TIER} tier: the sweep executes and registers, and publishes no "
            f"official population under {BASELINE_POPULATION}."
        )
    else:
        _plan = OfficialPopulation.create(
            _writable,
            name=BASELINE_POPULATION,
            member_kind="backtest",
            members=[row["backtest_hash"] for row in planned],
            supersedes=population_supersedes(
                _writable,
                name=BASELINE_POPULATION,
                declared=SUPERSEDES_BASELINE_POPULATIONS.get(BASELINE_POPULATION),
            ),
        )
        # Opened here, before a single member executes, so a run that dies part-way leaves an
        # attempt with no attestation behind it and the freeze declines. See
        # `sweep_attestation_name` for why the record is per attempt rather than per plan.
        _attempt = open_sweep_attempt(_writable, _plan)
        print(
            f"Baseline plan {BASELINE_POPULATION}: {_plan.hash}, {len(planned)} planned, "
            f"attempt {_attempt}"
        )

# %%
t0 = time.time()
completed = failed = skipped = 0
failures: list[str] = []
existing_hashes = load_existing_backtest_hashes(CASE_STUDY_ID, stage="signal")
print(f"Existing equal-weight baseline hashes in registry: {len(existing_hashes):,}")

for pred_hash, group in planned_by_prediction:
    predictions = normalize_prediction_columns(read_predictions(CASE_STUDY_ID, pred_hash))
    for row in group:
        backtest_hash = row["backtest_hash"]
        is_registered = backtest_hash in existing_hashes
        if is_registered and _artifact_matches_prediction_window(backtest_hash, predictions):
            skipped += 1
            continue
        _, error = _execute_one(pred_hash, row["spec"], predictions, is_registered)
        if error:
            failed += 1
            failures.append(f"{backtest_hash} ({row['source']} / {row['scheme']['name']}): {error}")
            print(f"  FAILED {row['source']} / {row['scheme']['name']}: {error}")
        else:
            completed += 1
            existing_hashes.add(backtest_hash)

        processed = completed + failed + skipped
        if processed % 20 == 0 or processed == total_backtests:
            elapsed = time.time() - t0
            rate = processed / elapsed if elapsed > 0 else 0
            print(f"  [{processed}/{total_backtests}] {rate:.1f} bt/s | failed: {failed}")

elapsed = time.time() - t0
print(f"\nSweep complete in {elapsed:.0f}s: {reuse_disclosure(completed, skipped, failed)}")

# A failure here stops the notebook rather than being counted and printed. `require_complete`
# below cannot be relied on to catch one: a member that failed because its registered artifact
# disagrees with the current prediction window is still a *registered, internally complete*
# backtest, so the plan passes while the row it admits was produced over a different window.
# That is the state that reaches the freeze and gets sealed into an immutable set. Every
# failure is printed above with its cause; this says the run does not continue past them.
if failures:
    raise RuntimeError(
        f"{len(failures)} of the {total_backtests} planned baseline backtests failed and the "
        "plan is not accepted. The causes are printed above. A failure whose backtest is "
        "already registered does not show up as an incomplete plan - the registered row is "
        "complete on its own terms - so it has to stop the run here or it reaches the freeze "
        "as a current result. `immutable backtest artifact conflict` means an artifact "
        "directory exists for an identity the registry does not hold: check for a "
        "`backtest_runs` row, and where there is none the directory is debris from an "
        "interrupted sweep and belongs in a quarantine directory, not in the way of the "
        f"re-run. First five: {failures[:5]}"
    )

if _plan is not None:
    _plan.require_complete()
    # Reaching this line is the statement being recorded: the sweep raised above on any
    # failure, so an attestation exists only for a run that executed its whole plan. The
    # plan alone cannot say it - a member that failed against a stale registered artifact
    # leaves the plan complete - and the freeze runs later, in another notebook, where the
    # exception above is not visible. See `sweep_attestation_name`.
    _attestation = attest_sweep(_writable, _plan, _attempt)
    print(f"\nBaseline plan {BASELINE_POPULATION} complete: {len(planned)} backtests")
    print(f"Sweep attested as {_attestation.name}")

# %% [markdown]
# ## 3. Signal Evaluation
#
# This section is **read-only**. It evaluates the full registry across all five
# labels, while the default sweep above covers only `fwd_ret_5d`. Eligibility
# requires each prediction set to match the maximum validation-date coverage
# within its family and label. This prevents a checkpoint scored on fewer dates
# from winning on a different evaluation sample.
#
# Every rank, interval and selection-adjustment statistic in this notebook is
# computed on validation data, and the holdout stays available for one
# evaluation after the strategy is fixed.

# %%
explorer = BacktestExplorer(CASE_STUDY_ID)
print(repr(explorer))

# %% [markdown]
# ### Eligible Leaders
#
# Requiring maximum decision-date coverage within each family and label removes
# partial prediction histories before ranking. Which configuration comes top is read from the
# table below rather than named here: it is a property of the run and it moves whenever the
# populations are refitted. What the section is for is the shape - whether the families separate
# at all, and by how much - not the identity of the row at the top.

# %%
all_baselines = explorer.best(
    stage="signal",
    top_n=9999,
    prediction_hashes=sorted(CURRENT_MEMBERS) if CURRENT_MEMBERS is not None else None,
)

# An empty read here is a statement about the registry, not a result to summarise. The
# division below is the first thing to touch it, so without this the failure arrives as
# `ZeroDivisionError: division by zero` and reads like a defect in the analysis - which is
# how ml4t/agent-workspace#1086 presented. The population is in force and its members are
# registered; what none of them has is a backtest, and only the counts say so.
if CURRENT_MEMBERS is not None and all_baselines.is_empty():
    _signal_rows = explorer.specs("signal").height
    raise RuntimeError(
        f"The populations in force publish {len(CURRENT_MEMBERS):,} prediction sets and not "
        f"one of them carries a signal backtest, so there is no baseline to rank. "
        f"Populations: {', '.join(CURRENT_POPULATIONS) or '(none named)'}. "
        f"The registry holds {_signal_rows:,} signal backtest(s), all against prediction "
        f"sets outside the populations, so this is a population and a set of backtests "
        f"describing different generations rather than an empty registry. A fixture "
        f"generated through a stage earlier than the backtest stage produces exactly this: "
        f"the model notebooks declare the population and nothing goes on to backtest its "
        f"members (ml4t/agent-workspace#1086)."
    )

search_context = {
    "total": len(all_baselines),
    "median_sharpe": all_baselines["sharpe"].median(),
    "pct_positive": 100 * all_baselines.filter(pl.col("sharpe") > 0).height / len(all_baselines),
}
top = all_baselines.head(10)

top_k = [
    explorer.inspect(backtest_hash).spec["strategy"]["signal"]["top_k"]
    for backtest_hash in top["backtest_hash"]
]
top = top.with_columns(pl.Series("top_k", top_k, dtype=pl.Int64))

leader_predictions = normalize_prediction_columns(
    read_predictions(CASE_STUDY_ID, top["prediction_hash"][0])
)
print(
    f"Eligible baselines: {search_context['total']:,}; "
    f"median Sharpe: {search_context['median_sharpe']:.3f}; "
    f"positive: {search_context['pct_positive']:.1f}%"
)
print(
    f"Leader coverage: {leader_predictions['symbol'].n_unique()} symbols across "
    f"{leader_predictions['timestamp'].n_unique()} validation dates"
)

top.select(
    "source",
    "label",
    "top_k",
    pl.col("sharpe").round(3),
    pl.col("ic_mean_daily").round(4).alias("daily_ic"),
    pl.col("cagr").round(3),
    pl.col("max_drawdown").round(3),
)

# %% [markdown]
# Downstream selection is label-specific and admits one checkpoint per distinct model
# configuration. The primary label carries its own lineage: whichever configuration leads the
# ranking above on another label does not displace it, because a strategy is built against one
# target rather than against the strongest number in the sweep.

# %%
primary_advancing = resolve_best_predictions(
    CASE_STUDY_ID,
    BACKTEST_LABEL,
    split=SPLIT,
    top_n=10,
    stage="signal",
    prediction_hashes=CURRENT_MEMBERS,
)
primary_advancing.select(
    "family",
    "config_name",
    "checkpoint_value",
    pl.col("sharpe").round(3),
)

# %% [markdown]
# ### Family Dispersion and IC Translation
#
# Family medians occupy a narrow Sharpe band while the maxima spread much wider, so which family
# holds the highest single configuration is a weaker statement than the bars make it look - and
# a family publishing ten checkpoints has ten chances at that maximum where one publishing a
# single checkpoint has one. Across all eligible baselines, daily-pooled IC has only a moderate
# rank association with portfolio Sharpe. Concentration, turnover and return-path differences
# therefore remain material after model ranking quality is known. The figure reports the current
# coefficient from the registry.

# %%
families = (
    all_baselines.group_by("family")
    .agg(
        n=pl.len(),
        sharpe_median=pl.col("sharpe").median(),
        sharpe_max=pl.col("sharpe").max(),
        sharpe_q75=pl.col("sharpe").quantile(0.75),
        pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
    )
    .sort("sharpe_median", descending=True)
)
ic_sharpe_rho = all_baselines.select(
    pl.corr("ic_mean_daily", "sharpe", method="spearman").alias("rho")
).item()

family_plot = families.sort("sharpe_median")
y_pos = list(range(family_plot.height))
family_labels = [name.replace("_", " ").title() for name in family_plot["family"]]

# %%
fig, ax = plt.subplots(figsize=FIGSIZE["single"], constrained_layout=True)
ax.barh(y_pos, family_plot["sharpe_median"], color=COLORS["blue"], label="Median")
ax.scatter(
    family_plot["sharpe_max"],
    y_pos,
    color=COLORS["amber"],
    marker="D",
    s=24,
    label="Maximum",
    zorder=3,
)
ax.set_yticks(y_pos, family_labels)
ax.set_xlabel("Validation Sharpe")
ax.set_ylabel("Model family")
zero_line(ax, axis="x")
ax.legend(frameon=False, loc="lower right")
add_message_title(
    ax,
    "Family medians sit close together; the maxima do not",
    "Median bars and maximum diamonds; eligible equal-weight baselines",
)
show_with_alt(
    fig,
    "Horizontal bars of each model family's median validation Sharpe, with a marker for that "
    "family's best single run.",
)

# %% [markdown]
# The same evaluation surface shows why prediction and portfolio diagnostics
# cannot substitute for one another: similar IC values map to a wide Sharpe
# range after tail selection and trading mechanics.

# %%
fig, ax = plt.subplots(figsize=FIGSIZE["single"], constrained_layout=True)
ax.scatter(
    all_baselines["ic_mean_daily"],
    all_baselines["sharpe"],
    color=COLORS["blue"],
    alpha=0.35,
    s=12,
)
ax.scatter(
    top["ic_mean_daily"][0],
    top["sharpe"][0],
    color=COLORS["amber"],
    edgecolor=COLORS["blue"],
    s=45,
    zorder=3,
    label=f"leader: {top['family'][0]}/{top['config_name'][0]}",
)
zero_line(ax, axis="x")
zero_line(ax, axis="y")
ax.set_xlabel("Daily-pooled Spearman IC")
ax.set_ylabel("Validation Sharpe")
ax.legend(frameon=False, loc="lower right")
add_message_title(
    ax,
    "IC explains only part of Sharpe dispersion",
    f"Spearman rho = {ic_sharpe_rho:.3f}; {len(all_baselines):,} eligible baselines",
)
show_with_alt(
    fig,
    "Scatter of validation Sharpe against daily-pooled Spearman IC, one faint point per eligible "
    "baseline, with the leading configuration drawn larger and labelled.",
)

# %% [markdown]
# ### Selection-Adjusted Uncertainty
#
# A single-strategy bootstrap asks whether each return path clears zero. The
# effective-rank deflated Sharpe ratio (DSR) additionally accounts for the
# correlated variants tested within each family and label:
#
# $$DSR = \Phi\left[\frac{(\hat{SR} - SR^*) \sqrt{T-1}}{\sqrt{1 - \hat{\gamma}_3 \hat{SR} + \frac{\hat{\gamma}_4 - 1}{4} \hat{SR}^2}}\right].$$
#
# Read three columns together, because each answers a different question. The **block-bootstrap
# Sharpe interval** asks whether one return path clears zero on its own. The **effective-rank
# DSR** asks whether it still clears once the correlated variants tried within its own family
# and label are counted, which is the number that accounts for the search. **PBO** asks how often
# the in-sample leader underperforms out of sample; on two validation folds it has very low
# resolution. These diagnostics support a validation candidate, not an out-of-sample claim.
#
# **Two of those three come from `cohort_metrics`, and this stage does not populate it.**
# `selection_adjusted_leader_table` LEFT JOINs that table, so `dsr_pvalue`, `k_variants` and
# `pbo` arrive null until a stage that computes cohort metrics has run - which for this case
# study is `20_strategy_analysis` via `compute_cohort_metrics`. The guard below says so rather
# than letting three empty columns print as though the adjustment had been made.

# %%
family_leaders = selection_adjusted_leader_table(
    CASE_STUDY_ID,
    stage="signal",
    prediction_hashes=CURRENT_MEMBERS,
)
_adjustment_cols = ["dsr_pvalue", "k_variants", "pbo"]
_absent = [c for c in _adjustment_cols if family_leaders[c].null_count() == family_leaders.height]
if _absent:
    print(
        f"selection adjustment not available at this stage: {', '.join(_absent)} are entirely "
        f"null because cohort_metrics holds no rows for {CASE_STUDY_ID} yet. The bootstrap "
        "interval below stands on its own; the search-adjusted reading arrives with "
        "20_strategy_analysis."
    )
family_leaders.select(
    "family",
    "config_name",
    "label",
    pl.col("sharpe").round(3),
    pl.col("sharpe_ci95_lo").round(3).alias("ci_lo"),
    pl.col("sharpe_ci95_hi").round(3).alias("ci_hi"),
    *[pl.col(c).round(4) if c == "dsr_pvalue" else pl.col(c) for c in _adjustment_cols],
)

# %% [markdown]
# The interval plot shows which family leaders clear zero on their own return path and which do
# not. It is a bootstrap reading only: the search adjustment that would narrow it is the one the
# guard above reports as unavailable at this stage.

# %%
leader_plot = family_leaders.sort("sharpe")
leader_labels = [family.replace("_", " ") for family in leader_plot["family"]]

fig, ax = plt.subplots(figsize=FIGSIZE["single_tall"], constrained_layout=True)
for i, row in enumerate(leader_plot.iter_rows(named=True)):
    focal = row["family"] == "latent_factors"
    color = COLORS["amber"] if focal else COLORS["blue"]
    ax.errorbar(
        row["sharpe"],
        i,
        xerr=[
            [row["sharpe"] - row["sharpe_ci95_lo"]],
            [row["sharpe_ci95_hi"] - row["sharpe"]],
        ],
        fmt="o",
        color=color,
        capsize=3,
        markersize=5,
    )
ax.set_yticks(range(leader_plot.height), leader_labels)
ax.set_xlabel("Annualized validation Sharpe (95% block-bootstrap CI)")
ax.set_ylabel("Family leader")
zero_line(ax, axis="x")
add_message_title(
    ax,
    "Which family leaders clear zero on their own return path",
    "Best eligible equal-weight baseline per model family across five labels",
)
show_with_alt(
    fig,
    "One row per family leader, each an annualized validation Sharpe point with its 95% "
    "block-bootstrap interval drawn as a horizontal whisker.",
)

# %% [markdown]
# ### Downstream Preview
#
# Each row is the strongest registered variant at that pipeline layer for the same prediction
# set. **That makes it a preview of what each layer can reach, not an attribution of the gain
# to the layer**: every row has been selected, so a rise from one layer to the next mixes what
# the layer contributes with what selecting over its variants contributes. A fixed
# specification carried through every layer is what would separate the two, and it is not what
# this table does.

# %%
stage_labels = {
    "signal": "equal-weight baseline",
    "allocation": "allocation",
    "cost_sensitivity": "cost sensitivity",
    "risk_overlay": "risk overlay",
}
best_pred = top["prediction_hash"][0]
progression = explorer.progression(best_pred)
progression.with_columns(
    pl.col("stage").replace_strict(stage_labels, default=pl.col("stage")).alias("pipeline_layer")
).select(
    "pipeline_layer",
    pl.col("sharpe").round(3),
    pl.col("cagr").round(3),
    pl.col("max_drawdown").round(3),
)

# %% [markdown]
# ## Key Takeaways
#
# 1. **Coverage-aware ranking comes before comparison.** A prediction set measured on fewer
#    validation dates than its rivals is not comparable to them, so partial histories are
#    removed before any Sharpe is ranked rather than after.
#
# 2. **Selection is on validation backtest Sharpe, and it happens here.** IC ranked nothing
#    upstream; every checkpoint of every model reached this notebook, and what advances is
#    decided on money after costs rather than on rank correlation.
#
# 3. **A Sharpe that clears zero on its own has not yet cleared the search.** The
#    effective-rank DSR is the column that counts the correlated variants tried within a family
#    and label; two-fold PBO is reported beside it and is too coarse on this sample to support
#    a stability claim either way.
#
# 4. **Each label keeps its own lineage.** The primary label's leader is chosen from the
#    primary label's candidates, and a stronger number on another label does not displace it.
#    Mixing them would select the label as well as the model, on the same validation data.
#
# 5. **IC and Sharpe are related and not interchangeable.** Portfolio construction reorders
#    candidates that IC ranked one way, which is the entire reason selection waits until here.
#    The measured coefficient is in the figure rather than frozen into this sentence.
#
# 6. **Everything here is validation data, on a current-constituent roster.** These are
#    validation results on a universe that embeds survivorship bias, so they are research
#    evidence rather than a prospective performance estimate.
#
# **Next:** [`15_portfolio_management`](15_portfolio_management.ipynb) compares the declared
# allocators within each label-specific lineage.

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.