S&P 500 주식 예측 포트폴리오 배분 방식 비교
코드 Machine Learning for Trading
요약
이 노트북은 동일 가중 신호 단계 기준선에서 검증 예측 설정을 추린 뒤, 같은 예측에 다른 자산 배분 방법을 테스트합니다. 점수 기반 및 적합도 가중치와 역변동성, 리스크 패리티, 평균·분산 최적화, 계층적 리스크 패리티를 비교하고 보유 종목 수를 바꿔봅니다. 검증 샤프 지수로 설정을 추려, 체크포인트가 촘촘한 구성이 후보 목록에서 지나치게 큰 비중을 차지하지 않도록 각 구성에서 전체 구간을 포괄하는 최적 체크포인트를 사용합니다.
이 비교는 각 계보의 동일 가중 결과를 기준으로 삼아 예측 신호와 자산 배분의 기여를 분리하도록 설계했습니다. 배분 방식의 성과는 바스켓 크기에 따라 달라질 수 있고 적합도 가중치는 예측 구간의 품질에 좌우된다고 강조합니다. 모든 결과는 검증 단계에서 얻었고 현재 구성 종목 유니버스를 사용해 생존 편향이 있으며, 표본 외 효용을 입증하지 않습니다. 금요일 종가 후 형성된 신호는 다음 이용 가능한 시가에 실행하며 홀드아웃은 참고하지 않습니다.
핵심 아이디어
- 예측을 고정한 상태에서 자산 배분 방법을 비교해 포지션 크기 효과와 신호 품질 변화를 구분합니다.
- 자산 배분의 기여도를 측정할 때 동일 가중 성과를 기준선으로 사용합니다.
- 집중도가 결과에 영향을 줄 수 있으므로 여러 바스켓 크기에서 배분 방식을 평가합니다.
- 적합도 가중치는 예측 구간 폭에 따라 달라지며 구간 보정 오류를 물려받을 수 있습니다.
- 현재 지수 구성 종목에서 얻은 검증 성과에는 생존 편향이 있으며 미래 효용을 입증하지 않습니다.
태그
전문
# 15_portfolio_management.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # S&P 500 Equity+Options: Portfolio Allocation
#
# This notebook advances the highest-ranked primary-label model configurations
# from the equal-weight baseline and tests the point-in-time allocation methods
# `setup.yaml` declares. It asks whether portfolio sizing improves validation
# performance without changing the prediction model or consulting the holdout.
#
# **Learning objectives**
#
# 1. Apply the top-ten, one-checkpoint-per-configuration selection funnel.
# 2. Compare score weighting, conformal weighting, inverse volatility, risk
# parity, MVO and HRP on the same validation predictions.
# 3. Measure how the number of selected stocks changes allocation performance.
# 4. Separate a validation improvement from evidence of out-of-sample efficacy.
#
# **Book reference:** Chapter 17, Sections 17.2-17.8.
#
# **Prerequisites:** `14_backtest` and its registry-backed equal-weight
# baselines. Signals form after Friday's close and execute at the next available
# open, normally Monday. Every result here is validation data. The
# current-constituent universe retains survivorship bias, so results describe
# this retrospective roster rather than historical S&P 500 membership or a
# prospective index population.
# %%
"""S&P 500 Equity+Options: portfolio allocation sweep."""
import sqlite3
import time
import matplotlib.pyplot as plt
import polars as pl
from case_studies.research import (
OfficialPopulation,
Study,
attest_sweep,
open_study,
open_sweep_attempt,
planned_backtests,
population_supersedes,
predictions_identity,
sweep_plan_name,
upstream_plan_hashes,
)
from case_studies.utils.backtest_loaders import (
get_backtest_config,
load_backtest_prices_for,
warmup_periods_for,
)
from case_studies.utils.backtest_presets import (
build_backtest_spec,
traded_universe_declaration,
)
from case_studies.utils.backtest_runner import run_backtest
from case_studies.utils.conformal import ensure_conformal_calibration_identity
from case_studies.utils.notebook_contracts import prediction_members_in_force
from case_studies.utils.registry import (
backtest_hash_from_parts,
read_predictions,
resolve_best_predictions,
)
from case_studies.utils.sweep_config import (
get_allocators,
get_checkpoints_per_config,
get_top_k_values_for,
get_top_n_predictions,
top_n_cap,
)
from utils.paths import get_case_study_dir
from utils.style import COLORS, FIGSIZE, add_message_title, show_with_alt
# %% tags=["parameters"]
CASE_STUDY_ID = "sp500_equity_option_analytics"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
LABEL = ""
MAX_SYMBOLS = 0
SKIP_EXPENSIVE_ALLOC = False
TOP_N_PREDICTIONS = None
SUPERSEDES_ALLOCATION_POPULATIONS: dict[str, str] | None = None
# %% [markdown]
# ### What is asked for, and what it resolves to
#
# The parameters above are the request; the values this notebook runs on are resolved here under
# different names, so a resolved value can never overwrite the request that produced it. An
# injected parameter wins; otherwise the case study's own declaration does.
# %%
# A run given a workspace reads and registers there rather than in the released case
# directory, and `open_study` is what activates that root. Activation rewrites
# `ML4T_OUTPUT_DIR` for the rest of the process, so it has to happen before the first
# `get_case_study_dir` rather than beside the registry read further down: `CASE_DIR` has to
# already answer for the workspace.
#
# `WORKSPACE` is read at both tiers. It used to be read on the preview branch only, so a
# canonical run that passed one was answered with `Study.regenerate` and registered its
# backtests in the published store while its caller read from the workspace it asked for -
# no exception, no warning, and an exit status that said the run had refused (#1100). A
# canonical run with a workspace is the same full-fidelity sweep writing to that root, which
# is what a rehearsal against a private registry needs. A preview still requires one, because
# a preview with no workspace has nowhere of its own to write.
_workspace_study = None
if EXECUTION_TIER == "preview" and not WORKSPACE:
raise ValueError("preview execution requires WORKSPACE")
if WORKSPACE or EXECUTION_TIER == "preview":
_workspace_study = open_study(
CASE_STUDY_ID,
execution_tier=EXECUTION_TIER,
workspace=WORKSPACE or None,
entry_point="15_portfolio_management",
)
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
bt_config = get_backtest_config(CASE_STUDY_ID)
TOP_N = (
TOP_N_PREDICTIONS
if TOP_N_PREDICTIONS is not None
else get_top_n_predictions(CASE_STUDY_ID, "allocation")
)
TOP_N_CAP = top_n_cap(TOP_N)
CHECKPOINTS_PER_CONFIG = get_checkpoints_per_config(CASE_STUDY_ID)
ALLOCATION_LABEL = LABEL or bt_config.primary_label
print(
f"Case study: {CASE_STUDY_ID}; label: {ALLOCATION_LABEL}; "
f"top configs: {TOP_N}; checkpoints/config: {CHECKPOINTS_PER_CONFIG}"
)
# `Study.at` is the read-only form: one root, no activation. These notebooks only read the
# populations - their backtests reach the registry by their own paths - and every other way in
# ends in `activate()`, which rewrites `ML4T_OUTPUT_DIR` process-wide. `open_study` at the
# canonical tier with no workspace routes to `Study.regenerate`, which refuses unless
# `features`, `labels` and `run_log` are symlinks: true in a maintainer worktree, false in
# every clean clone and CI run. Given a workspace it routes to `Study.open` instead, which is
# why the branch above opens one whenever `WORKSPACE` is set rather than only for a preview.
# `CASE_DIR` is already the directory this notebook resolved, including under a workspace, so
# asking it directly answers for the registry the rest of the notebook reads.
_study = (
_workspace_study
if _workspace_study is not None
else Study.at(CASE_DIR, case_study=CASE_STUDY_ID, entry_point="15_portfolio_management")
)
_members, _population_notes = prediction_members_in_force(_study, CASE_DIR)
for _note in _population_notes:
print(_note)
CURRENT_MEMBERS = _members
# %% [markdown]
# ## 1. Advance the leading baselines
#
# Selection uses validation Sharpe and counts distinct `(family, config_name)`
# pairs. Each configuration enters through its best full-coverage checkpoint,
# so a long checkpoint grid cannot crowd other model families out of the
# allocation round.
# %% [markdown]
# The baseline sweep this ranks is required first. Its plan says which backtests
# the current grid contains, and its attestation says the run that filled them
# finished; without both, the leading configurations here are whatever the
# registry happens to hold, which is how three of ten configurations advanced on
# an unfinished baseline once already.
# %%
UPSTREAM_PLANS = upstream_plan_hashes(
_study,
case_study=CASE_STUDY_ID,
label=ALLOCATION_LABEL,
stage="allocation",
prediction_hashes=CURRENT_MEMBERS,
)
BASELINE_GRID = planned_backtests(
_study,
case_study=CASE_STUDY_ID,
label=ALLOCATION_LABEL,
stage="signal",
prediction_hashes=CURRENT_MEMBERS,
)
# %%
top_preds = resolve_best_predictions(
CASE_STUDY_ID,
ALLOCATION_LABEL,
split="validation",
stage="signal",
top_n=TOP_N,
checkpoints_per_config=CHECKPOINTS_PER_CONFIG,
prediction_hashes=CURRENT_MEMBERS,
backtest_hashes=BASELINE_GRID,
)
# The unit here is a configuration, not a row: `resolve_best_predictions` returns
# `checkpoints_per_config` rows per advancing config, so `len(top_preds)` counts configurations
# only while that is 1. A width of 0 asks for every configuration, as `top_n_predictions.signal`
# does in this setup.yaml, and then there is no count to promise - only that something advanced.
advancing_configs = top_preds.select("family", "config_name").n_unique()
if TOP_N_CAP is None:
if not advancing_configs:
raise RuntimeError("No configuration advanced to the allocation stage")
elif advancing_configs != TOP_N_CAP:
raise RuntimeError(f"Expected {TOP_N} advancing configurations, found {advancing_configs}")
selected_hashes = top_preds["prediction_hash"].to_list()
top_preds.select("source", "prediction_hash", "sharpe")
# %% [markdown]
# The table above is authoritative: each row is one current configuration's
# best full-coverage checkpoint. These are inputs to the allocation comparison,
# not conclusions from it.
# %%
prices = load_backtest_prices_for(
CASE_STUDY_ID,
ALLOCATION_LABEL,
split="validation",
warmup_periods=warmup_periods_for(CASE_STUDY_ID),
max_symbols=MAX_SYMBOLS,
)
n_assets = prices["symbol"].n_unique()
print(f"Price support: {len(prices):,} rows across {n_assets} historical symbols")
# `MAX_SYMBOLS` reduces the price panel, and until the run says so in its own specification
# that reduction did not reach `backtest_hash`: a reduced run and the full run over the same
# predictions hashed alike, so the second was served the first's result and the reduction
# bought nothing (ml4t/agent-workspace#911). Declaring it here, before anything is hashed,
# gives a reduced run an identity of its own; `run_backtest` checks the panel against the
# declaration and narrows the predictions to it, so the sweep ranks the cross-section this
# says it ranks and `n_assets` above describes that same set. A full run declares nothing and
# is byte-identical to before.
# A reduced run is a preview run. Refused on the canonical tier so a narrowed result can
# never land in the registry the book's numbers come from, and so the two can never sit in
# one registry to be ranked against each other: `resolve_best_predictions` takes MAX(sharpe)
# over every backtest of a prediction, and a Sharpe earned over a handful of names would
# advance a configuration ahead of one earned over the whole panel. `us_equities_panel` 16
# through 19 already refuse the parameter this way, and `canonically_refused_parameters`
# reads the refusal out of the source, so the canonical fixture path drops the name rather
# than handing the notebook something its first cell raises on.
if EXECUTION_TIER == "canonical" and MAX_SYMBOLS:
raise ValueError(
"MAX_SYMBOLS narrows the universe this run trades, which makes it a different "
"portfolio from the declared one and gives it its own backtest identity "
"(ml4t/agent-workspace#911). A canonical run trades the declared universe: set "
"MAX_SYMBOLS=0, or run under EXECUTION_TIER='preview' with a WORKSPACE."
)
TRADED_UNIVERSE = traded_universe_declaration(prices) if MAX_SYMBOLS else None
# %% [markdown]
# ## 2. Sweep alternative allocators
#
# Equal weight is the baseline established in `14_backtest`, so it is not an
# allocation-stage method. The sweep combines every declared concentration level with every
# declared alternative allocator, both read from `setup.yaml` rather than counted here, so the
# grid follows the declaration. Existing hashes are reused; only missing combinations run.
# %%
TOP_K_VALUES = get_top_k_values_for(CASE_STUDY_ID, ALLOCATION_LABEL, n_assets)
ALLOC_CONFIGS = get_allocators(CASE_STUDY_ID)
if SKIP_EXPENSIVE_ALLOC:
ALLOC_CONFIGS = [
alloc for alloc in ALLOC_CONFIGS if alloc["method"] not in {"mvo_ledoit_wolf", "hrp"}
]
allocation_methods = [alloc["method"] for alloc in ALLOC_CONFIGS]
if "equal_weight" in allocation_methods:
raise RuntimeError("Equal weight is the baseline and cannot enter the allocation sweep")
print(f"TOP_K grid: {TOP_K_VALUES}; alternative allocators: {allocation_methods}")
# %% [markdown]
# Each planned row carries the same prediction set and execution convention as
# its parent baseline. Only the allocation method and concentration change.
# %%
planned = []
for pred_row in top_preds.iter_rows(named=True):
for top_k in TOP_K_VALUES:
for alloc in ALLOC_CONFIGS:
spec = build_backtest_spec(
CASE_STUDY_ID,
bt_config,
prices=prices,
traded_universe=TRADED_UNIVERSE,
prediction_hash=pred_row["prediction_hash"],
initial_cash=bt_config.initial_cash,
chapter="ch17",
signal={
"method": "equal_weight_top_k",
"top_k": top_k,
"long_short": bt_config.long_short,
},
allocation={**alloc, "top_k": top_k, "long_short": bt_config.long_short},
label=ALLOCATION_LABEL,
)
# run_backtest resolves the conformal calibration identity into the spec
# before registering, so hash the resolved spec or the cache never hits.
spec = ensure_conformal_calibration_identity(spec)
planned.append(
{
"prediction_hash": pred_row["prediction_hash"],
"source": pred_row["source"],
"top_k": top_k,
"allocator": alloc["method"],
"spec": spec,
"backtest_hash": backtest_hash_from_parts(pred_row["prediction_hash"], spec),
}
)
# %% [markdown]
# ### Record the grid before running it
#
# `planned` is every backtest this sweep intends to register, identified before any of them
# executes, and the loop below raises rather than dropping one. Publishing that list as an
# official population is what lets the freeze in `16_risk_management` tell an interrupted sweep
# from a finished one - which no reading of the registered rows can do, because an interruption
# leaves rows that look exactly like a smaller finished grid whether they are counted as rows,
# as model configurations, or as stages present.
#
# It is published *before* the sweep and checked with `require_complete` after, rather than
# written once at the end. Writing it at the end leaves a changed sweep represented by the
# previous generation: that generation is complete, so an interrupted re-run under a widened
# grid reports as finished on the strength of a plan it has already replaced. Publishing first
# means the name always describes the sweep in flight, and an interruption leaves a population
# whose members are not all registered - which `require_complete` reports and the freeze reads.
#
# Writing it activates the study, which rewrites `ML4T_OUTPUT_DIR` process-wide, so the guard
# below refuses if the study does not root at the directory the sweep resolves against. Doing
# that here also means the mismatch costs nothing, where at the end it cost the whole sweep. A
# reader's clean clone has no writable registry and reports that instead of failing; it has no
# field to freeze either.
# %%
# The name carries which prediction sets the sweep planned against, so "has this sweep run
# against the predictions in force" is a lookup rather than an inference over its members. An
# inference answers one direction only: comparing members against the current predictions
# catches one the refit removed and cannot see one it added, because the backtests riding a new
# prediction do not exist until this notebook runs again.
ALLOCATION_POPULATION = sweep_plan_name(
CASE_STUDY_ID, ALLOCATION_LABEL, "allocation", predictions_identity(CURRENT_MEMBERS)
)
# The generation this run retires, per population name. A plan that has grown - a new
# configuration advancing, a widened top-k grid - is a changed population under a live name and
# has to say which one it replaces; the refusal prints the current hash. Absent for a name this
# registry has never held, which is every clean clone and every first run of a label.
#
# All five moved on 2026-09-01. The generation each replaces was planned at 21:40 UTC, before
# 14_backtest published its baseline sweep at 22:34-22:46 and before the tabm_m, tabm_s and sae
# baselines were registered at 22:42. The top-ten those runs ranked was therefore taken over a
# baseline set that was still being produced, and three of the ten it named are not the ten the
# complete set gives. This is the state 16_risk_management declines to freeze over.
# Empty: this run reproduces the grid each name already holds, and `create` returns the
# recorded population unchanged when the member list matches. A declaration that names the
# generation currently in force is not a record of a retirement, it is standing permission to
# change that name's membership without saying so, and the refusal it pre-empts is the one
# thing that makes a changed grid visible. Add an entry when a run is actually refused, with
# the hash the refusal prints.
_DECLARED_SUPERSEDES_ALLOCATION_POPULATIONS: dict[str, str] = {}
# Resolved under a different name, per the convention stated at the parameters cell: an
# injected parameter wins, otherwise the case study's own declaration does. Until 2026-09-18
# the committed map above *was* the parameter name, and because it is assigned here rather
# than in the parameters cell it overwrote whatever papermill injected, before the
# `population_supersedes` call below ever read it. A run that declared the supersedes it was
# asked for was refused as though it had declared nothing. The two sibling notebooks guarded by the same freeze already take this
# as a parameter: cme_futures as SUPERSEDES_ALLOCATION_POPULATION, crypto_perps_funding as
# SUPERSEDES_ALLOCATION.
_supersedes_allocation_populations = (
_DECLARED_SUPERSEDES_ALLOCATION_POPULATIONS
if SUPERSEDES_ALLOCATION_POPULATIONS is None
else SUPERSEDES_ALLOCATION_POPULATIONS
)
_plan = None
try:
_writable = (
_workspace_study
if _workspace_study is not None
else open_study(CASE_STUDY_ID, entry_point="15_portfolio_management")
)
except PermissionError as exc:
print(f"Not recording the allocation plan here: {exc}")
else:
# `storage_root` is the registry this run writes, which is not always `root`: a preview's
# `root` stays the case directory while its writes go to `<workspace>/.preview/<case>`. A
# canonical run, with or without a workspace, writes its own root. Either way the guard
# compares it against the directory the sweep above read, so a run that reads one registry
# and records its plan in another is refused rather than recorded.
if _writable.storage_root(EXECUTION_TIER) != CASE_DIR:
raise RuntimeError(
f"15 ran its sweep against {CASE_DIR} but opened a study writing to {_writable.storage_root(EXECUTION_TIER)}. "
"Recording the plan there would describe a registry this run did not write."
)
if EXECUTION_TIER != "canonical":
# `OfficialPopulation.create` refuses a preview, and rightly: a published population is
# a durable claim about what this case study publishes, and a preview is discarded with
# its workspace. The sweep still executes and still registers. `_plan` stays None, which
# the attestation below already tests for - the same state a non-writable study leaves.
print(
f"{EXECUTION_TIER} tier: the sweep executes and registers, and publishes no "
f"official population under {ALLOCATION_POPULATION}."
)
else:
_plan = OfficialPopulation.create(
_writable,
name=ALLOCATION_POPULATION,
member_kind="backtest",
members=[row["backtest_hash"] for row in planned],
supersedes=population_supersedes(
_writable,
name=ALLOCATION_POPULATION,
declared=_supersedes_allocation_populations.get(ALLOCATION_POPULATION),
),
)
# Before any member executes; see `sweep_attestation_name`.
_attempt = open_sweep_attempt(_writable, _plan, UPSTREAM_PLANS)
print(
f"Allocation plan {ALLOCATION_POPULATION}: {_plan.hash}, {len(planned)} planned, "
f"attempt {_attempt}"
)
# %% [markdown]
# A production run fails if any planned backtest fails. The notebook does not
# silently drop expensive allocators based on elapsed time.
# %%
with sqlite3.connect(CASE_DIR / "run_log" / "registry.db") as db:
existing_hashes = {row[0] for row in db.execute("SELECT backtest_hash FROM backtest_runs")}
n_cached = sum(row["backtest_hash"] in existing_hashes for row in planned)
failures = []
started = time.monotonic()
print(f"Planned {len(planned)} allocation backtests; {n_cached} already complete")
for index, row in enumerate(planned, start=1):
if row["backtest_hash"] in existing_hashes:
continue
try:
result = run_backtest(
CASE_STUDY_ID,
row["prediction_hash"],
row["spec"],
prices=prices,
predictions=read_predictions(CASE_STUDY_ID, row["prediction_hash"]),
label=ALLOCATION_LABEL,
register=True,
initial_cash=bt_config.initial_cash,
calendar=bt_config.calendar,
)
existing_hashes.add(row["backtest_hash"])
print(
f"[{index}/{len(planned)}] {row['source']} k={row['top_k']} "
f"{row['allocator']}: Sharpe={result.metrics['sharpe']:.3f}",
flush=True,
)
except Exception as exc: # noqa: BLE001
failures.append(f"{row['backtest_hash']} {row['source']} {row['allocator']}: {exc}")
if failures:
raise RuntimeError("Allocation sweep failures:\n" + "\n".join(failures))
print(f"Allocation surface complete in {(time.monotonic() - started):.1f}s")
if _plan is not None:
_plan.require_complete()
# Only a run that raised on nothing reaches this. See `sweep_attestation_name`.
_attestation = attest_sweep(_writable, _plan, _attempt, UPSTREAM_PLANS)
print(f"Allocation plan {ALLOCATION_POPULATION} complete: {len(planned)} backtests")
print(f"Sweep attested as {_attestation.name}")
# %% [markdown]
# ## 3. Compare the active allocation surface
#
# The analysis is restricted to the primary label, maximum-coverage prediction
# sets, and the ten configurations advanced above. Accumulated rows from other
# labels or earlier funnels cannot enter these summaries.
# %%
from case_studies.utils.backtest_explorer import BacktestExplorer
explorer = BacktestExplorer(CASE_STUDY_ID)
alloc_comparison = explorer.compare_allocators(
label=ALLOCATION_LABEL,
prediction_hashes=selected_hashes,
).filter(pl.col("allocator").is_in(allocation_methods))
alloc_comparison
# %% [markdown]
# The table pairs each allocator's mean Sharpe across its prediction-by-concentration
# combinations with its single strongest one. Read the pair, not either column alone:
# an allocator can lead on the mean while another owns the peak, and a mean over a
# handful of combinations moves on one of them.
# %%
plot_alloc = alloc_comparison.sort("avg_sharpe")
fig, ax = plt.subplots(figsize=FIGSIZE["single"], constrained_layout=True)
y = range(len(plot_alloc))
_ALLOCATOR_NAMES = {
"inverse_vol": "Inverse volatility",
"mvo_ledoit_wolf": "MVO (Ledoit-Wolf)",
"hrp": "HRP",
}
def allocator_label(method: str) -> str:
"""Chart label for an allocator declared in `setup.yaml`."""
return _ALLOCATOR_NAMES.get(method, method.replace("_", " ").capitalize())
labels = [allocator_label(name) for name in plot_alloc["allocator"].to_list()]
ax.barh(y, plot_alloc["avg_sharpe"], color=COLORS["blue"], alpha=0.82, label="Mean")
ax.scatter(plot_alloc["best_sharpe"], y, color=COLORS["amber"], s=48, zorder=3, label="Best")
ax.set_yticks(list(y), labels)
ax.set_xlabel("Annualized validation Sharpe")
ax.legend(frameon=False, loc="lower right")
add_message_title(
ax,
"Mean and peak Sharpe for each declared allocator",
"Bars: mean across primary-label combinations; points: strongest single one",
)
show_with_alt(
fig,
"Horizontal bars of mean validation Sharpe for each allocator, ordered by that mean, with a "
"separate marker for the allocator's best single configuration.",
)
# %% [markdown]
# ## 4. Inspect the leading allocation
#
# The table keeps the highest-Sharpe rows visible, and the figure after it shows whether that
# result depends on one concentration choice. **Read the two together.** A row that leads at one
# `top_k` and disappears at the next is a concentration artefact rather than an allocator that
# suits this signal, and the sweep runs every level precisely so that is visible rather than
# assumed.
# %%
top_rows = explorer.best(
stage="allocation",
top_n=10,
label=ALLOCATION_LABEL,
prediction_hashes=selected_hashes,
)
winner = explorer.inspect(top_rows["backtest_hash"][0])
winner_strategy = winner.spec["strategy"]
winner_allocator = winner_strategy["allocation"]["method"]
winner_top_k = winner_strategy["signal"]["top_k"]
baseline_sharpe = top_preds.filter(pl.col("prediction_hash") == winner.prediction_hash)["sharpe"][0]
allocation_delta = winner.metrics["sharpe"] - baseline_sharpe
print(
f"Selected allocation: {winner.source}; allocator={winner_allocator}; "
f"top_k={winner_top_k}; validation Sharpe={winner.metrics['sharpe']:.3f}"
)
print(f"Equal-weight baseline={baseline_sharpe:.3f}; allocation delta={allocation_delta:+.3f}")
top_rows.select("source", "prediction_hash", "sharpe", "cagr", "max_drawdown")
# %% [markdown]
# The point estimate is conditional on selecting this row from the full allocation sweep. An
# ordinary interval for this one return path would omit that search, so it is not reported as
# uncertainty about the selected allocation.
# %% [markdown]
# The curve below asks whether the allocation result depends on how many names are
# held. A Sharpe that falls as the basket widens is what dilution of the
# cross-sectional ranking looks like; one that is flat says the allocator, not the
# concentration, is doing the work.
# %%
winner_curve = explorer.concentration_curve(winner.prediction_hash).filter(
pl.col("allocator").is_in(allocation_methods)
)
palette = [
COLORS["blue"],
COLORS["amber"],
COLORS["positive"],
COLORS["copper"],
COLORS["slate"],
COLORS["neutral"],
][: len(allocation_methods)]
fig, ax = plt.subplots(figsize=FIGSIZE["single"], constrained_layout=True)
for color, method in zip(palette, allocation_methods, strict=True):
curve = winner_curve.filter(pl.col("allocator") == method).sort("top_k")
ax.plot(
curve["top_k"],
curve["sharpe"],
marker="o",
linewidth=1.8,
color=color,
label=allocator_label(method),
)
ax.axhline(baseline_sharpe, color=COLORS["neutral"], linestyle="--", linewidth=1.2)
ax.set_xticks(TOP_K_VALUES)
ax.set_xlabel("Selected stocks per rebalance")
ax.set_ylabel("Annualized validation Sharpe")
ax.legend(
frameon=False,
ncol=3,
fontsize=8,
loc="upper center",
bbox_to_anchor=(0.5, -0.16),
)
add_message_title(
ax,
"How each allocator's Sharpe moves as the basket widens",
f"Dashed line: equal-weight baseline Sharpe {baseline_sharpe:.3f}",
)
show_with_alt(
fig,
"One line per allocator of annualized validation Sharpe against the number of stocks selected "
"per rebalance, with a dashed horizontal line at the equal-weight baseline.",
)
# %% [markdown]
# ## Key takeaways
#
# 1. **Equal weight is the baseline, not a competitor.** It is the signal stage's own weighting,
# so it is excluded from the allocator menu and the notebook raises if it reappears there.
# Every Sharpe here is read against it rather than ranked alongside it.
#
# 2. **An allocator is judged across concentration levels, not at one.** The method averages and
# the per-`top_k` figure answer different questions: which allocator suits this signal on
# average, and whether its leading row persists when the number of names held changes.
#
# 3. **The gain from allocation is measured against that lineage's own baseline.** The
# difference between a lineage's equal-weight Sharpe and its highest allocated Sharpe is
# what allocation contributed. That comparison holds the predictions fixed, which is what
# isolates the allocator's effect from the signal's.
#
# 4. **The declared allocators read three different things.** Score weighting reads the
# point prediction, `conformal_weighted` reads the width of its interval, and inverse
# volatility, risk parity, MVO and HRP weight by a moment of returns. Where the conformal
# intervals under-cover out of time - which `13_model_analysis` measures - only the
# interval-width allocator inherits that miscalibration, so read its result beside that
# coverage rather than on its own.
#
# 5. **These are selection-stage results on a current-constituent universe**, so they carry
# survivorship bias and establish no out-of-sample edge. The holdout is untouched here.
#
# **Next:** [`16_risk_management`](16_risk_management.ipynb) applies friction to the leading validation lineage. See
# Chapter 18 for the transaction-cost framework.
```출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.