使用自助法不确定性和基准评估ETF策略
代码 《交易机器学习》
总结
此笔记本汇总ETF案例研究中已登记的回测,对策略进行考虑不确定性的评估。它读取带有分块自助法置信区间的表现指标,并用配对自助法比较连续的流程阶段、策略变体和留出结果。流程阶段涵盖信号选择、资产配置、交易成本和风险叠加规则。它还将选定策略与等权跨资产ETF基准在验证期和留出期的表现进行比较。
分析还提供相对于基准的阿尔法、贝塔和信息比率诊断,并使用含动量因子的五因子股票模型和安慰剂投资组合进行因子归因。笔记本说明,跨资产投资范围可能限制标准股票因子的解释力。它报告不同回测队列的选择偏差指标,并在有相关结果时记录留出期衰减。结论取决于已登记的运行记录和选定的实时预测集;这是对现有回测的评估,并非新的训练或回测。
核心观点
- 使用分块自助法区间描述回测指标的不确定性。
- 通过对收益序列进行配对重抽样,比较策略和流程阶段。
- 相对于等权ETF投资范围基准,评估验证期和留出期表现。
- 对于横跨股票、债券、商品和货币的投资组合,因子归因的解释力较弱。
- 结合预测数据沿革、队列选择以及留出阶段是否已运行,解读报告结果。
标签
全文
# 20_strategy_analysis.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # ETFs - Strategy Analysis
#
# This notebook reads the ETF case study's whole registered pipeline - every backtest from the
# signal sweep through the risk overlay - and turns it into one strategy assessment. Every metric
# carries a block-bootstrap confidence interval, every comparison between two strategies goes
# through a paired bootstrap rather than a difference of point estimates, and the holdout closure
# is read the same way. Comparison across case studies is Chapter 20's subject.
#
# **Learning objectives**
#
# - Read uncertainty-aware backtest metrics - Sharpe with its interval, PSR, DSR - from the
# registry rather than transcribing point estimates.
# - Trace one strategy configuration through the four pipeline stages
# (signal to allocation → cost → risk) with paired-bootstrap stage
# transitions.
# - Use the equal-weight ETF universe benchmark, sliced into validation and holdout periods, for
# both the equity-curve overlay and the holdout strategy-versus-benchmark paired test.
# - Layer 1 + Layer 2 benchmark-aware diagnostics: PortfolioAnalysis
# alpha/beta/IR against the equal-weight universe, plus an FF5+MOM
# factor attribution with placebo-portfolio control. The cross-asset
# ETF universe (equities / bonds / commodities / currencies) means
# factor R² is structurally limited; the placebo benchmark separates
# universe-driven from selection-driven factor exposure.
#
# **Book reference**: Chapter 20, §20.1 (the §9 handoff feeds Ch20's
# cross-case-study aggregation).
#
# **Prerequisites**: case-study pipeline through `19_holdout_backtest`;
# the locked registry (`case_studies/etfs/run_log/registry.db`).
#
# **Scope**: no training and no re-backtesting. It does write two derived tables,
# `cohort_metrics` and `backtest_paired_metrics`, and that is a deliberate change from the
# read-only scope this notebook used to declare.
#
# Both tables are derived from backtests that already exist - selection-bias statistics over the
# cohorts, and paired-bootstrap comparisons between registered return series. Nothing is refitted
# and no backtest is added. They were previously produced by a chapter-20 notebook looping over
# every case study, which made a case study's own strategy analysis unreadable until a later
# chapter had been run, and left both tables empty for any reader working the case study in order.
# A stage that cannot be read without running a chapter that comes after it is not a stage. So the
# notebook that has every stage in front of it produces them, and re-running it recomputes only
# what is missing.
# %%
"""ETFs - Strategy Analysis."""
import json
import sqlite3
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl
# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch # noqa: F401
import yaml
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
BacktestReportMetadata,
generate_tearsheet_from_run_artifacts,
)
from case_studies.research import open_study, split_unpublished_members
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.factor_attribution import (
compute_bootstrap_ci,
compute_rolling_exposures,
format_attribution_summary,
load_factor_data,
plot_attribution_waterfall,
plot_rolling_exposures,
run_factor_regression,
run_placebo_benchmark,
)
from case_studies.utils.paired_metrics import populate_paired_metrics
from case_studies.utils.registry import (
load_backtest_fold_metrics,
load_backtest_metrics,
load_paired_metrics,
load_prediction_index,
)
from case_studies.utils.strategy_analysis import (
ci_status,
compute_operating_profile,
fmt_gate,
gate1_validation_sharpe_geq_zero,
gate2_holdout_diff_not_excludes_zero_negatively,
gate_passes,
plot_concentration_curve,
plot_equity_drawdown,
plot_sharpe_waterfall,
resolve_holdout_self_backtest,
resolve_solvent_carrier,
write_strategy_assessment,
)
from case_studies.utils.uncertainty import STAGE_SEQUENCE, descends_from
from utils.paths import get_case_study_dir, get_output_dir
from utils.style import show_with_alt
# %% tags=["parameters"]
# MAX_SYMBOLS is gone. Nothing below read it, and a declared cap the run does not apply is
# worse than no cap: a reader who sets it gets a full run and no warning. The two names that
# remain are read - `_declares_tier_and_workspace` (tests/pm_helpers.py) looks for exactly
# this pair - which is why they stay bound although nothing below references them. Without
# them the canonical branch regenerates in place, which needs symlinks a CI checkout has not
# got.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
# %% [markdown]
# The study is opened before any path or registry read. Under the preview tier, opening it
# activates a workspace and rewrites `ML4T_OUTPUT_DIR` process-wide; a `CASE_DIR` or a
# `BacktestExplorer` built first would address the released registry while everything after it
# reads the preview one.
# %%
CASE_STUDY = "etfs"
PRIMARY_LABEL = "fwd_ret_21d"
PERIODS_PER_YEAR = 252 # NYSE calendar, daily bars
study = open_study(CASE_STUDY, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
CASE_DIR = get_case_study_dir(CASE_STUDY)
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
with open(CASE_DIR / "config" / "setup.yaml") as f:
setup = yaml.safe_load(f)
explorer = BacktestExplorer(CASE_STUDY)
print(explorer)
# %% [markdown]
# **Which prediction sets their publishers still stand behind.** A refit publishes a second
# generation under the same population name, and the generation it replaced stays in the registry:
# complete, current under a schema version that has not moved, and carrying every backtest the
# previous sweep registered for it. The configuration this notebook describes is chosen by
# backtest Sharpe over that pool, so without the lineage the analysis can be about a strategy the
# case study no longer publishes.
# %%
LIVE_PREDICTIONS = (
split_unpublished_members(
study,
load_prediction_index(CASE_STUDY, split="validation"),
)
.live["prediction_hash"]
.to_list()
)
if not LIVE_PREDICTIONS:
raise RuntimeError(
f"no live prediction sets for {CASE_STUDY}/validation; run the model stage and "
"14_backtest first"
)
print(f"Live prediction sets: {len(LIVE_PREDICTIONS):,}")
# Both derived tables fill in two waves - stage transitions once allocation, cost and risk have
# run, holdout kinds only once the holdout has been evaluated - so the predicate is whether every
# kind this notebook reads is present, not whether anything is.
PAIRED_KINDS = (
"signal_leader",
"allocation_leader",
"cost_sensitivity_leader",
"val_rank1_self",
"equal_weight_holdout_side_artifact",
)
COHORT_TYPES = ("family", "stagelabel", "label")
def _derived_table_state() -> tuple[set[str], set[str]]:
"""Which paired kinds and cohort granularities the registry already holds."""
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as db:
tables = {r[0] for r in db.execute("SELECT name FROM sqlite_master WHERE type='table'")}
kinds = (
{
r[0]
for r in db.execute("SELECT DISTINCT benchmark_kind FROM backtest_paired_metrics")
}
if "backtest_paired_metrics" in tables
else set()
)
cohorts = (
{r[0] for r in db.execute("SELECT DISTINCT cohort_type FROM cohort_metrics")}
if "cohort_metrics" in tables
else set()
)
return kinds, cohorts
def _stale_derived_rows(live: list[str]) -> tuple[int, int]:
"""Derived rows built over anything this notebook does not report as live.
Presence of every enum value says the tables were built; it does not say they were built
over the population this notebook reports. A refit retires the generation a previous run
led with, and its cohort and paired rows survive intact under those same enum values.
A leader and a challenger are the visible halves. A cohort is also stale when a retired
prediction is one of the variants the correction was computed over - the leader can be
live while the trial count and the deflated Sharpe are not - and a pair is also stale
when its benchmark side is retired, which is the side the difference is measured
against. `member_digest` records the cohort's members, so a cohort whose digest is
absent cannot be shown to be live either and is recomputed rather than trusted.
Cohort membership is not queryable here - `backtest_runs` carries neither label nor
family - so the membership test is that the cohort's stage still holds a retired
backtest at all. `compute_and_register` refreshes the whole table rather than one row,
so triggering it too readily costs a recompute and can never report a stale number.
"""
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as db:
tables = {r[0] for r in db.execute("SELECT name FROM sqlite_master WHERE type='table'")}
payload = json.dumps(live)
stale_cohort = (
db.execute(
"""
SELECT COUNT(*) FROM cohort_metrics cm
WHERE cm.member_digest IS NULL
OR cm.leader_hash IN (
SELECT backtest_hash FROM backtest_runs
WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
OR EXISTS (
SELECT 1 FROM backtest_runs r
WHERE r.stage IS cm.stage
AND r.prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
""",
(payload,),
).fetchone()[0]
if "cohort_metrics" in tables
else 0
)
stale_paired = (
db.execute(
"""
SELECT COUNT(*) FROM backtest_paired_metrics pm
WHERE pm.challenger_hash IN (
SELECT backtest_hash FROM backtest_runs
WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
OR pm.benchmark_hash IN (
SELECT backtest_hash FROM backtest_runs
WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
""",
(payload,),
).fetchone()[0]
if "backtest_paired_metrics" in tables
else 0
)
return stale_cohort, stale_paired
# §1 reports the configuration the validation stages ranked first, and the paired-metrics
# producer below has to be told what it is. Left to itself the producer ranks the registry on
# raw Sharpe, which is a second selector beside this one - it applies neither the common-support
# re-ranking a conformal candidate forces nor the restrictions the resolver holds. The two agree
# on this registry today, which is why the published rows are right; etfs now has a second fully
# backtested label, so agreement is a fact about this registry rather than a property of either
# ranking. So the selection is resolved here, before anything is written, and the two rankings
# are required to agree rather than assumed to.
# %%
_HOLDOUT_STAGES = ("signal", "allocation", "risk_overlay")
_FIELD = (
pl.concat(
[
explorer.best(stage=s, top_n=2000, prediction_hashes=LIVE_PREDICTIONS)
for s in _HOLDOUT_STAGES
],
how="diagonal_relaxed",
)
.filter(pl.col("family") != "benchmark")
.sort("sharpe", descending=True)
.unique(subset=["prediction_hash"], keep="first", maintain_order=True)
)
ADMITTED = frozenset(_FIELD["backtest_hash"].to_list())
top_signal = _FIELD.head(1)
TOP_HASH = top_signal.row(0, named=True)["backtest_hash"]
TOP_PHASH = top_signal.row(0, named=True)["prediction_hash"]
RANK1_FAMILY = top_signal.row(0, named=True)["family"]
RANK1_CONFIG = top_signal.row(0, named=True)["config_name"]
# Every label the case study declares is backtested equal-weight and competes here; what wins
# decides the label everything below is keyed to. Reading `labels.primary` instead was only ever
# right by coincidence, and etfs has a second fully backtested label, so the coincidence is not
# one to rely on. The primary is kept as the declared value, to say when the two differ.
SELECTED_LABEL = top_signal.row(0, named=True)["label"]
# The resolver is given the same field, not the whole registry: it re-ranks on common timestamp
# support wherever a conformal candidate is present, so a row this notebook never admitted would
# otherwise decide how far the intersection reaches and therefore which admitted row wins.
#
# `resolve_solvent_carrier` rather than the bare lineage resolver, so a selected configuration whose
# equity reached zero is refused rather than reported. A long-short book with no margin call keeps
# compounding through zero, so every metric it reports after that point - including a Sharpe high
# enough to top a ranking - is arithmetic on a balance that no longer exists.
CARRIER = resolve_solvent_carrier(CASE_STUDY, admitted=ADMITTED)
if CARRIER["val_backtest_hash"] != TOP_HASH:
raise RuntimeError(
"this notebook's ranking and the canonical resolver disagree on the selected configuration: "
f"{TOP_HASH} against {CARRIER['val_backtest_hash']}. Everything below reports the "
"first and the paired rows would be written against the second, so the decay would "
"compare the right holdout against a different strategy."
)
# %%
have_kinds, have_cohorts = _derived_table_state()
stale_cohort, stale_paired = _stale_derived_rows(LIVE_PREDICTIONS)
missing_kinds = sorted(set(PAIRED_KINDS) - have_kinds)
missing_cohorts = sorted(set(COHORT_TYPES) - have_cohorts)
if missing_cohorts or stale_cohort:
reason = (
f"{len(missing_cohorts)} granularity(ies) missing"
if missing_cohorts
else f"{stale_cohort} row(s) led by a retired prediction"
)
counts = compute_and_register(CASE_STUDY, prediction_hashes=LIVE_PREDICTIONS)
print(
f"cohort_metrics: recomputed {sum(counts.values())} rows across {sorted(counts)} ({reason})"
)
else:
print(
f"cohort_metrics: {len(have_cohorts)} granularities present and every leader is live, "
"nothing recomputed"
)
if missing_kinds or stale_paired:
reason = (
f"missing {', '.join(missing_kinds)}"
if missing_kinds
else f"{stale_paired} pair(s) challenged by a retired prediction"
)
# `replace_all=False` is additive: the pairs this call does not produce stay. That is
# what this notebook has always done, and it is stated now because the argument
# decides what the table a reader loads below contains.
rows = populate_paired_metrics(
CASE_STUDY,
prediction_hashes=LIVE_PREDICTIONS,
carrier=CARRIER,
replace_all=False,
)
written = sum(1 for r in rows if "skip" not in r)
print(f"backtest_paired_metrics: wrote {written} pairs ({reason})")
else:
print(
"backtest_paired_metrics: every kind this notebook reads is present and every "
"challenger is live, nothing recomputed"
)
# Both producers are additive: they write the rows this run's population produces and leave
# every other row where it is. So the rebuild above cannot remove a cohort or a pair built over
# a generation that has since been retired, and re-checking only which *kinds* are present would
# report a table that still holds them as clean. The staleness check is run again and its
# residual stated.
#
# Neither producer is asked to prune. `_prune_paired_metrics` deletes the complement of what a
# run wrote, and a run scoped to one label's live predictions has not written the pairs Chapter
# 20 registered for the other labels - pruning to that partial set would delete them. The rows
# are inert for this notebook either way: every reader below resolves a pair by an explicit
# challenger hash taken from the live lineage, never by scanning the table, so a stale row has
# nothing here that would read it. It is stated rather than ignored because that is a property
# of this notebook's readers and not of the table.
residual_cohort, residual_paired = _stale_derived_rows(LIVE_PREDICTIONS)
if residual_cohort or residual_paired:
print(
f"derived tables still hold {residual_cohort} cohort row(s) and {residual_paired} "
"pair(s) from retired generations; no reader below resolves either"
)
else:
print("derived tables hold no rows from retired generations")
have_kinds, have_cohorts = _derived_table_state()
STILL_MISSING_KINDS = sorted(set(PAIRED_KINDS) - have_kinds)
if STILL_MISSING_KINDS:
# Named rather than left to surface as an empty frame eight cells later. The holdout kinds
# are absent until the holdout has been evaluated, which is a stage that has not run rather
# than a failure of this one.
print(f" still unavailable: {', '.join(STILL_MISSING_KINDS)}")
def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
"""Compact `point [lo, hi]` formatter with NULL-safety."""
if point is None:
return "-"
p = format(point, fmt)
if lo is None or hi is None:
return f"{p} [-, -]"
return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"
def _fmt(val: float | None, fmt: str = ".4f") -> str:
return "-" if val is None else format(val, fmt)
# The stage transitions below are read off `champion_lineage`, which takes the highest-Sharpe
# backtest at each stage independently. Two consecutive entries therefore share a prediction and
# nothing else: `populate_paired_metrics` says so where it builds them - "the pair is a stage
# comparison, not a demonstrated parent and child". Nothing in `backtest_paired_metrics` records
# which axes moved, so a difference produced by three simultaneous changes and one produced by a
# single change are stored identically and read alike.
#
# The cost stage makes that concrete rather than theoretical. `cost_sensitivity` is a monotone
# grid - the same strategy priced at seventeen cost levels - so its Sharpe maximum is the
# zero-cost point by construction, in this case study and in every other. Taking it as "the cost
# stage's leader" makes the allocation-to-cost transition a comparison of the same returns with
# friction switched off, and reports the saving as a gain the cost model contributed. Its interval
# is tight and its p-value is zero because the two series are nearly identical, which is a
# property of the comparison rather than evidence for it.
#
# So each transition prints the axes that actually differ, and the reading is qualified when more
# than one of them moved.
_COST_KEYS = ("commission", "slippage")
def _axes(backtest_hash: str) -> dict:
"""The comparable axes of one registered backtest, read from its stored specification."""
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as db:
row = db.execute(
"SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?", (backtest_hash,)
).fetchone()
if row is None or not row[0]:
return {}
spec = json.loads(row[0])
config = spec.get("backtest_config", {})
strategy = spec.get("strategy", {})
axes = {
"allocator": strategy.get("allocation", {}).get("method"),
"concentration": strategy.get("signal", {}).get("top_k"),
"risk overlay": (strategy.get("risk") or {}).get("name"),
}
for key in _COST_KEYS:
axes[key] = json.dumps(config.get(key, {}), sort_keys=True)
return axes
def _priced(axes: dict) -> bool:
"""Whether this backtest charges anything at all to trade."""
for key in _COST_KEYS:
model = json.loads(axes.get(key) or "{}")
if any(type(v) in (int, float) and v > 0 for v in model.values()):
return True
return False
def _changed(challenger_hash: str, benchmark_hash: str) -> list[str]:
"""Which axes differ between a transition's two sides, in reading order."""
chal, bench = _axes(challenger_hash), _axes(benchmark_hash)
if not chal or not bench:
return []
moved = [name for name in chal if chal[name] != bench[name]]
# The two cost keys always move together here and name one decision, so they read as one.
if set(_COST_KEYS) <= set(moved):
moved = [m for m in moved if m not in _COST_KEYS] + ["cost model"]
return moved
# %% [markdown]
# ## §1 What the strategy phase inherits
#
# The strategy phase does not choose a model. It receives one: **the configuration with the
# highest validation backtest Sharpe** across the holdout-eligible stages, where a configuration
# is the whole package - model, label, feature set, backtest settings and any risk overlay - and
# the checkpoint is part of it. Everything below describes that one configuration rather than
# comparing candidates. Note which metric selected it: not the information coefficient, which
# orders nothing here, and not a chapter's headline figure. [`13_model_analysis`](13_model_analysis.ipynb)
# is where the population was described; the choosing happened in the backtest stages.
#
# The prediction-side information coefficient printed below is the upstream prior on everything
# that follows. If the ranking the strategy trades on is not credibly different from zero, a
# strategy Sharpe that looks good is a fact about the portfolio construction and the window rather
# than about the signal - and the interval on the IC is what says which case this is.
# %%
_db = CASE_DIR / "run_log" / "registry.db"
with sqlite3.connect(str(_db)) as _con:
_row = _con.execute(
"SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
"ic_hac_lag, ic_pct_positive "
"FROM prediction_metrics WHERE prediction_hash = ?",
(TOP_PHASH,),
).fetchone()
ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row
print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={SELECTED_LABEL}")
print(f" prediction_hash={TOP_PHASH}, backtest_hash={TOP_HASH}")
print()
print("Daily-pooled IC (validation):")
print(f" IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')} (HAC, lag={int(ic_lag)})")
print(f" t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3f}")
print(f" n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f" CI status: {ci_status(ic_lo, ic_hi)}")
# %% [markdown]
# **Read the interval, not the point.** A daily-pooled IC is an average of many daily rank
# correlations on an overlapping label, so consecutive days are dependent and the HAC correction is
# what makes its interval mean anything. Whether that interval clears zero is the question; the
# magnitude is small on this panel either way, because the instruments are themselves diversified
# and there is little idiosyncratic variation left to rank.
#
# **What follows from each answer.** An interval clearing zero sets the expectation that the
# strategy Sharpe in §3 should also be credible, and makes it worth asking what went wrong if it is
# not. An interval spanning zero says the opposite: whatever §3 reports has no upstream support,
# and a good Sharpe there needs explaining rather than celebrating.
#
# **Kill conditions are not declared in `setup.yaml`.** §9 evaluates two universal gates: the
# validation Sharpe interval's lower bound against zero, and the holdout strategy-versus-equal-
# weight paired interval against zero on the negative side. Both are reported as pass, partial or
# fail, and neither is a judgement on the strategy.
# %% [markdown]
# ## §2 Where the leader sits in the search that produced it
#
# A leader is the maximum of a search, and the maximum of a search is not the same quantity as the
# performance of a strategy chosen in advance. The table below puts it back in context: how many
# backtests the signal stage produced, where their Sharpe ratios fell, and how far above the middle
# of that distribution the leader sits.
#
# The wider that distribution and the more configurations in it, the more of the leader's margin is
# attributable to having looked. §9's selection-adjusted statistics price that directly; this
# section is where the shape it prices becomes visible.
# %%
ctx = explorer.search_context("signal", prediction_hashes=LIVE_PREDICTIONS)
search_table = pl.DataFrame(
[
{"metric": "Total signal backtests", "value": f"{ctx['total']:,}"},
{"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
{"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
{"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
{"metric": "% positive Sharpe", "value": f"{ctx['pct_positive']:.1f}%"},
{"metric": "Top-by-Sharpe in this sweep", "value": f"{ctx['champion_sharpe']:.3f}"},
{"metric": "Top-by-Sharpe percentile", "value": f"{ctx['champion_percentile']:.1f}%"},
]
)
print("Signal-stage search context:")
print(search_table)
# %%
with sqlite3.connect(str(_db)) as _con:
_famdf = pl.DataFrame(
_con.execute(
"""
SELECT
t.family,
bm.sharpe,
bm.sharpe_ci95_lo,
bm.sharpe_ci95_hi
FROM backtest_metrics bm
JOIN backtest_runs b ON bm.backtest_hash = b.backtest_hash
JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
JOIN training_runs t ON p.training_hash = t.training_hash
WHERE b.stage = 'signal'
AND p.split = 'validation'
AND bm.sharpe IS NOT NULL
AND (bm.num_trades IS NULL OR bm.num_trades > 0)
AND b.prediction_hash IN (SELECT value FROM json_each(?))
""",
(json.dumps(LIVE_PREDICTIONS),),
).fetchall(),
schema=["family", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"],
orient="row",
)
family_summary = (
_famdf.group_by("family")
.agg(
n=pl.len(),
sharpe_median=pl.col("sharpe").median(),
sharpe_q25=pl.col("sharpe").quantile(0.25),
sharpe_q75=pl.col("sharpe").quantile(0.75),
sharpe_max=pl.col("sharpe").max(),
pct_positive=((pl.col("sharpe") > 0).sum() / pl.len() * 100),
)
.sort("sharpe_median", descending=True)
)
print("Family-level signal-stage Sharpe summary:")
print(family_summary)
# %%
fig, ax = plt.subplots(figsize=(9, 4))
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()
ax.errorbar(
medians,
y,
xerr=[medians - q25, q75 - medians],
fmt="o",
color="#1565C0",
ecolor="#5B9BD5",
elinewidth=2.0,
capsize=4,
label="median ±IQR",
)
ax.scatter(maxima, y, marker="x", color="#C62828", s=60, label="max", zorder=5)
ax.axvline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
ax.set_title("Baseline Sharpe by family: interquartile range and maximum")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
fig.tight_layout()
show_with_alt(
fig,
"Validation Sharpe by model family, one row per family: a marker at the median with a bar "
"spanning the interquartile range, a cross at the family maximum, and a dashed line at zero.",
)
# %% [markdown]
# **The median and the maximum answer different questions, which is why both are drawn.** A
# family's median says what a configuration drawn from it typically does; its maximum says what its
# strongest one did, and the strongest is what a search returns. A family whose maximum stands far above
# its own median is a family whose leader is mostly a draw from a wide distribution.
#
# **Where the families sit relative to each other matters less than how much they overlap.** Read
# the interquartile bars: where they cover each other, the ordering between those families is not
# something the signal stage decided, and the selected configuration's family may hold that
# position because of the stages that came after rather than because of the signal.
#
# **The lineage waterfall below is where that is settled.** It tracks one prediction through
# signal, allocation, cost and risk, so a leader that arrives in the lead late - lifted by portfolio
# construction rather than by its ranking - is visible as a rising line rather than as a high
# starting point. Neither shape is better; they are different claims about where the performance
# came from.
# %%
lineage = explorer.champion_lineage(TOP_PHASH)
ci_lo: dict[str, float] = {}
ci_hi: dict[str, float] = {}
for stage_name, info in lineage.items():
bm = load_backtest_metrics(CASE_STUDY, backtest_hash=info["backtest_hash"])
if not bm.is_empty():
row = bm.row(0, named=True)
ci_lo[stage_name] = row.get("sharpe_ci95_lo")
ci_hi[stage_name] = row.get("sharpe_ci95_hi")
print("Lineage stages present for rank-1 prediction:")
for s, info in lineage.items():
lo, hi = ci_lo.get(s), ci_hi.get(s)
print(f" {s}: hash={info['backtest_hash']}, Sharpe={_fmt_ci(info['sharpe'], lo, hi)}")
# %%
fig = plot_sharpe_waterfall(lineage, ci_lo=ci_lo, ci_hi=ci_hi)
show_with_alt(
fig,
"One bar per stage of the locked lineage, from the baseline backtest through allocation, cost "
"and risk overlay, each carrying its block-bootstrap interval as an error bar.",
)
# %%
# Stage-transition deltas via load_paired_metrics — never recompute paired
# metrics inline. ETFs is one of the few CSs whose rank-1 prediction has
# all four pipeline stages registered against it, so each transition has
# a populated paired row.
# The pairs are the consecutive stages this prediction actually has, taken in
# STAGE_SEQUENCE order, which is the order the backtests run: size positions,
# apply risk controls, then measure what realistic costs take off the winner.
present = [s for s in STAGE_SEQUENCE if lineage.get(s)]
for stage_name in STAGE_SEQUENCE:
if lineage.get(stage_name) is None:
print(f"Stage {stage_name}: not run for this prediction")
for prev_stage, stage_name in zip(present, present[1:]):
kind = f"{prev_stage}_leader"
stage_info = lineage[stage_name]
# champion_lineage selects each stage independently, so two entries can be
# siblings rather than parent and child. The producer writes a paired row
# only where the later stage carries the earlier one's whole strategy
# prefix; asking for one it declined to write is normal, and the reason is
# worth naming rather than reporting as a missing row.
if not descends_from(
stage_info.get("_strategy", {}),
lineage[prev_stage].get("_strategy", {}),
prev_stage,
):
print(
f"Stage {stage_name}: not paired against {kind} - this run does not "
f"descend from the {prev_stage} leader, so the two are siblings and "
f"the difference between them is not a stage effect."
)
continue
pair = load_paired_metrics(
CASE_STUDY,
challenger_hash=stage_info["backtest_hash"],
benchmark_kind=kind,
)
if pair.is_empty():
print(f"Stage {stage_name}: paired row missing for kind={kind}")
continue
r = pair.row(0, named=True)
print(f"{stage_name} challenger vs {kind}:")
print(
f" sharpe_diff = "
f"{_fmt_ci(r['sharpe_diff'], r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}"
)
print(f" p_value = {r['p_value']:.3f}")
print(f" prob_challenger_wins = {r['prob_challenger_wins']:.3f}")
print(f" CI status: {ci_status(r['sharpe_diff_ci95_lo'], r['sharpe_diff_ci95_hi'])}")
moved = _changed(stage_info["backtest_hash"], r["benchmark_hash"])
print(f" axes that differ: {', '.join(moved) if moved else 'none'}")
chal_axes, bench_axes = _axes(stage_info["backtest_hash"]), _axes(r["benchmark_hash"])
if "cost model" in moved and _priced(bench_axes) and not _priced(chal_axes):
# Read the other way round, which is the direction the comparison supports.
print(
f" the challenger is the frictionless member of the cost grid and the benchmark is "
f"priced, so what this measures is the price of friction: {r['sharpe_diff']:.3f} "
f"Sharpe, not a gain the cost stage produced"
)
if len(moved) > 1:
print(
f" {len(moved)} axes moved at once, so none of the difference is attributable to "
f"{stage_name} specifically"
)
print()
# %% [markdown]
# Read the transitions printed above against the order the backtests run:
# positions are sized, risk controls are applied to the sized strategy, and
# costs are charged against the strategy that survives both. Each printed row
# compares a stage against the leader of the stage before it, so a positive
# `sharpe_diff` is what that one step added and nothing else.
#
# Two readings need care. A confidence interval that straddles zero means the
# step's contribution is not resolved at this sample - the point estimate still
# says which way it leaned, and `prob_challenger_wins` says how often it led
# across bootstrap draws, but neither is a rejection. And a transition reported
# as siblings rather than a pair is not a gap in the evidence: the two stages
# were selected independently and the later one does not carry the earlier
# one's configuration, so their difference mixes the stage with everything else
# that differs between them, and no paired row is written for it.
# %%
# The stage is named rather than defaulted. This notebook plots one line per
# allocator, so it wants the allocation stage, and it happens to be what the old
# default gave it - but most predictions in this registry hold signal rows and no
# allocation rows, so a carrier that had not been through the allocator menu used
# to make this cell raise with the stage nowhere in the call.
conc_df = explorer.concentration_curve(TOP_PHASH, stage="allocation")
if not conc_df.is_empty():
fig = plot_concentration_curve(conc_df)
show_with_alt(
fig,
"Validation Sharpe against the number of positions held, one marker per top-k with the "
"best allocator at that k annotated beside it and the best k highlighted.",
)
best_per_k = conc_df.sort("sharpe", descending=True).group_by("top_k").first().sort("top_k")
print("Allocation: best Sharpe by top_k:")
print(best_per_k.select("top_k", "allocator", "sharpe", "max_drawdown"))
else:
print("No concentration data - allocation stage absent for this prediction.")
# %% [markdown]
# **Concentration is the portfolio decision the ranking does not make.** Holding fewer funds uses
# more of the signal's confidence and less of the universe's diversification; holding more does the
# reverse. On a cross-asset universe that trade has a second edge, because the funds are drawn from
# equities, bonds, commodities and currencies, and a tight selection can end up inside one of those
# rather than across them.
#
# The curve is read for where it flattens rather than for its maximum. A maximum picked off this
# scan is a selection over the same validation window everything else was selected on, and the
# concentration the lineage actually uses is the one the allocation stage registered.
# %% [markdown]
# ## §3 What the strategy earned, with its uncertainty
#
# One strategy, one validation window, every figure with the interval the block bootstrap gives it.
# The block bootstrap rather than an independent one because daily strategy returns are serially
# dependent, and resampling them independently would produce an interval far tighter than the data
# support.
#
# The equity overlay puts the strategy against the equal-weight ETF universe over the same dates.
# That benchmark is the honest comparator for this case study: it holds the same instruments with
# no model at all, so the gap between the two curves is what the ranking bought.
# %%
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)
spec_block = {
"case_study": CASE_STUDY,
"family": RANK1_FAMILY,
"config_name": RANK1_CONFIG,
"label": SELECTED_LABEL,
"signal_method": lineage["signal"].get("signal_method"),
"top_k": lineage["signal"].get("top_k"),
"allocation": lineage.get("allocation", {}).get("allocator"),
"cost_assumption": "no costs at signal stage; sensitivity in §5",
"risk_overlay": lineage.get("risk_overlay", {}).get("risk_name"),
"validation_window_periods": int(full["n_periods"]),
"num_trades": int(full["num_trades"]) if full["num_trades"] is not None else None,
"avg_turnover": full.get("avg_turnover"),
"bootstrap_block_length": int(full["bootstrap_block_length"]),
"bootstrap_n": int(full["bootstrap_n"]),
}
print("The selected configuration specification (signal stage, validation window):")
for k, v in spec_block.items():
print(f" {k}: {v}")
# Audit: bootstrap_block_length resolves from the label horizon (21 trading
# days for fwd_ret_21d). Rebalance cadence in setup.yaml is monthly
# month-end (~21 bars) - the two encode the same autocorrelation scale.
_block = int(full["bootstrap_block_length"])
print(
f" audit: bootstrap_block_length={_block} days "
f"(rebalance cadence={setup['decision']['cadence']} ≈ 21 bars)"
)
# %%
sharpe_status = ci_status(full["sharpe_ci95_lo"], full["sharpe_ci95_hi"])
sortino_status = ci_status(full["sortino_ci95_lo"], full["sortino_ci95_hi"])
ann_status = ci_status(full["ann_return_ci95_lo"], full["ann_return_ci95_hi"])
mdd_status = ci_status(full["max_dd_ci95_lo"], full["max_dd_ci95_hi"])
calmar_status = ci_status(full["calmar_ci95_lo"], full["calmar_ci95_hi"])
def _hrow(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}
# The selection-adjusted columns join in from cohort_metrics and carry no per-row interval, so
# their CI cells read as unavailable rather than as a computed bound.
headline = pl.DataFrame(
[
_hrow(
"Sharpe",
_fmt(full["sharpe"]),
_fmt(full["sharpe_ci95_lo"]),
_fmt(full["sharpe_ci95_hi"]),
sharpe_status,
),
_hrow(
"Sortino",
_fmt(full["sortino"]),
_fmt(full["sortino_ci95_lo"]),
_fmt(full["sortino_ci95_hi"]),
sortino_status,
),
_hrow(
"Annualized return",
_fmt(full["cagr"]),
_fmt(full["ann_return_ci95_lo"]),
_fmt(full["ann_return_ci95_hi"]),
ann_status,
),
_hrow(
"Max drawdown",
_fmt(full["max_drawdown"]),
_fmt(full["max_dd_ci95_lo"]),
_fmt(full["max_dd_ci95_hi"]),
mdd_status,
),
_hrow(
"Calmar",
_fmt(full["calmar"]),
_fmt(full["calmar_ci95_lo"]),
_fmt(full["calmar_ci95_hi"]),
calmar_status,
),
{
"metric": "PSR p-value (H0: SR≤0)",
"point": _fmt(full["psr_pvalue"]),
"ci95_lo": "-",
"ci95_hi": "-",
"status": "n/a",
},
{
"metric": "DSR (selection-adjusted)",
"point": _fmt(full["dsr"]),
"ci95_lo": "-",
"ci95_hi": "-",
"status": "n/a",
},
{
"metric": "Expected max Sharpe",
"point": _fmt(full["expected_max_sharpe"]),
"ci95_lo": "-",
"ci95_hi": "-",
"status": "n/a",
},
{
"metric": "PBO",
"point": _fmt(full["pbo"]),
"ci95_lo": "-",
"ci95_hi": "-",
"status": "n/a",
},
]
)
print("The selected configuration headline metrics with 95% CIs:")
print(headline)
# %%
# Forest plot: rank-1 metrics with CI bars + reference lines
ew_val = load_benchmark_metrics(CASE_STUDY, SELECTED_LABEL, period="validation")
forest_metrics = [
("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
("Ann. return", full["cagr"], full["ann_return_ci95_lo"], full["ann_return_ci95_hi"]),
]
fig, ax = plt.subplots(figsize=(8, 4))
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.errorbar(
points,
y,
xerr=[points - los, his - points],
fmt="o",
color="#1565C0",
ecolor="#5B9BD5",
elinewidth=2.0,
capsize=4,
markersize=7,
)
ax.axvline(0, color="#9E9E9E", linestyle="--", linewidth=0.8)
# `load_benchmark_metrics` returns None when the benchmark JSON is absent, which its docstring
# states and which is the ordinary case in a workspace that holds only what this run wrote. The
# reference line is a comparison against the equal-weight baseline, so without it there is
# nothing to draw - and drawing the rest of the forest is still worth doing. Same shape as the
# risk-overlay line below, which has always been conditional.
if ew_val is not None:
ax.axvline(
ew_val["sharpe"],
color="#43A047",
linestyle=":",
linewidth=1.0,
label=f"EW validation Sharpe ({ew_val['sharpe']:.2f})",
)
risk_sharpe = lineage.get("risk_overlay", {}).get("sharpe")
if risk_sharpe is not None:
ax.axvline(
risk_sharpe,
color="#E53935",
linestyle=":",
linewidth=1.0,
label=f"Risk-overlay leading Sharpe ({risk_sharpe:.2f})",
)
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Value")
ax.set_title("The selected configuration Headline Metrics with 95% CIs")
ax.legend(loc="lower right", fontsize=8, frameon=False)
fig.tight_layout()
show_with_alt(
fig,
"One row per headline metric, each a point estimate with a bar spanning its 95% interval, "
"against a dashed line at zero and dotted reference lines for the equal-weight and risk- "
"overlay comparisons where those are available.",
)
# %%
# Equity-curve overlay vs validation EW benchmark
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
pl.read_parquet(strat_returns_path)
.sort("timestamp")
.with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
.select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)
bench_val = (
load_benchmark_returns(CASE_STUDY, SELECTED_LABEL, period="validation")
.with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
.select(pl.col("ts"), pl.col("ew_return").alias("benchmark"))
)
aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)
cum_strat = np.cumprod(1 + aligned["strategy"].to_numpy()) - 1
cum_bench = np.cumprod(1 + aligned["benchmark"].to_numpy()) - 1
fig, ax = plt.subplots(figsize=(10, 4.2))
ax.plot(
aligned["ts"],
cum_strat,
color="#1565C0",
linewidth=1.2,
label="The selected configuration strategy",
)
ax.plot(aligned["ts"], cum_bench, color="#43A047", linewidth=1.2, label="EW universe")
ax.axhline(0, color="#9E9E9E", linewidth=0.6, linestyle="--")
ax.set_ylabel("Cumulative return")
ax.set_title("Validation-window cumulative return: rank-1 strategy vs EW universe")
ax.legend(loc="best", frameon=False)
fig.tight_layout()
show_with_alt(
fig,
"Cumulative return over the validation window, one line for the selected strategy and one for "
"the equal-weight universe, against a dashed line at zero.",
)
# %% [markdown]
# **The lower bound of the Sharpe interval is what the first gate reads**, and it answers a
# narrower question than the point estimate: not how well the strategy did, but whether this window
# is consistent with it having no edge at all. A point estimate cannot answer that; an interval can.
#
# **The probabilistic Sharpe ratio is a second, parametric answer to the same question.** It
# corrects for the skew and kurtosis a Sharpe ratio assumes away. Where the two disagree, the
# bootstrap interval is the one that made fewer assumptions.
#
# **The drawdown interval is not a test.** A drawdown is negative by construction, so an interval
# excluding zero says nothing. What it bounds is the magnitude, and the width is what to size
# against rather than the point.
#
# **The selection-adjusted statistics are what price the search**, and they read as unavailable
# rather than as favourable when `cohort_metrics` holds no row for this lineage. A deflated Sharpe
# absent from a table is not a deflation of zero. Where they are absent the in-sample guard is the
# probabilistic Sharpe alone and the decisive evidence moves to §6 - a weaker position, and worth
# saying so.
# %% [markdown]
# ## §4 Risk and drawdown analysis
#
# Risk metrics use the validation-window strategy returns paired
# against the validation EW benchmark. The drawdown panel surfaces the
# worst episode and recovery; rolling Sharpe and rolling beta locate
# when the strategy decoupled from the universe. ETFs are among the
# most liquid instruments in this book, so drawdowns reflect signal
# decay or cross-asset rotation reversals rather than execution
# slippage.
# %%
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()
pa = PortfolioAnalysis(
returns=strat_arr,
benchmark=bench_arr,
dates=ts_arr,
periods_per_year=PERIODS_PER_YEAR,
)
dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)
# %%
fig = plot_equity_drawdown(strat_returns_path)
show_with_alt(
fig,
"Two panels sharing a date axis: cumulative return of the strategy above, and its peak-to- "
"trough drawdown below.",
)
# %%
# Rolling Sharpe + rolling beta (window 126 ~ 6 months)
roll = pa.compute_rolling_metrics(windows=[126], metrics=["sharpe", "beta"])
print("Rolling-window keys:")
print({k: type(v).__name__ for k, v in roll.items()} if isinstance(roll, dict) else roll)
# Tail risk straight from registry
tail_table = pl.DataFrame(
[
{"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
{"metric": "VaR 95% (daily)", "value": f"{full['var_95']:.4f}"},
{"metric": "CVaR 95% (daily)", "value": f"{full['cvar_95']:.4f}"},
{"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
{"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
{"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
]
)
print()
print("Tail risk profile:")
print(tail_table)
# %%
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
if fold_df.height > 0:
print(f"Per-fold breakdown ({fold_df.height} folds):")
print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
print()
print(f"Fold Sharpe range: [{fold_df['sharpe'].min():.3f}, {fold_df['sharpe'].max():.3f}]")
print(f"Fold Sharpe std: {fold_df['sharpe'].std():.3f}")
else:
print(
"Per-fold metrics not populated for this backtest_hash. "
"ETFs rank-1 was bootstrapped against the consolidated validation "
"window rather than expanding-window folds; the strategy headline "
"uses the consolidated CI, and §6's val→ho paired test substitutes "
"for an explicit per-fold stability check."
)
# %% [markdown]
# **Depth and duration are different risks and only one is in the Sharpe.** A strategy can recover
# quickly from a deep fall or sit underwater for years after a shallow one, and the second is what
# ends a mandate. The panel reports both.
#
# **Kurtosis decides whether the interval above can be trusted.** Heavy tails mean the observed
# Sharpe rests on fewer effective observations than the sample size suggests, and a risk overlay
# clipping the left tail harder than the right shows up here as skew moving with the overlay rather
# than with the signal.
#
# **Rolling Sharpe and rolling beta locate when the strategy stopped tracking its universe.** A
# cross-asset strategy that is long equities most of the time has a beta near one and no
# diversifying behaviour to show for itself. Where beta falls is where the model rotated into bonds
# or commodities, and whether that helped is in the rolling Sharpe over the same dates.
# %% [markdown]
# ## §5 How much friction the strategy tolerates
#
# ETFs are among the most liquid instruments in this book: the spread on the largest funds is a
# fraction of a basis point, while sector, international and commodity funds cost several. The cost
# stage walked a per-leg grid over the selected configuration; the curve below is how its Sharpe
# responds, with bands marking the most-liquid end of the universe and the typical one.
#
# The reading that matters is the distance between the declared cost and where the curve crosses
# zero, not the Sharpe at any single level. [`17_costs`](17_costs.ipynb) computes that crossing
# directly and reports it as a bound when the grid does not reach it.
# %%
with sqlite3.connect(str(_db)) as _con:
cost_df = pl.DataFrame(
_con.execute(
"""
SELECT
b.spec_json,
bm.sharpe,
bm.sharpe_ci95_lo,
bm.sharpe_ci95_hi,
bm.max_drawdown
FROM backtest_runs b
JOIN backtest_metrics bm ON bm.backtest_hash = b.backtest_hash
JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
WHERE b.stage = 'cost_sensitivity'
AND p.split = 'validation'
AND bm.sharpe IS NOT NULL
AND (bm.num_trades IS NULL OR bm.num_trades > 0)
"""
).fetchall(),
schema=[
"spec_json",
"sharpe",
"sharpe_ci95_lo",
"sharpe_ci95_hi",
"max_drawdown",
],
orient="row",
)
def _cost_bps(spec_str: str) -> float:
"""Per-leg cost (bps) extracted from locked spec.
Cost rates live under backtest_config.commission.rate +
backtest_config.slippage.rate as decimal fractions; sum × 10_000
is total per-leg cost in bps.
"""
spec = json.loads(spec_str)
bc = spec.get("backtest_config", {})
comm = bc.get("commission", {}) or {}
slip = bc.get("slippage", {}) or {}
return float((comm.get("rate", 0) + slip.get("rate", 0)) * 10000)
cost_df = cost_df.with_columns(
pl.col("spec_json").map_elements(_cost_bps, return_dtype=pl.Float64).alias("cost_bps")
)
cost_curve = (
cost_df.group_by("cost_bps")
.agg(
sharpe_max=pl.col("sharpe").max(),
sharpe_min=pl.col("sharpe").min(),
sharpe_median=pl.col("sharpe").median(),
sharpe_ci_lo=pl.col("sharpe_ci95_lo").min(),
sharpe_ci_hi=pl.col("sharpe_ci95_hi").max(),
n=pl.len(),
)
.sort("cost_bps")
)
print("Cost sensitivity curve (validation, all configs):")
print(cost_curve)
# %%
fig, ax = plt.subplots(figsize=(9, 4))
xs = cost_curve["cost_bps"].to_numpy()
ax.fill_between(
xs,
cost_curve["sharpe_ci_lo"].to_numpy(),
cost_curve["sharpe_ci_hi"].to_numpy(),
alpha=0.18,
color="#5B9BD5",
label="best–worst CI envelope across configs",
)
ax.plot(
xs,
cost_curve["sharpe_median"].to_numpy(),
color="#1565C0",
linewidth=1.4,
label="median Sharpe",
)
ax.plot(
xs,
cost_curve["sharpe_max"].to_numpy(),
color="#43A047",
linewidth=1.0,
linestyle="--",
label="best-config Sharpe",
)
ax.axhline(0, color="#9E9E9E", linewidth=0.8, linestyle="--")
# Realistic ETF friction:
# - Most-liquid ETF spread (SPY/QQQ/IWM): 0.2-1 bps round-trip
# - Typical ETF spread: 2-5 bps round-trip
ax.axvspan(0.2, 1.0, color="#43A047", alpha=0.10, label="most-liquid ETF (0.2–1 bps)")
ax.axvspan(2.0, 5.0, color="#FB8C00", alpha=0.10, label="typical ETF (2 to 5 bps)")
ax.set_xlabel("Per-leg cost (bps)")
ax.set_ylabel("Sharpe (validation)")
ax.set_title("Cost sensitivity - etfs (validation, baseline+allocation+cost stages)")
ax.legend(loc="best", fontsize=8, frameon=False)
fig.tight_layout()
show_with_alt(
fig,
"Sharpe against per-leg cost in basis points: the median across configurations as a line, the "
"best configuration dashed, the best-to-worst envelope shaded, a dashed line at zero, and "
"shaded bands for the most-liquid and typical ETF spread ranges.",
)
# %%
# Breakeven cost: where the best-config Sharpe lower bound crosses zero.
crossing_rows = cost_curve.filter(pl.col("sharpe_ci_lo") > 0)
if not crossing_rows.is_empty():
breakeven = crossing_rows["cost_bps"].max()
print(f"Sharpe CI lower bound stays > 0 up to: {breakeven:.0f} bps")
else:
print("Sharpe CI lower bound never exceeds 0 across the cost grid.")
# Setup-encoded cost configuration
cost_config = setup.get("costs", {})
print()
print("Realistic ETF friction (per setup.yaml):")
for k, v in cost_config.items():
print(f" {k}: {v}")
print("See Chapter 18 for the transaction-cost framework.")
# %% [markdown]
# **Read the dispersion across configurations against the slope of the cost curve.** Where the
# spread between configurations at one cost level is wider than the effect of moving several
# levels, friction is not what decides this strategy's fate and the choice of configuration is.
# That is the ordinary situation for a monthly strategy, and it is a statement about the cadence
# rather than about the model. [`17_costs`](17_costs.ipynb) computes the crossing directly.
# %% [markdown]
# ## §6 The holdout, opened once
#
# Everything above is measured on the window the strategy was selected on. This section is the only
# out-of-sample evidence in the case study, and the holdout may be spent once: it is read here, and
# nothing that follows may be used to choose anything.
#
# Two paired tests. The first asks whether the edge held between the window that selected the
# strategy and the window that did not. The second asks whether it beat the equal-weight universe
# *inside* the holdout - the same instruments, the same dates, no model. Both come from
# `backtest_paired_metrics`, never from subtracting one Sharpe from another.
# %% [markdown]
# The anchor is the holdout backtest that replays the selected configuration's **strategy**, not
# the highest-Sharpe holdout backtest sharing its training hash. Matching on the strategy keeps the
# anchor on the lineage that was actually selected, even where an experimental side-channel
# allocator shares the holdout prediction set and posts a higher holdout Sharpe. The
# `val_rank1_self` pair is written against the canonical lineage's holdout hash, so it is findable
# only under that match.
#
# When no such backtest exists the section reports which state the registry is in and computes
# nothing further. The ordinary case is that the holdout has not been evaluated yet, which a
# reader working the case study in order meets before it has, and that is a stage that has not
# run rather than a failure of this one.
# %%
holdout_replay = resolve_holdout_self_backtest(CASE_STUDY, TOP_HASH)
HOLDOUT_AVAILABLE = holdout_replay.found
HO_HASH = holdout_replay.backtest_hash
print(f"Validation rank-1 hash: {TOP_HASH}")
if HOLDOUT_AVAILABLE:
print(f"Holdout rank-1 hash: {HO_HASH}")
else:
print(f"Holdout closure unavailable: {holdout_replay.reason}")
val_full = full
ho_full = (
load_backtest_metrics(CASE_STUDY, backtest_hash=HO_HASH).row(0, named=True)
if HOLDOUT_AVAILABLE
else None
)
# %% [markdown]
# A missing paired row is reported as missing rather than filled with NaN. A NaN decay propagates
# into the table below, into the section-9 gate, and into the assessment artifact, where it prints
# as a dash that a reader cannot distinguish from a computed zero - and the gate would then be
# evaluated on a comparison that was never made. Nothing here is computed from a substitute.
# %%
vh = None
if HOLDOUT_AVAILABLE:
val_ho_pair = load_paired_metrics(
CASE_STUDY, challenger_hash=HO_HASH, benchmark_kind="val_rank1_self"
)
if val_ho_pair.is_empty():
print(
"The holdout replay is registered but carries no val_rank1_self pair, so the "
"validation-to-holdout decay cannot be computed. The populator writes that pair "
"only when both series overlap enough to bootstrap."
)
else:
vh = val_ho_pair.row(0, named=True)
VAL_HO_DECAY_AVAILABLE = vh is not None
def _diff_row(
label: str,
v: float,
h: float,
diff: float | None,
lo: float | None,
hi: float | None,
p: float | None,
) -> dict:
return {
"metric": label,
"validation": _fmt(v, ".4f") if v is not None else "-",
"holdout": _fmt(h, ".4f") if h is not None else "-",
"diff (h-v)": _fmt(diff, ".4f") if diff is not None else "-",
"diff CI95": (f"[{_fmt(lo, '.4f')}, {_fmt(hi, '.4f')}]" if lo is not None else "-"),
"p-value": _fmt(p, ".4f") if p is not None else "-",
}
if not VAL_HO_DECAY_AVAILABLE:
print("No validation-to-holdout decay to report.")
else:
val_ho_table = pl.DataFrame(
[
_diff_row(
"Sharpe",
val_full["sharpe"],
ho_full["sharpe"],
vh["sharpe_diff"],
vh["sharpe_diff_ci95_lo"],
vh["sharpe_diff_ci95_hi"],
vh["p_value"],
),
_diff_row(
"Annualized return",
val_full["cagr"],
ho_full["cagr"],
vh["ret_diff"],
vh["ret_diff_ci95_lo"],
vh["ret_diff_ci95_hi"],
None,
),
_diff_row(
"Max drawdown",
val_full["max_drawdown"],
ho_full["max_drawdown"],
vh["max_dd_diff"],
vh["max_dd_diff_ci95_lo"],
vh["max_dd_diff_ci95_hi"],
None,
),
_diff_row(
"Information ratio",
None,
None,
vh["info_ratio"],
vh["info_ratio_ci95_lo"],
vh["info_ratio_ci95_hi"],
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。