コンテンツへスキップ
ライブラリの全資料

ETFポートフォリオ配分の階層的リスクパリティ

コード Machine Learning for Trading

サマリー

このノートブックでは、ノイズを含む共分散推定値への感度を抑えることを目的としたポートフォリオ構築手法、階層的リスクパリティ(HRP)を説明します。3段階の手順を紹介します。相関距離で資産をクラスタリングし、得られた階層に従って並べ替え、クラスタを再帰的に分割して分散の小さい側により大きな比率を割り当てます。また、共分散行列の逆行列計算が推定誤差を増幅し得る平均・分散最適化や、逆行列計算の前に共分散推定を安定化させる縮小法と、HRPを比較します。

例では、固定されたETFユニバースに各手法を適用し、ウォークフォワード分析で配分方法を比較します。このユニバースは事後的に選ばれており、この比較から一般にどの手法が優れているかは立証されないと注意しています。HRPも推定相関と分散に依存し、分散の小さい資産に保有が集中する可能性があります。また、観測数に対して資産数が少ない場合には、意図された優位性が現れないことがあります。報告された回転率には銘柄の入れ替えと配分比率の変化の両方が含まれるため、コストとタイミングは別途分析が必要です。

主なアイデア

  • HRPは相関距離で資産をクラスタ化し、階層に従って並べ、再帰的二分割で配分します。
  • 平均・分散最適化は、共分散行列の逆行列計算が推定精度の低い固有値の誤差を増幅するため、不安定になる場合があります。
  • HRPは共分散行列の逆行列への依存を減らしますが、推定相関と分散には引き続き依存します。
  • リスクに基づく配分では、幅広い分散投資が実現するとは限らず、分散の小さい資産に比率が集中することがあります。
  • ETFの比較は事後的に選定されたユニバースを前提としており、HRPが有効とされる資産数が観測数に対して多い状況を検証していません。

タグ

全文
# 06_hierarchical_risk_parity.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Hierarchical Risk Parity (HRP)
#
# **Docker image**: `ml4t`
#
# This notebook demonstrates Hierarchical Risk Parity, a modern portfolio construction
# method developed by Marcos López de Prado that reduces sensitivity to noisy covariance
# estimates through clustering, quasi-diagonalization, and recursive bisection.
#
# **Learning Objectives**:
# - Understand why classical MVO can be fragile in practice (the "Markowitz Curse")
# - Implement the three steps of HRP: clustering, quasi-diagonalization, recursive bisection
# - Visualize the asset hierarchy with dendrograms
# - Run walk-forward backtests comparing HRP to shrinkage MVO and to heuristic allocators, and
#   read the ranking against the assets-to-observations ratio that decides when HRP helps
#
# **Book Reference**: Chapter 17, Section 17.6 (Optimizing for stability with Hierarchical
# Risk Parity)
#
# **Prerequisites**: `02_mean_variance_optimization`, ETF price data

# %% [markdown]
# ## 1. Setup and Imports

# %%
"""Hierarchical Risk Parity: cluster assets and allocate using inverse variance."""

import hashlib
import warnings

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import plotly.express as px
import plotly.graph_objects as go
import polars as pl
from IPython.display import HTML, display
from ml4t.backtest import (
    BacktestConfig,
    CommissionType,
    DataFeed,
    Engine,
    ExecutionMode,
    Strategy,
)
from ml4t.backtest.config import SlippageType
from ml4t.backtest.execution.rebalancer import RebalanceConfig, TargetWeightExecutor
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.visualization import create_portfolio_dashboard
from plotly.subplots import make_subplots
from scipy.cluster.hierarchy import dendrogram, leaves_list, linkage
from scipy.spatial.distance import squareform
from sklearn.covariance import LedoitWolf

from case_studies.utils.backtest_loaders import compute_allocator_metrics
from case_studies.utils.registry.queries import load_prediction_index
from data import load_etfs
from utils.paths import get_case_study_dir, get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, add_message_title, show_plotly_with_alt, show_with_alt

# %% tags=["parameters"]
# Production defaults; Papermill overrides these values for CI testing
START_DATE = "2010-01-01"
SEED = 42

# %%
set_global_seeds(SEED)

# Track allocator fallbacks across the walk-forward backtest.
fallback_count: dict[str, int] = {}

# %% [markdown]
# ## 2. Where Mean-Variance Optimization Is Fragile
#
# Mean-variance optimization inverts an estimated covariance matrix. Inversion amplifies the
# smallest eigenvalues, which are the ones the sample estimates worst, so small errors in the
# inputs can move the weights a long way. The fragility grows with the ratio of assets to
# observations: with $N$ assets and $T$ periods the sample covariance is singular once $N > T$
# and poorly conditioned well before that.
#
# HRP is one response - never invert. Shrinkage is another - keep inverting, but pull the
# estimate toward a well-conditioned target first. This notebook runs both, and section 12 reads
# the comparison against the ratio $N/T$. That ratio is one contributor to how hard the estimate
# is, not the thing that decides which response to use: conditioning also depends on how the assets
# co-move, and a low $N/T$ over highly correlated assets can be worse conditioned than a higher one
# over independent ones. A single window over one universe cannot establish when either allocator
# helps in general; what section 12 shows is which did better here. See Section 17.6.

# %% [markdown]
# ## 3. Data Acquisition
#
# The 15 named ETFs form a fixed teaching universe. It is not a point-in-time
# membership reconstruction, so the static examples and walk-forward comparison are
# conditional on this ex-post universe and must not be read as an unbiased universe-selection test.

# %%
# Diversified ETF universe
ETF_UNIVERSE = {
    # US Equity
    "SPY": "S&P 500",
    "QQQ": "NASDAQ 100",
    "IWM": "Russell 2000",
    # International Equity
    "EFA": "EAFE Developed",
    "EEM": "Emerging Markets",
    # Fixed Income
    "AGG": "US Aggregate Bond",
    "TLT": "Long Treasury",
    "HYG": "High Yield",
    # Alternatives
    "GLD": "Gold",
    "VNQ": "Real Estate",
    "DBC": "Commodities",
    # Sectors
    "XLF": "Financials",
    "XLE": "Energy",
    "XLK": "Technology",
    "XLV": "Healthcare",
}

SYMBOLS = list(ETF_UNIVERSE.keys())
END_DATE = "2024-01-01"

# %%
# Load data from canonical ETFs
print("Loading ETF data...")
etf_data = load_etfs()
etf_filtered = etf_data.filter(
    (pl.col("symbol").is_in(SYMBOLS))
    & (pl.col("timestamp") >= pl.lit(START_DATE).str.to_datetime())
    & (pl.col("timestamp") <= pl.lit(END_DATE).str.to_datetime())
)

close_prices = (
    etf_filtered.select(["timestamp", "symbol", "close"])
    .pivot(on="symbol", index="timestamp", values="close")
    .sort("timestamp")
    .to_pandas()
    .set_index("timestamp")
    .ffill()
    .dropna()
)
returns = close_prices.pct_change().dropna()
print(f"Loaded {len(returns):,} days for {close_prices.shape[1]} ETFs")

# %% [markdown]
# ## 4. The HRP Algorithm
#
# HRP works in three steps:
#
# ### Step 1: Tree Clustering
# Build a hierarchy of assets based on their correlation structure using
# agglomerative clustering.
#
# $$d_{ij} = \sqrt{\frac{1-\rho_{ij}}{2}}$$
#
# ### Step 2: Quasi-Diagonalization
# Reorder the covariance matrix according to the clustering hierarchy
# to group similar assets together.
#
# ### Step 3: Recursive Bisection
# Allocate risk by recursively splitting the portfolio in half,
# with weights inversely proportional to the cluster variance.


# %%
def correlation_distance(corr_matrix: np.ndarray) -> np.ndarray:
    """Convert a correlation matrix to a bounded distance matrix."""
    return np.sqrt(np.clip(0.5 * (1 - corr_matrix), 0.0, 1.0))


# %% [markdown]
# #### Step 1: Hierarchical Clustering


# %%
def cluster_assets(returns: pd.DataFrame, method: str = "ward") -> np.ndarray:
    """
    Perform hierarchical clustering on assets based on correlation distance.

    Args:
        returns: DataFrame of asset returns
        method: Linkage method ('ward', 'single', 'complete', 'average')

    Returns:
        linkage matrix
    """
    corr = returns.corr().values
    dist = correlation_distance(corr)
    # Convert to condensed distance matrix (upper triangle)
    dist_condensed = squareform(dist, checks=False)
    return linkage(dist_condensed, method=method)


# %% [markdown]
# #### Step 2: Leaf Ordering


# %%
def get_quasi_diagonal_order(link: np.ndarray) -> list[int]:
    """
    Get the quasi-diagonal ordering from linkage matrix.

    This reorders assets so that similar ones are adjacent.
    """
    return list(leaves_list(link))


# %% [markdown]
# The variance of a sub-portfolio held at inverse-variance weights, which is the number each
# split compares its two halves on.


# %%
def cluster_variance(cov: np.ndarray, indices: list[int]) -> float:
    """Compute cluster variance using inverse-variance weights within a subset."""
    if len(indices) == 1:
        return float(cov[indices[0], indices[0]])
    c = cov[np.ix_(indices, indices)]
    ivp = 1 / np.diag(c)
    ivp /= ivp.sum()
    return float(np.dot(ivp, np.dot(c, ivp)))


# %% [markdown]
# #### Step 3: Recursive Bisection Allocation


# %%
def recursive_bisection(
    cov: np.ndarray,
    sorted_idx: list[int],
) -> np.ndarray:
    """Allocate weights by recursively splitting ordered clusters."""
    n = len(sorted_idx)
    weights = np.ones(n)
    sorted_position = {original: position for position, original in enumerate(sorted_idx)}
    clusters = [sorted_idx]

    while clusters:
        new_clusters = []
        for cluster in clusters:
            if len(cluster) <= 1:
                continue

            mid = len(cluster) // 2
            left = cluster[:mid]
            right = cluster[mid:]

            left_var = cluster_variance(cov, left)
            right_var = cluster_variance(cov, right)

            alpha = 1 - left_var / (left_var + right_var)
            weights[[sorted_position[i] for i in left]] *= alpha
            weights[[sorted_position[i] for i in right]] *= 1 - alpha

            if len(left) > 1:
                new_clusters.append(left)
            if len(right) > 1:
                new_clusters.append(right)

        clusters = new_clusters

    final_weights = np.zeros(n)
    final_weights[np.asarray(sorted_idx)] = weights

    return final_weights / final_weights.sum()


# %% [markdown]
# The three steps in order, from a covariance matrix to a weight vector.


# %%
def hrp_portfolio(returns: pd.DataFrame) -> np.ndarray:
    """
    Compute HRP portfolio weights.

    Args:
        returns: DataFrame of asset returns

    Returns:
        Array of portfolio weights
    """
    # Step 1: Cluster
    link = cluster_assets(returns)

    # Step 2: Quasi-diagonalize
    sorted_idx = get_quasi_diagonal_order(link)

    # Step 3: Recursive bisection
    cov = returns.cov().values
    weights = recursive_bisection(cov, sorted_idx)

    return weights


# %% [markdown]
# ## 5. Visualizing the Asset Hierarchy

# %%
# Compute linkage for visualization
link = cluster_assets(returns)
branch_colors = [COLORS["blue"], COLORS["copper"], COLORS["positive"], COLORS["neutral"]]

# Create dendrogram
fig, ax = plt.subplots(figsize=(14, 8))
dendrogram(
    link,
    labels=[ETF_UNIVERSE.get(s, s) for s in returns.columns],
    leaf_rotation=45,
    leaf_font_size=10,
    link_color_func=lambda node_id: branch_colors[(node_id - len(returns.columns)) % 4],
    ax=ax,
)
ax.set_ylabel("Distance (based on correlation)")
ax.set_xlabel("ETF")
add_message_title(
    ax,
    "Hierarchical clustering of fifteen ETFs on correlation distance",
    subtitle="Ward linkage, fixed teaching universe",
)
show_with_alt(
    fig,
    "Dendrogram of fifteen ETFs on correlation distance, with the bond and gold funds joining low on one branch and the equity and sector funds on another.",
)

# %% [markdown]
# Read the tree from the leaves upward. Two ETFs joined low down moved together over this
# history. The height of a merge is SciPy's Ward distance, which is a monotone transformation of
# the increase in within-cluster sum of squares the merge costs, so it rises as the two groups
# being joined become less alike. It is not the correlation distance between them, which is what
# the leaves were measured on.
#
# Only the leaf order reaches the allocation. Step 2 takes the left-to-right sequence this
# tree implies and reorders the covariance matrix by it; step 3 then halves that sequence by
# count at every level, without consulting the heights again. The halves it forms therefore need
# not be the clusters the tree draws, which is what section 7 traces through the first two
# splits.

# %% [markdown]
# ## 6. Quasi-Diagonal Covariance Matrix
#
# After clustering, we reorder the covariance matrix to reveal the block structure.

# %%
# Get quasi-diagonal ordering
sorted_idx = get_quasi_diagonal_order(link)
sorted_symbols = [returns.columns[i] for i in sorted_idx]
sorted_labels = [ETF_UNIVERSE.get(s, s) for s in sorted_symbols]

# Reorder covariance matrix
cov_original = returns.cov()
cov_reordered = cov_original.iloc[sorted_idx, sorted_idx]
cov_colorscale = [
    [0.0, COLORS["negative"]],
    [0.5, COLORS["silver"]],
    [1.0, COLORS["blue"]],
]

# %%
fig = make_subplots(
    rows=1, cols=2, subplot_titles=["Original Covariance Matrix", "Quasi-Diagonal (Reordered)"]
)
fig.add_trace(
    go.Heatmap(
        z=cov_original.values,
        x=cov_original.columns,
        y=cov_original.index,
        colorscale=cov_colorscale,
        zmid=0,
        showscale=False,
    ),
    row=1,
    col=1,
)
fig.add_trace(
    go.Heatmap(
        z=cov_reordered.values,
        x=sorted_symbols,
        y=sorted_symbols,
        colorscale=cov_colorscale,
        zmid=0,
        showscale=True,
    ),
    row=1,
    col=2,
)
fig.update_layout(
    title="Covariance matrix in the original and the quasi-diagonal ordering",
    height=560,
    width=1100,
    margin=dict(l=90, r=90, b=110, t=100),
)
fig.update_xaxes(tickangle=45, tickfont_size=10, automargin=True)
fig.update_yaxes(tickfont_size=10, automargin=True)
fig.update_xaxes(title_text="ETF ticker", row=1, col=1)
fig.update_xaxes(title_text="ETF ticker", row=1, col=2)
fig.update_yaxes(title_text="ETF ticker", row=1, col=1)
show_plotly_with_alt(
    fig,
    "Two covariance heatmaps side by side, the original ordering on the left and the quasi-diagonal reordering on the right, in which the large values group into blocks along the diagonal.",
)

# %% [markdown]
# ## 7. Compute HRP Weights

# %%
# Compute HRP weights
hrp_weights = hrp_portfolio(returns)

# Create weights DataFrame
weights_df = pd.DataFrame(
    {
        "Symbol": returns.columns,
        "Name": [ETF_UNIVERSE.get(s, s) for s in returns.columns],
        "HRP Weight": hrp_weights,
    }
)
weights_df = weights_df.sort_values("HRP Weight", ascending=False)

# %%
# Visualization
fig = px.bar(
    weights_df,
    x="Name",
    y="HRP Weight",
    title="HRP weight per ETF, largest first",
    color="HRP Weight",
    color_continuous_scale=[COLORS["silver_muted"], COLORS["blue"]],
)
fig.update_layout(
    height=400,
    xaxis_tickangle=45,
    xaxis_title="ETF",
    yaxis_title="Portfolio weight",
    yaxis_tickformat=".0%",
)
show_plotly_with_alt(
    fig,
    "Bars of HRP weight per ETF, sorted from largest to smallest, with one bond fund holding the majority of the portfolio and the rest falling away sharply.",
)

# %% [markdown]
# ### Where That Concentration Comes From
#
# HRP is introduced as the answer to the concentrated portfolios that mean-variance optimization
# produces, so a single dominant weight deserves an explanation rather than a shrug. It follows
# from the algorithm as specified.
#
# Recursive bisection splits the *ordered list* in half by count at every level. It does not cut
# the tree at the linkage's own cluster boundaries, so the halves it forms need not be the
# clusters the dendrogram shows. Each half then receives capital inversely to its cluster
# variance. Two splits are enough to see the effect.


# %%
def trace_bisection(cov: np.ndarray, sorted_idx: list[int], columns, levels: int = 2) -> None:
    """Print the halves, their annualized cluster variances, and the resulting split share."""
    clusters, depth = [sorted_idx], 0
    while clusters and depth < levels:
        next_clusters = []
        for cluster in clusters:
            if len(cluster) <= 1:
                continue
            mid = len(cluster) // 2
            left, right = cluster[:mid], cluster[mid:]
            left_var, right_var = cluster_variance(cov, left), cluster_variance(cov, right)
            alpha = 1 - left_var / (left_var + right_var)
            print(f"  split at depth {depth}:")
            print(
                f"    {[columns[i] for i in left]}\n"
                f"      annualized cluster variance {left_var * 252:.4f} -> {alpha:.1%} of the branch"
            )
            print(
                f"    {[columns[i] for i in right]}\n"
                f"      annualized cluster variance {right_var * 252:.4f} -> {1 - alpha:.1%} of the branch"
            )
            if len(left) > 1:
                next_clusters.append(left)
            if len(right) > 1:
                next_clusters.append(right)
        clusters = next_clusters
        depth += 1


print("First two levels of the recursive bisection:")
trace_bisection(returns.cov().values, get_quasi_diagonal_order(link), list(returns.columns))
print()
print(f"Weight in the single largest holding: {weights_df['HRP Weight'].max():.1%}")
print(f"Weight in the three largest: {weights_df['HRP Weight'].nlargest(3).sum():.1%}")

# %% [markdown]
# The low-variance half takes most of the capital at both splits, and the shares multiply. Bonds and gold
# end up holding most of the portfolio, and within them the lowest-volatility holding takes most
# of what is left.
#
# This is inverse-variance allocation doing exactly what it is defined to do, not a bug. But it
# means HRP is not concentration-free: it concentrates on the *low-variance* assets rather than on
# whichever assets the covariance inverse happens to favor. Whether that is an improvement depends
# on whether low realized variance in the estimation window is a better guide to the future than
# the inverse covariance is - a question the walk-forward comparison below can address and this
# static example cannot.

# %% [markdown]
# ## 8. Comparison: HRP vs Other Methods


# %%
def inverse_volatility_weights(returns: pd.DataFrame) -> np.ndarray:
    """Inverse volatility (risk parity) weights."""
    vols = returns.std().values
    inv_vols = 1 / vols
    return inv_vols / inv_vols.sum()


# %% [markdown]
# #### Equal-Weight Baseline


# %%
def equal_weights(n: int) -> np.ndarray:
    """Equal weights."""
    return np.ones(n) / n


# %% [markdown]
# #### Minimum-Variance Baseline
#
# The long-only projection starts from
# $w = \Sigma^{-1}\mathbf{1}/(\mathbf{1}^{\top}\Sigma^{-1}\mathbf{1})$,
# clips negative weights, and renormalizes the remaining capital.


# %%
def minimum_variance_weights(returns: pd.DataFrame) -> np.ndarray:
    """Minimum variance portfolio (no expected returns input)."""
    cov = returns.cov().values
    n = len(returns.columns)

    try:
        cov_inv = np.linalg.inv(cov)
        ones = np.ones(n)
        weights = cov_inv @ ones
        weights /= weights.sum()
        # Long-only projection: clip negatives, then renormalize so weights sum to 1
        weights = np.clip(weights, 0, None)
        total = weights.sum()
        return weights / total if total > 0 else equal_weights(n)
    except np.linalg.LinAlgError:
        return equal_weights(n)


# %% [markdown]
# #### Ledoit-Wolf Shrinkage Baseline


# %%
def mvo_shrinkage_weights(returns: pd.DataFrame) -> np.ndarray:
    """MVO with Ledoit-Wolf shrinkage."""
    # Use Ledoit-Wolf for robust covariance
    lw = LedoitWolf().fit(returns)
    cov_shrunk = lw.covariance_

    # Minimum variance with shrunk covariance
    n = len(returns.columns)
    try:
        cov_inv = np.linalg.inv(cov_shrunk)
        ones = np.ones(n)
        weights = cov_inv @ ones
        weights /= weights.sum()
        # Long-only projection: clip negatives, then renormalize so weights sum to 1
        weights = np.clip(weights, 0, None)
        total = weights.sum()
        return weights / total if total > 0 else equal_weights(n)
    except np.linalg.LinAlgError:
        return equal_weights(n)


# %%
# Compute all weights
n_assets = len(returns.columns)

all_weights = pd.DataFrame(
    {
        "Symbol": returns.columns,
        "Name": [ETF_UNIVERSE.get(s, s) for s in returns.columns],
        "Equal": equal_weights(n_assets),
        "Inv Vol": inverse_volatility_weights(returns),
        "Min Var (LW)": mvo_shrinkage_weights(returns),
        "HRP": hrp_weights,
    }
)

# %%
# Visualization: Weight comparison
fig = go.Figure()

# One colour per allocator, in the order the four are always listed. Equal weight is the
# benchmark, so it takes the neutral grey; HRP is the subject of the notebook and takes the
# primary colour. The same four are reused for the equity curves in section 9.
methods = ["Equal", "Inv Vol", "Min Var (LW)", "HRP"]
ALLOCATOR_COLORS = [COLORS["neutral"], COLORS["amber"], COLORS["copper"], COLORS["blue"]]

for method, color in zip(methods, ALLOCATOR_COLORS, strict=True):
    fig.add_trace(
        go.Bar(
            name=method,
            x=all_weights["Name"],
            y=all_weights[method],
            marker_color=color,
        )
    )

fig.update_layout(
    title="Portfolio weight per ETF under four allocators",
    barmode="group",
    xaxis_tickangle=45,
    xaxis_title="ETF",
    yaxis_title="Portfolio weight",
    yaxis_tickformat=".0%",
    height=500,
    legend=dict(orientation="h", yanchor="bottom", y=1.02),
)
show_plotly_with_alt(
    fig,
    "Grouped bars of portfolio weight per ETF for equal weight, inverse volatility, shrinkage "
    "minimum variance and HRP. The two variance-based allocators put the bulk of the "
    "portfolio in the aggregate bond fund, while the equal-weight bars sit at the same "
    "height across the universe.",
)

# %% [markdown]
# ## 9. Walk-Forward Backtest with ML Predictions
#
# The previous sections demonstrated *how* HRP allocates capital given a
# covariance matrix. We now apply it inside a realistic trading strategy:
#
# 1. Load the ETF case-study GBM walk-forward predictions (Chapter 12).
# 2. At each rebalance date, use the latest available prediction to *select* the
#    top-$N$ assets.
# 3. Apply HRP (and the other methods) to the selected subset using only
#    historical returns for covariance estimation.
#
# Asset selection is driven by a walk-forward validation signal, while the
# allocator only sees information available at the rebalance date. The target
# becomes effective on the following bar.

# %% [markdown]
# Asset selection uses the ETF case study's best-IC GBM walk-forward validation predictions. HRP
# itself needs only a covariance matrix, so the predictions are not part of the allocator; they
# decide *which* assets each allocator sizes, and every allocator sees the same selection.
#
# The highest-validation-IC GBM checkpoint is resolved at runtime from
# `case_studies/etfs/run_log/registry.db`. No `prediction_hash` is baked in, so re-running the
# GBM sweep changes which predictions feed the comparison without an edit here.
#
# The registry ranks and evaluates on validation, so what follows demonstrates allocation
# behavior on a sample the selection step has already seen. It is not an out-of-sample estimate.

# %%
etf_case_dir = get_case_study_dir("etfs")
etf_registry_path = etf_case_dir / "run_log" / "registry.db"
etf_registry_sha256 = hashlib.sha256(etf_registry_path.read_bytes()).hexdigest()
prediction_index = load_prediction_index(
    "etfs",
    label="fwd_ret_21d",
    split="validation",
    family="gbm",
    case_dir=etf_case_dir,
)
if prediction_index.is_empty():
    raise RuntimeError("No registered ETF GBM validation predictions are available")
best_ic = prediction_index["ic_mean"][0]
if best_ic is None or not np.isfinite(best_ic):
    raise RuntimeError("The leading ETF GBM validation prediction has no finite IC")
leaders = prediction_index.filter((pl.col("ic_mean") - best_ic).abs() <= 1e-12)
if leaders.height != 1:
    hashes = leaders["prediction_hash"].to_list()
    raise RuntimeError(f"Ambiguous best-IC ETF GBM predictions: {hashes}")
best_gbm = leaders.row(0, named=True)
ETF_GBM_PRED_HASH = best_gbm["prediction_hash"]

# %% [markdown]
# Validate the selected prediction parquet and record immutable input hashes.

# %%
PRED_PATH = etf_case_dir / "run_log" / "predictions" / ETF_GBM_PRED_HASH / "predictions.parquet"
if not PRED_PATH.exists():
    raise FileNotFoundError(
        f"Resolved best-IC GBM hash {ETF_GBM_PRED_HASH} (config={best_gbm['config_name']}) "
        f"has no predictions parquet at {PRED_PATH}. Re-run case_studies/etfs/07_gbm.py."
    )
print(
    f"Resolved best-IC GBM: hash={ETF_GBM_PRED_HASH}, config={best_gbm['config_name']}, "
    f"IC={best_gbm['ic_mean']:.4f}"
)
print(f"ETF registry SHA-256: {etf_registry_sha256}")
print(f"Prediction parquet SHA-256: {hashlib.sha256(PRED_PATH.read_bytes()).hexdigest()}")

upstream_preds = pl.read_parquet(PRED_PATH).select("timestamp", "symbol", "prediction")
print(
    f"Loaded ETF GBM predictions: {upstream_preds.height:,} rows, "
    f"{upstream_preds['symbol'].n_unique()} symbols, "
    f"{upstream_preds['timestamp'].min()} to {upstream_preds['timestamp'].max()}"
)


# %% [markdown]
# ### Asset Selection Rule
#
# At each rebalance date we take the most recent prediction available for each
# symbol (as-of join), restrict the ranking to the fixed teaching universe, and
# select the top $N$. If no prediction is available
# yet for the date (e.g., before the first walk-forward fold), the rebalance
# is skipped by returning an empty list.


# %%
def select_top_assets(
    date: pd.Timestamp,
    full_returns: pd.DataFrame,
    predictions: pl.DataFrame,
    top_n: int,
) -> list[str]:
    """Select top-N assets by latest available GBM prediction at `date`.

    Polars `group_by` does not guarantee row order, and ties at the top-N
    cutoff would otherwise produce non-deterministic selections across runs.
    We secondary-sort by symbol so the selection is reproducible.
    """
    as_of = predictions.filter(pl.col("timestamp") <= date)
    if as_of.is_empty():
        return []
    latest = (
        as_of.sort("timestamp")
        .group_by("symbol")
        .agg(pl.col("prediction").last())
        .filter(pl.col("symbol").is_in(full_returns.columns.to_list()))
        .filter(pl.col("prediction").is_finite())
        .sort(["prediction", "symbol"], descending=[True, False])
    )
    return latest.head(top_n)["symbol"].to_list()


# %% [markdown]
# ### Guarding against a covariance the allocator cannot use
#
# Validate each allocator's long-only weights before mapping them to the full universe. A singular
# covariance or invalid weight vector triggers a visible equal-weight fallback and increments the
# method's fallback counter.


# %%
def allocate_selected_assets(
    full_columns: pd.Index,
    top_assets: list[str],
    selected_returns: pd.DataFrame,
    allocation_fn,
    method: str = "allocator",
    date: pd.Timestamp | None = None,
) -> pd.Series:
    """Map valid selected weights to the full universe, with a reported equal-weight fallback."""
    weights = pd.Series(0.0, index=full_columns)
    try:
        alloc_weights = np.asarray(allocation_fn(selected_returns), dtype=float)
        if alloc_weights.shape != (len(top_assets),):
            raise ValueError(
                f"allocator returned shape {alloc_weights.shape}, expected {(len(top_assets),)}"
            )
        if not np.isfinite(alloc_weights).all() or (alloc_weights < 0).any():
            raise ValueError("allocator returned non-finite or negative weights")
        total = alloc_weights.sum()
        if total <= 0:
            raise ValueError("allocator returned non-positive total weight")
        alloc_weights = alloc_weights / total
        for asset, weight in zip(top_assets, alloc_weights, strict=False):
            weights[asset] = weight
    except (np.linalg.LinAlgError, ValueError) as exc:
        warnings.warn(
            f"{method} singular/invalid for {len(top_assets)} assets at {date}: {exc}",
            RuntimeWarning,
            stacklevel=2,
        )
        fallback_count[method] = fallback_count.get(method, 0) + 1
        for asset in top_assets:
            weights[asset] = 1.0 / len(top_assets)
    return weights


# %% [markdown]
# ### Rebalance Decision
#
# Build a target from trailing returns and the latest registered prediction. Returning
# `None` leaves the prior target unchanged and prevents a pre-signal baseline from entering results.


# %%
def rebalance_target(
    returns: pd.DataFrame,
    predictions: pl.DataFrame,
    date: pd.Timestamp,
    allocation_fn,
    top_n: int,
    lookback: int,
    method: str,
    min_history: int,
) -> pd.Series | None:
    """Return the target decided at `date`, or None when history/signal is unavailable."""
    loc = returns.index.get_loc(date)
    hist_returns = returns.iloc[max(0, loc - lookback) : loc]
    if len(hist_returns) < min_history:
        return None
    top_assets = select_top_assets(date, returns, predictions, top_n)
    if len(top_assets) < 2:
        return None
    return allocate_selected_assets(
        returns.columns,
        top_assets,
        hist_returns[top_assets],
        allocation_fn,
        method=method,
        date=date,
    )


# %% [markdown]
# ### Walk-Forward Backtest Engine


# %%
def walk_forward_backtest(
    returns: pd.DataFrame,
    allocation_fn,
    predictions: pl.DataFrame,
    lookback: int,
    rebalance_freq: str,
    top_n: int,
    min_history: int,
    method: str = "allocator",
) -> tuple[pd.DataFrame, pd.DataFrame]:
    """Hold positions between month-end decisions and let weights drift with returns."""
    rebalance_dates = set(returns.groupby(returns.index.to_period(rebalance_freq)).tail(1).index)
    portfolio_returns, target_history = [], []
    current_weights = pending_target = None
    for date in returns.index:
        # A target decided at the prior close becomes the next session's opening allocation.
        if pending_target is not None:
            current_weights = pending_target
            pending_target = None
        if current_weights is not None:
            asset_returns = returns.loc[date]
            daily_ret = float((asset_returns * current_weights).sum())
            portfolio_returns.append({"date": date, "return": daily_ret})
            end_values = current_weights * (1.0 + asset_returns)
            current_weights = end_values / end_values.sum()
        if date in rebalance_dates:
            target = rebalance_target(
                returns, predictions, date, allocation_fn, top_n, lookback, method, min_history
            )
            if target is not None:
                pending_target = target
                target_history.append(
                    {"date": date, **{f"w_{s}": target[s] for s in returns.columns}}
                )

    return (
        pd.DataFrame(portfolio_returns).set_index("date"),
        pd.DataFrame(target_history).set_index("date"),
    )


# %% [markdown]
# Four settings decide what the backtest does, and every allocator gets the same four.
# `LOOKBACK` is how much trailing history each covariance is estimated on - one year of
# sessions, which with five selected names leaves the sample covariance comfortably
# over-determined. `TOP_N` is how many of the fifteen the prediction picks each month, and it is
# the number that decides whether the ratio of assets to observations is anywhere near the
# regime HRP is designed for. `REBALANCE_FREQ` is month-end, so the targets are decided twelve
# times a year and held in between. `MIN_HISTORY` is the shortest trailing window an allocator
# is allowed to size from at all; below it the rebalance is skipped rather than estimated on too
# few rows.

# %%
LOOKBACK = 252
TOP_N = 5
REBALANCE_FREQ = "M"
MIN_HISTORY = 60

allocation_methods = {
    "Equal Weight": lambda r: equal_weights(len(r.columns)),
    "Inverse Volatility": inverse_volatility_weights,
    "Min Variance (LW)": mvo_shrinkage_weights,
    "HRP": hrp_portfolio,
}

results = {}
weights_all = {}

print("Running walk-forward backtests...")
for name, alloc_fn in allocation_methods.items():
    print(f"  {name}...")
    ret_df, w_df = walk_forward_backtest(
        returns=returns,
        allocation_fn=alloc_fn,
        predictions=upstream_preds,
        lookback=LOOKBACK,
        rebalance_freq=REBALANCE_FREQ,
        top_n=TOP_N,
        method=name,
        min_history=MIN_HISTORY,
    )
    results[name] = ret_df["return"]
    weights_all[name] = w_df

print("All backtests complete")

# Combine into DataFrame
portfolio_returns = pd.DataFrame(results)


# %% [markdown]
# ### Gross Backtest Metrics
#
# The vectorized comparison isolates allocation effects and is gross of implementation costs.
# Turnover is reported beside performance. The execution-aware bridge below applies its declared
# commission and slippage assumptions and uses next-bar fills.


# %%
def _weights_wide_to_long(weights_df: pd.DataFrame) -> pl.DataFrame:
    """Convert dense target weights to canonical timestamp/symbol/weight format."""
    wide = weights_df.reset_index()
    ts_col = wide.columns[0]
    wide = wide.rename(columns={ts_col: "timestamp"})
    long = wide.melt(id_vars=["timestamp"], var_name="symbol", value_name="weight")
    long["symbol"] = long["symbol"].str.replace("w_", "", regex=False)
    return pl.from_pandas(long[["timestamp", "symbol", "weight"]]).sort(["timestamp", "symbol"])


# %% [markdown]
# Compute allocator metrics for each method. The dense month-end target state preserves both new
# positions and explicit zero-weight liquidations before the per-symbol turnover difference.

# %%
allocator_metrics: dict[str, dict] = {}
metrics_list = []
for name in portfolio_returns.columns:
    returns_arr = portfolio_returns[name].to_numpy()
    monthly_targets = weights_all[name]
    weights_long = _weights_wide_to_long(monthly_targets)
    m = compute_allocator_metrics(
        pl.Series("returns", returns_arr),
        weights_df=weights_long,
        ann_factor=np.sqrt(252),
    )
    allocator_metrics[name] = m
    finite = returns_arr[np.isfinite(returns_arr)]
    annual_vol = float(np.nanstd(finite, ddof=1) * np.sqrt(252)) if finite.size > 1 else 0.0
    metrics_list.append(
        {
            "Method": name,
            "Annual Return": m["annual_return"],
            "Annual Vol": annual_vol,
            "Sharpe Ratio": m["sharpe"],
            "Max Drawdown": m["max_drawdown"],
            "Calmar Ratio": m["calmar"],
            "Avg Turnover": m["avg_turnover"],
        }
    )

# %%
metrics_df = pd.DataFrame(metrics_list)
metrics_df = metrics_df[
    [
        "Method",
        "Annual Return",
        "Annual Vol",
        "Sharpe Ratio",
        "Max Drawdown",
        "Calmar Ratio",
        "Avg Turnover",
    ]
]
metrics_df.round(4)

# %% [markdown]
# ### Scheduled Weight Strategy
#
# This adapter converts the notebook's target-weight schedule into the
# event-driven strategy interface expected by `ml4t-backtest`.


# %%
class ScheduledWeightStrategy(Strategy):
    def __init__(self, weights_long: pl.DataFrame, allow_short: bool):
        self.executor = TargetWeightExecutor(
            config=RebalanceConfig(
                min_trade_value=100.0,
                min_weight_change=0.001,
                allow_fractional=True,
                allow_short=allow_short,
            )
        )
        self._targets_by_ts: dict[pd.Timestamp, dict[str, float]] = {}
        for row in weights_long.iter_rows(named=True):
            ts = pd.Timestamp(row["timestamp"]).tz_localize(None)
            self._targets_by_ts.setdefault(ts, {})
            self._targets_by_ts[ts][str(row["symbol"])] = float(row["weight"])

    def on_data(self, timestamp, data, context, broker):
        ts = pd.Timestamp(timestamp).tz_localize(None)
        targets = self._targets_by_ts.get(ts)
        if not targets:
            return
        targets = {asset: weight for asset, weight in targets.items() if asset in data}
        if targets:
            self.executor.execute(targets, data, broker)


# %% [markdown]
# ### Restating the schedule in the form the engine reads
#
# Convert the wide monthly target schedule into the long-form price and target tables required by
# the execution engine. Zero targets remain present so liquidations are explicit.

# %%
bridge_method = "HRP" if "HRP" in weights_all else next(iter(weights_all))
weights_long = _weights_wide_to_long(weights_all[bridge_method]).with_columns(
    pl.col("timestamp").cast(pl.Datetime("us"))
)
allow_short_engine = (
    weights_long.filter(pl.col("weight") < 0).height > 0 if not weights_long.is_empty() else False
)

prices_panel = pl.from_pandas(close_prices.reset_index())
ts_col = prices_panel.columns[0]
if ts_col != "timestamp":
    prices_panel = prices_panel.rename({ts_col: "timestamp"})
prices_long = (
    prices_panel.unpivot(index="timestamp", variable_name="symbol", value_name="close")
    .with_columns(
        [
            pl.col("timestamp").cast(pl.Datetime("us")),
            pl.col("close").alias("open"),
            pl.col("close").alias("high"),
            pl.col("close").alias("low"),
            pl.lit(1_000_000).alias("volume"),
        ]
    )
    .sort(["timestamp", "symbol"])
)

# %% [markdown]
# ### Run the Execution-Aware Backtest
#
# Replay the same target weights with commissions, slippage, and next-bar fills. Synthetic
# OHLC fields equal the daily close, so this is an accounting and timing bridge rather than an
# intraday fill-quality model.

# %%
engine = Engine(
    feed=DataFeed(prices_df=prices_long),
    strategy=ScheduledWeightStrategy(weights_long, allow_short=allow_short_engine),
    config=BacktestConfig(
        initial_cash=100_000.0,
        execution_mode=ExecutionMode.NEXT_BAR,
        commission_type=CommissionType.PERCENTAGE,
        commission_rate=0.0005,
        slippage_type=SlippageType.PERCENTAGE,
        slippage_rate=0.0005,
        allow_short_selling=allow_short_engine,
    ),
)

# %% [markdown]
# ### Compare Engine and Vectorized Returns
#
# The bridge checks whether the event-driven backtest preserves the ranking and
# risk shape implied by the vectorized walk-forward simulation.

# %%
engine_daily = (
    engine.run()
    .to_daily_pnl()
    .select(
        pl.col("date").cast(pl.Datetime("us")).alias("date"),
        pl.col("return_pct").alias("engine_return"),
    )
)
vectorized_daily = pl.DataFrame(
    {
        "date": pl.Series(portfolio_returns.index.to_list()).cast(pl.Datetime("us")),
        "vectorized_return": portfolio_returns[bridge_method].to_numpy(),
    }
)

bridge = (
    vectorized_daily.join(engine_daily, on="date", how="inner")
    .drop_nulls(["vectorized_return", "engine_return"])
    .sort("date")
)

vec_summary = PortfolioAnalysis(returns=bridge["vectorized_return"], periods_per_year=252)
eng_summary = PortfolioAnalysis(returns=bridge["engine_return"], periods_per_year=252)
vec_stats = vec_summary.compute_summary_stats()
eng_stats = eng_summary.compute_summary_stats()

print(f"\nExecution bridge ({bridge_method} walk-forward):")
print(
    f"  Vectorized Sharpe={vec_stats.sharpe_ratio:.3f}, Engine Sharpe={eng_stats.sharpe_ratio:.3f}"
)
print(f"  Vectorized MaxDD={vec_stats.max_drawdown:.2%}, Engine MaxDD={eng_stats.max_drawdown:.2%}")

# %%
# Equity curves
cumulative = (1 + portfolio_returns).cumprod()

fig = go.Figure()

for col, color in zip(cumulative.columns, ALLOCATOR_COLORS, strict=True):
    fig.add_trace(
        go.Scatter(
            x=cumulative.index,
            y=cumulative[col],
            mode="lines",
            name=col,
            line=dict(color=color, width=2 if col == "HRP" else 1.5),
        )
    )

fig.add_hline(y=1.0, line_dash="dot", line_color=COLORS["neutral"])

fig.update_layout(
    title="Four allocators over the same monthly selection, gross of costs",
    xaxis_title="Date",
    yaxis_title="Growth of $1",
    height=500,
)
show_plotly_with_alt(
    fig,
    "Four growth-of-one-dollar paths from the walk-forward backtest, one per allocation "
    "method. They track each other until about 2018 and then separate into two pairs, equal "
    "weight and inverse volatility above, shrinkage minimum variance and HRP below, each "
    "falling sharply in early 2020.",
)

# %% [markdown]
# ## 10. The HRP series on its own terms
#
# The table above compares allocators to each other. The dashboard below looks at one of them -
# HRP - the way an investor would: cumulative growth against SPY, the underwater curve, rolling
# Sharpe, and the monthly return grid. Its metrics come from the same return series as the table,
# so the two agree to rounding; what it adds is the path behind the summary statistics.

# %%
hrp_returns_series = portfolio_returns["HRP"].dropna()
spy_returns_series = returns["SPY"].reindex(hrp_returns_series.index).dropna()
common_idx = hrp_returns_series.index.intersection(spy_returns_series.index)

hrp_analysis = PortfolioAnalysis(
    returns=hrp_returns_series.loc[common_idx].to_numpy(),
    benchmark=spy_returns_series.loc[common_idx].to_numpy(),
    dates=common_idx,
    risk_free=0.0,
    periods_per_year=252,
)

hrp_tear_sheet = create_portfolio_dashboard(hrp_analysis)

for dashboard_figure in hrp_tear_sheet.figures.values():
    dashboard_figure.update_layout(
        paper_bgcolor=COLORS["bg_light"],
        plot_bgcolor=COLORS["bg_light"],
        font_color=COLORS["neutral"],
    )

rolling_beta_figure = hrp_tear_sheet.figures["Rolling Beta"]
rolling_beta_figure.update_layout(margin=dict(l=60, r=90, t=40, b=40))
_ = rolling_beta_figure.update_annotations(x=0.995, xanchor="right")

# %% [markdown]
# `hrp_tear_sheet.show()` would render the metrics block and then loop `fig.show()` over the
# nine figures, which publishes each PNG with no alt text and leaves a screen reader with
# nothing. Displaying the same content a figure at a time is what lets each one carry a
# description of what it plots.

# %%
DASHBOARD_ALT = {
    "Cumulative Returns": (
        "Cumulative return of the HRP portfolio and of the SPY benchmark against date, the "
        "benchmark drawn dashed and finishing well above the portfolio."
    ),
    "Drawdown": (
        "The HRP portfolio's underwater curve against date, filled to zero, showing the "
        "percentage below its own running peak, with the deepest point marked in March 2020."
    ),
    "Rolling Sharpe Ratio": (
        "Two lines of rolling Sharpe ratio against date, over sixty-three and two hundred and "
        "fifty-two sessions, the shorter window swinging more widely than the longer one."
    ),
    "Rolling Volatility": (
        "Three lines of annualized rolling volatility against date, over twenty-one, "
        "sixty-three and two hundred and fifty-two sessions."
    ),
    "Rolling Beta": (
        "Rolling beta of the HRP portfolio against SPY, plotted against date and shaded down "
        "to zero, with a dashed reference line at the market's own beta."
    ),
    "Annual Returns": (
        "Bars of the HRP portfolio's annual return by calendar year, with the benchmark's "
        "annual return marked as points and a line at zero."
    ),
    "Monthly Returns Heatmap": (
        "Heatmap of monthly return, years down the vertical axis and calendar months across "
        "with a compounded annual column at the right, each cell labelled with its return and "
        "coloured from red for losses to green for gains."
    ),
    "Returns Distribution": (
        "Histogram of the HRP portfolio's daily returns with a fitted normal density drawn "
        "over it and vertical reference lines in the left tail."
    ),
    "Top Drawdowns": (
        "Horizontal bars of drawdown depth for the five deepest episodes, one bar per episode, "
        "ordered deepest at the top and annotated with the depth reached."
    ),
}

display(HTML(f"<pre>{hrp_tear_sheet.metrics.summary()}</pre>"))
for figure_name, dashboard_figure in hrp_tear_sheet.figures.items():
    show_plotly_with_alt(dashboard_figure, DASHBOARD_ALT[figure_name])

# %%
# HTML delivery: the same content as a single self-contained file.
output_dir = get_output_dir(17, "hrp")
output_dir.mkdir(parents=True, exist_ok=True)
hrp_tear_sheet_path = output_dir / "hrp_tear_sheet.html"
hrp_tear_sheet.save_html(hrp_tear_sheet_path, include_plotlyjs="cdn")
print(f"HRP tear sheet saved: {output_dir.name}/{hrp_tear_sheet_path.name}")
print(f"  Figures embedded: {list(hrp_tear_sheet.figures.keys())}")

# %% [markdown]
# ## 11. Weight Evolution in Strategy Context
#
# Analyze how the submitted HRP targets change over the walk-forward backtest. The portfolio holds
# each allocation until the next month-end decision, so realized weights drift between targets.

# %%
# Plot HRP weight evolution from actual backtest
hrp_weights_hist = weights_all.get("HRP", pd.DataFrame())

if not hrp_weights_hist.empty:
    # Get weight columns
    weight_cols = [c for c in hrp_weights_hist.columns if c.startswith("w_")]

    fig = go.Figure()

    # Show top assets by average weight
    avg_weights = hrp_weights_hist[weight_cols].mean().sort_values(ascending=False)
    top_cols = avg_weights.head(6).index.tolist()

    for col in top_cols:
        symbol = col.replace("w_", "")
        fig.add_trace(
            go.Scatter(
                x=hrp_weights_hist.index,
                y=hrp_weights_hist[col],
                mode="lines",
                name=ETF_UNIVERSE.get(symbol, symbol),
                line_shape="hv",
            )
        )

    fig.update_layout(
        title="Monthly HRP target weight for the six largest average holdings",
        xaxis_title="Date",
        yaxis_title="Weight",
        height=450,
    )
    show_plotly_with_alt(
        fig,
        "Step lines of the six largest average HRP target weights against date, each holding flat between month-end decisions and jumping when the selected set changes.",
    )
else:
    print("No HRP weights available")

# %% [markdown]
# The step lines show submitted targets only. Actual portfolio weights drift with asset returns
# between month-end decisions; neither the vectorized path nor the execution engine resets them
# during the month.

# %% [markdown]
# ## 12. Reading the Result
#
# The comparison above is the notebook's evidence. What follows reads it rather than restating
# what HRP is supposed to achieve, and it reads two columns rather than one: the Sharpe ratio the
# allocators are ranked on, and the turnover each of them needed to get there.

# %%
ranked = metrics_df.sort_values("Sharpe Ratio", ascending=False).reset_index(drop=True)
ranked.index = pd.RangeIndex(1, len(ranked) + 1, name="Rank by Sharpe")

print(f"Assets in the teaching universe: {n_assets}")
print(f"Estimation window each allocator sizes from: {LOOKBACK} sessions")
print(f"Assets selected each month: {TOP_N}")
ranked[["Method", "Sharpe Ratio", "Annual Vol", "Avg Turnover"]].round(4)

# %% [markdown] tags=["results"]
# Read the ranking against what each allocator had to estimate, because that is the axis the four
# differ on. Equal weight estimates nothing. Inverse volatility estimates one variance per asset.
# HRP reads the correlations to build the tree and order the assets, and each split then compares
# its two halves on their cluster variances, which `cluster_variance` computes as
# $w^\top \Sigma w$ over the block - so the off-diagonal entries reach the weights. Minimum
# variance with Ledoit-Wolf shrinkage reads and inverts the whole matrix.
#
# Whether estimating more pays depends on how good the estimate is, and the ratio that decides
# that is not favourable to the elaborate methods here. The argument for HRP is that inverting a
# noisy covariance matrix amplifies estimation error, and it has force when the estimate is badly
# under-determined - when the number of assets approaches or exceeds the number of observations.
# This comparison is the opposite case: five selected assets estimated over 252 daily
# observations, where the sample covariance is well conditioned and its inverse is not dominated
# by noise. The regime HRP was designed for is not the regime tested here, so wherever it lands
# in this table, the table is not evidence about that regime.
#
# What HRP never does is invert the matrix, and that is the property it is chosen for: it returns
# weights whatever the conditioning, where inversion has no unique solution at all once the
# assets outnumber the observations.
#
# The turnover column carries a second reading, and it is the one that survives a different
# sample. Clustering is often described as producing more stable allocations, and the column is
# where that claim would be tested - but it measures two things at once. Every row pays for the
# monthly re-selection of five names out of fifteen, and every row then pays for whatever its
# own sizing rule does with the names it keeps. The equal-weight row shows what the schedule
# costs an allocator that makes no sizing decision at all, which is a reference point and not a
# quantity the other three sit above: what a replacement costs depends on the weights being
# replaced, so a concentrated allocator's re-selection cost is a different number rather than
# the equal-weight one plus a margin.
#
# ### What is true regardless of the ranking
#
# These are properties of the algorithm, and they hold wherever it lands in a given sample:
#
# - It never inverts a covariance matrix, so it returns weights for any $N$ and $T$, including
#   $N > T$ where minimum-variance optimization has no unique solution at all.
# - It needs no expected-return forecast. Every input is second-moment.
# - Its weights are determined by the correlation hierarchy and the variances, so they can be
#   traced back to a specific split, as section 7 traced them.
#
# ### What this notebook cannot tell you
#
# One universe, one sample, one estimation window, and a selection step evaluated on validation
# data. A ranking under those conditions is a fact about this run. The regime where HRP was
# designed to help - many assets relative to observations - is not the regime tested here, so this
# result is not evidence against HRP in that regime either.


# %% [markdown]
# ### How often an allocator had to fall back
#
# When the allocation function raises (e.g., singular covariance), the
# walk-forward backtest falls back to equal weights and increments a counter.

# %%
print(f"Assets in universe: {n_assets}")
print(
    f"Invested backtest period: {portfolio_returns.index[0].date()} "
    f"to {portfolio_returns.index[-1].date()}"
)
print(f"Asset selection: ETF GBM walk-forward predictions (hash {ETF_GBM_PRED_HASH})")
if fallback_count:
    print("Allocator fallbacks (singular/invalid covariance):")
    for method_name, count in sorted(fallback_count.items()):
        print(f"  {method_name}: {count}")
else:
    print("Allocator fallbacks: none triggered")

# %% [markdown]
# ## Key Takeaways
#
# 1. **HRP replaces inversion with ordering and recursive splits.** It still depends on estimated
#    correlations and variances, so it reduces the exposure to estimation error rather than
#    removing it.
# 2. **Avoiding inversion pays off when the estimate is under-determined.** With five assets and
#    252 observations it is not, so nothing in this comparison exercises the property HRP is
#    chosen for. The case for it is strongest when the number of assets approaches or exceeds the
#    sample length, and that regime is not tested here.
# 3. **HRP concentrates too, on the low-variance assets.** Section 7 traces a majority weight in a
#    single bond fund to two successive inverse-variance splits. Risk-based is not the same thing
#    as diversified.
# 4. **Turnover measures re-selection and sizing together, and no column separates them.** Every
#    allocator pays for replacing names the monthly signal drops, and then pays again for
#    whatever its own sizing rule does with the names it keeps. Equal weight shows what the
#    schedule costs an allocator that makes no sizing decision, which is a reference point and
#    not a quantity the others sit above: what a replacement costs depends on the weights being
#    replaced. So HRP's turnover against equal weight's is suggestive about stability and is not
#    a measurement of it.
# 5. **The comparison is conditional in three ways.** A fixed 15-ETF ex-post universe, gross of
#    costs in the vectorized path, and a selection signal ranked on validation data. It shows
#    allocation behavior, not an out-of-sample estimate.
# 6. **Timing and costs need separate evidence.** Targets use returns through each decision date
#    and become effective on the next bar; the execution bridge shows what commissions and
#    slippage do to the vectorized result.
#
# **Next**: [`07_conformal_position_sizing`](07_conformal_position_sizing.ipynb) sizes positions
# from the width of a prediction interval rather than from a covariance matrix.

```

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。