Usar anchuras de predicción conformal para dimensionar posiciones en ETF y futuros
Resumen
Este cuaderno compara asignaciones con pesos iguales, ponderadas por puntuación positiva y por anchura conformal para los ETF y futuros CME mejor clasificados. Estima anchuras de intervalos de conformidad por partición específicas de cada entidad a partir de residuos de particiones de validación anteriores, excluyendo observaciones cuyas etiquetas de rendimientos futuros aún no se habrían conocido. Dentro de cada cartera seleccionada, los pesos son proporcionales al inverso de la anchura del intervalo, tratando los rangos históricos de residuos más estrechos como indicador de confianza.
La comparación presenta medidas de cartera como Sharpe, rendimiento, drawdown y rotación, junto con diagnósticos del coeficiente de información de la señal predictiva. Todos los métodos comparten las mismas selecciones top-K, por lo que las diferencias se atribuyen a los pesos de asignación. Sin embargo, los conjuntos de predicciones GBM se seleccionaron mediante el IC de validación del mismo registro, así que estos resultados son evidencia de validación condicionada por la selección, no una muestra de prueba imparcial. La cobertura conformal no garantiza estabilidad futura de los residuos, y dimensionar posiciones según el inverso de la anchura solo ayuda si la dispersión pasada de los residuos sigue siendo informativa. El análisis no incluye costes de trading, por lo que la rotación y las diferencias aparentes de rendimiento no demuestran un beneficio neto.
Ideas clave
- Las anchuras conformales específicas de cada entidad se estiman usando solo residuos anteriores, con embargo del horizonte.
- La asignación basada en el inverso de la anchura da más peso a entidades con rangos históricos de residuos más estrechos.
- Todas las reglas de dimensionamiento usan las mismas selecciones top-K para aislar el efecto de los pesos.
- El dimensionamiento basado en la incertidumbre no puede generar una dirección predictiva si la señal subyacente es débil.
- La selección de modelos basada en validación y la omisión de costes de trading limitan las conclusiones sobre el rendimiento fuera de muestra.
Etiquetas
Texto completo
# 07_conformal_position_sizing.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Conformal Prediction Position Sizing: Two Case Studies
#
# **Docker image**: `ml4t`
#
# This notebook builds uncertainty-based position sizing on top of registered
# walk-forward validation predictions for ETFs (`fwd_ret_21d`) and CME futures
# (`fwd_ret_5d`), then
# compares three allocation rules at the same top-K selection: equal weight,
# conformal-width-weighted ($w_i \propto 1/\Delta_i$), and score-weighted
# ($w_i \propto \hat y_i$ for positive scores).
#
# **Learning Objectives**:
# - Compute per-entity Mondrian split-conformal interval widths from strictly
# prior, horizon-embargoed validation residuals
# - Translate widths into position sizes via the inverse-width transform
# - Compare uncertainty-weighted, score-weighted, and equal-weighted top-K
# allocations on the same prediction panel
# - Identify when uncertainty-based sizing helps versus hurts
#
# **Book Reference**: Chapter 17, Section 17.4 (Defining baseline allocators)
#
# **Prerequisites**: `02_mean_variance_optimization`; conformal prediction from
# Chapter 11 (`06_conformal_prediction`); registered GBM predictions for the
# ETFs and CME futures case studies.
# %%
"""Compare conformal, score, and equal-weight sizing on registered GBM predictions."""
import sqlite3
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from ml4t.diagnostic.evaluation.portfolio_analysis import annual_return, max_drawdown
from ml4t.diagnostic.metrics import sharpe_ratio, sortino_ratio
from ml4t.diagnostic.metrics.ic_inference import compute_ic_hac_stats
from ml4t.diagnostic.signal.signal_ic import extract_signal_ic_series
from utils.paths import get_case_study_dir, registry_readonly_uri
from utils.style import COLORS, FIGSIZE, add_message_title, ml4t_palette, show_with_alt, zero_line
# %% tags=["parameters"]
# Production defaults - Papermill overrides for CI testing
CONFORMAL_ALPHA = 0.20 # 80% prediction interval
TOP_K_ETF = 20 # number of ETFs in the portfolio
TOP_K_CME = 10 # number of CME products in the portfolio
HORIZON_ETF = 21 # fwd_ret_21d → 21 trading days between non-overlapping rebalances
HORIZON_CME = 5 # fwd_ret_5d → 5 trading days
TRADING_DAYS_PER_YEAR = 252
# %% [markdown]
# ## 1. The Conformal-to-Confidence Transform
#
# Conformal prediction produces prediction intervals
# $[\hat{y}_i - q, \hat{y}_i + q]$ where the width $2q$ reflects model uncertainty.
# Here the calibration sample is strictly prior and horizon-embargoed. Narrower
# intervals indicate lower historical residual dispersion for that entity and model,
# which we use as a confidence proxy rather than a guarantee of future coverage.
#
# The confidence-weighted position for entity $i$ at time $t$ within a top-$K$
# selection is:
#
# $$w_{i,t}^{conf} = \frac{1 / \Delta_{i,t}}{\sum_{j \in \text{top-}K} 1 / \Delta_{j,t}}$$
#
# where $\Delta_{i,t}$ is the per-entity interval width. To produce cross-sectional
# variation in $\Delta_i$, we use **Mondrian split-conformal calibration**:
# for each fold $k$, the conformal quantile $q_i$ is computed separately for each
# entity from residuals whose forward-return labels were observable before fold $k$
# began. Entities with historically wider residuals get larger $\Delta_i$ and
# therefore smaller weight.
# %% [markdown]
# ## 2. Load Registered Predictions
#
# For each case study, the GBM validation prediction set with the highest recorded IC is loaded
# from its registry. During publication verification, `ML4T_OUTPUT_DIR` points to the
# immutable teaching-registry overlay; readers can omit it to use their local run logs.
# %% [markdown]
# Both panels name their entity column `symbol`. That is a property of the prediction artifact
# rather than of the case study: CME futures call the entity `product` in their labels and
# features, and the training step writes the same product roots (`ES`, `CL`, `6E`, and so on)
# under `symbol` when it records a prediction set.
#
# Selection is on the validation IC recorded in the registry, so everything measured below is
# conditioned on that selection and is not a holdout estimate.
# %%
REGISTRY_ROOTS = {
"etfs": get_case_study_dir("etfs", create=False) / "run_log",
"cme_futures": get_case_study_dir("cme_futures", create=False) / "run_log",
}
BEST_GBM = {
"etfs": {"label": "fwd_ret_21d", "id_col": "symbol"},
"cme_futures": {"label": "fwd_ret_5d", "id_col": "symbol"},
}
# %% [markdown]
# Registry reads are opened read-only, which is what stops this notebook writing to a
# registry the case studies own. The query excludes prediction sets with a degenerate
# fold before ranking by the daily-pooled IC.
# %%
def resolve_best_prediction(case_study: str, label: str) -> dict[str, str | float]:
"""Resolve the top validation GBM prediction set without writing to the registry."""
registry = REGISTRY_ROOTS[case_study] / "registry.db"
connection = sqlite3.connect(registry_readonly_uri(registry), uri=True)
try:
metric_columns = {
row[1] for row in connection.execute("PRAGMA table_info(prediction_metrics)")
}
ic_expr = (
"COALESCE(pm.ic_mean_daily, pm.ic_mean)"
if "ic_mean_daily" in metric_columns
else "pm.ic_mean"
)
row = connection.execute(
f"""
SELECT p.prediction_hash, t.config_name, {ic_expr} AS ic_mean
FROM training_runs t
JOIN prediction_sets p ON t.training_hash = p.training_hash
JOIN prediction_metrics pm ON p.prediction_hash = pm.prediction_hash
WHERE t.family = 'gbm' AND t.label = ? AND p.split = 'validation'
AND {ic_expr} IS NOT NULL
AND p.prediction_hash NOT IN
(SELECT prediction_hash FROM fold_metrics WHERE ic IS NULL)
ORDER BY ic_mean DESC LIMIT 1
""",
(label,),
).fetchone()
finally:
connection.close()
if row is None:
raise RuntimeError(f"No eligible GBM validation predictions for {case_study}/{label}")
return {"prediction_hash": row[0], "config_name": row[1], "ic_mean": float(row[2])}
# %%
for _case_study, _config in BEST_GBM.items():
_best = resolve_best_prediction(_case_study, _config["label"])
_config.update(_best)
print(
f"{_case_study}: validation GBM {_best['config_name']} "
f"({_best['prediction_hash']}), selection IC={_best['ic_mean']:.4f}"
)
# %%
def load_predictions(case_study: str) -> pl.DataFrame:
"""Load one canonical validation prediction panel and reject malformed keys."""
cfg = BEST_GBM[case_study]
path = (
REGISTRY_ROOTS[case_study] / "predictions" / cfg["prediction_hash"] / "predictions.parquet"
)
df = pl.read_parquet(path)
id_col = cfg["id_col"]
required = ["timestamp", id_col, "fold", "prediction", "actual"]
missing = set(required) - set(df.columns)
if missing or df.select(required).null_count().sum_horizontal().item() > 0:
raise ValueError(f"Malformed {case_study} prediction panel: missing={sorted(missing)}")
if df.select(id_col, "timestamp", "fold").is_duplicated().any():
raise ValueError(f"Duplicate canonical keys in {case_study} prediction panel")
return df.select(required)
# %%
preds = {cs: load_predictions(cs) for cs in REGISTRY_ROOTS}
for cs, df in preds.items():
cfg = BEST_GBM[cs]
print(
f"{cs}: {df.height:,} rows, {df[cfg['id_col']].n_unique()} entities, "
f"{df['fold'].n_unique()} validation folds, "
# strftime, not .date(): a prediction artifact's timestamp is Date in some case
# studies and Datetime in others - etfs ships both - and `datetime.date` has no
# .date(). Both types format the same way.
f"{df['timestamp'].min():%Y-%m-%d} to {df['timestamp'].max():%Y-%m-%d}"
)
# %% [markdown]
# ## 3. Signal Strength of the Underlying GBM Predictions
#
# Conformal position sizing only adds value when the underlying signal has
# meaningful cross-sectional predictive power. If IC is near zero, weighting
# *anything* by inverse uncertainty cannot rescue the strategy - the
# conformal widths just rescale noise.
#
# We summarize daily cross-sectional Spearman IC and use Newey-West inference
# with at least $h-1$ lags. That horizon-aware correction matters because daily
# IC observations share overlapping 5-day or 21-day forward returns. These are
# descriptive diagnostics on the same validation panel used to select the GBM,
# so they do not establish an unbiased out-of-sample edge.
# %%
ic_summaries = {}
for cs, df in preds.items():
id_col = BEST_GBM[cs]["id_col"]
horizon = HORIZON_ETF if cs == "etfs" else HORIZON_CME
sig_df = df.rename(
{
"prediction": "factor",
"actual": f"{horizon}D_fwd_return",
}
).select(["timestamp", id_col, "factor", f"{horizon}D_fwd_return"])
_, ic_values = extract_signal_ic_series(
sig_df,
period=horizon,
date_col="timestamp",
asset_col=id_col,
)
hac = compute_ic_hac_stats(np.asarray(ic_values), label_horizon=horizon)
ic_summaries[cs] = {
"mean": hac["mean_ic"],
"t_stat": hac["t_stat"],
"p_value": hac["p_value"],
"pct_positive": float(np.mean(np.asarray(ic_values) > 0)),
"effective_lags": hac["effective_lags"],
"horizon": horizon,
"n_dates": len(ic_values),
}
for cs, s in ic_summaries.items():
print(
f"{cs}: mean IC={s['mean']:+.4f}, HAC t={s['t_stat']:+.2f} "
f"(lags={s['effective_lags']}), positive on {s['pct_positive']:.1%} "
f"of {s['n_dates']:,} validation dates"
)
# %% [markdown]
# Mean IC and the share of positive dates describe whether the selected signal
# ranks returns at all; the HAC statistic prevents overlapping labels from
# manufacturing precision. The result remains selection-conditioned validation
# evidence. The next section asks a separate question: whether residual widths
# vary enough across entities for inverse-width sizing to differ from equal weight.
# %% [markdown]
# ## 4. Per-Entity Mondrian Widths Without Future-Fold Calibration
#
# For each fold $k$ and entity $i$, we compute the conformal order statistic from
# absolute residuals on chronologically prior folds. We also remove the trailing
# $h$ timestamps whose labels would not yet be observable when fold $k$ begins.
# The width applied to fold-$k$ predictions is
#
# $$\Delta_{i,k} = 2 \cdot r_{(\lceil(n+1)(1-\alpha)\rceil)},$$
#
# where $r_{(j)}$ is the $j$th ordered residual and the rank is capped at $n$.
# The earliest fold has no valid prior calibration pool and is excluded rather
# than borrowing future residuals. Consequently, all evaluated folds obey the
# same walk-forward contract.
#
# Per-entity calibration produces the cross-sectional width variation that
# drives $1/\Delta_i$ weighting. Pooled calibration
# split conformal would produce a single width per fold, collapsing the
# conformal weights to equal weights.
# %% [markdown]
# The finite-sample rank uses the standard split-conformal $(n+1)$ correction.
# This small helper is also exercised against a hand-calculated oracle before execution.
# %%
def conformal_order_stat(residuals: np.ndarray, alpha: float) -> float:
"""Return the finite-sample split-conformal residual order statistic."""
ordered = np.sort(np.asarray(residuals, dtype=float))
if ordered.size == 0:
raise ValueError("Conformal calibration requires at least one residual")
rank = min(int(np.ceil((ordered.size + 1) * (1.0 - alpha))), ordered.size)
return float(ordered[rank - 1])
# %% [markdown]
# Fold order is derived from timestamps rather than numeric fold IDs. A residual
# enters fold $k$'s calibration pool only when its origin plus the label horizon
# is strictly earlier than fold $k$'s first timestamp.
# %%
def mondrian_widths(df: pl.DataFrame, id_col: str, horizon: int, alpha: float) -> pl.DataFrame:
"""Compute strictly prior, horizon-embargoed widths for each evaluable fold."""
calendar = df.select("timestamp").unique().sort("timestamp").with_row_index("time_index")
panel = df.join(calendar, on="timestamp").with_columns(
abs_resid=(pl.col("actual") - pl.col("prediction")).abs()
)
fold_starts = (
panel.group_by("fold").agg(start_index=pl.col("time_index").min()).sort("start_index")
)
rows: list[dict[str, str | int | float]] = []
for fold, start_index in fold_starts.iter_rows():
if start_index == fold_starts["start_index"].min():
continue
current_ids = panel.filter(pl.col("fold") == fold)[id_col].unique()
calibration = panel.filter(
(pl.col("time_index") + horizon < start_index)
& pl.col(id_col).is_in(current_ids.to_list())
)
for entity_key, group in calibration.group_by(id_col):
entity = entity_key[0]
quantile = conformal_order_stat(group["abs_resid"].to_numpy(), alpha)
if quantile <= 0:
raise ValueError(f"Non-positive conformal width for {entity}, fold {fold}")
rows.append(
{
id_col: entity,
"fold": int(fold),
"width": 2.0 * quantile,
"n_calibration": group.height,
}
)
return pl.DataFrame(rows)
# %%
widths = {}
for cs, df in preds.items():
config = BEST_GBM[cs]
horizon = HORIZON_ETF if cs == "etfs" else HORIZON_CME
widths[cs] = mondrian_widths(df, config["id_col"], horizon, CONFORMAL_ALPHA)
for cs, w in widths.items():
print(
f"{cs}: {w.height} entity-fold widths across {w['fold'].n_unique()} evaluable folds; "
f"median={w['width'].median():.4f}, range={w['width'].min():.4f} to "
f"{w['width'].max():.4f}"
)
# %% [markdown]
# Because the two labels have different horizons, raw widths are not directly
# comparable. Normalizing each width by its case-study median isolates relative
# cross-sectional dispersion, the quantity that makes inverse-width weights depart
# from equal weight.
# %%
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
for (cs, w), color in zip(widths.items(), ml4t_palette(2, categorical=True), strict=True):
relative_width = np.sort(w["width"].to_numpy() / w["width"].median())
cumulative_share = np.arange(1, relative_width.size + 1) / relative_width.size
ax.step(relative_width, cumulative_share, where="post", label=cs, color=color)
ax.axvline(1.0, color=COLORS["neutral"], linewidth=0.8, linestyle="--")
ax.set_xscale("log")
ax.set_xlabel("Conformal width / case-study median (log scale)")
ax.set_ylabel("Cumulative share of entity-fold widths")
ax.legend(title="Validation panel")
add_message_title(
ax,
"Conformal widths relative to each panel's own median",
subtitle="Finite-sample Mondrian widths; strictly prior, horizon-embargoed calibration",
)
show_with_alt(
fig,
"Two step curves of the cumulative share of entity-fold conformal widths against width divided by the case study's own median, on a logarithmic horizontal axis, with a dashed line at the median. The CME curve is the flatter of the two.",
)
# %% [markdown]
# ## 5. Allocation Rules
#
# At each timestamp $t$ we (a) select the top-$K$ entities by signed
# $\hat y_{i,t}$, restricted to rows whose prediction is positive (long-only
# return-score allocation), then (b) form weights using each rule:
#
# - **Equal weight**: $w_i = 1/K$.
# - **Conformal-weighted**: $w_i = (1/\Delta_{i,t}) / \sum_{j \in \text{top-}K}(1/\Delta_{j,t})$.
# - **Score-weighted**: $w_i = \hat y_{i,t} / \sum_{j \in \text{top-}K}\hat y_{j,t}$.
#
# All weights are non-negative and sum to one - long-only by construction. The
# realized portfolio return at $t$ is $r_t = \sum_i w_{i,t}\,y_{i,t}$, where
# $y_{i,t}$ is the realized forward return label aligned to the prediction
# timestamp (`actual` column in the parquet). If fewer than $K$ entities have
# a positive predicted return at $t$, the selection drops to that smaller
# size; if none are positive the timestamp is skipped.
# %% [markdown]
# A single global schedule is built after the uncalibrated first fold is removed.
# It never resets at fold boundaries, so adjacent folds cannot reintroduce
# overlapping forward-return windows.
# %%
def nonoverlap_schedule(df: pl.DataFrame, horizon: int) -> pl.DataFrame:
"""Select every horizon-th timestamp once across the full evaluation panel."""
return (
df.select("timestamp")
.unique()
.sort("timestamp")
.with_row_index("time_index")
.filter(pl.col("time_index") % horizon == 0)
.select("timestamp")
)
# %% [markdown]
# The three allocators share the same positive-score top-$K$ selection. Only the
# within-selection weights differ, which isolates the sizing decision.
# %%
def build_portfolio_returns(
df: pl.DataFrame,
widths: pl.DataFrame,
id_col: str,
top_k: int,
horizon: int,
) -> tuple[pl.DataFrame, pl.DataFrame]:
"""Build non-overlapping validation returns and native-identifier weights."""
panel = df.join(
widths.select(id_col, "fold", "width"), on=[id_col, "fold"], how="inner"
).drop_nulls(subset=["prediction", "actual", "width"])
panel = panel.join(nonoverlap_schedule(panel, horizon), on="timestamp", how="inner")
panel = panel.filter(pl.col("prediction") > 0).with_columns(
rank=pl.col("prediction").rank(method="ordinal", descending=True).over("timestamp")
)
top = panel.filter(pl.col("rank") <= top_k)
inv_w = 1.0 / pl.col("width")
score = pl.col("prediction")
top = top.with_columns(
w_ew=pl.lit(1.0) / pl.len().over("timestamp"),
w_conf=inv_w / inv_w.sum().over("timestamp"),
w_score=score / score.sum().over("timestamp"),
)
realized = (
top.group_by("timestamp")
.agg(
ret_ew=(pl.col("w_ew") * pl.col("actual")).sum(),
ret_conf=(pl.col("w_conf") * pl.col("actual")).sum(),
ret_score=(pl.col("w_score") * pl.col("actual")).sum(),
)
.sort("timestamp")
)
weights_long = top.select("timestamp", id_col, "w_ew", "w_conf", "w_score").sort(
"timestamp", id_col
)
return realized, weights_long
# %% [markdown]
# One-way turnover is half the $L_1$ change in weights. The factor of one-half
# avoids counting every sale and offsetting purchase as two separate reallocations.
# %%
def compute_turnover(weights_long: pl.DataFrame, id_col: str, weight_col: str) -> float:
"""Return mean one-way turnover across consecutive rebalances."""
w = weights_long.pivot(
index="timestamp", on=id_col, values=weight_col, aggregate_function="first"
).fill_null(0.0)
arr = w.drop("timestamp").to_numpy()
if arr.shape[0] < 2:
return float("nan")
return float(0.5 * np.mean(np.abs(np.diff(arr, axis=0)).sum(axis=1)))
# %% [markdown]
# The four statistics come from `ml4t.diagnostic` rather than being written out here, and each
# is passed the number of non-overlapping horizon periods in a year instead of the daily 252:
# the return series has one observation per rebalance, not one per session, so annualizing it as
# if it were daily would inflate every ratio. `annual_return` compounds the realized path rather
# than the arithmetic mean return, which is what makes it consistent with the drawdown computed
# from the same path.
# %%
def metric_block(
rets: pl.DataFrame, weights: pl.DataFrame, id_col: str, horizon: int
) -> dict[str, dict[str, float]]:
"""Compute path-consistent validation metrics for the three allocators."""
periods_per_year = TRADING_DAYS_PER_YEAR / horizon
out: dict[str, dict[str, float]] = {}
for name, ret_col, w_col in [
("baseline_equal_weight", "ret_ew", "w_ew"),
("conformal_weighted", "ret_conf", "w_conf"),
("score_weighted", "ret_score", "w_score"),
]:
r = rets[ret_col].to_numpy()
out[name] = {
"sharpe": sharpe_ratio(r, periods_per_year=periods_per_year),
"sortino": sortino_ratio(r, periods_per_year=periods_per_year),
"max_drawdown": float(max_drawdown(r)),
"win_rate": float((r > 0).mean()),
"avg_turnover": compute_turnover(weights, id_col, w_col),
"n_obs": int(r.shape[0]),
"ann_return": annual_return(r, periods_per_year=periods_per_year),
}
return out
# %% [markdown]
# **On rebalancing and annualization.** The prediction panel is daily but each
# row's `actual` is a forward 5-day or 21-day return. To avoid the overlap that
# would inflate Sharpe and break drawdown accounting, one global schedule keeps
# every `HORIZON`th trading date without restarting at fold boundaries. The
# Sharpe is annualized by $\sqrt{252/h}$ - for ETFs with $h=21$ this is
# $\sqrt{252/21}$; for CME with $h=5$ it is $\sqrt{252/5}$. The realized portfolio
# return at each rebalance date $t$ is $\sum_i w_{i,t}\,y_{i,t}$, where $y_{i,t}$
# is the $h$-day forward return and the cohort is held to maturity (no
# intra-period rebalancing).
# %% [markdown]
# One table shape for both case studies, so the two sections are read the same way.
# %%
def allocator_table(metrics: dict[str, dict[str, float]]) -> pl.DataFrame:
"""One row per sizing rule, carrying the four measures section 8 draws."""
labels = {
"baseline_equal_weight": "Equal weight",
"conformal_weighted": "Conformal weighted",
"score_weighted": "Score weighted",
}
return pl.DataFrame(
[
{
"sizing rule": label,
"annualized Sharpe": metrics[key]["sharpe"],
"annualized return": metrics[key]["ann_return"],
"maximum drawdown": metrics[key]["max_drawdown"],
"mean one-way turnover": metrics[key]["avg_turnover"],
}
for key, label in labels.items()
]
)
# %% [markdown]
# ## 6. ETFs (`fwd_ret_21d`)
#
# The first validation fold is excluded because it has no strictly prior
# calibration residuals. The remaining folds share one global 21-session
# rebalance schedule. These metrics compare allocators on the selected GBM's
# validation panel; they are not final holdout performance.
# %%
etf_rets, etf_weights = build_portfolio_returns(
preds["etfs"], widths["etfs"], "symbol", TOP_K_ETF, HORIZON_ETF
)
etf_metrics = metric_block(etf_rets, etf_weights, "symbol", HORIZON_ETF)
print(f"ETFs: {etf_rets.height} non-overlapping validation rebalances")
# %%
allocator_table(etf_metrics)
# %% [markdown]
# ## 7. CME Futures (`fwd_ret_5d`)
#
# The same protocol excludes the first fold and uses a global five-session
# schedule across the rest. The entity values are CME product roots; the
# prediction artifact stores them under `symbol`, as it does for every case
# study.
# %%
cme_rets, cme_weights = build_portfolio_returns(
preds["cme_futures"], widths["cme_futures"], "symbol", TOP_K_CME, HORIZON_CME
)
cme_metrics = metric_block(cme_rets, cme_weights, "symbol", HORIZON_CME)
print(f"CME futures: {cme_rets.height} non-overlapping validation rebalances")
# %%
allocator_table(cme_metrics)
# %% [markdown]
# ## 8. Allocation Trade-offs on Validation
#
# Each panel uses the same case-study axis and allocator colors. Sharpe and
# annual return measure reward, max drawdown measures path risk, and one-way
# turnover shows how aggressively each sizing rule changes positions.
# %%
case_metrics = {"ETFs": etf_metrics, "CME futures": cme_metrics}
method_keys = ["baseline_equal_weight", "conformal_weighted", "score_weighted"]
method_labels = ["Equal weight", "Conformal", "Score"]
metric_specs = [
("sharpe", "Annualized Sharpe"),
("ann_return", "Annualized return"),
("max_drawdown", "Maximum drawdown"),
("avg_turnover", "Mean one-way turnover"),
]
fig, axes = plt.subplots(2, 2, figsize=FIGSIZE["dashboard_2x2"])
x = np.arange(len(case_metrics))
bar_width = 0.24
colors = ml4t_palette(3, categorical=True)
legend_handles = []
for panel_index, (ax, (metric, label)) in enumerate(zip(axes.flat, metric_specs, strict=True)):
for offset, (method, method_label, color) in enumerate(
zip(method_keys, method_labels, colors, strict=True)
):
values = [metrics[method][metric] for metrics in case_metrics.values()]
bars = ax.bar(
x + (offset - 1) * bar_width, values, bar_width, label=method_label, color=color
)
if panel_index == 0:
legend_handles.append(bars)
zero_line(ax)
ax.set_xticks(x, case_metrics.keys())
ax.set_ylabel(label)
if metric in {"ann_return", "max_drawdown"}:
ax.yaxis.set_major_formatter(lambda value, _: f"{value:.0%}")
fig.suptitle("Four measures of the same three sizing rules, on two panels")
fig.legend(legend_handles, method_labels, loc="outside lower center", ncol=3)
show_with_alt(
fig,
"Four panels comparing equal-weight, conformal and score-weighted sizing on the ETF and CME panels: annualized Sharpe, annualized return, maximum drawdown and mean one-way turnover, three grouped bars per case study in each panel.",
)
# %% [markdown] tags=["results"]
# ### What this run produced
#
# Two panels, three sizing rules, four measures each - tabulated in sections 6 and 7 and drawn
# together in the four-panel figure. The three rules share one selection at every rebalance, so
# any difference between them comes from the weights alone.
#
# Read the Sharpe panel against the turnover panel rather than on its own. A sizing rule that
# improves Sharpe while turning over more has not been shown to be worth using: this comparison
# is gross of the cost of that turnover, and Chapter 18 measures what it would take out. The two
# case studies also differ in how strong the underlying signal is, which section 3 reports as a
# HAC t-statistic on the mean IC - a sizing rule applied to a signal that does not rank returns
# is rescaling noise, whatever its Sharpe ratio comes out at.
# %% [markdown]
# ## Key Takeaways
#
# 1. **Per-entity calibration creates the sizing signal.** Pooled calibration
# would produce one width per fold and collapse inverse-width sizing to equal
# weight. The finite-sample Mondrian widths use only prior residuals whose
# labels are available before the next fold; the uncalibrated first fold is
# excluded rather than filled from the future.
#
# 2. **Inverse-width sizing is a bet on the residual dispersion being stable.** It puts more
# capital where the model's past residuals were narrow. That helps only if an entity whose
# residuals were narrow in the calibration window stays that way in the evaluation window;
# nothing in conformal prediction guarantees it, and the printed comparison is what says
# whether it held here.
#
# 3. **Width dispersion is necessary but not sufficient.** Inverse-width weights depart from
# equal weight only in proportion to how much the widths differ across entities, so a panel
# with tight dispersion cannot produce a materially different portfolio. Dispersion is what
# makes the rule able to act; it is not what makes the action pay. Uncertainty changes
# position size and cannot manufacture predictive direction.
#
# 4. **Treat these results as validation diagnostics.** The GBM was selected by
# validation IC from the same registry panel. A holdout that no selection step has seen is
# required before any allocator ranking counts as out-of-sample evidence.
#
# **Next**: [`08_library_comparison`](08_library_comparison.ipynb) puts four allocation
# libraries on the same problem and compares what each one produces.
#
# **Book**: Section 17.4 lists conformal-weighted allocation alongside the
# inverse-volatility, score-weighted, and equal-weight baselines.
```Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.