Inferencia HAC y bootstrap por bloques para coeficientes de información
Resumen
Este cuaderno evalúa si el coeficiente de información (IC) de una señal se distingue del ruido cuando las observaciones sucesivas son dependientes. Ilustra el problema con una señal de momentum y datos de ETF, y explica cómo los rendimientos futuros superpuestos, la persistencia de las clasificaciones de señales y los regímenes de mercado pueden hacer que los valores diarios de IC presenten autocorrelación. En ese caso, los estadísticos t ingenuos pueden subestimar la incertidumbre y exagerar la significación.
El cuaderno compara los errores estándar de Newey–West, consistentes ante heterocedasticidad y autocorrelación, con un bootstrap por bloques que remuestrea observaciones contiguas. La elección del ancho de banda se vincula al horizonte de las etiquetas y se contrasta con el patrón de autocorrelación observado; el ejemplo muestra que una regla basada solo en el tamaño de la muestra puede pasar por alto la dependencia. También analiza el tamaño efectivo de la muestra, los intervalos de confianza, la significación práctica después de costes y la cantidad de datos necesaria para evaluar una señal.
El análisis es un ejemplo práctico, no una prescripción universal del ancho de banda. La superposición por sí sola no garantiza dependencia serial en los IC, y la dependencia observada también puede reflejar clasificaciones persistentes o regímenes. El cuaderno recomienda comunicar la incertidumbre que tiene en cuenta la dependencia junto con la significación estadística y práctica.
Ideas clave
- La autocorrelación en una serie de IC puede hacer que un estadístico t iid exagere la evidencia a favor de una señal.
- Los errores estándar de Newey–West ajustan la incertidumbre mediante autocovarianzas ponderadas dentro de un ancho de banda de retardos elegido.
- El ancho de banda debe tener en cuenta el horizonte de las etiquetas y también contrastarse con el patrón de dependencia observado.
- Un bootstrap por bloques conserva la dependencia local al remuestrear tramos contiguos de la serie IC.
- El tamaño efectivo de la muestra y los costes de trading importan al juzgar si la evidencia estadística es útil en la práctica.
Etiquetas
Texto completo
# 06_ic_inference.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # IC Inference: HAC Adjustment and Block Bootstrap
#
# **Docker image**: `ml4t`
#
# **Chapter 7: Defining the Learning Task**
# **Section Reference**: 7.3 - Feature and Label Evaluation as Triage
#
# ## Purpose
#
# This notebook addresses **statistical inference for IC**: how confident should we
# be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated
# IC series and block bootstrap for robust confidence intervals.
#
# ## Learning Objectives
#
# 1. Understand why naive t-statistics fail for IC inference
# 2. Apply HAC (Newey-West) adjustment for proper standard errors
# 3. Implement block bootstrap for distribution-free confidence intervals
# 4. Distinguish practical vs statistical significance
# 5. Plan track records using HAC-adjusted effective sample size
#
# ## Data Policy
#
# Examples use **real ETF data** from the case study store.
#
# ## Prerequisites
#
# - `05_signal_evaluation` - produces the IC time series whose inference we
# formalize here.
# - Familiarity with autocorrelation, the Newey-West HAC estimator, and
# block bootstrap.
# %%
"""IC Inference - statistical testing and confidence intervals for information coefficients."""
from __future__ import annotations
import json
import logging
import warnings
import numpy as np
import plotly.graph_objects as go
import polars as pl
from arch.bootstrap import StationaryBootstrap
from IPython.display import display
from ml4t.diagnostic.evaluation.autocorrelation import analyze_autocorrelation
from ml4t.diagnostic.metrics import compute_ic_hac_stats
from ml4t.diagnostic.signal import analyze_signal
from plotly.subplots import make_subplots
from scipy import stats
from data import load_etfs
from utils.reproducibility import set_global_seeds
from utils.style import ( # importing utils.style activates the ml4t Plotly template
COLORS,
show_plotly_with_alt,
)
# Quiet ml4t library INFO logging. Child loggers set their own level, so raising the
# parent alone does not silence them - set every already-created ml4t logger explicitly.
for _lg in list(logging.root.manager.loggerDict):
if _lg.startswith("ml4t"):
logging.getLogger(_lg).setLevel(logging.WARNING)
# %% tags=["parameters"]
SEED = 42
START_DATE = ""
N_BOOT = 2000
# Forward-return horizon, in trading days. Every dependence-aware choice in this
# notebook is derived from it: the momentum lookback, the IC evaluation period,
# the HAC lag truncation, and the bootstrap block length.
LABEL_HORIZON = 21
# %%
set_global_seeds(SEED)
# %% [markdown]
# ## The Autocorrelation Problem
#
# IC time series exhibit **autocorrelation** due to:
#
# 1. **Overlapping forward returns** at longer horizons
# 2. **Persistent signal values** (momentum doesn't flip daily)
# 3. **Regime persistence** (markets stay in trends/ranges)
#
# Naive t-statistics assume iid observations and **underestimate standard errors**,
# leading to inflated significance.
# %%
# Load real ETF data and compute momentum signal
etfs = load_etfs()
if START_DATE:
etfs = etfs.filter(pl.col("timestamp") >= pl.lit(START_DATE).str.to_date())
# Compute momentum over the label horizon
factor_df = (
etfs.sort(["symbol", "timestamp"])
.with_columns(
[
(pl.col("close") / pl.col("close").shift(LABEL_HORIZON).over("symbol") - 1).alias(
"factor"
)
]
)
.filter(pl.col("factor").is_not_null())
.select(["timestamp", "symbol", "factor"])
)
prices_df = etfs.select(["timestamp", "symbol", "close"]).rename({"close": "price"})
# Run signal analysis to get IC series
result = analyze_signal(
factor_df,
prices_df,
periods=(LABEL_HORIZON,),
quantiles=5,
ic_method="spearman",
date_col="timestamp",
asset_col="symbol",
)
ic_series = np.array(result.ic_series.get(f"{LABEL_HORIZON}D", []))
print(f"IC series length: {len(ic_series)}")
print(f"Mean IC: {np.mean(ic_series):.4f}")
# %% [markdown]
# ### Autocorrelation in IC Series
#
# Let's examine the autocorrelation structure directly.
# %%
# Compute autocorrelation function
def compute_acf(series, nlags=20):
"""Compute autocorrelation function."""
n = len(series)
series = series - np.mean(series)
acf = []
for lag in range(nlags + 1):
if lag == 0:
acf.append(1.0)
else:
acf.append(np.corrcoef(series[lag:], series[:-lag])[0, 1])
return np.array(acf)
n_acf_lags = LABEL_HORIZON - 1
acf_values = compute_acf(ic_series, nlags=n_acf_lags)
# Significance bounds for white noise
n = len(ic_series)
sig_bound = 1.96 / np.sqrt(n)
fig = go.Figure()
n_significant_lags = sum(abs(acf_values[1:]) > sig_bound)
fig.add_trace(
go.Bar(
x=list(range(1, n_acf_lags + 1)),
y=acf_values[1:],
name="ACF",
marker_color=[
COLORS["amber"] if abs(v) > sig_bound else COLORS["blue"] for v in acf_values[1:]
],
)
)
fig.add_hline(
y=sig_bound, line_dash="dash", line_color=COLORS["negative"], annotation_text="95% band"
)
fig.add_hline(y=-sig_bound, line_dash="dash", line_color=COLORS["negative"])
fig.add_hline(y=0, line_color=COLORS["neutral"])
fig.update_layout(
title=f"Autocorrelation of the daily IC series, lags 1 to {LABEL_HORIZON - 1}",
xaxis_title="Lag (days)",
yaxis_title="Autocorrelation",
template="ml4t",
height=350,
)
show_plotly_with_alt(
fig,
alt=(
"A bar chart of the autocorrelation of the daily IC series against lag, from one "
"day out to one day short of the label horizon, with dashed red lines marking the "
"white-noise band either side of zero. The bars start just above 0.8 at lag one "
"and fall away in a smooth, almost linear decline, reaching the band around lag "
"sixteen. Every bar up to that point is amber, marking it outside the band; only "
"the last few lags are inside it and drawn in navy."
),
)
print(f"\nSignificant autocorrelation lags: {n_significant_lags}/{n_acf_lags}")
# %% [markdown]
# The decline is smooth and reaches the white-noise band near the label horizon, which is
# consistent with overlap driving it: two IC values computed $k$ days apart share $h - k$
# days of return, so the shared window shrinks as $k$ grows. It is consistent with
# persistent signal ranks as well, and the two are not separated here. Nor can this figure
# say what happens *past* the horizon - it stops at $h-1$, and a series of rank
# correlations can stay dependent beyond the overlap through regimes or a slow-moving
# signal. What the figure does establish is enough for the section's purpose: a naive
# t-statistic treats these values as independent draws, and they are not.
# %% [markdown]
# ### Library Autocorrelation Analysis
#
# The library version includes PACF and the **Ljung-Box portmanteau test**,
# which tests the joint null that all autocorrelations up to lag $L$ are zero.
# %%
acf_analysis = analyze_autocorrelation(ic_series, max_lags=LABEL_HORIZON - 1, alpha=0.05)
print("=== ml4t-diagnostic Autocorrelation Analysis ===\n")
print(f"Significant ACF lags: {acf_analysis.significant_acf_lags}")
print(f"Significant PACF lags: {acf_analysis.significant_pacf_lags}")
print(f"Is white noise: {acf_analysis.is_white_noise}")
print(f"Suggested ARIMA order: {acf_analysis.suggested_arima_order}")
# %% [markdown]
# ## HAC (Newey-West) Adjustment
#
# HAC (Heteroskedasticity and Autocorrelation Consistent) standard errors account
# for serial correlation in the IC series.
#
# The Newey-West estimator uses a weighted sum of autocovariances:
#
# $$\hat{\sigma}^2_{HAC} = \hat{\gamma}_0 + 2\sum_{j=1}^{L} w_j \hat{\gamma}_j$$
#
# Where $w_j = 1 - j/(L+1)$ (Bartlett kernel) and $L$ is the lag truncation.
#
# **$L$ should not be left to the sample size here.** The usual automatic rule
# $L = \lfloor 4(T/100)^{2/9} \rfloor$ reads only $T$, and on this series it
# returns 9. The labels are $h$-day *overlapping* forward returns, so date $t$ and
# date $t+j$ share part of their outcome window for every $j < h$ - and the signal
# is 21-day momentum, whose cross-sectional ranks change slowly. Overlapping
# outcomes measured against persistent ranks give ICs that move together over
# roughly the label horizon.
#
# **Both halves of that sentence are doing work, and the overlap alone is not
# enough.** The IC is a rank correlation, not a return: if the signal ranks were
# redrawn independently every day, the ICs could be serially uncorrelated no matter
# how much the return windows overlap. So $h-1$ is not a theorem about overlapping
# labels - it is a **conservative horizon-aware bandwidth for a persistent signal**,
# and the evidence that it is the right order of magnitude is the measured ACF in
# the figure above, which stays outside the white-noise band until about lag 16.
#
# What is not defensible is $L = 9$, chosen by a rule that never looked at the
# labels or the signal. Passing `label_horizon` makes the library take
# $L = \max(h-1, \text{auto rule})$; the ACF is what justifies that choice, and the
# next section measures what it is worth.
# %% [markdown]
# ### What the Bandwidth Is Worth
#
# Worth measuring rather than asserting, so the cell below runs the estimator
# both ways on the same IC series. The block bootstrap below resamples contiguous blocks
# of length $h$, so it already respects the overlap. If HAC is given a bandwidth that does
# not, the two will not agree - and the gap is then a property of the bandwidth, not of the
# estimator families, which is what the side-by-side is there to establish. The bootstrap
# section shows the two agreeing once both are horizon-aware.
#
# Calling the estimator without `label_horizon` raises a `UserWarning` saying the automatic
# bandwidth may be anti-conservative for overlapping labels. That warning is the library
# doing its job, and the call below triggers it on purpose; it is silenced for that one
# call so it does not read as a defect in the run.
# %%
# Compute HAC-adjusted statistics
hac_result = compute_ic_hac_stats(ic_series, label_horizon=LABEL_HORIZON)
# The same estimator with the bandwidth left to the sample-size rule. Omitting
# label_horizon is deliberate here and the library warns about it, correctly; the warning
# is silenced for this one call, by category and message, so it does not read as a defect.
with warnings.catch_warnings():
warnings.filterwarnings(
"ignore",
category=UserWarning,
message="label_horizon was not provided.*",
)
hac_auto = compute_ic_hac_stats(ic_series)
print("=== Bandwidth: sample-size rule vs label-horizon aware ===\n")
print(
f" auto rule L={hac_auto['effective_lags']:>3} "
f"SE={hac_auto['hac_se']:.6f} t={hac_auto['t_stat']:.4f} p={hac_auto['p_value']:.4f}"
)
print(
f" label_horizon={LABEL_HORIZON:<8} L={hac_result['effective_lags']:>3} "
f"SE={hac_result['hac_se']:.6f} t={hac_result['t_stat']:.4f} p={hac_result['p_value']:.4f}"
)
print(f"\n SE ratio (horizon-aware / auto): {hac_result['hac_se'] / hac_auto['hac_se']:.3f}x")
# Compare naive vs HAC
naive_se = np.std(ic_series, ddof=1) / np.sqrt(len(ic_series))
naive_t = np.mean(ic_series) / naive_se
print("=== Naive vs HAC Inference ===\n")
print(
f"HAC lag truncation: {hac_result['effective_lags']} "
f"(horizon-aware choice: label horizon {LABEL_HORIZON} -> L = {LABEL_HORIZON - 1}, "
f"consistent with the measured ACF above)"
)
print(f"Mean IC: {hac_result['mean_ic']:.4f}")
print("\nStandard Errors:")
print(f" Naive SE: {naive_se:.4f}")
print(f" HAC SE: {hac_result['hac_se']:.4f}")
print(f" Inflation: {hac_result['hac_se'] / naive_se:.2f}x")
print("\nt-Statistics:")
print(f" Naive t: {naive_t:.2f}")
print(f" HAC t: {hac_result['t_stat']:.2f}")
print("\np-Values:")
print(f" Naive p: {2 * stats.t.sf(abs(naive_t), len(ic_series) - 1):.4f}")
print(f" HAC p: {hac_result['p_value']:.4f}")
# Significance conclusion
alpha = 0.05
naive_sig = "Yes" if abs(naive_t) > stats.t.ppf(1 - alpha / 2, len(ic_series) - 1) else "No"
hac_sig = "Yes" if hac_result["p_value"] < alpha else "No"
print("\nSignificant at α=0.05?")
print(f" Naive: {naive_sig}")
print(f" HAC: {hac_sig}")
if naive_sig == "Yes" and hac_sig == "No":
print("\n WARNING: Naive test finds significance that HAC adjustment removes!")
# %% [markdown]
# ### Effective Sample Size
#
# The HAC SE inflation factor tells us how much the autocorrelation reduces our
# effective sample size:
#
# $$T_{eff} = T \times \left(\frac{\sigma_{naive}}{\sigma_{HAC}}\right)^2$$
# %%
# Compute effective sample size
inflation_factor = hac_result["hac_se"] / naive_se
effective_n = len(ic_series) / (inflation_factor**2)
print("=== Effective Sample Size ===\n")
print(f"Actual observations: {len(ic_series)}")
print(f"HAC inflation factor: {inflation_factor:.2f}x")
print(f"Effective sample size: {effective_n:.0f}")
print(f"Efficiency loss: {(1 - effective_n / len(ic_series)) * 100:.0f}%")
# %% [markdown]
# ## Block Bootstrap for Robust CI
#
# When IC series has autocorrelation, **iid bootstrap is invalid**. We use
# **block bootstrap** which preserves the dependence structure by resampling
# contiguous blocks of observations.
#
# Two common approaches:
# 1. **Moving Block Bootstrap (MBB)**: Fixed block length
# 2. **Stationary Bootstrap**: Random block lengths (geometric distribution)
#
# We use the `arch` library's `StationaryBootstrap` for robust inference.
# %%
# Block bootstrap implementation
def block_bootstrap_ic_ci(
ic_series: np.ndarray,
n_bootstrap: int = 1000,
block_length: int = 21,
alpha: float = 0.05,
random_state: int = 42,
) -> dict:
"""
Compute block bootstrap confidence interval for mean IC.
Uses the Politis-Romano stationary bootstrap from `arch.bootstrap`:
random block lengths drawn from a geometric distribution with
expected length equal to ``block_length``.
Parameters
----------
ic_series : array
Time series of IC values
n_bootstrap : int
Number of bootstrap replications
block_length : int
Expected block length (mean of the geometric distribution)
alpha : float
Significance level (default 0.05 for 95% CI)
random_state : int
Random seed for reproducibility
Returns
-------
dict with mean, CI bounds, and bootstrap distribution
"""
ic_series = np.array(ic_series)
bs = StationaryBootstrap(block_length, ic_series, seed=random_state)
bootstrap_means = np.array([np.mean(data[0][0]) for data in bs.bootstrap(n_bootstrap)])
# Percentile confidence interval
ci_lower = np.percentile(bootstrap_means, 100 * alpha / 2)
ci_upper = np.percentile(bootstrap_means, 100 * (1 - alpha / 2))
return {
"mean": np.mean(ic_series),
"ci_lower": ci_lower,
"ci_upper": ci_upper,
"bootstrap_std": np.std(bootstrap_means),
"bootstrap_means": bootstrap_means,
"method": "Stationary Bootstrap",
"block_length": block_length,
}
# %%
# Run block bootstrap
n_boot = N_BOOT
block_result = block_bootstrap_ic_ci(
ic_series, n_bootstrap=n_boot, block_length=LABEL_HORIZON, random_state=SEED
)
# Compare with HAC CI
hac_ci_lower = hac_result["mean_ic"] - 1.96 * hac_result["hac_se"]
hac_ci_upper = hac_result["mean_ic"] + 1.96 * hac_result["hac_se"]
print("=== Block Bootstrap vs HAC Confidence Intervals ===\n")
print(f"Method: {block_result['method']} (block length={block_result['block_length']})")
print(f"Bootstrap replications: {n_boot}")
print(f"\nMean IC: {block_result['mean']:.4f}")
print("\n95% Confidence Intervals:")
print(f" Block Bootstrap: [{block_result['ci_lower']:.4f}, {block_result['ci_upper']:.4f}]")
print(f" HAC (Newey-West): [{hac_ci_lower:.4f}, {hac_ci_upper:.4f}]")
print("\nStandard Errors:")
print(f" Block Bootstrap: {block_result['bootstrap_std']:.4f}")
print(f" HAC: {hac_result['hac_se']:.4f}")
# Does CI include zero?
ci_includes_zero = block_result["ci_lower"] <= 0 <= block_result["ci_upper"]
print(f"\nBlock Bootstrap CI includes zero: {'Yes' if ci_includes_zero else 'No'}")
# %% [markdown]
# ### Library vs Manual Bootstrap: When to Use Each
#
# The `ml4t-diagnostic` library provides `stationary_bootstrap_ic()` for
# **cross-sectional** bootstrap: given arrays of predictions and returns for a
# single date, it resamples asset pairs and recomputes Spearman IC. This is useful
# when you want a confidence interval for a single cross-section's IC.
#
# For **time-series** inference on the IC series itself - "is the mean IC over
# many dates significantly different from zero?" - the correct tool is the block bootstrap
# used above, which preserves temporal dependence.
#
# ```python
# # Cross-sectional bootstrap (library) - CI for one date's IC
# from ml4t.diagnostic.evaluation.stats.hac_standard_errors import stationary_bootstrap_ic
# result = stationary_bootstrap_ic(signals_t, returns_t, n_samples=1000)
#
# # Time-series bootstrap (manual) - CI for mean IC across dates
# block_bootstrap_ic_ci(ic_series, n_bootstrap=2000, block_length=21)
# ```
#
# Throughout this notebook we use the manual block bootstrap because we are
# testing whether the *time-averaged* IC is significantly positive.
# %%
print("=== Block Bootstrap Summary (recommended for IC time series) ===\n")
print(f" IC: {block_result['mean']:.4f}")
print(f" Bootstrap SE: {block_result['bootstrap_std']:.4f}")
print(f" 95% CI: [{block_result['ci_lower']:.4f}, {block_result['ci_upper']:.4f}]")
print(f" HAC CI: [{hac_ci_lower:.4f}, {hac_ci_upper:.4f}]")
# %%
# Visualize bootstrap distribution
fig = make_subplots(
rows=1,
cols=2,
subplot_titles=["Bootstrap distribution of mean IC", "95% CI: block bootstrap vs HAC"],
)
# Bootstrap distribution
fig.add_trace(
go.Histogram(
x=block_result["bootstrap_means"],
nbinsx=50,
name="Bootstrap means",
marker_color=COLORS["blue"],
opacity=0.7,
),
row=1,
col=1,
)
# Mark CIs
fig.add_vline(
x=block_result["ci_lower"],
line_dash="dash",
line_color=COLORS["negative"],
annotation_text="2.5%",
row=1,
col=1,
)
fig.add_vline(
x=block_result["ci_upper"],
line_dash="dash",
line_color=COLORS["negative"],
annotation_text="97.5%",
row=1,
col=1,
)
fig.add_vline(
x=0, line_dash="dot", line_color=COLORS["neutral"], annotation_text="Null", row=1, col=1
)
# CI comparison
methods = ["Block Bootstrap", "HAC"]
ci_lowers = [block_result["ci_lower"], hac_ci_lower]
ci_uppers = [block_result["ci_upper"], hac_ci_upper]
means = [block_result["mean"], hac_result["mean_ic"]]
for i, method in enumerate(methods):
# CI as error bars
fig.add_trace(
go.Scatter(
x=[method],
y=[means[i]],
error_y=dict(
type="data",
symmetric=False,
array=[ci_uppers[i] - means[i]],
arrayminus=[means[i] - ci_lowers[i]],
),
mode="markers",
marker=dict(size=12, color=[COLORS["blue"], COLORS["amber"]][i]),
name=method,
),
row=1,
col=2,
)
fig.add_hline(y=0, line_dash="dot", line_color=COLORS["neutral"], row=1, col=2)
fig.update_layout(
title="Mean IC: block-bootstrap distribution and the two interval estimates",
height=400,
template="ml4t",
showlegend=False,
)
fig.update_xaxes(title_text="Mean IC", row=1, col=1)
fig.update_yaxes(title_text="Frequency", row=1, col=1)
fig.update_yaxes(title_text="Mean IC", row=1, col=2)
show_plotly_with_alt(
fig,
alt=(
"Two panels. The left panel is a histogram of the block-bootstrap distribution of "
"the mean IC, a symmetric bell centred on zero and running out to roughly plus and "
"minus 0.045, with dashed red lines marking the 2.5th and 97.5th percentiles and a "
"dotted line at the null, which falls on the mode. "
"The right panel plots the two interval estimates side by side as a point with "
"error bars: the block bootstrap in navy and HAC in amber. Both point estimates "
"sit on the dotted zero line and both intervals span it, reaching roughly the "
"same distance above and below, so the two methods are hard to tell apart."
),
)
# %% [markdown]
# The two intervals are close enough to overlay, which is the result the bandwidth section
# predicted: once HAC is given a lag truncation matched to the label horizon, it agrees
# with a bootstrap that resamples blocks of that same length. They agree about the answer
# as well - both intervals contain zero, so the mean IC is not distinguishable from no
# skill at this sample size, whichever dependence-aware method is used.
# %% [markdown]
# ## Practical vs Statistical Significance
#
# Statistical significance of an IC time-series mean is a separate question from whether
# the signal is economically tradeable.
#
# There is no fixed IC magnitude that counts as detectable. Detectability is set by the
# standard error, and this notebook has already estimated the one that applies here - the
# HAC standard error on this series, at this sample size, with this label horizon. The
# cell below uses it rather than a table of conventions: it reports the smallest mean IC
# this series could distinguish from zero at the stated confidence, and then sweeps a
# range of candidate magnitudes through the same test.
#
# Net P&L after costs is a separate calculation, handled by the break-even analysis in
# `05_signal_evaluation` and by the case-study cost models in Chapters 16 to 18. A factor
# can clear the test below and still lose money.
# %% tags=["results"]
IC_CANDIDATES = (0.01, 0.02, 0.03, 0.05, 0.10)
CONFIDENCE = 0.95
z_crit = stats.norm.ppf(0.5 + CONFIDENCE / 2)
hac_se = hac_result["hac_se"]
min_detectable_ic = z_crit * hac_se
print(f"Sample: {len(ic_series):,} daily IC values, label horizon {LABEL_HORIZON}")
print(f"HAC standard error: {hac_se:.4f} (lag truncation {hac_result['effective_lags']})")
print(f"Smallest |mean IC| separable from zero at {CONFIDENCE:.0%}: {min_detectable_ic:.4f}")
print()
print(f"{'candidate IC':>13}{'t vs 0':>9}{'separable':>11}")
print("-" * 33)
for candidate in IC_CANDIDATES:
t_stat = candidate / hac_se
print(f"{candidate:>13.2f}{t_stat:>9.2f}{'yes' if t_stat > z_crit else 'no':>11}")
print()
print(f"This series' observed mean IC: {hac_result['mean_ic']:.4f} (t={hac_result['t_stat']:.2f})")
# %%
# Practical significance analysis
print("=== Practical Significance Analysis ===\n")
observed_ic = np.mean(ic_series)
print(f"Observed IC: {observed_ic:.4f}")
print(f"HAC t-stat: {hac_result['t_stat']:.2f}")
print(f"HAC p-value: {hac_result['p_value']:.4f}")
# Cost-adjusted interpretation
# Simplified model: Net IC ≈ IC - cost_factor × turnover
# Assuming ~50% monthly turnover and 10bps round-trip cost
turnover_assumption = 0.5 # 50% monthly turnover
cost_per_trade = 0.001 # 10bps
cost_impact = turnover_assumption * cost_per_trade * 12 # Annualized
print("\nCost assumptions:")
print(f" Monthly turnover: {turnover_assumption:.0%}")
print(f" Round-trip cost: {cost_per_trade * 10000:.0f} bps")
print(f" Annual cost drag: {cost_impact:.2%}")
# Rough IR approximation: IR ≈ IC × √(252) × some multiplier
# This is a simplified Grinold approximation
ir_approx = observed_ic * np.sqrt(252)
print(f"\nGrinold IR approximation (raw): {ir_approx:.2f}")
print(" (This is a raw, pre-cost upper bound - net IR depends on turnover and costs;")
print(" the break-even cost analysis in `05_signal_evaluation` evaluates feasibility.)")
# %% [markdown]
# ## Track Record Planning
#
# How many observations do we need to be confident in our IC estimate?
#
# Using HAC-adjusted effective sample size:
#
# $$T_{required} = \left(\frac{z_{\alpha/2} \times \sigma_{IC}}{IC_{target}}\right)^2 \times \text{HAC inflation}^2$$
# %%
def min_track_record_ic(
target_ic: float,
ic_std: float,
confidence: float = 0.95,
hac_inflation: float = 1.5,
) -> int:
"""
Minimum track record to confirm IC with given confidence.
Uses HAC-adjusted effective sample size.
Parameters
----------
target_ic : float
Minimum IC we want to detect
ic_std : float
Standard deviation of IC series
confidence : float
Required confidence level
hac_inflation : float
HAC SE inflation factor, measured as ``hac_se / naive_se`` on the
IC series in hand rather than assumed
Returns
-------
int : minimum periods needed (actual, not effective)
"""
z = stats.norm.ppf((1 + confidence) / 2)
# Base formula without HAC
base_n = (z * ic_std / target_ic) ** 2
# Adjust for autocorrelation
actual_n = base_n * (hac_inflation**2)
return int(np.ceil(actual_n))
# %%
# Track record requirements
ic_std_observed = np.std(ic_series)
hac_inflation = hac_result["hac_se"] / naive_se
print("=== Minimum Track Record for IC Detection ===\n")
print(f"Observed IC std: {ic_std_observed:.4f}")
print(f"HAC inflation factor: {hac_inflation:.2f}x")
print()
target_ics = [0.01, 0.02, 0.03, 0.05]
track_record_rows = []
for target in target_ics:
min_90 = min_track_record_ic(target, ic_std_observed, 0.90, hac_inflation)
min_95 = min_track_record_ic(target, ic_std_observed, 0.95, hac_inflation)
min_99 = min_track_record_ic(target, ic_std_observed, 0.99, hac_inflation)
track_record_rows.append(
{
"target_ic": target,
"days_90pct": min_90,
"years_90pct": round(min_90 / 252, 1),
"days_95pct": min_95,
"years_95pct": round(min_95 / 252, 1),
"days_99pct": min_99,
"years_99pct": round(min_99 / 252, 1),
}
)
track_record_df = pl.DataFrame(track_record_rows)
print("Days of cross-sectional IC observations needed at each confidence level:")
display(track_record_df)
# %% [markdown]
# ### Interpretation
#
# The track record requirements are substantial because:
#
# 1. **IC is noisy**: daily cross-sectional IC has high variance.
# 2. **Autocorrelation reduces the effective sample**: overlapping forward returns inflate
# the standard error by the factor printed above.
# 3. **Small effects need large samples**: the smallest separable mean IC printed earlier
# is the bar this series sets, and it moves with the sample and the dependence rather
# than sitting at a conventional number.
#
# **Implication**: a "predictive" factor claimed from one or two years of data deserves
# skepticism. The test to apply is its own standard error, not a threshold read off a
# table - which is what the cell above demonstrates on this series.
# %% [markdown]
# ## IC Inference Report
#
# Export a structured report for downstream use.
# %%
# Build IC inference report
inference_report = {
"signal_name": f"momentum_{LABEL_HORIZON}d",
"n_observations": len(ic_series),
"mean_ic": round(float(np.mean(ic_series)), 4),
"ic_std": round(float(np.std(ic_series)), 4),
"inference": {
"naive": {
"se": round(float(naive_se), 4),
"t_stat": round(float(naive_t), 2),
"p_value": round(float(2 * stats.t.sf(abs(naive_t), len(ic_series) - 1)), 4),
},
"hac": {
"se": round(float(hac_result["hac_se"]), 4),
"t_stat": round(float(hac_result["t_stat"]), 2),
"p_value": round(float(hac_result["p_value"]), 4),
"inflation_factor": round(float(hac_inflation), 2),
},
"block_bootstrap": {
"method": block_result["method"],
"block_length": block_result["block_length"],
"n_replications": n_boot,
"se": round(float(block_result["bootstrap_std"]), 4),
"ci_95_lower": round(float(block_result["ci_lower"]), 4),
"ci_95_upper": round(float(block_result["ci_upper"]), 4),
},
},
"effective_sample_size": round(float(effective_n), 0),
"ci_includes_zero": bool(ci_includes_zero),
"autocorrelation_lags_significant": int(n_significant_lags),
}
print("=== IC Inference Report ===\n")
print(json.dumps(inference_report, indent=2))
# %% [markdown]
# ## Summary
#
# ### Key Concepts
#
# | Concept | Description |
# |---------|-------------|
# | **HAC SE** | Newey-West adjustment for autocorrelated IC |
# | **Block Bootstrap** | Preserves dependence structure via block resampling |
# | **Effective N** | True sample size after accounting for autocorrelation |
# | **Track Record** | Minimum observations to detect a given IC |
#
# ### Best Practices
#
# 1. **Always use HAC-adjusted t-statistics** - naive t-stats overstate significance
# 2. **Report block bootstrap CIs** - distribution-free and robust
# 3. **Consider practical significance** - IC must exceed cost drag
# 4. **Plan for adequate track records** - 1-2 years rarely sufficient for small IC
# 5. **Check effective sample size** - may be much smaller than actual observations
#
# ### API Reference
#
# ```python
# from ml4t.diagnostic.metrics import compute_ic_hac_stats
#
# # HAC-adjusted statistics
# hac = compute_ic_hac_stats(ic_series, label_horizon=21)
# print(f"HAC t-stat: {hac['t_stat']:.2f}")
# print(f"HAC p-value: {hac['p_value']:.4f}")
#
# # Block bootstrap (requires arch library)
# from arch.bootstrap import StationaryBootstrap
# bs = StationaryBootstrap(21, ic_series) # 21-day blocks
# ci = bs.conf_int(np.mean, 1000, method='percentile')
# ```
#
# ### Next Notebooks
#
# - [`07_multiple_testing`](07_multiple_testing.ipynb) - FDR control when evaluating many factors
# - [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) - Causal falsification tests
```Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.