コンテンツへスキップ
ライブラリの全資料

HACとブロックブートストラップによるシグナル情報係数の推論

コード Machine Learning for Trading

サマリー

このノートブックでは、連続する観測値に依存関係がある場合、シグナルの情報係数(IC)がノイズと区別できるかを評価します。モメンタムシグナルとETFデータを例に、先渡しリターンの重複、持続性のあるシグナル順位、市場レジームによって日次のIC値に自己相関が生じ得ることを説明します。その場合、単純なt統計量は不確実性を過小評価し、有意性を過大評価する可能性があります。

ノートブックでは、不均一分散と自己相関に頑健なNewey–West標準誤差と、連続した観測値を再標本化するブロックブートストラップを比較します。帯域幅はラベル期間に合わせ、観測された自己相関パターンと照らして確認します。例では、標本サイズだけに基づくルールでは依存関係を見落とす場合が示されています。また、実効標本サイズ、信頼区間、コスト控除後の実務上の有意性、シグナル評価に必要なデータ量も論じています。

この分析は手順を示す例であり、あらゆる場合に使える帯域幅の処方箋ではありません。重複だけでICの系列依存が必ず生じるわけではなく、観測された依存関係は順位の持続性やレジームを反映している可能性もあります。統計的有意性や実務上の有意性とともに、依存関係を考慮した不確実性を報告するよう勧めています。

主なアイデア

  • IC系列の自己相関により、独立同分布を仮定したt統計量はシグナルの根拠を過大評価する可能性があります。
  • Newey–West標準誤差は、選択したラグ帯域幅内の加重自己共分散を使って不確実性を調整します。
  • ラベル期間を参考に帯域幅を決め、観測された依存パターンとの整合性も評価します。
  • ブロックブートストラップはIC系列の連続区間を再標本化し、局所的な依存関係を保ちます。
  • 統計的根拠が実務に役立つかを判断する際は、実効標本サイズと取引コストが重要です。

タグ

全文
# 06_ic_inference.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # IC Inference: HAC Adjustment and Block Bootstrap
#
# **Docker image**: `ml4t`
#
# **Chapter 7: Defining the Learning Task**
# **Section Reference**: 7.3 - Feature and Label Evaluation as Triage
#
# ## Purpose
#
# This notebook addresses **statistical inference for IC**: how confident should we
# be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated
# IC series and block bootstrap for robust confidence intervals.
#
# ## Learning Objectives
#
# 1. Understand why naive t-statistics fail for IC inference
# 2. Apply HAC (Newey-West) adjustment for proper standard errors
# 3. Implement block bootstrap for distribution-free confidence intervals
# 4. Distinguish practical vs statistical significance
# 5. Plan track records using HAC-adjusted effective sample size
#
# ## Data Policy
#
# Examples use **real ETF data** from the case study store.
#
# ## Prerequisites
#
# - `05_signal_evaluation` - produces the IC time series whose inference we
#   formalize here.
# - Familiarity with autocorrelation, the Newey-West HAC estimator, and
#   block bootstrap.

# %%
"""IC Inference - statistical testing and confidence intervals for information coefficients."""

from __future__ import annotations

import json
import logging
import warnings

import numpy as np
import plotly.graph_objects as go
import polars as pl
from arch.bootstrap import StationaryBootstrap
from IPython.display import display
from ml4t.diagnostic.evaluation.autocorrelation import analyze_autocorrelation
from ml4t.diagnostic.metrics import compute_ic_hac_stats
from ml4t.diagnostic.signal import analyze_signal
from plotly.subplots import make_subplots
from scipy import stats

from data import load_etfs
from utils.reproducibility import set_global_seeds
from utils.style import (  # importing utils.style activates the ml4t Plotly template
    COLORS,
    show_plotly_with_alt,
)

# Quiet ml4t library INFO logging. Child loggers set their own level, so raising the
# parent alone does not silence them - set every already-created ml4t logger explicitly.
for _lg in list(logging.root.manager.loggerDict):
    if _lg.startswith("ml4t"):
        logging.getLogger(_lg).setLevel(logging.WARNING)

# %% tags=["parameters"]
SEED = 42
START_DATE = ""
N_BOOT = 2000
# Forward-return horizon, in trading days. Every dependence-aware choice in this
# notebook is derived from it: the momentum lookback, the IC evaluation period,
# the HAC lag truncation, and the bootstrap block length.
LABEL_HORIZON = 21

# %%
set_global_seeds(SEED)


# %% [markdown]
# ## The Autocorrelation Problem
#
# IC time series exhibit **autocorrelation** due to:
#
# 1. **Overlapping forward returns** at longer horizons
# 2. **Persistent signal values** (momentum doesn't flip daily)
# 3. **Regime persistence** (markets stay in trends/ranges)
#
# Naive t-statistics assume iid observations and **underestimate standard errors**,
# leading to inflated significance.

# %%
# Load real ETF data and compute momentum signal
etfs = load_etfs()

if START_DATE:
    etfs = etfs.filter(pl.col("timestamp") >= pl.lit(START_DATE).str.to_date())

# Compute momentum over the label horizon
factor_df = (
    etfs.sort(["symbol", "timestamp"])
    .with_columns(
        [
            (pl.col("close") / pl.col("close").shift(LABEL_HORIZON).over("symbol") - 1).alias(
                "factor"
            )
        ]
    )
    .filter(pl.col("factor").is_not_null())
    .select(["timestamp", "symbol", "factor"])
)

prices_df = etfs.select(["timestamp", "symbol", "close"]).rename({"close": "price"})

# Run signal analysis to get IC series
result = analyze_signal(
    factor_df,
    prices_df,
    periods=(LABEL_HORIZON,),
    quantiles=5,
    ic_method="spearman",
    date_col="timestamp",
    asset_col="symbol",
)

ic_series = np.array(result.ic_series.get(f"{LABEL_HORIZON}D", []))
print(f"IC series length: {len(ic_series)}")
print(f"Mean IC: {np.mean(ic_series):.4f}")

# %% [markdown]
# ### Autocorrelation in IC Series
#
# Let's examine the autocorrelation structure directly.


# %%
# Compute autocorrelation function
def compute_acf(series, nlags=20):
    """Compute autocorrelation function."""
    n = len(series)
    series = series - np.mean(series)
    acf = []
    for lag in range(nlags + 1):
        if lag == 0:
            acf.append(1.0)
        else:
            acf.append(np.corrcoef(series[lag:], series[:-lag])[0, 1])
    return np.array(acf)


n_acf_lags = LABEL_HORIZON - 1
acf_values = compute_acf(ic_series, nlags=n_acf_lags)

# Significance bounds for white noise
n = len(ic_series)
sig_bound = 1.96 / np.sqrt(n)

fig = go.Figure()

n_significant_lags = sum(abs(acf_values[1:]) > sig_bound)

fig.add_trace(
    go.Bar(
        x=list(range(1, n_acf_lags + 1)),
        y=acf_values[1:],
        name="ACF",
        marker_color=[
            COLORS["amber"] if abs(v) > sig_bound else COLORS["blue"] for v in acf_values[1:]
        ],
    )
)

fig.add_hline(
    y=sig_bound, line_dash="dash", line_color=COLORS["negative"], annotation_text="95% band"
)
fig.add_hline(y=-sig_bound, line_dash="dash", line_color=COLORS["negative"])
fig.add_hline(y=0, line_color=COLORS["neutral"])

fig.update_layout(
    title=f"Autocorrelation of the daily IC series, lags 1 to {LABEL_HORIZON - 1}",
    xaxis_title="Lag (days)",
    yaxis_title="Autocorrelation",
    template="ml4t",
    height=350,
)
show_plotly_with_alt(
    fig,
    alt=(
        "A bar chart of the autocorrelation of the daily IC series against lag, from one "
        "day out to one day short of the label horizon, with dashed red lines marking the "
        "white-noise band either side of zero. The bars start just above 0.8 at lag one "
        "and fall away in a smooth, almost linear decline, reaching the band around lag "
        "sixteen. Every bar up to that point is amber, marking it outside the band; only "
        "the last few lags are inside it and drawn in navy."
    ),
)
print(f"\nSignificant autocorrelation lags: {n_significant_lags}/{n_acf_lags}")

# %% [markdown]
# The decline is smooth and reaches the white-noise band near the label horizon, which is
# consistent with overlap driving it: two IC values computed $k$ days apart share $h - k$
# days of return, so the shared window shrinks as $k$ grows. It is consistent with
# persistent signal ranks as well, and the two are not separated here. Nor can this figure
# say what happens *past* the horizon - it stops at $h-1$, and a series of rank
# correlations can stay dependent beyond the overlap through regimes or a slow-moving
# signal. What the figure does establish is enough for the section's purpose: a naive
# t-statistic treats these values as independent draws, and they are not.

# %% [markdown]
# ### Library Autocorrelation Analysis
#
# The library version includes PACF and the **Ljung-Box portmanteau test**,
# which tests the joint null that all autocorrelations up to lag $L$ are zero.

# %%
acf_analysis = analyze_autocorrelation(ic_series, max_lags=LABEL_HORIZON - 1, alpha=0.05)

print("=== ml4t-diagnostic Autocorrelation Analysis ===\n")
print(f"Significant ACF lags:  {acf_analysis.significant_acf_lags}")
print(f"Significant PACF lags: {acf_analysis.significant_pacf_lags}")
print(f"Is white noise:        {acf_analysis.is_white_noise}")
print(f"Suggested ARIMA order: {acf_analysis.suggested_arima_order}")

# %% [markdown]
# ## HAC (Newey-West) Adjustment
#
# HAC (Heteroskedasticity and Autocorrelation Consistent) standard errors account
# for serial correlation in the IC series.
#
# The Newey-West estimator uses a weighted sum of autocovariances:
#
# $$\hat{\sigma}^2_{HAC} = \hat{\gamma}_0 + 2\sum_{j=1}^{L} w_j \hat{\gamma}_j$$
#
# Where $w_j = 1 - j/(L+1)$ (Bartlett kernel) and $L$ is the lag truncation.
#
# **$L$ should not be left to the sample size here.** The usual automatic rule
# $L = \lfloor 4(T/100)^{2/9} \rfloor$ reads only $T$, and on this series it
# returns 9. The labels are $h$-day *overlapping* forward returns, so date $t$ and
# date $t+j$ share part of their outcome window for every $j < h$ - and the signal
# is 21-day momentum, whose cross-sectional ranks change slowly. Overlapping
# outcomes measured against persistent ranks give ICs that move together over
# roughly the label horizon.
#
# **Both halves of that sentence are doing work, and the overlap alone is not
# enough.** The IC is a rank correlation, not a return: if the signal ranks were
# redrawn independently every day, the ICs could be serially uncorrelated no matter
# how much the return windows overlap. So $h-1$ is not a theorem about overlapping
# labels - it is a **conservative horizon-aware bandwidth for a persistent signal**,
# and the evidence that it is the right order of magnitude is the measured ACF in
# the figure above, which stays outside the white-noise band until about lag 16.
#
# What is not defensible is $L = 9$, chosen by a rule that never looked at the
# labels or the signal. Passing `label_horizon` makes the library take
# $L = \max(h-1, \text{auto rule})$; the ACF is what justifies that choice, and the
# next section measures what it is worth.

# %% [markdown]
# ### What the Bandwidth Is Worth
#
# Worth measuring rather than asserting, so the cell below runs the estimator
# both ways on the same IC series. The block bootstrap below resamples contiguous blocks
# of length $h$, so it already respects the overlap. If HAC is given a bandwidth that does
# not, the two will not agree - and the gap is then a property of the bandwidth, not of the
# estimator families, which is what the side-by-side is there to establish. The bootstrap
# section shows the two agreeing once both are horizon-aware.
#
# Calling the estimator without `label_horizon` raises a `UserWarning` saying the automatic
# bandwidth may be anti-conservative for overlapping labels. That warning is the library
# doing its job, and the call below triggers it on purpose; it is silenced for that one
# call so it does not read as a defect in the run.

# %%
# Compute HAC-adjusted statistics
hac_result = compute_ic_hac_stats(ic_series, label_horizon=LABEL_HORIZON)

# The same estimator with the bandwidth left to the sample-size rule. Omitting
# label_horizon is deliberate here and the library warns about it, correctly; the warning
# is silenced for this one call, by category and message, so it does not read as a defect.
with warnings.catch_warnings():
    warnings.filterwarnings(
        "ignore",
        category=UserWarning,
        message="label_horizon was not provided.*",
    )
    hac_auto = compute_ic_hac_stats(ic_series)
print("=== Bandwidth: sample-size rule vs label-horizon aware ===\n")
print(
    f"  auto rule            L={hac_auto['effective_lags']:>3}  "
    f"SE={hac_auto['hac_se']:.6f}  t={hac_auto['t_stat']:.4f}  p={hac_auto['p_value']:.4f}"
)
print(
    f"  label_horizon={LABEL_HORIZON:<8} L={hac_result['effective_lags']:>3}  "
    f"SE={hac_result['hac_se']:.6f}  t={hac_result['t_stat']:.4f}  p={hac_result['p_value']:.4f}"
)
print(f"\n  SE ratio (horizon-aware / auto): {hac_result['hac_se'] / hac_auto['hac_se']:.3f}x")

# Compare naive vs HAC
naive_se = np.std(ic_series, ddof=1) / np.sqrt(len(ic_series))
naive_t = np.mean(ic_series) / naive_se

print("=== Naive vs HAC Inference ===\n")
print(
    f"HAC lag truncation: {hac_result['effective_lags']} "
    f"(horizon-aware choice: label horizon {LABEL_HORIZON} -> L = {LABEL_HORIZON - 1}, "
    f"consistent with the measured ACF above)"
)
print(f"Mean IC: {hac_result['mean_ic']:.4f}")
print("\nStandard Errors:")
print(f"  Naive SE:    {naive_se:.4f}")
print(f"  HAC SE:      {hac_result['hac_se']:.4f}")
print(f"  Inflation:   {hac_result['hac_se'] / naive_se:.2f}x")

print("\nt-Statistics:")
print(f"  Naive t:     {naive_t:.2f}")
print(f"  HAC t:       {hac_result['t_stat']:.2f}")

print("\np-Values:")
print(f"  Naive p:     {2 * stats.t.sf(abs(naive_t), len(ic_series) - 1):.4f}")
print(f"  HAC p:       {hac_result['p_value']:.4f}")

# Significance conclusion
alpha = 0.05
naive_sig = "Yes" if abs(naive_t) > stats.t.ppf(1 - alpha / 2, len(ic_series) - 1) else "No"
hac_sig = "Yes" if hac_result["p_value"] < alpha else "No"

print("\nSignificant at α=0.05?")
print(f"  Naive:       {naive_sig}")
print(f"  HAC:         {hac_sig}")

if naive_sig == "Yes" and hac_sig == "No":
    print("\n  WARNING: Naive test finds significance that HAC adjustment removes!")

# %% [markdown]
# ### Effective Sample Size
#
# The HAC SE inflation factor tells us how much the autocorrelation reduces our
# effective sample size:
#
# $$T_{eff} = T \times \left(\frac{\sigma_{naive}}{\sigma_{HAC}}\right)^2$$

# %%
# Compute effective sample size
inflation_factor = hac_result["hac_se"] / naive_se
effective_n = len(ic_series) / (inflation_factor**2)

print("=== Effective Sample Size ===\n")
print(f"Actual observations: {len(ic_series)}")
print(f"HAC inflation factor: {inflation_factor:.2f}x")
print(f"Effective sample size: {effective_n:.0f}")
print(f"Efficiency loss: {(1 - effective_n / len(ic_series)) * 100:.0f}%")

# %% [markdown]
# ## Block Bootstrap for Robust CI
#
# When IC series has autocorrelation, **iid bootstrap is invalid**. We use
# **block bootstrap** which preserves the dependence structure by resampling
# contiguous blocks of observations.
#
# Two common approaches:
# 1. **Moving Block Bootstrap (MBB)**: Fixed block length
# 2. **Stationary Bootstrap**: Random block lengths (geometric distribution)
#
# We use the `arch` library's `StationaryBootstrap` for robust inference.


# %%
# Block bootstrap implementation
def block_bootstrap_ic_ci(
    ic_series: np.ndarray,
    n_bootstrap: int = 1000,
    block_length: int = 21,
    alpha: float = 0.05,
    random_state: int = 42,
) -> dict:
    """
    Compute block bootstrap confidence interval for mean IC.

    Uses the Politis-Romano stationary bootstrap from `arch.bootstrap`:
    random block lengths drawn from a geometric distribution with
    expected length equal to ``block_length``.

    Parameters
    ----------
    ic_series : array
        Time series of IC values
    n_bootstrap : int
        Number of bootstrap replications
    block_length : int
        Expected block length (mean of the geometric distribution)
    alpha : float
        Significance level (default 0.05 for 95% CI)
    random_state : int
        Random seed for reproducibility

    Returns
    -------
    dict with mean, CI bounds, and bootstrap distribution
    """
    ic_series = np.array(ic_series)

    bs = StationaryBootstrap(block_length, ic_series, seed=random_state)
    bootstrap_means = np.array([np.mean(data[0][0]) for data in bs.bootstrap(n_bootstrap)])

    # Percentile confidence interval
    ci_lower = np.percentile(bootstrap_means, 100 * alpha / 2)
    ci_upper = np.percentile(bootstrap_means, 100 * (1 - alpha / 2))

    return {
        "mean": np.mean(ic_series),
        "ci_lower": ci_lower,
        "ci_upper": ci_upper,
        "bootstrap_std": np.std(bootstrap_means),
        "bootstrap_means": bootstrap_means,
        "method": "Stationary Bootstrap",
        "block_length": block_length,
    }


# %%
# Run block bootstrap
n_boot = N_BOOT
block_result = block_bootstrap_ic_ci(
    ic_series, n_bootstrap=n_boot, block_length=LABEL_HORIZON, random_state=SEED
)

# Compare with HAC CI
hac_ci_lower = hac_result["mean_ic"] - 1.96 * hac_result["hac_se"]
hac_ci_upper = hac_result["mean_ic"] + 1.96 * hac_result["hac_se"]

print("=== Block Bootstrap vs HAC Confidence Intervals ===\n")
print(f"Method: {block_result['method']} (block length={block_result['block_length']})")
print(f"Bootstrap replications: {n_boot}")
print(f"\nMean IC: {block_result['mean']:.4f}")
print("\n95% Confidence Intervals:")
print(f"  Block Bootstrap: [{block_result['ci_lower']:.4f}, {block_result['ci_upper']:.4f}]")
print(f"  HAC (Newey-West): [{hac_ci_lower:.4f}, {hac_ci_upper:.4f}]")
print("\nStandard Errors:")
print(f"  Block Bootstrap: {block_result['bootstrap_std']:.4f}")
print(f"  HAC:             {hac_result['hac_se']:.4f}")

# Does CI include zero?
ci_includes_zero = block_result["ci_lower"] <= 0 <= block_result["ci_upper"]
print(f"\nBlock Bootstrap CI includes zero: {'Yes' if ci_includes_zero else 'No'}")

# %% [markdown]
# ### Library vs Manual Bootstrap: When to Use Each
#
# The `ml4t-diagnostic` library provides `stationary_bootstrap_ic()` for
# **cross-sectional** bootstrap: given arrays of predictions and returns for a
# single date, it resamples asset pairs and recomputes Spearman IC. This is useful
# when you want a confidence interval for a single cross-section's IC.
#
# For **time-series** inference on the IC series itself - "is the mean IC over
# many dates significantly different from zero?" - the correct tool is the block bootstrap
# used above, which preserves temporal dependence.
#
# ```python
# # Cross-sectional bootstrap (library) - CI for one date's IC
# from ml4t.diagnostic.evaluation.stats.hac_standard_errors import stationary_bootstrap_ic
# result = stationary_bootstrap_ic(signals_t, returns_t, n_samples=1000)
#
# # Time-series bootstrap (manual) - CI for mean IC across dates
# block_bootstrap_ic_ci(ic_series, n_bootstrap=2000, block_length=21)
# ```
#
# Throughout this notebook we use the manual block bootstrap because we are
# testing whether the *time-averaged* IC is significantly positive.

# %%
print("=== Block Bootstrap Summary (recommended for IC time series) ===\n")
print(f"  IC: {block_result['mean']:.4f}")
print(f"  Bootstrap SE: {block_result['bootstrap_std']:.4f}")
print(f"  95% CI: [{block_result['ci_lower']:.4f}, {block_result['ci_upper']:.4f}]")
print(f"  HAC CI: [{hac_ci_lower:.4f}, {hac_ci_upper:.4f}]")

# %%
# Visualize bootstrap distribution
fig = make_subplots(
    rows=1,
    cols=2,
    subplot_titles=["Bootstrap distribution of mean IC", "95% CI: block bootstrap vs HAC"],
)

# Bootstrap distribution
fig.add_trace(
    go.Histogram(
        x=block_result["bootstrap_means"],
        nbinsx=50,
        name="Bootstrap means",
        marker_color=COLORS["blue"],
        opacity=0.7,
    ),
    row=1,
    col=1,
)

# Mark CIs
fig.add_vline(
    x=block_result["ci_lower"],
    line_dash="dash",
    line_color=COLORS["negative"],
    annotation_text="2.5%",
    row=1,
    col=1,
)
fig.add_vline(
    x=block_result["ci_upper"],
    line_dash="dash",
    line_color=COLORS["negative"],
    annotation_text="97.5%",
    row=1,
    col=1,
)
fig.add_vline(
    x=0, line_dash="dot", line_color=COLORS["neutral"], annotation_text="Null", row=1, col=1
)

# CI comparison
methods = ["Block Bootstrap", "HAC"]
ci_lowers = [block_result["ci_lower"], hac_ci_lower]
ci_uppers = [block_result["ci_upper"], hac_ci_upper]
means = [block_result["mean"], hac_result["mean_ic"]]

for i, method in enumerate(methods):
    # CI as error bars
    fig.add_trace(
        go.Scatter(
            x=[method],
            y=[means[i]],
            error_y=dict(
                type="data",
                symmetric=False,
                array=[ci_uppers[i] - means[i]],
                arrayminus=[means[i] - ci_lowers[i]],
            ),
            mode="markers",
            marker=dict(size=12, color=[COLORS["blue"], COLORS["amber"]][i]),
            name=method,
        ),
        row=1,
        col=2,
    )

fig.add_hline(y=0, line_dash="dot", line_color=COLORS["neutral"], row=1, col=2)

fig.update_layout(
    title="Mean IC: block-bootstrap distribution and the two interval estimates",
    height=400,
    template="ml4t",
    showlegend=False,
)
fig.update_xaxes(title_text="Mean IC", row=1, col=1)
fig.update_yaxes(title_text="Frequency", row=1, col=1)
fig.update_yaxes(title_text="Mean IC", row=1, col=2)

show_plotly_with_alt(
    fig,
    alt=(
        "Two panels. The left panel is a histogram of the block-bootstrap distribution of "
        "the mean IC, a symmetric bell centred on zero and running out to roughly plus and "
        "minus 0.045, with dashed red lines marking the 2.5th and 97.5th percentiles and a "
        "dotted line at the null, which falls on the mode. "
        "The right panel plots the two interval estimates side by side as a point with "
        "error bars: the block bootstrap in navy and HAC in amber. Both point estimates "
        "sit on the dotted zero line and both intervals span it, reaching roughly the "
        "same distance above and below, so the two methods are hard to tell apart."
    ),
)

# %% [markdown]
# The two intervals are close enough to overlay, which is the result the bandwidth section
# predicted: once HAC is given a lag truncation matched to the label horizon, it agrees
# with a bootstrap that resamples blocks of that same length. They agree about the answer
# as well - both intervals contain zero, so the mean IC is not distinguishable from no
# skill at this sample size, whichever dependence-aware method is used.

# %% [markdown]
# ## Practical vs Statistical Significance
#
# Statistical significance of an IC time-series mean is a separate question from whether
# the signal is economically tradeable.
#
# There is no fixed IC magnitude that counts as detectable. Detectability is set by the
# standard error, and this notebook has already estimated the one that applies here - the
# HAC standard error on this series, at this sample size, with this label horizon. The
# cell below uses it rather than a table of conventions: it reports the smallest mean IC
# this series could distinguish from zero at the stated confidence, and then sweeps a
# range of candidate magnitudes through the same test.
#
# Net P&L after costs is a separate calculation, handled by the break-even analysis in
# `05_signal_evaluation` and by the case-study cost models in Chapters 16 to 18. A factor
# can clear the test below and still lose money.

# %% tags=["results"]
IC_CANDIDATES = (0.01, 0.02, 0.03, 0.05, 0.10)
CONFIDENCE = 0.95

z_crit = stats.norm.ppf(0.5 + CONFIDENCE / 2)
hac_se = hac_result["hac_se"]
min_detectable_ic = z_crit * hac_se

print(f"Sample: {len(ic_series):,} daily IC values, label horizon {LABEL_HORIZON}")
print(f"HAC standard error: {hac_se:.4f}  (lag truncation {hac_result['effective_lags']})")
print(f"Smallest |mean IC| separable from zero at {CONFIDENCE:.0%}: {min_detectable_ic:.4f}")
print()
print(f"{'candidate IC':>13}{'t vs 0':>9}{'separable':>11}")
print("-" * 33)
for candidate in IC_CANDIDATES:
    t_stat = candidate / hac_se
    print(f"{candidate:>13.2f}{t_stat:>9.2f}{'yes' if t_stat > z_crit else 'no':>11}")
print()
print(f"This series' observed mean IC: {hac_result['mean_ic']:.4f} (t={hac_result['t_stat']:.2f})")

# %%
# Practical significance analysis
print("=== Practical Significance Analysis ===\n")

observed_ic = np.mean(ic_series)
print(f"Observed IC: {observed_ic:.4f}")
print(f"HAC t-stat: {hac_result['t_stat']:.2f}")
print(f"HAC p-value: {hac_result['p_value']:.4f}")

# Cost-adjusted interpretation
# Simplified model: Net IC ≈ IC - cost_factor × turnover
# Assuming ~50% monthly turnover and 10bps round-trip cost
turnover_assumption = 0.5  # 50% monthly turnover
cost_per_trade = 0.001  # 10bps
cost_impact = turnover_assumption * cost_per_trade * 12  # Annualized

print("\nCost assumptions:")
print(f"  Monthly turnover: {turnover_assumption:.0%}")
print(f"  Round-trip cost: {cost_per_trade * 10000:.0f} bps")
print(f"  Annual cost drag: {cost_impact:.2%}")

# Rough IR approximation: IR ≈ IC × √(252) × some multiplier
# This is a simplified Grinold approximation
ir_approx = observed_ic * np.sqrt(252)
print(f"\nGrinold IR approximation (raw): {ir_approx:.2f}")
print("  (This is a raw, pre-cost upper bound - net IR depends on turnover and costs;")
print("   the break-even cost analysis in `05_signal_evaluation` evaluates feasibility.)")

# %% [markdown]
# ## Track Record Planning
#
# How many observations do we need to be confident in our IC estimate?
#
# Using HAC-adjusted effective sample size:
#
# $$T_{required} = \left(\frac{z_{\alpha/2} \times \sigma_{IC}}{IC_{target}}\right)^2 \times \text{HAC inflation}^2$$


# %%
def min_track_record_ic(
    target_ic: float,
    ic_std: float,
    confidence: float = 0.95,
    hac_inflation: float = 1.5,
) -> int:
    """
    Minimum track record to confirm IC with given confidence.

    Uses HAC-adjusted effective sample size.

    Parameters
    ----------
    target_ic : float
        Minimum IC we want to detect
    ic_std : float
        Standard deviation of IC series
    confidence : float
        Required confidence level
    hac_inflation : float
        HAC SE inflation factor, measured as ``hac_se / naive_se`` on the
        IC series in hand rather than assumed

    Returns
    -------
    int : minimum periods needed (actual, not effective)
    """
    z = stats.norm.ppf((1 + confidence) / 2)

    # Base formula without HAC
    base_n = (z * ic_std / target_ic) ** 2

    # Adjust for autocorrelation
    actual_n = base_n * (hac_inflation**2)

    return int(np.ceil(actual_n))


# %%
# Track record requirements
ic_std_observed = np.std(ic_series)
hac_inflation = hac_result["hac_se"] / naive_se

print("=== Minimum Track Record for IC Detection ===\n")
print(f"Observed IC std: {ic_std_observed:.4f}")
print(f"HAC inflation factor: {hac_inflation:.2f}x")
print()

target_ics = [0.01, 0.02, 0.03, 0.05]
track_record_rows = []
for target in target_ics:
    min_90 = min_track_record_ic(target, ic_std_observed, 0.90, hac_inflation)
    min_95 = min_track_record_ic(target, ic_std_observed, 0.95, hac_inflation)
    min_99 = min_track_record_ic(target, ic_std_observed, 0.99, hac_inflation)
    track_record_rows.append(
        {
            "target_ic": target,
            "days_90pct": min_90,
            "years_90pct": round(min_90 / 252, 1),
            "days_95pct": min_95,
            "years_95pct": round(min_95 / 252, 1),
            "days_99pct": min_99,
            "years_99pct": round(min_99 / 252, 1),
        }
    )

track_record_df = pl.DataFrame(track_record_rows)
print("Days of cross-sectional IC observations needed at each confidence level:")
display(track_record_df)

# %% [markdown]
# ### Interpretation
#
# The track record requirements are substantial because:
#
# 1. **IC is noisy**: daily cross-sectional IC has high variance.
# 2. **Autocorrelation reduces the effective sample**: overlapping forward returns inflate
#    the standard error by the factor printed above.
# 3. **Small effects need large samples**: the smallest separable mean IC printed earlier
#    is the bar this series sets, and it moves with the sample and the dependence rather
#    than sitting at a conventional number.
#
# **Implication**: a "predictive" factor claimed from one or two years of data deserves
# skepticism. The test to apply is its own standard error, not a threshold read off a
# table - which is what the cell above demonstrates on this series.

# %% [markdown]
# ## IC Inference Report
#
# Export a structured report for downstream use.

# %%
# Build IC inference report
inference_report = {
    "signal_name": f"momentum_{LABEL_HORIZON}d",
    "n_observations": len(ic_series),
    "mean_ic": round(float(np.mean(ic_series)), 4),
    "ic_std": round(float(np.std(ic_series)), 4),
    "inference": {
        "naive": {
            "se": round(float(naive_se), 4),
            "t_stat": round(float(naive_t), 2),
            "p_value": round(float(2 * stats.t.sf(abs(naive_t), len(ic_series) - 1)), 4),
        },
        "hac": {
            "se": round(float(hac_result["hac_se"]), 4),
            "t_stat": round(float(hac_result["t_stat"]), 2),
            "p_value": round(float(hac_result["p_value"]), 4),
            "inflation_factor": round(float(hac_inflation), 2),
        },
        "block_bootstrap": {
            "method": block_result["method"],
            "block_length": block_result["block_length"],
            "n_replications": n_boot,
            "se": round(float(block_result["bootstrap_std"]), 4),
            "ci_95_lower": round(float(block_result["ci_lower"]), 4),
            "ci_95_upper": round(float(block_result["ci_upper"]), 4),
        },
    },
    "effective_sample_size": round(float(effective_n), 0),
    "ci_includes_zero": bool(ci_includes_zero),
    "autocorrelation_lags_significant": int(n_significant_lags),
}

print("=== IC Inference Report ===\n")
print(json.dumps(inference_report, indent=2))

# %% [markdown]
# ## Summary
#
# ### Key Concepts
#
# | Concept | Description |
# |---------|-------------|
# | **HAC SE** | Newey-West adjustment for autocorrelated IC |
# | **Block Bootstrap** | Preserves dependence structure via block resampling |
# | **Effective N** | True sample size after accounting for autocorrelation |
# | **Track Record** | Minimum observations to detect a given IC |
#
# ### Best Practices
#
# 1. **Always use HAC-adjusted t-statistics** - naive t-stats overstate significance
# 2. **Report block bootstrap CIs** - distribution-free and robust
# 3. **Consider practical significance** - IC must exceed cost drag
# 4. **Plan for adequate track records** - 1-2 years rarely sufficient for small IC
# 5. **Check effective sample size** - may be much smaller than actual observations
#
# ### API Reference
#
# ```python
# from ml4t.diagnostic.metrics import compute_ic_hac_stats
#
# # HAC-adjusted statistics
# hac = compute_ic_hac_stats(ic_series, label_horizon=21)
# print(f"HAC t-stat: {hac['t_stat']:.2f}")
# print(f"HAC p-value: {hac['p_value']:.4f}")
#
# # Block bootstrap (requires arch library)
# from arch.bootstrap import StationaryBootstrap
# bs = StationaryBootstrap(21, ic_series)  # 21-day blocks
# ci = bs.conf_int(np.mean, 1000, method='percentile')
# ```
#
# ### Next Notebooks
#
# - [`07_multiple_testing`](07_multiple_testing.ipynb) - FDR control when evaluating many factors
# - [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) - Causal falsification tests

```

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。