コンテンツへスキップ
ライブラリの全資料

自己相関する情報係数の推論

ノートブック Machine Learning for Trading

サマリー

このノートブックでは、日次のIC観測値が依存している場合、通常のt検定が情報係数(IC)の強さを過大評価する理由を説明します。ETFデータからモメンタムシグナルを構築し、ICの自己相関を調べ、Newey–Westの不均一分散・自己相関一貫性標準誤差を、信頼区間用のブロックブートストラップと併せて紹介します。ブートストラップは依存性を保つため連続した観測値を再標本化し、HAC推定値は重み付き自己共分散を使って不確実性を調整します。

この例では、サンプルサイズから選んだ帯域幅とラベル期間を考慮した帯域幅を比較し、有効サンプルサイズ、実務上の意義、実績期間の計画について説明します。自己相関の図では、およそラグ16まで依存性が持続しており、重複するリターンウィンドウや持続的なシグナル順位と整合します。このパターンはここでは期間を考慮した帯域幅を保守的な選択として支持しますが、重複があるだけでICの自己相関が保証されるわけではありません。この分析では重複の影響と順位の持続性を区別しておらず、依存性が図示したラグの先まで続く可能性にも注意を促しています。

主なアイデア

  • 素朴なt統計量は観測値の独立性を仮定するため、自己相関のあるIC系列では不確実性を過小評価することがあります。
  • Newey–WestのHAC標準誤差は、選択したラグの重み付き自己共分散を使って推論を調整します。
  • ブロックブートストラップは、IC系列から連続区間を再標本化して局所的な依存性を保ちます。
  • ラベル期間はHACの帯域幅を決める参考になりますが、持続的なシグナルでは実測した自己相関も選択の根拠になります。
  • 統計的有意性は、実務上のコスト負担と有効サンプルサイズも併せて検討します。

タグ

全文
# IC Inference: HAC Adjustment and Block Bootstrap


# IC Inference: HAC Adjustment and Block Bootstrap

**Docker image**: `ml4t`

**Chapter 7: Defining the Learning Task**
**Section Reference**: 7.3 - Feature and Label Evaluation as Triage

## Purpose

This notebook addresses **statistical inference for IC**: how confident should we
be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated
IC series and block bootstrap for robust confidence intervals.

## Learning Objectives

1. Understand why naive t-statistics fail for IC inference
2. Apply HAC (Newey-West) adjustment for proper standard errors
3. Implement block bootstrap for distribution-free confidence intervals
4. Distinguish practical vs statistical significance
5. Plan track records using HAC-adjusted effective sample size

## Data Policy

Examples use **real ETF data** from the case study store.

## Prerequisites

- `05_signal_evaluation` - produces the IC time series whose inference we
  formalize here.
- Familiarity with autocorrelation, the Newey-West HAC estimator, and
  block bootstrap.

```python
"""IC Inference - statistical testing and confidence intervals for information coefficients."""

from __future__ import annotations

import json
import logging
import warnings

import numpy as np
import plotly.graph_objects as go
import polars as pl
from arch.bootstrap import StationaryBootstrap
from IPython.display import display
from ml4t.diagnostic.evaluation.autocorrelation import analyze_autocorrelation
from ml4t.diagnostic.metrics import compute_ic_hac_stats
from ml4t.diagnostic.signal import analyze_signal
from plotly.subplots import make_subplots
from scipy import stats

from data import load_etfs
from utils.reproducibility import set_global_seeds
from utils.style import (  # importing utils.style activates the ml4t Plotly template
    COLORS,
    show_plotly_with_alt,
)

# Quiet ml4t library INFO logging. Child loggers set their own level, so raising the
# parent alone does not silence them - set every already-created ml4t logger explicitly.
for _lg in list(logging.root.manager.loggerDict):
    if _lg.startswith("ml4t"):
        logging.getLogger(_lg).setLevel(logging.WARNING)
```

```python
SEED = 42
START_DATE = ""
N_BOOT = 2000
# Forward-return horizon, in trading days. Every dependence-aware choice in this
# notebook is derived from it: the momentum lookback, the IC evaluation period,
# the HAC lag truncation, and the bootstrap block length.
LABEL_HORIZON = 21
```

```python
set_global_seeds(SEED)
```

## The Autocorrelation Problem

IC time series exhibit **autocorrelation** due to:

1. **Overlapping forward returns** at longer horizons
2. **Persistent signal values** (momentum doesn't flip daily)
3. **Regime persistence** (markets stay in trends/ranges)

Naive t-statistics assume iid observations and **underestimate standard errors**,
leading to inflated significance.

```python
# Load real ETF data and compute momentum signal
etfs = load_etfs()

if START_DATE:
    etfs = etfs.filter(pl.col("timestamp") >= pl.lit(START_DATE).str.to_date())

# Compute momentum over the label horizon
factor_df = (
    etfs.sort(["symbol", "timestamp"])
    .with_columns(
        [
            (pl.col("close") / pl.col("close").shift(LABEL_HORIZON).over("symbol") - 1).alias(
                "factor"
            )
        ]
    )
    .filter(pl.col("factor").is_not_null())
    .select(["timestamp", "symbol", "factor"])
)

prices_df = etfs.select(["timestamp", "symbol", "close"]).rename({"close": "price"})

# Run signal analysis to get IC series
result = analyze_signal(
    factor_df,
    prices_df,
    periods=(LABEL_HORIZON,),
    quantiles=5,
    ic_method="spearman",
    date_col="timestamp",
    asset_col="symbol",
)

ic_series = np.array(result.ic_series.get(f"{LABEL_HORIZON}D", []))
print(f"IC series length: {len(ic_series)}")
print(f"Mean IC: {np.mean(ic_series):.4f}")
```

### Autocorrelation in IC Series

Let's examine the autocorrelation structure directly.

```python
# Compute autocorrelation function
def compute_acf(series, nlags=20):
    """Compute autocorrelation function."""
    n = len(series)
    series = series - np.mean(series)
    acf = []
    for lag in range(nlags + 1):
        if lag == 0:
            acf.append(1.0)
        else:
            acf.append(np.corrcoef(series[lag:], series[:-lag])[0, 1])
    return np.array(acf)


n_acf_lags = LABEL_HORIZON - 1
acf_values = compute_acf(ic_series, nlags=n_acf_lags)

# Significance bounds for white noise
n = len(ic_series)
sig_bound = 1.96 / np.sqrt(n)

fig = go.Figure()

n_significant_lags = sum(abs(acf_values[1:]) > sig_bound)

fig.add_trace(
    go.Bar(
        x=list(range(1, n_acf_lags + 1)),
        y=acf_values[1:],
        name="ACF",
        marker_color=[
            COLORS["amber"] if abs(v) > sig_bound else COLORS["blue"] for v in acf_values[1:]
        ],
    )
)

fig.add_hline(
    y=sig_bound, line_dash="dash", line_color=COLORS["negative"], annotation_text="95% band"
)
fig.add_hline(y=-sig_bound, line_dash="dash", line_color=COLORS["negative"])
fig.add_hline(y=0, line_color=COLORS["neutral"])

fig.update_layout(
    title=f"Autocorrelation of the daily IC series, lags 1 to {LABEL_HORIZON - 1}",
    xaxis_title="Lag (days)",
    yaxis_title="Autocorrelation",
    template="ml4t",
    height=350,
)
show_plotly_with_alt(
    fig,
    alt=(
        "A bar chart of the autocorrelation of the daily IC series against lag, from one "
        "day out to one day short of the label horizon, with dashed red lines marking the "
        "white-noise band either side of zero. The bars start just above 0.8 at lag one "
        "and fall away in a smooth, almost linear decline, reaching the band around lag "
        "sixteen. Every bar up to that point is amber, marking it outside the band; only "
        "the last few lags are inside it and drawn in navy."
    ),
)
print(f"\nSignificant autocorrelation lags: {n_significant_lags}/{n_acf_lags}")
```

The decline is smooth and reaches the white-noise band near the label horizon, which is
consistent with overlap driving it: two IC values computed $k$ days apart share $h - k$
days of return, so the shared window shrinks as $k$ grows. It is consistent with
persistent signal ranks as well, and the two are not separated here. Nor can this figure
say what happens *past* the horizon - it stops at $h-1$, and a series of rank
correlations can stay dependent beyond the overlap through regimes or a slow-moving
signal. What the figure does establish is enough for the section's purpose: a naive
t-statistic treats these values as independent draws, and they are not.

### Library Autocorrelation Analysis

The library version includes PACF and the **Ljung-Box portmanteau test**,
which tests the joint null that all autocorrelations up to lag $L$ are zero.

```python
acf_analysis = analyze_autocorrelation(ic_series, max_lags=LABEL_HORIZON - 1, alpha=0.05)

print("=== ml4t-diagnostic Autocorrelation Analysis ===\n")
print(f"Significant ACF lags:  {acf_analysis.significant_acf_lags}")
print(f"Significant PACF lags: {acf_analysis.significant_pacf_lags}")
print(f"Is white noise:        {acf_analysis.is_white_noise}")
print(f"Suggested ARIMA order: {acf_analysis.suggested_arima_order}")
```

## HAC (Newey-West) Adjustment

HAC (Heteroskedasticity and Autocorrelation Consistent) standard errors account
for serial correlation in the IC series.

The Newey-West estimator uses a weighted sum of autocovariances:

$$\hat{\sigma}^2_{HAC} = \hat{\gamma}_0 + 2\sum_{j=1}^{L} w_j \hat{\gamma}_j$$

Where $w_j = 1 - j/(L+1)$ (Bartlett kernel) and $L$ is the lag truncation.

**$L$ should not be left to the sample size here.** The usual automatic rule
$L = \lfloor 4(T/100)^{2/9} \rfloor$ reads only $T$, and on this series it
returns 9. The labels are $h$-day *overlapping* forward returns, so date $t$ and
date $t+j$ share part of their outcome window for every $j < h$ - and the signal
is 21-day momentum, whose cross-sectional ranks change slowly. Overlapping
outcomes measured against persistent ranks give ICs that move together over
roughly the label horizon.

**Both halves of that sentence are doing work, and the overlap alone is not
enough.** The IC is a rank correlation, not a return: if the signal ranks were
redrawn independently every day, the ICs could be serially uncorrelated no matter
how much the return windows overlap. So $h-1$ is not a theorem about overlapping
labels - it is a **conservative horizon-aware bandwidth for a persistent signal**,
and the evidence that it is the right order of magnitude is the measured ACF in
the figure above, which stays outside the white-noise band until about lag 16.

What is not defensible is $L = 9$, chosen by a rule that never looked at the
labels or the signal. Passing `label_horizon` makes the library take
$L = \max(h-1, \text{auto rule})$; the ACF is what justifies that choice, and the
next section measures what it is worth.

### What the Bandwidth Is Worth

Worth measuring rather than asserting, so the cell below runs the estimator
both ways on the same IC series. The block bootstrap below resamples contiguous blocks
of length $h$, so it already respects the overlap. If HAC is given a bandwidth that does
not, the two will not agree - and the gap is then a property of the bandwidth, not of the
estimator families, which is what the side-by-side is there to establish. The bootstrap
section shows the two agreeing once both are horizon-aware.

Calling the estimator without `label_horizon` raises a `UserWarning` saying the automatic
bandwidth may be anti-conservative for overlapping labels. That warning is the library
doing its job, and the call below triggers it on purpose; it is silenced for that one
call so it does not read as a defect in the run.

```python
# Compute HAC-adjusted statistics
hac_result = compute_ic_hac_stats(ic_series, label_horizon=LABEL_HORIZON)

# The same estimator with the bandwidth left to the sample-size rule. Omitting
# label_horizon is deliberate here and the library warns about it, correctly; the warning
# is silenced for this one call, by category and message, so it does not read as a defect.
with warnings.catch_warnings():
    warnings.filterwarnings(
        "ignore",
        category=UserWarning,
        message="label_horizon was not provided.*",
    )
    hac_auto = compute_ic_hac_stats(ic_series)
print("=== Bandwidth: sample-size rule vs label-horizon aware ===\n")
print(
    f"  auto rule            L={hac_auto['effective_lags']:>3}  "
    f"SE={hac_auto['hac_se']:.6f}  t={hac_auto['t_stat']:.4f}  p={hac_auto['p_value']:.4f}"
)
print(
    f"  label_horizon={LABEL_HORIZON:<8} L={hac_result['effective_lags']:>3}  "
    f"SE={hac_result['hac_se']:.6f}  t={hac_result['t_stat']:.4f}  p={hac_result['p_value']:.4f}"
)
print(f"\n  SE ratio (horizon-aware / auto): {hac_result['hac_se'] / hac_auto['hac_se']:.3f}x")

# Compare naive vs HAC
naive_se = np.std(ic_series, ddof=1) / np.sqrt(len(ic_series))
naive_t = np.mean(ic_series) / naive_se

print("=== Naive vs HAC Inference ===\n")
print(
    f"HAC lag truncation: {hac_result['effective_lags']} "
    f"(horizon-aware choice: label horizon {LABEL_HORIZON} -> L = {LABEL_HORIZON - 1}, "
    f"consistent with the measured ACF above)"
)
print(f"Mean IC: {hac_result['mean_ic']:.4f}")
print("\nStandard Errors:")
print(f"  Naive SE:    {naive_se:.4f}")
print(f"  HAC SE:      {hac_result['hac_se']:.4f}")
print(f"  Inflation:   {hac_result['hac_se'] / naive_se:.2f}x")

print("\nt-Statistics:")
print(f"  Naive t:     {naive_t:.2f}")
print(f"  HAC t:       {hac_result['t_stat']:.2f}")

print("\np-Values:")
print(f"  Naive p:     {2 * stats.t.sf(abs(naive_t), len(ic_series) - 1):.4f}")
print(f"  HAC p:       {hac_result['p_value']:.4f}")

# Significance conclusion
alpha = 0.05
naive_sig = "Yes" if abs(naive_t) > stats.t.ppf(1 - alpha / 2, len(ic_series) - 1) else "No"
hac_sig = "Yes" if hac_result["p_value"] < alpha else "No"

print("\nSignificant at α=0.05?")
print(f"  Naive:       {naive_sig}")
print(f"  HAC:         {hac_sig}")

if naive_sig == "Yes" and hac_sig == "No":
    print("\n  WARNING: Naive test finds significance that HAC adjustment removes!")
```

### Effective Sample Size

The HAC SE inflation factor tells us how much the autocorrelation reduces our
effective sample size:

$$T_{eff} = T \times \left(\frac{\sigma_{naive}}{\sigma_{HAC}}\right)^2$$

```python
# Compute effective sample size
inflation_factor = hac_result["hac_se"] / naive_se
effective_n = len(ic_series) / (inflation_factor**2)

print("=== Effective Sample Size ===\n")
print(f"Actual observations: {len(ic_series)}")
print(f"HAC inflation factor: {inflation_factor:.2f}x")
print(f"Effective sample size: {effective_n:.0f}")
print(f"Efficiency loss: {(1 - effective_n / len(ic_series)) * 100:.0f}%")
```

## Block Bootstrap for Robust CI

When IC series has autocorrelation, **iid bootstrap is invalid**. We use
**block bootstrap** which preserves the dependence structure by resampling
contiguous blocks of observations.

Two common approaches:
1. **Moving Block Bootstrap (MBB)**: Fixed block length
2. **Stationary Bootstrap**: Random block lengths (geometric distribution)

We use the `arch` library's `StationaryBootstrap` for robust inference.

```python
# Block bootstrap implementation
def block_bootstrap_ic_ci(
    ic_series: np.ndarray,
    n_bootstrap: int = 1000,
    block_length: int = 21,
    alpha: float = 0.05,
    random_state: int = 42,
) -> dict:
    """
    Compute block bootstrap confidence interval for mean IC.

    Uses the Politis-Romano stationary bootstrap from `arch.bootstrap`:
    random block lengths drawn from a geometric distribution with
    expected length equal to ``block_length``.

    Parameters
    ----------
    ic_series : array
        Time series of IC values
    n_bootstrap : int
        Number of bootstrap replications
    block_length : int
        Expected block length (mean of the geometric distribution)
    alpha : float
        Significance level (default 0.05 for 95% CI)
    random_state : int
        Random seed for reproducibility

    Returns
    -------
    dict with mean, CI bounds, and bootstrap distribution
    """
    ic_series = np.array(ic_series)

    bs = StationaryBootstrap(block_length, ic_series, seed=random_state)
    bootstrap_means = np.array([np.mean(data[0][0]) for data in bs.bootstrap(n_bootstrap)])

    # Percentile confidence interval
    ci_lower = np.percentile(bootstrap_means, 100 * alpha / 2)
    ci_upper = np.percentile(bootstrap_means, 100 * (1 - alpha / 2))

    return {
        "mean": np.mean(ic_series),
        "ci_lower": ci_lower,
        "ci_upper": ci_upper,
        "bootstrap_std": np.std(bootstrap_means),
        "bootstrap_means": bootstrap_means,
        "method": "Stationary Bootstrap",
        "block_length": block_length,
    }
```

```python
# Run block bootstrap
n_boot = N_BOOT
block_result = block_bootstrap_ic_ci(
    ic_series, n_bootstrap=n_boot, block_length=LABEL_HORIZON, random_state=SEED
)

# Compare with HAC CI
hac_ci_lower = hac_result["mean_ic"] - 1.96 * hac_result["hac_se"]
hac_ci_upper = hac_result["mean_ic"] + 1.96 * hac_result["hac_se"]

print("=== Block Bootstrap vs HAC Confidence Intervals ===\n")
print(f"Method: {block_result['method']} (block length={block_result['block_length']})")
print(f"Bootstrap replications: {n_boot}")
print(f"\nMean IC: {block_result['mean']:.4f}")
print("\n95% Confidence Intervals:")
print(f"  Block Bootstrap: [{block_result['ci_lower']:.4f}, {block_result['ci_upper']:.4f}]")
print(f"  HAC (Newey-West): [{hac_ci_lower:.4f}, {hac_ci_upper:.4f}]")
print("\nStandard Errors:")
print(f"  Block Bootstrap: {block_result['bootstrap_std']:.4f}")
print(f"  HAC:             {hac_result['hac_se']:.4f}")

# Does CI include zero?
ci_includes_zero = block_result["ci_lower"] <= 0 <= block_result["ci_upper"]
print(f"\nBlock Bootstrap CI includes zero: {'Yes' if ci_includes_zero else 'No'}")
```

### Library vs Manual Bootstrap: When to Use Each

The `ml4t-diagnostic` library provides `stationary_bootstrap_ic()` for
**cross-sectional** bootstrap: given arrays of predictions and returns for a
single date, it resamples asset pairs and recomputes Spearman IC. This is useful
when you want a confidence interval for a single cross-section's IC.

For **time-series** inference on the IC series itself - "is the mean IC over
many dates significantly different from zero?" - the correct tool is the block bootstrap
used above, which preserves temporal dependence.

```python
# Cross-sectional bootstrap (library) - CI for one date's IC
from ml4t.diagnostic.evaluation.stats.hac_standard_errors import stationary_bootstrap_ic
result = stationary_bootstrap_ic(signals_t, returns_t, n_samples=1000)

# Time-series bootstrap (manual) - CI for mean IC across dates
block_bootstrap_ic_ci(ic_series, n_bootstrap=2000, block_length=21)
```

Throughout this notebook we use the manual block bootstrap because we are
testing whether the *time-averaged* IC is significantly positive.

```python
print("=== Block Bootstrap Summary (recommended for IC time series) ===\n")
print(f"  IC: {block_result['mean']:.4f}")
print(f"  Bootstrap SE: {block_result['bootstrap_std']:.4f}")
print(f"  95% CI: [{block_result['ci_lower']:.4f}, {block_result['ci_upper']:.4f}]")
print(f"  HAC CI: [{hac_ci_lower:.4f}, {hac_ci_upper:.4f}]")
```

```python
# Visualize bootstrap distribution
fig = make_subplots(
    rows=1,
    cols=2,
    subplot_titles=["Bootstrap distribution of mean IC", "95% CI: block bootstrap vs HAC"],
)

# Bootstrap distribution
fig.add_trace(
    go.Histogram(
        x=block_result["bootstrap_means"],
        nbinsx=50,
        name="Bootstrap means",
        marker_color=COLORS["blue"],
        opacity=0.7,
    ),
    row=1,
    col=1,
)

# Mark CIs
fig.add_vline(
    x=block_result["ci_lower"],
    line_dash="dash",
    line_color=COLORS["negative"],
    annotation_text="2.5%",
    row=1,
    col=1,
)
fig.add_vline(
    x=block_result["ci_upper"],
    line_dash="dash",
    line_color=COLORS["negative"],
    annotation_text="97.5%",
    row=1,
    col=1,
)
fig.add_vline(
    x=0, line_dash="dot", line_color=COLORS["neutral"], annotation_text="Null", row=1, col=1
)

# CI comparison
methods = ["Block Bootstrap", "HAC"]
ci_lowers = [block_result["ci_lower"], hac_ci_lower]
ci_uppers = [block_result["ci_upper"], hac_ci_upper]
means = [block_result["mean"], hac_result["mean_ic"]]

for i, method in enumerate(methods):
    # CI as error bars
    fig.add_trace(
        go.Scatter(
            x=[method],
            y=[means[i]],
            error_y=dict(
                type="data",
                symmetric=False,
                array=[ci_uppers[i] - means[i]],
                arrayminus=[means[i] - ci_lowers[i]],
            ),
            mode="markers",
            marker=dict(size=12, color=[COLORS["blue"], COLORS["amber"]][i]),
            name=method,
        ),
        row=1,
        col=2,
    )

fig.add_hline(y=0, line_dash="dot", line_color=COLORS["neutral"], row=1, col=2)

fig.update_layout(
    title="Mean IC: block-bootstrap distribution and the two interval estimates",
    height=400,
    template="ml4t",
    showlegend=False,
)
fig.update_xaxes(title_text="Mean IC", row=1, col=1)
fig.update_yaxes(title_text="Frequency", row=1, col=1)
fig.update_yaxes(title_text="Mean IC", row=1, col=2)

show_plotly_with_alt(
    fig,
    alt=(
        "Two panels. The left panel is a histogram of the block-bootstrap distribution of "
        "the mean IC, a symmetric bell centred on zero and running out to roughly plus and "
        "minus 0.045, with dashed red lines marking the 2.5th and 97.5th percentiles and a "
        "dotted line at the null, which falls on the mode. "
        "The right panel plots the two interval estimates side by side as a point with "
        "error bars: the block bootstrap in navy and HAC in amber. Both point estimates "
        "sit on the dotted zero line and both intervals span it, reaching roughly the "
        "same distance above and below, so the two methods are hard to tell apart."
    ),
)
```

The two intervals are close enough to overlay, which is the result the bandwidth section
predicted: once HAC is given a lag truncation matched to the label horizon, it agrees
with a bootstrap that resamples blocks of that same length. They agree about the answer
as well - both intervals contain zero, so the mean IC is not distinguishable from no
skill at this sample size, whichever dependence-aware method is used.

## Practical vs Statistical Significance

Statistical significance of an IC time-series mean is a separate question from whether
the signal is economically tradeable.

There is no fixed IC magnitude that counts as detectable. Detectability is set by the
standard error, and this notebook has already estimated the one that applies here - the
HAC standard error on this series, at this sample size, with this label horizon. The
cell below uses it rather than a table of conventions: it reports the smallest mean IC
this series could distinguish from zero at the stated confidence, and then sweeps a
range of candidate magnitudes through the same test.

Net P&L after costs is a separate calculation, handled by the break-even analysis in
`05_signal_evaluation` and by the case-study cost models in Chapters 16 to 18. A factor
can clear the test below and still lose money.

```python
IC_CANDIDATES = (0.01, 0.02, 0.03, 0.05, 0.10)
CONFIDENCE = 0.95

z_crit = stats.norm.ppf(0.5 + CONFIDENCE / 2)
hac_se = hac_result["hac_se"]
min_detectable_ic = z_crit * hac_se

print(f"Sample: {len(ic_series):,} daily IC values, label horizon {LABEL_HORIZON}")
print(f"HAC standard error: {hac_se:.4f}  (lag truncation {hac_result['effective_lags']})")
print(f"Smallest |mean IC| separable from zero at {CONFIDENCE:.0%}: {min_detectable_ic:.4f}")
print()
print(f"{'candidate IC':>13}{'t vs 0':>9}{'separable':>11}")
print("-" * 33)
for candidate in IC_CANDIDATES:
    t_stat = candidate / hac_se
    print(f"{candidate:>13.2f}{t_stat:>9.2f}{'yes' if t_stat > z_crit else 'no':>11}")
print()
print(f"This series' observed mean IC: {hac_result['mean_ic']:.4f} (t={hac_result['t_stat']:.2f})")
```

```python
# Practical significance analysis
print("=== Practical Significance Analysis ===\n")

observed_ic = np.mean(ic_series)
print(f"Observed IC: {observed_ic:.4f}")
print(f"HAC t-stat: {hac_result['t_stat']:.2f}")
print(f"HAC p-value: {hac_result['p_value']:.4f}")

# Cost-adjusted interpretation
# Simplified model: Net IC ≈ IC - cost_factor × turnover
# Assuming ~50% monthly turnover and 10bps round-trip cost
turnover_assumption = 0.5  # 50% monthly turnover
cost_per_trade = 0.001  # 10bps
cost_impact = turnover_assumption * cost_per_trade * 12  # Annualized

print("\nCost assumptions:")
print(f"  Monthly turnover: {turnover_assumption:.0%}")
print(f"  Round-trip cost: {cost_per_trade * 10000:.0f} bps")
print(f"  Annual cost drag: {cost_impact:.2%}")

# Rough IR approximation: IR ≈ IC × √(252) × some multiplier
# This is a simplified Grinold approximation
ir_approx = observed_ic * np.sqrt(252)
print(f"\nGrinold IR approximation (raw): {ir_approx:.2f}")
print("  (This is a raw, pre-cost upper bound - net IR depends on turnover and costs;")
print("   the break-even cost analysis in `05_signal_evaluation` evaluates feasibility.)")
```

## Track Record Planning

How many observations do we need to be confident in our IC estimate?

Using HAC-adjusted effective sample size:

$$T_{required} = \left(\frac{z_{\alpha/2} \times \sigma_{IC}}{IC_{target}}\right)^2 \times \text{HAC inflation}^2$$

```python
def min_track_record_ic(
    target_ic: float,
    ic_std: float,
    confidence: float = 0.95,
    hac_inflation: float = 1.5,
) -> int:
    """
    Minimum track record to confirm IC with given confidence.

    Uses HAC-adjusted effective sample size.

    Parameters
    ----------
    target_ic : float
        Minimum IC we want to detect
    ic_std : float
        Standard deviation of IC series
    confidence : float
        Required confidence level
    hac_inflation : float
        HAC SE inflation factor, measured as ``hac_se / naive_se`` on the
        IC series in hand rather than assumed

    Returns
    -------
    int : minimum periods needed (actual, not effective)
    """
    z = stats.norm.ppf((1 + confidence) / 2)

    # Base formula without HAC
    base_n = (z * ic_std / target_ic) ** 2

    # Adjust for autocorrelation
    actual_n = base_n * (hac_inflation**2)

    return int(np.ceil(actual_n))
```

```python
# Track record requirements
ic_std_observed = np.std(ic_series)
hac_inflation = hac_result["hac_se"] / naive_se

print("=== Minimum Track Record for IC Detection ===\n")
print(f"Observed IC std: {ic_std_observed:.4f}")
print(f"HAC inflation factor: {hac_inflation:.2f}x")
print()

target_ics = [0.01, 0.02, 0.03, 0.05]
track_record_rows = []
for target in target_ics:
    min_90 = min_track_record_ic(target, ic_std_observed, 0.90, hac_inflation)
    min_95 = min_track_record_ic(target, ic_std_observed, 0.95, hac_inflation)
    min_99 = min_track_record_ic(target, ic_std_observed, 0.99, hac_inflation)
    track_record_rows.append(
        {
            "target_ic": target,
            "days_90pct": min_90,
            "years_90pct": round(min_90 / 252, 1),
            "days_95pct": min_95,
            "years_95pct": round(min_95 / 252, 1),
            "days_99pct": min_99,
            "years_99pct": round(min_99 / 252, 1),
        }
    )

track_record_df = pl.DataFrame(track_record_rows)
print("Days of cross-sectional IC observations needed at each confidence level:")
display(track_record_df)
```

### Interpretation

The track record requirements are substantial because:

1. **IC is noisy**: daily cross-sectional IC has high variance.
2. **Autocorrelation reduces the effective sample**: overlapping forward returns inflate
   the standard error by the factor printed above.
3. **Small effects need large samples**: the smallest separable mean IC printed earlier
   is the bar this series sets, and it moves with the sample and the dependence rather
   than sitting at a conventional number.

**Implication**: a "predictive" factor claimed from one or two years of data deserves
skepticism. The test to apply is its own standard error, not a threshold read off a
table - which is what the cell above demonstrates on this series.

## IC Inference Report

Export a structured report for downstream use.

```python
# Build IC inference report
inference_report = {
    "signal_name": f"momentum_{LABEL_HORIZON}d",
    "n_observations": len(ic_series),
    "mean_ic": round(float(np.mean(ic_series)), 4),
    "ic_std": round(float(np.std(ic_series)), 4),
    "inference": {
        "naive": {
            "se": round(float(naive_se), 4),
            "t_stat": round(float(naive_t), 2),
            "p_value": round(float(2 * stats.t.sf(abs(naive_t), len(ic_series) - 1)), 4),
        },
        "hac": {
            "se": round(float(hac_result["hac_se"]), 4),
            "t_stat": round(float(hac_result["t_stat"]), 2),
            "p_value": round(float(hac_result["p_value"]), 4),
            "inflation_factor": round(float(hac_inflation), 2),
        },
        "block_bootstrap": {
            "method": block_result["method"],
            "block_length": block_result["block_length"],
            "n_replications": n_boot,
            "se": round(float(block_result["bootstrap_std"]), 4),
            "ci_95_lower": round(float(block_result["ci_lower"]), 4),
            "ci_95_upper": round(float(block_result["ci_upper"]), 4),
        },
    },
    "effective_sample_size": round(float(effective_n), 0),
    "ci_includes_zero": bool(ci_includes_zero),
    "autocorrelation_lags_significant": int(n_significant_lags),
}

print("=== IC Inference Report ===\n")
print(json.dumps(inference_report, indent=2))
```

## Summary

### Key Concepts

| Concept | Description |
|---------|-------------|
| **HAC SE** | Newey-West adjustment for autocorrelated IC |
| **Block Bootstrap** | Preserves dependence structure via block resampling |
| **Effective N** | True sample size after accounting for autocorrelation |
| **Track Record** | Minimum observations to detect a given IC |

### Best Practices

1. **Always use HAC-adjusted t-statistics** - naive t-stats overstate significance
2. **Report block bootstrap CIs** - distribution-free and robust
3. **Consider practical significance** - IC must exceed cost drag
4. **Plan for adequate track records** - 1-2 years rarely sufficient for small IC
5. **Check effective sample size** - may be much smaller than actual observations

### API Reference

```python
from ml4t.diagnostic.metrics import compute_ic_hac_stats

# HAC-adjusted statistics
hac = compute_ic_hac_stats(ic_series, label_horizon=21)
print(f"HAC t-stat: {hac['t_stat']:.2f}")
print(f"HAC p-value: {hac['p_value']:.4f}")

# Block bootstrap (requires arch library)
from arch.bootstrap import StationaryBootstrap
bs = StationaryBootstrap(21, ic_series)  # 21-day blocks
ci = bs.conf_int(np.mean, 1000, method='percentile')
```

### Next Notebooks

- [`07_multiple_testing`](07_multiple_testing.ipynb) - FDR control when evaluating many factors
- [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) - Causal falsification tests
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。