본문으로 건너뛰기
라이브러리 문서 전체

겹치는 레이블의 패널 자기상관과 유효 표본 크기

코드 Machine Learning for Trading

요약

이 문서는 미래 수익률 구간이 겹치는 레이블을 위한 두 가지 진단 지표, 즉 통합 패널 자기상관과 레이블 고유성에 기반한 유효 표본 크기를 정의합니다. 자기상관은 원래 바 그리드에서 요청한 거리만큼 떨어진 관측값을 같은 개체 안에서만 짝짓습니다. 통합하기 전에 각 개체 내에서 값을 평균 중심화해 개체별 평균 차이로 생기는 허위 지속성을 방지합니다. 유효한 쌍이 없는 시차도 NaN으로 표시해 둡니다.

유효 표본 크기는 각 레이블의 미래 수익률 구간 중 동시에 존재하는 다른 레이블이 사용하지 않는 비율을 가중치로 삼고, 개체별 가중치를 합산합니다. 고정 기간이나 행마다 다른 기간을 지정해 지속 시간이 달라지는 이벤트 레이블을 처리할 수 있습니다. 문서는 h개 바에 걸친 미래 수익률이 h개의 수익률 구간을 차지하며, 기준점을 포함해 h+1개 바를 차지하는 것이 아니라고 강조합니다. 한 세션 기간에서는 연속 수익률 구간이 서로 겹치지 않아야 합니다. 삭제된 행이 거래 중단이나 상장 공백을 감출 수 있으므로 원래 그리드 위치를 보존하는 것이 중요합니다. 이 지표들은 의존성과 정보 중복을 진단하지만 레이블을 독립적으로 만들거나 다른 추정 오차를 보정하지는 않습니다.

핵심 아이디어

  • 각 개체 안에서 평균을 제거한 뒤, 개체 간 관측값을 통합해 자기상관을 계산합니다.
  • 누락된 행이 잘못된 인접 관계를 만들지 않도록 원래 그리드 위치로 관측값을 짝짓습니다.
  • 시차 축의 의미를 보존하도록 추정할 수 없는 시차 값은 NaN으로 둡니다.
  • 동시에 존재하는 미래 수익률 구간을 바탕으로 개체별 레이블 고유성을 계산합니다.
  • h개의 수익률 구간에 걸친 레이블은 기준점을 포함한 h+1개 바가 아니라 h개 단위를 차지합니다.

태그

전문
# label_diagnostics.py


```py
"""Panel diagnostics for overlapping labels, shared across the case studies.

Both statistics here answer the same question - how much independent information a
per-bar label with a multi-bar horizon actually carries - and both are wrong in the
same three ways when computed carelessly: on one entity rather than the panel, with
the concurrency of overlapping windows ignored, or with the frame's row order
mistaken for the grid the horizon is counted in.

The third is why both take `bar_col`. A diagnostics frame usually holds only rows
with a non-null label, and where a bar is missing - an outage, a settlement an
exchange skipped, a symbol that had not listed - the surviving rows close over the
hole. Counting positions among survivors then makes the two rows either side of a
hole adjacent, so windows that share nothing appear to overlap and windows `lag`
apart on the grid are pooled with windows further apart. `bar_col` names each row's
position on the grid the label's horizon is measured in, which the caller builds
from the frame the label was built on, before any row was dropped. Only differences
within an entity are read, so any affine origin will do.
"""

from __future__ import annotations

import numpy as np
import polars as pl
from ml4t.engineer.labeling import calculate_label_uniqueness


def panel_autocorrelation(
    frame: pl.DataFrame,
    column: str,
    *,
    max_lag: int,
    bar_col: str,
    entity_col: str = "symbol",
) -> np.ndarray:
    """Autocorrelation of *column* at lags 1..max_lag, pooled across entities.

    A pair is kept only if both rows belong to the same entity and their `bar_col`
    positions differ by exactly the lag, so no pair spans two entities and none
    spans a hole in the grid. The column is demeaned within its entity before
    pooling: without the demeaning a panel whose entities sit at different levels
    reports that level dispersion as persistence, and a series that is constant
    inside every entity - so with no autocorrelation to speak of - would come back
    at 1.0.

    A single-entity estimate is a claim about that entity, and the two disagree
    most at the lag that matters - the label horizon. A lag with no surviving pair
    is reported as NaN rather than dropped, so the returned array always has
    `max_lag` entries and the lag axis of a figure drawn from it stays honest.
    """
    centred = frame.select(
        entity_col,
        pl.col(bar_col).alias("_bar"),
        (pl.col(column) - pl.col(column).mean().over(entity_col)).alias("_centred"),
    )
    out = []
    for lag in range(1, max_lag + 1):
        lagged = centred.select(
            entity_col,
            (pl.col("_bar") - lag).alias("_bar"),
            pl.col("_centred").alias("_lagged"),
        )
        pairs = centred.join(lagged, on=[entity_col, "_bar"], how="inner")
        value = pairs.select(pl.corr("_centred", "_lagged")).item() if pairs.height else None
        out.append(np.nan if value is None else value)
    return np.array(out, dtype=float)


def effective_sample_size(
    frame: pl.DataFrame,
    *,
    bar_col: str,
    horizon: int | None = None,
    horizon_col: str | None = None,
    entity_col: str = "symbol",
) -> tuple[int, float]:
    """Return (rows, N_eff) for a label sampled every bar over *horizon* bars.

    Pass ``horizon_col`` instead of ``horizon`` where the window is not the same length
    for every row - an event label that resolves when a barrier is hit or when a contract
    expires. The column holds each row's window in the same units as ``bar_col``, and a
    single ``horizon`` is the special case where every row carries the same value. A
    median window standing in for a variable one prices the overlap of a label none of
    the rows has.

    ``N_eff`` is Chapter 7.2's average-uniqueness sum: each row is weighted by the
    share of its forward window no concurrent label also spans. Concurrency is a
    property of one entity's overlapping windows, so the weights are computed per
    entity and summed, over the entity's own grid positions - a window that starts
    on the far side of a hole is concurrent with nothing on the near side.

    **What a label occupies is ``horizon`` return intervals, not ``horizon + 1``
    bars.** The label at bar *i* is $P_{i+h}/P_i - 1$, so it consumes the returns
    realised over bars $i{+}1 \\ldots i{+}h$ - *h* of them - and the label at *i+1*
    shares $h-1$ of those, which is the overlap the audit record prints. Passing a
    closed bar interval ``[i, i+h]`` instead counts the anchor bar as consumed and
    makes every label span ``h+1`` units, so consecutive labels appear to share one
    interval even when they share none.

    The one-session horizon is the case that settles it: consecutive one-day
    forward returns are built from disjoint returns and are fully independent, so
    every weight must be 1 and ``N_eff`` must equal ``N``. The closed-bar form
    returns ``N/2`` there. On a gapless grid average uniqueness converges to
    ``1/h``, so ``N_eff`` tends to ``N/h`` - the reference value the stage standard
    cites - and a grid with holes sits above it, because a hole ends an overlap
    early.

    *frame* is expected to hold only rows with a non-null label, so every row has a
    complete forward window even though the bars closing the last few are not
    themselves rows of *frame*; the endpoints are left uncapped and the concurrency
    array extended past the last window's end rather than truncated, which would
    shorten exactly those windows.
    """
    if (horizon is None) == (horizon_col is None):
        raise ValueError("pass exactly one of horizon and horizon_col")
    # `maintain_order=True` is what makes the total reproducible. Summing floats is not
    # associative, and polars does not fix the order groups come back in, so the same frame
    # summed twice differs in the last bits. Printed as an integer that lands on either side
    # of a rounding boundary: sp500_options' fwd_ret_10d reported N_eff 39,746 on one run and
    # 39,747 on the next, from identical inputs and an unchanged label digest.
    rows, weight = 0, 0.0
    for _, group in frame.group_by([entity_col], maintain_order=True):
        bars = group[bar_col].to_numpy()
        order = np.argsort(bars)
        events = bars[order] - bars.min()
        windows = horizon if horizon_col is None else group[horizon_col].to_numpy()[order]
        ends = events + windows - 1
        weights = calculate_label_uniqueness(events, ends, n_bars=int(ends.max()) + 1)
        rows += group.height
        weight += float(weights.sum())
    return rows, weight

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.