跳至正文
返回文库全部文档

构建月度收益标签:US公司特征

代码 《交易机器学习》

总结

本笔记为横截面 US 公司特征研究定义并验证月度收益标签。笔记检查数据提供方如何将每行特征与收益结果配对,并强调面板数据已经对齐,不应再次移位。主要标签是月度总收益;缩尾收益和月内分类变体则对同一结果进行转换。月度横截面阈值使用其适用月份的信息,而时间序列阈值则需要在每个训练期内估算。

笔记还将缺失结果与各公司的最后观测期进行核对,检查序列相关性和有效样本量,并使用预先指定的信号建立基准比较。笔记记录了收盘到收盘的标签惯例,同时指出策略在下一开盘时交易,因此留下未计量的隔夜缺口。其他限制包括匿名公司历史数据、沿用数据提供方关于股息和退市的惯例,以及仅基于一个特征的基准。所得标签和折分边界旨在为后续特征构建和回测阶段提供依据。

核心观点

  • 应先核验预先对齐的特征面板与其观测时间是否一致,再决定是否需要进一步移位收益。
  • 月度收益、缩尾收益和类别标签,是对同一结果进行不同转换的表示。
  • 在每个月内计算横截面阈值,可避免用后续月份定义该月标签。
  • 标签覆盖范围应与观测到的结果可用性一致;若结果窗口重叠,行数统计也应考虑这一点。
  • 收盘到收盘标签可能不同于下一开盘执行收益;若源数据无法核实,仍会沿用数据提供方的惯例。

标签

全文
# 02_labels.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # US Firm Characteristics: Label Engineering
#
# Every model in this case study is trained to predict the label defined here, so an error
# in it is silent where it is made and reaches every metric and every backtest after it.
# This notebook fixes what the label is and when it is known, checks against the data that
# the panel is paired the way the provider documents, measures how much independent
# information the rows carry, measures what the simplest signal the study already names
# earns against the label, and writes the label files the rest of the pipeline is built on.
#
# ## Learning objectives
#
# - Work out from the data which month's return a pre-built research panel has already
#   paired with each row, instead of taking the provider's description on trust
# - Check that a label is missing on exactly the rows whose outcome was never observed - a
#   count of how many rows carry a label cannot tell you that
# - Decide whether a rule that sorts firms into classes needs re-estimating inside each
#   training period, or whether the way it is built already keeps later data out of it
# - Measure how much of one row's outcome the next row's label repeats, and turn a row
#   count into the number of independent observations it is worth
# - Measure what a signal the study already names earns against the label, so that a
#   feature built afterwards is compared against a number fixed before it existed
#
# ## Book reference, prerequisites and artifacts
#
# Chapter 7, Section 7.2. Reads the Chen-Pelger-Zhu panel through
# `load_firm_characteristics()`, whose breadth and cost feasibility
# [`01_feasibility_analysis`](01_feasibility_analysis.ipynb) establishes, and
# `config/setup.yaml`, which declares the label set, the buffer and the holdout boundary.
# Writes `labels/prices.parquet`, which the backtest reads to price positions; one parquet
# per declared label; a small JSON record beside each of those files describing what is in
# it; and `config/cv_config.json`, the committed record of the fold boundaries these labels
# imply.

# %%
"""US Firm Characteristics: Label Engineering."""

import json
from datetime import date

import matplotlib.pyplot as plt
import numpy as np
import polars as pl
import yaml
from matplotlib.ticker import PercentFormatter
from ml4t.diagnostic.metrics import compute_ic_hac_stats, cross_sectional_ic_series

from case_studies.utils.artifact_digest import value_digest, write_artifact
from case_studies.utils.artifact_quality import quality_report, render_quality_report
from case_studies.utils.label_diagnostics import effective_sample_size, panel_autocorrelation
from case_studies.utils.warning_policy import apply_notebook_warning_policy
from data import load_firm_characteristics
from utils.artifact_specs import resolve_label_buffer, resolve_label_horizon
from utils.cv_splits import generate_cv_splits
from utils.paths import display_path, get_case_study_dir
from utils.style import COLORS, FIGSIZE, add_message_title, show_with_alt

apply_notebook_warning_policy()

CASE_STUDY_ID = "us_firm_characteristics"
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
LABELS_DIR = CASE_DIR / "labels"

# %% [markdown]
# The two parameters bound the study window and are both read below. The release starts two
# decades earlier than the window used here; the folds `setup.yaml` declares need a decade
# of training before the first validation year, and this window leaves room for that
# without reaching back into a period whose cross-section is half its later size.

# %% tags=["parameters"]
START_DATE = "1990-01-01"
END_DATE = "2016-12-31"

# %% [markdown]
# ## Configuration
#
# Everything that defines a label is declared in `config/setup.yaml` and bound here. A
# horizon or a boundary typed into a cell is a second copy of a value the rest of the
# pipeline reads from the file, and the two drift apart the first time either is edited.
#
# Three durations are declared, and conflating any two of them is the mistake this panel
# invites. `labels.buffer` is the gap left between a training window and the validation window
# that follows it. `labels.horizons` is how far past its own timestamp a label's outcome
# resolves, which is what the splitter seals the last validation fold on. The span the label
# measures over is a third thing again, and it is what a Newey-West lag and an effective sample
# size are counted in. Here the buffer is one month, the outcome horizon is zero because the
# release dates each row by the month its return was earned in, and the span is the one month
# the buffer was sized to. Section D measures the alignment the middle one rests on.

# %%
SETUP = yaml.safe_load((CASE_DIR / "config" / "setup.yaml").read_text())

PRIMARY_LABEL = SETUP["labels"]["primary"]
LABEL_NAMES = [PRIMARY_LABEL, *SETUP["labels"].get("variants", [])]
WINSORIZED_LABEL, CLASS_LABEL = LABEL_NAMES[1], LABEL_NAMES[2]
OUTCOME_HORIZON = resolve_label_horizon(CASE_STUDY_ID, PRIMARY_LABEL, SETUP)
LABEL_BUFFER = resolve_label_buffer(CASE_STUDY_ID, PRIMARY_LABEL, SETUP)
LABEL_SPAN_MONTHS = int(str(LABEL_BUFFER).rstrip("Mm"))
HOLDOUT_START = date.fromisoformat(str(SETUP["evaluation"]["holdout_start"]))
BASELINE_SIGNAL = SETUP["causal"]["treatment"]
KEYS = ["timestamp", "symbol"]

print(
    f"Labels: {LABEL_NAMES}, primary {PRIMARY_LABEL}, span {LABEL_SPAN_MONTHS} month, "
    f"buffer {LABEL_BUFFER}, outcome horizon {OUTCOME_HORIZON}\n"
    f"Study window {START_DATE} to {END_DATE}, holdout opens {HOLDOUT_START}"
)

# %% [markdown]
# ## A. The learning task
#
# The hypothesis is cross-sectional. Firms differ along accounting ratios, price-based
# measures and turnover proxies, and the claim is that a firm ranked high on those measures
# out-earns a firm ranked low over the following month - ranked against the other firms
# trading that month, not judged in isolation. The label is therefore a firm's monthly total
# return, and the strategy that consumes it is long-short across the cross-section.
#
# The decision cadence comes from `setup.yaml`: the book is set at the month-end close and
# executed at the next open, so one month is both the interval a position is held for and
# the interval an outcome is measured over. Accounting variables are refreshed once a year
# at the end of June and price-based ones every month-end, which is what makes a monthly
# rebalance the fastest cadence the release supports.
#
# Two variants transform the same monthly outcome rather than defining a second one. The
# winsorized return exists because a monthly cross-section of individual stocks carries
# returns of several hundred percent, and a squared-error loss fitted against those spends
# its capacity on them. The classification label turns the same return into a rank question
# - is this firm in the better half of this month - which is closer to what a long-short
# book acts on than the return's magnitude is.

# %% [markdown]
# ## B. Preparation before the label
#
# **The panel is already aligned, and the mistake this dataset invites is aligning it
# again.** Chen, Pelger and Zhu publish each row as a pair: the characteristics a firm
# carried at the end of one month, and the return that firm went on to earn over the
# following month. The row is dated by the month the return was earned in, so the
# information a model reads is a month older than the outcome it predicts, and shifting
# `ret` forward once more would pair a firm's characteristics with a return earned two
# months later. Section D reads that alignment out of the data rather than accepting it
# here.
#
# Nothing else is transformed. The characteristics arrive cross-sectionally rank-transformed
# by the provider, and the return is the raw monthly total return. Which firms are in the
# universe is decided by the provider's completeness rule - a firm-month is kept only where
# every characteristic is available - so that screen runs before this notebook, and neither
# this notebook nor stage 03 drops a further row.
#
# Where a label is built by shifting a price series rather than read from a column, the order
# of those two steps matters and getting it wrong is silent. A shift counts rows, so dropping
# ineligible rows first makes the shift step over whatever was removed: a one-month label then
# spans however long the gap was, and still looks like a full column. Applying the screen after
# the label, or expressing the horizon as a length of time so that a gap produces a null,
# avoids it.

# %%
window = pl.col("timestamp").is_between(
    pl.lit(START_DATE).str.to_date(), pl.lit(END_DATE).str.to_date()
)
firm_chars = load_firm_characteristics(split="all").filter(window).sort(["symbol", "timestamp"])
CHARACTERISTICS = [c for c in firm_chars.columns if c not in {*KEYS, "ret", "split"}]

# Recorded as the `inputs` of every artifact written below: a re-run against a refreshed
# release is otherwise indistinguishable from this one.
MARKET_DATA_DIGEST = value_digest(firm_chars, [*KEYS, "ret"])

print(f"{firm_chars['symbol'].n_unique():,} firms, {firm_chars.height:,} firm-months")
print(
    f"Month-ends {firm_chars['timestamp'].min()} to {firm_chars['timestamp'].max()}, "
    f"{len(CHARACTERISTICS)} characteristics, {firm_chars['ret'].null_count()} rows without a return"
)
print(f"market_data digest: {MARKET_DATA_DIGEST}")

# %% [markdown]
# ## C. Label construction
#
# One execution convention, written once and inherited by all three labels:
#
# $$r_{i,t} = \frac{P_{i,t}}{P_{i,t-1}} - 1$$
#
# where $P_{i,t}$ is firm $i$'s adjusted month-end close in month $t$, and the row carrying
# $r_{i,t}$ also carries the characteristics observed at the close of month $t-1$. That is
# Chapter 7.2's close-to-close convention, with the decision taken at the earlier of the two
# closes. The strategy that consumes the label fills at the next open instead, as
# `setup.yaml` declares, so the label omits whatever the price moves overnight between the
# decision close and that open. Nothing in this release can measure that gap - it publishes
# no opening price - so it is inherited as a known difference between the label and the
# traded return rather than corrected here.
#
# The provider computes $r_{i,t}$, so the labelling library has nothing to shift and is not
# called: `fixed_time_horizon_labels` divides a price column by its own lag, and this release
# publishes no price column. The two variants are cross-sectional transforms of the primary
# label, each computed inside one month.
#
# **Both thresholds are read off the cross-section of the month they apply to**, which is
# what makes them point in time. Chapter 7.2 draws the line here: a percentile taken across
# the firms trading in one month uses only information that month's close already carries,
# while a percentile taken along the time axis reads later months and has to be fitted inside
# the training fold. A median split also holds its class proportions near one half by
# construction, which Section E measures rather than assumes.
#
# A label resolves on the month-end that dates the row carrying it, because the return the
# row reports was earned over the month ending there. The date a label resolves on is what
# decides whether it may be looked at: a row whose outcome lands on or after
# `evaluation.holdout_start` falls inside the period held back for the final test, so it is
# written to the label files but excluded from every measurement below. The files themselves
# keep every row, because the later stages need the held-back months to score on.
#
# The `month` column numbers the panel's own month-end grid. Sections D and F count on it
# rather than on position among the rows, so that two rows either side of a month a firm
# skipped are not read as consecutive observations.

# %%
month_ends = firm_chars.select("timestamp").unique().sort("timestamp").with_row_index("month")
bounds = firm_chars.group_by("timestamp").agg(
    pl.col("ret").median().alias("_median"),
    pl.col("ret").quantile(0.01).alias("_lo"),
    pl.col("ret").quantile(0.99).alias("_hi"),
)
labels_df = (
    firm_chars.join(month_ends, on="timestamp", how="left")
    .join(bounds, on="timestamp", how="left")
    .with_columns(
        pl.col("timestamp").alias("_label_end"),
        pl.col("ret").alias(PRIMARY_LABEL),
        pl.col("ret").clip(pl.col("_lo"), pl.col("_hi")).alias(WINSORIZED_LABEL),
        pl.when(pl.col("ret").is_not_null())
        .then(pl.col("ret") > pl.col("_median"))
        .cast(pl.Int32)
        .alias(CLASS_LABEL),
    )
)
dev = labels_df.filter(pl.col("_label_end") < HOLDOUT_START)
print(
    f"Constructed {', '.join(LABEL_NAMES)} on {labels_df.height:,} firm-months, "
    f"{dev.height:,} of them development rows through {dev['timestamp'].max()}: "
    f"{dev['timestamp'].n_unique()} month-ends, {dev['symbol'].n_unique():,} firms"
)

# %% [markdown]
# ## D. Window validity
#
# A column of the right length always arrives; the question is whether what it holds is the
# quantity the label claims. Each property below fails silently and leaves plausible numbers
# behind, so each is asserted rather than described.
#
# The first assertion is the one that catches a fabricated tail. A firm's last month in the
# panel closes its label window, and a construction that shifted a price series would leave
# that month with no outcome - which has to be a null, never a value. The rest check the two
# monthly joins: a duplicated month-end in either would multiply the panel's rows, and a
# clipped return falling outside its own month's percentiles would mean a join had matched
# the wrong month.

# %%
missing = labels_df["ret"].is_null()
inside = pl.col(WINSORIZED_LABEL).is_between(pl.col("_lo"), pl.col("_hi"))
for name in LABEL_NAMES:
    # 1. Null exactly where the outcome was not observed.
    assert (labels_df[name].is_null() == missing).all(), name
# 2. One row per firm-month, so neither monthly join fanned the panel out.
assert labels_df.select(pl.struct(KEYS).n_unique()).item() == labels_df.height
# 3. Each label is a transform of its own row's return, inside its own month.
assert labels_df.filter(pl.col(PRIMARY_LABEL) != pl.col("ret")).height == 0
assert labels_df.filter(~inside).height == 0
# 4. No discrete label is derived from a null return.
assert labels_df.filter(missing & pl.col(CLASS_LABEL).is_not_null()).height == 0

clipped = labels_df.filter(pl.col(PRIMARY_LABEL) != pl.col(WINSORIZED_LABEL)).height
skipped = labels_df.select((pl.col("month").diff().over("symbol") > 1).sum()).item()
print(
    f"{clipped:,} of {labels_df.height:,} returns clipped by their own month's percentiles; "
    f"{skipped:,} places where a firm skips a month-end inside its own history"
)

# %% [markdown]
# Section B's claim about the alignment is what every number downstream rests on, and it can
# be read out of the release itself. `ST_REV` is the short-term-reversal characteristic: a
# firm's own most recent monthly return, ranked across firms, as of the date the
# characteristics were observed. The rank correlation between `ST_REV` and the label is
# therefore a test of which row's return the characteristics were recorded after, and only
# one lag can carry it. Pairs are restricted to rows one month-end apart, so a firm that
# skips a month contributes none.

# %%
step = pl.col("month").diff().over("symbol")
alignment = dev.with_columns(
    pl.when(step == 1).then(pl.col(PRIMARY_LABEL).shift(1).over("symbol")).alias("_previous"),
    pl.when(step.shift(-1).over("symbol") == 1)
    .then(pl.col(PRIMARY_LABEL).shift(-1).over("symbol"))
    .alias("_next"),
)
for described, column in (
    ("the previous month's", "_previous"),
    ("this row's own", PRIMARY_LABEL),
    ("the next month's", "_next"),
):
    paired = alignment.drop_nulls(["ST_REV", column])
    value = paired.select(pl.corr("ST_REV", column, method="spearman")).item()
    print(f"ST_REV against {described:>21} return: rank correlation {value:+.3f}")

# %% [markdown]
# Position zero below is each firm's last month in the panel. Every label has to be present
# there, because the outcome that month reports was earned before the row was written. The
# grey line is the same panel under one forward shift, and it is the failure this figure
# exists to make visible: a scalar count of valid rows reports both shapes as complete.
#
# The horizontal axis counts month-ends on the panel's own grid, not rows, so a firm that
# is absent for a month reads as two month-ends apart and not as one.

# %%
profile = (
    labels_df.with_columns(
        (pl.col("month").max().over("symbol") - pl.col("month")).alias("from_end"),
        pl.col(PRIMARY_LABEL).shift(-LABEL_SPAN_MONTHS).over("symbol").alias("_shifted"),
    )
    .filter(pl.col("from_end") <= LABEL_SPAN_MONTHS + 4)
    .group_by("from_end")
    .agg(pl.col(c).is_not_null().mean() for c in [*LABEL_NAMES, "_shifted"])
    .sort("from_end")
)

fig, ax = plt.subplots(figsize=FIGSIZE["single"])
back = profile["from_end"]
markers = zip(LABEL_NAMES, (COLORS["blue"], COLORS["amber"], COLORS["copper"]), "os^", strict=True)
for name, color, marker in markers:
    ax.plot(back, profile[name], marker, ms=8, mfc="none", c=color, label=name)
ax.plot(
    back, profile["_shifted"], "--", ds="steps-mid", lw=1.4, c=COLORS["neutral"], label="shifted"
)
ax.set(
    xlabel="Month-ends back from each firm's last month in the panel",
    ylabel="Share of firms carrying a label",
    ylim=(-0.08, 1.12),
)
add_message_title(
    ax,
    "Every firm's last month in the panel carries a label",
    subtitle="All three labels sit on one another; one forward shift empties position zero",
)
ax.legend(loc="lower right", frameon=False, ncol=2)
show_with_alt(fig, "Non-null label rate by month-ends back from each firm's last observation.")

# %% [markdown]
# ## E. Distribution and base rate
#
# What scale is the label, and does it mean the same thing across regimes?
#
# Both continuous labels go on one axis with identical bins and a logarithmic count axis.
# The claim the figure has to support is about shape rather than width - two standard
# deviations would carry the width - and the shape here is that the two series are
# indistinguishable through the body and separate only in the tails, where the winsorized
# label stops at the widest boundary any month produced and the raw one runs on. The axis
# has to reach past that widest boundary or the figure is cut off exactly where the two
# labels part; it still stops well short of the raw label's largest return, and the rows
# past it are counted underneath rather than drawn.

# %%
bins = np.linspace(-1.0, 2.0, 91)
histograms = {
    PRIMARY_LABEL: dict(color=COLORS["neutral"], histtype="step", lw=2, zorder=2),
    WINSORIZED_LABEL: dict(color=COLORS["amber"], alpha=0.8, zorder=1),
}
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
for name, style in histograms.items():
    series = dev[name]
    tag = f"{name}, std {series.std():.3f}, kurtosis {series.kurtosis():.1f}"
    ax.hist(series.to_numpy(), bins=bins, label=tag, **style)
ax.axvline(0, color=COLORS["neutral"], linestyle="--", lw=0.8)
ax.set(xlabel="One-month total return", ylabel="Firm-months per bin, log scale", yscale="log")
ax.xaxis.set_major_formatter(PercentFormatter(1.0))
add_message_title(
    ax,
    "Winsorizing at the monthly percentiles pulls in only the tails",
    subtitle="Identical bins, development window; rows beyond the axis are counted below",
)
ax.legend(loc="upper right", frameon=False, fontsize=8)
show_with_alt(fig, "Histograms of the raw and winsorized labels on identical bins, log counts.")

for name in (PRIMARY_LABEL, WINSORIZED_LABEL):
    series, drawn = dev[name], dev[name].is_between(bins[0], bins[-1]).sum()
    print(
        f"{name}: mean {series.mean():+.5f}, std {series.std():.5f}, "
        f"kurtosis {series.kurtosis():.2f}, range {series.min():.3f} to {series.max():.3f}, "
        f"{series.len() - drawn:,} rows beyond the axis"
    )

# %% [markdown]
# Chapter 7.2 asks for the scale and the base rate to be tracked through time. For the
# continuous labels the quantity that has to hold up is the spread the model ranks within:
# where it narrows, the same rank correlation buys less return. That spread is taken across
# firms on each month-end first and only then averaged over the year, because pooling every
# firm-month in a year into one standard deviation adds the market's own move from month to
# month to the spread across firms, and a ranking model is scored on the second alone. For
# the classification label the quantity is the share of firms in the upper class, and the
# lower panel is what shows whether the median split delivered the balance it promises.

# %%
monthly = dev.group_by("timestamp").agg(
    pl.col(PRIMARY_LABEL).std().alias("dispersion"),
    pl.col(CLASS_LABEL).mean().alias("upper_share"),
    (pl.col("ret") == pl.col("_median")).mean().alias("tied"),
)
annual = (
    monthly.group_by(pl.col("timestamp").dt.year().alias("year"))
    .agg(pl.col("dispersion").mean(), pl.col("upper_share").mean())
    .sort("year")
)
fig, axes = plt.subplots(2, 1, sharex=True, figsize=FIGSIZE["dual_v"])
panels = (
    ("dispersion", COLORS["blue"], annual["dispersion"].median(), "Cross-firm return std"),
    ("upper_share", COLORS["amber"], 0.5, "Share in the upper class"),
)
for ax, (column, color, reference, ylabel) in zip(axes, panels, strict=True):
    ax.plot(annual["year"], annual[column], "o-", ms=4, lw=1.8, c=color)
    ax.axhline(reference, color=COLORS["neutral"], ls="--", lw=1.1)
    ax.set_ylabel(ylabel)
    ax.yaxis.set_major_formatter(PercentFormatter(1.0))
# Bound the lower panel symmetrically about the even split, by the series it actually draws.
# Centring keeps the distance from balance readable as a distance; taking the bound from the
# monthly range instead would leave most of the panel blank and flatten the annual series.
reach = (annual["upper_share"] - 0.5).abs().max() * 1.3
axes[1].set(ylim=(0.5 - reach, 0.5 + reach), xlabel="Year")
add_message_title(
    axes[0],
    "Dispersion shifts across regimes while the median split stays balanced",
    subtitle="Annual means of the monthly cross-section; dashed lines mark the typical year's "
    "dispersion and an even split",
)
show_with_alt(fig, "Annual mean cross-firm return dispersion above, upper-class share below.")

peak, trough = (annual.sort("dispersion", descending=d).row(0, named=True) for d in (True, False))
print(
    f"dispersion peaks at {peak['dispersion']:.1%} in {peak['year']:.0f} against "
    f"{trough['dispersion']:.1%} in {trough['year']:.0f}; median year "
    f"{annual['dispersion'].median():.1%}\n"
    f"upper class runs {monthly['upper_share'].min():.3f} to {monthly['upper_share'].max():.3f} "
    f"per month, mean {monthly['upper_share'].mean():.3f}; the shortfall is returns tied at the "
    f"median, up to {monthly['tied'].max():.3f} of a month's firms"
)

# %% [markdown] tags=["results"]
# On the development window the raw label has a standard deviation of 0.17430 and a kurtosis
# of 335.36, against 0.14850 and 6.47 for the winsorized variant: clipping at each month's
# own percentiles touches a fiftieth of the rows and removes almost all of the fourth moment.
# Cross-firm dispersion is not stable, peaking at 22.4% in 2000 against 11.8% in 2013 and a
# median year of 15.3%, while the classification base rate is, running from 0.438 to 0.500 a
# month around a mean of 0.498. The shortfall below an even split is returns tied at the
# month's median, which reach 0.084 of the firms in a month.

# %% [markdown]
# ## F. Overlap and effective sample size
#
# Where a label is sampled more often than the period it measures over, consecutive rows
# report overlapping stretches of the same price path, and the row count then claims more
# evidence than the data holds. Here a firm is observed once a month and each label measures
# one month, so consecutive labels are built from returns that share nothing. Two
# measurements say so in different units: the autocorrelation profile, and the row count
# after each row is discounted by how much of its window another row also covers - the
# average-uniqueness weighting of Chapter 7.2.
#
# Both read `month`, each row's position on the panel's own month-end grid, rather than its
# position among the rows that reach this point. A firm absent for a month would otherwise
# have the rows either side of the gap treated as consecutive, pairing windows that in fact
# share nothing.

# %%
MAX_LAG = 12
acf = {n: panel_autocorrelation(dev, n, max_lag=MAX_LAG, bar_col="month") for n in LABEL_NAMES}

fig, ax = plt.subplots(figsize=FIGSIZE["single"])
lags = np.arange(1, MAX_LAG + 1)
palette = (COLORS["blue"], COLORS["amber"], COLORS["copper"])
for name, color in zip(LABEL_NAMES, palette, strict=True):
    ax.plot(lags, acf[name], "o-", ms=3, lw=1.6, c=color, label=name)
ax.axhline(0, color=COLORS["neutral"], lw=0.8)
ax.set(xlabel="Lag in month-ends", ylabel="Panel autocorrelation")
add_message_title(
    ax,
    "Monthly labels carry weak reversal rather than an overlap decay",
    subtitle="Consecutive labels share no return interval, so every lag drawn is already "
    "overlap-free",
)
ax.legend(loc="lower right", frameon=False)
show_with_alt(fig, "Panel autocorrelation of all three labels against lag in month-ends.")

# A horizon-h label consumes the h returns realised over its window and its neighbour one
# period later shares h-1 of them, so at a one-period horizon nothing is shared at all.
n_rows, n_eff = effective_sample_size(dev, horizon=LABEL_SPAN_MONTHS, bar_col="month")
assert n_eff == n_rows, "a one-month label overlaps nothing, so every row must weigh one"
print(
    f"{PRIMARY_LABEL}: N={n_rows:,}, N_eff={n_eff:,.0f}, ratio {n_eff / n_rows:.4f}; "
    f"autocorrelation {acf[PRIMARY_LABEL][0]:+.4f} at lag one, "
    f"{acf[PRIMARY_LABEL][-1]:+.4f} at lag {MAX_LAG}"
)

# %% [markdown] tags=["results"]
# The development window's 780,882 firm-months carry 780,882 effective observations: a
# one-month label sampled monthly shares no return interval with its neighbour, so average
# uniqueness is one and the ratio is 1.0000. What autocorrelation remains is the return
# process rather than the label construction - -0.0227 at one month and -0.0045 at twelve -
# and it is the short-term reversal `ST_REV` already carries as a characteristic. The purge
# gap a fold needs is set by the forward window, which is one month here.

# %% [markdown]
# ## G. Baseline floor
#
# One signal, measured against the primary label over the development months only and with
# no feature engineering: the raw momentum characteristic `setup.yaml` names as the treatment
# whose effect `09_causal_dml.py` estimates, taken exactly as the provider ranked it. This is
# the number a feature built in the next notebook has to beat, and it is only worth beating
# if it was fixed before any feature existed.
#
# The information coefficient is the rank correlation between the signal and the outcome
# across the firms trading in one month, computed for each month-end and then averaged. That
# is the quantity a ranking model is scored on; pooling every firm-month into one correlation
# instead answers a different question, mixing how firms rank against each other with how the
# market moved from month to month. The library call returns the series ordered by time,
# which the standard error below depends on.
#
# A month with too thin a cross-section produces a rank correlation too noisy to average in,
# so months below a minimum count are left out. That minimum is set at half the median
# month's firm count rather than at a fixed number, so the same rule transfers to a universe
# of a different size. The standard error is Newey-West adjusted, which allows for one
# month's information coefficient being related to the next month's; the count the statistic
# reports is the number of month-ends that actually entered it.

# %%
baseline = dev.drop_nulls([BASELINE_SIGNAL, PRIMARY_LABEL])
min_obs = int(baseline.group_by("timestamp").len()["len"].median() // 2)

ic = cross_sectional_ic_series(
    baseline,
    baseline,
    pred_col=BASELINE_SIGNAL,
    ret_col=PRIMARY_LABEL,
    date_col="timestamp",
    entity_col="symbol",
    min_obs=min_obs,
).sort("timestamp")  # HAC autocovariances are meaningless over a permutation of time
stats = compute_ic_hac_stats(ic, ic_col="ic", label_horizon=LABEL_SPAN_MONTHS)

# `ic` carries one row per month-end whatever its cross-section, with the correlation left
# null below `min_obs`; the statistic drops those, so its own count is what is reported.
print(
    f"Baseline: {BASELINE_SIGNAL} against {PRIMARY_LABEL}, minimum cross-section {min_obs} firms\n"
    f"  month-ends entering the statistic {stats['n_periods']:,} "
    f"of {dev['timestamp'].n_unique()}\n"
    f"  mean IC {stats['mean_ic']:+.5f}, HAC standard error {stats['hac_se']:.5f}\n"
    f"  HAC t {stats['t_stat']:+.2f} on {stats['effective_lags']} Bartlett lags, "
    f"naive t {stats['naive_t_stat']:+.2f}, p {stats['p_value']:.3g}"
)

# %% [markdown] tags=["results"]
# Raw momentum earns a mean information coefficient of +0.04398 against the monthly label
# over all 312 development month-ends, on a cross-section of at least 1293 firms. Under the
# naive standard error that is a t-statistic of +6.89; the Newey-West rule picks 5 lags here,
# above the none a one-month horizon requires on its own, and the HAC statistic is +6.14 with
# a standard error of 0.00716. So a feature is not compared against zero: it has to improve
# on a mean information coefficient of 0.04398 that the data already separates from zero,
# fixed here before any feature exists.

# %% [markdown]
# ## H. Artifacts and the audit record
#
# Every file written here gets a small JSON record saved next to it. Its job is to let a
# later reader tell one version of a file apart from another without opening either, and to
# say where the contents came from. It holds a digest - a short string computed from the
# values in the file, which changes if any of them change - together with the row count, the
# columns that identify a row, the notebook that wrote the file, and the digest of each file
# or dataset it was built from. That last field is what ties a label back to the release of
# the data it was computed from: without it, a re-run against a refreshed download looks
# exactly like a re-run against the same one.
#
# The characteristic panel goes out first, because the backtest prices this case study's
# positions from it and all three labels derive from its `ret` column; the two variants
# additionally carry the primary label's digest, which is what they transform.

# %%
prices = write_artifact(
    firm_chars.select([*KEYS, "ret", *CHARACTERISTICS]),
    LABELS_DIR / "prices.parquet",
    keys=KEYS,
    written_by="02_labels",
    inputs={"market_data": MARKET_DATA_DIGEST},
)
print(f"prices.parquet: {prices['n_rows']:,} rows, digest {prices['digest']}")

records: dict[str, dict] = {}
for name in LABEL_NAMES:
    derived = {} if name == PRIMARY_LABEL else {PRIMARY_LABEL: records[PRIMARY_LABEL]["digest"]}
    records[name] = write_artifact(
        labels_df.select([*KEYS, name]).drop_nulls(),
        LABELS_DIR / f"{name}.parquet",
        keys=KEYS,
        written_by="02_labels",
        inputs={"prices": prices["digest"], **derived},
    )
    print(f"{name}.parquet: {records[name]['n_rows']:,} rows, digest {records[name]['digest']}")

# %% [markdown]
# The training and validation periods the model stages use are worked out per label by
# `case_studies/utils/cv_window.py`, from `config/setup.yaml` and the range of dates in the
# label file written above - so which rows land in that file is what decides where the period
# boundaries fall. Every stage that needs those periods derives them the same way rather than
# reading a file an earlier notebook happened to write, so none of them depends on the order the
# pipeline was run in. The cell below runs the generator once more and saves the result to
# `config/cv_config.json` as a committed record of the geometry these labels imply, which is what
# makes a change in it show up in a diff.
#
# The two durations the generator is given are the ones separated above. `label_buffer` sets the
# gap between a training window and its validation window. `outcome_horizon` is how far past the
# last validation date an outcome is still unresolved, and it is what decides how much of the
# last fold has to be given back before the holdout opens - zero here, because a row's return is
# realised on the timestamp the row carries.

# %%
splits = generate_cv_splits(
    labels_df.select("timestamp"),
    case_study_id=CASE_STUDY_ID,
    label_buffer=LABEL_BUFFER,
    outcome_horizon=OUTCOME_HORIZON,
)
assert len(splits) == int(SETUP["evaluation"]["n_splits"])
assert max(split["val_end"] for split in splits).date() < HOLDOUT_START

BLOCKS = ("train_start", "train_end", "val_start", "val_end")
cv_config = {
    "case_study_id": CASE_STUDY_ID,
    "n_splits": len(splits),
    "train_size": str(SETUP["evaluation"]["train_size"]),
    "val_size": str(SETUP["evaluation"]["val_size"]),
    "holdout_start": HOLDOUT_START.isoformat(),
    "holdout_end": str(SETUP["evaluation"]["holdout_end"]),
    "splits": [
        {"fold": int(split["fold"]), "label_buffer": LABEL_BUFFER}
        | {key: split[key].date().isoformat() for key in BLOCKS}
        for split in splits
    ],
}
cv_path = CASE_DIR / "config" / "cv_config.json"
cv_path.write_text(json.dumps(cv_config, indent=2) + "\n")
last_val = max(split["val_end"] for split in splits).date()
print(f"Saved {display_path(cv_path)}: {len(splits)} folds, validation through {last_val}")

# %% [markdown]
# The record Chapter 7.2 requires to close a label definition, one row per label, built from
# the values computed above rather than written by hand.

# %%
shared = "cv_window.py, for this label's training and validation periods; the model stages"
readers = {PRIMARY_LABEL: f"{shared}; 03_financial_features.py, to check firm identities match"}
print("\nLabel audit record")
for name in LABEL_NAMES:
    series = dev[name]
    base_rate = (
        f"upper class {series.mean():.3f} of firms"
        if name == CLASS_LABEL
        else f"mean {series.mean():+.5f}, std {series.std():.5f}"
    )
    print(
        f"\n{name}\n  anchor       the adjusted close of the month before the one dating the row"
        f"\n  span         {LABEL_SPAN_MONTHS} month, the month the row is dated by"
        f"\n  resolution   fixed at that month's close, the timestamp the row itself carries;"
        f" monthly bars need no intraday tie-break"
        f"\n  overlap      {LABEL_SPAN_MONTHS - 1} months shared by consecutive rows"
        f"\n  base rate    {base_rate}"
        f"\n  consumed by  {readers.get(name, shared)}"
    )

# %% [markdown]
# ## What the labels hold, and what they owe
#
# Two questions about the files this stage just wrote, and the rows that are there answer only one
# of them. The first is what is in each column - nulls, how much sits at exactly zero, how far the
# extreme values are from the body, whether anything is constant. A threshold crossed there asks
# for a sentence and settles nothing on its own: a winsorized return is *supposed* to pile up at
# its bounds, and a class label is supposed to concentrate.
#
# The second is coverage, and it needs a denominator that is not the labels themselves. **The
# reference is `firm_chars`** - every firm-month the panel carries, which is the set a label could
# in principle have been written for. Comparing one label to the other two would hide any month
# where all three are absent together; comparing them to the panel cannot.
#
# What each label is entitled to be short by is unusually easy to state here, and unusually strict:
# **the outcome horizon is zero.** A row is dated by the month its return was earned in rather
# than by a month whose outcome is still ahead of it, so there is no forward window running off
# the end of the sample and no burn-in before a window fills. Every one of these labels owes a
# value on every firm-month the panel carries a return for, and anything missing is either a null
# return in the source or something this stage did.

# %%
expected_keys = firm_chars.select(KEYS).unique()
print(
    f"panel: {expected_keys.height:,} (firm, month) keys across "
    f"{expected_keys['symbol'].n_unique()} firms and "
    f"{expected_keys['timestamp'].n_unique()} months\n"
)
for name in LABEL_NAMES:
    written = labels_df.select([*KEYS, name]).drop_nulls()
    render_quality_report(
        quality_report(
            written,
            name=name,
            key_columns=KEYS,
            expected=expected_keys,
            keys=KEYS,
            entity="symbol",
            session="timestamp",
            expected_missing={
                "trailing": (0, "the outcome horizon is zero, so nothing runs off the end")
            },
        )
    )
    print()

# %% [markdown]
# ### Sign-off
#
# **All three labels cover the panel exactly: 804,530 of 804,530 firm-months, nothing missing and
# nothing unexpected, in every one of the three files.** That is the number a zero outcome horizon
# predicts, and it is worth printing precisely because it is the prediction. Every other case study
# in this book loses its horizon at the end of the sample and a burn-in at the start; this one
# dates each row by the month its return was earned in, so there is no forward window to run off
# the end and no window to fill. A single missing key here would mean a null return in the source
# or a row this stage dropped, and there are none.
#
# **No column crossed a distribution threshold in any of the three.** None is constant and none
# carries a non-finite value. The winsorized variant is a transform of the primary and shares its
# keys exactly, which the identical coverage confirms rather than assumes - a transform that lost
# rows of its own would show a different count. Its distribution is deliberately clipped at the
# bounds declared above, so its tails read tighter than the raw return's by construction; the
# class label concentrates because it is a discretisation, and the zero-share ceiling is loose
# enough not to fire on either.
#
# %% [markdown]
# ## Key takeaways
#
# 1. **Read a pre-built panel's alignment out of the data before relying on it.** A
#    characteristic that restates a past return, correlated against the label at each
#    candidate lag, says which row's outcome the characteristics were recorded before - and
#    shifting a panel that was already shifted fails silently.
# 2. **Assert that a label exists exactly where its outcome was observed.** A count of valid
#    rows reports a fabricated tail and a complete one identically; the null rate read
#    against each entity's own last period separates them.
# 3. **A threshold read across the firms trading in one period can be applied to the whole
#    sample at once; one read along the time axis has to be re-estimated inside each
#    training period.** A month's own median and percentiles use only information that month
#    already carries, so no later month can reach a row through them.
# 4. **A row count only overstates the evidence when forward windows overlap.** Observing a
#    label once per period it measures over leaves each row fully independent, so an
#    effective count must return the row count itself - a case worth checking, because it is
#    the one where the answer is known in advance.
# 5. **Set the comparison with the signal the design names, not the strongest one
#    available.** The number a later feature has to beat is only meaningful if it was fixed
#    before the features existed.
#
# **Known limitations.** The release is anonymized and its firm axis persists only inside
# each published block, so nothing here reconciles against a named security or a
# corporate-action record. The provider computes the return, so its conventions for delisting
# and dividend reinvestment are inherited rather than checked. The label measures close to
# close while the strategy fills at the next open, and this release carries no opening price
# to measure that difference with. The comparison signal is one characteristic of the several
# dozen the release carries.
#
# **Next**: `03_financial_features.py` builds value, quality, investment, momentum and risk
# features from these characteristics, and the composites and interactions over them;
# `04_evaluation.py` is where those features are scored against these labels.

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。