构建月度收益标签:US公司特征
代码 《交易机器学习》
总结
本笔记为横截面 US 公司特征研究定义并验证月度收益标签。笔记检查数据提供方如何将每行特征与收益结果配对,并强调面板数据已经对齐,不应再次移位。主要标签是月度总收益;缩尾收益和月内分类变体则对同一结果进行转换。月度横截面阈值使用其适用月份的信息,而时间序列阈值则需要在每个训练期内估算。
笔记还将缺失结果与各公司的最后观测期进行核对,检查序列相关性和有效样本量,并使用预先指定的信号建立基准比较。笔记记录了收盘到收盘的标签惯例,同时指出策略在下一开盘时交易,因此留下未计量的隔夜缺口。其他限制包括匿名公司历史数据、沿用数据提供方关于股息和退市的惯例,以及仅基于一个特征的基准。所得标签和折分边界旨在为后续特征构建和回测阶段提供依据。
核心观点
- 应先核验预先对齐的特征面板与其观测时间是否一致,再决定是否需要进一步移位收益。
- 月度收益、缩尾收益和类别标签,是对同一结果进行不同转换的表示。
- 在每个月内计算横截面阈值,可避免用后续月份定义该月标签。
- 标签覆盖范围应与观测到的结果可用性一致;若结果窗口重叠,行数统计也应考虑这一点。
- 收盘到收盘标签可能不同于下一开盘执行收益;若源数据无法核实,仍会沿用数据提供方的惯例。
标签
全文
# 02_labels.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # US Firm Characteristics: Label Engineering
#
# Every model in this case study is trained to predict the label defined here, so an error
# in it is silent where it is made and reaches every metric and every backtest after it.
# This notebook fixes what the label is and when it is known, checks against the data that
# the panel is paired the way the provider documents, measures how much independent
# information the rows carry, measures what the simplest signal the study already names
# earns against the label, and writes the label files the rest of the pipeline is built on.
#
# ## Learning objectives
#
# - Work out from the data which month's return a pre-built research panel has already
# paired with each row, instead of taking the provider's description on trust
# - Check that a label is missing on exactly the rows whose outcome was never observed - a
# count of how many rows carry a label cannot tell you that
# - Decide whether a rule that sorts firms into classes needs re-estimating inside each
# training period, or whether the way it is built already keeps later data out of it
# - Measure how much of one row's outcome the next row's label repeats, and turn a row
# count into the number of independent observations it is worth
# - Measure what a signal the study already names earns against the label, so that a
# feature built afterwards is compared against a number fixed before it existed
#
# ## Book reference, prerequisites and artifacts
#
# Chapter 7, Section 7.2. Reads the Chen-Pelger-Zhu panel through
# `load_firm_characteristics()`, whose breadth and cost feasibility
# [`01_feasibility_analysis`](01_feasibility_analysis.ipynb) establishes, and
# `config/setup.yaml`, which declares the label set, the buffer and the holdout boundary.
# Writes `labels/prices.parquet`, which the backtest reads to price positions; one parquet
# per declared label; a small JSON record beside each of those files describing what is in
# it; and `config/cv_config.json`, the committed record of the fold boundaries these labels
# imply.
# %%
"""US Firm Characteristics: Label Engineering."""
import json
from datetime import date
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
import yaml
from matplotlib.ticker import PercentFormatter
from ml4t.diagnostic.metrics import compute_ic_hac_stats, cross_sectional_ic_series
from case_studies.utils.artifact_digest import value_digest, write_artifact
from case_studies.utils.artifact_quality import quality_report, render_quality_report
from case_studies.utils.label_diagnostics import effective_sample_size, panel_autocorrelation
from case_studies.utils.warning_policy import apply_notebook_warning_policy
from data import load_firm_characteristics
from utils.artifact_specs import resolve_label_buffer, resolve_label_horizon
from utils.cv_splits import generate_cv_splits
from utils.paths import display_path, get_case_study_dir
from utils.style import COLORS, FIGSIZE, add_message_title, show_with_alt
apply_notebook_warning_policy()
CASE_STUDY_ID = "us_firm_characteristics"
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
LABELS_DIR = CASE_DIR / "labels"
# %% [markdown]
# The two parameters bound the study window and are both read below. The release starts two
# decades earlier than the window used here; the folds `setup.yaml` declares need a decade
# of training before the first validation year, and this window leaves room for that
# without reaching back into a period whose cross-section is half its later size.
# %% tags=["parameters"]
START_DATE = "1990-01-01"
END_DATE = "2016-12-31"
# %% [markdown]
# ## Configuration
#
# Everything that defines a label is declared in `config/setup.yaml` and bound here. A
# horizon or a boundary typed into a cell is a second copy of a value the rest of the
# pipeline reads from the file, and the two drift apart the first time either is edited.
#
# Three durations are declared, and conflating any two of them is the mistake this panel
# invites. `labels.buffer` is the gap left between a training window and the validation window
# that follows it. `labels.horizons` is how far past its own timestamp a label's outcome
# resolves, which is what the splitter seals the last validation fold on. The span the label
# measures over is a third thing again, and it is what a Newey-West lag and an effective sample
# size are counted in. Here the buffer is one month, the outcome horizon is zero because the
# release dates each row by the month its return was earned in, and the span is the one month
# the buffer was sized to. Section D measures the alignment the middle one rests on.
# %%
SETUP = yaml.safe_load((CASE_DIR / "config" / "setup.yaml").read_text())
PRIMARY_LABEL = SETUP["labels"]["primary"]
LABEL_NAMES = [PRIMARY_LABEL, *SETUP["labels"].get("variants", [])]
WINSORIZED_LABEL, CLASS_LABEL = LABEL_NAMES[1], LABEL_NAMES[2]
OUTCOME_HORIZON = resolve_label_horizon(CASE_STUDY_ID, PRIMARY_LABEL, SETUP)
LABEL_BUFFER = resolve_label_buffer(CASE_STUDY_ID, PRIMARY_LABEL, SETUP)
LABEL_SPAN_MONTHS = int(str(LABEL_BUFFER).rstrip("Mm"))
HOLDOUT_START = date.fromisoformat(str(SETUP["evaluation"]["holdout_start"]))
BASELINE_SIGNAL = SETUP["causal"]["treatment"]
KEYS = ["timestamp", "symbol"]
print(
f"Labels: {LABEL_NAMES}, primary {PRIMARY_LABEL}, span {LABEL_SPAN_MONTHS} month, "
f"buffer {LABEL_BUFFER}, outcome horizon {OUTCOME_HORIZON}\n"
f"Study window {START_DATE} to {END_DATE}, holdout opens {HOLDOUT_START}"
)
# %% [markdown]
# ## A. The learning task
#
# The hypothesis is cross-sectional. Firms differ along accounting ratios, price-based
# measures and turnover proxies, and the claim is that a firm ranked high on those measures
# out-earns a firm ranked low over the following month - ranked against the other firms
# trading that month, not judged in isolation. The label is therefore a firm's monthly total
# return, and the strategy that consumes it is long-short across the cross-section.
#
# The decision cadence comes from `setup.yaml`: the book is set at the month-end close and
# executed at the next open, so one month is both the interval a position is held for and
# the interval an outcome is measured over. Accounting variables are refreshed once a year
# at the end of June and price-based ones every month-end, which is what makes a monthly
# rebalance the fastest cadence the release supports.
#
# Two variants transform the same monthly outcome rather than defining a second one. The
# winsorized return exists because a monthly cross-section of individual stocks carries
# returns of several hundred percent, and a squared-error loss fitted against those spends
# its capacity on them. The classification label turns the same return into a rank question
# - is this firm in the better half of this month - which is closer to what a long-short
# book acts on than the return's magnitude is.
# %% [markdown]
# ## B. Preparation before the label
#
# **The panel is already aligned, and the mistake this dataset invites is aligning it
# again.** Chen, Pelger and Zhu publish each row as a pair: the characteristics a firm
# carried at the end of one month, and the return that firm went on to earn over the
# following month. The row is dated by the month the return was earned in, so the
# information a model reads is a month older than the outcome it predicts, and shifting
# `ret` forward once more would pair a firm's characteristics with a return earned two
# months later. Section D reads that alignment out of the data rather than accepting it
# here.
#
# Nothing else is transformed. The characteristics arrive cross-sectionally rank-transformed
# by the provider, and the return is the raw monthly total return. Which firms are in the
# universe is decided by the provider's completeness rule - a firm-month is kept only where
# every characteristic is available - so that screen runs before this notebook, and neither
# this notebook nor stage 03 drops a further row.
#
# Where a label is built by shifting a price series rather than read from a column, the order
# of those two steps matters and getting it wrong is silent. A shift counts rows, so dropping
# ineligible rows first makes the shift step over whatever was removed: a one-month label then
# spans however long the gap was, and still looks like a full column. Applying the screen after
# the label, or expressing the horizon as a length of time so that a gap produces a null,
# avoids it.
# %%
window = pl.col("timestamp").is_between(
pl.lit(START_DATE).str.to_date(), pl.lit(END_DATE).str.to_date()
)
firm_chars = load_firm_characteristics(split="all").filter(window).sort(["symbol", "timestamp"])
CHARACTERISTICS = [c for c in firm_chars.columns if c not in {*KEYS, "ret", "split"}]
# Recorded as the `inputs` of every artifact written below: a re-run against a refreshed
# release is otherwise indistinguishable from this one.
MARKET_DATA_DIGEST = value_digest(firm_chars, [*KEYS, "ret"])
print(f"{firm_chars['symbol'].n_unique():,} firms, {firm_chars.height:,} firm-months")
print(
f"Month-ends {firm_chars['timestamp'].min()} to {firm_chars['timestamp'].max()}, "
f"{len(CHARACTERISTICS)} characteristics, {firm_chars['ret'].null_count()} rows without a return"
)
print(f"market_data digest: {MARKET_DATA_DIGEST}")
# %% [markdown]
# ## C. Label construction
#
# One execution convention, written once and inherited by all three labels:
#
# $$r_{i,t} = \frac{P_{i,t}}{P_{i,t-1}} - 1$$
#
# where $P_{i,t}$ is firm $i$'s adjusted month-end close in month $t$, and the row carrying
# $r_{i,t}$ also carries the characteristics observed at the close of month $t-1$. That is
# Chapter 7.2's close-to-close convention, with the decision taken at the earlier of the two
# closes. The strategy that consumes the label fills at the next open instead, as
# `setup.yaml` declares, so the label omits whatever the price moves overnight between the
# decision close and that open. Nothing in this release can measure that gap - it publishes
# no opening price - so it is inherited as a known difference between the label and the
# traded return rather than corrected here.
#
# The provider computes $r_{i,t}$, so the labelling library has nothing to shift and is not
# called: `fixed_time_horizon_labels` divides a price column by its own lag, and this release
# publishes no price column. The two variants are cross-sectional transforms of the primary
# label, each computed inside one month.
#
# **Both thresholds are read off the cross-section of the month they apply to**, which is
# what makes them point in time. Chapter 7.2 draws the line here: a percentile taken across
# the firms trading in one month uses only information that month's close already carries,
# while a percentile taken along the time axis reads later months and has to be fitted inside
# the training fold. A median split also holds its class proportions near one half by
# construction, which Section E measures rather than assumes.
#
# A label resolves on the month-end that dates the row carrying it, because the return the
# row reports was earned over the month ending there. The date a label resolves on is what
# decides whether it may be looked at: a row whose outcome lands on or after
# `evaluation.holdout_start` falls inside the period held back for the final test, so it is
# written to the label files but excluded from every measurement below. The files themselves
# keep every row, because the later stages need the held-back months to score on.
#
# The `month` column numbers the panel's own month-end grid. Sections D and F count on it
# rather than on position among the rows, so that two rows either side of a month a firm
# skipped are not read as consecutive observations.
# %%
month_ends = firm_chars.select("timestamp").unique().sort("timestamp").with_row_index("month")
bounds = firm_chars.group_by("timestamp").agg(
pl.col("ret").median().alias("_median"),
pl.col("ret").quantile(0.01).alias("_lo"),
pl.col("ret").quantile(0.99).alias("_hi"),
)
labels_df = (
firm_chars.join(month_ends, on="timestamp", how="left")
.join(bounds, on="timestamp", how="left")
.with_columns(
pl.col("timestamp").alias("_label_end"),
pl.col("ret").alias(PRIMARY_LABEL),
pl.col("ret").clip(pl.col("_lo"), pl.col("_hi")).alias(WINSORIZED_LABEL),
pl.when(pl.col("ret").is_not_null())
.then(pl.col("ret") > pl.col("_median"))
.cast(pl.Int32)
.alias(CLASS_LABEL),
)
)
dev = labels_df.filter(pl.col("_label_end") < HOLDOUT_START)
print(
f"Constructed {', '.join(LABEL_NAMES)} on {labels_df.height:,} firm-months, "
f"{dev.height:,} of them development rows through {dev['timestamp'].max()}: "
f"{dev['timestamp'].n_unique()} month-ends, {dev['symbol'].n_unique():,} firms"
)
# %% [markdown]
# ## D. Window validity
#
# A column of the right length always arrives; the question is whether what it holds is the
# quantity the label claims. Each property below fails silently and leaves plausible numbers
# behind, so each is asserted rather than described.
#
# The first assertion is the one that catches a fabricated tail. A firm's last month in the
# panel closes its label window, and a construction that shifted a price series would leave
# that month with no outcome - which has to be a null, never a value. The rest check the two
# monthly joins: a duplicated month-end in either would multiply the panel's rows, and a
# clipped return falling outside its own month's percentiles would mean a join had matched
# the wrong month.
# %%
missing = labels_df["ret"].is_null()
inside = pl.col(WINSORIZED_LABEL).is_between(pl.col("_lo"), pl.col("_hi"))
for name in LABEL_NAMES:
# 1. Null exactly where the outcome was not observed.
assert (labels_df[name].is_null() == missing).all(), name
# 2. One row per firm-month, so neither monthly join fanned the panel out.
assert labels_df.select(pl.struct(KEYS).n_unique()).item() == labels_df.height
# 3. Each label is a transform of its own row's return, inside its own month.
assert labels_df.filter(pl.col(PRIMARY_LABEL) != pl.col("ret")).height == 0
assert labels_df.filter(~inside).height == 0
# 4. No discrete label is derived from a null return.
assert labels_df.filter(missing & pl.col(CLASS_LABEL).is_not_null()).height == 0
clipped = labels_df.filter(pl.col(PRIMARY_LABEL) != pl.col(WINSORIZED_LABEL)).height
skipped = labels_df.select((pl.col("month").diff().over("symbol") > 1).sum()).item()
print(
f"{clipped:,} of {labels_df.height:,} returns clipped by their own month's percentiles; "
f"{skipped:,} places where a firm skips a month-end inside its own history"
)
# %% [markdown]
# Section B's claim about the alignment is what every number downstream rests on, and it can
# be read out of the release itself. `ST_REV` is the short-term-reversal characteristic: a
# firm's own most recent monthly return, ranked across firms, as of the date the
# characteristics were observed. The rank correlation between `ST_REV` and the label is
# therefore a test of which row's return the characteristics were recorded after, and only
# one lag can carry it. Pairs are restricted to rows one month-end apart, so a firm that
# skips a month contributes none.
# %%
step = pl.col("month").diff().over("symbol")
alignment = dev.with_columns(
pl.when(step == 1).then(pl.col(PRIMARY_LABEL).shift(1).over("symbol")).alias("_previous"),
pl.when(step.shift(-1).over("symbol") == 1)
.then(pl.col(PRIMARY_LABEL).shift(-1).over("symbol"))
.alias("_next"),
)
for described, column in (
("the previous month's", "_previous"),
("this row's own", PRIMARY_LABEL),
("the next month's", "_next"),
):
paired = alignment.drop_nulls(["ST_REV", column])
value = paired.select(pl.corr("ST_REV", column, method="spearman")).item()
print(f"ST_REV against {described:>21} return: rank correlation {value:+.3f}")
# %% [markdown]
# Position zero below is each firm's last month in the panel. Every label has to be present
# there, because the outcome that month reports was earned before the row was written. The
# grey line is the same panel under one forward shift, and it is the failure this figure
# exists to make visible: a scalar count of valid rows reports both shapes as complete.
#
# The horizontal axis counts month-ends on the panel's own grid, not rows, so a firm that
# is absent for a month reads as two month-ends apart and not as one.
# %%
profile = (
labels_df.with_columns(
(pl.col("month").max().over("symbol") - pl.col("month")).alias("from_end"),
pl.col(PRIMARY_LABEL).shift(-LABEL_SPAN_MONTHS).over("symbol").alias("_shifted"),
)
.filter(pl.col("from_end") <= LABEL_SPAN_MONTHS + 4)
.group_by("from_end")
.agg(pl.col(c).is_not_null().mean() for c in [*LABEL_NAMES, "_shifted"])
.sort("from_end")
)
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
back = profile["from_end"]
markers = zip(LABEL_NAMES, (COLORS["blue"], COLORS["amber"], COLORS["copper"]), "os^", strict=True)
for name, color, marker in markers:
ax.plot(back, profile[name], marker, ms=8, mfc="none", c=color, label=name)
ax.plot(
back, profile["_shifted"], "--", ds="steps-mid", lw=1.4, c=COLORS["neutral"], label="shifted"
)
ax.set(
xlabel="Month-ends back from each firm's last month in the panel",
ylabel="Share of firms carrying a label",
ylim=(-0.08, 1.12),
)
add_message_title(
ax,
"Every firm's last month in the panel carries a label",
subtitle="All three labels sit on one another; one forward shift empties position zero",
)
ax.legend(loc="lower right", frameon=False, ncol=2)
show_with_alt(fig, "Non-null label rate by month-ends back from each firm's last observation.")
# %% [markdown]
# ## E. Distribution and base rate
#
# What scale is the label, and does it mean the same thing across regimes?
#
# Both continuous labels go on one axis with identical bins and a logarithmic count axis.
# The claim the figure has to support is about shape rather than width - two standard
# deviations would carry the width - and the shape here is that the two series are
# indistinguishable through the body and separate only in the tails, where the winsorized
# label stops at the widest boundary any month produced and the raw one runs on. The axis
# has to reach past that widest boundary or the figure is cut off exactly where the two
# labels part; it still stops well short of the raw label's largest return, and the rows
# past it are counted underneath rather than drawn.
# %%
bins = np.linspace(-1.0, 2.0, 91)
histograms = {
PRIMARY_LABEL: dict(color=COLORS["neutral"], histtype="step", lw=2, zorder=2),
WINSORIZED_LABEL: dict(color=COLORS["amber"], alpha=0.8, zorder=1),
}
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
for name, style in histograms.items():
series = dev[name]
tag = f"{name}, std {series.std():.3f}, kurtosis {series.kurtosis():.1f}"
ax.hist(series.to_numpy(), bins=bins, label=tag, **style)
ax.axvline(0, color=COLORS["neutral"], linestyle="--", lw=0.8)
ax.set(xlabel="One-month total return", ylabel="Firm-months per bin, log scale", yscale="log")
ax.xaxis.set_major_formatter(PercentFormatter(1.0))
add_message_title(
ax,
"Winsorizing at the monthly percentiles pulls in only the tails",
subtitle="Identical bins, development window; rows beyond the axis are counted below",
)
ax.legend(loc="upper right", frameon=False, fontsize=8)
show_with_alt(fig, "Histograms of the raw and winsorized labels on identical bins, log counts.")
for name in (PRIMARY_LABEL, WINSORIZED_LABEL):
series, drawn = dev[name], dev[name].is_between(bins[0], bins[-1]).sum()
print(
f"{name}: mean {series.mean():+.5f}, std {series.std():.5f}, "
f"kurtosis {series.kurtosis():.2f}, range {series.min():.3f} to {series.max():.3f}, "
f"{series.len() - drawn:,} rows beyond the axis"
)
# %% [markdown]
# Chapter 7.2 asks for the scale and the base rate to be tracked through time. For the
# continuous labels the quantity that has to hold up is the spread the model ranks within:
# where it narrows, the same rank correlation buys less return. That spread is taken across
# firms on each month-end first and only then averaged over the year, because pooling every
# firm-month in a year into one standard deviation adds the market's own move from month to
# month to the spread across firms, and a ranking model is scored on the second alone. For
# the classification label the quantity is the share of firms in the upper class, and the
# lower panel is what shows whether the median split delivered the balance it promises.
# %%
monthly = dev.group_by("timestamp").agg(
pl.col(PRIMARY_LABEL).std().alias("dispersion"),
pl.col(CLASS_LABEL).mean().alias("upper_share"),
(pl.col("ret") == pl.col("_median")).mean().alias("tied"),
)
annual = (
monthly.group_by(pl.col("timestamp").dt.year().alias("year"))
.agg(pl.col("dispersion").mean(), pl.col("upper_share").mean())
.sort("year")
)
fig, axes = plt.subplots(2, 1, sharex=True, figsize=FIGSIZE["dual_v"])
panels = (
("dispersion", COLORS["blue"], annual["dispersion"].median(), "Cross-firm return std"),
("upper_share", COLORS["amber"], 0.5, "Share in the upper class"),
)
for ax, (column, color, reference, ylabel) in zip(axes, panels, strict=True):
ax.plot(annual["year"], annual[column], "o-", ms=4, lw=1.8, c=color)
ax.axhline(reference, color=COLORS["neutral"], ls="--", lw=1.1)
ax.set_ylabel(ylabel)
ax.yaxis.set_major_formatter(PercentFormatter(1.0))
# Bound the lower panel symmetrically about the even split, by the series it actually draws.
# Centring keeps the distance from balance readable as a distance; taking the bound from the
# monthly range instead would leave most of the panel blank and flatten the annual series.
reach = (annual["upper_share"] - 0.5).abs().max() * 1.3
axes[1].set(ylim=(0.5 - reach, 0.5 + reach), xlabel="Year")
add_message_title(
axes[0],
"Dispersion shifts across regimes while the median split stays balanced",
subtitle="Annual means of the monthly cross-section; dashed lines mark the typical year's "
"dispersion and an even split",
)
show_with_alt(fig, "Annual mean cross-firm return dispersion above, upper-class share below.")
peak, trough = (annual.sort("dispersion", descending=d).row(0, named=True) for d in (True, False))
print(
f"dispersion peaks at {peak['dispersion']:.1%} in {peak['year']:.0f} against "
f"{trough['dispersion']:.1%} in {trough['year']:.0f}; median year "
f"{annual['dispersion'].median():.1%}\n"
f"upper class runs {monthly['upper_share'].min():.3f} to {monthly['upper_share'].max():.3f} "
f"per month, mean {monthly['upper_share'].mean():.3f}; the shortfall is returns tied at the "
f"median, up to {monthly['tied'].max():.3f} of a month's firms"
)
# %% [markdown] tags=["results"]
# On the development window the raw label has a standard deviation of 0.17430 and a kurtosis
# of 335.36, against 0.14850 and 6.47 for the winsorized variant: clipping at each month's
# own percentiles touches a fiftieth of the rows and removes almost all of the fourth moment.
# Cross-firm dispersion is not stable, peaking at 22.4% in 2000 against 11.8% in 2013 and a
# median year of 15.3%, while the classification base rate is, running from 0.438 to 0.500 a
# month around a mean of 0.498. The shortfall below an even split is returns tied at the
# month's median, which reach 0.084 of the firms in a month.
# %% [markdown]
# ## F. Overlap and effective sample size
#
# Where a label is sampled more often than the period it measures over, consecutive rows
# report overlapping stretches of the same price path, and the row count then claims more
# evidence than the data holds. Here a firm is observed once a month and each label measures
# one month, so consecutive labels are built from returns that share nothing. Two
# measurements say so in different units: the autocorrelation profile, and the row count
# after each row is discounted by how much of its window another row also covers - the
# average-uniqueness weighting of Chapter 7.2.
#
# Both read `month`, each row's position on the panel's own month-end grid, rather than its
# position among the rows that reach this point. A firm absent for a month would otherwise
# have the rows either side of the gap treated as consecutive, pairing windows that in fact
# share nothing.
# %%
MAX_LAG = 12
acf = {n: panel_autocorrelation(dev, n, max_lag=MAX_LAG, bar_col="month") for n in LABEL_NAMES}
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
lags = np.arange(1, MAX_LAG + 1)
palette = (COLORS["blue"], COLORS["amber"], COLORS["copper"])
for name, color in zip(LABEL_NAMES, palette, strict=True):
ax.plot(lags, acf[name], "o-", ms=3, lw=1.6, c=color, label=name)
ax.axhline(0, color=COLORS["neutral"], lw=0.8)
ax.set(xlabel="Lag in month-ends", ylabel="Panel autocorrelation")
add_message_title(
ax,
"Monthly labels carry weak reversal rather than an overlap decay",
subtitle="Consecutive labels share no return interval, so every lag drawn is already "
"overlap-free",
)
ax.legend(loc="lower right", frameon=False)
show_with_alt(fig, "Panel autocorrelation of all three labels against lag in month-ends.")
# A horizon-h label consumes the h returns realised over its window and its neighbour one
# period later shares h-1 of them, so at a one-period horizon nothing is shared at all.
n_rows, n_eff = effective_sample_size(dev, horizon=LABEL_SPAN_MONTHS, bar_col="month")
assert n_eff == n_rows, "a one-month label overlaps nothing, so every row must weigh one"
print(
f"{PRIMARY_LABEL}: N={n_rows:,}, N_eff={n_eff:,.0f}, ratio {n_eff / n_rows:.4f}; "
f"autocorrelation {acf[PRIMARY_LABEL][0]:+.4f} at lag one, "
f"{acf[PRIMARY_LABEL][-1]:+.4f} at lag {MAX_LAG}"
)
# %% [markdown] tags=["results"]
# The development window's 780,882 firm-months carry 780,882 effective observations: a
# one-month label sampled monthly shares no return interval with its neighbour, so average
# uniqueness is one and the ratio is 1.0000. What autocorrelation remains is the return
# process rather than the label construction - -0.0227 at one month and -0.0045 at twelve -
# and it is the short-term reversal `ST_REV` already carries as a characteristic. The purge
# gap a fold needs is set by the forward window, which is one month here.
# %% [markdown]
# ## G. Baseline floor
#
# One signal, measured against the primary label over the development months only and with
# no feature engineering: the raw momentum characteristic `setup.yaml` names as the treatment
# whose effect `09_causal_dml.py` estimates, taken exactly as the provider ranked it. This is
# the number a feature built in the next notebook has to beat, and it is only worth beating
# if it was fixed before any feature existed.
#
# The information coefficient is the rank correlation between the signal and the outcome
# across the firms trading in one month, computed for each month-end and then averaged. That
# is the quantity a ranking model is scored on; pooling every firm-month into one correlation
# instead answers a different question, mixing how firms rank against each other with how the
# market moved from month to month. The library call returns the series ordered by time,
# which the standard error below depends on.
#
# A month with too thin a cross-section produces a rank correlation too noisy to average in,
# so months below a minimum count are left out. That minimum is set at half the median
# month's firm count rather than at a fixed number, so the same rule transfers to a universe
# of a different size. The standard error is Newey-West adjusted, which allows for one
# month's information coefficient being related to the next month's; the count the statistic
# reports is the number of month-ends that actually entered it.
# %%
baseline = dev.drop_nulls([BASELINE_SIGNAL, PRIMARY_LABEL])
min_obs = int(baseline.group_by("timestamp").len()["len"].median() // 2)
ic = cross_sectional_ic_series(
baseline,
baseline,
pred_col=BASELINE_SIGNAL,
ret_col=PRIMARY_LABEL,
date_col="timestamp",
entity_col="symbol",
min_obs=min_obs,
).sort("timestamp") # HAC autocovariances are meaningless over a permutation of time
stats = compute_ic_hac_stats(ic, ic_col="ic", label_horizon=LABEL_SPAN_MONTHS)
# `ic` carries one row per month-end whatever its cross-section, with the correlation left
# null below `min_obs`; the statistic drops those, so its own count is what is reported.
print(
f"Baseline: {BASELINE_SIGNAL} against {PRIMARY_LABEL}, minimum cross-section {min_obs} firms\n"
f" month-ends entering the statistic {stats['n_periods']:,} "
f"of {dev['timestamp'].n_unique()}\n"
f" mean IC {stats['mean_ic']:+.5f}, HAC standard error {stats['hac_se']:.5f}\n"
f" HAC t {stats['t_stat']:+.2f} on {stats['effective_lags']} Bartlett lags, "
f"naive t {stats['naive_t_stat']:+.2f}, p {stats['p_value']:.3g}"
)
# %% [markdown] tags=["results"]
# Raw momentum earns a mean information coefficient of +0.04398 against the monthly label
# over all 312 development month-ends, on a cross-section of at least 1293 firms. Under the
# naive standard error that is a t-statistic of +6.89; the Newey-West rule picks 5 lags here,
# above the none a one-month horizon requires on its own, and the HAC statistic is +6.14 with
# a standard error of 0.00716. So a feature is not compared against zero: it has to improve
# on a mean information coefficient of 0.04398 that the data already separates from zero,
# fixed here before any feature exists.
# %% [markdown]
# ## H. Artifacts and the audit record
#
# Every file written here gets a small JSON record saved next to it. Its job is to let a
# later reader tell one version of a file apart from another without opening either, and to
# say where the contents came from. It holds a digest - a short string computed from the
# values in the file, which changes if any of them change - together with the row count, the
# columns that identify a row, the notebook that wrote the file, and the digest of each file
# or dataset it was built from. That last field is what ties a label back to the release of
# the data it was computed from: without it, a re-run against a refreshed download looks
# exactly like a re-run against the same one.
#
# The characteristic panel goes out first, because the backtest prices this case study's
# positions from it and all three labels derive from its `ret` column; the two variants
# additionally carry the primary label's digest, which is what they transform.
# %%
prices = write_artifact(
firm_chars.select([*KEYS, "ret", *CHARACTERISTICS]),
LABELS_DIR / "prices.parquet",
keys=KEYS,
written_by="02_labels",
inputs={"market_data": MARKET_DATA_DIGEST},
)
print(f"prices.parquet: {prices['n_rows']:,} rows, digest {prices['digest']}")
records: dict[str, dict] = {}
for name in LABEL_NAMES:
derived = {} if name == PRIMARY_LABEL else {PRIMARY_LABEL: records[PRIMARY_LABEL]["digest"]}
records[name] = write_artifact(
labels_df.select([*KEYS, name]).drop_nulls(),
LABELS_DIR / f"{name}.parquet",
keys=KEYS,
written_by="02_labels",
inputs={"prices": prices["digest"], **derived},
)
print(f"{name}.parquet: {records[name]['n_rows']:,} rows, digest {records[name]['digest']}")
# %% [markdown]
# The training and validation periods the model stages use are worked out per label by
# `case_studies/utils/cv_window.py`, from `config/setup.yaml` and the range of dates in the
# label file written above - so which rows land in that file is what decides where the period
# boundaries fall. Every stage that needs those periods derives them the same way rather than
# reading a file an earlier notebook happened to write, so none of them depends on the order the
# pipeline was run in. The cell below runs the generator once more and saves the result to
# `config/cv_config.json` as a committed record of the geometry these labels imply, which is what
# makes a change in it show up in a diff.
#
# The two durations the generator is given are the ones separated above. `label_buffer` sets the
# gap between a training window and its validation window. `outcome_horizon` is how far past the
# last validation date an outcome is still unresolved, and it is what decides how much of the
# last fold has to be given back before the holdout opens - zero here, because a row's return is
# realised on the timestamp the row carries.
# %%
splits = generate_cv_splits(
labels_df.select("timestamp"),
case_study_id=CASE_STUDY_ID,
label_buffer=LABEL_BUFFER,
outcome_horizon=OUTCOME_HORIZON,
)
assert len(splits) == int(SETUP["evaluation"]["n_splits"])
assert max(split["val_end"] for split in splits).date() < HOLDOUT_START
BLOCKS = ("train_start", "train_end", "val_start", "val_end")
cv_config = {
"case_study_id": CASE_STUDY_ID,
"n_splits": len(splits),
"train_size": str(SETUP["evaluation"]["train_size"]),
"val_size": str(SETUP["evaluation"]["val_size"]),
"holdout_start": HOLDOUT_START.isoformat(),
"holdout_end": str(SETUP["evaluation"]["holdout_end"]),
"splits": [
{"fold": int(split["fold"]), "label_buffer": LABEL_BUFFER}
| {key: split[key].date().isoformat() for key in BLOCKS}
for split in splits
],
}
cv_path = CASE_DIR / "config" / "cv_config.json"
cv_path.write_text(json.dumps(cv_config, indent=2) + "\n")
last_val = max(split["val_end"] for split in splits).date()
print(f"Saved {display_path(cv_path)}: {len(splits)} folds, validation through {last_val}")
# %% [markdown]
# The record Chapter 7.2 requires to close a label definition, one row per label, built from
# the values computed above rather than written by hand.
# %%
shared = "cv_window.py, for this label's training and validation periods; the model stages"
readers = {PRIMARY_LABEL: f"{shared}; 03_financial_features.py, to check firm identities match"}
print("\nLabel audit record")
for name in LABEL_NAMES:
series = dev[name]
base_rate = (
f"upper class {series.mean():.3f} of firms"
if name == CLASS_LABEL
else f"mean {series.mean():+.5f}, std {series.std():.5f}"
)
print(
f"\n{name}\n anchor the adjusted close of the month before the one dating the row"
f"\n span {LABEL_SPAN_MONTHS} month, the month the row is dated by"
f"\n resolution fixed at that month's close, the timestamp the row itself carries;"
f" monthly bars need no intraday tie-break"
f"\n overlap {LABEL_SPAN_MONTHS - 1} months shared by consecutive rows"
f"\n base rate {base_rate}"
f"\n consumed by {readers.get(name, shared)}"
)
# %% [markdown]
# ## What the labels hold, and what they owe
#
# Two questions about the files this stage just wrote, and the rows that are there answer only one
# of them. The first is what is in each column - nulls, how much sits at exactly zero, how far the
# extreme values are from the body, whether anything is constant. A threshold crossed there asks
# for a sentence and settles nothing on its own: a winsorized return is *supposed* to pile up at
# its bounds, and a class label is supposed to concentrate.
#
# The second is coverage, and it needs a denominator that is not the labels themselves. **The
# reference is `firm_chars`** - every firm-month the panel carries, which is the set a label could
# in principle have been written for. Comparing one label to the other two would hide any month
# where all three are absent together; comparing them to the panel cannot.
#
# What each label is entitled to be short by is unusually easy to state here, and unusually strict:
# **the outcome horizon is zero.** A row is dated by the month its return was earned in rather
# than by a month whose outcome is still ahead of it, so there is no forward window running off
# the end of the sample and no burn-in before a window fills. Every one of these labels owes a
# value on every firm-month the panel carries a return for, and anything missing is either a null
# return in the source or something this stage did.
# %%
expected_keys = firm_chars.select(KEYS).unique()
print(
f"panel: {expected_keys.height:,} (firm, month) keys across "
f"{expected_keys['symbol'].n_unique()} firms and "
f"{expected_keys['timestamp'].n_unique()} months\n"
)
for name in LABEL_NAMES:
written = labels_df.select([*KEYS, name]).drop_nulls()
render_quality_report(
quality_report(
written,
name=name,
key_columns=KEYS,
expected=expected_keys,
keys=KEYS,
entity="symbol",
session="timestamp",
expected_missing={
"trailing": (0, "the outcome horizon is zero, so nothing runs off the end")
},
)
)
print()
# %% [markdown]
# ### Sign-off
#
# **All three labels cover the panel exactly: 804,530 of 804,530 firm-months, nothing missing and
# nothing unexpected, in every one of the three files.** That is the number a zero outcome horizon
# predicts, and it is worth printing precisely because it is the prediction. Every other case study
# in this book loses its horizon at the end of the sample and a burn-in at the start; this one
# dates each row by the month its return was earned in, so there is no forward window to run off
# the end and no window to fill. A single missing key here would mean a null return in the source
# or a row this stage dropped, and there are none.
#
# **No column crossed a distribution threshold in any of the three.** None is constant and none
# carries a non-finite value. The winsorized variant is a transform of the primary and shares its
# keys exactly, which the identical coverage confirms rather than assumes - a transform that lost
# rows of its own would show a different count. Its distribution is deliberately clipped at the
# bounds declared above, so its tails read tighter than the raw return's by construction; the
# class label concentrates because it is a discretisation, and the zero-share ceiling is loose
# enough not to fire on either.
#
# %% [markdown]
# ## Key takeaways
#
# 1. **Read a pre-built panel's alignment out of the data before relying on it.** A
# characteristic that restates a past return, correlated against the label at each
# candidate lag, says which row's outcome the characteristics were recorded before - and
# shifting a panel that was already shifted fails silently.
# 2. **Assert that a label exists exactly where its outcome was observed.** A count of valid
# rows reports a fabricated tail and a complete one identically; the null rate read
# against each entity's own last period separates them.
# 3. **A threshold read across the firms trading in one period can be applied to the whole
# sample at once; one read along the time axis has to be re-estimated inside each
# training period.** A month's own median and percentiles use only information that month
# already carries, so no later month can reach a row through them.
# 4. **A row count only overstates the evidence when forward windows overlap.** Observing a
# label once per period it measures over leaves each row fully independent, so an
# effective count must return the row count itself - a case worth checking, because it is
# the one where the answer is known in advance.
# 5. **Set the comparison with the signal the design names, not the strongest one
# available.** The number a later feature has to beat is only meaningful if it was fixed
# before the features existed.
#
# **Known limitations.** The release is anonymized and its firm axis persists only inside
# each published block, so nothing here reconciles against a named security or a
# corporate-action record. The provider computes the return, so its conventions for delisting
# and dividend reinvestment are inherited rather than checked. The label measures close to
# close while the strategy fills at the next open, and this release carries no opening price
# to measure that difference with. The comparison signal is one characteristic of the several
# dozen the release carries.
#
# **Next**: `03_financial_features.py` builds value, quality, investment, momentum and risk
# features from these characteristics, and the composites and interactions over them;
# `04_evaluation.py` is where those features are scored against these labels.
```在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。