ऑप्शन संकेतों से फ़ॉरवर्ड इक्विटी रिटर्न लेबल तैयार करना
सारांश
यह नोटबुक उस रणनीति के लिए फ़ॉरवर्ड-रिटर्न लेबल बनाती है जो सूचीबद्ध ऑप्शन संकेत पढ़ती है और अंतर्निहित शेयरों में ट्रेड करती है। यह दैनिक शेयर कीमतों को स्प्लिट और लाभांश के लिए समायोजित करती है, स्थायी प्रतिभूति पहचानकर्ता का उपयोग करती है और इच्छित प्रवेश तथा निकास पर निष्पादन योग्य कीमतों के बीच रिटर्न परिभाषित करती है। यह जाँचती है कि हर फ़ॉरवर्ड विंडो पूर्ण हो, एक ही प्रतिभूति के भीतर रहे और उपलब्ध पंक्तियों पर केवल आगे खिसकने के बजाय बाज़ार सत्र गिने। लेबल सेट में अलग-अलग होल्डिंग अवधियों के लिए रिटर्न, जोखिम-स्केल किए गए और दिशात्मक विकल्प शामिल हैं।
नोटबुक निदान गणनाओं को उन परिणामों तक सीमित रखती है जो होल्डआउट से पहले तय हो चुके हों, प्रभावी सैंपल आकार के अनुमान में लगातार फ़ॉरवर्ड विंडो के ओवरलैप का हिसाब रखती है और कच्चे ऑप्शन-आधारित संकेत की लेबल से तुलना करती है। यूनिवर्स ऑप्शन डेटा वाली प्रतिभूतियों तक सीमित है और समय-बिंदु पात्रता आगे की प्रक्रिया में संभाली जाती है। ओवरलैप करती पंक्तियाँ अधिक अवलोकन देती हैं, लेकिन कम स्वतंत्र जानकारी। लेबल लाभांश सहित कुल शेयर रिटर्न दर्शाते हैं, और आधाररेखा संकेत हर स्टॉक की सामान्य वोलैटिलिटी के लिए समायोजित नहीं है; ये विकल्प तय करते हैं कि बाद के मॉडल परिणामों की व्याख्या कैसे करें।
मुख्य विचार
- फ़ॉरवर्ड रिटर्न लेबल में स्प्लिट और लाभांश शामिल करने के लिए समायोजित शेयर कीमतों का उपयोग करें।
- हर लेबल विंडो को एक स्थायी प्रतिभूति पहचानकर्ता के भीतर रखें और उसकी अवधि ट्रेडिंग सत्रों में मापें।
- दिशात्मक वर्ग बनाते समय अधूरे फ़ॉरवर्ड रिटर्न से बचाव करें, ताकि अनुपस्थित परिणाम बिना लेबल रहें।
- विकास-चरण के निदान को इतना पहले समाप्त करें कि होल्डआउट शुरू होने से पहले हर लेबल का परिणाम ज्ञात हो।
- ओवरलैप करती फ़ॉरवर्ड विंडो कच्ची पंक्ति संख्या में निहित स्वतंत्र जानकारी घटाती हैं।
टैग
पूरा पाठ
# S&P 500 Equity + Option Analytics: Label Engineering
# S&P 500 Equity + Option Analytics: Label Engineering
The strategy reads a signal off a name's listed options and holds the share, so what every
model here predicts is a forward return on the share. This notebook builds the price series
that return is measured on, states which two prices a return is taken between, checks that
every labelled row has a complete forward window inside one security, measures how much
independent information those rows carry once consecutive windows overlap, measures what the
simplest available signal already earns against the label, and writes one file per label.
## Learning objectives
- Turn raw daily bars into a series a return can be taken on, using the adjustment factor
that carries splits and dividends, and the identifier that stays with a company when its
ticker changes
- State a forward return as the pair of prices a strategy could have traded at, then check in
code that every window is complete and stays inside one security
- Derive an up-or-down label from a return without inventing a class for the rows the return
itself could not fill
- Cut a diagnostic off at the date a label's outcome is known rather than the date its signal
is read, so that nothing measured here has seen the holdout
- Measure what one unmodified signal earns against the label, under a standard error that
allows for consecutive labels sharing most of their forward window
## Book reference, prerequisites and artifacts
Chapter 7, Section 7.2. Reads the daily share bars and the option surface through
`load_sp500_daily_bars()` and `load_sp500_options_surface()`, whose coverage
[`01_feasibility_analysis`](01_feasibility_analysis.ipynb) establishes, and
`config/setup.yaml`, which declares the label set, each label's horizon and buffer, and the
holdout boundary. Writes one parquet per declared label into `labels/`, each with a
`.digest.json` sidecar beside it. [`05_evaluation`](05_evaluation.ipynb) reads the primary
label and scores the engineered features against it, the model stages read whichever label
they train on, and `case_studies/utils/cv_window.py` reads each file's own timeline to place
the folds those models are trained and validated on.
```python
"""S&P 500 Equity + Option Analytics - Label Engineering."""
from datetime import date
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
import yaml
from ml4t.diagnostic.metrics import compute_ic_hac_stats, cross_sectional_ic_series
from case_studies.utils.artifact_digest import value_digest, write_artifact
from case_studies.utils.artifact_quality import quality_report, render_quality_report
from case_studies.utils.label_diagnostics import effective_sample_size, panel_autocorrelation
from case_studies.utils.warning_policy import apply_notebook_warning_policy
from data import load_sp500_daily_bars, load_sp500_options_surface
from utils.artifact_specs import resolve_label_buffer, resolve_label_horizon
from utils.paths import get_case_study_dir
from utils.style import COLORS, FIGSIZE, add_message_title, show_with_alt
apply_notebook_warning_policy()
```
The share bars and the option surface are licensed extracts covering these five years, so the
defaults span the whole of the data. Narrowing the window shortens a run, at the cost of less
history behind the trailing volatility in Section C and fewer years in the development window
that everything from Section E on is computed over.
```python
START_DATE = "2017-01-01"
END_DATE = "2021-12-31"
```
## Configuration
Everything that defines a label is declared in `config/setup.yaml` and bound here. A horizon
or a boundary typed into a cell is a second copy of a value the rest of the pipeline reads
from the file, and the two drift apart the first time either is edited.
Two quantities are declared under separate keys here, and most case studies hold them equal.
The **outcome horizon** is how long a position is held, and it decides which two prices a
label is measured between. The **buffer** is the gap the walk-forward split leaves between the
end of a training fold and the start of the validation fold after it, counted in trading
sessions because the evaluation calendar is NYSE. A buffer shorter than the horizon would put
a training row whose outcome is still unresolved next to the validation window that resolves
it. The primary label carries ten sessions of buffer against a five-session horizon, twice
what independence requires; each variant carries exactly its own horizon.
`resolve_label_horizon` and `resolve_label_buffer` read the two from their separate keys, and
Section H records both against every label.
```python
CASE_STUDY_ID = "sp500_equity_option_analytics"
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
LABELS_DIR = CASE_DIR / "labels"
setup = yaml.safe_load((CASE_DIR / "config" / "setup.yaml").read_text())
PRIMARY_LABEL = setup["labels"]["primary"]
LABEL_NAMES = [PRIMARY_LABEL, *setup["labels"]["variants"]]
HORIZONS = {
n: int(resolve_label_horizon(CASE_STUDY_ID, n, setup).rstrip("Dd")) for n in LABEL_NAMES
}
BUFFERS = {n: resolve_label_buffer(CASE_STUDY_ID, n, setup) for n in LABEL_NAMES}
DIRECTION_SOURCE = setup["labels"]["classification_eval_label"]
SCALED_LABEL = next(n for n in LABEL_NAMES if "risk_adj" in n)
PLAIN_RETURNS = [n for n in LABEL_NAMES if n.startswith("fwd_ret_") and n != SCALED_LABEL]
PRIMARY_HORIZON = HORIZONS[PRIMARY_LABEL]
HOLDOUT_START = date.fromisoformat(setup["evaluation"]["holdout_start"])
PERIODS_PER_YEAR = setup["evaluation"]["periods_per_year"]
ENTITY = "sec_id"
IV_COL = "iv_30_atm"
IV_LAG = int(setup["decision"]["iv_feature_lag"].split("_")[0])
RV_WINDOW = 20
print(f"{PRIMARY_LABEL} is the primary label, and every later stage defaults to it. Each label:")
for name in LABEL_NAMES:
print(
f" {name} is held {HORIZONS[name]} trading sessions, and a training fold ends "
f"{int(BUFFERS[name].rstrip('Dd'))} sessions before its validation fold begins"
)
print(
f"The holdout opens {HOLDOUT_START}; each development window ends a label's own horizon of "
f"sessions earlier, so nothing measured below resolves inside the holdout"
)
```
## A. The learning task
The hypothesis is cross-sectional, and it is about a disagreement between two views of the
same company. The option market prices a distribution of outcomes for a name over the coming
month; the share price says what the equity market pays for it today. The claim is that names
whose options are priced richly against what their shares go on to do can be ranked against
each other, and that the ranking pays over the following week. The label is therefore a
forward return on the **share** and not on any option: the strategy reads options and holds
equity.
The decision cadence comes from `setup.yaml`. A Friday close is observed, the option
quantities are lagged one session behind it - the surface summary for a session is stamped at
that session's close and cannot be acted on until the next one - and the resulting portfolio is
entered at Monday's open. That fixes the primary horizon at one trading week. The two-week
variant asks whether the same signal still pays when the portfolio turns over half as often, which
is a question about cost and turnover rather than a second hypothesis. The volatility-scaled
variant asks the same question of a target made comparable across names of very different
volatility. The two direction variants reframe the weekly and the two-weekly return as
classification, which is what the logistic configurations in the model stages train on.
Labels are sampled every session rather than only on Fridays: that buys five times the rows at
the price of overlap between consecutive windows, and Section F measures what those rows are
worth.
## B. Preparation before the label
**These are raw traded prices, and a forward return taken on them is not a return.** Each row
carries `open` and `close` as they printed on the day, and `adj_factor`, the cumulative factor
that puts a price on a comparable footing with the rest of that security's history. Splits,
reverse splits, demergers and cash dividends all move it. Multiplying gives a series in which a
four-for-one split is not a three-quarter loss and a dividend is money the holder received
rather than a fall in the price. Because the factor carries dividends as well as splits, the
label below is a **total return** to a holder of the share.
The entity a label may not cross is the **security**, and the column that identifies it is
`sec_id` rather than the ticker. A ticker gets reassigned after a merger or a spin-off and
`adj_factor` restarts with the new security, so a window that steps across the change reads a
price from one company against a price from another. How much of that this extract carries is
counted below, and most of it arrives with no gap in the series to notice it by.
`session` numbers the market's own trading sessions, and Sections C, D and F count in it rather
than in rows: a name that stops trading for a fortnight and returns has consecutive rows
spanning a hole, and only a market-wide session index shows that. No eligibility filter runs
before the shift, because once rows are dropped from inside a series a shift counts survivors,
the horizon stops being measured in trading sessions, and the window silently spans whatever
was removed. Point-in-time eligibility is applied downstream in `03_financial_features`.
**The universe is bounded before any label is built, and the bound is the one
`setup.yaml::universe` declares.** `eligibility_rule` is `sp500_with_options`, and what makes a
name satisfy it is having a row in the option-surface summary: this strategy reads an option to
rank a share, so a share with no listed options is not a candidate however liquid it is. The
share-bar extract is wider than that, and the loader takes a `symbols=` argument that nothing
was passing, so every count below and every label file would otherwise have covered names the
strategy can never hold.
The roster is derived from the surface rather than typed out, and every roster name is then
checked against the share bars: a name that can be ranked must have a price to trade at.
**The roster is read off the whole extract, not off the requested window**, because which names
have listed options is a property of the dataset. Deriving it after `START_DATE` would give a
roster missing every name not yet listed, and the bound would then vary with a parameter the
notebook documents as free to change.
**How many names that comes to is a fact about the extract, so it is reported rather than
asserted.** `universe.n_assets` describes the production dataset, and the reduced extract CI
runs on carries a handful of names by design; an assertion on the count fails there for being
small rather than for being wrong. `tests/test_eoa_universe_roster.py` holds the declaration to
the production data, which is the only place the comparison means anything.
The session grid is taken from the **unbounded** extract, because a session is a property of the
market rather than of the universe, and the horizon is counted on it.
```python
full_surface = load_sp500_options_surface()
full_bars = load_sp500_daily_bars()
ROSTER = sorted(full_surface["symbol"].unique().to_list())
assert ROSTER, "the option-surface extract carries no names to rank"
priced = set(full_bars["symbol"].unique().to_list())
assert not set(ROSTER) - priced, f"no share bars for {sorted(set(ROSTER) - priced)}"
outside = sorted(priced - set(ROSTER))
window = pl.col("timestamp").is_between(
pl.lit(START_DATE).str.to_date(), pl.lit(END_DATE).str.to_date()
)
surface = full_surface.filter(window)
extract = full_bars.filter(window)
sessions = extract.select("timestamp").unique().sort("timestamp").with_row_index("session")
sessions = sessions.with_columns(pl.col("session").cast(pl.Int64))
bars = extract.filter(pl.col("symbol").is_in(ROSTER))
prices = (
bars.join(sessions, on="timestamp")
.with_columns(
(pl.col("open") * pl.col("adj_factor")).alias("adj_open"),
(pl.col("close") * pl.col("adj_factor")).alias("adj_close"),
)
.sort([ENTITY, "session"])
)
# Every label's `inputs`: without it a re-run against a refreshed download is invisible. It is
# taken over the bounded panel, which is what the labels are actually built from.
PRICE_COLS = ["symbol", "sec_id", "timestamp", "open", "close", "adj_factor"]
MARKET_DATA_DIGEST = value_digest(bars, PRICE_COLS)
print(f"market_data digest: {MARKET_DATA_DIGEST}")
DECLARED = setup["universe"]["n_assets"]
against = (
"as universe.n_assets declares"
if len(ROSTER) == DECLARED
else f"a reduced extract; universe.n_assets declares {DECLARED}"
)
print(
f"Universe {len(ROSTER)} names with an option surface ({against}); "
f"{len(outside)} priced names carry no surface and are excluded ({', '.join(outside)})"
)
pairs = prices.select("symbol", ENTITY).unique()
shared = {
key: pairs.group_by(key).len().filter(pl.col("len") > 1).height for key in ("symbol", ENTITY)
}
moves = (
prices.sort(["symbol", "session"])
.with_columns(
(pl.col(ENTITY) != pl.col(ENTITY).shift(1).over("symbol")).alias("_moved"),
pl.col("session").diff().over("symbol").alias("_gap"),
)
.filter(pl.col("_moved") & pl.col("_gap").is_not_null())
)
print(f"{prices['symbol'].n_unique()} tickers over {sessions.height:,} sessions")
print(f"{prices.height:,} name-sessions in {prices[ENTITY].n_unique()} securities")
print(
f"{shared['symbol']} tickers cover more than one security and {shared[ENTITY]} securities "
f"trade under more than one ticker; {moves.filter(pl.col('_gap') == 1).height} of "
f"{moves.height} changeovers land on the very next session"
)
```
## C. Label construction
One execution convention, written once and applied at both horizons:
$$r^{(h)}_{i,t} = \frac{C_{i,t+h}}{O_{i,t+1}} - 1$$
where $O$ and $C$ are the adjusted open and close of security $i$, and $t+h$ counts $h$
**trading sessions**. The signal is read at the close of $t$, the position is entered at the
next open, and it is exited $h$ sessions later at the close. That is the convention
`setup.yaml` declares under `decision`, and it is deliberately not Chapter 7.2's close-to-close
form: a label anchored on the close of $t$ credits the strategy with the overnight move between
the signal and the first price it can trade at.
The library's `fixed_time_horizon_labels` cannot express this convention: it divides by the
price at $t$ in the same column it takes the numerator from, and here the numerator is a close
and the denominator an open. The construction is written out below instead, under the rule the
library would have applied - a window counts only if every price in it exists and its two ends
are exactly $h$ sessions apart.
Three working columns are computed on the complete series, because each is only meaningful
before a row is dropped: `from_end`, how far back a row sits from the last session its security
trades; `session`, the market's own session index; and the one-session return the scaled
label's trailing volatility is built from.
```python
def forward_return(df: pl.DataFrame, horizon: int, name: str) -> pl.DataFrame:
"""Next-open-to-close return over `horizon` sessions, null unless the window is complete."""
entry = pl.col("adj_open").shift(-1).over(ENTITY)
exit_price = pl.col("adj_close").shift(-horizon).over(ENTITY)
spans_horizon = pl.col("session").shift(-horizon).over(ENTITY) - pl.col("session") == horizon
return df.with_columns(
pl.when(spans_horizon & (entry > 0) & (exit_price > 0))
.then(exit_price / entry - 1)
.alias(name)
)
ONE_SESSION = pl.col("session").diff().over(ENTITY) == 1
labels_df = prices.with_columns(
(pl.len().over(ENTITY) - 1 - pl.int_range(pl.len()).over(ENTITY)).alias("from_end"),
pl.when(ONE_SESSION).then(pl.col("adj_close").pct_change().over(ENTITY)).alias("_daily_ret"),
)
for label_name in PLAIN_RETURNS:
labels_df = forward_return(labels_df, HORIZONS[label_name], label_name)
```
The scaled variant divides the weekly return by the volatility realized over the previous
twenty sessions, about a month of trading, annualized on the sessions per year `setup.yaml`
declares, so that a five percent week
in a quiet utility and a five percent week in a semiconductor are not the same target value.
The denominator is guarded twice. A daily return is only a daily return between two adjacent
sessions, so it is null wherever the security missed the session before, and the rolling window
stays null until it holds a full month of consecutive returns rather than closing over the gap
as though nothing had been missed. And a security that did not move at all
over the lookback has an undefined ratio, which is left undefined rather than floored at a
constant that would answer with a very large number instead.
The direction variants are the sign of the return they come from, and
`setup.yaml::labels.classification_eval_label` says which return that is. **The guard on the
outer `when` is the whole point of the cell.** In Polars a null predicate is not true, so
`when(ret > 0).then(1).otherwise(0)` writes a confident zero wherever the return is null - in
every row of the last week of every security, which have no forward window at all. Those rows
cannot be removed afterwards, because the value is not null, it is wrong.
```python
trailing_rv = pl.col("_daily_ret").rolling_std(RV_WINDOW).over(ENTITY) * PERIODS_PER_YEAR**0.5
labels_df = labels_df.with_columns(trailing_rv.alias("_trailing_rv"))
labels_df = labels_df.with_columns(
pl.when(pl.col("_trailing_rv") > 0)
.then(pl.col(PRIMARY_LABEL) / pl.col("_trailing_rv"))
.alias(SCALED_LABEL)
)
for label_name, source in DIRECTION_SOURCE.items():
sign = pl.when(pl.col(source).is_not_null()).then((pl.col(source) > 0).cast(pl.Int32))
labels_df = labels_df.with_columns(sign.alias(label_name))
```
## D. Window validity
A shift always returns something; the question is whether what it returns is the quantity the
label claims. Each property below fails silently and leaves plausible numbers behind, so each
is asserted rather than described.
The second assertion is a full reconciliation rather than a bound. Every row carrying no label
is attributed to exactly one cause - the tail of its security's series, a hole between the two
ends of the window, or a missing price at one end - and the counts have to sum to the height of
the frame. A label crossing a security boundary, or a short label masked by a longer one's null
set, breaks that identity.
The third assertion bounds the calendar span a labelled window may cover, and it is what tests
the session count against the calendar instead of against itself: $h$ trading sessions span
about $7h/5$ calendar days on a five-session week, plus a week for exchange holidays. A
security that stops trading for a year and returns would produce windows far wider than any
holiday pattern accounts for.
```python
for label_name in [*PLAIN_RETURNS, SCALED_LABEL]:
horizon = HORIZONS[label_name]
checked = labels_df.with_columns(
(pl.col("timestamp").shift(-horizon).over(ENTITY) - pl.col("timestamp"))
.dt.total_days()
.alias("_span")
)
tail = pl.col("from_end") < horizon
holed = pl.col("session").shift(-horizon).over(ENTITY) - pl.col("session") != horizon
unfilled = ~tail & ~holed & pl.col(label_name).is_null()
warm_up = unfilled & pl.col(PRIMARY_LABEL).is_not_null() & (label_name == SCALED_LABEL)
causes = {"no forward window": tail, "a hole in the window": ~tail & holed}
causes |= {"no trailing volatility yet": warm_up, "a missing price": unfilled & ~warm_up}
labelled = checked.drop_nulls(label_name)
# 1. An incomplete forward window is null, never a value.
assert checked.filter(tail)[label_name].null_count() == checked.filter(tail).height
# 2. Labelled rows plus the causes account for every row, each cause once.
counts = {cause: checked.filter(cond).height for cause, cond in causes.items()}
assert labelled.height + sum(counts.values()) == checked.height, (label_name, counts)
# 3. No labelled window spans more calendar days than holidays alone can explain.
tolerance = -(-horizon * 7 // 5) + 7
assert labelled.filter(pl.col("_span") > tolerance).height == 0, label_name
unlabelled = ", ".join(f"{n:,} with {cause}" for cause, n in counts.items() if n)
print(
f"{label_name}: {labelled.height:,} labelled, spans up to {labelled['_span'].max()}d "
f"against a {tolerance}d tolerance; unlabelled {unlabelled}"
)
# 4. No direction label is derived from a null return.
for label_name, source in DIRECTION_SOURCE.items():
fabricated = labels_df.filter(pl.col(source).is_null() & pl.col(label_name).is_not_null())
assert fabricated.height == 0, (label_name, fabricated.height)
print(f"{label_name}: the sign of {source}, and null wherever {source} is")
```
Position zero below is each security's last session. The non-null rate has to fall to zero over
exactly the last `horizon` positions and sit flat beyond them, and each direction label has to
fall exactly where the return it comes from does. A scalar count of valid rows shows neither
failure this catches: a tail written as a confident zero reads as fully valid, and a short
label masked by a longer one's null set reads as the longer one's count. The figure reads only
which rows are null and never a label value, so it covers the whole extract rather than the
development window that Sections E to G are restricted to.
```python
profile = (
labels_df.filter(pl.col("from_end") <= max(HORIZONS.values()) + 3)
.group_by("from_end")
.agg([pl.col(n).is_not_null().mean().alias(n) for n in LABEL_NAMES])
.sort("from_end")
)
family = {5: COLORS["blue"], 10: COLORS["amber"]}
styles = {n: dict(lw=1.8, color=family[HORIZONS[n]]) for n in LABEL_NAMES}
styles |= {n: dict(marker="o", ms=7, mfc="none", ls="none", **styles[n]) for n in DIRECTION_SOURCE}
styles[SCALED_LABEL] = dict(lw=1.4, ls="--", color=COLORS["copper"])
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
for label_name in LABEL_NAMES:
ax.plot(profile["from_end"], profile[label_name], label=label_name, **styles[label_name])
for horizon, colour in family.items():
ax.axvline(horizon - 0.5, color=colour, linestyle=":", lw=1)
ax.set_xlabel("Sessions from the end of each security's own series")
ax.set_ylabel("Share of securities with a non-null label")
ax.set_ylim(-0.05, 1.15)
add_message_title(
ax,
"Every label nulls exactly its own horizon of trailing sessions",
subtitle="Dotted lines mark each horizon; a fabricated tail would sit flat across it",
)
ax.legend(loc="center left", frameon=False, fontsize=8)
show_with_alt(fig, "Non-null rate of each label by position from the end of a security's series.")
```
## E. Distribution and base rate
What scale is each label, and does it mean the same thing across names and across regimes?
Everything from here through Section G is computed on the development window only, and that
window is cut at the date each label's outcome is **known** rather than at the date its signal
is read. A row observed shortly before the holdout still resolves inside it, so a filter on the
observation date leaves holdout outcomes in the diagnostic. Each label is cut at its own
endpoint, because the horizons differ. The label files still carry every row: the cut governs
what this notebook looks at, not what it writes.
```python
dev = {
name: labels_df.with_columns(
pl.col("timestamp").shift(-horizon).over(ENTITY).alias("_label_end")
)
.filter(pl.col("_label_end") < HOLDOUT_START)
.drop_nulls(name)
for name, horizon in HORIZONS.items()
}
for label_name, frame in dev.items():
print(f"{label_name}: {frame.height:,} development rows through {frame['timestamp'].max()}")
```
Both return labels go on one axis with identical bins and a logarithmic count axis. The claim
the figure has to support is about shape rather than width - two dispersion scalars carry the
width - and the shape is that the longer horizon adds mass across the whole body rather than
only in the tails. The axis is symmetric and narrower than either label's range, so rows
outside it are counted below rather than drawn.
```python
bins = np.linspace(-0.15, 0.15, 81)
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
hist_styles = (dict(color=COLORS["amber"]), dict(color=COLORS["blue"], histtype="step", lw=2))
for label_name, style in zip(PLAIN_RETURNS[::-1], hist_styles, strict=True):
series = dev[label_name][label_name]
ax.hist(series.to_numpy(), bins=bins, label=f"{label_name}, std {series.std():.3f}", **style)
ax.axvline(0, color=COLORS["neutral"], linestyle="--", lw=0.8)
ax.set_yscale("log")
ax.set_xlabel("Forward total return on the adjusted share price")
ax.set_ylabel("Rows per bin, log scale")
add_message_title(
ax,
"The longer horizon adds mass across the whole body, not only the tails",
subtitle="Identical bins, development window; rows beyond the axis are counted below",
)
ax.legend(loc="upper left", frameon=False)
show_with_alt(fig, "Histograms of the two forward-return labels on identical bins, log counts.")
for label_name in [*PLAIN_RETURNS, SCALED_LABEL]:
series, on_axis = dev[label_name][label_name], label_name != SCALED_LABEL
beyond = f", {int((series.abs() > bins[-1]).sum()):,} rows beyond the axis" if on_axis else ""
print(f"{label_name}: std {series.std():.5f}, kurtosis {series.kurtosis():.2f}{beyond}")
short, long_ = (dev[n][n].std() for n in PLAIN_RETURNS)
root_h = (HORIZONS[PLAIN_RETURNS[1]] / PRIMARY_HORIZON) ** 0.5
print(f"width ratio {long_ / short:.2f} against {root_h:.2f} under square-root-of-horizon scaling")
```
The scaled variant lives on a different axis, a multiple of the name's own volatility rather
than a return, so it is not drawn on the return axis above and its width is printed instead.
What the division does is visible in the kurtosis: a five percent week is an ordinary week in
one name and an extreme one in another, and pooling them unscaled puts the ordinary weeks of
the most volatile names where a model reads them as the extremes of the target.
Chapter 7.2 asks for the base rate to be tracked through time. For a continuous label ranked
across a cross-section, the quantity that has to be stable is the spread the model ranks
within: where it is not, the same rank correlation buys a different amount of return. The
spread is taken across names on each session first and only then averaged over the year.
Pooling every name-session in a year into one standard deviation instead adds the movement of
the market's own mean from session to session to the spread across names on a session, and a
ranking model is scored on the second alone. For a direction label the same question is the
share of rows in the positive class, which is what a classifier's threshold sits against.
```python
def annual_profile(name: str) -> pl.DataFrame:
"""Per year: the spread across names on a session for a return, the positive share for a sign."""
frame = dev[name].with_columns(pl.col("timestamp").dt.year().alias("year"))
if name in DIRECTION_SOURCE:
return frame.group_by("year").agg(pl.col(name).mean().alias("value")).sort("year")
return (
frame.group_by("timestamp", "year")
.agg(pl.col(name).std().alias("value"))
.group_by("year")
.agg(pl.col("value").mean())
.sort("year")
)
annual = {n: annual_profile(n) for n in [*PLAIN_RETURNS, *DIRECTION_SOURCE]}
```
The two panels share a year axis so the same regime can be found in both. The dashed line on the
right marks an even split, which is where a direction label carries no information at all before
a model sees it. A share that sits that close to its reference needs an axis that does not start
at zero, so it is drawn as a series against the reference rather than as a bar whose baseline
would have to carry the comparison.
```python
fig, axes = plt.subplots(1, 2, figsize=FIGSIZE["dual_h_tall"], layout="tight")
for offset, name in zip((-0.2, 0.2), PLAIN_RETURNS, strict=True):
table = annual[name]
axes[0].bar(
table["year"] + offset, table["value"], 0.4, color=family[HORIZONS[name]], label=name
)
for name in DIRECTION_SOURCE:
table = annual[name]
axes[1].plot(
table["year"], table["value"], "o-", ms=4, color=family[HORIZONS[name]], label=name
)
axes[1].axhline(0.5, color=COLORS["neutral"], linestyle="--", lw=0.8)
axes[1].set_ylim(0.45, 0.65)
axes[0].set_ylabel("Cross-name spread, mean over sessions")
axes[1].set_ylabel("Share of rows in the positive class")
for ax, loc in zip(axes, ("upper left", "lower left"), strict=True):
ax.set_xticks(annual[PRIMARY_LABEL]["year"].to_list())
ax.tick_params(axis="x", labelsize=8)
ax.legend(loc=loc, frameon=False, fontsize=8)
add_message_title(
axes[0],
"Neither the spread nor the positive-class share holds still",
subtitle="Development window; the dashed line marks an even split",
)
fig.tight_layout()
show_with_alt(fig, "Annual cross-name dispersion and annual positive-class share, by label.")
for label_name, table in annual.items():
peak, low = (table.sort("value", descending=d).row(0, named=True) for d in (True, False))
print(
f"{label_name}: {low['value']:.4f} in {low['year']} rising to {peak['value']:.4f} in {peak['year']}"
)
```
On the development window the weekly label has a standard deviation of 0.04697 and the
two-weekly one 0.06540, a ratio of 1.39 against the 1.41 that square-root-of-horizon scaling
implies; dividing the weekly label by each name's trailing volatility cuts its kurtosis from
16.48 to 7.16. Neither the spread nor the base rate is stable: the daily cross-name spread of
the weekly label runs from 0.0285 in 2017 to 0.0476 in 2020, and the share of weeks that end
higher from 0.5070 in 2018 to 0.5783 in 2019.
## F. Overlap and effective sample size
Sampling a multi-session label at every session makes consecutive rows share most of their
forward window, so the row count overstates the evidence. Two measurements answer that in
different units: how fast the overlap decays, and what the rows are worth once it is priced in.
`effective_sample_size` applies Chapter 7.2's average-uniqueness weighting per security,
because concurrency is a property of one security's own overlapping windows.
Both are counted on `session`, the grid the label was built on, rather than on position among
the rows that survive into the development frame. Closing over a security's own trading gap
would pair windows that share nothing and report the overlap as larger than it is.
```python
max_lag = max(HORIZONS.values()) + 4
acf = {
name: panel_autocorrelation(
dev[name], name, max_lag=max_lag, bar_col="session", entity_col=ENTITY
)
for name in PLAIN_RETURNS
}
fig, ax = plt.subplots(figsize=FIGSIZE["single"])
lags = np.arange(1, max_lag + 1)
for label_name in PLAIN_RETURNS:
colour = family[HORIZONS[label_name]]
ax.plot(lags, acf[label_name], "o-", ms=3, c=colour, lw=1.8, label=label_name)
ax.axvline(HORIZONS[label_name], color=colour, linestyle=":", lw=1.5)
ax.axhline(0, color=COLORS["neutral"], lw=0.8)
ax.set_xlabel("Lag in trading sessions")
ax.set_ylabel("Panel autocorrelation")
add_message_title(
ax,
"The overlap in each label decays to zero at its own horizon",
subtitle="Dotted lines mark each horizon; what remains past one is not overlap",
)
ax.legend(loc="upper right", frameon=False)
show_with_alt(fig, "Panel autocorrelation of both return labels against lag in trading sessions.")
```
The second measurement turns the same overlap into a row count. Average uniqueness weights each
row by the share of its forward window no concurrent label also spans, and summing those weights
gives the number of independent observations the frame is worth. A horizon-h label consumes the
h returns realized over its window and its neighbour one session later shares h-1 of them, so
average uniqueness converges to 1/h and the ratio below is what it converges to.
```python
for label_name, horizon in HORIZONS.items():
n_rows, n_eff = effective_sample_size(
dev[label_name], horizon=horizon, bar_col="session", entity_col=ENTITY
)
print(
f"{label_name}: N={n_rows:,}, N_eff={n_eff:,.0f}, ratio {n_eff / n_rows:.4f} against "
f"{1 / horizon:.4f} for windows overlapping this fully"
)
for label_name in PLAIN_RETURNS:
decay, horizon = acf[label_name], HORIZONS[label_name]
print(
f"{label_name}: autocorrelation {decay[0]:.3f} at lag one, {decay[horizon - 1]:.3f} at lag {horizon}"
)
```
The weekly label's 503,079 development rows carry 101,099 effective observations, a ratio of
0.2010 against the 0.2000 a fully overlapped five-session window implies; the two-weekly
label's 500,062 rows carry 50,548, a ratio of 0.1011 against 0.1000. Both sit just above the
reference value, because a security that stops trading ends an overlap early. Panel
autocorrelation falls from 0.772 at lag one to -0.051 at lag five for the weekly label, and
from 0.891 to 0.005 at lag ten for the two-weekly one. The purge gap a fold needs is set by the
forward window itself, not by these counts.
## G. Baseline floor
One signal against the primary label on the development window, with no feature
engineering: the thirty-day at-the-money implied volatility whose persistence
[`01_feasibility_analysis`](01_feasibility_analysis.ipynb) measured, lagged by the session
`setup.yaml::decision.iv_feature_lag` declares, and read raw. Every option-derived feature in
`03_financial_features` is built out of that quantity, so it is the floor those features have
to clear, and measuring it before building them is what makes a later improvement mean
something.
The lag is a market session inside the security, not one row of the surface. Coverage of the
surface swings on a monthly cycle, so on the narrow weeks a name's previous quoted row can be
a fortnight back, and a shift taken over the surface's own rows would hand the model a
fortnight-old volatility while calling it yesterday's. Joining the surface onto the dense
price panel first makes the missing sessions visible as nulls, and the guard drops them.
The information coefficient is the cross-sectional rank correlation on each session, averaged
over sessions, which is the quantity a ranking model is scored on; pooling every name-session
instead mixes a cross-sectional claim with a time-series one. The series is put back in time
order before the standard error reads it, because that standard error is built out of the
correlations between one session and the next.
A correlation over a handful of names is noise, so a session is scored only where at least a
minimum number of names are quoted. That minimum is half the median cross-section rather than
a fixed count, so it means the same thing on a universe of another size, and the sessions that
fall below it are dropped from the statistic - the count printed below is how many of the
window's sessions were scored. The standard error is HAC-adjusted, because the series of
per-session correlations inherits the label's overlap and a naive statistic would treat five
consecutive sessions of one week's return as five separate pieces of evidence.
```python
signal_panel = (
prices.select("timestamp", "symbol", ENTITY, "session")
.join(surface.select("timestamp", "symbol", IV_COL), on=["timestamp", "symbol"], how="left")
.sort([ENTITY, "session"])
.with_columns(
pl.when(ONE_SESSION).then(pl.col(IV_COL).shift(IV_LAG).over(ENTITY)).alias("signal")
)
.drop_nulls("signal")
)
baseline = dev[PRIMARY_LABEL].join(
signal_panel.select("timestamp", "symbol", "signal"), on=["timestamp", "symbol"], how="inner"
)
min_obs = int(baseline.group_by("timestamp").len()["len"].median() // 2)
ic = cross_sectional_ic_series(
baseline,
baseline,
pred_col="signal",
ret_col=PRIMARY_LABEL,
date_col="timestamp",
entity_col="symbol",
min_obs=min_obs,
).sort("timestamp") # a HAC autocovariance over a permutation of the time axis means nothing
stats = compute_ic_hac_stats(ic, ic_col="ic", label_horizon=PRIMARY_HORIZON)
print(
f"Baseline: {IV_COL} lagged {IV_LAG} session against {PRIMARY_LABEL}, "
f"{baseline.height:,} rows, minimum cross-section {min_obs} names\n"
f" sessions scored {stats['n_periods']:,} of the {ic.height:,} in the window, "
f"mean IC {stats['mean_ic']:.4f}\n"
f" HAC t {stats['t_stat']:.2f} on {stats['effective_lags']} Bartlett lags, "
f"naive t {stats['naive_t_stat']:.2f}, p {stats['p_value']:.3g}"
)
```
The lagged at-the-money volatility earns a mean information coefficient of -0.0072 against the
weekly label. Every one of the 1,001 sessions in the development window quotes at least the 130
names the minimum cross-section asks for, so all 1,001 are scored. Under the naive standard
error that is a t-statistic of -0.93; the Newey-West rule picks 6 lags here, above the four the
horizon alone requires, and the HAC statistic is -0.50 with a p-value of 0.614. The floor a
feature has to clear is a mean IC the data cannot separate from zero, so the sign of it carries
nothing either.
## H. Artifacts and the audit record
Beside each label file the notebook writes a second, small JSON file, `<label>.digest.json`.
It is there so that a later stage can tell whether the label it is reading still holds the
values it was built against. It does that with a **digest**: a short hash of the values in the
file, which changes if any value changes and stays the same if the same values are written
again. The sidecar records that digest, the row count, the columns, the key columns, the
notebook that wrote the file, and the digest of the price data the label was built from. The
last of those is what ties a label to a data vintage - without it, a re-run against a refreshed
download looks exactly like a re-run against the same one.
The key is `timestamp` and `symbol`, the pair every file downstream joins on. `sec_id` is the
entity the label was built inside and is not carried into the file.
The folds that train models are derived per label by `case_studies/utils/cv_window.py` from
`config/setup.yaml` and the timeline of the label parquet written here, so which rows land in
these files is what sets where the fold boundaries fall.
```python
KEYS = ["timestamp", "symbol"]
for label_name in LABEL_NAMES:
record = write_artifact(
labels_df.select([*KEYS, label_name]).drop_nulls(),
LABELS_DIR / f"{label_name}.parquet",
keys=KEYS,
written_by="02_labels",
inputs={"market_data": MARKET_DATA_DIGEST},
)
print(f"{label_name}.parquet: {record['n_rows']:,} rows, digest {record['digest']}")
```
The record Chapter 7.2 requires to close a label definition, one row per label, built from the
values computed above rather than written by hand. The horizon and the buffer are printed
separately, because this is the case study that separates them.
```python
readers = {
PRIMARY_LABEL: "05_evaluation.py, which scores the features against it, and the model stages"
}
print("\nLabel audit record")
for label_name, horizon in HORIZONS.items():
frame, source = dev[label_name], DIRECTION_SOURCE.get(label_name)
anchor = f"the sign of {source}" if source else "the adjusted open of the session after t"
buffer_sessions = int(BUFFERS[label_name].rstrip("Dd"))
print(
f"\n{label_name}\n anchor {anchor}"
f"\n horizon {horizon} trading sessions, against a {buffer_sessions}-session fold buffer"
f"\n resolution fixed at the close of t+{horizon}; daily bars need no tie-break"
f"\n overlap {horizon - 1} sessions shared by consecutive rows"
f"\n base rate mean {frame[label_name].mean():+.5f}, std {frame[label_name].std():.5f}"
f"\n consumed by {readers.get(label_name, 'the model stages, as a variant')}"
)
```
## What the labels hold, and what they owe
The sections above check each label against its own definition - that the sign matches the
return, that no value is fabricated where the source is null, that the last sessions are
withheld. This one asks the two questions that are about the file rather than the rule.
The first is what each label's distribution looks like: how much is null, how much is exactly
zero, how far the extreme values sit from the body. A concentration at zero is a defect in a
return and is what a balanced direction label is supposed to look like, so the check reports the
number and this notebook says which it is.
The second is coverage, and it needs a denominator that is not the labels themselves. **The
reference is `bars`** - every session a security in the roster actually traded, which is the set
a forward return could in principle have been computed on. Comparing a label to the union of the
other labels would hide any session where all five are absent together; comparing it to the
price panel cannot.
A percentage on its own decides nothing, so each label also declares *where* it is entitled to
be short and by how much, and the report is read against that declaration rather than against
the raw number. A forward return owes no value in the last `horizon` sessions of a security's
history, because the price that would resolve it is past the end of the sample; the
risk-adjusted variant additionally owes none until its trailing volatility window has filled.
Both are counted per security, so a name that enters late pays its own burn-in on its own dates.
What the sign-off then has to speak to is not the shortfall but the residual: keys missing
inside a security's own span, where no window length explains them, and securities that spend
more than the mechanism allows.
```python
expected_keys = bars.select(KEYS).unique()
print(
f"price panel: {expected_keys.height:,} traded (symbol, session) keys across "
f"{expected_keys['symbol'].n_unique()} securities and "
f"{expected_keys['timestamp'].n_unique()} sessions\n"
)
for label_name in LABEL_NAMES:
written = labels_df.select([*KEYS, label_name]).drop_nulls()
budget: dict[str, tuple[int | None, str]] = {
"trailing": (
HORIZONS[label_name],
f"{HORIZONS[label_name]}-session forward window past the end of the sample",
)
}
if label_name == SCALED_LABEL:
budget["leading"] = (
RV_WINDOW,
f"{RV_WINDOW}-session trailing realized volatility in the denominator",
)
render_quality_report(
quality_report(
written,
name=label_name,
key_columns=KEYS,
expected=expected_keys,
keys=KEYS,
entity="symbol",
session="timestamp",
expected_missing=budget,
)
)
print()
```
### Sign-off
**The price panel offers 632,602 traded keys over 633 securities and 1,259 sessions, every label
emits keys the panel has and nothing else, and no column crossed a distribution threshold.** The
three coverage figures differ from each other by amounts each label's own definition predicts,
which is the point of declaring the budget before printing the number: 97.53% and 99.50% are
both complete, and a reader who only saw the percentages could not know that.
- **`fwd_ret_5d` and `fwd_dir_5d` reach 99.50%.** Of the 3,159 keys they lack, 3,071 sit after a
security's last session, a median of 5 apiece against a declared horizon of 5. The forward
window runs past the end of the sample and there is no future price to return; withholding
them is the seal working, and a value there would have to be invented.
- **`fwd_ret_10d` and `fwd_dir_10d` reach 99.01%**, and lose 6,087 keys the same way at a median
of 10 apiece against a horizon of 10. Twice the horizon costs twice the history, which is the
trade the horizon choice makes and is worth seeing rather than assuming.
- **`fwd_ret_risk_adj_5d` reaches 97.53%**, and it is the only label that pays at *both* ends.
It loses 3,070 keys to its five-session horizon like the plain return, and a further 12,187 to
the front of each security's history at a median of exactly 20 sessions - the trailing
realized volatility in its denominator, undefined until its 20-session window fills. Only two
of 608 securities spend more than the 20, and 27 keys between them. **This is the label the
model stages select on**, so a reader should know it is measured on 2% fewer decisions than the
plain five-session return beside it, and that the decisions it lacks are the earliest ones.
**What no mechanism accounts for is 350 keys at the worst label, and 88 at the primary.** Two
shapes make it up, and neither reaches a scale that moves anything. A handful of securities -
seven at the ten-session horizon, three at the five - carry no label at all, because their entire
life in this panel is shorter than the horizon: a name with four traded sessions has no fifth to
return to. The rest sit inside a security's own span, a median of one horizon's worth each, which
is what a break in a security's traded sessions costs when the forward window steps across it.
At 0.06% of the panel at the extreme, neither changes a fold, a fit or a ranking, and the reason
to print them is that a residual nobody states is a residual nobody would notice growing.
**The direction labels sit near half at each value and nothing flagged them.** A binary down/up
indicator is supposed to; a share far from half would say the sign convention or the null
handling had gone wrong, which is why the number is printed rather than assumed. It is a little
below half in both, which is the equity drift showing through - more sessions up than down over
this sample. The zero-share ceiling exists to catch a *return* that is mostly zero, and it is
deliberately loose enough not to fire on a label that is meant to be half zeros.
**The return labels carry heavy tails and are almost never exactly zero.** That is what a forward
equity return looks like at daily frequency, and it is the reason the risk-adjusted variant
exists at all. Nothing is winsorized here; how a model handles the tail is a modelling choice
made downstream.
**No label column is constant and none carries a non-finite value.** Both are flagged
unconditionally above, and neither appears in any of the five.
## Key takeaways
1. **On a raw equity panel, rebuild the tradable price series before writing any label.** A
split, a spin-off, or a ticker reassigned to another company all leave the raw prices intact
and the return between them meaningless, and none of them raises anything.
2. **Check the window in code, and account for every unlabelled row.** An incomplete window,
a hole in a security's series and a window crossing a change of security all fail without
raising, and a reconciliation that has to balance catches what a row count passes over.
3. **Derive a discrete label under an explicit null guard.** A comparison against a null return
is false rather than null, so the naive form writes a confident class into every row the
return could not fill, and no later `drop_nulls` can find it.
4. **Cut a diagnostic off at the label's endpoint, not at its observation date.** A row
observed before the holdout whose outcome resolves inside it is a holdout row, so each
label's usable boundary is the holdout date minus its own horizon, counted in trading
sessions.
5. **A row count overstates the evidence when forward windows overlap.** The effective count
says by how much, and it does not set the purge gap - the forward window does.
**Known limitations.** The label is a total return to a holder of the share, because the
adjustment factor carries dividends as well as splits; a strategy credited only with price
returns would be measured against a slightly different target. The universe is every name
carrying an option surface, which is what `universe.eligibility_rule` declares, and not a
liquidity screen applied point in time: a name enters on having listed options, not on trading
enough of them. Point-in-time eligibility is `03_financial_features`'. The hole rule tests each
security's own position on the market calendar, so it finds an absent session and not a session
on which the name barely traded. The baseline is one signal, read raw off the surface with no
adjustment for the level of volatility a name usually carries.
**Next**: `03_financial_features.py` builds the volatility-surface, variance-premium, skew and
equity-momentum features from the same two sources. `05_evaluation.py` is where those features
are first scored against the primary label written here.



स्रोत के लाइसेंस के तहत श्रेय सहित पूरा पाठ दिखाया गया है। लाइसेंस: MIT
यह सारांश मूल स्रोत के आधार पर Stratmill के शोध एजेंट ने लिखा है; यह स्रोत की प्रति नहीं है।