检验期权隐含特征能否对未来股票收益进行排序
代码 《交易机器学习》
总结
本笔记评估期权市场指标能否对不同股票的未来收益进行排序。笔记使用期权衍生特征拟合已声明的线性模型,采用滚动验证,并比较岭回归、套索回归和弹性网络的惩罚强度。特征集包括各期限的隐含波动率、波动率偏斜、期限结构斜率和方差风险溢价等指标。正则化用于处理相关特征之间的冗余;套索回归和弹性网络还可以将部分系数缩减为零。
笔记强调,期权价格反映对未来分布的预期,尤其是波动率,而其主要标签关注收益水平或方向。笔记将信息系数作为排序诊断,而非盈利交易的证据或选择规则;策略选择留待之后的回测进行。验证历史较短,只有两个折,其中包括2020年的波动冲击,因此结果可能反映了不寻常的时期。除非明确提供相应特征,线性模型也无法捕捉交互作用。文档未能证明期权特征是否能成功预测波动率,或是否有助于期权交易策略。
核心观点
- 期权衍生特征代表市场预期,与历史收益特征不同。
- 滚动比较应让每种模型配置使用相同的股票、日期和数据折。
- 岭回归会收缩系数,套索回归可将部分系数设为零;弹性网络结合了两种惩罚。
- 信息系数评估排序,而非扣除成本后的交易表现或策略盈利能力。
- 验证折只有两个,且包括一个波动异常的年份,因此结论的说服力有限。
标签
全文
# 06_linear.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Option analytics: can what the options market prices rank the stock
#
# Every other case study in this book builds its features from the history of the thing it is
# trying to predict - past returns, past volatility, past volume. This one does not. Its features
# come from the options written on each stock, and an option price is not a record of what
# happened. It is what someone was willing to pay today for a claim on what happens next. Implied
# volatility, the skew between puts and calls, the slope of the term structure, the gap between
# implied and realized variance: each is a statement the market is making about the future.
#
# That makes the question here sharper than "do these features work". The features are already
# forecasts, made by people with money at stake. If a forecast that good does not rank the
# cross-section, the interesting part is understanding what it is a forecast *of*.
#
# The design matrix has the redundancy every case study in this book has - implied volatility
# appears at three tenors and again as a rank, a percentile and two z-scores; the variance risk
# premium appears as a level, a rank and a z-score - so the penalty sweep has the same job here as
# elsewhere. **Regularization** adds a penalty on coefficient size to the fitting objective, and
# how much to apply is an empirical question this notebook answers by trying ten orders of
# magnitude of it.
#
# Two folds, and one of them validates on 2020. The usable history of this option analytics
# dataset is short, so `05_evaluation` set a walk-forward schedule whose validation windows are
# 2019 and 2020. Whatever the models find or fail to find, they are being asked about a period
# containing the fastest volatility shock in the sample.
#
# **Learning objectives.** By the end of this notebook you will be able to:
#
# - Read the set of models a case study has declared for a label, and say which estimator and
# which hyperparameters each declared name resolves to.
# - Bind those declarations to the data on disk and check, before anything is fitted, that every
# configuration will be measured on the same stocks, the same dates and the same folds.
# - Fit a population of models on walk-forward folds and publish one complete set of validation
# predictions per configuration.
# - Say what a feature set of option-implied quantities is a forecast of, and why that is not the
# same thing as the quantity this notebook asks it to rank.
# - Read a grid that sits entirely on one side of zero on one label and entirely on the other
# side on another, and say what that is and is not evidence for.
# - Run configurations of your own into a private copy of the run log, and have them compared on
# the same footing as the ones shipped here.
#
# **Book reference**: Chapter 11 (The ML Pipeline). Chapter 6, Section 6.7 (Search accounting and
# run logging) introduces the run log this notebook writes to.
#
# **Prerequisites**: [`03_financial_features`](03_financial_features.ipynb) and
# [`04_model_based_features`](04_model_based_features.ipynb) have written the feature matrices,
# and [`05_evaluation`](05_evaluation.ipynb) has established the walk-forward folds and screened
# the individual features.
#
# **What it writes**: one training run and one complete validation prediction set per
# configuration, in `run_log/registry.db` and under `run_log/training/` and
# `run_log/predictions/`, grouped under a named population.
# [`14_backtest`](14_backtest.ipynb) reads that population, runs every member against the
# equal-weight baseline, and selects on validation backtest Sharpe. **Selection happens there,
# not here.** This notebook ranks configurations by information coefficient to show what
# regularization does to this feature set; that ranking decides nothing.
# %%
"""Fit the declared option-analytics linear-model population on the walk-forward folds."""
import re
import numpy as np
import plotly.graph_objects as go
import polars as pl
from plotly.subplots import make_subplots
from case_studies.research import (
declared_labels,
load_model_configs,
model_requests,
narrows_declared_catalog,
open_study,
primary_label,
resolved_model_plan,
run_model_population,
)
from utils.style import COLORS, show_plotly_with_alt
# %% tags=["parameters"]
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
CONFIG_NAMES: list[str] = []
POPULATION_NAME = ""
SUPERSEDES_POPULATION: str = ""
# %%
study = open_study(
"sp500_equity_option_analytics",
execution_tier=EXECUTION_TIER,
workspace=WORKSPACE or None,
entry_point="06_linear",
)
# %% [markdown]
# ## 1. Which label, and which models
#
# A label is the thing being predicted. This case study defines five in `config/setup.yaml`.
# `fwd_ret_5d`, the stock's total return over the five trading days after the decision date, is
# the primary one - the horizon the strategy chapters trade. `fwd_ret_10d` is the same idea over
# a longer horizon, `fwd_ret_risk_adj_5d` divides the move by a volatility estimate, and
# `fwd_dir_5d` and `fwd_dir_10d` are classification variants that ask for the direction rather
# than the size.
#
# **This notebook fits every declared label in one run.** Each label has its own training menu at
# `config/training/{label}.yaml`, listing family by family the named configurations to fit for
# that label. A label with no menu file has nothing declared and nothing to fit; the ones below
# are those that declare linear models. `LABELS` narrows the run to a subset, which is a
# diagnostic rather than the canonical population.
# %%
declared_labels(study, "linear")
# %% [markdown]
# Each name in the menu resolves to a preset file in the shared directory
# `case_studies/config/{model_type}/`, which holds that configuration's hyperparameters. The
# frame below is the menu for every declared label, with each name resolved to the estimator class it names
# and the arguments that class is constructed with. To change what runs, edit the menu or the
# presets rather than this notebook.
#
# The grid covers the two shapes a penalty can take:
#
# - **Ridge** penalizes the sum of squared coefficients. It shrinks correlated coefficients
# towards each other and keeps every feature, at a strength set by `alpha`. The grid steps
# `alpha` by powers of ten across ten orders of magnitude, because the useful value depends on
# the scale and the collinearity of the design matrix and neither is known in advance.
# - **Lasso** penalizes the sum of absolute coefficients, which drives some of them exactly to
# zero: it selects features rather than shrinking them. **ElasticNet** mixes the two.
#
# Lasso and ElasticNet are parameterized here by `alpha_frac` rather than a raw penalty. For any
# fold there is a threshold penalty $\alpha_{\max}$ - the smallest one that zeros every
# coefficient - which is computed from that fold's own data. `alpha_frac` is the fraction of it
# to apply, so one declared `alpha_frac` means the same thing on every fold, while a fixed raw
# penalty would mean something different on each.
# %%
configs = load_model_configs(
study,
"linear",
labels=LABELS or None,
config_names=CONFIG_NAMES or None,
)
configs
# %% [markdown]
# `LABELS` and `CONFIG_NAMES` both narrow what is fitted, and a narrowed run declares a different
# set of members than the canonical population does. A population is immutable once written, so
# such a run must publish under its own name: on a fresh workspace it would otherwise register an
# incomplete snapshot under the canonical one, and where the full population already exists the
# registry refuses it. Comparing the loaded rows against the complete declared catalog catches
# either knob, and says so here rather than several cells later in a message about hashes.
# %%
if narrows_declared_catalog(study, "linear", configs) and not POPULATION_NAME:
raise ValueError(
f"this run declares {configs.height} label-configuration pairs, which is not the "
"complete declared catalog, so it cannot publish the canonical population; pass "
"POPULATION_NAME to give it its own"
)
# %% [markdown]
# ## 2. Binding the declarations to the data
#
# A menu entry says which estimator to fit. It does not say which feature columns exist today,
# where the walk-forward folds fall, or which stock-date pairs have both a feature row and a
# label. **Resolving** a request is the step that goes and finds all of that: it reads the label
# and feature files, computes the fold boundaries from the walk-forward parameters in
# `config/setup.yaml`, works out the exact set of rows each fit is expected to predict, and turns
# any data-dependent hyperparameter into the number it will actually use - each fold's own
# $\alpha_{\max}$ times `alpha_frac`, in the case of Lasso.
#
# Resolving reads the inputs and fits nothing, so the plan below can be inspected before any
# computation starts. The three things to check in it:
#
# - **`feature_count`, `eligible_entities` and `eligible_rows` agree across every row.** They are
# the width of the design matrix, the number of stocks with a usable option surface, and the
# number of stock-date pairs to be predicted. Every configuration here reads the same feature
# matrix, so a row that differs is a configuration being measured on a different sample from
# its neighbours, and its results are not comparable with theirs.
# - **`folds` is the same everywhere**, and equals the number of walk-forward splits
# `05_evaluation` established.
# - **`validation_start` and `validation_end` bracket the development sample.** The held-out tail
# must not appear here: it is scored once, at the end of the case study, and any of it visible
# in this window would mean it had been used to choose something.
#
# Each row also carries a `training_hash`: the identity of that computation, derived from
# everything that can change its result. [`RUN_LOG.md`](../RUN_LOG.md#identity) sets out what goes
# into one and what follows from it.
# %%
requests = model_requests(
study,
configs,
execution_tier=EXECUTION_TIER,
preview_reductions=PREVIEW_REDUCTIONS,
)
resolved = tuple(request.resolve() for request in requests)
plan = resolved_model_plan(resolved)
plan.select(
"config_name",
"feature_count",
"eligible_entities",
"eligible_rows",
"folds",
"validation_start",
"validation_end",
)
# %% [markdown]
# ## 3. Fitting the population
#
# `run_model_population` fits every resolved request. For one request it walks the folds, and on
# each one:
#
# 1. takes the rows inside that fold's training window,
# 2. fills missing feature values with the training window's median for that column, then
# standardizes each column to zero mean and unit variance - both fitted on training rows only
# and then applied to the validation rows, so nothing from the validation window reaches the
# fit,
# 3. fits the estimator with that fold's resolved parameters,
# 4. predicts the fold's validation rows.
#
# The fold predictions are concatenated into one series covering the whole validation period,
# which is what a walk-forward prediction set is: each date predicted by a model that saw only
# data before it. The run then writes a `training_runs` row and the fitted coefficients, a
# `prediction_sets` row and the predictions themselves, and the metrics computed from them. It
# does this per configuration rather than once at the end, so an interruption costs the
# configuration in flight and nothing else.
#
# Every case study that fits linear models calls this same runner, which is what makes their
# results comparable. Unlike gradient boosting or a neural network, a linear model has no
# intermediate states worth scoring: there is one fit and therefore one checkpoint per
# configuration, and no learning curve to plot.
#
# **What the call publishes is a population**: a named, immutable list of the prediction sets it
# is going to produce. The list is computed from the resolved specifications before the first fit
# and written down, and afterwards every member must exist and be complete. That is what makes
# the downstream comparison well defined - `14_backtest` backtests this population, not whatever
# predictions happen to be in the registry - and it is why a configuration that raises fails the
# whole call rather than publishing a population one member short. Everything that finished stays
# registered, and re-running fits only what is missing.
# %%
population_name = POPULATION_NAME or "sp500_equity_option_analytics-linear-validation-v1"
execution, population = run_model_population(
study, resolved, population_name=population_name, supersedes=SUPERSEDES_POPULATION or None
)
fitted = sum(len(item["fitted_folds"]) for item in execution.diagnostics)
reused = sum(len(item["reused_folds"]) for item in execution.diagnostics)
print(f"{len(execution.runs)} configurations: {fitted} folds fitted, {reused} reused")
print(f"population {population.name}: {len(population.members)} prediction sets")
# %% [markdown]
# `reused` is not zero on a second run. Every identity is re-derived from the inputs, the
# registry already holds the matching rows, and the runner returns the stored result rather than
# fitting again - so re-running this notebook unchanged costs the time it takes to read the data.
#
# ### Running configurations of your own
#
# The published run log is read-only. To add runs, open the study against a workspace, which
# holds its own registry and artifacts and reads the same labels and features:
#
# ```python
# study = open_study("sp500_equity_option_analytics", workspace="~/ml4t-experiments")
# configs = load_model_configs(
# study, "linear", labels=["fwd_ret_5d"], config_names=["ols", "ridge_a1.0", "ridge_a3.0"]
# )
# requests = model_requests(study, configs)
# resolved = tuple(request.resolve() for request in requests)
# execution, population = run_model_population(study, resolved, population_name="my-linear-v1")
# ```
#
# `CONFIG_NAMES` fits a subset of what the menu already declares; a name the menu does not
# declare raises rather than quietly fitting fewer models than you asked for. To fit something
# new, add a preset at `case_studies/config/ridge/ridge_a3.0.yaml` and list `ridge_a3.0` under
# `linear:` in the label's menu. Editing an existing preset changes that configuration's
# identity, so its result registers as a new row beside the old one instead of replacing it.
#
# Give the run its own `population_name`: a name refers to one set of members permanently, and
# reusing it for a different set raises. Everything downstream reads the registry rather than the
# notebook, so predictions produced this way are selected and backtested on the same footing as
# the ones shipped here, inside your workspace.
# [`RUN_LOG.md`](../RUN_LOG.md#running-your-own-configurations) covers the rest, including how to
# rehearse on a reduced universe first.
# %% [markdown]
# ## 4. What came out
#
# One row per configuration and label, read back from the registry. `ic_mean` is the **information
# coefficient**: on each validation date, rank the stocks by the model's prediction, rank them by
# the return they went on to earn, correlate the two rankings, and average that daily correlation
# over the validation period. It measures whether the model ranks stocks correctly, on a scale
# where zero is no relationship, positive means the ranking points the right way, and negative
# means it points the wrong way.
#
# The catalog is joined to the menu on **both** `label` and `config_name`. A configuration name is
# unique within a label's menu and not across them: `ridge_a1.0` is declared by every regression
# label here, so joining on the name alone would multiply each result row by the number of labels
# that declare it and fill the table with copies carrying identical ICs.
#
# `ic_n_days` is how many validation dates produced a defined correlation, and coverage is judged
# against **each label's own** maximum. The labels do not offer the same number of dates to begin
# with - a ten-day forward window runs out earlier than a five-day one - so a single global
# maximum would mark a whole label incomplete for a reason that has nothing to do with any model.
# Within a label, the check is doing what it was built for: a model whose coefficients collapse to
# one or two features predicts nearly the same value for every stock on some dates, a constant has
# no rank correlation with anything, and its IC is then an average over a sample it selected
# itself. `full_coverage` marks the configurations measured on all of their label's dates.
# %% tags=["results"]
catalog = (
execution.catalog_rows.select(
"config_name",
"label",
"task",
"complete",
"ic_mean",
"ic_std",
"ic_n_days",
"auc_mean_daily",
"direction_label",
"n_folds",
"training_hash",
"prediction_hash",
)
.sort(["label", "ic_mean"], descending=[False, True])
.join(
configs.select("label", "config_name", "model_class", "params"),
on=["label", "config_name"],
how="left",
)
)
if catalog.filter(~pl.col("complete")).height:
raise RuntimeError("linear execution returned a partial prediction set")
catalog = catalog.with_columns(
full_coverage=pl.col("ic_n_days") == pl.col("ic_n_days").max().over("label")
)
primary = primary_label(study)
present = sorted(set(catalog.get_column("label")))
# The primary label leads when it was fitted. A subset run that leaves it out orders the panels by
# whichever label it did fit rather than by one that is not there.
panel_labels = [label for label in [primary] if label in present] + [
label for label in present if label != primary
]
print(f"{catalog.height} candidate models across {len(panel_labels)} labels")
catalog.select(
"label",
"config_name",
"model_class",
"params",
"ic_mean",
"ic_std",
"ic_n_days",
"full_coverage",
)
# %% [markdown]
# ### What the grid does on each label
#
# The frame below is the comparison this notebook exists to make, and it is only available because
# every declared label was fitted in one run. The features are the same and the folds are the
# same throughout. The grid is the same across the three regression labels, which share one menu
# of 28 configurations, so down those three rows the only thing that changes is what is being
# predicted. The two `fwd_dir_*` labels declare their own menu of 10 classifiers, because a
# ridge regression has nothing to say about a binary outcome; read those rows against each other
# and against their own regression sibling, not as two more members of the same sweep. The
# `configurations` column is what tells them apart.
#
# Read `n_positive` against `configurations`. A label where the whole grid sits on one side of zero
# is saying something about the label; a label where the grid straddles zero is saying that the
# spread across configurations is larger than anything the features contribute.
#
# Every row can carry `auc_mean_daily`, the within-date reading of how well a score separates
# the stocks that went up from those that did not, and `direction_label` says what it was scored
# against. A classification row scores its own label and leaves that column null. A regression
# row has no classes of its own, so its predicted return is scored as a ranking signal against a
# declared direction sibling: `fwd_ret_5d` against `fwd_dir_5d`, `fwd_ret_10d` against
# `fwd_dir_10d`. That makes a regression row and its sibling classification row comparable on
# one number - the same dates, the same outcome, two ways of getting at it. `fwd_ret_risk_adj_5d`
# declares no direction sibling, so it carries no AUC at all; null there means not computed, not
# zero.
# %% tags=["results"]
by_label = (
catalog.filter("full_coverage")
.group_by("label")
.agg(
task=pl.col("task").first(),
configurations=pl.len(),
scored_dates=pl.col("ic_n_days").max(),
best_ic=pl.col("ic_mean").max(),
worst_ic=pl.col("ic_mean").min(),
n_positive=(pl.col("ic_mean") > 0).sum(),
best_auc_daily=pl.col("auc_mean_daily").max(),
auc_scored_against=pl.col("direction_label").drop_nulls().first(),
)
.sort("best_ic", descending=True)
)
by_label
# %% [markdown]
# ### How the penalty grid ranks
#
# One panel per label on one shared vertical scale, so the labels are compared rather than
# each one rescaled to fill its own panel. Only configurations measured on all of their label's dates are
# charted. The zero line is the reference that matters: a bar below it is a model whose ranking
# pointed the wrong way out of sample.
#
# The configurations are held in one order across the panels - their ranking on the primary label
# - so a panel that does not descend is a label that orders the grid differently.
# %%
def compact(params: str) -> str:
"""Render declared parameters for a label: `alpha=1000000.0` reads as `alpha=1e+06`."""
return re.sub(r"\d+\.?\d*(?:[eE][+-]?\d+)?", lambda m: f"{float(m.group()):g}", params)
full = catalog.filter("full_coverage")
config_order = (
full.filter(pl.col("label") == panel_labels[0])
.sort("ic_mean", descending=True)
.get_column("config_name")
.to_list()
)
# `shared_yaxes` matches axes across columns, so with one column it does nothing and each
# panel would be rescaled to fill itself. Matching every row to the first is what puts the
# labels on one vertical scale, which is what stacking them is for.
fig_ic = make_subplots(
rows=len(panel_labels),
cols=1,
shared_xaxes=True,
vertical_spacing=0.04,
subplot_titles=[
f"{label} ({'primary' if label == primary else 'variant'})" for label in panel_labels
],
)
for row, label in enumerate(panel_labels, start=1):
panel = full.filter(pl.col("label") == label)
# A label whose menu declares different configurations - the classification ones do - keeps
# its own members and appends them after the shared order rather than being dropped.
order = [name for name in config_order if name in set(panel.get_column("config_name"))]
order += [name for name in panel.get_column("config_name").to_list() if name not in order]
panel = panel.with_columns(
rank=pl.col("config_name").replace_strict(
{name: index for index, name in enumerate(order)}, return_dtype=pl.Int32
)
).sort("rank")
best = panel.get_column("ic_mean").max()
fig_ic.add_trace(
go.Bar(
x=panel.get_column("config_name").to_list(),
y=panel.get_column("ic_mean").to_list(),
marker_color=[
COLORS["amber"] if value == best else COLORS["blue"]
for value in panel.get_column("ic_mean")
],
showlegend=False,
),
row=row,
col=1,
)
fig_ic.add_hline(
y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"], row=row, col=1
)
fig_ic.update_yaxes(title_text="Mean IC (validation)", row=row, col=1)
if row > 1:
fig_ic.update_yaxes(matches="y", row=row, col=1)
fig_ic.update_xaxes(
title_text=f"Configuration, ordered by rank on {panel_labels[0]}",
tickangle=-45,
row=len(panel_labels),
col=1,
)
fig_ic.update_layout(
title="Which side of zero the grid sits on depends on the label, not the penalty",
height=260 * len(panel_labels),
width=1100,
margin=dict(t=90),
)
# Which labels clear zero is a fact about the frame, so the alt text reads it rather than
# asserting it. Describing every bar as negative was true of the one label this notebook used to
# fit and is false of the set it fits now.
side_text = "; ".join(
f"{row['label']} has {row['n_positive']} of {row['configurations']} above zero"
for row in by_label.sort("label").iter_rows(named=True)
)
# Whether the panels overlap is also a fact about the frame. Stating a magnitude - "the spread
# within a panel is small next to the gap between panels" - is the thing that goes stale on the
# next run, so the two quantities are read and compared here.
leader = by_label.sort("best_ic", descending=True).row(0, named=True)
rest_best = by_label.filter(pl.col("label") != leader["label"]).get_column("best_ic").max()
if rest_best is None:
# A one-label `LABELS` run has no second panel, and `max()` over the empty selection is null
# rather than a number: comparing it raises, and formatting it would publish alt text about a
# comparison that was never made.
separation_text = (
f"there is one panel: this run fitted {leader['label']} alone, whose configurations "
f"span {leader['worst_ic']:.4f} to {leader['best_ic']:.4f}, so no comparison across "
"labels is drawn"
)
elif leader["worst_ic"] > rest_best:
separation_text = (
f"the panels do not overlap: {leader['label']}'s weakest configuration scores "
f"{leader['worst_ic']:.4f}, above the {rest_best:.4f} that is the best any other label "
"reaches"
)
else:
separation_text = (
f"the panels overlap: {leader['label']} leads on its best configuration "
f"({leader['best_ic']:.4f}) but its weakest ({leader['worst_ic']:.4f}) falls inside "
f"the range other labels reach, whose best is {rest_best:.4f}"
)
show_plotly_with_alt(
fig_ic,
"Bar charts of mean validation information coefficient for every full-coverage linear "
"configuration, one panel per declared label sharing a vertical scale, each panel's highest "
"bar in amber and the rest in dark navy, with a dashed zero line across each. The bars are "
"held in the primary label's ranking order in every panel, so a panel that does not descend "
f"is a label that orders the grid differently. Counted from the frame: {side_text}. Read "
f"off the same frame, {separation_text}.",
)
# %% [markdown]
# ### What shrinkage does on its own
#
# The bar chart mixes three estimators. Tracing IC across the Ridge penalty alone isolates the
# effect of shrinkage, with the estimator, the features and the folds all held fixed and only
# `alpha` moving. The alpha is read from each configuration's declared parameters rather than
# parsed out of its name, so the curve plots what was fitted. One line per label, because the
# question is whether the penalty behaves the same way regardless of what is being predicted.
# %%
ridge = (
catalog.filter(pl.col("model_class") == "Ridge")
.with_columns(alpha=pl.col("params").str.extract(r"alpha=([0-9.eE+-]+)").cast(pl.Float64))
.drop_nulls("alpha")
.sort("label", "alpha")
)
if ridge.height:
# One colour per label, and amber is not among them: the peak marker is amber, so a line
# drawn in it would swallow its own ring. Wrapping a short palette would give two labels
# the same colour and two identical legend swatches, so this raises instead.
line_colors = [
COLORS["blue"],
COLORS["copper"],
COLORS["positive"],
COLORS["negative"],
COLORS["slate"],
COLORS["recede"],
]
if len(panel_labels) > len(line_colors):
raise ValueError(
f"{len(panel_labels)} labels declared but only {len(line_colors)} distinct line "
"colours; add colours rather than letting two labels share one"
)
fig_alpha = go.Figure()
for index, label in enumerate(panel_labels):
series = ridge.filter(pl.col("label") == label)
if not series.height:
continue
log_alpha = np.log10(series.get_column("alpha").to_numpy())
values = series.get_column("ic_mean").to_numpy()
peak = int(np.argmax(values))
color = line_colors[index]
fig_alpha.add_trace(
go.Scatter(
x=log_alpha,
y=values,
mode="lines+markers",
name=label,
line=dict(color=color, width=2),
marker=dict(size=7, color=color),
)
)
fig_alpha.add_trace(
go.Scatter(
x=[log_alpha[peak]],
y=[values[peak]],
mode="markers",
marker=dict(size=13, color=COLORS["amber"], symbol="circle-open", line_width=3),
showlegend=False,
)
)
fig_alpha.add_hline(y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"])
fig_alpha.update_layout(
title="Ridge IC against penalty strength, over ten orders of magnitude",
height=520,
width=950,
margin=dict(t=70),
legend=dict(title_text="Label"),
)
fig_alpha.update_xaxes(title_text="log₁₀(α) (Ridge penalty strength)", zeroline=False)
fig_alpha.update_yaxes(title_text="Mean cross-sectional IC (validation)")
# Whether the label or the penalty dominates is the claim this figure is here to support,
# so it is measured rather than described: the widest sweep any one line covers, against
# the narrowest vertical gap between two lines at a penalty they both carry.
sweep_span = float(
ridge.group_by("label")
.agg(span=pl.col("ic_mean").max() - pl.col("ic_mean").min())
.get_column("span")
.max()
)
shared_alpha = (
ridge.group_by("alpha")
.agg(
n_labels=pl.col("label").n_unique(),
spread=pl.col("ic_mean").sort().diff().abs().min(),
)
.filter(pl.col("n_labels") > 1)
)
if not shared_alpha.height:
# One Ridge label carries no gap between lines to measure. Computing one gave nan, which
# fell through to the else branch and published "the closest two lines come within nan of
# each other" over a chart showing one line.
dominance = (
f"the one line drawn covers {sweep_span:.4f} across the sweep; a second label "
"carrying Ridge at the same penalties is what this chart would compare it against"
)
else:
gap = float(shared_alpha.get_column("spread").min())
dominance = (
f"the closest two lines come within {gap:.4f} of each other at any shared penalty, "
f"wider than the {sweep_span:.4f} the most penalty-sensitive line covers over the "
"whole sweep, so the label a line belongs to matters more than where on the line "
"it sits"
if gap > sweep_span
else (
f"the most penalty-sensitive line covers {sweep_span:.4f} across the sweep "
f"while the closest two lines come within {gap:.4f} of each other, so the "
"penalty moves a line by as much as the label separates them"
)
)
show_plotly_with_alt(
fig_alpha,
"Line chart of mean validation information coefficient against the base-ten logarithm of "
"the Ridge penalty, one line per label over ten orders of magnitude, each line's highest "
"point ringed in amber, against a dashed zero line, and no two lines sharing a colour. "
f"Measured from the frame: {dominance}.",
)
else:
print(
"No declared label carries a Ridge configuration, so there is no penalty sweep to "
"trace. Which estimators this section can show is decided by the per-label menus at "
"config/training/{label}.yaml."
)
# %% [markdown]
# ## 5. What to notice
#
# **What is being predicted decides which side of zero the grid sits on; the penalty does not.**
# `fwd_ret_risk_adj_5d` has every one of its configurations above zero and `fwd_ret_5d` has none
# of them there, and the two share a menu, a feature set and a fold schedule. Nothing inside the
# grid produces a separation of that kind: the alt text under the penalty sweep measures how far
# the most penalty-sensitive line moves against how close two lines ever come, and the sign of a
# label's whole grid is not something ten orders of magnitude of shrinkage reverses. What this
# does not say is that the gap between labels exceeds the spread inside every one of them: read
# `best_ic` against `worst_ic` in `by_label` and the leading label's own grid is wider than the
# distance from it to the next label. The claim is about which side of zero, not about
# magnitude. The top row of a grid sorted across labels
# reports which label was easiest, dressed as a model comparison, so the table below is grouped
# by label and never sorted across them.
#
# **These features are a forecast of dispersion, not of direction, and the labels test that
# directly.** That is a property of what an option price is. Implied volatility says how wide the
# market expects the distribution to be; skew says how asymmetric; the term structure says how
# that changes with horizon; the variance risk premium says how much the market charges for
# bearing it. None of them is a statement about the *mean* of the distribution, which is what a
# forward return is. A risk-adjusted return divides that mean by a measure of width, so a feature
# set that forecasts width well and direction badly should rank the risk-adjusted label better
# than the raw one. Read the `by_label` frame with that expectation in hand: it is a prediction
# the mechanism makes about which row leads, and it is the reason to fit more than one label
# rather than a decoration on having done so. `05_evaluation` screens these features one at a time
# and reaches the same place before any model is fitted.
#
# **Two folds, and one of them is 2020.** Half the validation evidence comes from a year in which
# implied volatility across every name moved together and by more than in the rest of the sample
# combined. A cross-sectional ranking asks which stock will out-return which, and that question
# gets harder when one factor is moving everything. Nothing above separates a weak feature set
# from an unrepresentative window, and with two folds nothing can. This applies to every panel,
# including whichever one leads.
#
# **What this does not rule out.** It says nothing about whether these features forecast the
# *volatility* of these stocks, which is the quantity they are actually about and which the risk
# chapters use them for. It says nothing about a strategy that trades the options rather than the
# underlying, which is the question the `sp500_options` case study asks. And a linear model can
# only represent a weighted sum of these columns, so it cannot express something like "skew
# matters when the term structure is inverted" unless someone builds that column first.
#
# **None of this selects anything.** IC measures whether predictions rank stocks correctly, not
# whether a strategy trading them makes money after costs and turnover. Selection is on validation
# backtest Sharpe over the population this notebook just published, and it happens in
# [`14_backtest`](14_backtest.ipynb).
#
# **Known limitations.** The IC here is an average of daily rank correlations with no adjustment
# for the serial dependence that overlapping multi-day returns create, so it is a ranking
# diagnostic rather than a test. The grid is a one-dimensional sweep of penalty strength at fixed
# features and fixed folds. And every number here is measured on the validation folds, which have
# been read many times over by the time a case study reaches this notebook.
#
# **Next**: [`07_gbm`](07_gbm.ipynb) asks whether gradient boosting finds structure a linear model
# cannot represent at all. The interaction named above is the concrete version of that question,
# and it is worth being clear in advance what a better result there would mean: with this many
# features and two folds, a tree ensemble that clears zero where the linear grid did not has
# either found a real interaction or fitted the 2020 window, and telling those apart is what
# [`13_model_analysis`](13_model_analysis.ipynb) is for.
```在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。