跳至正文
返回文库全部文档

检验期权隐含特征能否对未来股票收益进行排序

代码 《交易机器学习》

总结

本笔记评估期权市场指标能否对不同股票的未来收益进行排序。笔记使用期权衍生特征拟合已声明的线性模型,采用滚动验证,并比较岭回归、套索回归和弹性网络的惩罚强度。特征集包括各期限的隐含波动率、波动率偏斜、期限结构斜率和方差风险溢价等指标。正则化用于处理相关特征之间的冗余;套索回归和弹性网络还可以将部分系数缩减为零。

笔记强调,期权价格反映对未来分布的预期,尤其是波动率,而其主要标签关注收益水平或方向。笔记将信息系数作为排序诊断,而非盈利交易的证据或选择规则;策略选择留待之后的回测进行。验证历史较短,只有两个折,其中包括2020年的波动冲击,因此结果可能反映了不寻常的时期。除非明确提供相应特征,线性模型也无法捕捉交互作用。文档未能证明期权特征是否能成功预测波动率,或是否有助于期权交易策略。

核心观点

  • 期权衍生特征代表市场预期,与历史收益特征不同。
  • 滚动比较应让每种模型配置使用相同的股票、日期和数据折。
  • 岭回归会收缩系数,套索回归可将部分系数设为零;弹性网络结合了两种惩罚。
  • 信息系数评估排序,而非扣除成本后的交易表现或策略盈利能力。
  • 验证折只有两个,且包括一个波动异常的年份,因此结论的说服力有限。

标签

全文
# 06_linear.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Option analytics: can what the options market prices rank the stock
#
# Every other case study in this book builds its features from the history of the thing it is
# trying to predict - past returns, past volatility, past volume. This one does not. Its features
# come from the options written on each stock, and an option price is not a record of what
# happened. It is what someone was willing to pay today for a claim on what happens next. Implied
# volatility, the skew between puts and calls, the slope of the term structure, the gap between
# implied and realized variance: each is a statement the market is making about the future.
#
# That makes the question here sharper than "do these features work". The features are already
# forecasts, made by people with money at stake. If a forecast that good does not rank the
# cross-section, the interesting part is understanding what it is a forecast *of*.
#
# The design matrix has the redundancy every case study in this book has - implied volatility
# appears at three tenors and again as a rank, a percentile and two z-scores; the variance risk
# premium appears as a level, a rank and a z-score - so the penalty sweep has the same job here as
# elsewhere. **Regularization** adds a penalty on coefficient size to the fitting objective, and
# how much to apply is an empirical question this notebook answers by trying ten orders of
# magnitude of it.
#
# Two folds, and one of them validates on 2020. The usable history of this option analytics
# dataset is short, so `05_evaluation` set a walk-forward schedule whose validation windows are
# 2019 and 2020. Whatever the models find or fail to find, they are being asked about a period
# containing the fastest volatility shock in the sample.
#
# **Learning objectives.** By the end of this notebook you will be able to:
#
# - Read the set of models a case study has declared for a label, and say which estimator and
#   which hyperparameters each declared name resolves to.
# - Bind those declarations to the data on disk and check, before anything is fitted, that every
#   configuration will be measured on the same stocks, the same dates and the same folds.
# - Fit a population of models on walk-forward folds and publish one complete set of validation
#   predictions per configuration.
# - Say what a feature set of option-implied quantities is a forecast of, and why that is not the
#   same thing as the quantity this notebook asks it to rank.
# - Read a grid that sits entirely on one side of zero on one label and entirely on the other
#   side on another, and say what that is and is not evidence for.
# - Run configurations of your own into a private copy of the run log, and have them compared on
#   the same footing as the ones shipped here.
#
# **Book reference**: Chapter 11 (The ML Pipeline). Chapter 6, Section 6.7 (Search accounting and
# run logging) introduces the run log this notebook writes to.
#
# **Prerequisites**: [`03_financial_features`](03_financial_features.ipynb) and
# [`04_model_based_features`](04_model_based_features.ipynb) have written the feature matrices,
# and [`05_evaluation`](05_evaluation.ipynb) has established the walk-forward folds and screened
# the individual features.
#
# **What it writes**: one training run and one complete validation prediction set per
# configuration, in `run_log/registry.db` and under `run_log/training/` and
# `run_log/predictions/`, grouped under a named population.
# [`14_backtest`](14_backtest.ipynb) reads that population, runs every member against the
# equal-weight baseline, and selects on validation backtest Sharpe. **Selection happens there,
# not here.** This notebook ranks configurations by information coefficient to show what
# regularization does to this feature set; that ranking decides nothing.

# %%
"""Fit the declared option-analytics linear-model population on the walk-forward folds."""

import re

import numpy as np
import plotly.graph_objects as go
import polars as pl
from plotly.subplots import make_subplots

from case_studies.research import (
    declared_labels,
    load_model_configs,
    model_requests,
    narrows_declared_catalog,
    open_study,
    primary_label,
    resolved_model_plan,
    run_model_population,
)
from utils.style import COLORS, show_plotly_with_alt

# %% tags=["parameters"]
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_REDUCTIONS: dict = {}
CONFIG_NAMES: list[str] = []
POPULATION_NAME = ""
SUPERSEDES_POPULATION: str = ""

# %%
study = open_study(
    "sp500_equity_option_analytics",
    execution_tier=EXECUTION_TIER,
    workspace=WORKSPACE or None,
    entry_point="06_linear",
)

# %% [markdown]
# ## 1. Which label, and which models
#
# A label is the thing being predicted. This case study defines five in `config/setup.yaml`.
# `fwd_ret_5d`, the stock's total return over the five trading days after the decision date, is
# the primary one - the horizon the strategy chapters trade. `fwd_ret_10d` is the same idea over
# a longer horizon, `fwd_ret_risk_adj_5d` divides the move by a volatility estimate, and
# `fwd_dir_5d` and `fwd_dir_10d` are classification variants that ask for the direction rather
# than the size.
#
# **This notebook fits every declared label in one run.** Each label has its own training menu at
# `config/training/{label}.yaml`, listing family by family the named configurations to fit for
# that label. A label with no menu file has nothing declared and nothing to fit; the ones below
# are those that declare linear models. `LABELS` narrows the run to a subset, which is a
# diagnostic rather than the canonical population.

# %%
declared_labels(study, "linear")

# %% [markdown]
# Each name in the menu resolves to a preset file in the shared directory
# `case_studies/config/{model_type}/`, which holds that configuration's hyperparameters. The
# frame below is the menu for every declared label, with each name resolved to the estimator class it names
# and the arguments that class is constructed with. To change what runs, edit the menu or the
# presets rather than this notebook.
#
# The grid covers the two shapes a penalty can take:
#
# - **Ridge** penalizes the sum of squared coefficients. It shrinks correlated coefficients
#   towards each other and keeps every feature, at a strength set by `alpha`. The grid steps
#   `alpha` by powers of ten across ten orders of magnitude, because the useful value depends on
#   the scale and the collinearity of the design matrix and neither is known in advance.
# - **Lasso** penalizes the sum of absolute coefficients, which drives some of them exactly to
#   zero: it selects features rather than shrinking them. **ElasticNet** mixes the two.
#
# Lasso and ElasticNet are parameterized here by `alpha_frac` rather than a raw penalty. For any
# fold there is a threshold penalty $\alpha_{\max}$ - the smallest one that zeros every
# coefficient - which is computed from that fold's own data. `alpha_frac` is the fraction of it
# to apply, so one declared `alpha_frac` means the same thing on every fold, while a fixed raw
# penalty would mean something different on each.

# %%
configs = load_model_configs(
    study,
    "linear",
    labels=LABELS or None,
    config_names=CONFIG_NAMES or None,
)
configs

# %% [markdown]
# `LABELS` and `CONFIG_NAMES` both narrow what is fitted, and a narrowed run declares a different
# set of members than the canonical population does. A population is immutable once written, so
# such a run must publish under its own name: on a fresh workspace it would otherwise register an
# incomplete snapshot under the canonical one, and where the full population already exists the
# registry refuses it. Comparing the loaded rows against the complete declared catalog catches
# either knob, and says so here rather than several cells later in a message about hashes.

# %%
if narrows_declared_catalog(study, "linear", configs) and not POPULATION_NAME:
    raise ValueError(
        f"this run declares {configs.height} label-configuration pairs, which is not the "
        "complete declared catalog, so it cannot publish the canonical population; pass "
        "POPULATION_NAME to give it its own"
    )


# %% [markdown]
# ## 2. Binding the declarations to the data
#
# A menu entry says which estimator to fit. It does not say which feature columns exist today,
# where the walk-forward folds fall, or which stock-date pairs have both a feature row and a
# label. **Resolving** a request is the step that goes and finds all of that: it reads the label
# and feature files, computes the fold boundaries from the walk-forward parameters in
# `config/setup.yaml`, works out the exact set of rows each fit is expected to predict, and turns
# any data-dependent hyperparameter into the number it will actually use - each fold's own
# $\alpha_{\max}$ times `alpha_frac`, in the case of Lasso.
#
# Resolving reads the inputs and fits nothing, so the plan below can be inspected before any
# computation starts. The three things to check in it:
#
# - **`feature_count`, `eligible_entities` and `eligible_rows` agree across every row.** They are
#   the width of the design matrix, the number of stocks with a usable option surface, and the
#   number of stock-date pairs to be predicted. Every configuration here reads the same feature
#   matrix, so a row that differs is a configuration being measured on a different sample from
#   its neighbours, and its results are not comparable with theirs.
# - **`folds` is the same everywhere**, and equals the number of walk-forward splits
#   `05_evaluation` established.
# - **`validation_start` and `validation_end` bracket the development sample.** The held-out tail
#   must not appear here: it is scored once, at the end of the case study, and any of it visible
#   in this window would mean it had been used to choose something.
#
# Each row also carries a `training_hash`: the identity of that computation, derived from
# everything that can change its result. [`RUN_LOG.md`](../RUN_LOG.md#identity) sets out what goes
# into one and what follows from it.

# %%
requests = model_requests(
    study,
    configs,
    execution_tier=EXECUTION_TIER,
    preview_reductions=PREVIEW_REDUCTIONS,
)
resolved = tuple(request.resolve() for request in requests)

plan = resolved_model_plan(resolved)
plan.select(
    "config_name",
    "feature_count",
    "eligible_entities",
    "eligible_rows",
    "folds",
    "validation_start",
    "validation_end",
)

# %% [markdown]
# ## 3. Fitting the population
#
# `run_model_population` fits every resolved request. For one request it walks the folds, and on
# each one:
#
# 1. takes the rows inside that fold's training window,
# 2. fills missing feature values with the training window's median for that column, then
#    standardizes each column to zero mean and unit variance - both fitted on training rows only
#    and then applied to the validation rows, so nothing from the validation window reaches the
#    fit,
# 3. fits the estimator with that fold's resolved parameters,
# 4. predicts the fold's validation rows.
#
# The fold predictions are concatenated into one series covering the whole validation period,
# which is what a walk-forward prediction set is: each date predicted by a model that saw only
# data before it. The run then writes a `training_runs` row and the fitted coefficients, a
# `prediction_sets` row and the predictions themselves, and the metrics computed from them. It
# does this per configuration rather than once at the end, so an interruption costs the
# configuration in flight and nothing else.
#
# Every case study that fits linear models calls this same runner, which is what makes their
# results comparable. Unlike gradient boosting or a neural network, a linear model has no
# intermediate states worth scoring: there is one fit and therefore one checkpoint per
# configuration, and no learning curve to plot.
#
# **What the call publishes is a population**: a named, immutable list of the prediction sets it
# is going to produce. The list is computed from the resolved specifications before the first fit
# and written down, and afterwards every member must exist and be complete. That is what makes
# the downstream comparison well defined - `14_backtest` backtests this population, not whatever
# predictions happen to be in the registry - and it is why a configuration that raises fails the
# whole call rather than publishing a population one member short. Everything that finished stays
# registered, and re-running fits only what is missing.

# %%
population_name = POPULATION_NAME or "sp500_equity_option_analytics-linear-validation-v1"
execution, population = run_model_population(
    study, resolved, population_name=population_name, supersedes=SUPERSEDES_POPULATION or None
)

fitted = sum(len(item["fitted_folds"]) for item in execution.diagnostics)
reused = sum(len(item["reused_folds"]) for item in execution.diagnostics)
print(f"{len(execution.runs)} configurations: {fitted} folds fitted, {reused} reused")
print(f"population {population.name}: {len(population.members)} prediction sets")

# %% [markdown]
# `reused` is not zero on a second run. Every identity is re-derived from the inputs, the
# registry already holds the matching rows, and the runner returns the stored result rather than
# fitting again - so re-running this notebook unchanged costs the time it takes to read the data.
#
# ### Running configurations of your own
#
# The published run log is read-only. To add runs, open the study against a workspace, which
# holds its own registry and artifacts and reads the same labels and features:
#
# ```python
# study = open_study("sp500_equity_option_analytics", workspace="~/ml4t-experiments")
# configs = load_model_configs(
#     study, "linear", labels=["fwd_ret_5d"], config_names=["ols", "ridge_a1.0", "ridge_a3.0"]
# )
# requests = model_requests(study, configs)
# resolved = tuple(request.resolve() for request in requests)
# execution, population = run_model_population(study, resolved, population_name="my-linear-v1")
# ```
#
# `CONFIG_NAMES` fits a subset of what the menu already declares; a name the menu does not
# declare raises rather than quietly fitting fewer models than you asked for. To fit something
# new, add a preset at `case_studies/config/ridge/ridge_a3.0.yaml` and list `ridge_a3.0` under
# `linear:` in the label's menu. Editing an existing preset changes that configuration's
# identity, so its result registers as a new row beside the old one instead of replacing it.
#
# Give the run its own `population_name`: a name refers to one set of members permanently, and
# reusing it for a different set raises. Everything downstream reads the registry rather than the
# notebook, so predictions produced this way are selected and backtested on the same footing as
# the ones shipped here, inside your workspace.
# [`RUN_LOG.md`](../RUN_LOG.md#running-your-own-configurations) covers the rest, including how to
# rehearse on a reduced universe first.

# %% [markdown]
# ## 4. What came out
#
# One row per configuration and label, read back from the registry. `ic_mean` is the **information
# coefficient**: on each validation date, rank the stocks by the model's prediction, rank them by
# the return they went on to earn, correlate the two rankings, and average that daily correlation
# over the validation period. It measures whether the model ranks stocks correctly, on a scale
# where zero is no relationship, positive means the ranking points the right way, and negative
# means it points the wrong way.
#
# The catalog is joined to the menu on **both** `label` and `config_name`. A configuration name is
# unique within a label's menu and not across them: `ridge_a1.0` is declared by every regression
# label here, so joining on the name alone would multiply each result row by the number of labels
# that declare it and fill the table with copies carrying identical ICs.
#
# `ic_n_days` is how many validation dates produced a defined correlation, and coverage is judged
# against **each label's own** maximum. The labels do not offer the same number of dates to begin
# with - a ten-day forward window runs out earlier than a five-day one - so a single global
# maximum would mark a whole label incomplete for a reason that has nothing to do with any model.
# Within a label, the check is doing what it was built for: a model whose coefficients collapse to
# one or two features predicts nearly the same value for every stock on some dates, a constant has
# no rank correlation with anything, and its IC is then an average over a sample it selected
# itself. `full_coverage` marks the configurations measured on all of their label's dates.

# %% tags=["results"]
catalog = (
    execution.catalog_rows.select(
        "config_name",
        "label",
        "task",
        "complete",
        "ic_mean",
        "ic_std",
        "ic_n_days",
        "auc_mean_daily",
        "direction_label",
        "n_folds",
        "training_hash",
        "prediction_hash",
    )
    .sort(["label", "ic_mean"], descending=[False, True])
    .join(
        configs.select("label", "config_name", "model_class", "params"),
        on=["label", "config_name"],
        how="left",
    )
)

if catalog.filter(~pl.col("complete")).height:
    raise RuntimeError("linear execution returned a partial prediction set")

catalog = catalog.with_columns(
    full_coverage=pl.col("ic_n_days") == pl.col("ic_n_days").max().over("label")
)

primary = primary_label(study)
present = sorted(set(catalog.get_column("label")))
# The primary label leads when it was fitted. A subset run that leaves it out orders the panels by
# whichever label it did fit rather than by one that is not there.
panel_labels = [label for label in [primary] if label in present] + [
    label for label in present if label != primary
]
print(f"{catalog.height} candidate models across {len(panel_labels)} labels")
catalog.select(
    "label",
    "config_name",
    "model_class",
    "params",
    "ic_mean",
    "ic_std",
    "ic_n_days",
    "full_coverage",
)

# %% [markdown]
# ### What the grid does on each label
#
# The frame below is the comparison this notebook exists to make, and it is only available because
# every declared label was fitted in one run. The features are the same and the folds are the
# same throughout. The grid is the same across the three regression labels, which share one menu
# of 28 configurations, so down those three rows the only thing that changes is what is being
# predicted. The two `fwd_dir_*` labels declare their own menu of 10 classifiers, because a
# ridge regression has nothing to say about a binary outcome; read those rows against each other
# and against their own regression sibling, not as two more members of the same sweep. The
# `configurations` column is what tells them apart.
#
# Read `n_positive` against `configurations`. A label where the whole grid sits on one side of zero
# is saying something about the label; a label where the grid straddles zero is saying that the
# spread across configurations is larger than anything the features contribute.
#
# Every row can carry `auc_mean_daily`, the within-date reading of how well a score separates
# the stocks that went up from those that did not, and `direction_label` says what it was scored
# against. A classification row scores its own label and leaves that column null. A regression
# row has no classes of its own, so its predicted return is scored as a ranking signal against a
# declared direction sibling: `fwd_ret_5d` against `fwd_dir_5d`, `fwd_ret_10d` against
# `fwd_dir_10d`. That makes a regression row and its sibling classification row comparable on
# one number - the same dates, the same outcome, two ways of getting at it. `fwd_ret_risk_adj_5d`
# declares no direction sibling, so it carries no AUC at all; null there means not computed, not
# zero.

# %% tags=["results"]
by_label = (
    catalog.filter("full_coverage")
    .group_by("label")
    .agg(
        task=pl.col("task").first(),
        configurations=pl.len(),
        scored_dates=pl.col("ic_n_days").max(),
        best_ic=pl.col("ic_mean").max(),
        worst_ic=pl.col("ic_mean").min(),
        n_positive=(pl.col("ic_mean") > 0).sum(),
        best_auc_daily=pl.col("auc_mean_daily").max(),
        auc_scored_against=pl.col("direction_label").drop_nulls().first(),
    )
    .sort("best_ic", descending=True)
)
by_label

# %% [markdown]
# ### How the penalty grid ranks
#
# One panel per label on one shared vertical scale, so the labels are compared rather than
# each one rescaled to fill its own panel. Only configurations measured on all of their label's dates are
# charted. The zero line is the reference that matters: a bar below it is a model whose ranking
# pointed the wrong way out of sample.
#
# The configurations are held in one order across the panels - their ranking on the primary label
# - so a panel that does not descend is a label that orders the grid differently.


# %%
def compact(params: str) -> str:
    """Render declared parameters for a label: `alpha=1000000.0` reads as `alpha=1e+06`."""
    return re.sub(r"\d+\.?\d*(?:[eE][+-]?\d+)?", lambda m: f"{float(m.group()):g}", params)


full = catalog.filter("full_coverage")
config_order = (
    full.filter(pl.col("label") == panel_labels[0])
    .sort("ic_mean", descending=True)
    .get_column("config_name")
    .to_list()
)

# `shared_yaxes` matches axes across columns, so with one column it does nothing and each
# panel would be rescaled to fill itself. Matching every row to the first is what puts the
# labels on one vertical scale, which is what stacking them is for.
fig_ic = make_subplots(
    rows=len(panel_labels),
    cols=1,
    shared_xaxes=True,
    vertical_spacing=0.04,
    subplot_titles=[
        f"{label} ({'primary' if label == primary else 'variant'})" for label in panel_labels
    ],
)
for row, label in enumerate(panel_labels, start=1):
    panel = full.filter(pl.col("label") == label)
    # A label whose menu declares different configurations - the classification ones do - keeps
    # its own members and appends them after the shared order rather than being dropped.
    order = [name for name in config_order if name in set(panel.get_column("config_name"))]
    order += [name for name in panel.get_column("config_name").to_list() if name not in order]
    panel = panel.with_columns(
        rank=pl.col("config_name").replace_strict(
            {name: index for index, name in enumerate(order)}, return_dtype=pl.Int32
        )
    ).sort("rank")
    best = panel.get_column("ic_mean").max()
    fig_ic.add_trace(
        go.Bar(
            x=panel.get_column("config_name").to_list(),
            y=panel.get_column("ic_mean").to_list(),
            marker_color=[
                COLORS["amber"] if value == best else COLORS["blue"]
                for value in panel.get_column("ic_mean")
            ],
            showlegend=False,
        ),
        row=row,
        col=1,
    )
    fig_ic.add_hline(
        y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"], row=row, col=1
    )
    fig_ic.update_yaxes(title_text="Mean IC (validation)", row=row, col=1)
    if row > 1:
        fig_ic.update_yaxes(matches="y", row=row, col=1)
fig_ic.update_xaxes(
    title_text=f"Configuration, ordered by rank on {panel_labels[0]}",
    tickangle=-45,
    row=len(panel_labels),
    col=1,
)
fig_ic.update_layout(
    title="Which side of zero the grid sits on depends on the label, not the penalty",
    height=260 * len(panel_labels),
    width=1100,
    margin=dict(t=90),
)
# Which labels clear zero is a fact about the frame, so the alt text reads it rather than
# asserting it. Describing every bar as negative was true of the one label this notebook used to
# fit and is false of the set it fits now.
side_text = "; ".join(
    f"{row['label']} has {row['n_positive']} of {row['configurations']} above zero"
    for row in by_label.sort("label").iter_rows(named=True)
)
# Whether the panels overlap is also a fact about the frame. Stating a magnitude - "the spread
# within a panel is small next to the gap between panels" - is the thing that goes stale on the
# next run, so the two quantities are read and compared here.
leader = by_label.sort("best_ic", descending=True).row(0, named=True)
rest_best = by_label.filter(pl.col("label") != leader["label"]).get_column("best_ic").max()
if rest_best is None:
    # A one-label `LABELS` run has no second panel, and `max()` over the empty selection is null
    # rather than a number: comparing it raises, and formatting it would publish alt text about a
    # comparison that was never made.
    separation_text = (
        f"there is one panel: this run fitted {leader['label']} alone, whose configurations "
        f"span {leader['worst_ic']:.4f} to {leader['best_ic']:.4f}, so no comparison across "
        "labels is drawn"
    )
elif leader["worst_ic"] > rest_best:
    separation_text = (
        f"the panels do not overlap: {leader['label']}'s weakest configuration scores "
        f"{leader['worst_ic']:.4f}, above the {rest_best:.4f} that is the best any other label "
        "reaches"
    )
else:
    separation_text = (
        f"the panels overlap: {leader['label']} leads on its best configuration "
        f"({leader['best_ic']:.4f}) but its weakest ({leader['worst_ic']:.4f}) falls inside "
        f"the range other labels reach, whose best is {rest_best:.4f}"
    )
show_plotly_with_alt(
    fig_ic,
    "Bar charts of mean validation information coefficient for every full-coverage linear "
    "configuration, one panel per declared label sharing a vertical scale, each panel's highest "
    "bar in amber and the rest in dark navy, with a dashed zero line across each. The bars are "
    "held in the primary label's ranking order in every panel, so a panel that does not descend "
    f"is a label that orders the grid differently. Counted from the frame: {side_text}. Read "
    f"off the same frame, {separation_text}.",
)

# %% [markdown]
# ### What shrinkage does on its own
#
# The bar chart mixes three estimators. Tracing IC across the Ridge penalty alone isolates the
# effect of shrinkage, with the estimator, the features and the folds all held fixed and only
# `alpha` moving. The alpha is read from each configuration's declared parameters rather than
# parsed out of its name, so the curve plots what was fitted. One line per label, because the
# question is whether the penalty behaves the same way regardless of what is being predicted.

# %%
ridge = (
    catalog.filter(pl.col("model_class") == "Ridge")
    .with_columns(alpha=pl.col("params").str.extract(r"alpha=([0-9.eE+-]+)").cast(pl.Float64))
    .drop_nulls("alpha")
    .sort("label", "alpha")
)
if ridge.height:
    # One colour per label, and amber is not among them: the peak marker is amber, so a line
    # drawn in it would swallow its own ring. Wrapping a short palette would give two labels
    # the same colour and two identical legend swatches, so this raises instead.
    line_colors = [
        COLORS["blue"],
        COLORS["copper"],
        COLORS["positive"],
        COLORS["negative"],
        COLORS["slate"],
        COLORS["recede"],
    ]
    if len(panel_labels) > len(line_colors):
        raise ValueError(
            f"{len(panel_labels)} labels declared but only {len(line_colors)} distinct line "
            "colours; add colours rather than letting two labels share one"
        )
    fig_alpha = go.Figure()
    for index, label in enumerate(panel_labels):
        series = ridge.filter(pl.col("label") == label)
        if not series.height:
            continue
        log_alpha = np.log10(series.get_column("alpha").to_numpy())
        values = series.get_column("ic_mean").to_numpy()
        peak = int(np.argmax(values))
        color = line_colors[index]
        fig_alpha.add_trace(
            go.Scatter(
                x=log_alpha,
                y=values,
                mode="lines+markers",
                name=label,
                line=dict(color=color, width=2),
                marker=dict(size=7, color=color),
            )
        )
        fig_alpha.add_trace(
            go.Scatter(
                x=[log_alpha[peak]],
                y=[values[peak]],
                mode="markers",
                marker=dict(size=13, color=COLORS["amber"], symbol="circle-open", line_width=3),
                showlegend=False,
            )
        )
    fig_alpha.add_hline(y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"])
    fig_alpha.update_layout(
        title="Ridge IC against penalty strength, over ten orders of magnitude",
        height=520,
        width=950,
        margin=dict(t=70),
        legend=dict(title_text="Label"),
    )
    fig_alpha.update_xaxes(title_text="log₁₀(α)  (Ridge penalty strength)", zeroline=False)
    fig_alpha.update_yaxes(title_text="Mean cross-sectional IC (validation)")
    # Whether the label or the penalty dominates is the claim this figure is here to support,
    # so it is measured rather than described: the widest sweep any one line covers, against
    # the narrowest vertical gap between two lines at a penalty they both carry.
    sweep_span = float(
        ridge.group_by("label")
        .agg(span=pl.col("ic_mean").max() - pl.col("ic_mean").min())
        .get_column("span")
        .max()
    )
    shared_alpha = (
        ridge.group_by("alpha")
        .agg(
            n_labels=pl.col("label").n_unique(),
            spread=pl.col("ic_mean").sort().diff().abs().min(),
        )
        .filter(pl.col("n_labels") > 1)
    )
    if not shared_alpha.height:
        # One Ridge label carries no gap between lines to measure. Computing one gave nan, which
        # fell through to the else branch and published "the closest two lines come within nan of
        # each other" over a chart showing one line.
        dominance = (
            f"the one line drawn covers {sweep_span:.4f} across the sweep; a second label "
            "carrying Ridge at the same penalties is what this chart would compare it against"
        )
    else:
        gap = float(shared_alpha.get_column("spread").min())
        dominance = (
            f"the closest two lines come within {gap:.4f} of each other at any shared penalty, "
            f"wider than the {sweep_span:.4f} the most penalty-sensitive line covers over the "
            "whole sweep, so the label a line belongs to matters more than where on the line "
            "it sits"
            if gap > sweep_span
            else (
                f"the most penalty-sensitive line covers {sweep_span:.4f} across the sweep "
                f"while the closest two lines come within {gap:.4f} of each other, so the "
                "penalty moves a line by as much as the label separates them"
            )
        )
    show_plotly_with_alt(
        fig_alpha,
        "Line chart of mean validation information coefficient against the base-ten logarithm of "
        "the Ridge penalty, one line per label over ten orders of magnitude, each line's highest "
        "point ringed in amber, against a dashed zero line, and no two lines sharing a colour. "
        f"Measured from the frame: {dominance}.",
    )
else:
    print(
        "No declared label carries a Ridge configuration, so there is no penalty sweep to "
        "trace. Which estimators this section can show is decided by the per-label menus at "
        "config/training/{label}.yaml."
    )

# %% [markdown]
# ## 5. What to notice
#
# **What is being predicted decides which side of zero the grid sits on; the penalty does not.**
# `fwd_ret_risk_adj_5d` has every one of its configurations above zero and `fwd_ret_5d` has none
# of them there, and the two share a menu, a feature set and a fold schedule. Nothing inside the
# grid produces a separation of that kind: the alt text under the penalty sweep measures how far
# the most penalty-sensitive line moves against how close two lines ever come, and the sign of a
# label's whole grid is not something ten orders of magnitude of shrinkage reverses. What this
# does not say is that the gap between labels exceeds the spread inside every one of them: read
# `best_ic` against `worst_ic` in `by_label` and the leading label's own grid is wider than the
# distance from it to the next label. The claim is about which side of zero, not about
# magnitude. The top row of a grid sorted across labels
# reports which label was easiest, dressed as a model comparison, so the table below is grouped
# by label and never sorted across them.
#
# **These features are a forecast of dispersion, not of direction, and the labels test that
# directly.** That is a property of what an option price is. Implied volatility says how wide the
# market expects the distribution to be; skew says how asymmetric; the term structure says how
# that changes with horizon; the variance risk premium says how much the market charges for
# bearing it. None of them is a statement about the *mean* of the distribution, which is what a
# forward return is. A risk-adjusted return divides that mean by a measure of width, so a feature
# set that forecasts width well and direction badly should rank the risk-adjusted label better
# than the raw one. Read the `by_label` frame with that expectation in hand: it is a prediction
# the mechanism makes about which row leads, and it is the reason to fit more than one label
# rather than a decoration on having done so. `05_evaluation` screens these features one at a time
# and reaches the same place before any model is fitted.
#
# **Two folds, and one of them is 2020.** Half the validation evidence comes from a year in which
# implied volatility across every name moved together and by more than in the rest of the sample
# combined. A cross-sectional ranking asks which stock will out-return which, and that question
# gets harder when one factor is moving everything. Nothing above separates a weak feature set
# from an unrepresentative window, and with two folds nothing can. This applies to every panel,
# including whichever one leads.
#
# **What this does not rule out.** It says nothing about whether these features forecast the
# *volatility* of these stocks, which is the quantity they are actually about and which the risk
# chapters use them for. It says nothing about a strategy that trades the options rather than the
# underlying, which is the question the `sp500_options` case study asks. And a linear model can
# only represent a weighted sum of these columns, so it cannot express something like "skew
# matters when the term structure is inverted" unless someone builds that column first.
#
# **None of this selects anything.** IC measures whether predictions rank stocks correctly, not
# whether a strategy trading them makes money after costs and turnover. Selection is on validation
# backtest Sharpe over the population this notebook just published, and it happens in
# [`14_backtest`](14_backtest.ipynb).
#
# **Known limitations.** The IC here is an average of daily rank correlations with no adjustment
# for the serial dependence that overlapping multi-day returns create, so it is a ranking
# diagnostic rather than a test. The grid is a one-dimensional sweep of penalty strength at fixed
# features and fixed folds. And every number here is measured on the validation folds, which have
# been read many times over by the time a case study reaches this notebook.
#
# **Next**: [`07_gbm`](07_gbm.ipynb) asks whether gradient boosting finds structure a linear model
# cannot represent at all. The interaction named above is the concrete version of that question,
# and it is worth being clear in advance what a better result there would mean: with this many
# features and two folds, a tree ensemble that clears zero where the linear grid did not has
# either found a real interaction or fitted the 2020 window, and telling those apart is what
# [`13_model_analysis`](13_model_analysis.ipynb) is for.

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。