Zum Inhalt springen
Alle Bibliotheksdokumente

Optionsmerkmale mit Rang IC, Abhängigkeitskorrekturen und Mehrfachtestkontrolle bewerten

Code Machine Learning for Trading

Zusammenfassung

Dieses Notebook prüft mögliche Merkmale für eine Strategie, die Straddles am Geld verkauft und bis zum Verfall hält. Für jeden Handelstag vergleicht es die Querschnittsrangfolge verfügbarer Titel anhand eines Merkmals mit der Rangfolge ihrer späteren Straddle-Ergebnisse und mittelt die Rangübereinstimmung anschließend über Validierungsperioden. Es passt die Unsicherheit an überlappende Positionen an und kontrolliert Falschentdeckungen, da viele Kandidaten gemeinsam getestet werden.

Der Ablauf prüft außerdem Abdeckung und Aktualität, vergleicht die Beständigkeit über Zeiträume hinweg und hält Evidenz und Behandlungsentscheidungen in einem Triage-Protokoll fest. Er unterscheidet eine statistische Vorauswahl von der Modellauswahl: Der univariate Test kann Merkmale übersehen, die nur in Kombination nützlich sind, während spätere Modelle das vollständige Panel statt nur der ausgewählten Kandidaten verwenden. Das Notebook belegt keine Rentabilität nach Optionshandelskosten, und zwei Validierungsperioden liefern nur begrenzte Evidenz zur Stabilität; die Schwellen für die Auswahl enthalten Ermessensentscheidungen.

Kernaussagen

  • Messen Sie den Nutzen von Merkmalen anhand der täglichen Übereinstimmung von Querschnittsrängen und realisierten Ergebnissen.
  • Berücksichtigen Sie Abhängigkeiten, weil an aufeinanderfolgenden Tagen eröffnete Positionen überlappen.
  • Passen Sie Signifikanzbehauptungen daran an, dass viele Kandidaten zugleich getestet werden.
  • Betrachten Sie die Auswahl als univariaten Triage-Schritt, nicht als Nachweis eines Nutzens im kombinierten Modell.
  • Trennen Sie prognostischen Zusammenhang von Rentabilität nach Handelskosten.

Schlagwörter

Volltext
# 05_evaluation.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # S&P 500 Options: Feature Evaluation
#
# The strategy this case study builds sells a *straddle* - one call option and one put
# option struck at the money on the same name and expiring on the same day - about a
# month before expiry, and holds the position until the options expire. `03_financial_features`
# and `04_model_based_features` between them built a panel of candidates: quantities known
# about a name on the day the position would be opened. This notebook asks one question
# of each candidate on its own. On the days the strategy would trade, does that
# quantity's ordering of the names available line up with the ordering of what those
# names' straddles went on to earn?
#
# Asking it one candidate at a time is a filter, not a decision. It cannot see a quantity
# that carries information only in combination with another, and it says nothing about
# whether the strategy earns anything once the cost of trading options is paid. The
# modelling notebooks that follow do the first; Chapters 16 to 18 do the second.
#
# **What it reads**: `features/financial.parquet`, `features/model_based.parquet`, and
# the label files under `labels/`.
#
# **What it writes**: `evaluation/triage_ledger.parquet`, one row per candidate carrying
# the evidence gathered here and the handling decision that evidence earned, which the
# Chapter 20 cross-case-study notebook reads; and `evaluation/ic_timeseries.parquet`, the
# day-by-day agreement series Section 2 plots. The modelling notebooks read neither of
# them - they take the whole feature panel.
#
# **Learning Objectives**
#
# - Measure how well one quantity's ranking of the assets available on a day agrees with
#   the ranking of what those assets went on to earn, and average that agreement over a
#   period.
# - Widen the uncertainty around that average to allow for positions opened on
#   consecutive days overlapping in time, so that their outcomes are not independent.
# - Raise the bar a single result has to clear, to account for having asked the same
#   question of every candidate at once.
# - Check whether an association held over both of the periods it was measured on, or
#   came out of one of them.
# - Record, for every candidate, the decision it earned and the evidence behind it, in
#   the file a later chapter reads.
#
# **Book Reference**: Chapter 7, Section 7.3 (Univariate feature-label evaluation) and
# Section 7.4 (Search accounting and multiple testing), with Chapter 8, Section 8.6
# (Combining features and controlling search) as the secondary reference for search
# control.
#
# **Prerequisites**: [`03_financial_features`](03_financial_features.ipynb) and
# [`04_model_based_features`](04_model_based_features.ipynb).

# %%
"""Feature Evaluation - S&P 500 Options.

Consolidated evaluation of Ch8 financial features and Ch9 temporal features
against forward return labels. Produces a diagnostic triage ledger.
"""

from datetime import date, datetime

import numpy as np
import plotly.graph_objects as go
import polars as pl
import yaml
from IPython.display import display
from ml4t.diagnostic.evaluation.stats import benjamini_hochberg_fdr
from ml4t.diagnostic.metrics import compute_ic_hac_stats, compute_ic_uncertainty
from plotly.subplots import make_subplots

from case_studies.utils.feature_engineering import (
    assign_families,
    families_from_config,
    quantile_profile,
)
from utils import style
from utils.artifact_specs import resolve_label_buffer_unit
from utils.cv_splits import generate_cv_splits, load_evaluation_config
from utils.data_quality import top_entities
from utils.paths import get_case_study_dir
from utils.style import (  # COLORS registers the ml4t Plotly template on import
    COLORS,
    show_plotly_with_alt,
)

# %% tags=["parameters"]
# Production defaults (Papermill overrides for testing)
# Zero means all symbols.
MAX_SYMBOLS = 0

# %% [markdown]
# ### The one-month holding period, read from the configuration
#
# The straddle is sold about a month before it expires and held to expiry, so a position
# opened today and one opened tomorrow are alive at the same time over almost all of
# their lives. Three things in this notebook have to span exactly one holding period: the
# window the daily agreement series is smoothed over, the number of lags the uncertainty
# calculation has to allow for, and the horizon the second measurement uses. All three
# are counted from the one number `config/setup.yaml` declares, so that changing the
# strategy's holding period moves them together.

# %%
CASE_STUDY_ID = "sp500_options"
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
EVAL_DIR = CASE_DIR / "evaluation"
EVAL_DIR.mkdir(exist_ok=True)

JOIN_COLS = ["timestamp", "symbol"]
DATE_COL = "timestamp"

_setup = yaml.safe_load((CASE_DIR / "config" / "setup.yaml").read_text())
HOLD_SESSIONS = int(_setup["features"]["hold_sessions"])
IC_ROLLING_WINDOW = HOLD_SESSIONS

# %% [markdown]
# ### The five thresholds this notebook screens on
#
# None of these is estimated from the data. Each is a judgment about what the notebook is
# willing to pass on to the modelling chapters, and each one is stated here rather than
# written into the comparison that uses it, so that a reader can change one and re-run.
#
# | Setting | What it decides |
# |---|---|
# | `COVERAGE_MIN` | A candidate is dropped unless it has a value on at least this share of the panel's rows. Below it, whatever the screen measures comes from a subset of names and dates the notebook has not characterised. |
# | `STALENESS_MAX` | A candidate is dropped if it repeats the previous day's value on more than this share of rows. A quantity that rarely changes cannot re-rank the names between one trade and the next, whatever it correlates with. |
# | `FDR_ALPHA` | The share of the candidates called significant that the notebook accepts being wrong about, once every candidate is tested at once. |
# | `SIGN_CONSISTENCY_MIN` | The share of validation periods in which a candidate has to point the same way before the second promotion route will take it. |
# | `IC_THRESHOLD` | The size of agreement a consistent candidate also has to reach on that route. It is set at the resolution the screen can distinguish rather than at anything the strategy needs: an agreement smaller than this is not separable from zero on a two-period panel of this length. |
#
# `REDUNDANCY_CUT` is not a screen - nothing is dropped for it. It is the level above
# which Section 5 calls a pair of candidates the same evidence twice.

# %%
COVERAGE_MIN = 0.70
STALENESS_MAX = 0.50
FDR_ALPHA = 0.05
SIGN_CONSISTENCY_MIN = 0.60
IC_THRESHOLD = 0.01
REDUNDANCY_CUT = 0.70

# The columns financial.parquet carries so that a position can be priced, which are
# inputs to the strategy rather than candidates for it.
META_COLS = {
    "timestamp",
    "symbol",
    "instrument_id",
    "underlying_price",
    "instr_mid",
    "instr_bid",
    "instr_ask",
}

# %% [markdown]
# ## 0. The panel, and the period it may be measured over
#
# Two of the candidates are the output of a model rather than an arithmetic transform of
# a price: `04_model_based_features` fits a GARCH model and a stochastic-volatility model
# to each name's history and writes what each one says the name's volatility was. A model
# fitted on a stretch of history that includes the day being scored has seen the answer,
# so that notebook estimates its parameters on a **refit schedule**: a burn-in, then a fit
# on everything available at that point, which speaks for the following month and no
# earlier day, then a refit. Every value is therefore produced by parameters that stopped
# before the day it describes, on training days as well as validation days, and the file
# carries one row per name and day rather than one per name, day and fold.
#
# What this section keeps is the days that fall inside some fold's **validation period** -
# a period separated from that fold's training window by a gap wide enough that no training
# outcome is still unresolved when it starts. It is the days that are selected here, not an
# estimator: there is only one.
#
# The last year of the sample is a **holdout**: a stretch of data set aside untouched, so
# that the assessment several chapters later is made on prices nothing in the research
# has been chosen against. `config/setup.yaml` declares where it starts. Nothing here may
# read it - not the correlations, not the quantile profiles, not the screens - and
# because the outcome of a straddle sold before the holdout begins is not known until it
# expires, the boundary has to be applied to the day the position closes rather than to
# the day it opens.


# %% [markdown]
# ### Putting fold boundaries and panel dates on the same schema
#
# The split library reports its boundaries as timestamps while the panel is on daily
# dates. Normalising once here keeps every comparison below on one schema.


# %%
def _as_date(value: date | datetime) -> date:
    """Normalize pandas and Python boundary values to dates."""
    return value.date() if hasattr(value, "date") else value


# %% [markdown]
# A row is identified by its date and its symbol, and by nothing else. It used to carry a
# fold as well, because `04_model_based_features` fitted one estimator per fold and the fold
# said which estimator produced the value; it now fits on a refit schedule, so a date and a
# symbol have exactly one value whichever fold reads them. A missing or repeated key would
# leave that ambiguous, so the check runs before any value enters the panel rather than after.


# %%
def _validate_temporal_keys(temporal: pl.DataFrame) -> None:
    """Reject incomplete or ambiguous temporal keys, and a fold column that came back."""
    required = {"timestamp", "symbol"}
    missing = required.difference(temporal.columns)
    if missing:
        raise ValueError(f"Temporal artifact is missing required columns: {sorted(missing)}")

    if "fold" in temporal.columns:
        raise ValueError(
            "Temporal artifact carries a `fold` column. It is written on a refit schedule, so "
            "a date and symbol have one value; a fold column means it was produced by the "
            "per-fold design, under which a training row's parameters had read its own future."
        )

    null_keys = temporal.select(
        pl.any_horizontal([pl.col(key).is_null() for key in sorted(required)]).sum()
    ).item()
    if null_keys:
        raise ValueError(f"Temporal artifact has {null_keys} rows with null alignment keys")

    duplicate_keys = (
        temporal.group_by(["timestamp", "symbol"]).len().filter(pl.col("len") > 1).height
    )
    if duplicate_keys:
        raise ValueError(f"Temporal artifact has {duplicate_keys} duplicate date-symbol keys")


# %% [markdown]
# Keep the rows that fall inside some fold's validation period. There is no longer a wrong
# estimator to exclude: `04_model_based_features` produces one value per date and symbol,
# from parameters fitted on sessions ending before that date, so a value on a validation date
# is out of sample for every fold that scores it.
#
# What the old version had to do here, and no longer does, was exclude the pass fitted on
# everything up to the holdout - an estimator that had seen every validation date and whose
# values were therefore in sample. That pass does not exist under a refit schedule.
#
# The check that remains is the one that is not structural: a fold with no rows in its own
# validation period, which would mean the artifact was built against a different fold
# configuration.


# %%
def build_validation_temporal_panel(
    temporal: pl.DataFrame,
    cv_folds: list[dict],
) -> pl.DataFrame:
    """Select exactly the out-of-sample temporal estimate for each validation row."""
    _validate_temporal_keys(temporal)

    windows = []
    for split in cv_folds:
        fold_id = int(split["fold"])
        train_end = _as_date(split["train_end"])
        val_start = _as_date(split["val_start"])
        val_end = _as_date(split["val_end"])
        if train_end >= val_start:
            raise ValueError(f"Fold {fold_id} training ends inside its validation window")

        if temporal.filter(
            pl.col("timestamp").is_between(val_start, val_end, closed="both")
        ).is_empty():
            raise ValueError(
                f"Temporal artifact has no validation rows for fold {fold_id} "
                f"({val_start} through {val_end})"
            )
        windows.append((val_start, val_end))

    # A union, not a concatenation. Two folds' validation windows can overlap, and under the
    # per-fold artifact the overlap produced two rows for one date and symbol - one per fold -
    # which is why this function used to check for multiply assigned keys. There is one value
    # now, so an overlapping date is scored once.
    return temporal.filter(
        pl.any_horizontal(
            [pl.col("timestamp").is_between(start, end, closed="both") for start, end in windows]
        )
    )


# %% [markdown]
# Most case studies close a position a fixed number of days after opening it, so the day
# the outcome is known is the trading day plus a constant. This one holds to expiry, and
# the straddle picked on a given day is whichever listed expiry sits nearest a month out,
# so the gap runs from a few days short of a month to a few days over. The label file
# carries that gap per row as `dte_calendar`, and adding it to the trading day gives the
# day each individual outcome is known. That is what the holdout boundary is applied to.


# %%
def keep_outcomes_resolved_before_holdout(
    labels: pl.DataFrame, holdout_start: date
) -> pl.DataFrame:
    """Keep only labels whose option expires strictly before the holdout begins."""
    if "dte_calendar" not in labels.columns:
        raise ValueError("Primary label artifact is missing dte_calendar")
    if labels["dte_calendar"].null_count():
        raise ValueError("Primary label artifact has null expiry horizons")
    return labels.with_columns(
        (pl.col(DATE_COL) + pl.duration(days=pl.col("dte_calendar"))).alias("_label_end")
    ).filter(pl.col("_label_end") < holdout_start)


# %% [markdown]
# The other labels this notebook reads close a fixed number of trading sessions after
# entry rather than at an expiry, so their last day is found by counting positions along
# the panel's own session grid. Sessions, not calendar days: ten sessions run about a
# fortnight, so a calendar count places the exit too early and can pass a position that
# in fact settles inside the holdout.


# %%
def latest_exit_session(
    signal_dates: pl.Series, horizon_sessions: int, session_grid: pl.Series
) -> date:
    """The last session on which a position opened on one of *signal_dates* is closed."""
    exits = pl.DataFrame({DATE_COL: signal_dates.unique().sort()}).join(
        pl.DataFrame({DATE_COL: session_grid}).with_columns(
            session_grid.shift(-horizon_sessions - 1).alias("_exit")
        ),
        on=DATE_COL,
        how="left",
    )
    if exits["_exit"].null_count():
        raise ValueError(
            f"A signal date has no exit session {horizon_sessions} sessions ahead of it"
        )
    return exits["_exit"].max()


# %% [markdown]
# ### The two label definitions this notebook uses
#
# `setup.yaml` names the hold-to-expiry short-straddle return as the label the case study
# trades. Section 6 measures a second one alongside it - the return on the same position
# with the directional exposure hedged away daily and the position closed after ten
# sessions - to see how much of any agreement is about volatility rather than about
# direction. Which second label to use is this notebook's choice rather than something
# the configuration states, so the name is checked against the label set `setup.yaml`
# does declare instead of being trusted because a file with that name exists.

# %%
features = pl.read_parquet(CASE_DIR / "features" / "financial.parquet")
temporal = pl.read_parquet(CASE_DIR / "features" / "model_based.parquet")

_primary_name = _setup["labels"]["primary"]
_secondary_name = "fwd_ret_dh_10d"
_declared_variants = set(_setup["labels"]["variant_buffers"])
if _secondary_name not in _declared_variants:
    raise ValueError(
        f"Secondary label {_secondary_name!r} is not among the labels setup.yaml declares: "
        f"{sorted(_declared_variants)}"
    )
primary_label_df = pl.read_parquet(CASE_DIR / "labels" / f"{_primary_name}.parquet")
secondary_label_df = pl.read_parquet(CASE_DIR / "labels" / f"{_secondary_name}.parquet")

primary_label_col = [
    c for c in primary_label_df.columns if c not in META_COLS and c != "instrument_id"
][0]
secondary_label_col = [
    c for c in secondary_label_df.columns if c not in META_COLS and c != "instrument_id"
][0]

print(f"Primary label: {primary_label_col}")
print(f"Secondary label: {secondary_label_col}")

# %% [markdown]
# ### Folds, and the boundary the outcomes are held behind
#
# The fold boundaries come from one call to the shared split generator, given the same
# label gap `04_model_based_features` gave it, so the periods scored here are exactly the
# periods those estimators were out of sample over. Then the labels are cut back to the
# outcomes that had resolved before the holdout began. Every quantity computed below -
# the agreement measurements, the significance adjustment, the shape profiles, the
# decisions - descends from this frame, so holding the boundary once here holds it for
# all of them.

# %%
cv_folds = generate_cv_splits(
    features.select(DATE_COL),
    case_study_id=CASE_STUDY_ID,
    label_buffer=str(_setup["labels"]["buffer"]),
    # `ret_to_expiry` resolves on an expiration date, so its buffer is calendar days and
    # `setup.yaml` declares it. The splitter defaults to sessions, and a caller that does not
    # pass the declaration derives boundaries that disagree with the ones the models were
    # fitted on - the folds this notebook evaluates over have to be those folds.
    buffer_unit=resolve_label_buffer_unit(CASE_STUDY_ID, str(_setup["labels"]["primary"]), _setup),
)
evaluation_config = load_evaluation_config(CASE_STUDY_ID)
HOLDOUT_START = pl.Series([str(evaluation_config["holdout_start"])]).str.to_date().item()
temporal = build_validation_temporal_panel(temporal, cv_folds)
primary_selection_df = keep_outcomes_resolved_before_holdout(primary_label_df, HOLDOUT_START)

# %% [markdown]
# ### Joining features, model outputs and labels into one frame
#
# The join is declared one-to-one in both directions, so a duplicated key raises here
# rather than quietly multiplying a name's contribution to every statistic that follows.
# Being an inner join, it also drops any feature row for which no model output exists,
# which is what confines the frame to the validation periods.
#
# A row inside a validation period can now carry a null model-based value, and it is worth
# being precise about why. `04_model_based_features` fits on a refit schedule, so a name has
# no fitted volatility until its own burn-in has passed - and on this panel a name is only
# quoted while a straddle near the target maturity exists, so many names are short and reach
# that point late or never. Under the per-fold design the column was never null, because each
# fold fitted on its whole training window and then emitted backwards across it: completeness
# there was a symptom of the leak, not evidence against one.
#
# Those rows are dropped rather than kept, and counted rather than dropped quietly. Keeping
# them would score the stage-03 features on rows where the model-based ones cannot be
# evaluated, and this section exists to compare the two on the same rows.

# %%
financial_cols = [c for c in features.columns if c not in META_COLS]
temporal_cols = [c for c in temporal.columns if c not in JOIN_COLS]

eval_panel = features.join(temporal, on=JOIN_COLS, how="inner", validate="1:1")
eval_panel = eval_panel.join(
    primary_selection_df.select(JOIN_COLS + [primary_label_col, "_label_end"]),
    on=JOIN_COLS,
    how="inner",
    validate="1:1",
)

_before_burnin_drop = eval_panel.height
eval_panel = eval_panel.filter(~pl.any_horizontal([pl.col(col).is_null() for col in temporal_cols]))
_dropped = _before_burnin_drop - eval_panel.height
print(
    f"Rows dropped for carrying no model-based value: {_dropped:,} of "
    f"{_before_burnin_drop:,} ({_dropped / _before_burnin_drop:.1%})"
)
if eval_panel.is_empty():
    raise ValueError(
        "No validation row carries a model-based value. Either the burn-in in "
        "`setup.yaml::model_based` exceeds every segment's history, or "
        "`04_model_based_features` did not write the columns this section screens."
    )

# The holdout boundary, asserted on the frame rather than trusted from the filter that
# produced it: the day each straddle expires must fall before the holdout begins.
if eval_panel.filter(pl.col("_label_end") >= HOLDOUT_START).height:
    raise ValueError("A straddle held past the holdout boundary reached the evaluation panel")
eval_panel = eval_panel.drop("_label_end")

all_feature_cols = financial_cols + temporal_cols

if MAX_SYMBOLS > 0:
    # `top_entities` rather than a local reduction. The rule is the same - the most-observed
    # symbols - but the tie-break is not: this sorted on row count alone, and where counts tie the
    # winner came from frame order, which is not stable across runs or across callers. Two stages
    # reducing the same panel to the same size have to get the same universe, or a symbol only one
    # of them chose carries null features on the other and the run answers wrongly while looking
    # clean.
    eval_panel = eval_panel.filter(
        pl.col("symbol").is_in(top_entities(eval_panel, MAX_SYMBOLS, entity_col="symbol"))
    )

n_rows = len(eval_panel)
n_symbols = eval_panel["symbol"].n_unique()
n_dates = eval_panel[DATE_COL].n_unique()

# %% [markdown]
# ### How many names have to be quoted before a day's ranking means anything
#
# The measurement below is a correlation between two rankings of the names available on
# one day. On a day with a handful of names, that correlation takes a small number of
# values and none of them is informative, so days below a floor are dropped rather than
# averaged in. Thirty is the floor for the full universe. It is written as the smaller of
# thirty and the number of names actually loaded so that a reduced run - the smoke test
# uses five - narrows the floor with the universe instead of discarding every day and
# measuring nothing.

# %%
MIN_CROSS_SECTION = min(30, n_symbols)

print(
    f"Evaluation panel: {n_rows:,} rows, {n_symbols} names, {n_dates} trading days, "
    f"{eval_panel[DATE_COL].min()} to {eval_panel[DATE_COL].max()}"
)
print(f"Holdout begins {HOLDOUT_START} and is not read here")
print(f"A day enters the measurement with at least {MIN_CROSS_SECTION} names quoted")

# %% [markdown]
# ### What is in the candidate set
#
# `config/setup.yaml` groups the features `03_financial_features` writes into families,
# each with the reasoning that put it there, and that grouping is read rather than
# restated: a family named twice is a family that can disagree with itself. The two
# volatility models `04_model_based_features` fits are not in that register - it covers
# the arithmetic feature matrix - so they are named here, where they enter the screen.
#
# The table below is what the rest of the notebook runs on: how many candidates each
# family contributes, and the share of panel rows on which they carry a value. The
# families differ in that share by construction. A quantity read off the option quote
# exists on every row that has a quote; one standardised against a year of a name's own
# history exists only once that year has accumulated.

# %%
TEMPORAL_FAMILIES = {"garch_": "garch_volatility", "sv_": "stochastic_volatility"}

families = assign_families(financial_cols, families_from_config(_setup))
for column in temporal_cols:
    families[column] = next(
        family for prefix, family in TEMPORAL_FAMILIES.items() if column.startswith(prefix)
    )

family_register = (
    pl.DataFrame(
        {
            "feature": all_feature_cols,
            "family": [families[f] for f in all_feature_cols],
            "observed": [eval_panel[f].drop_nulls().len() / n_rows for f in all_feature_cols],
        }
    )
    .group_by("family")
    .agg(
        pl.len().alias("candidates"),
        pl.col("observed").min().round(3).alias("least observed"),
        pl.col("observed").max().round(3).alias("most observed"),
        pl.col("feature").sort().str.join(", ").alias("columns"),
    )
    # Families holding the same number of candidates would otherwise be ordered by
    # whatever the grouping emitted, which differs between runs, so the table printed
    # here would not be the one a reader re-running the notebook sees.
    .sort(["candidates", "family"], descending=[True, False])
)
with pl.Config(fmt_str_lengths=400, tbl_width_chars=220):
    display(family_register)

# %% [markdown]
# ### Candidates that barely move within a fold's validation period
#
# An estimator refitted monthly can still hand back a quantity that barely moves inside a
# fold's validation period while differing between names - a fitted long-run volatility
# level behaves that way. That is not the same thing as a feed that has gone stale, and it
# does not stop the quantity from ranking names against each other, so it is detected here
# and read differently by the staleness screen below.

# %%
fold_constant_features = set()
for feat in temporal_cols:
    unique_per_sym = eval_panel.group_by("symbol").agg(
        pl.col(feat).drop_nulls().n_unique().alias("n_unique")
    )["n_unique"]
    if unique_per_sym.mean() <= 3:
        fold_constant_features.add(feat)

print(f"Constant within a fold for most names: {sorted(fold_constant_features) or 'none'}")

# %% [markdown]
# ## 1. Can the candidate be read at all?
#
# Two things disqualify a candidate before any question about prediction is worth asking.
#
# The first is **coverage**: the share of the panel's rows on which it has a value at
# all. A candidate that is missing on a third of the panel is still measurable on the
# rows it has, but what it measures is an association over whichever names and dates
# happened to supply those rows, and that subset is not the universe the strategy trades.
#
# The second is **staleness**: the share of rows on which it repeats the value it had for
# the same name on the previous day. A quantity that almost never changes cannot re-order
# the names between one trade and the next, so it cannot drive a strategy that re-ranks
# them weekly, however well it correlates. The two quality flags the feature matrix
# carries are the clearest case - they are there so that a model can be checked for
# leaning on them, and they are the same value on nearly every row.

# %%
coverage = {}
staleness = {}

for feat in all_feature_cols:
    coverage[feat] = eval_panel[feat].drop_nulls().len() / n_rows

    df_sorted = eval_panel.select(JOIN_COLS + [feat]).sort(JOIN_COLS)
    unchanged = df_sorted.with_columns(
        (pl.col(feat) == pl.col(feat).shift(1).over("symbol")).alias("_same")
    )["_same"].sum()
    staleness[feat] = float(unchanged) / max(n_rows - n_symbols, 1)

readable = {
    feat: coverage[feat] >= COVERAGE_MIN and staleness[feat] <= STALENESS_MAX
    for feat in all_feature_cols
}

n_pass = sum(readable.values())
print(f"{n_pass} of {len(readable)} candidates clear both screens")

# %% [markdown]
# The candidates that do not clear them, with the two measurements that decided it. Read
# the two columns against `COVERAGE_MIN` and `STALENESS_MAX`: only one of them has to be
# on the wrong side.

# %%
screen_failures = pl.DataFrame(
    {
        "feature": [f for f, ok in readable.items() if not ok],
        "family": [families[f] for f, ok in readable.items() if not ok],
        "coverage": [round(coverage[f], 3) for f, ok in readable.items() if not ok],
        "staleness": [round(staleness[f], 3) for f, ok in readable.items() if not ok],
    }
)
display(screen_failures)

# %% [markdown]
# ## 2. Does the candidate rank the names the way the outcomes did?
#
# On each trading day, rank the names available by the candidate, rank the same names by
# what their straddles went on to earn, and correlate the two rankings. That correlation
# is the **information coefficient**, or IC. Ranks rather than values, because the
# strategy acts on an ordering - it sells the top of the list - and because a single
# straddle return can be many times the size of a typical one, which would let a handful
# of days set a correlation computed on values.
#
# One day gives one number. Repeating it over every day in the validation periods gives a
# series, and the average of that series is what the rest of the notebook screens on.
#
# Averaging a series is only as informative as the series is independent, and here it is
# not. A straddle sold on Monday and one sold on Tuesday are both alive for almost the
# same month, so they share almost all of the price path that decides them, and
# consecutive days' correlations move together. Treating them as independent draws would
# make the average look far more precisely measured than it is. The correction for this
# is a **heteroskedasticity- and autocorrelation-consistent**, or HAC, standard error: it
# estimates how much consecutive observations move together and widens the error bar
# accordingly. It has to be told how far that dependence reaches, which here is one
# holding period.
#
# Two kinds of candidate cannot have an IC and are separated out rather than scored as
# zero. One is a quantity that takes the same value for every name on a day - it orders
# nothing, though it may still be worth conditioning on in a model. The other is a
# candidate too sparsely observed to leave enough names on enough days.

# %%
evaluable_features = [f for f in all_feature_cols if readable[f]]

cs_std_df = eval_panel.group_by(DATE_COL).agg(
    [pl.col(f).std().alias(f) for f in evaluable_features]
)
date_level_features = set()
for feat in evaluable_features:
    mean_std = cs_std_df[feat].drop_nulls().mean()
    if mean_std is not None and mean_std < 1e-10:
        date_level_features.add(feat)

cs_features = [f for f in evaluable_features if f not in date_level_features]
print(f"Same for every name on a date, so unrankable: {sorted(date_level_features) or 'none'}")
print(f"{len(cs_features)} candidates go to the daily measurement")


# %% [markdown]
# The floor on the cross-section applies to each candidate separately, on the pairs it
# actually has. A day can carry hundreds of quoted names and still offer only a handful
# on which one sparse candidate and the outcome are both present, and a correlation
# computed from that handful would otherwise enter the average at the same weight as one
# computed from hundreds.


# %%
def daily_rank_ic(panel: pl.DataFrame, label_col: str, feature_cols: list[str]) -> pl.DataFrame:
    """One rank correlation per trading day per feature, over its own non-null pairs."""
    pairs = {
        f: (pl.col(f).is_not_null() & pl.col(label_col).is_not_null()).sum() for f in feature_cols
    }
    return (
        panel.group_by(DATE_COL)
        .agg(
            [
                pl.when(pairs[f] >= MIN_CROSS_SECTION)
                .then(pl.corr(f, label_col, method="spearman"))
                .alias(f)
                for f in feature_cols
            ]
            + [pl.len().alias("n_obs")]
        )
        .sort(DATE_COL)
    )


ic_wide = daily_rank_ic(eval_panel, primary_label_col, cs_features)
print(f"{len(cs_features)} candidates measured across {len(ic_wide):,} trading days")

# %% [markdown]
# The bandwidth is passed as the holding period rather than as a lag count, so the call
# says why it is what it is: the library takes the wider of one holding period and its
# own sample-size rule. A candidate whose series is shorter than twenty days is left
# unscored - an average over fewer days than that carries no useful error bar on a panel
# with this much overlap.

# %%
MIN_IC_DAYS = 20

ic_results = {}
ic_timeseries = {}
for feat in cs_features:
    ic_df = (
        ic_wide.select([DATE_COL, pl.col(feat).alias("ic"), "n_obs"])
        .drop_nulls(subset=["ic"])
        .filter(pl.col("ic").is_finite())
    )
    if len(ic_df) < MIN_IC_DAYS:
        continue
    ic_results[feat] = compute_ic_hac_stats(ic_df, ic_col="ic", label_horizon=HOLD_SESSIONS)
    ic_timeseries[feat] = ic_df

print(f"{len(ic_results)} candidates have a measurable daily series")

# %% [markdown]
# ### The series the averages come from
#
# Everything above is an average of a daily series, and two things an average cannot show
# live in the series itself: an association that comes from one stretch of the period
# rather than from all of it, and one that reverses direction inside it. Both change what
# the average means, and neither is visible in it. The series is the object; Section 7
# writes it to `evaluation/ic_timeseries.parquet`.
#
# The left panel is the daily series for the candidate with the largest average, with a
# smoothed line over one holding period, a zero line, and a dotted rule where one
# validation period ends and the next begins.
#
# The right panel puts three error bars on each of the leading candidates. The first
# assumes the daily measurements are independent, which they are not; it is drawn as the
# baseline a reader would get without thinking about the overlap. The second is the HAC
# interval described above. The third resamples the series in contiguous blocks at least
# one holding period long, which keeps the overlap intact in the resampled series instead
# of modelling it, and is the check on whether the second is doing its job.

# %%
leading_features = sorted(ic_results, key=lambda f: abs(ic_results[f]["mean_ic"]), reverse=True)[
    :10
]

uncertainty = {
    feat: compute_ic_uncertainty(ic_timeseries[feat], horizon=HOLD_SESSIONS, ic_col="ic")
    for feat in leading_features
}

# %% [markdown]
# The intervals and the reported significance are only comparable if they were computed
# over the same number of lags. Both calls derive that number from the holding period, so
# they agree by construction - which is exactly the kind of agreement worth asserting
# rather than assuming, because it would break silently if either call were changed.

# %%
mismatched_lags = {
    feat: (u["hac_lag"], ic_results[feat]["effective_lags"])
    for feat, u in uncertainty.items()
    if u["hac_lag"] != ic_results[feat]["effective_lags"]
}
if mismatched_lags:
    raise ValueError(f"Interval and t-statistic bandwidths disagree: {mismatched_lags}")

# %%
series_feature = leading_features[0]
series = (
    ic_timeseries[series_feature]
    .sort(DATE_COL)
    .with_columns(
        pl.col("ic").rolling_mean(IC_ROLLING_WINDOW, min_samples=IC_ROLLING_WINDOW).alias("ic_roll")
    )
)

fig = make_subplots(
    rows=1,
    cols=2,
    subplot_titles=[
        f"Daily rank IC of {series_feature}, and its {IC_ROLLING_WINDOW}-session mean",
        "Three ways of putting an error bar on one average",
    ],
    horizontal_spacing=0.14,
    column_widths=[0.55, 0.45],
)
fig.add_trace(
    go.Scatter(
        x=series[DATE_COL].to_list(),
        y=series["ic"].to_list(),
        mode="lines",
        line=dict(color=COLORS["silver_muted"], width=1),
        name="Daily IC",
    ),
    row=1,
    col=1,
)
fig.add_trace(
    go.Scatter(
        x=series[DATE_COL].to_list(),
        y=series["ic_roll"].to_list(),
        mode="lines",
        line=dict(color=COLORS["amber"], width=2.5),
        name="Rolling mean",
    ),
    row=1,
    col=1,
)
fig.add_hline(y=0, line=dict(color=COLORS["neutral"], width=1), row=1, col=1)
series_start, series_end = series[DATE_COL].min(), series[DATE_COL].max()
for split in cv_folds:
    boundary = _as_date(split["val_start"])
    if series_start < boundary < series_end:
        fig.add_vline(
            x=str(boundary),
            line=dict(color=COLORS["neutral"], width=1, dash="dot"),
            row=1,
            col=1,
        )

# %% [markdown]
# The right panel takes the same candidates in order of average agreement and draws all
# three intervals for each. Drawn on one row per candidate they would sit on top of each
# other and the widest would hide the rest, so each is offset vertically: the difference
# between the three is what the panel is for.

# %%
interval_order = list(reversed(leading_features))
BAND_OFFSET = 0.24
for rank, (band, color, label) in enumerate(
    [
        ("naive", COLORS["silver_muted"], "Naive 95%"),
        ("hac", COLORS["copper"], "HAC 95%"),
        ("boot", COLORS["blue"], "Block bootstrap 95%"),
    ]
):
    lower = [uncertainty[f][f"ci_{band}_lower"] for f in interval_order]
    upper = [uncertainty[f][f"ci_{band}_upper"] for f in interval_order]
    means = [uncertainty[f]["mean_ic"] for f in interval_order]
    fig.add_trace(
        go.Scatter(
            x=means,
            y=[i + (rank - 1) * BAND_OFFSET for i in range(len(interval_order))],
            mode="markers",
            marker=dict(color=color, size=6),
            error_x=dict(
                type="data",
                symmetric=False,
                array=[u - m for u, m in zip(upper, means, strict=True)],
                arrayminus=[m - lo for m, lo in zip(means, lower, strict=True)],
                color=color,
                thickness=1.4,
                width=3,
            ),
            text=interval_order,
            name=label,
        ),
        row=1,
        col=2,
    )
fig.add_vline(x=0, line=dict(color=COLORS["neutral"], width=1), row=1, col=2)

fig.update_layout(
    template="ml4t",
    height=560,
    width=1150,
    title={
        "text": (
            "Daily rank agreement, and how precisely its average is known"
            "<br><sup>Rank IC against the hold-to-expiry short-straddle return, over the "
            "validation periods only; the dotted rule marks where one validation period "
            "ends and the next begins</sup>"
        )
    },
    margin=dict(l=60, r=30, t=110, b=60),
    legend=dict(orientation="h", y=-0.16, x=0),
)
fig.update_xaxes(title_text="Validation date", row=1, col=1)
fig.update_xaxes(title_text="Average daily rank IC", row=1, col=2)
fig.update_yaxes(title_text="Daily rank IC", row=1, col=1)
fig.update_yaxes(
    title_text="Candidate",
    tickmode="array",
    tickvals=list(range(len(interval_order))),
    ticktext=interval_order,
    tickfont=dict(size=9),
    row=1,
    col=2,
)
show_plotly_with_alt(
    fig,
    "Two panels. On the left, the daily rank IC of ret_21d against the hold-to-expiry "
    "short-straddle return across the validation periods, drawn as a noisy grey daily series "
    "with a 21-session rolling mean over it; the rolling mean oscillates between about -0.15 "
    "and +0.18, crossing zero repeatedly, and a dotted vertical rule marks where one "
    "validation period ends and the next begins. On the right, the average daily rank IC for "
    "each of ten candidates, each with three error bars from a naive, a HAC and a block "
    "bootstrap interval. ret_21d and ret_10d have the largest positive averages at about 0.035 "
    "and 0.033, instr_delta the largest negative at about -0.018, and for every candidate the "
    "naive interval is the narrowest of the three.",
)

# %% [markdown]
# ### Did it hold over both periods, or over one?
#
# An average over the whole span can be carried by a single stretch of it. Splitting the
# same daily series by validation period and averaging each separately is the cheapest
# check on that, and it is the one Chapter 7 asks for. This case study is configured with
# two validation periods, so each candidate gets two averages. That is enough to see a
# sign change and not enough to support a quartile view across periods, which is why the
# per-period values are drawn rather than summarised.
#
# Direction agreement records the share of those periods in which the candidate points the
# same way as it does over the window as a whole. A candidate that ranks names inversely -
# low value, high straddle return - is as usable as one that ranks them directly, since a
# ranking is read in whichever direction it works, so what matters is whether the direction
# holds from period to period and not which direction it is. It is one of the two routes to
# promotion in Section 7, and the figure below shows the per-period averages behind it.
#
# The weakest and strongest periods written to the ledger follow the same rule: the weakest
# is the period furthest against the candidate's own direction, which for an inverse
# candidate is its algebraic maximum rather than its minimum.

# %%
fold_stats = {}
for feat in ic_results:
    fold_ics = {}
    ts = ic_timeseries[feat]
    for split in cv_folds:
        fold_start = _as_date(split["val_start"])
        fold_end = _as_date(split["val_end"])
        fold_ic = ts.filter(pl.col(DATE_COL).is_between(fold_start, fold_end, closed="both"))
        if len(fold_ic) >= 5:
            fold_ics[int(split["fold"])] = float(fold_ic["ic"].mean())

    if fold_ics:
        values = list(fold_ics.values())
        direction = 1.0 if (ic_results[feat]["mean_ic"] or 0.0) >= 0 else -1.0
        signed = [ic * direction for ic in values]
        fold_stats[feat] = {
            "n_folds": len(fold_ics),
            "fold_ics": fold_ics,
            "sign_consistency": sum(1 for s in signed if s > 0) / len(values),
            "worst_fold_ic": values[int(np.argmin(signed))],
            "best_fold_ic": values[int(np.argmax(signed))],
            "median_fold_ic": float(np.median(values)),
        }

n_consistent = sum(1 for s in fold_stats.values() if s["sign_consistency"] >= SIGN_CONSISTENCY_MIN)
print(
    f"{n_consistent} of {len(fold_stats)} candidates hold one direction in at least "
    f"{SIGN_CONSISTENCY_MIN:.0%} of the validation periods"
)

# %% [markdown]
# One bar per validation period, for the candidates with the largest average agreement. A
# candidate whose two bars point the same way carried one direction through both periods.
# A candidate whose bars point opposite ways has an average taken across a sign change,
# which is exactly what the average hides and what this figure exists to show.
#
# The split generator numbers periods from the most recent backwards, so the legend gives
# each one its dates rather than its number alone.

# %%
fold_plot_features = [f for f in leading_features if f in fold_stats]

fold_windows = {
    int(split["fold"]): (_as_date(split["val_start"]), _as_date(split["val_end"]))
    for split in cv_folds
}

fig = go.Figure()
plot_fold_ids = sorted({fid for f in fold_plot_features for fid in fold_stats[f]["fold_ics"]})
for position, fold_id in enumerate(plot_fold_ids):
    val_start, val_end = fold_windows[fold_id]
    fig.add_trace(
        go.Bar(
            x=[fold_stats[f]["fold_ics"].get(fold_id) for f in fold_plot_features],
            y=fold_plot_features,
            orientation="h",
            marker_color=[COLORS["blue"], COLORS["amber"]][position % 2],
            name=f"Fold {fold_id}: {val_start:%b %Y} - {val_end:%b %Y}",
        )
    )
fig.add_vline(x=0, line=dict(color=COLORS["neutral"], width=1))
fig.update_layout(
    template="ml4t",
    height=520,
    width=900,
    barmode="group",
    title={
        "text": (
            "Some of the leading candidates change direction between periods"
            "<br><sup>Average daily rank IC inside each validation period, for the candidates "
            "with the largest average over the two periods combined</sup>"
        )
    },
    margin=dict(l=140, r=30, t=110, b=60),
    legend=dict(orientation="h", y=-0.12, x=0),
)
fig.update_xaxes(title_text="Average daily rank IC within the period")
fig.update_yaxes(title_text="Candidate", autorange="reversed")
show_plotly_with_alt(
    fig,
    "Grouped horizontal bar chart of the average daily rank IC inside each validation period, "
    "two bars per candidate, for the ten candidates with the largest average over the two "
    "periods combined. Seven keep their side in both periods: ret_21d, ret_10d, ret_5d, "
    "instr_ret_5d and the two iv_rv_ratio columns positive, instr_delta negative. Three do not - "
    "iv_mom_10d, vrp_mom_10d and vrp_mom_5d are each negative in fold 0 and positive in fold 1, "
    "iv_mom_10d swinging from about -0.032 to about +0.005. So the three that change direction "
    "between the periods are all momentum terms, and ret_5d, while it keeps its sign, falls from "
    "about 0.036 to about 0.011.",
)

# %% [markdown]
# ## 3. What the search cost
#
# A p-value is a statement about one test. Test fifty of them at a threshold of one in
# twenty and two or three will cross it with nothing behind them, so a result cannot be
# read without knowing how many questions were asked to get it. That is why the set of
# things tested has to be stated before any of the p-values are read.
#
# **The set searched here** is every candidate that cleared the two screens in Section 1
# and had enough of a cross-section to measure, tested against one label - the
# hold-to-expiry return the case study trades. Section 4 measures two shorter horizons
# and Section 6 a hedged variant, but neither feeds a decision: a candidate is promoted
# or not on the primary label alone, so those measurements do not widen the set. Nothing
# was tested and dropped from the count, and the candidate set was fixed by
# `03_financial_features` and `04_model_based_features` before any of it was measured.
#
# **The Benjamini-Hochberg procedure** adjusts for that. It sorts the p-values, and
# accepts the largest one whose rank-scaled threshold it still clears, along with
# everything below it. What it controls is the *false discovery rate*: of the candidates
# it calls significant, the expected share that are not is at most `FDR_ALPHA`. That is a
# weaker and more useful guarantee than requiring no false positive at all, which on a
# panel this size would leave nothing.
#
# The three counts printed below - candidates that clear a plain threshold, that clear
# the overlap-aware one, and that clear the family-wide one - are nested, and the ratio
# between the first and the others is how much of an apparent finding is an artefact of
# not correcting.

# %%
feature_names = list(ic_results.keys())
p_values = [ic_results[f]["p_value"] for f in feature_names]

fdr_result = benjamini_hochberg_fdr(p_values, alpha=FDR_ALPHA, return_details=True)

eval_summary = pl.DataFrame(
    {
        "feature": feature_names,
        "source": ["temporal" if f in temporal_cols else "financial" for f in feature_names],
        "ic_mean": [ic_results[f]["mean_ic"] for f in feature_names],
        "hac_se": [ic_results[f]["hac_se"] for f in feature_names],
        "hac_t": [ic_results[f]["t_stat"] for f in feature_names],
        "hac_p": p_values,
        "fdr_p": [float(p) for p in fdr_result["adjusted_p_values"]],
        "fdr_sig": [bool(r) for r in fdr_result["rejected"]],
        "naive_t": [ic_results[f]["naive_t_stat"] for f in feature_names],
    },
    schema_overrides={
        "ic_mean": pl.Float64,
        "hac_se": pl.Float64,
        "hac_t": pl.Float64,
        "hac_p": pl.Float64,
        "fdr_p": pl.Float64,
        "fdr_sig": pl.Boolean,
        "naive_t": pl.Float64,
    },
).sort(pl.col("ic_mean").cast(pl.Float64, strict=False).abs(), descending=True)

# %%
n_significant_naive = sum(
    1 for feature in feature_names if abs(ic_results[feature]["naive_t_stat"]) > 1.96
)
n_significant_hac = sum(1 for p_value in p_values if p_value < FDR_ALPHA)
n_significant_fdr = int(fdr_result["n_rejected"])

inflation_hac = n_significant_naive / max(n_significant_hac, 1)
inflation_fdr = n_significant_naive / n_significant_fdr if n_significant_fdr else float("inf")

print(f"Candidates tested: {len(feature_names)}")
print(f"  clearing a plain two-sided threshold at |t| > 1.96: {n_significant_naive}")
print(f"  still clearing it once the overlap is allowed for:  {n_significant_hac}")
print(f"  clearing the family-wide adjustment at {FDR_ALPHA}:        {n_significant_fdr}")
print(f"Allowing for the overlap alone removes a factor of {inflation_hac:.2f}")
if np.isfinite(inflation_fdr):
    print(f"Allowing for both removes a factor of {inflation_fdr:.2f}")
else:
    print("The family-wide adjustment leaves nothing, so the second factor is undefined")

# %% [markdown] tags=["results"]
# **Screen result.** 49 of the 51 candidates reach the daily measurement; the two dropped
# are the solver-quality flags, which hold the same value on every panel row. 13 of the
# 49 clear a plain threshold, 3 still clear it once the overlap between consecutive
# hold-to-expiry positions is allowed for, and none clears the family-wide adjustment.
# Allowing for the overlap alone removes a factor of 4.33, and the family-wide adjustment
# then removes what is left. The largest average daily rank agreement in the screen is
# 0.0349, on the 21-session return of the underlying, whose adjusted p-value is 0.263.

# %% [markdown]
# The left panel ranks candidates by the size of their average agreement, drawn
# horizontally so the names stay readable. A bar is coloured when that candidate's
# average clears the overlap-aware threshold on its own; the colour says nothing about
# the family-wide decision, which the right panel and Section 7 carry.

# %%
top_n = min(15, len(eval_summary))
top = eval_summary.head(top_n).sort("ic_mean")

fig = make_subplots(
    rows=1,
    cols=2,
    subplot_titles=[
        "Largest average agreements, by candidate",
        "How far the overlap correction moves each one",
    ],
    horizontal_spacing=0.16,
)

for is_sig, label, color in [
    (False, "Not individually significant", COLORS["silver_muted"]),
    (True, "Individually significant, overlap allowed for", COLORS["blue"]),
]:
    subset = top.filter((pl.col("hac_p") < FDR_ALPHA) == is_sig)
    fig.add_trace(
        go.Bar(
            x=subset["ic_mean"].to_list(),
            y=subset["feature"].to_list(),
            orientation="h",
            marker_color=color,
            name=label,
        ),
        row=1,
        col=1,
    )

# %% [markdown]
# The right panel puts each candidate's plain t-statistic against its overlap-aware one.
# Points below the diagonal in the upper half, and above it in the lower half, are
# candidates the correction has pulled towards zero; the distance from the diagonal is
# how much of the plain figure came from treating overlapping positions as independent.

# %%
for is_sig, label, color in [
    (False, "Not individually significant", COLORS["silver_muted"]),
    (True, "Individually significant, overlap allowed for", COLORS["blue"]),
]:
    subset = eval_summary.filter((pl.col("hac_p") < FDR_ALPHA) == is_sig)
    fig.add_trace(
        go.Scatter(
            x=subset["naive_t"].to_list(),
            y=subset["hac_t"].to_list(),
            mode="markers",
            marker=dict(color=color, size=8),
            text=subset["feature"].to_list(),
            showlegend=False,
        ),
        row=1,
        col=2,
    )

# %%
max_t = (
    max(
        float(eval_summary["naive_t"].abs().max() or 1.0),
        float(eval_summary["hac_t"].abs().max() or 1.0),
    )
    * 1.1
)
fig.add_trace(
    go.Scatter(
        x=[-max_t, max_t],
        y=[-max_t, max_t],
        mode="lines",
        line=dict(dash="dash", color=COLORS["neutral"]),
        showlegend=False,
    ),
    row=1,
    col=2,
)

fig.update_layout(
    template="ml4t",
    height=540,
    width=1100,
    title={
        "text": (
            "Allowing for the overlap pulls the strongest evidence towards zero"
            "<br><sup>Average daily rank IC against the hold-to-expiry short-straddle return, "
            "and the t-statistic on that average before and after the correction</sup>"
        )
    },
    margin=dict(l=110, r=30, t=110, b=60),
    legend=dict(orientation="h", y=-0.16, x=0),
)
fig.update_xaxes(title_text="Average daily rank IC", zeroline=True, row=1, col=1)
fig.update_xaxes(title_text="t-statistic, overlap ignored", row=1, col=2)
fig.update_yaxes(title_text="Candidate", row=1, col=1)
fig.update_yaxes(title_text="t-statistic, overlap allowed for", row=1, col=2)
show_plotly_with_alt(
    fig,
    "Two panels. On the left, a horizontal bar chart of the average daily rank IC by "
    "candidate, ordered by size, with only ret_21d, ret_10d and instr_delta shaded as "
    "individually significant once the overlap is allowed for and the other twelve left pale. "
    "On the right, a scatter of the t-statistic with the overlap ignored against the "
    "t-statistic with it allowed for, one point per candidate, with a dashed diagonal marking "
    "where the two would agree. Every point sits inside the diagonal, so each t-statistic "
    "shrinks towards zero; the three largest move the most, ret_21d from 6.0 to 2.4 and ret_10d "
    "from 5.7 to 2.7 on the positive side, instr_delta from -4.6 to -2.4 on the negative.",
)

# %% [markdown]
# The gap between what the two panels support is the point of the section. Some candidates
# still stand out individually once the overlap is allowed for; whether any of them
# clears the bar set by having asked the question of every candidate at once is a
# separate question, and the answer is what the promotion rule in Section 7 acts on.

# %% [markdown]
# ### How far ahead does the agreement reach?
#
# Everything above is measured against one outcome, held for about a month. A quantity
# whose agreement is concentrated at a few days and gone by ten cannot be traded at this
# cadence, and one still reaching a month out could be traded less often and more
# cheaply. Neither is visible from a single horizon, so the same measurement is repeated
# against the two shorter unhedged outcomes `02_labels` also writes - the return over the
# next five and the next ten sessions.
#
# These are diagnostic only. No decision in Section 7 reads them, so measuring them does
# not widen the set of tests the adjustment above has to cover. The traded outcome is
# plotted at the holding period `setup.yaml` declares, which is where a one-month expiry
# falls on average; its own horizon varies by a few sessions from name to name.
#
# The right panel divides each average by the standard deviation of its own daily series.
# That ratio says how large the average is against the variation it was drawn from, which
# is what decides whether a longer horizon is genuinely more reliable or just measured
# over a smoother series. The usual form of this ratio is taken across validation
# periods, which two of them cannot support.

# %%
HORIZON_LABELS = {
    "fwd_ret_5d": int(_setup["labels"]["variant_buffers"]["fwd_ret_5d"].rstrip("D")),
    "fwd_ret_10d": int(_setup["labels"]["variant_buffers"]["fwd_ret_10d"].rstrip("D")),
    _primary_name: HOLD_SESSIONS,
}
session_grid = features[DATE_COL].unique().sort()

horizon_ic = {}
for label_name, horizon in sorted(HORIZON_LABELS.items(), key=lambda item: item[1]):
    if label_name == _primary_name:
        horizon_panel, horizon_col = eval_panel, primary_label_col
    else:
        horizon_labels = pl.read_parquet(CASE_DIR / "labels" / f"{label_name}.parquet")
        horizon_col = [
            c for c in horizon_labels.columns if c not in META_COLS and c != "instrument_id"
        ][0]
        horizon_panel = eval_panel.join(
            horizon_labels.select(JOIN_COLS + [horizon_col]), on=JOIN_COLS, how="inner"
        )
        exit_max = latest_exit_session(horizon_panel[DATE_COL], horizon, session_grid)
        if exit_max >= HOLDOUT_START:
            raise ValueError(f"A {label_name} position closing {exit_max} reaches the holdout")

    wide = daily_rank_ic(horizon_panel, horizon_col, leading_features)
    horizon_ic[horizon] = {
        feat: (
            float(wide[feat].mean()),
            float(wide[feat].mean() / wide[feat].std()) if wide[feat].std() else float("nan"),
        )
        for feat in leading_features
    }

# %%
horizons = sorted(horizon_ic)
fig = make_subplots(
    rows=1,
    cols=2,
    subplot_titles=[
        "Average agreement by horizon",
        "The same, against its own daily variation",
    ],
    horizontal_spacing=0.12,
)
# The three largest are named and given their own colour; the rest are drawn as one grey
# group, because ten distinguishable colours is not a legend anyone reads.
EMPHASIS_COLORS = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]

for panel, index in [(1, 0), (2, 1)]:
    for rank, feat in enumerate(leading_features):
        emphasised = rank < len(EMPHASIS_COLORS)
        fig.add_trace(
            go.Scatter(
                x=horizons,
                y=[horizon_ic[h][feat][index] for h in horizons],
                mode="lines+markers" if emphasised else "lines",
                line=dict(
                    color=EMPHASIS_COLORS[rank] if emphasised else COLORS["recede"],
                    width=2.5 if emphasised else 1.0,
                ),
                marker=dict(size=6),
                opacity=1.0 if emphasised else 0.45,
                name=feat if emphasised else "Other leading candidates",
                legendgroup=feat if emphasised else "other",
                showlegend=panel == 1 and (emphasised or rank == len(EMPHASIS_COLORS)),
                hovertext=feat,
            ),
            row=1,
            col=panel,
        )
    fig.add_hline(y=0, line=dict(color=COLORS["neutral"], width=1), row=1, col=panel)

fig.update_layout(
    template="ml4t",
    height=520,
    width=1120,
    title={
        "text": (
            "Agreement holds out to the horizon the strategy actually trades"
            "<br><sup>Leading candidates measured against the five- and ten-session unhedged "
            "returns and against the hold-to-expiry return, which averages one month; the "
            "three largest are named</sup>"
        )
    },
    margin=dict(l=70, r=30, t=120, b=60),
    legend=dict(orientation="h", y=-0.18, x=0, font=dict(size=9)),
)
fig.update_xaxes(title_text="Sessions held", tickvals=horizons)
fig.update_yaxes(title_text="Average daily rank IC", row=1, col=1)
fig.update_yaxes(title_text="Average divided by daily standard deviation", row=1, col=2)
show_plotly_with_alt(
    fig,
    "Two panels sharing a horizontal axis of sessions held, at 5, 10 and 21. On the left the "
    "average daily rank IC and on the right the same average divided by its daily standard "
    "deviation. ret_21d, ret_10d and ret_5d are drawn as named coloured lines and the other "
    "leading candidates as pale grey lines. All three named lines dip at 10 sessions and rise "
    "to their highest value at 21, ret_21d reaching about 0.035 on the left and about 0.28 on "
    "the right. The grey lines mostly sit below zero and do not share that shape, so agreement "
    "holds out to the horizon the strategy actually trades for the named three.",
)

# %% [markdown]
# ## 4. Is the relationship one a ranking strategy can use?
#
# A correlation can be produced by a handful of extreme values while the middle of the
# distribution says nothing. That matters here because the strategy does not act on the
# correlation, it acts on the ordering: it sells the straddles at one end of the ranking.
# So the question is whether 

Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT

Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.