Passer au contenu
Tous les documents de la bibliothèque

Sélection comparable de modèles et analyse inter-études de cas

Code Machine Learning for Trading

Résumé

Le document décrit un processus d’analyse commun pour comparer des configurations de modèles entre études de cas. Il sélectionne d’abord une configuration par étude, puis établit des synthèses par pli et entre études à partir de cette sélection, afin que les analyses suivantes reposent sur un choix cohérent. Les candidats doivent avoir des métriques finies, couvrir les plis attendus et compter autant de jours d’évaluation que le candidat ayant la meilleure couverture ; parmi eux, celui dont le coefficient d’information agrégé quotidien est le plus élevé est sélectionné.

Le document explique aussi comment déduire les identifiants de plis attendus à partir des candidats qui déclarent le nombre de plis prévu. Ces candidats doivent s’accorder sur les identifiants ; en cas de désaccord, une erreur est déclenchée au lieu de choisir silencieusement une géométrie. Un contrôle distinct de complétude compare les plis sélectionnés à la grille de référence de l’étude. Le document décrit également la comparaison des coefficients d’information sur des horodatages communs et la recherche des paires de labels de régression et de direction binaire déclarées. Il s’agit d’infrastructure de recherche, et non de preuves de performance de trading ; certaines parties de la source sont omises, de sorte que l’extrait disponible ne montre pas toutes les méthodes d’analyse ou de visualisation.

Idées clés

  • Sélectionnez une configuration par étude de cas, puis basez les comparaisons ultérieures sur cette sélection fixe.
  • Exigez une couverture d’évaluation comparable avant de classer les candidats selon leur coefficient d’information quotidien.
  • Déduisez les identifiants de plis des candidats couvrant tous les plis et signalez les géométries contradictoires.
  • Comparez les coefficients d’information uniquement sur les horodatages communs aux deux séries.
  • N’incluez que les paires de labels déclarées ayant des prédictions enregistrées et un label de direction binaire.

Étiquettes

Texte intégral
# insight_chapter.py


```py
"""Cross-case-study aggregation for chapter insight notebooks (Ch11–Ch15).

Wraps the per-case-study helpers in `model_analysis` with a thin "collect across
N case studies" layer plus the plotters the insight chapters draw.

The chapters ask one question of nine case studies: which configuration ranked
first, and how did it behave across folds? So the flow is always *select once,
then derive* — `collect_rank1_per_cs` picks the winning configuration per case
study, and everything else takes that frame rather than running its own
selection and risking a different answer.

Selection compares only what is comparable. A configuration evaluated on three
folds and one evaluated on six do not have comparable ICs, so candidates are
first restricted to those covering every declared fold and the same number of
days; only then does the highest daily-pooled IC win.

Usage::

    from case_studies.utils.insight_chapter import (
        collect_rank1_per_cs,
        collect_fold_ic_per_cs,
        plot_cross_cs_forest,
    )
    from case_studies.utils.analytics import CASE_STUDY_IDS

    rank1 = collect_rank1_per_cs(CASE_STUDY_IDS, family="linear")
    folds = collect_fold_ic_per_cs(rank1)
    plot_cross_cs_forest(rank1, family="linear",
                         title="Linear: rank-1 per case study (primary label)")
"""

from __future__ import annotations

import contextlib
import json
import math
import sqlite3
from collections.abc import Callable, Iterable
from functools import lru_cache

import polars as pl

# ml4t.diagnostic dlopens cudart; load torch first so its bundled CUDA
# runtime wins. Same pattern as case_studies/utils/model_analysis.py.
import torch  # noqa: F401

from case_studies.research.population import retired_prediction_hashes
from case_studies.utils.analytics import PRIMARY_LABELS, SHORT_NAMES
from case_studies.utils.booster_paths import booster_dir
from case_studies.utils.conformal import (
    sizing_conformal_lag,
    walk_forward_conformal_coverage,
)
from case_studies.utils.cv_window import (
    IntradayFoldBoundaryError,
    modeling_fold_boundaries,
)
from case_studies.utils.gbm_importance import top_features_by_gain
from case_studies.utils.registry.specs import declared_fold_count
from utils.paths import get_case_study_dir

LabelResolver = Callable[[str], str | None]


class RegistrySelectionError(ValueError):
    """Raised when registry candidates cannot support a comparable rank-one selection."""


class IncomparableFoldGeometryError(RegistrySelectionError):
    """Raised when two candidates each cover the declared fold count, but different folds.

    Distinct from "nothing reaches the declared count", which is a case study with no
    complete candidate and is a reason to leave it out of a table. Two disagreeing
    full-length geometries are a registry that cannot be ranked at all, so every caller
    surfaces it rather than skipping the case study.
    """


def compare_ic_on_shared_timestamps(
    left: pl.DataFrame,
    right: pl.DataFrame,
) -> dict[str, float | int]:
    """Average two IC series over their exact evaluation-timestamp intersection."""
    required = {"date", "ic"}
    for name, frame in (("left", left), ("right", right)):
        missing = required - set(frame.columns)
        if missing:
            raise RegistrySelectionError(f"{name} daily IC is missing {sorted(missing)}")

    def _daily(frame: pl.DataFrame, value_name: str) -> pl.DataFrame:
        return (
            frame.select("date", "ic")
            .drop_nulls()
            .with_columns(pl.col("date").cast(pl.Datetime("ns")))
            .filter(pl.col("ic").is_finite())
            .group_by("date")
            .agg(pl.col("ic").mean().alias(value_name))
        )

    shared = _daily(left, "left_ic").join(_daily(right, "right_ic"), on="date", how="inner")
    if shared.is_empty():
        raise RegistrySelectionError("IC series have no shared evaluation timestamps")
    return {
        "left_ic": float(shared["left_ic"].mean()),
        "right_ic": float(shared["right_ic"].mean()),
        "n_timestamps": shared.height,
    }


def _resolve_label(case_study: str, label_resolver: LabelResolver | None) -> str:
    if label_resolver is None:
        return PRIMARY_LABELS[case_study]
    out = label_resolver(case_study)
    return out if out is not None else PRIMARY_LABELS[case_study]


def select_rank1(
    metrics: pl.DataFrame,
    folds: pl.DataFrame,
    *,
    expected_fold_ids: Iterable[int],
) -> dict:
    """Select the best configuration among those that are comparable.

    A candidate is comparable when it reports every declared fold with a finite
    IC and covers the same number of days as the best-covered candidate. Ranking
    IC across candidates that saw different folds or different days would reward
    a short evaluation window rather than a better model.
    """
    metric_columns = {
        "prediction_hash",
        "training_hash",
        "config_name",
        "ic_mean_daily",
        "ic_n_days",
        "ic_se_hac",
        "ic_ci_lo",
        "ic_ci_hi",
        "ic_t_hac",
        "ic_p_hac",
        "ic_hac_lag",
    }
    fold_columns = {"prediction_hash", "fold_id", "ic"}
    missing_metrics = metric_columns - set(metrics.columns)
    missing_folds = fold_columns - set(folds.columns)
    if missing_metrics or missing_folds:
        raise RegistrySelectionError(
            f"missing selection columns: metrics={sorted(missing_metrics)}, "
            f"folds={sorted(missing_folds)}"
        )
    if metrics.is_empty():
        raise RegistrySelectionError("no registry candidates")

    expected = set(expected_fold_ids)
    if not expected:
        raise RegistrySelectionError("expected fold set is empty")

    eligible: list[dict] = []
    finite_columns = [
        "ic_mean_daily",
        "ic_n_days",
        "ic_se_hac",
        "ic_ci_lo",
        "ic_ci_hi",
        "ic_t_hac",
        "ic_p_hac",
        "ic_hac_lag",
    ]
    for row in metrics.iter_rows(named=True):
        prediction_hash = row["prediction_hash"]
        if any(
            row[column] is None or not math.isfinite(float(row[column]))
            for column in finite_columns
        ):
            continue
        if float(row["ic_n_days"]) <= 0:
            continue

        candidate_folds = folds.filter(pl.col("prediction_hash") == prediction_hash)
        fold_ids = candidate_folds["fold_id"].to_list()
        if len(fold_ids) != len(set(fold_ids)) or set(fold_ids) != expected:
            continue
        if any(value is None or not math.isfinite(float(value)) for value in candidate_folds["ic"]):
            continue
        eligible.append(row)

    if not eligible:
        raise RegistrySelectionError("no finite candidate covers every declared fold")
    full_days = max(float(row["ic_n_days"]) for row in eligible)
    comparable = [row for row in eligible if float(row["ic_n_days"]) == full_days]
    return max(comparable, key=lambda row: float(row["ic_mean_daily"]))


def resolve_expected_fold_ids(folds: pl.DataFrame, n_folds: int) -> tuple[int, ...]:
    """The fold ids a comparable candidate has to report, read off the candidates.

    This answers comparability within one (case study, family, label) group and nothing
    else. Whether that group covers the case study's whole fold grid is a different
    question, answered by `canonical_fold_ids` and carried on the selected row as
    ``n_folds_canonical`` and ``covers_fold_grid``.

    Neither live spec shape declares *which* folds a run used. Identity v3 nests a fold
    count under ``computation.expected_prediction_keys`` and the v2 shape carries
    ``n_folds`` at the top level; `declared_fold_count` reads both and there is no
    fold-id list in either. Every caller used to substitute ``range(n_folds)``, which is
    right only where the ids run contiguously from zero, and one live group breaks it:
    ``us_equities_panel/deep_learning/fwd_ret_5d`` declares four folds and all 30 of its
    prediction sets report 0, 5, 10 and 15, so every candidate failed the comparison and
    `select_rank1` raised on a case study it had selected from before those rows existed.

    **Those ids are canonical and the run is a subsample, not a renumbering.**
    ``us_equities_panel/12_dl_weekly`` sets ``MAX_FOLDS = 4`` and takes four of the case
    study's sixteen modelling folds evenly spaced, keeping each one's canonical id. So a
    count comparison is self-referential here: the run declares the size of the subsample
    it chose and passes its own test. That is why completeness is measured against the
    grid instead, and why this function does not decide it.

    The ids come from the candidates that reach the declared count: those reporting
    exactly ``n_folds`` distinct folds must agree on which ones, and the agreed set is
    what the gate compares against. A candidate reporting fewer is still refused - a
    shorter evaluation is an easier one. Two disagreeing full-length geometries in one
    group are not comparable to each other either, so that raises rather than letting the
    more numerous one define the standard.
    """
    if n_folds <= 0:
        raise RegistrySelectionError("n_folds is not declared")
    if folds.is_empty() or "fold_id" not in folds.columns:
        raise RegistrySelectionError("no fold rows to resolve the declared fold ids from")
    per_candidate = (
        folds.select("prediction_hash", "fold_id")
        .unique()
        .group_by("prediction_hash")
        .agg(pl.col("fold_id").sort())
    )
    full_length = {
        tuple(int(fold_id) for fold_id in ids)
        for ids in per_candidate["fold_id"].to_list()
        if len(ids) == n_folds
    }
    if not full_length:
        raise RegistrySelectionError(f"no candidate reports all {n_folds} declared folds")
    if len(full_length) > 1:
        raise IncomparableFoldGeometryError(
            f"candidates disagree on which {n_folds} folds they cover: {sorted(full_length)}"
        )
    return next(iter(full_length))


@lru_cache(maxsize=256)
def canonical_fold_ids(case_study: str, label: str) -> tuple[int, ...] | None:
    """The case study's whole modelling fold grid for one label, or ``None``.

    ``None`` means the grid is not derivable here, and there are exactly two ways to get
    it. The first is boundaries carrying a time of day, which today means the two intraday
    case studies: `modeling_fold_boundaries` runs its boundaries through `fold_boundary_date`,
    which refuses a timestamp carrying a time of day - "the spans that read it are daily,
    so truncating it would move the fold". That is 28 of the corpus's 105 (case study,
    family, label) groups, all of them `crypto_perps_funding` at eight-hourly and
    `nasdaq100_microstructure` at five, fifteen and sixty minutes. The second is the label
    surface not being on disk, where `modeling_fold_boundaries` returns ``None`` outright:
    `case_studies/*/labels` is gitignored and the reader bundle ships `run_log/` alone, so
    a reader who downloaded artifacts to run the insight chapters without training gets
    this for every label, not just the intraday ones.

    Both are the same answer to the caller - completeness was not measured, which is not
    the same as a group measured and found complete - and the selected rows keep that apart
    from a measurement by carrying ``n_folds_canonical`` as null rather than as a number.
    A consumer that prints an exclusion list has to say how many cells were never measured,
    or its "none" reads as "every cell was checked".

    A misconfiguration is not swallowed. `_derive_modeling_splits` raises `ValueError`
    deliberately for config drift - a label with a parquet but no declared buffer, or one
    whose parquet carries neither ``timestamp`` nor ``date`` - and `cv_window` states that
    as a loud-fail contract. Caught here it would arrive as a third meaning of ``None`` and
    the cell would sit in a census under the rule written for the intraday case studies with
    nothing printed.
    """
    try:
        boundaries = modeling_fold_boundaries(case_study, label)
    except IntradayFoldBoundaryError:
        return None
    if not boundaries:
        return None
    return tuple(sorted(int(boundary["fold"]) for boundary in boundaries))


def _fold_grid_columns(
    case_study: str, label: str, expected_fold_ids: tuple[int, ...]
) -> dict[str, int | bool | None]:
    """How much of the case study's fold grid the selected candidate actually scored.

    Carried on every selected row so a consumer that claims completeness can honour it.
    `13_dl_time_series/12_case_study_insights`'s horizon figure is the one that has to:
    it draws a line per case study across labels against a shared "average daily IC"
    axis, and ``us_equities_panel/fwd_ret_5d`` is a Friday-resampled panel scored on four
    of sixteen folds, so joining it to the daily ``fwd_ret_1d`` point would draw two
    different quantities as one moving with horizon.
    """
    canonical = canonical_fold_ids(case_study, label)
    return {
        "n_folds_scored": len(expected_fold_ids),
        "n_folds_canonical": len(canonical) if canonical is not None else None,
        "covers_fold_grid": (
            None if canonical is None else set(expected_fold_ids) == set(canonical)
        ),
    }


def _raw_primary_candidates(
    case_study: str, family: str, label: str
) -> tuple[pl.DataFrame, pl.DataFrame, int]:
    """Candidate rows for one (case study, family, label), retired generations removed.

    Two things about this query are load-bearing and were both absent until 2026-09-18.

    It selects the ``auc_*`` columns beside the ``ic_*`` ones. They sit in the same
    ``prediction_metrics`` row and the chapters print them: Chapter 11's Table 11.6 has a
    "Native AUC" column and Chapter 12's Table 12.4 a "Classifier -> AUC" column, and
    neither notebook could reproduce its own published table because the query named the
    seven IC columns and none of the five AUC ones.

    It drops the generations a refit has retired. A registry keeps every generation - the
    record of what was superseded is evidence - and supersession is recorded one layer up,
    in ``official_populations``, so a reader that joins ``training_runs`` to
    ``prediction_metrics`` directly ranks history along with the present. That is not a
    neutral error: a superseded generation is often the one that was refitted *because* it
    was short or narrow, and a shorter or narrower sample is an easier one, so the stale
    rows run toward the top of the ranking. Measured on ``cme_futures`` 2026-09-18, the
    retired deep-learning generation covers 27,326 of 38,262 declared (entity, session)
    pairs, 71.4%, and outranked the complete refit that replaced it.
    """
    db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
    if not db_path.exists():
        return pl.DataFrame(), pl.DataFrame(), 0
    uri = f"file:{db_path}?mode=ro"
    with sqlite3.connect(uri, uri=True) as db:
        db.row_factory = sqlite3.Row
        rows = db.execute(
            """
            SELECT p.prediction_hash, t.training_hash, t.family, t.config_name, t.label,
                   p.checkpoint_value, p.checkpoint_kind, t.git_commit, t.spec_json,
                   t.created_at AS training_created_at, pm.computed_at,
                   pm.ic_mean, pm.ic_std, pm.ic_mean_daily, pm.ic_n_days,
                   pm.ic_se_hac, pm.ic_ci_lo, pm.ic_ci_hi, pm.ic_t_hac,
                   pm.ic_p_hac, pm.ic_hac_lag,
                   pm.auc_mean_daily, pm.auc_se_hac, pm.auc_ci_lo, pm.auc_ci_hi,
                   pm.auc_n_days
            FROM training_runs t
            JOIN prediction_sets p ON p.training_hash = t.training_hash
            JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
            WHERE t.family = ? AND t.label = ? AND p.split = 'validation'
            """,
            (family, label),
        ).fetchall()
        fold_rows = db.execute(
            """
            SELECT p.prediction_hash, fm.fold_id, fm.ic
            FROM training_runs t
            JOIN prediction_sets p ON p.training_hash = t.training_hash
            JOIN fold_metrics fm ON fm.prediction_hash = p.prediction_hash
            WHERE t.family = ? AND t.label = ? AND p.split = 'validation'
            """,
            (family, label),
        ).fetchall()
        retired = retired_prediction_hashes(db)
    rows = [row for row in rows if row["prediction_hash"] not in retired]
    # The fold frame is restricted to the prediction sets the metrics frame holds, and
    # that is not the same filter as dropping the retired ones. The metrics query joins
    # `prediction_metrics` and the fold query does not, and the two tables are written by
    # separate calls - `register_prediction_set` then `register_fold_metrics` - so a set
    # interrupted between them, or registered with `metrics=None`, has fold rows and no
    # metrics row. It is unrankable, and it used to reach `resolve_expected_fold_ids`
    # anyway: a full-length but differently numbered geometry from one raised
    # `IncomparableFoldGeometryError` and stopped a whole chapter over a candidate
    # `select_rank1` could never have returned. Restricted here rather than at each call
    # site because the three collectors that take this frame resolve a geometry from it
    # and only one of them had the guard. `collect_checkpoint_fold_trajectories` is a
    # fourth resolver and does not take this frame - it queries the registry itself, and
    # carries the same two filters inline.
    rankable = {row["prediction_hash"] for row in rows}
    fold_rows = [
        row
        for row in fold_rows
        if row["prediction_hash"] not in retired and row["prediction_hash"] in rankable
    ]
    if not rows:
        return pl.DataFrame(), pl.DataFrame(), 0
    metrics = pl.DataFrame([dict(row) for row in rows], infer_schema_length=None)
    fold_metrics = pl.DataFrame([dict(row) for row in fold_rows], infer_schema_length=None)
    declared_folds = []
    for spec_json in metrics["spec_json"].drop_nulls().to_list():
        with contextlib.suppress(json.JSONDecodeError, TypeError, ValueError):
            value = declared_fold_count(json.loads(spec_json))
            if value > 0:
                declared_folds.append(value)
    return metrics, fold_metrics, max(declared_folds, default=0)


def collect_rank1_per_cs(
    case_studies: Iterable[str],
    family: str,
    label_resolver: LabelResolver | None = None,
) -> pl.DataFrame:
    """Rank-one configuration per case study, by daily-pooled IC.

    A case study with no registry rows for this family is left out of the
    result, so an absent case study shows up as a missing row rather than an
    exception. Every other collector takes the frame this returns.
    """
    rows = []
    for case_study in case_studies:
        label = _resolve_label(case_study, label_resolver)
        metrics, folds, n_folds = _raw_primary_candidates(case_study, family, label)
        if metrics.is_empty():
            continue
        if n_folds <= 0:
            raise RegistrySelectionError(f"{case_study}/{family}/{label}: n_folds is not declared")
        try:
            expected_fold_ids = resolve_expected_fold_ids(folds, n_folds)
            row = select_rank1(metrics, folds, expected_fold_ids=expected_fold_ids)
        except IncomparableFoldGeometryError as exc:
            raise IncomparableFoldGeometryError(f"{case_study}/{family}/{label}: {exc}") from exc
        except RegistrySelectionError as exc:
            raise RegistrySelectionError(f"{case_study}/{family}/{label}: {exc}") from exc
        row.update(
            case_study=case_study,
            short_name=SHORT_NAMES.get(case_study, case_study),
            **_fold_grid_columns(case_study, label, expected_fold_ids),
        )
        rows.append(row)
    return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()


def collect_fold_ic_per_cs(rank1: pl.DataFrame) -> pl.DataFrame:
    """Per-fold IC for each configuration `collect_rank1_per_cs` selected."""
    rows = []
    for selected in rank1.iter_rows(named=True):
        db_path = get_case_study_dir(selected["case_study"]) / "run_log" / "registry.db"
        uri = f"file:{db_path}?mode=ro"
        with sqlite3.connect(uri, uri=True) as db:
            db.row_factory = sqlite3.Row
            fold_rows = db.execute(
                """
                SELECT fold_id, ic, ic_std, n_entities, rmse, mae
                FROM fold_metrics WHERE prediction_hash = ? ORDER BY fold_id
                """,
                (selected["prediction_hash"],),
            ).fetchall()
        for fold in fold_rows:
            rows.append(
                {
                    "case_study": selected["case_study"],
                    "short_name": selected["short_name"],
                    "family": selected["family"],
                    "config_name": selected["config_name"],
                    "label": selected["label"],
                    "prediction_hash": selected["prediction_hash"],
                    **dict(fold),
                }
            )
    return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()


def collect_checkpoint_fold_trajectories(rank1: pl.DataFrame) -> pl.DataFrame:
    """Per-fold IC at every checkpoint of each selected training run."""
    rows = []
    for selected in rank1.iter_rows(named=True):
        spec = json.loads(selected["spec_json"])
        n_folds = declared_fold_count(spec)
        if n_folds <= 0:
            raise RegistrySelectionError(
                f"{selected['case_study']}/{selected['training_hash']}: n_folds is not declared"
            )
        db_path = get_case_study_dir(selected["case_study"]) / "run_log" / "registry.db"
        uri = f"file:{db_path}?mode=ro"
        with sqlite3.connect(uri, uri=True) as db:
            db.row_factory = sqlite3.Row
            # Joined to `prediction_metrics` and filtered for retired hashes for the
            # same reasons `_raw_primary_candidates` does both, and this is the fourth
            # site that resolves a fold geometry. A checkpoint with fold rows and no
            # metrics row is unrankable and would still set the standard here, and a
            # retired generation under this same `training_hash` would too. Chapter 13
            # calls this collector immediately after rank-one selection, so without the
            # join the chapter dies one cell past the one the restriction fixed, saying
            # the fold geometries disagree rather than that a registration is stray.
            # Measured 2026-09-19 across all nine canonical registries: no metric-less
            # validation prediction set and no duplicate (training_hash,
            # checkpoint_value) group, so nothing reaches it today.
            checkpoint_rows = db.execute(
                """
                SELECT p.prediction_hash, p.checkpoint_value, fm.fold_id, fm.ic
                FROM prediction_sets p
                JOIN fold_metrics fm ON fm.prediction_hash = p.prediction_hash
                JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
                WHERE p.training_hash = ? AND p.split = 'validation'
                ORDER BY p.checkpoint_value, fm.fold_id
                """,
                (selected["training_hash"],),
            ).fetchall()
            retired = retired_prediction_hashes(db)
        checkpoint_rows = [r for r in checkpoint_rows if r["prediction_hash"] not in retired]
        checkpoints = pl.DataFrame([dict(row) for row in checkpoint_rows], infer_schema_length=None)
        if checkpoints.is_empty():
            raise RegistrySelectionError(
                f"{selected['case_study']}/{selected['training_hash']}: no checkpoint folds"
            )
        try:
            expected_folds = set(resolve_expected_fold_ids(checkpoints, n_folds))
        except IncomparableFoldGeometryError as exc:
            raise IncomparableFoldGeometryError(
                f"{selected['case_study']}/{selected['training_hash']}: {exc}"
            ) from exc
        except RegistrySelectionError as exc:
            raise RegistrySelectionError(
                f"{selected['case_study']}/{selected['training_hash']}: {exc}"
            ) from exc
        for checkpoint in checkpoints["checkpoint_value"].unique().sort().to_list():
            current = checkpoints.filter(pl.col("checkpoint_value") == checkpoint)
            fold_ids = current["fold_id"].to_list()
            if len(fold_ids) != len(set(fold_ids)) or set(fold_ids) != expected_folds:
                raise RegistrySelectionError(
                    f"{selected['case_study']}/{selected['training_hash']}/{checkpoint}: "
                    "checkpoint fold coverage is incomplete or duplicate"
                )
            if any(value is None or not math.isfinite(float(value)) for value in current["ic"]):
                raise RegistrySelectionError(
                    f"{selected['case_study']}/{selected['training_hash']}/{checkpoint}: "
                    "checkpoint IC is non-finite"
                )
            for row in current.iter_rows(named=True):
                rows.append(
                    {
                        "case_study": selected["case_study"],
                        "short_name": selected["short_name"],
                        "config_name": selected["config_name"],
                        "training_hash": selected["training_hash"],
                        **row,
                    }
                )
    return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()


def conformal_coverage_for_selected_prediction(
    selected: dict,
    *,
    levels: tuple[float, ...] = (0.80, 0.90, 0.95),
    embargo_steps: int | None = None,
) -> pl.DataFrame:
    """Realised coverage of the sizing widths, for one exact selected prediction artifact.

    Measured on the estimator `conformal_weighted` allocates with - see
    :func:`~case_studies.utils.conformal.walk_forward_conformal_coverage`, which also states
    why the figure is a diagnostic of residual dispersion rather than a guarantee.

    ``embargo_steps`` defaults to :func:`~case_studies.utils.conformal.sizing_conformal_lag`
    for this row's case study and label.
    The label is taken from the selected row where it carries one and from the training spec
    otherwise - `us_equities_panel/15_model_analysis` attaches it to the returned frame rather
    than to the row it passes, and both know the same label.
    """
    required = {
        "case_study",
        "family",
        "config_name",
        "prediction_hash",
        "spec_json",
    }
    missing = required - set(selected)
    if missing:
        raise RegistrySelectionError(f"selected row missing conformal fields: {sorted(missing)}")

    spec = json.loads(selected["spec_json"])
    n_folds = declared_fold_count(spec)
    if n_folds < 2:
        raise RegistrySelectionError(
            f"{selected['case_study']}/{selected['prediction_hash']}: "
            "conformal coverage requires at least two declared folds"
        )

    prediction_path = (
        get_case_study_dir(selected["case_study"])
        / "run_log"
        / "predictions"
        / selected["prediction_hash"]
        / "predictions.parquet"
    )
    if not prediction_path.exists():
        raise RegistrySelectionError(f"missing prediction artifact: {prediction_path}")
    predictions = pl.read_parquet(prediction_path)
    required_columns = {"timestamp", "y_true", "y_score", "fold_id"}
    legacy_columns = {"actual": "y_true", "prediction": "y_score", "fold": "fold_id"}
    present_legacy = {
        legacy: canonical
        for legacy, canonical in legacy_columns.items()
        if legacy in predictions.columns and canonical not in predictions.columns
    }
    if present_legacy:
        predictions = predictions.rename(present_legacy)
    missing_columns = required_columns - set(predictions.columns)
    if missing_columns:
        raise RegistrySelectionError(
            f"{selected['case_study']}/{selected['prediction_hash']}: "
            f"prediction artifact missing {sorted(missing_columns)}"
        )
    usable = predictions.drop_nulls(required_columns)
    for column in ("y_true", "y_score"):
        usable = usable.filter(pl.col(column).cast(pl.Float64, strict=False).is_finite())
    # The artifact has to carry every fold the spec declared. It is compared by count
    # rather than against ``range(n_folds)`` because a run may number its folds any way
    # it likes and one live sweep does: ``us_equities_panel/deep_learning/fwd_ret_5d``
    # declares four folds and writes 0, 5, 10 and 15.
    #
    # The exact geometry is available - ``fold_metrics`` is keyed by ``prediction_hash``
    # and records the ids this run scored - and the count is used instead because nothing
    # downstream reads an id at all. `walk_forward_widths` selects ``fold_id`` and carries
    # it to the output untouched; precedence comes from the ``step`` index built off the
    # unique timestamps, so relabelling every fold changes no width. What would break the
    # calibration is a fold that is missing, and that is what the count catches.
    fold_ids = sorted(usable["fold_id"].unique().to_list())
    if len(fold_ids) != n_folds:
        raise RegistrySelectionError(
            f"{selected['case_study']}/{selected['prediction_hash']}: "
            f"expected {n_folds} declared folds, observed {fold_ids}"
        )
    if embargo_steps is None:
        label = selected.get("label") or spec.get("label")
        if not label:
            raise RegistrySelectionError(
                f"{selected['case_study']}/{selected['prediction_hash']}: neither the selected "
                "row nor its training spec names a label, so no reviewed conformal embargo "
                "can be resolved"
            )
        embargo_steps = sizing_conformal_lag(selected["case_study"], label)
    try:
        coverage_rows = walk_forward_conformal_coverage(
            predictions, levels=levels, embargo_steps=embargo_steps
        )
    except ValueError as error:
        raise RegistrySelectionError(
            f"{selected['case_study']}/{selected['prediction_hash']}: {error}"
        ) from error
    rows = [
        {
            "case_study": selected["case_study"],
            "family": selected["family"],
            "config_name": selected["config_name"],
            "prediction_hash": selected["prediction_hash"],
            **row,
        }
        for row in coverage_rows
    ]
    return pl.DataFrame(rows)


def collect_grid_per_cs(
    case_studies: Iterable[str],
    family: str,
    label_resolver: LabelResolver | None = None,
    config_parser: Callable[[str], dict] | None = None,
) -> pl.DataFrame:
    """Best checkpoint of every configuration, per case study.

    `collect_rank1_per_cs` keeps one row per case study; this keeps the whole
    grid, so a chapter can show the spread the winner came from.
    """
    rows = []
    for case_study in case_studies:
        label = _resolve_label(case_study, label_resolver)
        metrics, folds, n_folds = _raw_primary_candidates(case_study, family, label)
        if metrics.is_empty() or n_folds <= 0:
            continue
        # Resolved once over the whole group: a configuration whose candidates are all
        # short would otherwise define its own standard and rank against complete ones.
        try:
            expected_fold_ids = resolve_expected_fold_ids(folds, n_folds)
        except IncomparableFoldGeometryError as exc:
            raise IncomparableFoldGeometryError(f"{case_study}/{family}/{label}: {exc}") from exc
        except RegistrySelectionError:
            continue
        selected = []
        # `Series.unique` does not define an order, so this loop set `dl_grid`'s row
        # order and every group order downstream of it. Sorted, the frame is a function
        # of the registry.
        for config_name in sorted(metrics["config_name"].unique().to_list()):
            config_metrics = metrics.filter(pl.col("config_name") == config_name)
            try:
                selected.append(
                    select_rank1(config_metrics, folds, expected_fold_ids=expected_fold_ids)
                )
            except RegistrySelectionError:
                continue
        if not selected:
            continue
        full_days = max(float(row["ic_n_days"]) for row in selected)
        # Loop-invariant, and not free: it reads the label parquet's time column.
        grid_columns = _fold_grid_columns(case_study, label, expected_fold_ids)
        for selected_row in selected:
            if float(selected_row["ic_n_days"]) != full_days:
                continue
            row = {
                **selected_row,
                "case_study": case_study,
                "short_name": SHORT_NAMES.get(case_study, case_study),
                **grid_columns,
            }
            if config_parser is not None:
                row.update(config_parser(row["config_name"]))
            rows.append(row)
    return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()


def collect_multi_label_per_cs(
    case_studies: Iterable[str],
    family: str,
    labels: list[str] | Callable[[str], list[str]],
) -> pl.DataFrame:
    """Rank-one configuration per case study and label.

    `labels` is either one list applied to every case study, or a callable
    `case_study -> list[label]` where the label sets differ.
    """
    rows = []
    for case_study in case_studies:
        case_labels = labels(case_study) if callable(labels) else labels
        for label in case_labels:
            metrics, folds, n_folds = _raw_primary_candidates(case_study, family, label)
            if metrics.is_empty():
                continue
            if n_folds <= 0:
                raise RegistrySelectionError(
                    f"{case_study}/{family}/{label}: n_folds is not declared"
                )
            try:
                expected_fold_ids = resolve_expected_fold_ids(folds, n_folds)
                selected = select_rank1(metrics, folds, expected_fold_ids=expected_fold_ids)
            except IncomparableFoldGeometryError as exc:
                raise IncomparableFoldGeometryError(
                    f"{case_study}/{family}/{label}: {exc}"
                ) from exc
            except RegistrySelectionError as exc:
                raise RegistrySelectionError(f"{case_study}/{family}/{label}: {exc}") from exc
            selected.update(
                case_study=case_study,
                short_name=SHORT_NAMES.get(case_study, case_study),
                **_fold_grid_columns(case_study, label, expected_fold_ids),
            )
            rows.append(selected)
    return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()


def collect_gbm_checkpoint_trajectories(rank1: pl.DataFrame) -> pl.DataFrame:
    """Learning curve for each selected GBM training run."""
    from case_studies.utils.registry import get_training_dir

    frames = []
    for row in rank1.iter_rows(named=True):
        spec = json.loads(row["spec_json"])
        path = get_training_dir(row["case_study"], spec) / "learning_curves.parquet"
        if not path.exists():
            raise RegistrySelectionError(f"missing learning curve for {row['training_hash']}")
        curve = pl.read_parquet(path).filter(pl.col("config") == row["config_name"])
        if curve.is_empty() or "iteration" not in curve.columns:
            raise RegistrySelectionError(f"incomplete learning curve for {row['training_hash']}")
        trajectory = (
            curve.group_by("iteration")
            .agg(
                pl.col("ic_mean").mean().alias("ic_mean"),
                pl.col("ic_std").mean().alias("ic_std"),
            )
            .sort("iteration")
            .with_columns(
                pl.lit(row["case_study"]).alias("case_study"),
                pl.lit(row["short_name"]).alias("short_name"),
                pl.lit(row["config_name"]).alias("config_name"),
                pl.lit(row["training_hash"]).alias("training_hash"),
            )
        )
        frames.append(trajectory)
    return pl.concat(frames, how="diagonal_relaxed") if frames else pl.DataFrame()


def load_gbm_feature_importance(
    case_study: str,
    training_hash: str,
    config_name: str,
    *,
    top_n: int,
    num_iteration: int | None = None,
) -> pl.DataFrame:
    """Load booster importance from one selected GBM training identity.

    Training saves every boosting round, while a configuration is selected at one
    checkpoint along that trajectory. Pass that checkpoint as `num_iteration` and
    the importance is measured over the first that many trees, which is the model
    the selection actually chose; leave it None to read the whole booster.
    """
    import lightgbm as lgb

    case_dir = get_case_study_dir(case_study)
    run_booster_dir = booster_dir(case_dir, training_hash)
    if run_booster_dir is None:
        return pl.DataFrame()

    rows = []
    for booster_file in sorted(run_booster_dir.glob("*.txt")):
        fold_text = booster_file.stem.split("fold")[-1].lstrip("_")
        with contextlib.suppress(ValueError):
            fold_id = int(fold_text)
            model = lgb.Booster(model_file=str(booster_file))
            if num_iteration is not None and num_iteration < model.num_trees():
                model = lgb.Booster(model_str=model.model_to_string(num_iteration=num_iteration))
            rows.extend(
                {
                    "config_name": config_name,
                    "training_hash": training_hash,
                    "fold_id": fold_id,
                    "feature": feature,
                    "importance": float(importance),
                }
                for feature, importance in zip(
                    model.feature_name(),
                    model.feature_importance(importance_type="gain"),
                    strict=False,
                )
            )
    if not rows:
        return pl.DataFrame()
    result = pl.DataFrame(rows).with_columns(
        fold_max=pl.col("importance").max().over(["config_name", "fold_id"])
    )
    if result.filter(pl.col("fold_max") <= 0).height:
        raise RegistrySelectionError(f"{case_study}/{training_hash}: zero feature importance")
    result = result.with_columns(importance_norm=pl.col("importance") / pl.col("fold_max")).drop(
        "fold_max"
    )
    return result.filter(pl.col("feature").is_in(top_features_by_gain(result, top_n)))


def plot_cross_cs_forest(
    df: pl.DataFrame,
    family: str,
    title: str,
    *,
    sort_by: str = "ic_mean_daily",
    figsize: tuple[float, float] | None = None,
    sig_t: float = 2.0,
):
    """Forest of rank-1 daily-pooled IC ± HAC 95% CI per case study.

    Marker style encodes whether the HAC CI excludes zero:
      - filled circle: |t_hac| > sig_t (distinguishable from zero)
      - open circle:  |t_hac| ≤ sig_t (overlaps zero)

    Y-axis order is ascending by `sort_by`, so the largest value sits at top.

    Returns (fig, ax).
    """
    import matplotlib.pyplot as plt
    import numpy as np

    from utils.style import COLORS

    if df.is_empty():
        fig, ax = plt.subplots(figsize=(7, 2.5))
        ax.text(0.5, 0.5, f"No {family} runs in registry", ha="center", va="center")
        ax.set_title(title)
        ax.set_axis_off()
        return fig, ax

    d = df.sort(sort_by, descending=False, nulls_last=False).to_pandas()
    n = len(d)
    if figsize is None:
        figsize = (7.5, max(2.5, 0.45 * n + 1.2))
    fig, ax = plt.subplots(figsize=figsize, layout="tight")
    y = np.arange(n)
    ic = d["ic_mean_daily"].to_numpy()
    lo = d["ic_ci_lo"].to_numpy()
    hi = d["ic_ci_hi"].to_numpy()
    t_hac = d["ic_t_hac"].to_numpy() if "ic_t_hac" in d.columns else np.full(n, np.nan)

    ax.errorbar(
        ic,
        y,
        xerr=[ic - lo, hi - ic],
        fmt="none",
        color=COLORS["neutral"],
        lw=1.0,
        capsize=3,
    )
    t_arr = np.asarray(t_hac, dtype=float)
    sig = (np.abs(t_arr) > sig_t) & ~np.isnan(t_arr)
    if sig.any():
        ax.scatter(
            ic[sig],
            y[sig],
            s=60,
            marker="o",
            facecolor=COLORS["blue"],
            edgecolor=COLORS["blue"],
            zorder=3,
            label=f"|t_hac| > {sig_t:g}",
        )
    if (~sig).any():
        ax.scatter(
            ic[~sig],
            y[~sig],
            s=60,
            marker="o",
            facecolor="white",
            edgecolor=COLORS["blue"],
            zorder=3,
            label=f"|t_hac| ≤ {sig_t:g}",
        )

    ax.axvline(0, color=COLORS["neutral"], lw=0.7, linestyle="--")
    ax.set_yticks(y)
    ax.set_yticklabels(d["short_name"].tolist())
    ax.set_xlabel("Daily-pooled IC (HAC 95% CI)")
    ax.set_title(title)
    ax.legend(loc="lower right", fontsize=8, frameon=False)
    if fig.get_layout_engine() is None:
        fig.tight_layout()
    return fig, ax


def plot_per_fold_violin(
    fold_df: pl.DataFrame,
    order: list[str],
    *,
    title: str,
    figsize: tuple[float, float] | None = None,
    jitter_color: str = "#3B82F6",
):
    """Box-plus-scatter of per-fold IC across case studies.

    `fold_df` must carry columns `short_name` and `ic`. `order` is the CS
    display order (left → right). CSs absent from `fold_df` are skipped.
    """
    import matplotlib.pyplot as plt
    import numpy as np

    present = [c for c in order if c in fold_df["short_name"].unique().to_list()]
    if not present:
        fig, ax = plt.subplots(figsize=(7, 2.5))
        ax.text(0.5, 0.5, "No fold IC data", ha="center", va="center")
        ax.set_axis_off()
        return fig, ax

    if figsize is None:
        figsize = (max(6.5, 1.0 * len(present) + 2), 4.5)
    fig, ax = plt.subplots(figsize=figsize, layout="tight")
    data = [fold_df.filter(pl.col("short_name") == cs)["ic"].to_numpy() for cs in present]
    positions = np.arange(len(present))
    ax.boxplot(data, positions=positions, widths=0.55, showfliers=True)
    for i, arr in enumerate(data):
        if len(arr):
            ax.scatter(np.full(len(arr), i), arr, alpha=0.5, s=14, color=jitter_color)
    ax.axhline(0, color="gray", linewidth=0.7, linestyle="--")
    ax.set_xticks(positions)
    ax.set_xticklabels(present, rotation=30, ha="right")
    ax.set_ylabel("Per-fold Spearman IC")
    ax.set_title(title)
    if fig.get_layout_engine() is None:
        fig.tight_layout()
    return fig, ax


def _rank1_full_coverage_hash(case_study: str, label: str) -> str | None:
    """Validation prediction_hash of the highest daily-IC linear config with NO dropped fold.

    Selection is coverage-aware: a configuration whose path zeroes out on some
    fold leaves that fold's IC NULL and pools its daily IC over fewer days,
    which can make a naive ``MAX(ic_mean_daily)`` crown a config scored on a
    non-comparable subset. We therefore rank only among configs whose
    fold_metrics carry no NULL IC. Returns None if the registry has no
    full-coverage linear row for that label.
    """
    db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
    if not db_path.exists():
        return None
    with sqlite3.connect(db_path) as db:
        rows = db.execute(
            """
            SELECT p.prediction_hash, pm.ic_mean_daily, pm.ic_n_days,
                   (SELECT COUNT(*) FROM fold_metrics fm
                    WHERE fm.prediction_hash = p.prediction_hash AND fm.ic IS NULL) AS n_null
            FROM training_runs t
            JOIN prediction_sets p ON p.training_hash = t.training_hash AND p.split = 'validation'
            JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
            WHERE t.family = 'linear' AND t.label = ? AND pm.ic_mean_daily IS NOT NULL
            """,
            (label,),
        ).fetchall()
    full = [r for r in rows if r[2] is not None and r[3] == 0]
    if not full:
        return None
    max_n_days = max(r[2] for r in full)
    full = [r for r in full if r[2] == max_n_days]
    full.sort(key=lambda r: -r[1])
    return full[0][0]


def plot_rolling_daily_ic(
    case_studies: Iterable[str],
    *,
    window: int = 63,
    label_resolver: LabelResolver | None = None,
    common_window: bool = True,
    title: str = "Persistence of linear ranking signal (rolling daily IC)",
    figsize: tuple[float, float] = (10, 4),
    colors: list[str] | None = None,
    selected_prediction_hashes: dict[str, str] | None = None,
):
    """Rolling-mean daily-IC persistence chart for the rank-1 linear fit per case study.

    For each case study, take the coverage-aware rank-1 linear configuration
    (highest daily IC with no dropped fold), load its per-day IC series from
    ``daily_metrics.parquet``, average across overlapping folds per calendar
    day, and plot a ``window``-day rolling mean (63 ≈ three trading months).

    Case studies cover different validation windows, so they cannot share a
    calendar axis unless their periods overlap. With ``common_window=True``
    the series are clipped to the intersection of all input case studies'
    spans — pass only case studies whose windows overlap (e.g. ``etfs`` and
    ``fx_pairs`` over 2016–2023). Case studies without a daily-IC series are
    skipped.

    Returns (fig, ax).
    """
    import matplotlib.pyplot as plt

    from case_studies.utils.model_analysis import load_daily_metrics_series

    if colors is None:
        from utils.style import COLORS

        colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"], COLORS["slate"]]

    series: dict[str, pl.DataFrame] = {}
    for cs in case_studies:
        label = _resolve_label(cs, label_resolver)
        h = (
            selected_prediction_hashes.get(cs)
            if selected_prediction_hashes is not None
            else _rank1_full_coverage_hash(cs, label)
        )
        if h is None:
            continue
        d = load_daily_metrics_series(cs, h)
        if d.is_empty() or "date" not in d.columns:
            continue
        roll = (
            d.group_by("date")
            .agg(pl.col("ic").mean())
            .sort("date")
            .with_columns(
                pl.col("ic").rolling_mean(window_size=window, min_periods=window).alias("roll")
            )
            .drop_nulls("roll")
        )
        if not roll.is_empty():
            series[cs] = roll

    fig, ax = plt.subplots(figsize=figsize, layout="tight")
    if not series:
        ax.text(0.5, 0.5, "No daily-IC series available", ha="center", va="center")
        ax.set_axis_off()
        return fig, ax

    lo = max(s["date"].min() for s in series.values()) if common_window else None
    hi = min(s["date"].max() for s in series.values()) if common_window else None
    for idx, (cs, roll) in enumerate(series.items()):
        sub = roll
        if common_window:
            sub = roll.filter((pl.col("date") >= lo) & (pl.col("date") <= hi))
        ax.plot(
            sub["date"].to_numpy(),
            sub["roll"].to_numpy(),
            color=colors[idx % len(colors)],
            linewidth=1.6,
            label=SHORT_NAMES.get(cs, cs),
        )
    from utils.style import COLORS

    ax.axhline(0, color=COLORS["neutral"], linewidth=0.8, linestyle="--")
    ax.set_ylabel(f"{window}-day rolling daily IC")
    ax.set_xlabel("Validation date")
    ax.set_title(title)
    ax.legend(loc="upper right", frameon=False)
    if fig.get_layout_engine() is None:
        fig.tight_layout()
    return fig, ax


# Horizon → trading-day mapping shared by all chapter insight notebooks.
HORIZON_DAYS: dict[str, float] = {
    "fwd_ret_5m": 5 / (6.5 * 60),
    "fwd_ret_15m": 15 / (6.5 * 60),
    "fwd_ret_60m": 60 / (6.5 * 60),
    "fwd_ret_8h": 1.0 / 3,
    "fwd_ret_24h": 1.0,
    "fwd_ret_1d": 1.0,
    "fwd_ret_5d": 5.0,
    "fwd_ret_10d": 10.0,
    "fwd_ret_21d": 21.0,
    "fwd_ret_1m": 21.0,
    "fwd_ret_3m": 63.0,
    "fwd_ret_1m_win": 21.0,
    "fwd_ret_risk_adj_5d": 5.0,
}


def plot_multi_label_horizon(
    horizon_df: pl.DataFrame,
    *,
    title: str,
    min_labels_per_cs: int = 2,
    figsize: tuple[float, float] = (10, 5),
    palette: list[str] | None = None,
):
    """Faceted horizon plot: daily-pooled IC vs trading-day horizon per CS.

    `horizon_df` must carry `short_name`, `label`, `ic_mean_daily`, `ic_ci_lo`,
    `ic_ci_hi`. CSs with fewer than `min_labels_per_cs` mapped horizons are
    omitted from the figure (a coverage fact, not a defect).
    """
    import matplotlib.pyplot as plt

    plot_df = horizon_df.with_columns(
        horizon_days=pl.col("label").replace_strict(HORIZON_DAYS, default=None).cast(pl.Float64),
    ).filter(pl.col("horizon_days").is_not_null())
    multi_cs = (
        plot_df.group_by("short_name")
        .len()
        .filter(pl.col("len") >= min_labels_per_cs)["short_name"]
        .to_list()
    )
    plot_df = plot_df.filter(pl.col("short_name").is_in(multi_cs))
    if plot_df.is_empty():
        fig, ax = plt.subplots(figsize=(7, 2.5))
        ax.text(0.5, 0.5, "No multi-horizon coverage", ha="center", va="center")
        ax.set_axis_off()
        return fig, ax

    fig, ax = plt.subplots(figsize=figsize, layout="tight")
    if palette is None:
        from utils.style import COLORS

        palette = [
            COLORS["blue"],
            COLORS["amber"],
            COLORS["copper"],
            COLORS["positive"],
            COLORS["negative"],
            COLORS["neutral"],
            COLORS["slate"],
            COLORS["amber_light"],
        ]
    cs_sorted = sorted(plot_df["short_name"].unique().to_list())
    markers = ["o", "s", "D", "^", "v", "P", "X", "*"]
    linestyles = ["-", "--", "-.", ":", "-", "--", "-.", ":"]
    for idx, cs in enumerate(cs_sorted):
        # Sorted on (horizon_days, label). HORIZON_DAYS maps several labels onto one
        # value - fwd_ret_1d and fwd_ret_24h onto 1.0, fwd_ret_5d and
        # fwd_ret_risk_adj_5d onto 5.0, fwd_ret_21d, fwd_ret_1m and fwd_ret_1m_win onto
        # 21.0 - so a case study carrying both members of a pair has two points at one x
        # and the tie order decides which the line reaches first. The tie is live, not
        # hypothetical: measured 2026-09-19, SP500 Eq+Opt has fwd_ret_5d and
        # fwd_ret_risk_adj_5d in the deep_learning census, and gbm, linear and tabular_dl
        # each add US Firms at fwd_ret_1m against fwd_ret_1m_win. Sorting on the horizon
        # alone left the drawn order a property of the frame; it happens to match today,
        # so pinning it moves no committed figure.
        sub = plot_df.filter(pl.col("short_name") == cs).sort(["horizon_days", "label"])
        if sub.height < 2:
            continue
        x = sub["horizon_days"].to_numpy()
        ic = sub["ic_mean_daily"].to_numpy()
        lo = sub["ic_ci_lo"].to_numpy()
        hi = sub["ic_ci_hi"].to_numpy()
        color = palette[idx % len(palette)]
        ax.fill_between(x, lo, hi, color=color, alpha=0.12)
        ax.plot(
            x,
            ic,
            marker=markers[idx % len(markers)],
            linestyle=linestyles[idx % len(linestyles)],
            color=color,
            label=cs,
            linewidth=1.6,
            markersize=6,
            alpha=0.9,
        )
    ax.set_xscale("log")
    ax.set_xlabel("Horizon (trading days, log scale)")
    ax.set_ylabel("Daily-pooled IC (HAC 95 % CI band)")
    ax.axhline(0, color="gray", linewidth=0.7, linestyle="--")
    ax.set_title(title)
    ax.legend(loc="best", frameon=False, fontsize=8, ncol=2)
    if fig.get_layout_engine() is None:
        fig.tight_layout()
    return fig, ax


def parse_gbm_config(config: str) -> dict:
    """Decode a GBM `config_name` into profile / loss / leaves / objective_kind.

    Examples
    --------
    >>> parse_gbm_config("leaves_31_huber")
    {"profile": "leaves_31", "loss": "huber", "leaves": 31, "objective_kind": "regression"}
    >>> parse_gbm_config("default_binary")
    {"profile": "default", "loss": "binary", "leaves": None, "objective_kind": "classification"}
    """
    out = {"profile": config, "loss": "unknown", "leaves": None, "objective_kind": "regression"}
    parts = config.rsplit("_", 1)
    if len(parts) == 2 and parts[1] in ("mse", "mae", "huber"):
        out["profile"], out["loss"] = parts
    elif config.endswith("_binary"):
        out["loss"] = "binary"
        out["objective_kind"] = "classification"
        out["profile"] = config.removesuffix("_binary")
    if "leaves_" in out["profile"]:
        with contextlib.suppress(ValueError, IndexError):
            out["leaves"] = int(out["profile"].split("_")[1])
    return out


SYMMETRY_TABLE_SCHEMA: dict[str, pl.DataType] = {
    "short_name": pl.Utf8,
    "reg_label": pl.Utf8,
    "dir_label": pl.Utf8,
    "cls_config": pl.Utf8,
    "cls_score_ic": pl.Float64,
    "cls_score_ic_lo": pl.Float64,
    "cls_score_ic_hi": pl.Float64,
    "cls_score_ic_t": pl.Float64,
    "cls_score_auc": pl.Float64,
    "cls_score_auc_lo": pl.Float64,
    "cls_score_auc_hi": pl.Float64,
    "reg_config": pl.Utf8,
    "reg_score_auc": pl.Float64,
    "n_b": pl.Int64,
    "n_b_days": pl.Int64,
}
"""The columns of the Chapter 11 and 12 symmetry tables, declared once.

A full schema and not ``schema_overrides``: `discover_symmetry_pairs` can legitimately
return nothing - every declared pair skipped, or a registry with no classification runs
for this family - and a frame built from an empty list with partial overrides has no
string columns at all, so selecting ``short_name`` raises ``ColumnNotFoundError`` and the
skip reasons, which on that path are the entire output, never reach the reader.

Shared rather than written out in each notebook because two copies of a column list drift,
and because a test can only pin the contract if there is one thing to pin. Chapter 11
carries ``n_b_days`` and Chapter 12 does not; both tolerate the extra column in an empty
frame, which is cheaper than two schemas.
"""


# The pairing of a regression label with the binary direction label it is scored against
# used to be a hand-written literal in each chapter that draws the cross-evaluation. It
# shrank without saying so: `us_firm_characteristics: [("fwd_ret_1m", "fwd_class_1m")]`
# was dropped from Chapter 12's copy by the 2026-07-31 chapter-tree restore and kept in
# Chapter 11's, so one chapter published four rows and the other three while both
# reported a full count of their own literal. A literal cannot report what is missing
# from it, so the pairs are discovered and the skips are named.


def _binary_label_domain(case_study: str, label: str) -> set[int] | None:
    """Distinct values of one label surface, or None when the file is absent."""
    path = get_case_study_dir(case_study) / "labels" / f"{label}.parquet"
    if not path.is_file():
        return None
    values = pl.scan_parquet(path).select(pl.col(label)).unique().collect().to_series().drop_nulls()
    return {int(value) for value in values.to_list()}


def _declared_classification_pairs(case_study: str) -> dict[str, str]:
    """``{direction_label: regression_label}`` as ``setup.yaml`` declares it.

    Read from ``labels.classification_eval_label`` rather than inferred from the label
    names. Two of the nine case studies would be got wrong by a naming rule: crypto maps
    both ``fwd_dir_8h`` and ``fwd_dir_8h_3c`` onto ``fwd_ret_8h``, so a rule that builds
    the direction name from the regression suffix never sees the ternary one and cannot
    report skipping it, which is the failure this whole function exists to avoid.
    """
    import yaml

    setup_path = get_case_study_dir(case_study) / "config" / "setup.yaml"
    if not setup_path.is_file():
        return {}
    with setup_path.open() as handle:
        setup = yaml.safe_load(handle)
    labels = (setup or {}).get("labels") or {}
    declared = labels.get("classification_eval_label") or {}
    return {str(k): str(v) for k, v in declared.items()}


def discover_symmetry_pairs(
    case_studies: Iterable[str], family: str
) -> tuple[dict[str, list[tuple[str, str]]], list[str]]:
    """Regression and binary-direction label pairs, per case study, from the declaration.

    A pair qualifies when ``setup.yaml`` declares the direction label's
    ``classification_eval_label``, this family has registered validation predictions for
    both labels, and the direction label's own surface is binary. A ternary direction
    label would need a multi-class AUC and is out of scope for this comparison, so it is
    skipped by measuring its domain rather than by being left out of a list.

    Returns ``(pairs, skipped)``. Every declared candidate that does not qualify appears
    in ``skipped`` with its reason, so a comparison that covers fewer case studies than
    the corpus says which ones and why instead of reporting a full count of itself.
    """
    pairs: dict[str, list[tuple[str, str]]] = {}
    skipped: list[str] = []
    for case_study in case_studies:
        declared = _declared_classification_pairs(case_study)
        if not declared:
            continue
        db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
        if not db_path.is_file():
            continue
        with sqlite3.connect(f"file:{db_path}?mode=ro", uri=True) as db:
            registered = {
                row[0]
                for row in db.execute(
                    "SELECT DISTINCT t.label FROM training_runs t "
                    "JOIN prediction_sets p ON p.training_hash = t.training_hash "
                    "WHERE t.family = ? AND p.split = 'validation'",
                    (family,),
                )
            }
        found: list[tuple[str, str]] = []
        for direction, regression in sorted(declared.items()):
            if direction not in registered or regression not in registered:
                # Not a defect and not worth a line: this family simply did not run one
                # of the two labels here. Only a declared pair the family DID run, and
                # that still does not qualify, is a skip worth naming.
                continue
            domain = _binary_label_domain(case_study, direction)
            if domain is None:
                skipped.append(f"{case_study}/{direction}: no label surface on disk")
            elif not domain.issubset({0, 1}):
                skipped.append(
                    f"{case_study}/{direction}: domain {sorted(domain)} is not binary, "
                    "a multi-class AUC is out of scope here"
                )
            else:
                found.append((regression, direction))
        if found:
            pairs[case_study] = found
    return pairs, skipped

```

Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT

Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.