Vergleichbare Modellauswahl und Fallstudienanalyse
Zusammenfassung
Das Dokument beschreibt einen gemeinsamen Analyseablauf zum Vergleich von Modellkonfigurationen über Fallstudien hinweg. Zunächst wählt es eine Konfiguration je Fallstudie aus und erstellt anschließend Fold- und fallstudienübergreifende Zusammenfassungen auf Grundlage dieser Auswahl, damit spätere Analysen eine konsistente Wahl verwenden. Kandidaten müssen endliche Kennzahlen aufweisen, die erwarteten Folds abdecken und dieselbe Anzahl an Auswertungstagen wie der am besten abgedeckte Kandidat haben. Aus diesen Kandidaten wird der höchste tägliche gepoolte Informationskoeffizient ausgewählt.
Das Dokument erläutert auch, wie erwartete Fold-Kennungen aus Kandidaten abgeleitet werden, die die angegebene Anzahl von Folds melden. Diese Kandidaten müssen sich auf die Kennungen einigen; bei Abweichungen wird ein Fehler ausgelöst, statt stillschweigend eine Fold-Geometrie auszuwählen. Eine separate Vollständigkeitsprüfung vergleicht die ausgewählten Folds mit dem kanonischen Raster der Fallstudie. Außerdem beschreibt das Dokument den Vergleich von Informationskoeffizienten für gemeinsame Zeitstempel und das Auffinden angegebener Paare aus Regressions- und binären Richtungslabels. Dies ist Forschungsinfrastruktur, kein Beleg für Trading-Performance. Da Teile der Quelle ausgelassen sind, zeigt der verfügbare Auszug nicht jede Analyse- oder Plotmethode.
Kernaussagen
- Wählen Sie einmal je Fallstudie eine Konfiguration aus und leiten Sie nachgelagerte Vergleiche aus dieser festen Auswahl ab.
- Verlangen Sie eine vergleichbare Auswertungsabdeckung, bevor Sie Kandidaten nach täglichem Informationskoeffizienten ordnen.
- Leiten Sie Fold-Kennungen aus Kandidaten mit vollständiger Fold-Anzahl ab und kennzeichnen Sie widersprüchliche Fold-Geometrien.
- Vergleichen Sie Informationskoeffizienten nur für Zeitstempel, die beide Reihen gemeinsam haben.
- Beziehen Sie nur angegebene Label-Paare mit registrierten Prognosen und einem binären Richtungslabel ein.
Schlagwörter
Volltext
# insight_chapter.py
```py
"""Cross-case-study aggregation for chapter insight notebooks (Ch11–Ch15).
Wraps the per-case-study helpers in `model_analysis` with a thin "collect across
N case studies" layer plus the plotters the insight chapters draw.
The chapters ask one question of nine case studies: which configuration ranked
first, and how did it behave across folds? So the flow is always *select once,
then derive* — `collect_rank1_per_cs` picks the winning configuration per case
study, and everything else takes that frame rather than running its own
selection and risking a different answer.
Selection compares only what is comparable. A configuration evaluated on three
folds and one evaluated on six do not have comparable ICs, so candidates are
first restricted to those covering every declared fold and the same number of
days; only then does the highest daily-pooled IC win.
Usage::
from case_studies.utils.insight_chapter import (
collect_rank1_per_cs,
collect_fold_ic_per_cs,
plot_cross_cs_forest,
)
from case_studies.utils.analytics import CASE_STUDY_IDS
rank1 = collect_rank1_per_cs(CASE_STUDY_IDS, family="linear")
folds = collect_fold_ic_per_cs(rank1)
plot_cross_cs_forest(rank1, family="linear",
title="Linear: rank-1 per case study (primary label)")
"""
from __future__ import annotations
import contextlib
import json
import math
import sqlite3
from collections.abc import Callable, Iterable
from functools import lru_cache
import polars as pl
# ml4t.diagnostic dlopens cudart; load torch first so its bundled CUDA
# runtime wins. Same pattern as case_studies/utils/model_analysis.py.
import torch # noqa: F401
from case_studies.research.population import retired_prediction_hashes
from case_studies.utils.analytics import PRIMARY_LABELS, SHORT_NAMES
from case_studies.utils.booster_paths import booster_dir
from case_studies.utils.conformal import (
sizing_conformal_lag,
walk_forward_conformal_coverage,
)
from case_studies.utils.cv_window import (
IntradayFoldBoundaryError,
modeling_fold_boundaries,
)
from case_studies.utils.gbm_importance import top_features_by_gain
from case_studies.utils.registry.specs import declared_fold_count
from utils.paths import get_case_study_dir
LabelResolver = Callable[[str], str | None]
class RegistrySelectionError(ValueError):
"""Raised when registry candidates cannot support a comparable rank-one selection."""
class IncomparableFoldGeometryError(RegistrySelectionError):
"""Raised when two candidates each cover the declared fold count, but different folds.
Distinct from "nothing reaches the declared count", which is a case study with no
complete candidate and is a reason to leave it out of a table. Two disagreeing
full-length geometries are a registry that cannot be ranked at all, so every caller
surfaces it rather than skipping the case study.
"""
def compare_ic_on_shared_timestamps(
left: pl.DataFrame,
right: pl.DataFrame,
) -> dict[str, float | int]:
"""Average two IC series over their exact evaluation-timestamp intersection."""
required = {"date", "ic"}
for name, frame in (("left", left), ("right", right)):
missing = required - set(frame.columns)
if missing:
raise RegistrySelectionError(f"{name} daily IC is missing {sorted(missing)}")
def _daily(frame: pl.DataFrame, value_name: str) -> pl.DataFrame:
return (
frame.select("date", "ic")
.drop_nulls()
.with_columns(pl.col("date").cast(pl.Datetime("ns")))
.filter(pl.col("ic").is_finite())
.group_by("date")
.agg(pl.col("ic").mean().alias(value_name))
)
shared = _daily(left, "left_ic").join(_daily(right, "right_ic"), on="date", how="inner")
if shared.is_empty():
raise RegistrySelectionError("IC series have no shared evaluation timestamps")
return {
"left_ic": float(shared["left_ic"].mean()),
"right_ic": float(shared["right_ic"].mean()),
"n_timestamps": shared.height,
}
def _resolve_label(case_study: str, label_resolver: LabelResolver | None) -> str:
if label_resolver is None:
return PRIMARY_LABELS[case_study]
out = label_resolver(case_study)
return out if out is not None else PRIMARY_LABELS[case_study]
def select_rank1(
metrics: pl.DataFrame,
folds: pl.DataFrame,
*,
expected_fold_ids: Iterable[int],
) -> dict:
"""Select the best configuration among those that are comparable.
A candidate is comparable when it reports every declared fold with a finite
IC and covers the same number of days as the best-covered candidate. Ranking
IC across candidates that saw different folds or different days would reward
a short evaluation window rather than a better model.
"""
metric_columns = {
"prediction_hash",
"training_hash",
"config_name",
"ic_mean_daily",
"ic_n_days",
"ic_se_hac",
"ic_ci_lo",
"ic_ci_hi",
"ic_t_hac",
"ic_p_hac",
"ic_hac_lag",
}
fold_columns = {"prediction_hash", "fold_id", "ic"}
missing_metrics = metric_columns - set(metrics.columns)
missing_folds = fold_columns - set(folds.columns)
if missing_metrics or missing_folds:
raise RegistrySelectionError(
f"missing selection columns: metrics={sorted(missing_metrics)}, "
f"folds={sorted(missing_folds)}"
)
if metrics.is_empty():
raise RegistrySelectionError("no registry candidates")
expected = set(expected_fold_ids)
if not expected:
raise RegistrySelectionError("expected fold set is empty")
eligible: list[dict] = []
finite_columns = [
"ic_mean_daily",
"ic_n_days",
"ic_se_hac",
"ic_ci_lo",
"ic_ci_hi",
"ic_t_hac",
"ic_p_hac",
"ic_hac_lag",
]
for row in metrics.iter_rows(named=True):
prediction_hash = row["prediction_hash"]
if any(
row[column] is None or not math.isfinite(float(row[column]))
for column in finite_columns
):
continue
if float(row["ic_n_days"]) <= 0:
continue
candidate_folds = folds.filter(pl.col("prediction_hash") == prediction_hash)
fold_ids = candidate_folds["fold_id"].to_list()
if len(fold_ids) != len(set(fold_ids)) or set(fold_ids) != expected:
continue
if any(value is None or not math.isfinite(float(value)) for value in candidate_folds["ic"]):
continue
eligible.append(row)
if not eligible:
raise RegistrySelectionError("no finite candidate covers every declared fold")
full_days = max(float(row["ic_n_days"]) for row in eligible)
comparable = [row for row in eligible if float(row["ic_n_days"]) == full_days]
return max(comparable, key=lambda row: float(row["ic_mean_daily"]))
def resolve_expected_fold_ids(folds: pl.DataFrame, n_folds: int) -> tuple[int, ...]:
"""The fold ids a comparable candidate has to report, read off the candidates.
This answers comparability within one (case study, family, label) group and nothing
else. Whether that group covers the case study's whole fold grid is a different
question, answered by `canonical_fold_ids` and carried on the selected row as
``n_folds_canonical`` and ``covers_fold_grid``.
Neither live spec shape declares *which* folds a run used. Identity v3 nests a fold
count under ``computation.expected_prediction_keys`` and the v2 shape carries
``n_folds`` at the top level; `declared_fold_count` reads both and there is no
fold-id list in either. Every caller used to substitute ``range(n_folds)``, which is
right only where the ids run contiguously from zero, and one live group breaks it:
``us_equities_panel/deep_learning/fwd_ret_5d`` declares four folds and all 30 of its
prediction sets report 0, 5, 10 and 15, so every candidate failed the comparison and
`select_rank1` raised on a case study it had selected from before those rows existed.
**Those ids are canonical and the run is a subsample, not a renumbering.**
``us_equities_panel/12_dl_weekly`` sets ``MAX_FOLDS = 4`` and takes four of the case
study's sixteen modelling folds evenly spaced, keeping each one's canonical id. So a
count comparison is self-referential here: the run declares the size of the subsample
it chose and passes its own test. That is why completeness is measured against the
grid instead, and why this function does not decide it.
The ids come from the candidates that reach the declared count: those reporting
exactly ``n_folds`` distinct folds must agree on which ones, and the agreed set is
what the gate compares against. A candidate reporting fewer is still refused - a
shorter evaluation is an easier one. Two disagreeing full-length geometries in one
group are not comparable to each other either, so that raises rather than letting the
more numerous one define the standard.
"""
if n_folds <= 0:
raise RegistrySelectionError("n_folds is not declared")
if folds.is_empty() or "fold_id" not in folds.columns:
raise RegistrySelectionError("no fold rows to resolve the declared fold ids from")
per_candidate = (
folds.select("prediction_hash", "fold_id")
.unique()
.group_by("prediction_hash")
.agg(pl.col("fold_id").sort())
)
full_length = {
tuple(int(fold_id) for fold_id in ids)
for ids in per_candidate["fold_id"].to_list()
if len(ids) == n_folds
}
if not full_length:
raise RegistrySelectionError(f"no candidate reports all {n_folds} declared folds")
if len(full_length) > 1:
raise IncomparableFoldGeometryError(
f"candidates disagree on which {n_folds} folds they cover: {sorted(full_length)}"
)
return next(iter(full_length))
@lru_cache(maxsize=256)
def canonical_fold_ids(case_study: str, label: str) -> tuple[int, ...] | None:
"""The case study's whole modelling fold grid for one label, or ``None``.
``None`` means the grid is not derivable here, and there are exactly two ways to get
it. The first is boundaries carrying a time of day, which today means the two intraday
case studies: `modeling_fold_boundaries` runs its boundaries through `fold_boundary_date`,
which refuses a timestamp carrying a time of day - "the spans that read it are daily,
so truncating it would move the fold". That is 28 of the corpus's 105 (case study,
family, label) groups, all of them `crypto_perps_funding` at eight-hourly and
`nasdaq100_microstructure` at five, fifteen and sixty minutes. The second is the label
surface not being on disk, where `modeling_fold_boundaries` returns ``None`` outright:
`case_studies/*/labels` is gitignored and the reader bundle ships `run_log/` alone, so
a reader who downloaded artifacts to run the insight chapters without training gets
this for every label, not just the intraday ones.
Both are the same answer to the caller - completeness was not measured, which is not
the same as a group measured and found complete - and the selected rows keep that apart
from a measurement by carrying ``n_folds_canonical`` as null rather than as a number.
A consumer that prints an exclusion list has to say how many cells were never measured,
or its "none" reads as "every cell was checked".
A misconfiguration is not swallowed. `_derive_modeling_splits` raises `ValueError`
deliberately for config drift - a label with a parquet but no declared buffer, or one
whose parquet carries neither ``timestamp`` nor ``date`` - and `cv_window` states that
as a loud-fail contract. Caught here it would arrive as a third meaning of ``None`` and
the cell would sit in a census under the rule written for the intraday case studies with
nothing printed.
"""
try:
boundaries = modeling_fold_boundaries(case_study, label)
except IntradayFoldBoundaryError:
return None
if not boundaries:
return None
return tuple(sorted(int(boundary["fold"]) for boundary in boundaries))
def _fold_grid_columns(
case_study: str, label: str, expected_fold_ids: tuple[int, ...]
) -> dict[str, int | bool | None]:
"""How much of the case study's fold grid the selected candidate actually scored.
Carried on every selected row so a consumer that claims completeness can honour it.
`13_dl_time_series/12_case_study_insights`'s horizon figure is the one that has to:
it draws a line per case study across labels against a shared "average daily IC"
axis, and ``us_equities_panel/fwd_ret_5d`` is a Friday-resampled panel scored on four
of sixteen folds, so joining it to the daily ``fwd_ret_1d`` point would draw two
different quantities as one moving with horizon.
"""
canonical = canonical_fold_ids(case_study, label)
return {
"n_folds_scored": len(expected_fold_ids),
"n_folds_canonical": len(canonical) if canonical is not None else None,
"covers_fold_grid": (
None if canonical is None else set(expected_fold_ids) == set(canonical)
),
}
def _raw_primary_candidates(
case_study: str, family: str, label: str
) -> tuple[pl.DataFrame, pl.DataFrame, int]:
"""Candidate rows for one (case study, family, label), retired generations removed.
Two things about this query are load-bearing and were both absent until 2026-09-18.
It selects the ``auc_*`` columns beside the ``ic_*`` ones. They sit in the same
``prediction_metrics`` row and the chapters print them: Chapter 11's Table 11.6 has a
"Native AUC" column and Chapter 12's Table 12.4 a "Classifier -> AUC" column, and
neither notebook could reproduce its own published table because the query named the
seven IC columns and none of the five AUC ones.
It drops the generations a refit has retired. A registry keeps every generation - the
record of what was superseded is evidence - and supersession is recorded one layer up,
in ``official_populations``, so a reader that joins ``training_runs`` to
``prediction_metrics`` directly ranks history along with the present. That is not a
neutral error: a superseded generation is often the one that was refitted *because* it
was short or narrow, and a shorter or narrower sample is an easier one, so the stale
rows run toward the top of the ranking. Measured on ``cme_futures`` 2026-09-18, the
retired deep-learning generation covers 27,326 of 38,262 declared (entity, session)
pairs, 71.4%, and outranked the complete refit that replaced it.
"""
db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
if not db_path.exists():
return pl.DataFrame(), pl.DataFrame(), 0
uri = f"file:{db_path}?mode=ro"
with sqlite3.connect(uri, uri=True) as db:
db.row_factory = sqlite3.Row
rows = db.execute(
"""
SELECT p.prediction_hash, t.training_hash, t.family, t.config_name, t.label,
p.checkpoint_value, p.checkpoint_kind, t.git_commit, t.spec_json,
t.created_at AS training_created_at, pm.computed_at,
pm.ic_mean, pm.ic_std, pm.ic_mean_daily, pm.ic_n_days,
pm.ic_se_hac, pm.ic_ci_lo, pm.ic_ci_hi, pm.ic_t_hac,
pm.ic_p_hac, pm.ic_hac_lag,
pm.auc_mean_daily, pm.auc_se_hac, pm.auc_ci_lo, pm.auc_ci_hi,
pm.auc_n_days
FROM training_runs t
JOIN prediction_sets p ON p.training_hash = t.training_hash
JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
WHERE t.family = ? AND t.label = ? AND p.split = 'validation'
""",
(family, label),
).fetchall()
fold_rows = db.execute(
"""
SELECT p.prediction_hash, fm.fold_id, fm.ic
FROM training_runs t
JOIN prediction_sets p ON p.training_hash = t.training_hash
JOIN fold_metrics fm ON fm.prediction_hash = p.prediction_hash
WHERE t.family = ? AND t.label = ? AND p.split = 'validation'
""",
(family, label),
).fetchall()
retired = retired_prediction_hashes(db)
rows = [row for row in rows if row["prediction_hash"] not in retired]
# The fold frame is restricted to the prediction sets the metrics frame holds, and
# that is not the same filter as dropping the retired ones. The metrics query joins
# `prediction_metrics` and the fold query does not, and the two tables are written by
# separate calls - `register_prediction_set` then `register_fold_metrics` - so a set
# interrupted between them, or registered with `metrics=None`, has fold rows and no
# metrics row. It is unrankable, and it used to reach `resolve_expected_fold_ids`
# anyway: a full-length but differently numbered geometry from one raised
# `IncomparableFoldGeometryError` and stopped a whole chapter over a candidate
# `select_rank1` could never have returned. Restricted here rather than at each call
# site because the three collectors that take this frame resolve a geometry from it
# and only one of them had the guard. `collect_checkpoint_fold_trajectories` is a
# fourth resolver and does not take this frame - it queries the registry itself, and
# carries the same two filters inline.
rankable = {row["prediction_hash"] for row in rows}
fold_rows = [
row
for row in fold_rows
if row["prediction_hash"] not in retired and row["prediction_hash"] in rankable
]
if not rows:
return pl.DataFrame(), pl.DataFrame(), 0
metrics = pl.DataFrame([dict(row) for row in rows], infer_schema_length=None)
fold_metrics = pl.DataFrame([dict(row) for row in fold_rows], infer_schema_length=None)
declared_folds = []
for spec_json in metrics["spec_json"].drop_nulls().to_list():
with contextlib.suppress(json.JSONDecodeError, TypeError, ValueError):
value = declared_fold_count(json.loads(spec_json))
if value > 0:
declared_folds.append(value)
return metrics, fold_metrics, max(declared_folds, default=0)
def collect_rank1_per_cs(
case_studies: Iterable[str],
family: str,
label_resolver: LabelResolver | None = None,
) -> pl.DataFrame:
"""Rank-one configuration per case study, by daily-pooled IC.
A case study with no registry rows for this family is left out of the
result, so an absent case study shows up as a missing row rather than an
exception. Every other collector takes the frame this returns.
"""
rows = []
for case_study in case_studies:
label = _resolve_label(case_study, label_resolver)
metrics, folds, n_folds = _raw_primary_candidates(case_study, family, label)
if metrics.is_empty():
continue
if n_folds <= 0:
raise RegistrySelectionError(f"{case_study}/{family}/{label}: n_folds is not declared")
try:
expected_fold_ids = resolve_expected_fold_ids(folds, n_folds)
row = select_rank1(metrics, folds, expected_fold_ids=expected_fold_ids)
except IncomparableFoldGeometryError as exc:
raise IncomparableFoldGeometryError(f"{case_study}/{family}/{label}: {exc}") from exc
except RegistrySelectionError as exc:
raise RegistrySelectionError(f"{case_study}/{family}/{label}: {exc}") from exc
row.update(
case_study=case_study,
short_name=SHORT_NAMES.get(case_study, case_study),
**_fold_grid_columns(case_study, label, expected_fold_ids),
)
rows.append(row)
return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()
def collect_fold_ic_per_cs(rank1: pl.DataFrame) -> pl.DataFrame:
"""Per-fold IC for each configuration `collect_rank1_per_cs` selected."""
rows = []
for selected in rank1.iter_rows(named=True):
db_path = get_case_study_dir(selected["case_study"]) / "run_log" / "registry.db"
uri = f"file:{db_path}?mode=ro"
with sqlite3.connect(uri, uri=True) as db:
db.row_factory = sqlite3.Row
fold_rows = db.execute(
"""
SELECT fold_id, ic, ic_std, n_entities, rmse, mae
FROM fold_metrics WHERE prediction_hash = ? ORDER BY fold_id
""",
(selected["prediction_hash"],),
).fetchall()
for fold in fold_rows:
rows.append(
{
"case_study": selected["case_study"],
"short_name": selected["short_name"],
"family": selected["family"],
"config_name": selected["config_name"],
"label": selected["label"],
"prediction_hash": selected["prediction_hash"],
**dict(fold),
}
)
return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()
def collect_checkpoint_fold_trajectories(rank1: pl.DataFrame) -> pl.DataFrame:
"""Per-fold IC at every checkpoint of each selected training run."""
rows = []
for selected in rank1.iter_rows(named=True):
spec = json.loads(selected["spec_json"])
n_folds = declared_fold_count(spec)
if n_folds <= 0:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['training_hash']}: n_folds is not declared"
)
db_path = get_case_study_dir(selected["case_study"]) / "run_log" / "registry.db"
uri = f"file:{db_path}?mode=ro"
with sqlite3.connect(uri, uri=True) as db:
db.row_factory = sqlite3.Row
# Joined to `prediction_metrics` and filtered for retired hashes for the
# same reasons `_raw_primary_candidates` does both, and this is the fourth
# site that resolves a fold geometry. A checkpoint with fold rows and no
# metrics row is unrankable and would still set the standard here, and a
# retired generation under this same `training_hash` would too. Chapter 13
# calls this collector immediately after rank-one selection, so without the
# join the chapter dies one cell past the one the restriction fixed, saying
# the fold geometries disagree rather than that a registration is stray.
# Measured 2026-09-19 across all nine canonical registries: no metric-less
# validation prediction set and no duplicate (training_hash,
# checkpoint_value) group, so nothing reaches it today.
checkpoint_rows = db.execute(
"""
SELECT p.prediction_hash, p.checkpoint_value, fm.fold_id, fm.ic
FROM prediction_sets p
JOIN fold_metrics fm ON fm.prediction_hash = p.prediction_hash
JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
WHERE p.training_hash = ? AND p.split = 'validation'
ORDER BY p.checkpoint_value, fm.fold_id
""",
(selected["training_hash"],),
).fetchall()
retired = retired_prediction_hashes(db)
checkpoint_rows = [r for r in checkpoint_rows if r["prediction_hash"] not in retired]
checkpoints = pl.DataFrame([dict(row) for row in checkpoint_rows], infer_schema_length=None)
if checkpoints.is_empty():
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['training_hash']}: no checkpoint folds"
)
try:
expected_folds = set(resolve_expected_fold_ids(checkpoints, n_folds))
except IncomparableFoldGeometryError as exc:
raise IncomparableFoldGeometryError(
f"{selected['case_study']}/{selected['training_hash']}: {exc}"
) from exc
except RegistrySelectionError as exc:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['training_hash']}: {exc}"
) from exc
for checkpoint in checkpoints["checkpoint_value"].unique().sort().to_list():
current = checkpoints.filter(pl.col("checkpoint_value") == checkpoint)
fold_ids = current["fold_id"].to_list()
if len(fold_ids) != len(set(fold_ids)) or set(fold_ids) != expected_folds:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['training_hash']}/{checkpoint}: "
"checkpoint fold coverage is incomplete or duplicate"
)
if any(value is None or not math.isfinite(float(value)) for value in current["ic"]):
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['training_hash']}/{checkpoint}: "
"checkpoint IC is non-finite"
)
for row in current.iter_rows(named=True):
rows.append(
{
"case_study": selected["case_study"],
"short_name": selected["short_name"],
"config_name": selected["config_name"],
"training_hash": selected["training_hash"],
**row,
}
)
return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()
def conformal_coverage_for_selected_prediction(
selected: dict,
*,
levels: tuple[float, ...] = (0.80, 0.90, 0.95),
embargo_steps: int | None = None,
) -> pl.DataFrame:
"""Realised coverage of the sizing widths, for one exact selected prediction artifact.
Measured on the estimator `conformal_weighted` allocates with - see
:func:`~case_studies.utils.conformal.walk_forward_conformal_coverage`, which also states
why the figure is a diagnostic of residual dispersion rather than a guarantee.
``embargo_steps`` defaults to :func:`~case_studies.utils.conformal.sizing_conformal_lag`
for this row's case study and label.
The label is taken from the selected row where it carries one and from the training spec
otherwise - `us_equities_panel/15_model_analysis` attaches it to the returned frame rather
than to the row it passes, and both know the same label.
"""
required = {
"case_study",
"family",
"config_name",
"prediction_hash",
"spec_json",
}
missing = required - set(selected)
if missing:
raise RegistrySelectionError(f"selected row missing conformal fields: {sorted(missing)}")
spec = json.loads(selected["spec_json"])
n_folds = declared_fold_count(spec)
if n_folds < 2:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['prediction_hash']}: "
"conformal coverage requires at least two declared folds"
)
prediction_path = (
get_case_study_dir(selected["case_study"])
/ "run_log"
/ "predictions"
/ selected["prediction_hash"]
/ "predictions.parquet"
)
if not prediction_path.exists():
raise RegistrySelectionError(f"missing prediction artifact: {prediction_path}")
predictions = pl.read_parquet(prediction_path)
required_columns = {"timestamp", "y_true", "y_score", "fold_id"}
legacy_columns = {"actual": "y_true", "prediction": "y_score", "fold": "fold_id"}
present_legacy = {
legacy: canonical
for legacy, canonical in legacy_columns.items()
if legacy in predictions.columns and canonical not in predictions.columns
}
if present_legacy:
predictions = predictions.rename(present_legacy)
missing_columns = required_columns - set(predictions.columns)
if missing_columns:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['prediction_hash']}: "
f"prediction artifact missing {sorted(missing_columns)}"
)
usable = predictions.drop_nulls(required_columns)
for column in ("y_true", "y_score"):
usable = usable.filter(pl.col(column).cast(pl.Float64, strict=False).is_finite())
# The artifact has to carry every fold the spec declared. It is compared by count
# rather than against ``range(n_folds)`` because a run may number its folds any way
# it likes and one live sweep does: ``us_equities_panel/deep_learning/fwd_ret_5d``
# declares four folds and writes 0, 5, 10 and 15.
#
# The exact geometry is available - ``fold_metrics`` is keyed by ``prediction_hash``
# and records the ids this run scored - and the count is used instead because nothing
# downstream reads an id at all. `walk_forward_widths` selects ``fold_id`` and carries
# it to the output untouched; precedence comes from the ``step`` index built off the
# unique timestamps, so relabelling every fold changes no width. What would break the
# calibration is a fold that is missing, and that is what the count catches.
fold_ids = sorted(usable["fold_id"].unique().to_list())
if len(fold_ids) != n_folds:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['prediction_hash']}: "
f"expected {n_folds} declared folds, observed {fold_ids}"
)
if embargo_steps is None:
label = selected.get("label") or spec.get("label")
if not label:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['prediction_hash']}: neither the selected "
"row nor its training spec names a label, so no reviewed conformal embargo "
"can be resolved"
)
embargo_steps = sizing_conformal_lag(selected["case_study"], label)
try:
coverage_rows = walk_forward_conformal_coverage(
predictions, levels=levels, embargo_steps=embargo_steps
)
except ValueError as error:
raise RegistrySelectionError(
f"{selected['case_study']}/{selected['prediction_hash']}: {error}"
) from error
rows = [
{
"case_study": selected["case_study"],
"family": selected["family"],
"config_name": selected["config_name"],
"prediction_hash": selected["prediction_hash"],
**row,
}
for row in coverage_rows
]
return pl.DataFrame(rows)
def collect_grid_per_cs(
case_studies: Iterable[str],
family: str,
label_resolver: LabelResolver | None = None,
config_parser: Callable[[str], dict] | None = None,
) -> pl.DataFrame:
"""Best checkpoint of every configuration, per case study.
`collect_rank1_per_cs` keeps one row per case study; this keeps the whole
grid, so a chapter can show the spread the winner came from.
"""
rows = []
for case_study in case_studies:
label = _resolve_label(case_study, label_resolver)
metrics, folds, n_folds = _raw_primary_candidates(case_study, family, label)
if metrics.is_empty() or n_folds <= 0:
continue
# Resolved once over the whole group: a configuration whose candidates are all
# short would otherwise define its own standard and rank against complete ones.
try:
expected_fold_ids = resolve_expected_fold_ids(folds, n_folds)
except IncomparableFoldGeometryError as exc:
raise IncomparableFoldGeometryError(f"{case_study}/{family}/{label}: {exc}") from exc
except RegistrySelectionError:
continue
selected = []
# `Series.unique` does not define an order, so this loop set `dl_grid`'s row
# order and every group order downstream of it. Sorted, the frame is a function
# of the registry.
for config_name in sorted(metrics["config_name"].unique().to_list()):
config_metrics = metrics.filter(pl.col("config_name") == config_name)
try:
selected.append(
select_rank1(config_metrics, folds, expected_fold_ids=expected_fold_ids)
)
except RegistrySelectionError:
continue
if not selected:
continue
full_days = max(float(row["ic_n_days"]) for row in selected)
# Loop-invariant, and not free: it reads the label parquet's time column.
grid_columns = _fold_grid_columns(case_study, label, expected_fold_ids)
for selected_row in selected:
if float(selected_row["ic_n_days"]) != full_days:
continue
row = {
**selected_row,
"case_study": case_study,
"short_name": SHORT_NAMES.get(case_study, case_study),
**grid_columns,
}
if config_parser is not None:
row.update(config_parser(row["config_name"]))
rows.append(row)
return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()
def collect_multi_label_per_cs(
case_studies: Iterable[str],
family: str,
labels: list[str] | Callable[[str], list[str]],
) -> pl.DataFrame:
"""Rank-one configuration per case study and label.
`labels` is either one list applied to every case study, or a callable
`case_study -> list[label]` where the label sets differ.
"""
rows = []
for case_study in case_studies:
case_labels = labels(case_study) if callable(labels) else labels
for label in case_labels:
metrics, folds, n_folds = _raw_primary_candidates(case_study, family, label)
if metrics.is_empty():
continue
if n_folds <= 0:
raise RegistrySelectionError(
f"{case_study}/{family}/{label}: n_folds is not declared"
)
try:
expected_fold_ids = resolve_expected_fold_ids(folds, n_folds)
selected = select_rank1(metrics, folds, expected_fold_ids=expected_fold_ids)
except IncomparableFoldGeometryError as exc:
raise IncomparableFoldGeometryError(
f"{case_study}/{family}/{label}: {exc}"
) from exc
except RegistrySelectionError as exc:
raise RegistrySelectionError(f"{case_study}/{family}/{label}: {exc}") from exc
selected.update(
case_study=case_study,
short_name=SHORT_NAMES.get(case_study, case_study),
**_fold_grid_columns(case_study, label, expected_fold_ids),
)
rows.append(selected)
return pl.DataFrame(rows, infer_schema_length=None) if rows else pl.DataFrame()
def collect_gbm_checkpoint_trajectories(rank1: pl.DataFrame) -> pl.DataFrame:
"""Learning curve for each selected GBM training run."""
from case_studies.utils.registry import get_training_dir
frames = []
for row in rank1.iter_rows(named=True):
spec = json.loads(row["spec_json"])
path = get_training_dir(row["case_study"], spec) / "learning_curves.parquet"
if not path.exists():
raise RegistrySelectionError(f"missing learning curve for {row['training_hash']}")
curve = pl.read_parquet(path).filter(pl.col("config") == row["config_name"])
if curve.is_empty() or "iteration" not in curve.columns:
raise RegistrySelectionError(f"incomplete learning curve for {row['training_hash']}")
trajectory = (
curve.group_by("iteration")
.agg(
pl.col("ic_mean").mean().alias("ic_mean"),
pl.col("ic_std").mean().alias("ic_std"),
)
.sort("iteration")
.with_columns(
pl.lit(row["case_study"]).alias("case_study"),
pl.lit(row["short_name"]).alias("short_name"),
pl.lit(row["config_name"]).alias("config_name"),
pl.lit(row["training_hash"]).alias("training_hash"),
)
)
frames.append(trajectory)
return pl.concat(frames, how="diagonal_relaxed") if frames else pl.DataFrame()
def load_gbm_feature_importance(
case_study: str,
training_hash: str,
config_name: str,
*,
top_n: int,
num_iteration: int | None = None,
) -> pl.DataFrame:
"""Load booster importance from one selected GBM training identity.
Training saves every boosting round, while a configuration is selected at one
checkpoint along that trajectory. Pass that checkpoint as `num_iteration` and
the importance is measured over the first that many trees, which is the model
the selection actually chose; leave it None to read the whole booster.
"""
import lightgbm as lgb
case_dir = get_case_study_dir(case_study)
run_booster_dir = booster_dir(case_dir, training_hash)
if run_booster_dir is None:
return pl.DataFrame()
rows = []
for booster_file in sorted(run_booster_dir.glob("*.txt")):
fold_text = booster_file.stem.split("fold")[-1].lstrip("_")
with contextlib.suppress(ValueError):
fold_id = int(fold_text)
model = lgb.Booster(model_file=str(booster_file))
if num_iteration is not None and num_iteration < model.num_trees():
model = lgb.Booster(model_str=model.model_to_string(num_iteration=num_iteration))
rows.extend(
{
"config_name": config_name,
"training_hash": training_hash,
"fold_id": fold_id,
"feature": feature,
"importance": float(importance),
}
for feature, importance in zip(
model.feature_name(),
model.feature_importance(importance_type="gain"),
strict=False,
)
)
if not rows:
return pl.DataFrame()
result = pl.DataFrame(rows).with_columns(
fold_max=pl.col("importance").max().over(["config_name", "fold_id"])
)
if result.filter(pl.col("fold_max") <= 0).height:
raise RegistrySelectionError(f"{case_study}/{training_hash}: zero feature importance")
result = result.with_columns(importance_norm=pl.col("importance") / pl.col("fold_max")).drop(
"fold_max"
)
return result.filter(pl.col("feature").is_in(top_features_by_gain(result, top_n)))
def plot_cross_cs_forest(
df: pl.DataFrame,
family: str,
title: str,
*,
sort_by: str = "ic_mean_daily",
figsize: tuple[float, float] | None = None,
sig_t: float = 2.0,
):
"""Forest of rank-1 daily-pooled IC ± HAC 95% CI per case study.
Marker style encodes whether the HAC CI excludes zero:
- filled circle: |t_hac| > sig_t (distinguishable from zero)
- open circle: |t_hac| ≤ sig_t (overlaps zero)
Y-axis order is ascending by `sort_by`, so the largest value sits at top.
Returns (fig, ax).
"""
import matplotlib.pyplot as plt
import numpy as np
from utils.style import COLORS
if df.is_empty():
fig, ax = plt.subplots(figsize=(7, 2.5))
ax.text(0.5, 0.5, f"No {family} runs in registry", ha="center", va="center")
ax.set_title(title)
ax.set_axis_off()
return fig, ax
d = df.sort(sort_by, descending=False, nulls_last=False).to_pandas()
n = len(d)
if figsize is None:
figsize = (7.5, max(2.5, 0.45 * n + 1.2))
fig, ax = plt.subplots(figsize=figsize, layout="tight")
y = np.arange(n)
ic = d["ic_mean_daily"].to_numpy()
lo = d["ic_ci_lo"].to_numpy()
hi = d["ic_ci_hi"].to_numpy()
t_hac = d["ic_t_hac"].to_numpy() if "ic_t_hac" in d.columns else np.full(n, np.nan)
ax.errorbar(
ic,
y,
xerr=[ic - lo, hi - ic],
fmt="none",
color=COLORS["neutral"],
lw=1.0,
capsize=3,
)
t_arr = np.asarray(t_hac, dtype=float)
sig = (np.abs(t_arr) > sig_t) & ~np.isnan(t_arr)
if sig.any():
ax.scatter(
ic[sig],
y[sig],
s=60,
marker="o",
facecolor=COLORS["blue"],
edgecolor=COLORS["blue"],
zorder=3,
label=f"|t_hac| > {sig_t:g}",
)
if (~sig).any():
ax.scatter(
ic[~sig],
y[~sig],
s=60,
marker="o",
facecolor="white",
edgecolor=COLORS["blue"],
zorder=3,
label=f"|t_hac| ≤ {sig_t:g}",
)
ax.axvline(0, color=COLORS["neutral"], lw=0.7, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(d["short_name"].tolist())
ax.set_xlabel("Daily-pooled IC (HAC 95% CI)")
ax.set_title(title)
ax.legend(loc="lower right", fontsize=8, frameon=False)
if fig.get_layout_engine() is None:
fig.tight_layout()
return fig, ax
def plot_per_fold_violin(
fold_df: pl.DataFrame,
order: list[str],
*,
title: str,
figsize: tuple[float, float] | None = None,
jitter_color: str = "#3B82F6",
):
"""Box-plus-scatter of per-fold IC across case studies.
`fold_df` must carry columns `short_name` and `ic`. `order` is the CS
display order (left → right). CSs absent from `fold_df` are skipped.
"""
import matplotlib.pyplot as plt
import numpy as np
present = [c for c in order if c in fold_df["short_name"].unique().to_list()]
if not present:
fig, ax = plt.subplots(figsize=(7, 2.5))
ax.text(0.5, 0.5, "No fold IC data", ha="center", va="center")
ax.set_axis_off()
return fig, ax
if figsize is None:
figsize = (max(6.5, 1.0 * len(present) + 2), 4.5)
fig, ax = plt.subplots(figsize=figsize, layout="tight")
data = [fold_df.filter(pl.col("short_name") == cs)["ic"].to_numpy() for cs in present]
positions = np.arange(len(present))
ax.boxplot(data, positions=positions, widths=0.55, showfliers=True)
for i, arr in enumerate(data):
if len(arr):
ax.scatter(np.full(len(arr), i), arr, alpha=0.5, s=14, color=jitter_color)
ax.axhline(0, color="gray", linewidth=0.7, linestyle="--")
ax.set_xticks(positions)
ax.set_xticklabels(present, rotation=30, ha="right")
ax.set_ylabel("Per-fold Spearman IC")
ax.set_title(title)
if fig.get_layout_engine() is None:
fig.tight_layout()
return fig, ax
def _rank1_full_coverage_hash(case_study: str, label: str) -> str | None:
"""Validation prediction_hash of the highest daily-IC linear config with NO dropped fold.
Selection is coverage-aware: a configuration whose path zeroes out on some
fold leaves that fold's IC NULL and pools its daily IC over fewer days,
which can make a naive ``MAX(ic_mean_daily)`` crown a config scored on a
non-comparable subset. We therefore rank only among configs whose
fold_metrics carry no NULL IC. Returns None if the registry has no
full-coverage linear row for that label.
"""
db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
if not db_path.exists():
return None
with sqlite3.connect(db_path) as db:
rows = db.execute(
"""
SELECT p.prediction_hash, pm.ic_mean_daily, pm.ic_n_days,
(SELECT COUNT(*) FROM fold_metrics fm
WHERE fm.prediction_hash = p.prediction_hash AND fm.ic IS NULL) AS n_null
FROM training_runs t
JOIN prediction_sets p ON p.training_hash = t.training_hash AND p.split = 'validation'
JOIN prediction_metrics pm ON pm.prediction_hash = p.prediction_hash
WHERE t.family = 'linear' AND t.label = ? AND pm.ic_mean_daily IS NOT NULL
""",
(label,),
).fetchall()
full = [r for r in rows if r[2] is not None and r[3] == 0]
if not full:
return None
max_n_days = max(r[2] for r in full)
full = [r for r in full if r[2] == max_n_days]
full.sort(key=lambda r: -r[1])
return full[0][0]
def plot_rolling_daily_ic(
case_studies: Iterable[str],
*,
window: int = 63,
label_resolver: LabelResolver | None = None,
common_window: bool = True,
title: str = "Persistence of linear ranking signal (rolling daily IC)",
figsize: tuple[float, float] = (10, 4),
colors: list[str] | None = None,
selected_prediction_hashes: dict[str, str] | None = None,
):
"""Rolling-mean daily-IC persistence chart for the rank-1 linear fit per case study.
For each case study, take the coverage-aware rank-1 linear configuration
(highest daily IC with no dropped fold), load its per-day IC series from
``daily_metrics.parquet``, average across overlapping folds per calendar
day, and plot a ``window``-day rolling mean (63 ≈ three trading months).
Case studies cover different validation windows, so they cannot share a
calendar axis unless their periods overlap. With ``common_window=True``
the series are clipped to the intersection of all input case studies'
spans — pass only case studies whose windows overlap (e.g. ``etfs`` and
``fx_pairs`` over 2016–2023). Case studies without a daily-IC series are
skipped.
Returns (fig, ax).
"""
import matplotlib.pyplot as plt
from case_studies.utils.model_analysis import load_daily_metrics_series
if colors is None:
from utils.style import COLORS
colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"], COLORS["slate"]]
series: dict[str, pl.DataFrame] = {}
for cs in case_studies:
label = _resolve_label(cs, label_resolver)
h = (
selected_prediction_hashes.get(cs)
if selected_prediction_hashes is not None
else _rank1_full_coverage_hash(cs, label)
)
if h is None:
continue
d = load_daily_metrics_series(cs, h)
if d.is_empty() or "date" not in d.columns:
continue
roll = (
d.group_by("date")
.agg(pl.col("ic").mean())
.sort("date")
.with_columns(
pl.col("ic").rolling_mean(window_size=window, min_periods=window).alias("roll")
)
.drop_nulls("roll")
)
if not roll.is_empty():
series[cs] = roll
fig, ax = plt.subplots(figsize=figsize, layout="tight")
if not series:
ax.text(0.5, 0.5, "No daily-IC series available", ha="center", va="center")
ax.set_axis_off()
return fig, ax
lo = max(s["date"].min() for s in series.values()) if common_window else None
hi = min(s["date"].max() for s in series.values()) if common_window else None
for idx, (cs, roll) in enumerate(series.items()):
sub = roll
if common_window:
sub = roll.filter((pl.col("date") >= lo) & (pl.col("date") <= hi))
ax.plot(
sub["date"].to_numpy(),
sub["roll"].to_numpy(),
color=colors[idx % len(colors)],
linewidth=1.6,
label=SHORT_NAMES.get(cs, cs),
)
from utils.style import COLORS
ax.axhline(0, color=COLORS["neutral"], linewidth=0.8, linestyle="--")
ax.set_ylabel(f"{window}-day rolling daily IC")
ax.set_xlabel("Validation date")
ax.set_title(title)
ax.legend(loc="upper right", frameon=False)
if fig.get_layout_engine() is None:
fig.tight_layout()
return fig, ax
# Horizon → trading-day mapping shared by all chapter insight notebooks.
HORIZON_DAYS: dict[str, float] = {
"fwd_ret_5m": 5 / (6.5 * 60),
"fwd_ret_15m": 15 / (6.5 * 60),
"fwd_ret_60m": 60 / (6.5 * 60),
"fwd_ret_8h": 1.0 / 3,
"fwd_ret_24h": 1.0,
"fwd_ret_1d": 1.0,
"fwd_ret_5d": 5.0,
"fwd_ret_10d": 10.0,
"fwd_ret_21d": 21.0,
"fwd_ret_1m": 21.0,
"fwd_ret_3m": 63.0,
"fwd_ret_1m_win": 21.0,
"fwd_ret_risk_adj_5d": 5.0,
}
def plot_multi_label_horizon(
horizon_df: pl.DataFrame,
*,
title: str,
min_labels_per_cs: int = 2,
figsize: tuple[float, float] = (10, 5),
palette: list[str] | None = None,
):
"""Faceted horizon plot: daily-pooled IC vs trading-day horizon per CS.
`horizon_df` must carry `short_name`, `label`, `ic_mean_daily`, `ic_ci_lo`,
`ic_ci_hi`. CSs with fewer than `min_labels_per_cs` mapped horizons are
omitted from the figure (a coverage fact, not a defect).
"""
import matplotlib.pyplot as plt
plot_df = horizon_df.with_columns(
horizon_days=pl.col("label").replace_strict(HORIZON_DAYS, default=None).cast(pl.Float64),
).filter(pl.col("horizon_days").is_not_null())
multi_cs = (
plot_df.group_by("short_name")
.len()
.filter(pl.col("len") >= min_labels_per_cs)["short_name"]
.to_list()
)
plot_df = plot_df.filter(pl.col("short_name").is_in(multi_cs))
if plot_df.is_empty():
fig, ax = plt.subplots(figsize=(7, 2.5))
ax.text(0.5, 0.5, "No multi-horizon coverage", ha="center", va="center")
ax.set_axis_off()
return fig, ax
fig, ax = plt.subplots(figsize=figsize, layout="tight")
if palette is None:
from utils.style import COLORS
palette = [
COLORS["blue"],
COLORS["amber"],
COLORS["copper"],
COLORS["positive"],
COLORS["negative"],
COLORS["neutral"],
COLORS["slate"],
COLORS["amber_light"],
]
cs_sorted = sorted(plot_df["short_name"].unique().to_list())
markers = ["o", "s", "D", "^", "v", "P", "X", "*"]
linestyles = ["-", "--", "-.", ":", "-", "--", "-.", ":"]
for idx, cs in enumerate(cs_sorted):
# Sorted on (horizon_days, label). HORIZON_DAYS maps several labels onto one
# value - fwd_ret_1d and fwd_ret_24h onto 1.0, fwd_ret_5d and
# fwd_ret_risk_adj_5d onto 5.0, fwd_ret_21d, fwd_ret_1m and fwd_ret_1m_win onto
# 21.0 - so a case study carrying both members of a pair has two points at one x
# and the tie order decides which the line reaches first. The tie is live, not
# hypothetical: measured 2026-09-19, SP500 Eq+Opt has fwd_ret_5d and
# fwd_ret_risk_adj_5d in the deep_learning census, and gbm, linear and tabular_dl
# each add US Firms at fwd_ret_1m against fwd_ret_1m_win. Sorting on the horizon
# alone left the drawn order a property of the frame; it happens to match today,
# so pinning it moves no committed figure.
sub = plot_df.filter(pl.col("short_name") == cs).sort(["horizon_days", "label"])
if sub.height < 2:
continue
x = sub["horizon_days"].to_numpy()
ic = sub["ic_mean_daily"].to_numpy()
lo = sub["ic_ci_lo"].to_numpy()
hi = sub["ic_ci_hi"].to_numpy()
color = palette[idx % len(palette)]
ax.fill_between(x, lo, hi, color=color, alpha=0.12)
ax.plot(
x,
ic,
marker=markers[idx % len(markers)],
linestyle=linestyles[idx % len(linestyles)],
color=color,
label=cs,
linewidth=1.6,
markersize=6,
alpha=0.9,
)
ax.set_xscale("log")
ax.set_xlabel("Horizon (trading days, log scale)")
ax.set_ylabel("Daily-pooled IC (HAC 95 % CI band)")
ax.axhline(0, color="gray", linewidth=0.7, linestyle="--")
ax.set_title(title)
ax.legend(loc="best", frameon=False, fontsize=8, ncol=2)
if fig.get_layout_engine() is None:
fig.tight_layout()
return fig, ax
def parse_gbm_config(config: str) -> dict:
"""Decode a GBM `config_name` into profile / loss / leaves / objective_kind.
Examples
--------
>>> parse_gbm_config("leaves_31_huber")
{"profile": "leaves_31", "loss": "huber", "leaves": 31, "objective_kind": "regression"}
>>> parse_gbm_config("default_binary")
{"profile": "default", "loss": "binary", "leaves": None, "objective_kind": "classification"}
"""
out = {"profile": config, "loss": "unknown", "leaves": None, "objective_kind": "regression"}
parts = config.rsplit("_", 1)
if len(parts) == 2 and parts[1] in ("mse", "mae", "huber"):
out["profile"], out["loss"] = parts
elif config.endswith("_binary"):
out["loss"] = "binary"
out["objective_kind"] = "classification"
out["profile"] = config.removesuffix("_binary")
if "leaves_" in out["profile"]:
with contextlib.suppress(ValueError, IndexError):
out["leaves"] = int(out["profile"].split("_")[1])
return out
SYMMETRY_TABLE_SCHEMA: dict[str, pl.DataType] = {
"short_name": pl.Utf8,
"reg_label": pl.Utf8,
"dir_label": pl.Utf8,
"cls_config": pl.Utf8,
"cls_score_ic": pl.Float64,
"cls_score_ic_lo": pl.Float64,
"cls_score_ic_hi": pl.Float64,
"cls_score_ic_t": pl.Float64,
"cls_score_auc": pl.Float64,
"cls_score_auc_lo": pl.Float64,
"cls_score_auc_hi": pl.Float64,
"reg_config": pl.Utf8,
"reg_score_auc": pl.Float64,
"n_b": pl.Int64,
"n_b_days": pl.Int64,
}
"""The columns of the Chapter 11 and 12 symmetry tables, declared once.
A full schema and not ``schema_overrides``: `discover_symmetry_pairs` can legitimately
return nothing - every declared pair skipped, or a registry with no classification runs
for this family - and a frame built from an empty list with partial overrides has no
string columns at all, so selecting ``short_name`` raises ``ColumnNotFoundError`` and the
skip reasons, which on that path are the entire output, never reach the reader.
Shared rather than written out in each notebook because two copies of a column list drift,
and because a test can only pin the contract if there is one thing to pin. Chapter 11
carries ``n_b_days`` and Chapter 12 does not; both tolerate the extra column in an empty
frame, which is cheaper than two schemas.
"""
# The pairing of a regression label with the binary direction label it is scored against
# used to be a hand-written literal in each chapter that draws the cross-evaluation. It
# shrank without saying so: `us_firm_characteristics: [("fwd_ret_1m", "fwd_class_1m")]`
# was dropped from Chapter 12's copy by the 2026-07-31 chapter-tree restore and kept in
# Chapter 11's, so one chapter published four rows and the other three while both
# reported a full count of their own literal. A literal cannot report what is missing
# from it, so the pairs are discovered and the skips are named.
def _binary_label_domain(case_study: str, label: str) -> set[int] | None:
"""Distinct values of one label surface, or None when the file is absent."""
path = get_case_study_dir(case_study) / "labels" / f"{label}.parquet"
if not path.is_file():
return None
values = pl.scan_parquet(path).select(pl.col(label)).unique().collect().to_series().drop_nulls()
return {int(value) for value in values.to_list()}
def _declared_classification_pairs(case_study: str) -> dict[str, str]:
"""``{direction_label: regression_label}`` as ``setup.yaml`` declares it.
Read from ``labels.classification_eval_label`` rather than inferred from the label
names. Two of the nine case studies would be got wrong by a naming rule: crypto maps
both ``fwd_dir_8h`` and ``fwd_dir_8h_3c`` onto ``fwd_ret_8h``, so a rule that builds
the direction name from the regression suffix never sees the ternary one and cannot
report skipping it, which is the failure this whole function exists to avoid.
"""
import yaml
setup_path = get_case_study_dir(case_study) / "config" / "setup.yaml"
if not setup_path.is_file():
return {}
with setup_path.open() as handle:
setup = yaml.safe_load(handle)
labels = (setup or {}).get("labels") or {}
declared = labels.get("classification_eval_label") or {}
return {str(k): str(v) for k, v in declared.items()}
def discover_symmetry_pairs(
case_studies: Iterable[str], family: str
) -> tuple[dict[str, list[tuple[str, str]]], list[str]]:
"""Regression and binary-direction label pairs, per case study, from the declaration.
A pair qualifies when ``setup.yaml`` declares the direction label's
``classification_eval_label``, this family has registered validation predictions for
both labels, and the direction label's own surface is binary. A ternary direction
label would need a multi-class AUC and is out of scope for this comparison, so it is
skipped by measuring its domain rather than by being left out of a list.
Returns ``(pairs, skipped)``. Every declared candidate that does not qualify appears
in ``skipped`` with its reason, so a comparison that covers fewer case studies than
the corpus says which ones and why instead of reporting a full count of itself.
"""
pairs: dict[str, list[tuple[str, str]]] = {}
skipped: list[str] = []
for case_study in case_studies:
declared = _declared_classification_pairs(case_study)
if not declared:
continue
db_path = get_case_study_dir(case_study) / "run_log" / "registry.db"
if not db_path.is_file():
continue
with sqlite3.connect(f"file:{db_path}?mode=ro", uri=True) as db:
registered = {
row[0]
for row in db.execute(
"SELECT DISTINCT t.label FROM training_runs t "
"JOIN prediction_sets p ON p.training_hash = t.training_hash "
"WHERE t.family = ? AND p.split = 'validation'",
(family,),
)
}
found: list[tuple[str, str]] = []
for direction, regression in sorted(declared.items()):
if direction not in registered or regression not in registered:
# Not a defect and not worth a line: this family simply did not run one
# of the two labels here. Only a declared pair the family DID run, and
# that still does not qualify, is a skip worth naming.
continue
domain = _binary_label_domain(case_study, direction)
if domain is None:
skipped.append(f"{case_study}/{direction}: no label surface on disk")
elif not domain.issubset({0, 1}):
skipped.append(
f"{case_study}/{direction}: domain {sorted(domain)} is not binary, "
"a multi-class AUC is out of scope here"
)
else:
found.append((regression, direction))
if found:
pairs[case_study] = found
return pairs, skipped
```Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT
Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.