Agregación y comparación de backtests en estudios de trading
Resumen
Este cuaderno consulta registros de backtests en nueve estudios de caso y prepara tablas comparables para análisis posteriores. Organiza los resultados por clase de activo y frecuencia de los datos, y registra el rendimiento en las etapas de señal, asignación, costes y riesgo, junto con información sobre familias de modelos y resultados de validación frente al periodo de reserva. El proceso de selección compara especificaciones completas de estrategia —la señal, el método de asignación y el componente de riesgo— y sigue la configuración seleccionada hasta una ejecución de reserva con reentrenamiento. Si esta no produce un backtest de reserva utilizable, recurre a otro candidato elegible según la clasificación de validación.
El cuaderno también mide cuán concentrados o estables parecen los mejores resultados de validación. Presenta la diferencia entre las configuraciones mejor clasificadas y las de menor rango, la variación entre segmentos y la frecuencia con que la configuración líder obtiene resultados positivos en los distintos segmentos. Una agregación completa se niega a escribir resultados para todo el capítulo si falta algún registro previsto o está vacío, para evitar que un conjunto parcial de datos parezca una comparación completa. Estos diagnósticos ayudan a matizar la evidencia entre casos, pero los resúmenes heredan los datos, costes, etiquetas y supuestos de backtest de cada estudio; las clasificaciones de validación y los resultados del periodo de reserva no deben interpretarse como rendimiento universal de una estrategia.
Ideas clave
- Compara backtests entre estudios de caso mediante registros y tablas resumen coherentes por etapa.
- Define una especificación de estrategia con su señal, método de asignación y componente de riesgo considerados en conjunto.
- Sigue la especificación de validación seleccionada hasta la evaluación del periodo de reserva y recurre a otra si su ejecución no es utilizable.
- Usa la amplitud del grupo superior y los resultados por segmento para evaluar cuánto puede depender una selección del ruido de clasificación.
- Exige todos los registros antes de publicar agregados de alcance completo, para que una cobertura incompleta no parezca completa.
Etiquetas
Texto completo
# 01_aggregate_synthesis.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Aggregate Synthesis
#
# **Docker image**: `ml4t`
#
# This notebook queries all 9 case study registries via `BacktestExplorer`
# and builds cross-dataset comparison DataFrames for the remaining Ch20 notebooks.
#
# **Data source**: `registry.db` per case study (no JSON files needed).
#
# **Learning Objectives**:
# - Query per-case-study backtest registries for signal, allocation, cost, and risk metrics
# - Build cross-dataset comparison tables
# - Export summary DataFrames for downstream notebooks (02–06)
#
# **Book Reference**: Chapter 20, Section 20.1 (First-pass results across nine case studies)
#
# **Prerequisites**: Case studies must have run Ch16–19 backtests.
# %%
"""Ch20 Aggregate Synthesis — query registries and compare all 9 case studies."""
import json
import sqlite3
from functools import cache
from pathlib import Path
import polars as pl
import yaml
from IPython.display import Markdown, display
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_returns
from case_studies.utils.strategy_analysis import (
allocation_method_of,
compute_cost_bps,
rank_one,
training_run_fitted_for_the_holdout,
)
from utils.paths import REPO_ROOT, get_case_study_dir, get_chapter_dir
# %% tags=["parameters"]
MAX_SYMBOLS = 0
# When non-empty, restricts the cross-CS iteration to the given subset.
# Used by the per-CS pipeline driver to populate `backtest_paired_metrics`
# for a single CS after its holdout has landed, without re-running the
# full 9-CS aggregation.
CASE_STUDIES: list[str] = []
# Test-only: in an isolated test registry, nasdaq's out-of-band cost-feasible
# selected configuration is absent, so its spine cannot resolve. Production leaves this False
# (a missing selection fails loudly); the test harness sets it True so cost/risk
# for such a case study are reported not-applicable instead of raising.
ALLOW_MISSING_SPINE = False
# A full-mode run overwrites the nine-case-study artifacts every downstream notebook and the
# chapter figures read. Refuse to start one unless all nine registries are present and carry
# backtests, so a partial aggregation cannot be published as a complete one. The test harness
# sets this False because its isolated registry is not the production store; a subset run
# (`CASE_STUDIES` non-empty) is the per-case-study driver path and is never full mode.
REQUIRE_ALL_REGISTRIES = True
# %%
OUTPUT_DIR = get_chapter_dir(20) / "output"
OUTPUT_DIR.mkdir(exist_ok=True)
ALL_CASE_STUDIES = [
"etfs",
"crypto_perps_funding",
"nasdaq100_microstructure",
"sp500_equity_option_analytics",
"us_firm_characteristics",
# FX rank-1 is linear/ridge_a100.0 on fwd_ret_21d (val Sharpe +0.048,
# holdout +0.194), resolved after the 2026-06-01 DL-lookback fix. The
# earlier deep_learning/tcn selection (val +0.108 / holdout -1.59) was an
# artifact of gappy validation folds (lookback=60 warmup consumed each
# fold's head); those sets were purged and the clean lineage re-resolved.
# See backtest_audit.md and project_registry_hash_collisions.
"fx_pairs",
"cme_futures",
"sp500_options",
"us_equities_panel",
]
if CASE_STUDIES:
ALL_CASE_STUDIES = [cs for cs in ALL_CASE_STUDIES if cs in set(CASE_STUDIES)]
DISPLAY_NAMES = {
"etfs": "ETFs",
"crypto_perps_funding": "Crypto",
"nasdaq100_microstructure": "NASDAQ-100",
"sp500_equity_option_analytics": "S&P 500 Eq+Opt",
"us_firm_characteristics": "US Firms",
"fx_pairs": "FX Pairs",
"cme_futures": "CME Futures",
"sp500_options": "S&P 500 Options",
"us_equities_panel": "US Equities",
}
# %%
ASSET_CLASS_MAP = {
"etfs": "equity_etf",
"crypto_perps_funding": "crypto",
"nasdaq100_microstructure": "equity_micro",
"sp500_equity_option_analytics": "equity_options",
"us_firm_characteristics": "equity_firm",
"fx_pairs": "fx",
"cme_futures": "futures",
"sp500_options": "options",
"us_equities_panel": "equity_panel",
}
FREQ_MAP = {
"etfs": "daily",
"crypto_perps_funding": "8h",
"nasdaq100_microstructure": "15min",
"sp500_equity_option_analytics": "daily",
"us_firm_characteristics": "monthly",
"fx_pairs": "daily",
"cme_futures": "daily",
"sp500_options": "daily",
"us_equities_panel": "daily",
}
# %% [markdown]
# ## Load Registries
#
# Create a `BacktestExplorer` for each case study that has a registry.
# %%
explorers: dict[str, BacktestExplorer] = {}
configs: dict[str, dict] = {}
unreadable_registries: list[str] = []
empty_registries: list[str] = []
for cs in ALL_CASE_STUDIES:
try:
explorers[cs] = BacktestExplorer(cs)
setup_path = get_case_study_dir(cs) / "config" / "setup.yaml"
if setup_path.exists():
configs[cs] = yaml.safe_load(setup_path.read_text())
else:
configs[cs] = {}
summary = explorers[cs].summary()
total = sum(summary.values())
if total == 0:
empty_registries.append(cs)
print(f" [EMPTY] {cs}: registry.db present, zero backtest runs")
else:
print(f" [OK] {cs}: {total} backtests ({summary})")
except FileNotFoundError:
unreadable_registries.append(cs)
print(f" [MISSING] {cs}: no registry.db")
print(f"\nLoaded: {len(explorers)}/{len(ALL_CASE_STUDIES)} case studies")
# %% [markdown]
# ### The full-mode precondition
#
# This notebook overwrites the nine-case-study artifacts that notebooks 02 through 08 and the
# chapter figures read. A run that finds only some of the nine registries produces an output
# indistinguishable from a complete one, and on 2026-08-28 that is exactly what happened: a
# synthesis run from a worktree carrying three registries stamped itself production and
# published a holdout Sharpe under prose describing nine case studies. The push gate caught it;
# the notebook did not.
#
# So a full-mode run refuses to continue unless every case study named above has a readable
# registry holding backtests. A subset run - `CASE_STUDIES` non-empty, the per-case-study
# driver path that repopulates one case study's paired metrics after its holdout lands - is not
# full mode and is not covered by the check.
# %%
def refuse_partial_full_mode(
*,
expected: list[str],
subset: list[str],
unreadable: list[str],
empty: list[str],
enforce: bool = True,
) -> None:
"""Raise unless a full-mode run can see every registry it claims to aggregate.
A subset run - ``subset`` non-empty - is the per-case-study driver path and is not
full mode, so it is never refused. ``enforce`` is the seeded-test-registry escape and
is False in exactly one place, ``tests/overrides.yaml``.
"""
if not enforce or subset or not (unreadable or empty):
return
raise RuntimeError(
f"Full-mode synthesis needs all {len(expected)} registries present and holding "
"backtests, and this checkout does not have them. Refusing before anything is "
"written, because the artifacts this notebook overwrites are read as a complete "
"set covering every case study.\n"
f" no registry.db: {unreadable or 'none'}\n"
f" zero backtest runs: {empty or 'none'}\n"
"Run this once every case study has registered its backtests and the fleet has "
"stopped writing to them. Passing CASE_STUDIES is not a way round this: a subset "
"run repopulates one case study's paired metrics in its own registry and writes "
"none of the chapter-wide artifacts."
)
refuse_partial_full_mode(
expected=ALL_CASE_STUDIES,
subset=CASE_STUDIES,
unreadable=unreadable_registries,
empty=empty_registries,
enforce=REQUIRE_ALL_REGISTRIES,
)
# %% [markdown]
# ## Top-Cluster Diagnostics
#
# Rather than pre-committing to a single selected configuration per case study, we inspect the
# *cluster* of top configurations on the validation split. A signal with genuine predictive
# structure shows a thick top of the distribution: many configurations sit within a
# fold-standard-error of the top-ranked Sharpe, and the implied pick is insensitive to small
# perturbations in the selection rule. A thin cluster - a large gap between the top-ranked
# configuration and the tenth - suggests the top result is closer to a tail draw than to a
# stable optimum.
#
# For each case study we report the top-ranked Sharpe, the tenth-ranked Sharpe where ten
# configurations exist, the spread between them, the mean per-fold Sharpe, and the number of
# folds in which the top-ranked configuration has positive Sharpe. These are measurements that
# feed the downstream narrative.
#
# ### The selection rule
#
# **A backtest's full strategy specification is the signal method, the allocation method and the
# risk overlay taken together.** Naming all three is what makes a validation result and a
# holdout result comparable, because it pins every stage rather than the signal alone.
#
# Each case study's selected configuration is the highest-Sharpe validation backtest across
# those three pipeline stages. The deployed holdout configuration is that same specification,
# retrained on holdout data. When the holdout retrain produces no usable backtest at it -
# degenerate predictions, a vol window that does not match the history available, a universe
# filter that rejects the sample, or another generation failure - the rule falls back to the
# next-highest validation Sharpe that does have a usable holdout, and so on until one succeeds.
# The helper implementing that walk feeds the holdout query and the lineage resolver, which pin
# each validation and holdout pair to one specification.
# %% [markdown]
# ### Two restrictions, and why the selection needs both
#
# **A label restriction**, so the cluster diagnostics and the Chapter 20 holdout retrain rank the
# same thing. sp500_options trains a hold-to-maturity label with coherent option costs alongside
# four fixed-horizon straddle labels priced through the vectorized path with a generic
# basis-point cost. The Chapter 20 narrative uses the hold-to-maturity label as its
# option-strategy reference, so restricting the cluster diagnostics to that same label keeps the
# §20.1 top-cluster numbers aligned with the §20.5 and §20.6 narrative.
#
# Those two section numbers are right, and they are recorded here because they will not look it.
# Eleven references in this chapter pointed at sections that exist and do not carry what was
# claimed, and a sweep for §20.5 in a chapter-20 notebook now finds this one and sees the same
# shape. It is not the same: §20.5's Table 20.6 carries an sp500_options allocator row, and which
# row that is depends on the label pinned here; §20.6 carries the option cost model, which is the
# hold-to-maturity accounting rather than the basis-point sweep the other four labels get. Both
# targets hold material this restriction decides. Do not retarget it.
#
# **An execution-regime restriction**, because sp500_options is evaluated under the
# O'Donovan-Yu (2025) cost-mitigation cascade, whose three rungs are a naive round trip, full
# hold-to-maturity, and hold-to-maturity restricted to the liquid bottom-spread quintile. The
# registered strategy is the third rung; the second is the demoted variant §18.8 discusses. The
# first two rungs both carry the same universe filter, so filtering on that column
# alone leaves `ORDER BY sharpe DESC LIMIT 1` free to pick whichever of the two happens to score
# higher in the current data. Pinning the universe filter *and* the exit rule together is what
# makes the selected row deterministic and coherent with hold-to-maturity. Case studies with no
# entry here skip the filter altogether.
# %%
# The rung pins are imported, not restated. This file used to carry its own copy of both
# predicates and of the dict around them, verbatim, and `paired_metrics.populate_paired_metrics`
# carries the other - both write `backtest_paired_metrics`, so a pin corrected on one side only
# would let one of them overwrite the other's rows with a differently-selected lineage. The
# duplication is how that drift happens, and the mirror keys beside each predicate
# (`universe_filter`, `exit_at_max_days`, `label`) exist for the SQL paths and `progression(...)`
# calls that cannot take a polars expression.
from case_studies.utils.paired_metrics import RUNG_PINS as _CLUSTER_RUNG_RESTRICTIONS # noqa: E402
from case_studies.utils.strategy_analysis import ( # noqa: E402
LABEL_RESTRICTIONS as _CLUSTER_LABEL_RESTRICTIONS,
)
from case_studies.utils.strategy_analysis import ( # noqa: E402
NoSelectableCandidates,
resolve_solvent_carrier,
selectable_validation_candidates,
)
@cache
def _canonical_carrier(cs: str) -> dict | None:
"""The configuration this case study reports, from the resolver that decides it.
This notebook built the same cross-stage rank-1 by hand in four places - concatenating
`explorer.best` over signal, allocation and risk_overlay, dropping benchmark families,
applying `LABEL_RESTRICTIONS` and `RUNG_PINS`, and taking the highest Sharpe. Three of
them wanted the winner and read it from here; the fourth, `_val_rank1_carrier`, walks the
whole field and takes it from `selectable_validation_candidates`, which is the same
ranking one step earlier. That is not the ranking the case studies report. `resolve_canonical_rank1_lineage` re-ranks the field
on exact common timestamp support whenever a conformal candidate is in it, because a
conformal allocator abstains until it is calibrated and books zeros over the abstention,
and it applies `UNIVERSE_RESTRICTIONS` and `CARRIER_PINS` besides. Measured 2026-09-18
against the nine canonical registries, the two rankings named different configurations on
two case studies: `fx_pairs` (`linear/ridge_a1000000.0` at Sharpe 0.4121
against `deep_learning/lstm_h64` at comparison Sharpe 0.3091) and
`nasdaq100_microstructure` (`deep_learning/nlinear` on fwd_ret_15m at 2.3001 against
`gbm/default_multiclass` on fwd_dir_15m at 2.4159). `spine_prediction_hash` is what
`05_portfolio_allocation` and Figure 20.7 pin their allocator comparison to, so on
`fx_pairs` the chapter compared allocators on a configuration the case study does not
report.
Returns None where the resolver finds nothing selectable, which is the state the
hand-built rankings reported as an empty frame. An insolvent or mis-calibrated carrier
still raises: it is a sweep to fix, not a case study to skip.
"""
try:
return resolve_solvent_carrier(cs)
except NoSelectableCandidates:
return None
def _retired(cs: str) -> frozenset[str]:
"""Identities a later generation retired, from the same helper `populate_paired_metrics`
uses. Both write `backtest_paired_metrics`, so a disagreement here would let a Chapter 20
run overwrite the corrected pairs with a retired lineage."""
from case_studies.utils.paired_metrics import _retired_prediction_hashes
return _retired_prediction_hashes(cs)
@cache
def _live_predictions(cs: str) -> list[str] | None:
"""What this case study currently publishes, or None when it declares no populations.
Membership, not the complement of retirement. A prediction no population ever listed has
not been retired by anyone, so ranking over "everything not retired" admits experimental
results the case study never published; ranking over the members in force does not.
Applied inside the query rather than to its result, because `best()` applies its SQL
`LIMIT top_n` first - a row filtered afterwards has already consumed a slot and can hide a
live candidate below the cut.
"""
from case_studies.research.population import published_members_at
published = published_members_at(get_case_study_dir(cs), member_kind="prediction")
if published is None:
return None
if not published:
# `best()` tests this argument for truthiness, so an empty list would read as "no
# filter" and rank everything. A study that declares populations and publishes
# nothing has nothing to report, which is a refusal rather than a wide-open ranking.
raise RuntimeError(f"{cs} declares populations but publishes no prediction identities")
return sorted(published)
def _best_live(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
"""`explorer.best` narrowed to what the case study still publishes."""
return explorer.best(stage=stage, top_n=top_n, prediction_hashes=_live_predictions(cs))
def _best_pinned(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
"""`explorer.best` for a stage, fetching enough rows that a rung-restricted
cohort survives the post-hoc predicate filter.
`best()` extracts `universe_filter` from `spec_json` in Python, after the
SQL `LIMIT top_n`. For nasdaq the pinned cost-feasible carrier sits below
the full-universe in-sample maxima, so a small `top_n` truncates it before
`_apply_rung_restriction` runs. Pull all rows for restricted case studies."""
live = _live_predictions(cs)
if cs in _CLUSTER_RUNG_RESTRICTIONS:
return explorer.best(stage=stage, top_n=1_000_000, prediction_hashes=live)
return explorer.best(stage=stage, top_n=top_n, prediction_hashes=live)
def _apply_rung_restriction(df: pl.DataFrame, cs: str) -> pl.DataFrame:
"""Filter `df` to the case study's pinned rung, if one is configured.
Returns the input untouched if no restriction applies. The helper
relies on `BacktestExplorer.best()` always emitting both
`universe_filter` and `exit_at_max_days` columns; if a future
schema regression drops them, the polars `filter` will raise a
column-not-found error rather than silently allowing the rank-1
selection to drift back to the cross-rung max."""
rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
if rung is None or df.is_empty():
return df
return df.filter(rung["predicate"])
# There is no selected configuration pin here, and there is no mechanism for one.
# `_CARRIER_PIN_PREDICATES` held `"us_firm_characteristics": pl.col("config_name") ==
# "default_huber"` until 2026-08-25, copied from `case_studies.utils.strategy_analysis.CARRIER_PINS`
# and translated into a config-name predicate, under a "keep in sync" comment doing the job a
# mechanism should.
#
# It had not been in sync for a rebuild. Against the current registry `default_huber` is the
# WEAKEST of the ten configs that reached the allocation stage (48 validation backtests, best
# Sharpe 2.128, 2.075 average - tenth of ten), while the documented rule selects `leaves_63_mse`
# (59 backtests, 3.116). So this restricted one case study to its worst advanced configuration
# while every notebook inside that case study reported its best.
#
# Worse than the hash pin removed from `CARRIER_PINS` the same day, because a hash pin dies
# loudly: every hash changes when a sweep is rebuilt, so it resolves to nothing and stops. A
# config-name predicate survives the rebuild and keeps selecting, silently and wrongly.
#
# The mapping stayed empty behind an `_apply_carrier_pin` that could no longer fire, which is a
# second implementation of a rule nothing applied. A selected configuration restriction needed here
# again is `carrier_pins.carrier_config_name(cs)`, which resolves an owner's pin to its config
# through the registry - the thing the copy existed to avoid, and the thing that would have failed
# loudly rather than filtering to the wrong config.
def _progression_for(
explorer: "BacktestExplorer",
pred_hash: str,
cs: str,
) -> pl.DataFrame:
"""Call `progression()` with the case study's rung scope, if any."""
rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
if rung is None:
return explorer.progression(pred_hash)
return explorer.progression(
pred_hash,
universe_filter=rung["universe_filter"],
exit_at_max_days=rung["exit_at_max_days"],
)
# Stages whose registry numbers should never be reported for a case study,
# either because the strategy makes the stage structurally meaningless (HTM
# short-straddle has no allocator choice, no bps cost sweep) or because the
# legacy registry contains deprecated entries that pre-date the strategy
# redesign. Consumed by both `build_backtest_rows` and the `synthesis_dict`
# sanitizer below so the in-notebook attrition funnel and the JSON artifact
# cannot drift on the same case study.
_STAGES_NOT_APPLICABLE: dict[str, set[str]] = {
"sp500_options": {"costs", "risk"},
"us_firm_characteristics": {"risk"},
"nasdaq100_microstructure": {"allocation", "costs", "risk"},
}
# Per-CS rationale strings for `not_applicable_reason` fields written into
# `synthesis_dict`. Keyed by (cs, stage) so two case studies that skip the
# same stage for different structural reasons render different prose.
_STAGE_NA_REASONS: dict[tuple[str, str], str] = {
("sp500_options", "allocation"): ("HTM short-straddle has fixed 1/n_roll cohort weighting"),
("sp500_options", "costs"): ("option costs use §18.8 bid-ask accounting, not bps sweep"),
("sp500_options", "risk"): ("HTM expiration structure sets risk profile"),
("us_firm_characteristics", "risk"): (
"vectorized-engine path; portfolio overlays purged 2026-05-17"
),
("nasdaq100_microstructure", "allocation"): (
"carrier is a signal-stage slot strategy; the slot mechanism is the sizing rule"
),
("nasdaq100_microstructure", "costs"): (
"timing-corrected broad carrier cost grid deferred to v3.1"
),
("nasdaq100_microstructure", "risk"): (
"timing-corrected broad carrier risk grid deferred to v3.1"
),
}
def _stage_applicable(cs: str, stage: str) -> bool:
"""Return False if `cs` has stage `stage` declared not applicable.
`stage` must be one of the canonical labels used by
``_STAGES_NOT_APPLICABLE`` itself (`allocation`, `costs`, `risk`).
Callers in this notebook always pass canonical literals."""
return stage not in _STAGES_NOT_APPLICABLE.get(cs, set())
# %%
cluster_rows = []
for cs, explorer in explorers.items():
top = _best_pinned(explorer, cs, "signal", 200)
if top.is_empty() or top["sharpe"][0] is None:
continue
if "family" in top.columns:
top = top.filter(pl.col("family") != "benchmark")
label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
if label_restriction and "label" in top.columns:
top = top.filter(pl.col("label").is_in(list(label_restriction)))
top = _apply_rung_restriction(top, cs)
if top.is_empty() or top["sharpe"][0] is None:
continue
rank1 = top["sharpe"][0]
rank10 = top["sharpe"][9] if top.height >= 10 else None
spread = (rank1 - rank10) if rank10 is not None else None
# Fold-level stability for the rank-1 backtest
try:
bt_hash = top["backtest_hash"][0]
fold_df = explorer.fold_performance(bt_hash)
if not fold_df.is_empty():
mean_fold_sh = float(fold_df["sharpe"].mean())
se_fold_sh = (
float(fold_df["sharpe"].std(ddof=1) / (fold_df.height**0.5))
if fold_df.height > 1
else None
)
n_folds_pos = int(fold_df.filter(pl.col("sharpe") > 0).height)
n_folds = fold_df.height
else:
mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0
except Exception:
mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0
cluster_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"cs_id": cs,
"rank1_sharpe": rank1,
"rank10_sharpe": rank10,
"rank1_rank10_spread": spread,
"mean_per_fold_sharpe": mean_fold_sh,
"fold_sharpe_se": se_fold_sh,
"n_folds_pos": n_folds_pos,
"n_folds": n_folds,
"n_configs": int(top.height),
}
)
cluster_df = pl.DataFrame(cluster_rows)
if not cluster_df.is_empty():
print("\n=== Rank-1 Cluster Diagnostics (validation) ===")
print(
cluster_df.select(
"case_study",
"rank1_sharpe",
"rank10_sharpe",
"rank1_rank10_spread",
"mean_per_fold_sharpe",
"fold_sharpe_se",
"n_folds_pos",
"n_folds",
"n_configs",
)
)
# %% [markdown]
# Read the table as a tuple: (rank1 Sharpe, rank10 Sharpe, spread, fold-SE,
# folds-positive). A small rank1→rank10 spread relative to the fold-SE signals
# a thick top-of-distribution. Folds-positive close to the total fold count
# signals temporal stability. Both can be read off per case study without
# collapsing the evidence into a single label.
# %% [markdown]
# ## Overview Table
#
# The nine case studies differ in asset class, rebalancing frequency, universe
# size, and the cost assumption each one carries. The table records those four
# properties so that a later result can be attributed to the setting it was
# measured in rather than to the model that produced it.
# %%
overview_rows = []
for cs, explorer in explorers.items():
setup = configs.get(cs, {})
cost_bps = compute_cost_bps(setup)
universe = setup.get("universe", {})
# A futures universe is sized in products rather than in assets, so
# cme_futures declares `n_products` where the others declare `n_assets`.
# Reading only the latter reported that case study as an empty universe.
n_assets = (
universe.get("n_assets")
or universe.get("n_products")
or len(universe.get("symbols", []))
or 0
)
primary_label = setup.get("labels", {}).get("primary", "")
families = explorer.compare_families(stage="signal")
overview_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"cs_id": cs,
"asset_class": ASSET_CLASS_MAP.get(cs, "unknown"),
"frequency": FREQ_MAP.get(cs, "daily"),
"universe": n_assets,
"primary_label": primary_label,
"cost_bps": cost_bps,
"n_model_families": len(families) if not families.is_empty() else 0,
}
)
overview_df = pl.DataFrame(overview_rows)
overview_df.select("case_study", "asset_class", "frequency", "universe", "cost_bps")
# %% [markdown]
# The test bed covers equity ETFs, crypto perpetuals, intraday microstructure, equity plus
# options, firm characteristics, FX, futures, pure options, and a broad equity panel. The
# `cost_bps` column of the table above records the transaction-cost assumption each one
# carries, so a later result can be read against the cost regime it was measured under.
# %% [markdown]
# ## Model IC Comparison
#
# Mean IC by model family across case studies, queried from `prediction_metrics`.
# Each cell shows the average IC across all configurations within a family,
# filtered to each case study's primary label.
# %%
ic_rows = []
for cs, explorer in explorers.items():
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
if not db_path.exists():
continue
# Filter by primary label so IC values match book prose
primary_label = configs.get(cs, {}).get("labels", {}).get("primary", "")
db = sqlite3.connect(str(db_path))
# Best IC per family on primary label only
# Exclude causal_dml: it estimates treatment effects, not predictive IC.
# NOTE: best_ic and best_ic_daily are independent per-family MAXes — they may
# come from *different* predictions. This is intentional ("best daily IC in
# the family"), not "the daily IC of the best-by-fold model".
query = """
SELECT t.family, MAX(pm.ic_mean) AS best_ic,
MAX(pm.ic_mean_daily) AS best_ic_daily,
AVG(pm.ic_mean) AS mean_ic, COUNT(*) AS n_preds
FROM training_runs t
JOIN prediction_sets p ON t.training_hash = p.training_hash
JOIN prediction_metrics pm ON p.prediction_hash = pm.prediction_hash
WHERE p.split != 'holdout'
AND pm.ic_mean IS NOT NULL
AND t.family != 'causal_dml'
"""
params: tuple = ()
if primary_label:
query += " AND t.label = ?\n"
params = (primary_label,)
query += " GROUP BY t.family"
rows = db.execute(query, params).fetchall()
db.close()
for family, best_ic, best_ic_daily, mean_ic, n_preds in rows:
ic_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"family": family,
"ic_mean": mean_ic,
"ic_best": best_ic,
"ic_best_daily": best_ic_daily,
"n_predictions": n_preds,
}
)
ic_df = pl.DataFrame(ic_rows)
# %%
if not ic_df.is_empty():
ic_pivot = ic_df.pivot(on="family", index="case_study", values="ic_mean").sort("case_study")
else:
ic_pivot = pl.DataFrame()
# %% [markdown]
# ### Mean IC by Model Family
# %%
ic_pivot
# %% [markdown]
# Each cell is the mean of `ic_mean` over that family's non-holdout prediction sets, taken at
# the primary label the case study declares in `setup.yaml`, with `causal_dml` excluded because
# its runs are not fit to predict. Families are comparable within a row, since the label is
# fixed across the row; they are not comparable across rows, because each case study declares a
# different primary label, from a fifteen-minute forward return to a twenty-one-day one, and
# prices a different instrument. A negative mean IC marks a case study where prediction is
# difficult under that label rather than a defect in the family. §20.3 carries this table as
# Table 20.4, and §20.6 works through how an option case study's raw IC translates into Sharpe
# once single-name execution costs are charged.
# %% [markdown]
# ## Backtest Comparison
#
# Cross-dataset comparison of pipeline outcomes. Each row takes the highest-Sharpe result at
# each stage **independently**, so the signal that tops one column may come from a different
# model than the allocation that tops the next.
# %%
def build_backtest_rows():
"""Build backtest comparison rows from all case study explorers."""
bt_rows = []
for cs, explorer in explorers.items():
summary = explorer.summary()
# Best signal-stage result — exclude benchmark families (equal_weight,
# etc.) since §20.4's model comparison is about trained models, not
# passive baselines. Also apply case-study label and universe-filter
# restrictions so the Ch20 rank-1 is HTM-coherent for sp500_options
# and pinned to the Rung-3 liquid subset, which is what
# `strategy_analysis.UNIVERSE_RESTRICTIONS` holds ({"sp500_options":
# "liquid"}). The Rung-2 full universe is retained only for the §18.8
# cascade comparison and never anchors the deployed carrier.
label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
signal_candidates = _best_pinned(explorer, cs, "signal", 200)
if not signal_candidates.is_empty() and "family" in signal_candidates.columns:
signal_candidates = signal_candidates.filter(pl.col("family") != "benchmark")
if (
label_restriction
and "label" in signal_candidates.columns
and not signal_candidates.is_empty()
):
signal_candidates = signal_candidates.filter(
pl.col("label").is_in(list(label_restriction))
)
signal_candidates = _apply_rung_restriction(signal_candidates, cs)
best_signal = signal_candidates.head(1)
signal_sharpe = best_signal["sharpe"][0] if not best_signal.is_empty() else None
best_source = best_signal["source"][0] if not best_signal.is_empty() else ""
# Carrier-pred pin for cost/risk. Case studies with a rung restriction
# (nasdaq cost-feasible ensemble) carry their headline cost/risk on the
# selected prediction only; the full-universe sweep rows are the
# Ch18/Ch19 cost-defeat demonstration and must not pool into the
# cross-case comparison. Other case studies pass None (no pin) and keep
# the registry-wide aggregation unchanged.
carrier_pred = (
best_signal["prediction_hash"][0]
if cs in _CLUSTER_RUNG_RESTRICTIONS and not best_signal.is_empty()
else None
)
# Best allocation-stage result (same filters). For case studies that
# declare the allocation stage not applicable (e.g. sp500_options HTM),
# the registry numbers come from deprecated runs, so report None to
# match the synthesis_dict sanitizer below.
if _stage_applicable(cs, "allocation"):
alloc_candidates = _best_live(explorer, cs, "allocation", 200)
if not alloc_candidates.is_empty() and "family" in alloc_candidates.columns:
alloc_candidates = alloc_candidates.filter(pl.col("family") != "benchmark")
if (
label_restriction
and "label" in alloc_candidates.columns
and not alloc_candidates.is_empty()
):
alloc_candidates = alloc_candidates.filter(
pl.col("label").is_in(list(label_restriction))
)
alloc_candidates = _apply_rung_restriction(alloc_candidates, cs)
best_alloc = alloc_candidates.head(1)
alloc_sharpe = best_alloc["sharpe"][0] if not best_alloc.is_empty() else None
# The allocator that produced `alloc_sharpe`, read from that row's own spec, so
# the name and the number describe one configuration. See
# `strategy_analysis.allocation_method_of`.
best_allocator = allocation_method_of(
cs, best_alloc["backtest_hash"][0] if not best_alloc.is_empty() else None
)
else:
alloc_sharpe = None
best_allocator = ""
# Cost sensitivity (gated by stage policy)
survives_costs = None
if _stage_applicable(cs, "costs"):
cost_df = explorer.cost_sensitivity(prediction_hash=carrier_pred)
if not cost_df.is_empty():
zero_cost = cost_df.filter(pl.col("cost_bps") == 0)
survives_costs = not zero_cost.is_empty() and zero_cost["sharpe"].max() > 0
# Risk overlay (gated by stage policy)
best_overlay = ""
managed_sharpe = None
if _stage_applicable(cs, "risk"):
risk_df = explorer.risk_impact(prediction_hash=carrier_pred)
if not risk_df.is_empty():
# rank_one, not a one-key sort: overlays that never trigger book the
# baseline Sharpe exactly, so ties at the top are ordinary here and a
# one-key sort would let frame order decide the name reported below.
best_risk_row = rank_one(risk_df, by="sharpe", name="risk_name")
best_overlay = best_risk_row["risk_name"][0]
managed_sharpe = best_risk_row["sharpe"][0]
# Spine rank-1 prediction_hash - the configuration the case study reports, taken
# from `_canonical_carrier` rather than ranked a second time here. Figure 20.7 and
# `05_portfolio_allocation` both read this value, and Ch20 prose Tables 20.5-20.7
# quote the resolver, so the two have to be one answer.
_carrier = _canonical_carrier(cs)
spine_pred_hash = _carrier["val_prediction_hash"] if _carrier else None
bt_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"case_study_id": cs,
"spine_prediction_hash": spine_pred_hash,
"n_signal": summary.get("signal", 0),
"n_allocation": summary.get("allocation", 0),
"n_cost": summary.get("cost_sensitivity", 0),
"n_risk": summary.get("risk_overlay", 0),
"best_source": best_source,
"signal_sharpe": signal_sharpe,
"best_allocator": best_allocator,
"alloc_sharpe": alloc_sharpe,
"survives_costs": survives_costs,
"best_overlay": best_overlay,
"managed_sharpe": managed_sharpe,
}
)
return bt_rows
# %%
bt_rows = build_backtest_rows()
# %%
bt_df = pl.DataFrame(bt_rows)
print("\nPipeline Comparison:")
print(
bt_df.select(
"case_study",
"signal_sharpe",
"alloc_sharpe",
"survives_costs",
"managed_sharpe",
)
)
# %% [markdown]
# Read the baseline column of the table above for how many case studies enter the pipeline with
# a positive baseline-stage Sharpe and which do not. Those counts move whenever a registry is
# rebuilt, which is why they are in the table rather than in this sentence.
#
# The lineage table below traces each selected prediction across the stages in the order the
# backtests run: baseline, allocation, risk overlay, then the cost sweep charged against
# whatever survived. A Sharpe that rises from one column to the next is what that stage added,
# and only where the later stage carries the earlier one's configuration - the paired rows above
# say which transitions meet that test. NASDAQ-100 is excluded from that comparison in v3.0
# because its timing-corrected broad cost and risk grids are deferred to v3.1.
# %% [markdown]
# ## Paired-Bootstrap Comparison vs Equal-Weight Benchmark
#
# Each case study's selected baseline-stage backtest, under the same label, universe-filter and
# rung restrictions used for the cluster diagnostics, is compared to its equal-weight benchmark
# using a **paired stationary block bootstrap on daily strategy returns**. Block length is derived from
# ``setup.yaml.labels.{label}.rebalance_step`` (falling back to the optimal
# block size, never below the label horizon). Reported quantities:
#
# - ``sharpe_diff`` with a bootstrap confidence interval
# - ``ret_diff``, the annualized return difference, with its confidence interval
# - ``info_ratio`` of the daily-return difference
# - ``prob_challenger_wins`` — bootstrap fraction in which challenger Sharpe
# exceeds the benchmark
# - ``p_value`` — two-sided bootstrap p-value for ``sharpe_diff = 0``
#
# Results land in ``backtest_paired_metrics`` (per case study) and roll up
# into the cross-dataset table below. Intervals are at the conventional confidence level the
# bootstrap call sets. This is the right unit of uncertainty for the headline claim about the
# selected configuration: the Sharpe **difference against the passive baseline that experienced
# the same market conditions**, rather than the Sharpe alone.
# %%
def _benchmark_returns_from_artifact(
cs: str, label: str, period: str = "overall"
) -> tuple[str, pl.DataFrame, str] | None:
"""Resolve the side-artifact equal-weight benchmark for ``(cs, label)``.
The benchmark is the daily-MTM EW reference series persisted by
``scripts/compute_vectorized_ew_benchmark.py`` at
``case_studies/{cs}/benchmark/{label}.parquet``. Single, well-defined
methodology per (cs, label) — no universe/rung/cadence ambiguity that
a registry-side ``family='benchmark'`` lookup would have to disambiguate.
``period`` selects the time window slice (``"overall"`` or ``"holdout"``)
applied by ``load_benchmark_returns``. Classification-label fallback to
the matching ``fwd_ret_*`` artifact applies in both periods.
Returns ``(synthetic_hash, returns_df, resolved_label)`` where
``synthetic_hash`` is a deterministic identifier safe to use as the PK
column in ``backtest_paired_metrics`` (which has no FK on
``benchmark_hash``) and ``resolved_label`` is the actual label whose
artifact was loaded — equal to ``label`` unless the classification
fallback fired, in which case it's the matching ``fwd_ret_*`` label.
Returns ``None`` if the artifact is missing.
"""
df = load_benchmark_returns(cs, label, period=period)
bench_label = label
if df.is_empty() or "ew_return" not in df.columns:
# Fallback: classification labels (``fwd_class_*``, ``fwd_dir_*``,
# ``fwd_tb_*``, ``fwd_carry_*``) share the same forecast window as
# their continuous counterpart (``fwd_ret_*``). The EW universe over
# the same window is identical regardless of the label being
# predicted, so map e.g. ``fwd_class_1m`` to ``fwd_ret_1m``.
fallback = None
for prefix in ("fwd_class_", "fwd_dir_", "fwd_tb_", "fwd_carry_"):
if label.startswith(prefix):
fallback = "fwd_ret_" + label[len(prefix) :]
break
if fallback is None:
return None
df = load_benchmark_returns(cs, fallback, period=period)
if df.is_empty() or "ew_return" not in df.columns:
return None
bench_label = fallback
suffix = "" if period == "overall" else f":{period}"
bench_hash = f"side_ew:{cs}:{bench_label}{suffix}"
return (
bench_hash,
df.select(
pl.col("timestamp").cast(pl.Date).alias("timestamp"),
pl.col("ew_return").cast(pl.Float64).alias("ret"),
),
bench_label,
)
def _aligned_returns(cs: str, h: str) -> pl.DataFrame | None:
"""Load and normalize a backtest's daily returns; columns ``[timestamp, ret]``."""
parquet = get_case_study_dir(cs) / "run_log" / "backtest" / h / "daily_returns.parquet"
if not parquet.exists():
return None
df = pl.read_parquet(parquet)
ret_col = next(
(c for c in ("daily_return", "ret", "return", "value") if c in df.columns),
df.columns[-1],
)
ts_col = next(
(c for c in ("timestamp", "date", "datetime") if c in df.columns),
df.columns[0],
)
return df.select(
pl.col(ts_col).cast(pl.Date).alias("timestamp"),
pl.col(ret_col).cast(pl.Float64).alias("ret"),
)
# %%
import numpy as np
from case_studies.utils.uncertainty import (
SIGNAL_BASELINE_BY_CASE_STUDY,
STAGE_SEQUENCE,
compute_independent_diff_uncertainty,
compute_paired_uncertainty,
descends_from,
joint_returns,
)
def _min_paired_n(ppy: int) -> int:
"""Minimum series length for paired-bootstrap stability, frequency-aware.
The ~21 floor was written for daily cadences (about a month of obs).
Monthly case studies (e.g. ``us_firm_characteristics``) have ~12 holdout
observations by design, and ``compute_paired_uncertainty`` runs cleanly
on n=12. Scale the floor with ``ppy`` so monthly/weekly CSs aren't
blocked by a daily-tuned guard.
"""
if ppy <= 12: # monthly
return 6
if ppy <= 52: # weekly
return 12
return 21 # daily / 8h / intraday
# Distinguish skipped CSs from real failures so empty cross-dataset rollups
# aren't indistinguishable from a silent crash.
paired_rows: list[dict] = []
paired_skips: list[dict] = []
for cs, explorer in explorers.items():
# The leader is the configuration the case study reports, not a ranking rebuilt here.
# `_canonical_carrier` documents why the two are not the same ordering.
carrier = _canonical_carrier(cs)
if carrier is None:
paired_skips.append({"case_study": cs, "reason": "no_selectable_candidates"})
continue
leader_hash = carrier["val_backtest_hash"]
leader_label = carrier["label"]
if not leader_label:
paired_skips.append({"case_study": cs, "reason": "no_label_on_leader"})
continue
bench_resolution = _benchmark_returns_from_artifact(cs, leader_label)
if not bench_resolution:
paired_skips.append(
{"case_study": cs, "reason": f"no_benchmark_artifact_for_label:{leader_label}"}
)
continue
benchmark_hash, base, resolved_bench_label = bench_resolution
chal = _aligned_returns(cs, leader_hash)
if chal is None:
paired_skips.append({"case_study": cs, "reason": "no_challenger_returns_parquet"})
continue
ppy = {"daily": 252, "weekly": 52, "monthly": 12, "8h": 1095}.get(
FREQ_MAP.get(cs, "daily"), 252
)
min_n = _min_paired_n(ppy)
aligned = chal.join(base, on="timestamp", how="inner", suffix="_b")
if aligned.height < min_n:
paired_skips.append(
{"case_study": cs, "reason": f"insufficient_overlap:n={aligned.height}"}
)
continue
# A strategy against a benchmark: the leader's leading flat run is warmup before its
# first signal, not a position it held, so the sample starts where both are trading.
c_arr, b_arr = joint_returns(aligned["ret"].to_numpy(), aligned["ret_b"].to_numpy())
if c_arr.size < min_n:
paired_skips.append(
{"case_study": cs, "reason": f"insufficient_after_coerce:n={c_arr.size}"}
)
continue
paired = compute_paired_uncertainty(
c_arr,
b_arr,
periods_per_year=ppy,
case_study=cs,
label=leader_label,
n_boot=2000,
seed=42,
)
if not paired:
paired_skips.append({"case_study": cs, "reason": "paired_uncertainty_empty"})
continue
# Side-artifact benchmark — deterministic across (cs, label), no
# universe/rung ambiguity, no fallback-by-recency.
benchmark_kind = f"{SIGNAL_BASELINE_BY_CASE_STUDY.get(cs, 'equal_weight')}_side_artifact"
paired_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"label": leader_label,
"benchmark_label": resolved_bench_label, # may differ from leader_label when classification fallback fired
"sharpe_diff": paired.get("sharpe_diff"),
"sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
"sharpe_diff_ci_hi": paired.get("sharpe_diff_ci95_hi"),
"ret_diff": paired.get("ret_diff"),
"info_ratio": paired.get("info_ratio"),
"p_value": paired.get("p_value"),
"prob_wins": paired.get("prob_challenger_wins"),
"block": paired.get("bootstrap_block_length"),
"n_boot": paired.get("bootstrap_n"),
}
)
paired_df = pl.DataFrame(paired_rows)
if not paired_df.is_empty():
print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
print(
paired_df.select(
"case_study",
"label",
"benchmark_label",
"sharpe_diff",
"sharpe_diff_ci_lo",
"sharpe_diff_ci_hi",
"info_ratio",
"prob_wins",
"p_value",
)
)
else:
print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
print("No paired-bootstrap rows produced — see skip table below for reasons.")
if paired_skips:
print("\nSkipped case studies:")
for s in paired_skips:
print(f" - {s['case_study']:<32} {s['reason']}")
# Loud invariant — a 0/N or all-skipped outcome is now obvious in the
# notebook output instead of buried under "no paired-bootstrap rows."
print(f"\npaired={len(paired_rows)}/{len(explorers)}, skipped={len(paired_skips)}/{len(explorers)}")
# %% [markdown]
# Read each row as the selected challenger's annualized Sharpe **minus** the equal-weight
# benchmark's, with a confidence interval from the paired stationary block bootstrap on the
# daily-return difference; the information ratio summarizes the
# excess-return-to-tracking-error ratio; ``prob_wins`` is the fraction of
# bootstrap resamples in which the challenger beat the benchmark; ``p_value``
# tests ``H0: sharpe_diff = 0``. A confident "the model adds skill over the
# passive baseline" claim requires (i) the CI excludes zero, (ii) ``prob_wins``
# close to 1, and (iii) a small ``p_value``. Cases where the CI straddles
# zero are not failures — they signal that the apparent Sharpe gap is within
# block-bootstrap sampling error and should be reported as such.
# %% [markdown]
# ## Paired metrics — full coverage for strategy-analysis notebook
#
# The block above populates the first pair type, the selected baseline signal against
# equal-weight over the whole window. The strategy-analysis notebook (per-CS strategy notebooks)
# requires five additional pair types per case study to render §2 (stage-
# transition waterfall), §6 (holdout decay + holdout-vs-benchmark) and §7
# (benchmark-aware diagnostics) without inline bootstrap recomputation.
#
# The pair set:
#
# 1. selected signal (overall) ↔ equal-weight (overall) — populated above
# 2. selected signal (holdout) ↔ equal-weight (holdout window)
# 3. the selected configuration on holdout ↔ the same configuration on
# validation (same lineage decay; min-length truncation since the
# windows are disjoint)
# 4-6. one pair per consecutive stage transition the prediction actually has,
# in ``STAGE_SEQUENCE`` order: allocation ↔ signal, risk-overlay ↔
# allocation, cost-sensitivity ↔ risk-overlay. A case study that did not
# run a stage yields fewer pairs, and a stage that does not carry the
# previous stage's configuration yields none for that transition - the two
# were selected independently and their difference is not a stage effect.
#
# Pair #3 truncates both series to ``min(len(val), len(ho))`` to satisfy
# ``compute_paired_uncertainty``'s equal-length precondition. The CI is
# interpreted as bootstrap resampling Sharpe in each window independently
# and taking the difference; the truncation is preserved in the
# ``benchmark_kind`` value (``val_rank1_self`` always carries the truncation
# caveat). All pairs use the same paired stationary block bootstrap helper.
# %%
def _full_strategy_spec_from_backtest(db: sqlite3.Connection, bt_hash: str) -> dict | None:
"""Pull the full strategy spec dict (signal + allocation + risk) from
`bt_hash`'s spec_json. Returns None if the row is missing or signal has
no `method` field.
A backtest's full specification is the tuple (signal, allocation, risk). Pinning
the val→holdout pair on this full spec keeps the comparison apples-to-
apples; pinning on signal alone allows MAX(sharpe) to surface a holdout
row with a different allocation (e.g. conformal_weighted) or risk overlay
than the validation rank-1 carrier.
"""
row = db.execute(
"SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
(bt_hash,),
).fetchone()
if not row:
return None
strat = json.loads(row[0]).get("strategy", {})
sig = strat.get("signal", {})
if not sig.get("method"):
return None
alloc = strat.get("allocation") or {}
risk = strat.get("risk") or {}
return {
"signal": {
"method": sig.get("method"),
"top_k": sig.get("top_k"),
"percentile": sig.get("percentile"),
},
"allocation": {
"method": alloc.get("method"),
"top_k": alloc.get("top_k"),
"long_short": alloc.get("long_short"),
},
"risk": {
"name": risk.get("name"),
},
}
def _val_rank1_carrier(cs: str) -> dict | None:
"""Return ``{'spec', 'prediction_hash'}`` for ``cs``'s validation rank-1 carrier.
The prediction hash is carried out alongside the spec because the holdout resolver
needs it: naming all three stages pins the configuration AND the checkpoint, and without
it a case study that registered several checkpoints against one strategy is ambiguous
and the resolver refuses. It was determinable all along - this walk had it in hand and
threw it away - so refusing there would have dropped a case study out of the
reader-facing holdout table for want of a value one line above.
The val rank-1 *full strategy* spec for ``cs`` — the
highest-Sharpe validation backtest across (signal, allocation,
risk_overlay) stages — walking candidates by val Sharpe descending until
one with a matching holdout backtest at the SAME full spec is found.
Implements the selection rule documented in §20.1: the deployed
holdout for each case study is the val rank-1 across all three pipeline
stages, retrained on holdout data; when retrain produces no usable
holdout at that full spec (degenerate predictions, vol-window mismatch,
universe filter rejection, etc.) the walk falls through to the next
candidate by val Sharpe. The first val candidate with a registered
holdout backtest at the same (signal, allocation, risk) tuple defines
the apples-to-apples carrier pair.
Returns None when no val candidate up to rank ~200 has a matching
holdout under the case study's label / rung restrictions.
"""
explorer = explorers.get(cs)
if explorer is None:
return None
# The field the resolver ranks, in the resolver's order, rather than a concat of
# `explorer.best` re-filtered here: the walk starts at the carrier `_canonical_carrier`
# names and falls through in the same order the resolver would. Every row is kept - no
# dedup by prediction_hash - because when the rank-1 configuration has no matching holdout
# retrain but a same-prediction lower-Sharpe variant (a different allocator or risk
# overlay) does, a dedup would jump to a different prediction instead of accepting the
# same-prediction variant as the apples-to-apples match.
try:
candidates = selectable_validation_candidates(cs)
except NoSelectableCandidates:
# The helper raises on an empty pool rather than returning one, and this walk's
# callers read `None` as "no holdout pair for this case study" - the state the
# hand-built ranking reported as an empty frame. A pool with nothing eligible in it
# is that state, not a reason to stop aggregating the other eight.
return None
label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
db = sqlite3.connect(str(db_path))
try:
for candidate in candidates[:200]:
bt_hash = candidate["backtest_hash"]
spec = _full_strategy_spec_from_backtest(db, bt_hash)
if spec is None:
continue
# The probe below asks whether THIS candidate has a holdout, so it matches the
# candidate's own configuration and checkpoint and not only its strategy spec.
#
# Matching the spec alone made the walk stop at a candidate whose own checkpoint
# had no holdout whenever a sibling checkpoint had one at the same spec. The
# resolver, handed that configuration, then finds nothing for it - and the walk has
# already stopped, so the case study reports no holdout while one exists for a
# later candidate. Advancing instead is what makes the fall-through the resolver
# no longer performs unnecessary rather than merely forbidden.
carrier_row = db.execute(
"""
SELECT t.family, t.config_name, t.label,
p.checkpoint_value, p.checkpoint_kind
FROM prediction_sets p
JOIN training_runs t ON t.training_hash = p.training_hash
WHERE p.prediction_hash = ?
""",
(candidate["prediction_hash"],),
).fetchone()
if carrier_row is None:
continue
spec_clauses, spec_params = _full_strategy_clauses(spec)
ho_clauses = ["p.split = 'holdout'"] + spec_clauses
if _retired(cs):
ho_clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
ho_params: list[object] = list(spec_params)
if _retired(cs):
ho_params.append(json.dumps(sorted(_retired(cs))))
if label_restriction:
placeholders = ",".join("?" for _ in label_restriction)
ho_clauses.append(f"t.label IN ({placeholders})")
ho_params.extend(sorted(label_restriction))
if rung is not None:
ho_clauses.append(
"COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
)
ho_params.append(rung["universe_filter"])
if rung["exit_at_max_days"] is None:
ho_clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
)
else:
ho_clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?"
)
ho_params.append(rung["exit_at_max_days"])
# `t.spec_json` rather than `1`, and no LIMIT: the probe has to apply the same
# eligibility test the resolver applies, and that test is not expressible in SQL.
#
# A model fitted on the validation folds can publish predictions over the holdout
# window, so `p.split = 'holdout'` with a non-null Sharpe is not enough to make a
# row a holdout result. The resolver drops those through
# `training_run_fitted_for_the_holdout`; a probe that admitted them would stop the
# walk at a candidate whose only holdout is validation-fitted, the resolver would
# then find nothing eligible for it, and the case study would report no holdout
# while a later candidate had a real one.
probe_rows = db.execute(
f"""
SELECT t.spec_json FROM prediction_sets p
JOIN training_runs t ON p.training_hash = t.training_hash
JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
AND b.stage IN ('signal','allocation','risk_overlay','holdout')
JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
WHERE {" AND ".join(ho_clauses)}
AND t.family = ?
AND t.config_name = ?
AND t.label = ?
AND p.checkpoint_value IS ?
AND p.checkpoint_kind IS ?
AND bm.sharpe IS NOT NULL
""",
ho_params + list(carrier_row),
).fetchall()
row = any(training_run_fitted_for_the_holdout(probe[0]) for probe in probe_rows)
if row:
return {"spec": spec, "prediction_hash": candidate["prediction_hash"]}
finally:
db.close()
return None
def _full_strategy_clauses(spec: dict | None) -> tuple[list[str], list[object]]:
"""Build SQL WHERE clauses + params that pin a backtest row to the full
strategy spec (signal + allocation + risk). Empty list returned when
spec is None (no constraint).
Pinning on the full spec ensures `MAX(sharpe)` over candidate holdout
backtests cannot surface a different allocator (e.g. conformal_weighted
when val rank-1 was score_weighted) or a different risk overlay than
the validation carrier — the val→holdout pair stays apples-to-apples
on the full pipelinSe muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.